Originally written in French. Translated by AI — the meaning has been preserved, not the prose.
Main idea
Tightening the rules of a system that automatically extracts knowledge makes you lose things without knowing it: what is no longer produced leaves no trace. You observe the quality going up, never the coverage going down.
The previous version's output, on the other hand, does leave a trace. Handing the tightened version the notes the POC had produced and asking which ones it would no longer write — and on which rule it discards them — makes the invisible legible. The corpus judged too noisy becomes a control set, and its value rests not on its quality but on its precedence.
The adjustment then bears only on the rules that discarded those cases, without undoing the classification rigour just acquired.
Why it matters
This gives a concrete reason never to throw away a prototype's output, including when you judge it poor: it is the only document attesting to what a more permissive version was able to see.
And the procedure transposes to any tightening of a filter, as soon as an earlier version has produced a comparable corpus: you do not measure a false negative in the absolute, you measure it against a version that did produce it.
Nuances and limits
The permissive version is not a standard of truth: it contains noise, and a note it had produced is not relevant merely because it has disappeared. The procedure flags candidates for re-examination, it does not return a verdict.
Above all, it bounds regression without measuring coverage: what neither version ever saw stays invisible to both.
Open questions
- What can spot something no successive version of the same rule set has ever detected?