Originally written in French. Translated by AI — the meaning has been preserved, not the prose.
Main idea
A research paper, a piece of technical documentation or an analyst report can contain five independent mechanisms. Two very different documents can support exactly the same idea. And a single source can mix a solid claim, a speculative assumption and an ambiguous definition.
Counting or summarising documents therefore doesn't tell you what a corpus knows. You need an intermediate layer where sources are turned into units of knowledge: atomic notes that each pin down one idea, thematic notes that connect several ideas, glossary entries that fix the meaning of a term when an ambiguity would change the reading.
It is on this layer, not on the pile of documents, that you can read how solid a claim is, whether it depends on a single source, or where two authors contradict each other.
Why it matters
It makes measurable things a summary erases: how many independent sources support an idea, which one has only one, where two sources clash.
And it lets you steer research from what has been understood rather than from what has been read.
Nuances and limits
Extraction has a cost, and it brings its own bias: whatever wasn't broken down into a note drops out of the diagnosis as surely as whatever wasn't read.
And a source can remain unusable for lack of reliable extraction: it is then in the corpus without being in the knowledge.
Open questions
- How do you spot what an extraction left out of a source, without rereading the whole source?