Originally written in French. Translated by AI — the meaning has been preserved, not the prose.
When you ask an AI to do some research, the reflex is usually the same:
Find me 20, 30 or 50 good sources on this subject.
The problem isn't the number of sources. The problem is deciding how many to look for before you know what is actually missing.
So you easily end up piling up documents that overlap, rehash the same concepts, cite the same work and give a false impression of coverage. Meanwhile, some angles stay invisible because they are poorly represented in the first results.
I use a different approach: start from a very small but carefully vetted corpus, extract structured knowledge from it, diagnose what this corpus already lets you understand — and above all what it doesn't yet — and then let that state decide the next search.
I call this method the Stateful Corpus Research Loop, or SCRL.
Its logic fits in one sentence:
Don't look for more sources. Look for the next source likely to change what you already know about a given mission.
1. Bulk research isn't the problem
Looking for lots of sources is obviously not bad in itself. Some research needs dozens, even hundreds, of documents.
The problem appears when the strategy is fixed from the outset:
scope → 30 sources → reading → synthesis
Volume then becomes a proxy for quality.
Yet thirty sources may in fact cover only a few repeated ideas. They may come from the same kind of player, cite the same work, rest on the same assumptions or miss the same contradiction.
The interesting question isn't: how many sources do I have?
It is: what does my corpus already know, what doesn't it know yet, and which search is most likely to change that state?
An SCRL therefore reverses the usual logic.
Instead of collecting first and understanding later, you understand enough to decide what to look for next.
2. It all starts with a mission and a scope
An iterative loop with no boundary can go on forever.
Every new source opens new questions. Every question leads to new concepts. And every concept can justify twenty more things to read.
That is why an SCRL starts with an explicit mission:
- a central question;
- what is in scope;
- what is out of scope;
- the themes to understand;
- the kind of knowledge expected.
Convergence never exists in absolute terms.
You can't conclude: "we now know everything there is to know about AI agents".
You can only conclude: for this mission, within this scope, new iterations are now unlikely to substantially change our current model.
The scope isn't red tape. It is what makes a stopping criterion possible.
3. Start from a small but carefully vetted seed
An SCRL doesn't start with exhaustive collecting. It starts with a seed: a few carefully chosen sources.
This first set has to be:
- reliable;
- varied;
- close to the heart of the subject;
- rich enough to bring out a first conceptual structure.
The seed isn't meant to cover the field.
Its role is more modest and more useful: to produce a first state of knowledge good enough to reveal what is missing.
It gets the loop going.
4. Turn sources into knowledge, not summaries
A raw document is a poor unit of reasoning.
One article can contain five independent mechanisms. Two very different documents can support exactly the same idea. Conversely, a single source can contain a solid claim, a speculative assumption and an ambiguous definition.
This means you need an intermediate layer between the sources and the deliverable.
In my system, that layer rests mainly on three types of note I have already described on this blog:
- an atomic note pins down a self-contained unit of knowledge, with a single intellectual centre of gravity; What is an atomic note?
- a thematic note connects several pieces of knowledge to produce a further level of understanding; What Is a Thematic Note?
- a glossary entry pins down the meaning of a term when an ambiguity could change how the corpus is read. What is a glossary entry?
An SCRL therefore doesn't steer the research straight from a pile of documents.
It steers it from knowledge that has already been extracted and structured.
5. The corpus becomes stateful
After the first ingestion, the corpus has a state.
Over time it can represent:
- what seems established;
- what remains thinly documented;
- what depends on a single source;
- what remains uncertain;
- the contradictions;
- terminological ambiguities;
- mechanisms still poorly understood;
- limits already identified;
- over-represented areas;
- areas that are almost absent;
- the relationships between different parts of the corpus.
That state becomes the input to the next decision.
Mission / scope → Vetted seed → Sources → Notes → State of the corpus → Diagnosis → Targeted search → New sources → Corpus update → Convergence test → new iteration
The corpus is no longer just the result of the research.
It becomes the engine of the next search.
6. What a new iteration can actually look for
A new iteration no longer means: "find me some more relevant articles".
It receives a goal produced by the diagnosis of the corpus.
An iteration can aim to:
- broaden — discover an important dimension that is still missing;
- reinforce — add independent evidence to an important piece of knowledge;
- fill a gap — document a mechanism, a step, a metric or a condition that is still missing;
- look for contradiction — deliberately find work that could invalidate or limit what the corpus already treats as solid;
- pin down the limits — establish where and when a claim stops being true;
- separate competing meanings — spot terms that different authors, products or disciplines use differently;
- move from description to mechanism — understand not just what happens, but why and how;
- diversify the evidence — mix theory, benchmarks, empirical studies, technical documentation, field feedback and real-world observation;
- test transferability — check whether an idea observed in one field still holds in another;
- identify dependencies — find the conditions the mechanism needs in order to work;
- look for side effects — costs, latency, bias, regressions, organisational constraints or governance risks;
- check for obsolescence — detect whether a technical change invalidates part of the corpus;
- connect separate areas — notice that the same mechanism runs through several parts of the subject;
- rebalance — correct a conclusion that only looks strong because one type of source is over-represented;
- look for what could break the model — explicitly formulate the search that, if it succeeded, would force you to revisit a central conclusion;
- measure saturation — test whether a new search still produces substantial knowledge.
This list matters because it changes what research is for.
The search engine's job is no longer to find "good sources". It has to find the type of evidence the corpus needs right now.
7. Bulk research vs SCRL
Conventional research often looks like this:
scope → large-scale collecting → reading → extraction → synthesis
An SCRL looks more like this:
scope → small seed → knowledge → diagnosis → next search decided by the diagnosis
The difference isn't 5 sources versus 50 sources.
An SCRL can perfectly well end up with 100 sources.
The difference is between:
collecting before you know precisely what you need
and
letting the knowledge you already have decide what to look for next.
8. A real experiment: Adobe Summit Paris 2026
I tested this approach while doing my own preparation for Adobe Summit Paris 2026.
The mission was deliberately limited to five themes:
- Content production and organisation in the agentic era — agents, orchestration, human oversight, workflows and industrialisation;
- Brand context and AI-driven content governance — brand representation, constraints, validation, persistent context and control;
- Queryable documents and knowledge systems — RAG, retrieval, grounding, citations, documentary contradictions and production from a corpus;
- The agentic web, discoverability and visibility in LLMs — GEO, citation, attribution, structured data, distribution and visibility measurement;
- Conversational interfaces and intent-based creation — natural language, creative control, state persistence, direct manipulation and translating intent into operations.
The research then moved forward in successive loops.
By the end of the experiment:
- 15 research iterations;
- 82 sources added to the corpus;
- 1 unusable source for lack of reliable extraction;
- 51 active notes;
- 34 atomic notes;
- 10 glossary entries;
- 7 thematic notes;
- 5 synthesis briefs, one per theme.
But the interesting number isn't 82.
It's the way the role of the iterations changed.
At the start: broaden
The first loops mostly serve to understand the terrain. They reveal the concepts, the mechanisms and the first boundaries of the subject.
In the middle: reinforce and structure
The following iterations look for better evidence, separate certain ideas, fill gaps and bring out relationships between notes.
At the end: attack the model
The last iterations change in nature.
Iterations 13, 14 and 15 mainly served as break tests.
The question was no longer: "what else can we add?"
It became: "what could we find that would force us to revisit what we think we have understood?"
At the fifteenth iteration, no new source important enough to change the conceptual model of the five themes was retained.
That obviously doesn't mean the subject was exhausted.
It indicated that, within the defined scope, the corpus was beginning to converge.
9. Convergence: knowing when to stop
This is probably the most interesting benefit of the approach.
In conventional research, the question "when have I searched enough?" often goes unanswered.
30 sources? 50? 100?
Those numbers are not an epistemic criterion.
An SCRL can use a different signal.
The research starts to converge when new iterations:
- no longer create new core knowledge;
- no longer significantly change the important notes;
- no longer reveal any major contradiction;
- no longer uncover any significant gap in the scope;
- mostly add confirmation or redundancy;
- leave the remaining uncertainties clearly identified.
So convergence doesn't mean "we know everything".
It means: the probability that a new iteration will substantially change our current representation has become low relative to the cost of that iteration.
And this criterion can be tested.
You can deliberately run searches designed to break the model. If they fail several times in a row, the saturation signal becomes far more credible.
10. What this changes for research agents
Many "deep research" systems still implicitly follow an architecture close to:
search → scrape → summarise → write
Even when they run several searches, these often remain branches of a plan set very early on.
A genuinely adaptive research loop needs more.
The agent has to maintain a state that includes, among other things:
- the scope;
- the stabilised knowledge;
- its provenance;
- its level of support;
- the contradictions;
- the uncertainties;
- the gaps identified;
- the searches already run;
- the strategies that yielded nothing;
- the areas considered saturated.
It can then decide:
Which search now maximises the chance of a useful gain in knowledge?
The architecture becomes:
scope → memory → state of the corpus → diagnosis → search strategy → new evidence → knowledge update → convergence test
Research is no longer a phase before generation.
It becomes a system that updates its own representation of what it knows and uses that representation to choose its next move.
Stateful rather than just deep
The phrase deep research puts the emphasis on depth.
But depth doesn't guarantee quality.
An agent can read a hundred documents while repeating the same biases. It can explore extensively without ever looking for the source that contradicts its central conclusion. It can accumulate a gigantic context without knowing clearly what remains uncertain.
The important shift may not be to search more deeply.
It is to search with a state.
That is the idea behind the Stateful Corpus Research Loop.
The corpus is no longer a pile of documents.
It becomes an evolving representation of the knowledge acquired on a given mission — and that representation is what decides the next search.
The next source needs a reason to exist
Good research probably shouldn't start with this any more:
What are the best articles on this subject?
It should gradually arrive at a more demanding question:
Given what we already know, what information would now be most likely to change our understanding?
It's a subtle difference.
But it completely transforms the role of research.
An SCRL doesn't try to find lots of good sources. It tries to work out, at every moment, which source could still change what we know about a given mission.