Originally written in French. Translated by AI — the meaning has been preserved, not the prose.
Main idea
The article describing the context engine ported onto ChatGPT was prepared inside that engine: ideas persisted, sources ingested, blind spots identified, tasks created and then completed, a first version produced, reread, then reworked into a V2 from the same context. The system was not tested by a scenario built for it; it was put in charge of the work it existed for.
The difference is not one of rigour but of the nature of the test. A demonstration picks its data and its path — it verifies that the intended route works. Real work brings what nobody anticipated: an idea arriving out of order, a contradiction between two notes, a pick-up several days later, a volume that was not planned for. Those are the situations that decide whether a system holds.
Why it matters
It changes what you can assert at the end. After a demonstration, you know the system does what you asked it to show. After real work carried through to a deliverable, the question "does it work" is closed, and the next question becomes one of scaling — which is progress, because it bears on a use and not on a capability.
And it supplies a practical discipline: pick as a system's first use the one you have to produce anyway, rather than a textbook case that costs nothing to abandon.
Nuances and limits
A single piece of real work does not exhaust the regimes of use. Producing an article over a few days says nothing about how a context behaves when fed for several weeks, nor about a dense week in the field — and the article acknowledges this explicitly.
And the test is biased by its operator: whoever built the system works around what they know to be fragile without thinking about it. The work is real, the user not quite.
Open questions
- How do you tell apart, in a system tested by its author, what holds because of the system and what holds because of the workarounds he makes without seeing them?