Originally written in French. Translated by AI — the meaning has been preserved, not the prose.
Angle
A language model retrieves poorly what has been buried in the middle of a long document, and this degradation through filling is a measured property, not a user impression: a window's stated capacity and its usable capacity do not coincide. The entire architecture of a workspace shared with an assistant follows from that single constraint — a short core loaded systematically, the rest called on demand, separate tiers for sources, memory and outputs, short linked notes rather than monolithic syntheses. And this partitioning, designed to fit inside a technical limit, produces three effects that have nothing to do with it any more: it makes conclusions traceable back to their sources, it allows a deliverable to be regenerated at any moment rather than at the end of a phase, and it turns open subjects that led nowhere into the raw material of the next ones.
Synthesis
The starting point looks like a volume problem. A context.md file grows until it no longer fits, and the obvious answer is to cut it. The first cut is made by addressee — one file for the assistant, one for the human — because the two needs really are incompatible: one wants markers for picking the work back up, the other wants to aggregate and arbitrate. That cut does not hold: both files swell in turn and each fills with a bit of everything.
What settles the matter is not volume, it is a property of the model. The window is finite, which we know, but what sits inside it is not equally available — information placed in the middle is markedly less well retrieved than information placed at the ends. Filling is therefore not neutral: every line added pushes towards the middle whatever was already there, and nothing in the answer signals what was missed.
Two rules follow directly. The first bears on loading: a short core — mission, current state, plan, decisions, open contradictions — goes in systematically, and everything else waits to be called, in an order written in advance rather than improvised when the session opens. The second bears on partitioning: it is done neither by volume nor by addressee but by nature of content, raw material on one side, what is retained from it in the middle, the outputs at the end. What distinguishes this architecture from mere tidying is that it has a mechanical reason.
And this is where the setup returns more than was asked of it. Separating the tiers lets you walk a conclusion backwards — from the synthesis to the checkpoint, from the checkpoint to the sessions, from the sessions to the sources — and therefore to answer a stakeholder with something other than "the assistant recommended it". Keeping the material up to date makes a deliverable cheaply reproducible, which strips the discovery-specs-delivery sequence of its economic justification. And because opening a subject becomes cheap, you open many of them, most of which come to nothing — except that they cross-fertilize, and a fourth subject is born the richer for three that went nowhere.
The grain of the writing obeys the same constraint and produces the same surplus. Short linked notes load in pieces, and they also let you hold two ideas side by side — which a ten-page synthesis forbids, since none of its forty ideas can be cited without the other thirty-nine. Simulated annealing, carried over from metallurgy to combinatorial optimization, does not come from a stock of knowledge but from bringing together two precise and distant fields; the grain of the material decides what can meet what.
There remains what the architecture does not capture. What makes a practitioner valuable is the depth of situation he has internalized, and that is exactly what an assistant does not have. Part of it can be written — that is the work of building the context, and writing it tests what you thought you had understood, since the machine takes literally what a colleague would have caught. Another part cannot be written: specific knowledge, the kind that leads to a conviction and to a bet. The partitioning moves a great deal of work towards what can be handed to a machine; by contrast, it makes visible what cannot be handed over.
Tensions / contradictions
One tension bears on selection itself. A short core is what makes the assistant operational, and it is also what silently decides what it will be ignorant of: what has not been loaded produces no visible error, exactly like what was buried in the middle. The partitioning therefore does not remove the flaw it treats, it moves its cause from the machine to whoever wrote the manifest.
Second tension, on forgetting. A long memory needs a ceiling that forces it to consolidate, failing which it becomes a journal nobody rereads; but it is precisely because a file system forgets nothing that walking a conclusion back to its sources stays possible. What consolidation gains in rereadability, it takes away from the chain of evidence.
Third tension, on grain. Partitioning into fine units is what makes distant connections possible, and it is also what produces a network in which a logical entailment, an antagonism and a vague kinship all take the same form. Nothing here says at what fineness the gain in connection turns into a loss of discernment.
Questions
- How do you spot that something left out of the core should have gone in, given that its absence produces no visible error?
- What is it, in a set of open subjects, that triggers the connection at the right moment rather than three months too late?
- How far does this architecture, designed for one person and their assistant, hold when several people write in the same workspace?