Idea

An agent enumerates better than a human and declares the work finished sooner than one

Info

Originally written in French. Translated by AI — the meaning has been preserved, not the prose.

Main idea

What an agent does remarkably well comes down to a few verbs: enumerate, cross-reference, count, keep a line-by-line record, and read the translations of a product's twelfth area as diligently as those of the first. That work is massively repetitive, and that is exactly what made it unaffordable until now.

What it does badly comes down to a single act: declaring the work finished. On the description of a product, three stopping criteria were proposed in turn as proof of completion, and each time they had to be refused. The bias is that of a tired human at the end of a project, without tiredness as an excuse: confusing reporting with delivering.

Human interventions then fall into three categories, readable in the reasons recorded in the project's decision log: widening a scope that had been implicitly narrowed — the agent was reading the core code and stopping there; fixing a structure that was going to diverge — the same rule copied into three documents, out of a desire to do things properly; refusing a stopping criterion that was too convenient.

Why it matters

This says where to place human attention on delegated work, and it is counter-intuitive: not on execution, which is better than ours, but on the boundaries — the scope at the start, the structure along the way, the ending on arrival.

It holds for any work handed over, to a machine as much as to another person, and it rests on four devices: a stopping criterion a third party can verify, a measurement you can re-run rather than an assertion, a written record at the granularity of the area and the place read, and enumeration before search.

Nuances and limits

The observation rests on traces, not on recollections: a decision log where each entry carries its reason, and a record of what the agent produced between two arbitrations. Without those two traces, the split of "it does this well, that badly" goes back to being an impression.

And excellence at enumerating is only worth something on an enumerable object: ask it to count what has no list and it returns a fabricated figure with the same assurance.

Open questions

  • Can an agent refuse its own stopping criterion if you hand it the distinction, or does human supervision remain structurally necessary at that point?