000%

NordNeuron

Initializing Intelligence Systems

AI Research

When Agents Can't Forget: What LedgerBench Found About Requirement Memory

Give an AI coding agent a permanent, append-only memory of every requirement it has ever met, and it stops regressing — but it also calcifies, defending rules that stopped being true weeks ago. A pre-registered benchmark measured exactly how much of each.

Pankaj Kumar•September 2026•8 min read

Real software changes its mind. A requirement you nailed in week one gets overturned in week six — "store tasks in JSON files" becomes "store them in SQLite." A human team absorbs that by updating its shared understanding. An AI coding agent working across many sessions has no shared understanding to update; it has whatever you chose to carry forward as its memory. The design of that memory turns out to matter more than almost anything else about the agent.

One tempting design is to never forget: compile every requirement into a runnable test and keep an append-only ledger of all of them, re-run on every change, so nothing is ever lost. It is an appealing idea — provably nothing slips — and it is the design I most wanted to believe in when I started building agent memory for operational systems. LedgerBench is the benchmark I built to find out whether it actually holds up, or whether that ledger quietly calcifies, so old, now-wrong tests block correct new work.

The setup is deliberately small and controlled: an AI coding agent runs through a fixed sequence of 24 tasks on a command-line to-do app, the storage requirement is flipped partway through (and a second requirement later), and five memory strategies are compared over 12 random seeds each — 60 runs in total. Crucially, the whole thing was pre-registered and frozen: every prediction was written down in advance, with no peeking until all 60 runs finished. That last detail is what makes the results worth reporting rather than just anecdotes, because it includes the predictions that turned out wrong.

Five ways to remember

The arms differ only in how the agent's memory works. Arm A is the control: no test ledger at all, just a rolling plain-text summary of previous sessions. Arm B is the pure ledger: every requirement becomes a test, the whole set replays on every change, and tests are never removed — even after a later requirement makes an old one wrong. Arm C adds governance: a separate judge can retire a test, but only when documented evidence — an official requirement change — justifies it. Arm C+ adds a docket that auto-flags any test failing repeatedly for the judge to review. And Arm D gives the coding agent its own voice: it can request that a test be retired and argue its case to the judge.

The headline score is M1@T24 — of every requirement introduced across the whole run, what fraction is the project still satisfying by the final task. Higher is better; it is a direct measure of how much the agent remembered without breaking. Alongside it, two failure counts matter: trap violations (planted requirements the agent was expected to regress on) and rewrite regressions (old, correct behavior that crept back to being broken). And, because none of this is free, tokens per run.

ArmRetention (M1@T24)Trap viol.Rewrite regr.Tokens/run
A baseline0.3061314384k
B pure ledger0.68400875k
C governed0.69830492k
C+ +docket0.67736480k
D +agent voice0.74027483k

Frozen campaign: 5 arms × 12 seeds × 24 tasks, deepseek-v4-flash, temperature 0.2. Formal significance testing was out of scope.

Finding 1: the ledger earns its keep

The clearest result in the table is also the one I was least worried about going in, and it is worth stating plainly because it is the robust one: keeping an automated memory of past requirements roughly doubles how much the agent still gets right at the end. End-state retention goes from 0.31 for the summary-only baseline to 0.68–0.74 for every arm that carries a ledger. Violations on the planted trap requirements fall by 77–100%.

This is the part of the append-only idea that works. A rolling text summary is lossy in exactly the way you would fear — by task 24 the baseline has quietly broken most of what it built earlier, and it walked into 13 of the traps. A ledger that re-checks old requirements on every change simply does not let that drift happen silently. If the only question were "does replaying past requirements help," the answer is an unambiguous yes, and the gap is large enough to survive the noise.

Finding 2: but a ledger that never forgets, calcifies

The pure ledger (Arm B) buys its zero regressions at a real price, and the benchmark was built to make that price visible. It cost 875k tokens per run — 2.3× the baseline — because it replays the entire accumulated test set on every single change, forever. And it produced nine coerced regressions: cases where a stale, now-invalid test literally forced the agent into making a wrong edit to satisfy it.

The mechanism is worth watching happen. After the storage requirement flips from JSON to SQLite, the agent writes a correct SQLite migration — and the old JSON-era tests, still in the permanent replay set, reject it on every iteration. The agent thrashes between backends, then names the contradiction itself. In one run it wrote:

"We have a genuine conflict between the earlier requirements that assume JSON file storage and the later requirement that mandates SQLite storage with no JSON file… it is impossible to satisfy both. You may need to adjust the test ledger…"

And then, having no way to adjust the ledger, it capitulated — reverting a correct migration back to JSON to satisfy the dead rules, ending on the wrong backend. That is calcification in the concrete: not a vague "technical debt" worry but a specific, reproducible failure where the memory system overrides a correct decision. The agent even diagnoses the missing piece — it asks for exactly the amendment mechanism the pure ledger, by design, refuses to have.

Finding 3: a judge recovers the benefit at half the cost

The most useful practical result is that the calcification is fixable without giving up the protection. Arm C — the ledger plus a judge that can retire a test when a documented requirement change justifies it — matches Arm B's protection almost exactly: the same doubled retention, the same zero rewrite regressions, comparable trap suppression. It does it while spending only 492k tokens, 56% of B's.

The reason it is cheaper as well as gentler is the same reason: letting a judge retire the three genuinely-obsolete storage tests, instead of replaying all of them forever, removes both the wasted token cost of re-checking dead rules and the coercion pressure they exert. The judge is deliberately isolated — it sees only the evidence artifact, the assertion in question, and the new requirement, never the agent's working context — and in the runs its supersession decisions were sound and sparse, concentrated at the moment of the requirement change. Governance here is not bureaucracy; it is the difference between a memory that accumulates and one that is maintained.

Finding 4: every degree of flexibility has a regression cost

This is where a pre-registered prediction was refuted, and the refutation is more interesting than the confirmations. Going in, I predicted that all the governed arms would stay well under the baseline's regression rate. B and C did — zero each. But the two more permissive arms went the other way: C+ committed 6 rewrite regressions and D committed 7, both over the line I had drawn in advance.

The cause is mechanical and, in hindsight, obvious: every time the governance layer retires a test, it also removes the guard that test was providing — and the old bug it was guarding against is free to creep back. The more freely an arm can drop checks, the more old regressions return. C+ (auto-flagging persistent failures) and D (letting the agent lobby for retirement) each loosen the gate, and each pays for it. Arm D posts the highest retention of any arm (0.740) and the most regressions — and, tellingly, 21 "capture" events, where the agent successfully argued the judge into dropping a check. Every other arm scored zero on that, because none of them gave the agent that channel. Give an agent a voice to remove its own tests and it will use it; sometimes that is correct pruning, and sometimes it is an agent talking its way out of a constraint it found inconvenient.

The honest reading of the middle of the table is restraint: the retention differences among the four ledgered arms (0.68–0.74) are small relative to seed-to-seed variance. The large, trustworthy effect is ledger-versus-baseline. The ranking of one ledger variant over another is not something I would build a claim on from this campaign — and saying so is part of the point.

What the pre-registration bought

A second prediction also failed: I expected the ledgered arms to cling to dead code — a config fallback made obsolete by the later requirement change — more than the baseline. Instead every arm, baseline included, left the deprecated fallback in place in 100% of runs. It is a flat ceiling, not a difference: at this scale the models essentially never remove a deprecated-but-still-working path, so the metric cannot tell the arms apart. Prediction unsupported, and only visible as a clean null because the expectation was written down first.

The caveats are load-bearing and I would rather state them than have them found. The campaign ran on a small, cheap model (deepseek-v4-flash); a strong-model replication is the obvious next step, and the harness and tasks are open for exactly that. It is one project, one 24-task sequence, twelve seeds per arm, and no formal statistics. Treat the numbers as the shape of an effect, not a leaderboard.

What I take from it, building agent systems that run for months rather than minutes: an append-only memory is the easy 80% and the dangerous last 20%. Replaying past requirements is what stops silent drift, and it is worth doing. But a memory that can only accumulate will eventually defend something that is no longer true, at real cost in both tokens and correctness. The missing institution is not more memory — it is a cheap, evidence-bound, isolated way to retire it, and a wary eye on who gets to pull that lever.

LedgerBench is open

Harness, tasks, frozen results, and the full pre-registration.

View on GitHub
The interesting question for long-running agents isn't how much they can remember — it's whether they can be trusted to let go of what stopped being true. Build the retirement path with the same care as the memory itself.
© 2026 Pankaj Kumar · Enterprise AI & Logistics Intelligence