Stale by Default: Why Agents Act on Superseded Data
Retrieval systems rank by similarity. Nothing in that ranking function knows which of two near-identical clauses is currently in force.
Most discussion of retrieval failure is about relevance: the wrong chunk, a missed domain term, the answer buried at position nine. Those problems are real, and broadly solvable with better engineering.
There is a second failure mode that is quieter and considerably worse: the system retrieves exactly the right document, and the document is no longer true.
Nothing looks wrong when this happens. Recall is high, the cited passage genuinely supports the answer, the trace is clean. The agent is reasoning confidently over a policy clause replaced four months ago, or a rate that expired last contract period, and no signal anywhere in the pipeline says so.
Three pieces of work from the last week converge on this from different directions: a preprint measuring how often agents check whether inherited memory is still valid, an engineering write-up on making memory invalidate itself, and a commercial dispute in US freight that quietly degraded a feed much of the industry depends on. Together they make a case worth stating plainly — temporal validity is not a data-quality concern to clean up later. It is part of the retrieval contract, and most of us are not modelling it.
The failure has a measurable shape
A preprint posted to arXiv on 26 August sets up a deliberately narrow experiment. An agent inherits memory containing a constraint that was accurate when written but has since been superseded by a newer authoritative record. It gets a fixed verification budget — two record inspections — and has to make a decision.
The agents inspected the constraint's provenance path in roughly one episode in five. When the constraint had in fact been superseded, they produced stale-consistent decisions in 77.3% of episodes in the primary run, and 74.7% in both a re-worded replication and a held-out domain. Three-quarters of decisions, made against a record the system itself could have discovered was obsolete.
The interesting half is the fix. Re-assigning just one of the two budget slots to the provenance path — not increasing the budget, just spending it differently — raised current-record-consistent decisions by 74.0, 72.7 and 61.3 points across the three runs, positive in six of six models, while leaving already-correct cases untouched. The agents had enough budget the whole time. They were spending all of it gathering more evidence and none of it asking whether the evidence they already had was still valid.
This is one study on a synthetic setup; I would not treat the exact percentages as load-bearing. The direction matches what I have seen: given a choice between retrieving more and verifying what it already has, a model retrieves more. Nothing in the default loop rewards checking.
Embeddings make supersession worse, not better
In the procurement decision engine I have been building, the policy agent runs FAISS retrieval over internal policy and contract documents. Semantic search is the right method there — unstructured prose, paraphrased queries, nothing that fits a table.
What took me longer to internalise is that supersession is adversarial to nearest-neighbour retrieval specifically. A revised clause is usually a rewording of the original: same subject, same terminology, same structure, one changed threshold or one added exception. In embedding space the two sit almost on top of each other — and the old one is often the closer match to a query phrased in the language people have used for years, because that is the language of the version they have been reading.
So the ranking function does not merely fail to prefer the current version — under realistic phrasing it can actively prefer the superseded one. A relevance metric scores that retrieval as a success, and an answer-support metric does too, because the passage genuinely supports the answer. Both are measuring the wrong thing.
There is no reranker configuration that fixes this. Recency is not a property the retriever can infer from the text; it is metadata, and if it is not attached to the chunk it does not exist as far as the pipeline is concerned.
The structured path got this right by accident
The rate agent in the same system does not use vectors at all. Historical rate data is structured — lane, equipment type, effective period, value — so the right method is a deterministic lookup against the nearest matching record. I have argued that on general grounds before: match the retrieval method to the shape of the data, and do not pay for vectors where a table lookup is exact.
The part I did not appreciate at the time is that this also solved the staleness problem for that agent, for free. Rate records carry validity windows because that is how commercial rate data is structured. A lookup keyed on lane and date cannot return an expired rate, because the date is part of the key rather than a hint the ranker may or may not weigh. Temporal correctness came bundled with determinism.
The vector path got no such guarantee, and I did not go looking for one. Two agents, one architecture, and only one of them had any notion of time — because in one case the data format enforced it and in the other nothing did.
I suspect that asymmetry is common. Teams inherit temporal discipline wherever they touch relational data and lose it the moment they move to a document store — and because both paths return plausible answers, nobody notices which one is checking.
Feeds degrade before they fail
A freight story from the same week makes the operational version concrete. In a dispute over data access, ELD provider Motive throttled the API access of Highway — a carrier-vetting platform sitting behind roughly 80% of US brokered loads, holding insurance certificates for more than 175,000 carriers. Highway refused the payment demand and withdrew its performance guarantee for carriers on that equipment.
The detail worth sitting with: this was a reduction in refresh frequency, not a disconnection. Carriers stayed visible. Every API call kept returning 200. What changed was how old the data behind those responses was.
A decision system built on that feed will not notice. Health checks pass, no exception is raised, no retry fires, and the agent evaluates a carrier against an insurance certificate hours or days behind reality. In freight that means a load tendered to a carrier whose coverage lapsed; in procurement, an award made against a rate renegotiated last week.
Availability monitoring answers whether the data arrived. It says nothing about whether the data is current. Those are different questions, and most stacks instrument only the first.
What to actually build
LangChain published an engineering note on 25 August describing how they handle this in OpenWiki, and the shape of their answer is the one I would reach for. Memory is stored as claims, each binding a statement to a specific versioned piece of evidence. Staleness detection is then a deterministic version comparison against the current source — no model calls — which keeps it cheap across thousands of claims. Only flagged claims go to a model for repair.
Their reported numbers on a 2,000-claim evaluation are 97.8% supported and 0.5% stale, against a 92.9% and 3.5% baseline. More telling is the recovery test: after a change invalidated 17% of claims, the next checkpoint was back to 0% stale. Unresolved claims persist as unresolved rather than being silently dropped, which is the right default.
In operational terms:
Make validity a filter, not a feature. Effective dates, supersession pointers and document version go into the metadata filter that runs before similarity search, in the same position as counterparty and document type. If a clause is out of force for the query date, it should never reach the ranker.
Separate the cheap check from the expensive one. Deciding whether a claim is stale is a version comparison. Deciding what to do about it needs a model. Conflating them means paying LLM prices to answer a question a string equality could have answered — the same argument for using static rules wherever they suffice.
Spend part of the verification budget on provenance. If your agent has a bounded number of lookups, one should establish whether the evidence is current rather than fetch a fourth supporting document. That was the highest-leverage change in the arXiv study, and it costs nothing extra.
Instrument age, not just availability. Log the age of every retrieved record in the trace. Set sanity bounds on it as you would any other domain quantity, and fail loudly when a feed that should refresh hourly has not moved in a day.
Bounding the model in time
The framing I keep coming back to for production agents is that you bound the model from both sides: schema validation and domain sanity bounds on what comes out, scoped and validated context on what goes in. Staleness is that same discipline applied to a dimension I had been treating as someone else's problem. The input side of that boundary has a time axis, and leaving it unbounded means the model is free to reason correctly over inputs that stopped being true.
None of this is exotic engineering. Effective dates, version pointers and freshness thresholds are things any data team already knows how to model. The gap is that retrieval pipelines were built to optimise relevance, evaluation harnesses to measure it, and a superseded document scores well on every metric in that chain.
An agent that retrieves the right document and acts on the wrong version of it is not making a retrieval error. It is making a decision the system had all the information to prevent, and chose not to check.