There is a specific kind of article that arrives every few weeks. A study has found that AI models cannot do some particular thing — reason reliably, plan more than a few steps ahead, avoid a certain category of mistake. The finding is real. The methodology is often good. The write-up is careful.
And it is describing a system that, by the time you read about it, has been replaced twice.
This is not a criticism of the researchers. It is a structural mismatch between two clocks that nobody designed to run together, and until you account for it, you will keep drawing confident conclusions from evidence that has quietly expired.
Two clocks
Serious research runs on a cycle measured in seasons. You design the study, you run it, you write it up, you submit it, you wait for review, you revise, it appears. Even a preprint — which skips most of that — carries months between the experiments and the moment anyone reads about them. Then journalism adds its own lag on top, and the study reaches a general audience later still.
Frontier model releases do not run on that cycle. A capable system can be superseded by a substantially different one inside a year, sometimes considerably faster, and the differences between generations are frequently concentrated in exactly the weaknesses the previous generation was criticised for.
So the pipeline reliably produces a specific artefact: a rigorous, well-conducted, carefully caveated finding about a system that is no longer the thing anyone is using.
This cuts in both directions
It is tempting to read the above as a defence of AI, and that is a misreading worth heading off, because the same lag produces inflated claims just as efficiently.
The optimistic version is the demo that circulates for months after the conditions that produced it stopped being reproducible. The benchmark result that was achieved with a scaffolding, a prompt strategy and a compute budget that no ordinary user has. The capability claim that was true of a research configuration and never true of the shipped product.
Stale evidence is not systematically biased toward pessimism or optimism. It is biased toward whatever was true a while ago, which in a fast-moving field means it is simply wrong in an unpredictable direction. That is worse than a known bias, because you cannot correct for it by leaning the other way.
The failure this actually causes
The practical damage is not that people believe one wrong thing. It is that people make decisions on a snapshot and then stop looking.
An organisation evaluates a model for a workflow, finds it insufficient, and files the question as answered. Two years later the answer has changed and nobody revisits it, because the evaluation felt rigorous at the time and rigour feels durable. The opposite happens too: a capability looked achievable in a demo, a project was committed to it, and the shipped reality never arrived.
Both are the same error. Both come from treating a measurement as a property of the technology rather than as a property of one system on one date.
What to actually do about it
I do not think the answer is to ignore research. It is the best evidence available and the alternative is vibes. But it needs handling differently from how most coverage handles it.
Read the date before the finding. Not the publication date — the date the experiments were run, which is usually buried in the methodology and is frequently much earlier. That is the date the claim is actually about.
Check which system was tested, specifically. "AI models struggle with X" is almost never what the paper says. The paper says a named version of a named model, under stated conditions, struggled with X. The generalisation is added afterwards, usually by someone else.
Distinguish claims about the ceiling from claims about the floor. "This architecture cannot in principle do X" is a strong claim that ages slowly, and is rarely what has been demonstrated. "This system did not do X" is a weak claim that ages in months. They get reported in the same sentence and they are not remotely the same thing.
Re-run your own evaluation on a schedule. If a capability question matters to a decision you have made, it deserves a calendar entry, not a conclusion. The cost of re-testing is a few hours. The cost of being two years stale on something load-bearing is considerably higher.
Treat your own experience as evidence, and date it too. Personal experience of these systems goes stale exactly as fast as published research does, and it feels far more reliable than it is. "I tried this and it could not do it" carries an implicit timestamp that people forget to attach.
The uncomfortable part
There is no version of this problem that gets solved by better journalism or faster review. The gap is structural. As long as capability changes faster than evidence can be produced about it, the public record will describe a past state of the world and present it in the present tense.
Which means the epistemic burden lands on you, and it is not a comfortable place for it to land. The honest position on most specific AI capability questions is that the last time anyone checked properly was a while ago, the answer has probably moved, and nobody can tell you in which direction without checking again.
That is a genuinely unsatisfying conclusion. It is also, as far as I can tell, the accurate one — and I would rather hand you an unsatisfying accurate answer than a confident stale one.
This piece expands on a video of the same argument. If working out what is actually true for your own situation is the problem, that is the thing I help with.