The Measurement Gap
IntroPart IPart IIPart IIIPart IVAppendices

The Measurement Gap

Situational Awareness, two years on, and what to build after ARC.

62.7%
99.9%

GPT-6 Astra on ARC-AGI-3, September 3, 2026. Same model, same weights, same games. The difference was the harness.

Srihari MysoreSeptember 2026Five parts, about 44 minutes

This essay was written with Claude Fable 5.1 as a research and drafting collaborator; the argument, the editing, and the responsibility for errors are the author's. Figures were checked against primary sources as of September 7, 2026. The benchmark proposed in Part III has not been run. Predictions in sections 3.8 and 4.5 are dated so they can be graded.

About 2 min

0. The week of September 1

In the first three days of September 2026, two things happened that had been predicted and one thing happened that had not.

The predicted things: Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 on September 1, the same model under two safeguard regimes. OpenAI shipped GPT-6 Astra on September 3, trained on more than a hundred thousand GPUs at its Stargate site in Texas, the first model it has ever classified at the Critical tier for cybersecurity under its own Preparedness Framework. Both labs used the phrase “most aligned model ever.” Both launches were slowed, by weeks or months, over the cyber capabilities of the models. Anyone who read Leopold Aschenbrenner’s Situational Awareness in June 2024 and took its central arithmetic seriously would have expected a September like this one. The compute arrived. The capabilities arrived. The security problem arrived.

The unpredicted thing was smaller and, I will argue, more important.

On September 3, ARC Prize published its evaluation of Astra on ARC-AGI-3, the interactive reasoning benchmark it had launched six months earlier as the successor to a series that had resisted frontier models for five years. At launch in March, every frontier model scored under one percent. On September 3, Astra scored 99.9 percent.

Except that it also scored 62.7 percent.

Same model. Same weights. Same games. The difference was the harness: the thin layer of software that decides what the model gets to remember between turns. Under ARC’s Standard harness, where the model carries forward only the notes it chooses to write, Astra at maximum reasoning effort solved 62.7 percent. Under a Provider Adapter harness, where OpenAI’s own opaque reasoning state persists between calls and long conversations get compacted, it solved 99.9 percent, faster, for less money.

OpenAI’s president said “Welcome to the AGI era.” ARC Prize, in the same news cycle, wrote that it was not claiming Astra is AGI. Both statements were made by careful people looking at the same model. They were not disagreeing about what Astra can do. They were disagreeing about what the word means, and about what the number means.

That is the thesis of this essay. For most of the last decade the interesting question about AI was a capability question: can the models do X yet? Situational Awareness was the most forceful statement of that framing, and on its own terms it mostly held up. But the capability question has now been answered often enough, and ambiguously enough, that it has turned into a different question. We no longer have a capability gap that a benchmark can measure. We have a measurement gap, and no benchmark that can close it.

The rest of this essay does three things. Part I re-reads Situational Awareness against September 2026 and argues that its arithmetic was right and its conclusion was underspecified. Part II takes apart what ARC-AGI-3 actually measured, using the Astra numbers, and derives a quantity I think every leaderboard should be forced to report. Part III proposes what to build next: a benchmark designed for the thing the current generation of models has not been shown to do, with a metric, a human baseline, and a prototype small enough for one person to run.