A Footnote Attaches to a Table Cell, Never to the Sentence Above It
OpenAI announced GPT-6 Astra on September 3, 2026, rolling it out first to customers in its cybersecurity program before opening it, over the following week, to ChatGPT’s paid plans and the API [8]. OpenAI’s own launch page for GPT-6 Astra makes three claims of total conquest in one paragraph, in identical unhedged language: “Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems in mathematics… Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score” [1]. Nothing in that sentence marks the ARC-AGI-3 figure as standing on different ground than the other two. The qualification exists, but OpenAI placed it several sections and roughly two hundred lines further down the same page, attached not to the sentence but to a single cell in a benchmark table: the row reads “ARC-AGI-3 99.9%” with a numeral-1 footnote marker sitting on the number itself, nowhere near the paragraph that already told the reader the benchmark was saturated [1]. Footnote 1, read in full: “On ARC-AGI-3, GPT-6 Astra was run with our responses API harness, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3” [1].
That footnote does something a little unusual for a hedge: it discloses that a setting was changed, and it states the change was not aimed at this particular test, but it never says what Astra scored without the change. A reader who stops at OpenAI’s own page has no way to know whether the unstated number is 99.8% or 40%. Verdict first: Astra’s 99.9% on ARC-AGI-3 is real. So is its 62.7%. Both are the same model taking the same benchmark; they differ only in which harness sat between the model and the test, and how much of that harness’s own memory OpenAI’s product settings were permitted to keep.
The missing number is not a secret — it just isn’t on OpenAI’s page. ARC Prize, the nonprofit that designs, scores, and maintains the ARC-AGI benchmark series and the only party with standing to publish an independent read on what Astra actually did, put it in a post of its own the same week, authored by ARC Prize’s Greg Kamradt and dated September 3, 2026: “With our Standard harness, OpenAI’s Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K. With the Provider Adapter harness, Astra (high) scores 99.9% for $19K. Both are state-of-the-art scores” [2]. Everything that follows in this piece stays inside those two organizations’ own published numbers — no aggregator, no press paraphrase, no number that doesn’t trace to arcprize.org or openai.com directly.
ARC-AGI-3 Was Built to Catch a Machine That Cannot Learn a New Room
The benchmark this argument runs on is not a trivia set. ARC Prize built ARC-AGI-3 as “a benchmark for studying agentic intelligence through novel, abstract, turn-based environments,” where “agents must explore, infer goals, and build internal models of environments to effectively plan actions without explicit instructions” [2, 6]. ARC Prize launched the benchmark on March 25, 2026, and its own opening scoreboard is the number every later percentage in this piece has to be read against: humans solved 100% of the environments, and frontier AI scored 0.51% [5]. It scores four separate capacities: exploration, since “information is rarely provided passively” and an agent has to obtain it by acting; modeling, turning raw observation into something that predicts what happens next; goal-setting from sparse reward alone; and planning and execution, mapping a route to a goal and correcting course as new information arrives [2]. Human testers, drawn from roughly five hundred members of the general public with no selection for puzzle-solving skill, solve 100% of the environments; the interesting variable is never whether a person can do it, but how efficiently [2].
ARC Prize is explicit that this is the third generation of a series built to track a specific, named quantity: “the goal of the ARC-AGI series is to measure the ‘residual gap’ between current artificial intelligence and AGI. We define AGI as a system’s ability to acquire any skill a human can, as efficiently as a human can” [2]. That framing matters for what “saturates” is supposed to mean here. A score near 100% is not a trivia record; it is ARC Prize’s own instrument reporting that the residual gap it was built to find, on this particular test, has mostly closed. Which is exactly why the instrument’s own calibration — what harness stood between the model and the test — carries more weight here than it would on a static knowledge quiz, and why ARC Prize itself, as the next section shows, built two different harnesses on purpose rather than one.
For a sense of what “efficiently as a human can” is actually asking, ARC Prize built its own baseline out of roughly five hundred untrained participants and priced their time directly: sessions paid “$115 per 90-minute session, plus $5 per game completed,” with participants attempting “approximately nine games per session, roughly $12.78 per attempted game before bonuses” [2]. That figure is ARC Prize’s own rough human-cost comparison point, not a per-game cost this piece derives for Astra — the dollar figures in the harness table below are aggregate totals across the full evaluation run, not a per-game rate, and the two should not be read against each other directly. What the human baseline does support is the efficiency claim ARC Prize separately reports for Astra under the Provider Adapter harness: “Astra (max) used fewer actions than the human baseline on 96.0% of levels and used 51.7% fewer actions per level on average” [2]. ARC Prize calls this “a material milestone,” distinct from and additional to the completion-rate scores this piece is built around.
Two Harnesses, Two Definitions, One Model
ARC Prize did not stumble into publishing two numbers. It ran Astra under two named configurations and defined both in its own post. The “Standard harness… enables a model to carry forward notes it chooses to keep with it throughout the environment” — a deliberately minimal, provider-neutral interface. The “Provider Adapter harness… preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work” [2]. Under the first, Astra tops out at 62.7% (at “max” reasoning effort, for a reported $26,098). Under the second, it reaches 99.9% (at “high” reasoning effort, for $18,817) [2]. ARC Prize’s own summary sentence pairs each harness’s single best result — max under Standard, high under Provider Adapter — which is a legitimate way to report “best observed,” but it is worth naming plainly: the headline 62.7%-versus-99.9% comparison is a best-versus-best pairing across two different reasoning-effort settings, not a like-for-like one. ARC Prize in fact tested six effort levels under each harness and published every cell:
| Reasoning effort | Standard harness | Provider Adapter harness |
|---|---|---|
| max | 62.7%, $26,098 | 98.6%, $17,332 |
| xhigh | 59.3%, $37,317 | 98.4%, $18,147 |
| high | 54.8%, $40,705 | 99.9%, $18,817 |
| medium | 38.6%, $48,090 | 98.4%, $19,285 |
| low | 17.5%, $38,166 | 98.0%, $21,298 |
| none | 35.2%, $49,791 | 96.7%, $23,457 |
[2]
ARC Prize’s own aggregate efficiency claim, stated across the full set of tested conditions rather than any single row: “Astra’s best observed score on ARC-AGI-3 Semi-Private increased from 62.7% to 99.9%. Looking across Public and Semi-Private and all reasoning levels, Provider Adapter runs were approximately 3.66x faster by aggregate recorded elapsed time and used 49% fewer total tokens across the 167 game-reasoning pairs both harnesses solved” [2]. That last clause is a real scope limit worth keeping in view: the speed and token figures are computed only over the subset of game-reasoning pairs both harnesses actually solved, not the full test battery — a reasonable choice for an apples-to-apples efficiency comparison, but not the same claim as “faster and cheaper across everything Astra attempted.”
The Adapter Wins at Every Setting, and the Margin Is Not Fixed
The six-row table lets a reader do the comparison ARC Prize’s headline sentence doesn’t: hold reasoning effort constant and read straight across. At every single one of the six levels, the Provider Adapter harness scores higher than the Standard harness — and the size of that gap is not fixed. At max effort it is 35.9 points (98.6% minus 62.7%). At xhigh, 39.1 points. At high — the pairing ARC Prize’s own headline sentence uses for its 99.9% figure — 45.1 points, wider than the 37.2-point gap the headline’s max-versus-high comparison implies. At medium, 59.8 points. At low, 80.5 points. At no reasoning effort at all, 61.5 points. The real, same-effort harness gap for Astra on ARC-AGI-3 runs from 35.9 to 80.5 points depending only on how much the model was asked to think — never zero, and never as narrow as the headline pairing’s 37.2-point figure at its narrowest.
Cost moves the same direction. At every one of the six effort levels, the Provider Adapter run also costs less than the Standard run at that same level — by $8,766 at max, up to $28,805 at medium. The harness that scores higher never once costs more. And within each harness on its own, cost does not rise monotonically with nominal reasoning effort: under the Standard harness, “none” ($49,791) costs more than “max” ($26,098), and under Provider Adapter, “none” ($23,457) costs more than “max” ($17,332) — because, as ARC Prize explains it, “at max reasoning effort, Astra solves games more efficiently, requiring fewer actions and therefore lowering total cost relative to the other reasoning-effort levels” [2]. A model that thinks more per action can still finish the job in fewer actions overall, and on this benchmark that arithmetic beats out the naive “more reasoning tokens equals more spend” intuition in both harnesses at once. None of this is OpenAI’s arithmetic or ARC Prize’s stated conclusion; it is what their own six published rows compute out to when read straight across rather than through the single best-versus-best pairing either party’s prose actually states.
The Gap Is About Memory, Not About Whether Astra Can Actually Play
Before treating the whole 35-to-80-point spread as evidence of harness engineering doing the work capability should be doing, it is worth reading what ARC Prize’s own qualitative review of Astra’s play logs found, independent of either harness’s completion percentage. Reviewing replays, ARC Prize reported that Astra “chose which strategy notes it would like to carry forward,” tracking “objects, coordinates, rules, and unfinished plans” through “a custom domain-specific language notation it generated for the environments” — condensing a game state into lines like “L8: hub q2 (8↓). Lengths: 14=1…” to record a level, a rotation index, and mechanism lengths in a compact, self-invented shorthand [2]. ARC Prize is careful to place this in context — “we’ve seen similar behavior in other models,” it notes, “but Astra’s notes stood out for their precision and information density” [2] — which is a claim about degree, not a claim that no other model has ever taken notes.
ARC Prize also ran a third, separate condition outside the two-harness comparison entirely: the “PRO-LONG” harness, built by an early ARC-AGI-3 red-teaming partner, which gave Astra a sandbox to execute its own code. There, reviewers watched Astra build small, game-specific software libraries on the fly — in one maze-like game with patrolling guards, it wrote a navigation module, then added a combat-rules module, then a patrol-prediction module, then a script to check its own predictions against what it observed [2]. ARC Prize is explicit that PRO-LONG is a different evaluation condition from its controlled human testing — “our testing participants did not have a code interpreter, scratch pad, etc., so PRO-LONG’s results should be understood as the combined performance of the model and its tools” [2] — so it does not belong in either the Standard-versus-Provider-Adapter comparison or the human-baseline efficiency figures above. What it does establish, independent of both harnesses and both headline scores, is that ARC Prize’s own qualitative read does not describe a model coasting on retained context alone: the notation and the tool-building are behaviors a harness change does not manufacture by itself. The harness gap this piece measures sits on top of real, independently observed capability, not in place of it.
OpenAI’s Defense of the Harness Predates the Model It Defends
Footnote 1’s claim — that the Provider Adapter settings “do not specifically target ARC-AGI-3” — rests on a separate OpenAI post: “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark” [3]. That post is not about Astra. It is about GPT-5.6 Sol, and corroborating discussion of its publication dates it to around July 31, 2026 [4] — roughly five weeks before Astra existed as a named, released model. OpenAI’s own account there: “GPT‑5.6 Sol scored just 7.8%” on ARC-AGI-3 overall, a figure that matches the 7.8% Sol row on Astra’s own later comparison table exactly, and on a separate “public” task subset, “with the official harness, GPT‑5.6 Sol scored 13.3%… with retained reasoning and compaction, it scored 38.3%” [3]. The two settings the post names — “retained reasoning” and “compaction” — are the same two OpenAI’s Astra footnote gestures at without naming.
OpenAI’s own framing of why those two settings help is direct: the official ARC-AGI-3 harness “discarded” all private reasoning after each action, so “with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking,” and separately used “a rolling truncation window, causing older actions to become invisible as the history grew” [3]. Turning both off — retaining reasoning across turns, replacing truncation with compaction — let Sol “spend less time thinking before each action” and “learn over time” instead of restarting its understanding of the game on every move [3]. OpenAI states plainly that neither setting was invented for this test: “This is how our models are trained, and also how they are deployed in ChatGPT and Codex” [3]. The post closes with an explicit recommendation to anyone benchmarking OpenAI’s models at all: “if you’re comparing models, we recommend relying on evals that use the settings above, which best match real-world use in ChatGPT and Codex” [3].
What the Explainer Actually Shows, and What Astra’s Footnote Borrows From It
Read precisely, the explainer demonstrates one thing directly and implies a second. Directly demonstrated: on GPT-5.6 Sol, on ARC-AGI-3’s public task set, turning on retained reasoning and compaction took the score from 13.3% to 38.3% — a roughly threefold jump, in line with OpenAI’s own “tripled scores” framing [3]. Not directly demonstrated: the same experiment, run on Astra, on the semi-private set that produced 62.7% and 99.9%. Astra’s footnote borrows the Sol finding by analogy rather than repeating it as a fresh result, and the analogy is a reasonable one — same company, same two named settings, same benchmark family, same Responses API mechanism — but it is an inference, not a rerun. Nothing this session could locate shows OpenAI independently reporting a 13.3%-to-38.3%-style before/after pair for Astra itself; the 62.7%-to-99.9% pair comes only from ARC Prize’s own separately conducted comparison, using ARC Prize’s harness definitions, not OpenAI’s original demonstration repeated on the newer model.
There is a second precision worth adding to “how our models are … deployed in ChatGPT and Codex.” OpenAI’s explainer specifies the mechanism: “for GPT‑5.6, passing the previous response ID automatically retains reasoning across tool calls and turns” [3]. OpenAI’s own Responses API documentation describes the identical mechanism in general terms, tied to no single model or benchmark: chain a previous_response_id through a multi-turn conversation and “the model will automatically have access to all previously produced reasoning items” [7]. Automatic, in that sentence, describes what happens once a caller uses the Responses API and chains a previous response ID through a conversation — which is how OpenAI’s own first-party products are built, and is not a claim that every API integration gets retained reasoning with zero configuration regardless of which endpoint or calling pattern it uses. The honest reading of “real-world performance” here is narrower than “how everyone uses the model”: it is how OpenAI’s own products use it, and how a developer following OpenAI’s own stated recommendation would use it, which is a real and defensible standard for “real-world” but not an unconditional one.
One question this piece cannot close, and flags rather than guesses at: whether an ordinary developer calling Astra through the API, with no special configuration, gets the Provider Adapter behavior by default, or has to deliberately opt into it by chaining response IDs and enabling compaction. Nothing fetched in this session — not the Astra announcement, not the explainer, not ARC Prize’s post — states a default harness for a bare API call the way each states a default for ChatGPT and Codex specifically. That gap matters because it is the exact fact that would settle whether “real-world performance” means most real Astra usage or means OpenAI’s own products specifically. Absent that fact, the more defensible claim is the narrower one already stated above, and this piece states no wider one.
ARC Prize Itself Refuses to Call Either Number the Fake One
The fair reading of all of this is not that OpenAI picked a flattering harness and hid the honest one. ARC Prize built the Standard harness on purpose, and defends it on its own terms: it “asks how models compare under the same minimal, provider-neutral interface… We believe a future AGI should be able to solve ARC-AGI-3 under these conditions. The shared interface also gives us a consistent, apples-to-apples comparison across providers” [2]. That is a real methodological commitment, not a strawman OpenAI’s footnote is quietly working around — a shared, minimal interface is the only way ARC Prize can compare Astra to a rival lab’s model without letting each vendor’s own product engineering decide the outcome. At the same time, ARC Prize labeled both of Astra’s scores “state-of-the-art” in the same sentence, and its going-forward policy is not to pick a winner between the two conditions but to keep reporting both: “we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled” [2]. ARC Prize is also explicit that none of this should be mistaken for a verdict on general intelligence: “when we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent ‘proof of achieving AGI’… we are not claiming that it is AGI” [2] — a restraint OpenAI’s own “saturates” language does not carry over from the benchmark’s maintainer to its own launch page.
Saturating a Benchmark Should Name Which Harness Did the Saturating
None of the two numbers here is manufactured, and neither party disputes the other’s arithmetic. What the record actually supports is narrower and more checkable than “OpenAI gamed a benchmark”: a single sentence on OpenAI’s own launch page states an unqualified 99.9%; the qualification that number needs lives four sections away, attached to a table cell rather than the sentence; the qualification names a mechanism without naming the number it changed; and the benchmark’s own maintainer, tracing the same run through six reasoning-effort settings, shows a gap that never closes and in fact widens at lower effort levels, alongside a cost advantage that runs the same direction at every single row. OpenAI’s defense of the harness choice is genuine and well-evidenced — just evidenced on a different, older model, published weeks before the one it is now cited to justify.
A reader who wants to know what “Astra saturates ARC-AGI-3” should mean, on the evidence gathered here, has a workable answer: under the harness OpenAI’s own products use and recommend, yes, close to saturated; under the harness ARC Prize built specifically so that no vendor’s product engineering could decide the comparison, not yet, and the distance depends heavily on how hard the model is asked to think. Both sentences are true. The one that belongs in a headline should say which harness it is describing, and as of this writing, on OpenAI’s own launch page, it still doesn’t.