A Stopwatch Claim With Three Words Missing

Here is the sentence, in full, from OpenAI’s own GPT-6 Astra announcement: “In latency simulations on OSWorld 2.0, Astra achieves higher computer-use performance in about 47% less time per task than GPT-5.6 Sol, scoring 72.6% at roughly 40 minutes per task, compared with 65.7% at roughly 75 minutes” [1]. Nothing in that sentence is fabricated. The arithmetic even checks out on its own terms: a drop from 75 minutes to 40 is a 46.7% reduction, which rounds cleanly to the “about 47%” the sentence claims.

Two stacks of index-sized task cards side by side, one visibly thicker than the other, a small mechanical stopwatch resting between them with its button mid-press
Figure 1. "About 47% less time per task" is arithmetic on two real numbers — 72.6% at roughly 40 minutes against 65.7% at roughly 75 — not a disclosed rate on any single, comparable unit of work [@openai-gpt-6-astra-announcement].Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Five lines below that sentence, in the announcement’s own benchmark table, the identical two numbers appear again — attached, this time, to three words the prose sentence never uses: “OSWorld 2.0 (v2026.08.08, offline set, partial score) 72.6% 65.7%” [1]. A version tag. A named subset. And a scoring method that is not “did the model finish the task,” which is what an unqualified reader would assume “scoring 72.6%” means, but something else entirely — a fraction of an average of 27.25 separate graded checkpoints per task, confirmed directly against OSWorld 2.0’s own published methodology [2] [3]. None of the three qualifiers is secret. All three are one table row away from the sentence that omits them. The question this piece actually audits is narrower and more checkable than “is 72.6% real”: it is what happens when you follow the citation number OpenAI itself attached to that sentence, and whether what you find there tells you what the missing words mean.

OSWorld 2.0 Was Built to Punish the Old Way of Scoring Computer Use

Start with what OSWorld 2.0 is actually built to do, because the benchmark’s own design is the reason “task” cannot mean what the headline sentence implies it means. The benchmark’s own project page describes 108 long-horizon workflows spanning “seven professional domains and 21 sub-categories, covering research, creative production, engineering, personal services, business and finance, administration and compliance, and healthcare workflows” [2]. These are not the short, single-goal tasks that made earlier computer-use benchmarks tractable to score as pass or fail. The project’s own numbers describe a median human completion time of about 1.6 hours per task, with 69.6% of tasks taking a skilled human longer than an hour, and Claude Opus 4.7, run with maximum thinking, needing an average of roughly 318 tool calls to attempt one — against roughly 30 tool calls for a task on the original OSWorld benchmark this one replaced [2] [4]. A benchmark built around hour-long, multi-application workflows cannot honestly report “did the model succeed” as a single bit without throwing away almost everything an evaluator actually learns from watching an agent work through most of a task and then stall.

ADVERTISEMENT

So OSWorld 2.0 does not report a single bit. Its own published methodology defines two distinct metrics side by side: a binary-completion rate, which requires every one of a task’s checkpoints to pass and which the benchmark’s authors themselves call the primary metric, and a partial score — “the fraction of checkpoints reached,” averaged across an average of 27.25 task-specific checkpoints per task [2] [3]. The arXiv paper’s own reported numbers show exactly why the second metric exists at all: on the benchmark’s official harness, Claude Opus 4.8 reaches 20.6% binary completion and 54.8% partial score on the same run [3] — a 34-point gap between “finished” and “made real, gradeable progress,” on the very model the paper’s own authors used to validate the benchmark. A tracker that logs public OSWorld 2.0 results states the same pattern from the other side: across every model it has recorded, the best publicly tracked binary-completion rate on the full official harness is 32.0%, while partial scores across the same field routinely clear 75% [4]. Binary completion on a 1.6-hour, 318-tool-call task is rare enough, across every lab’s model, that ranking by it would flatten almost the entire field to noise near zero. Partial credit, on this reading, is not a shortcut around a hard question. It is closer to the only metric on this particular benchmark that currently tells anyone anything.

That is a real, defensible design choice, and it means “72.6%” was never going to be a completion rate for anyone reading OSWorld 2.0’s own materials. It is a checkpoint-fraction average, on a task unit that OSWorld 2.0 itself treats as radically finer-grained than “task” ordinarily implies. The problem is not that OpenAI reported a partial score. The problem is what happens between the table that names it a partial score and the sentence built on top of that table.

The Footnote Bolted to the Headline Sentence Defines the Wrong Qualifier

OpenAI’s headline sentence carries a citation marker — a small numeral 3, linking to a footnote, positioned directly at the end of the clause making the claim [1]. A reader who follows that marker, the way a citation is supposed to be followed, reaches this text, verbatim, from OpenAI’s own footnotes section: “OSWorld V2-Offline is a subset of the original OSWorld V2 that works without internet access. Claude model performance on OSWorld-V2 Offline was reproduced by the authors on the official leaderboard. On OSWorld 2.0, the scores for Claude use the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card” [1]. That named document is real and checkable on its own terms: Anthropic’s own system card for the Fable 5.1 and Mythos 5.1 release states that its OSWorld 2.0 numbers run on “the benchmark authors’ August 2026 task release” and that, “because the task files differ, these results supersede the OSWorld 2.0 figures in the Claude Opus 5 System Card and are not directly comparable to results reported on earlier releases of the benchmark” [9] — confirmation, from Anthropic’s side of the same footnote, that the task set alone had already moved between one Claude release and the next. The same footnote number, footnote 3, is attached a second time to the table row two lines below — the row that actually carries the words “offline set” and “partial score” printed as a parenthetical next to the benchmark’s name.

Here is the audit, laid out plainly:

Where the number appears What it says What it discloses
Prose sentence (the claim) “scoring 72.6% at roughly 40 minutes per task” Nothing about version, subset, or scoring method
Table row “OSWorld 2.0 (v2026.08.08, offline set, partial score) 72.6%” Version tag, subset name, scoring method — all three, in one parenthetical
Footnote 3 (attached to both) Defines “offline” as no-internet-access; explains how Claude’s numbers in the same row were reproduced Defines exactly one of the three words in the table’s parenthetical — “offline” — and says nothing at all about what “partial score” means

The footnote does real work. It tells a careful reader that “offline” means the subset excludes tasks needing live internet access, and it discloses — honestly, and to OpenAI’s credit — that the Claude figures sharing that row were run under the benchmark’s own official settings rather than a modified version from Anthropic’s system card, a transparency move that matters and that a less careful announcement would have skipped. What it does not do, anywhere in its two sentences, is define or even gesture at “partial score.” A reader who does everything right — sees the citation marker, follows it, reads the footnote in full — comes away knowing what “offline” excludes and knowing Claude’s numbers are apples-to-apples on that axis. They come away with no more information than before about whether 72.6% means “Astra finished nearly three-quarters of these tasks” or “Astra averaged three-quarters of the graded checkpoints across tasks it may or may not have finished.” Those are very different claims about the same model, and the specific piece of paper OpenAI attached to the sentence to support it answers the wrong one of the two questions a reader would actually want answered.

ADVERTISEMENT
A single task card with a large printed score number on its face, its bottom corner lifted to reveal dense small print underneath that the flat face does not show
Figure 2. The table's own parenthetical — version, offline set, partial score — is real; it just is not the sentence built on top of it [@openai-gpt-6-astra-announcement].Image prompt and art direction by Brecht Corbeel; image generated to that direction.

An “Offline” Subset Is Not the Same Claim as an Easy One

There is a second, separate question sitting underneath the first: does “offline” quietly mean “easier”? It is a reasonable suspicion. A subset defined by excluding every task that requires live internet access is, on its face, excluding exactly the category of unpredictability — rate limits, layout drift, CAPTCHA walls, a page that has changed since the benchmark was built — that would seem hardest for an agent to handle reliably. If that suspicion is right, then reporting a headline number on the offline subset, without saying so in the prose, would inflate the number in a specific, predictable direction, on top of the units problem already established above.

OSWorld 2.0’s own materials do not settle this either way. Neither the project’s official page nor its arXiv paper describes the offline subset as easier, harder, or comparable in difficulty to the full task set — the term “offline” does not appear anywhere in the arXiv paper’s own text at all [3]; it is a distinction OpenAI’s own footnote introduces and defines for its own table, not a difficulty tier the benchmark’s authors have themselves characterized. So this piece went looking for the one thing that could settle it empirically instead of definitionally: a case where the same model has a publicly reported score under both conditions.

There is exactly one. A public benchmark tracker that records a primary source for every row it publishes lists Claude Opus 5 — an Anthropic model with no relationship to Astra beyond sharing this one table — twice on OSWorld 2.0: at 70.6%, self-reported by Anthropic as “a five-run first-attempt average at 1080p, 500 steps, Opus 4.8 grader” [4], a description this piece confirmed against Anthropic’s own system card, which reports Opus 5 achieving “an OSWorld 2.0 score of 70.57% (first-attempt success rate, averaged over five runs)” under the benchmark’s default 1080p resolution, a 500-action-step cap, and Opus 4.8 as grader [8] — and separately at 70.2%, which the same tracker identifies as “OpenAI’s launch table['s]… author-reproduced run on the offline subset” [4] — a figure this piece confirmed directly against OpenAI’s own table, where the Claude Opus 5 column on the same row carrying “72.6% 65.7%” for Astra and Sol reads 70.2%, with the same footnote 3 marker attached to it [1]. If the offline subset were meaningfully easier, the same model’s offline score should read higher than its official-settings score. It does not. It reads 0.4 percentage points lower.

This is one model, one comparison, and it deserves exactly the hedging that scale implies: Anthropic’s 70.6% and OpenAI’s 70.2% were not run as a controlled pair, may not share an identical tool-call budget or grader calibration beyond both citing OSWorld 2.0’s official settings, and a 0.4-point gap on a partial-credit metric with real run-to-run variance proves less than a large, clean gap would. It is not proof that the offline subset is equally hard, still less that it is harder. What it is is evidence — the only piece of evidence either OSWorld 2.0’s own materials or the public tracking of its results currently supplies — and what it shows is that the plain assumption behind the word “easiest,” including in this article’s own title, is not something OpenAI’s materials establish and not something the one real data point available confirms. The honest position is not “the offline subset is easy” or “the offline subset is hard.” It is that nobody, including this piece, has the evidence to say, and the word “offline” in OpenAI’s own footnote was never a claim about difficulty in the first place — only about internet access.

The experiment that would actually settle the “easiest subset” question is simple to describe and, as far as this piece could establish, nobody has published it: run one model, at one fixed reasoning effort, under one grader, once on OSWorld 2.0’s offline subset and once on its full official task set, and report both partial scores side by side. Neither OpenAI, Anthropic, nor OSWorld 2.0’s own authors appear to have published that specific paired result for any model as of this writing. Until someone does, “offline” should be read as a documented fact about internet access and an open, currently untestable question about difficulty — not as the quiet efficiency booster a skeptical reader might otherwise assume it to be.

A sharper, better-documented problem sits directly next to that unresolved one. The same tracker’s own methodology notes state plainly: “Release mixing is the main trap: the 06.24 and 08.08 releases change tasks and grading, OpenAI’s 08.08 numbers use the offline subset, and Anthropic states its Fable 5.1 run modified tasks and grading” [4]. OSWorld 2.0, as a benchmark, is not a fixed artifact. It has already been revised at least once in 2026, and that revision changed which tasks are on it and how they are graded — meaning “v2026.08.08” in OpenAI’s table parenthetical is not decorative version-control trivia. It names a specific, dated snapshot of a benchmark that had already changed once before that date, and comparing a “08.08” number to any number run under an earlier release, or under a different lab’s modified task set, is a documented trap according to the people who track these scores for a living. OpenAI’s own table does the responsible thing by printing the version tag at all. The headline sentence, once again, does not carry it forward.

ADVERTISEMENT
A shelf of identical ring-binders labelled by date, one older binder partway pulled out while a newer-dated binder is partway slid into the adjacent slot
Figure 3. A public benchmark tracker's own reading: OSWorld 2.0's June and August 2026 releases "change tasks and grading," so a bare version tag is doing real work in that table cell [@steel-osworld-leaderboard].Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The Mind2Web Sentence Fails the Same Audit, Plus an Admission About Edited Clips

A second computer-use claim in the same announcement fails a version of the identical audit, with an added wrinkle. “Alongside Astra, we are also updating the Codex harness to significantly improve the speed of computer use. Combined with Astra’s efficiency, this translates to a 1.9x faster task completion compared to the current GPT-5.6 Sol experience, on the Mind2Web benchmark” [1]. That sentence carries its own footnote marker, footnote 4. Follow it, and OpenAI’s own text says this, in full: “Model times are the reported elapsed times for the corresponding demonstration runs. The displayed clips are edited excerpts” [1].

Read plainly, that is OpenAI disclosing, in its own footnote, that the 1.9x figure comes from timing two demonstration runs — not a benchmark leaderboard score, not an averaged evaluation across a task suite, but the elapsed time of a specific recorded run of each model — and that the video evidence a viewer would actually watch alongside the claim has been edited. Neither of those facts makes the 1.9x number false. A single timed demonstration can be a real, honestly reported data point. But it is a categorically different kind of evidence than the OSWorld 2.0 table above it, and the prose sentence presents both claims — a checkpoint-fraction average across 108 long-horizon tasks, and the elapsed time of one edited demonstration clip — in the same register, back to back, with the same unqualified confidence.

There is a second gap here that OpenAI’s own materials never close. “Mind2Web” is not one fixed thing. The name originates with a 2023 static dataset and evaluation protocol for web agents, built from human demonstration traces across real websites [5], and has since been extended by other research groups into at least two further, materially different evaluation setups — an online variant that grades an agent against intermediate “key node” checkpoints on live sites rather than a single final answer [6], and a separate live-web variant, built specifically because static evaluation was found to overstate agent competence, that tests 300 tasks across 136 real websites still actively changing rather than cached [7]. OpenAI’s announcement text names none of these specifically. It says “the Mind2Web benchmark,” full stop, with no version number, no citation to which paper or leaderboard defines the specific task set Astra and Sol were timed against. This piece could not resolve, from OpenAI’s own materials or from any tier-1 coverage of the announcement, which of the several things now called Mind2Web that sentence refers to. That is not a claim this piece can make more precise than OpenAI’s own text allows — it is a gap in OpenAI’s own disclosure, stated here as exactly that: unresolved, and not filled in by inference.

A length of demonstration-reel film laid across an illuminated light table with a visible splice mark partway along it, a loupe resting nearby
Figure 4. OpenAI's own footnote on the Mind2Web speed claim: reported times are "elapsed times for the corresponding demonstration runs," and "the displayed clips are edited excerpts" [@openai-gpt-6-astra-announcement].Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Partial Credit Isn’t the Problem; an Unlabeled Unit Is

The fair version of the opposing case deserves to be made in full, because it is a real case and not a strawman. Partial-credit scoring on a long-horizon benchmark is not a loophole; it is, on the evidence above, closer to a necessity. A benchmark where every model’s binary-completion rate sits under a third, and where the best-performing systems separate cleanly on a checkpoint-fraction metric that binary completion cannot distinguish at all, is a benchmark where reporting only the pass/fail number would tell a reader less, not more. OSWorld 2.0’s own authors built partial scoring into the benchmark from the start, not as an afterthought to flatter weak results, and a public tracker with no stake in any vendor’s outcome ranks by the same metric for the same stated reason [3] [4]. None of the three disclosures this piece is auditing — version, subset, scoring method — were invented or hidden by OpenAI. All three sit in the same table, one paragraph below the sentence that omits them, available to anyone who reads past the headline.

That is exactly the distinction worth holding onto: this is not a story about a number that was faked, gamed, or concealed. It is a story about a number that was disclosed precisely, in one place, and then restated imprecisely, one paragraph later, in the sentence every reader who does not click through to the table will actually carry away. A partial score on a checkpoint-graded, version-tagged, internet-excluded subset of a 108-task benchmark is a specific, legitimate, well-defined measurement. “Scoring 72.6%… at roughly 40 minutes per task” is not that measurement. It is what that measurement sounds like once its three defining conditions have been quietly set down between the table and the prose.

What 72.6% Can Honestly Be Said to Measure

Here, then, is what survives the audit. Astra really does average a higher fraction of OSWorld 2.0’s per-task checkpoints, on the offline subset of the benchmark’s August 2026 release, in less simulated elapsed time, than GPT-5.6 Sol does under the same conditions — a real, specific, checkable result, sourced to OpenAI’s own table and confirmed against OSWorld 2.0’s own published methodology [1] [2]. What does not survive contact with that same table, footnote, and the wider record around the benchmark it cites is the unqualified sentence built on top of it: not because the sentence’s two numbers are wrong, but because “task,” left undefined, invites a reader to imagine a coin-flip unit — done or not done — that this specific benchmark does not use and was not designed to use, and because “offline” and “v2026.08.08” name conditions that this piece could not confirm make the number easier, only conditions that OpenAI’s own footnote and a public tracker’s methodology notes agree are worth knowing about before repeating the number anywhere else.

A workable reading convention falls directly out of that gap, and it is a narrow, proposed one rather than a sweeping rule: any computer-use benchmark claim that pairs a percentage with a stated time-per-task should be read as carrying, by default, whatever qualifiers sit in the source table’s own parenthetical next to the benchmark’s name — version, subset, and scoring method chief among them — whether or not the sentence stating the claim repeats them, and a reader who cannot find that parenthetical at all should treat its absence as the open question, not as an answer. Applied here, that convention does not overturn OpenAI’s claim. It resizes it — from “Astra completes computer-use tasks correctly 72.6% of the time” to something closer to “on one dated, internet-excluding slice of a 108-task, checkpoint-graded benchmark, Astra’s per-task average share of completed checkpoints outpaced Sol’s by a wide margin, in noticeably less simulated time” — true, specific, and considerably harder to misquote than the sentence OpenAI actually printed.

The stakes of that resizing are not academic. A reader who takes “72.6% at roughly 40 minutes per task” at face value is being handed a number shaped like a completion rate on ordinary office work — the kind of figure that ends up in a procurement slide or a cost projection two clicks away from this announcement. A reader who instead carries forward the resized version knows they are looking at checkpoint-fraction progress on a benchmark whose own task set was revised mid-year, on a subset nobody has shown is representative of the harder, internet-facing work most real computer-use deployments actually involve. Both readers cite the same footnote number. Only one of them has actually read what it says.