Then in LinkedIn: Write article → click into the body → paste (Ctrl+V). Headings, links and images come with it. The title usually pastes as the first line — cut it into LinkedIn's title field. back to the article

OpenAI Buries Its Own Asterisk Two Screens Below the Table It Belongs To

OpenAI's Astra launch page prints two of Fable 5.1's scores using a different Anthropic model's numbers, disclosed only in a footnote a page below the row it corrects. Anthropic's own table runs the same substitution inside the cell holding the number.

A large printed benchmark table on a lightbox, a small numbered adhesive flag reading 17 half-peeled from its backing strip beside one row, its corner still curling rather than pressed flat

The flag names which footnote a row belongs to; it does not say, at the row, what that footnote will turn out to mean [@openai-gpt-6-astra-launch]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Abstract

This article measures how far a reader travels, inside OpenAI's own GPT-6 Astra launch page, between two of Claude Fable 5.1's printed benchmark scores and the footnote disclosing that both numbers actually belong to Claude Mythos 5.1, the same model with its safeguards off. It walks footnote 17 and footnote 11 verbatim, counts the sections and footnote positions a reader crosses to reach each one, and sets that count against Anthropic's own launch-day table, which — as this outlet's Reading the Claude Fable 5.1 Scorecard already reported for the identical Terminal-Bench 4.0 case — ran the equivalent substitution inside the same table cell as the number itself. It does not re-argue that prior piece's account of Anthropic's disclosure, does not audit either company's benchmark numbers on their own scientific terms, and does not claim to know what any reader actually notices or remembers — only what a reader would have to cross, on the page as published, to find out. It also states directly why the distance between a footnote's existence and its distance from the number it qualifies is a judgment call this outlet is making, not a self-evident failure on OpenAI's part.

Two of Fable 5.1’s Numbers on OpenAI’s Table Aren’t Fable 5.1’s

OpenAI’s launch page for GPT-6 Astra carries nine stacked benchmark tables, one after another, under plain category headers: Computer Use, Professional, Coding, Academic, Science and Health, Cybersecurity, Alignment, Long Context, Abstract reasoning [1]. Read the third row of the first table and it says, without qualification, that in the column headed “Claude Fable 5.1,” ScreenSpot-Pro — a no-tools test of whether a model can locate the right element on a screen [6] — scored 92.7% for Astra against 87.3% for Fable 5.1. A small superscript “17” sits beside that 87.3%, the kind of mark a reader’s eye slides past on the way to the next column.

Keep reading. Four table-sections later, under “Science and Health,” the last row is HealthBench Professional, length-adjusted [8]. Three Claude figures run across the row — Fable 5.1 at 58.1%, Fable 5 at 60.9%, Opus 5 at 56.4% — and all three carry a different superscript, “11” [1]. One table-section further on, under “Cybersecurity,” the second row is ExploitGym — a test of whether a model can turn a known software vulnerability into a working exploit. Astra scores 42.4%. Two Claude columns sit beside it: “Claude Fable 5.1” at 30.4% and “Claude Fable 5” at 28.4%. Both carry the same superscript, “17” [1].

Both tables belong to the same launch week. Anthropic published Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026 [2] [4]; OpenAI published GPT-6 Astra, with this benchmark table, in the days immediately after [1] [5]. A table published into a rival’s own launch week, comparing a new model against that rival’s numbers, is doing something specific: it is asking a reader to treat its own columns as a fair, apples-to-apples record of what the rival model actually does. That is the only reason a “Claude Fable 5.1” column exists on OpenAI’s page at all — not as background context, but as the yardstick Astra’s own numbers are being measured against.

Nothing about any of those five cells announces itself as unusual. A reader who stops at the row — which is what a row is for — walks away believing Claude Fable 5.1 scored 87.3% locating screen elements with no tools, and 30.4% turning known vulnerabilities into exploits. Both of those numbers are true of some Claude model, tested by OpenAI, on OpenAI’s own hardware. Neither of them is true of Claude Fable 5.1.

What Footnote 17 and Footnote 11 Actually Say

The page’s footnotes are collected in one block, under one heading, “FOOTNOTES,” after all nine table-sections have run their course. Footnote 17, the last of the seventeen, reads in full: “For ScreenSpot-Pro and ExploitGym, the Fable scores we report come from Mythos, which is Fable with fewer safeguards” [1]. Claude Mythos 5.1 is not a separate model from Claude Fable 5.1 in the way GPT-6 Astra is a separate model from GPT-5.6 Sol — Anthropic’s own launch page states plainly that “Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards,” with Mythos 5.1 released only through Anthropic’s trusted-access programs to vetted institutions working in cybersecurity and the life sciences [2]. So the 87.3% and the 30.4% printed under “Claude Fable 5.1” on OpenAI’s table are not Fable 5.1’s own scores at all. They are what the same weights produce once Anthropic’s own safeguards — the ones that keep Fable 5.1 from developing exploits for the vulnerabilities it can already find — are switched off.

Footnote 11 does a related but more layered job. It reads: “We independently evaluated all Claude models following the intended HealthBench Professional procedure, using GPT-5.4 grading and length-adjusted, unclipped scores. For Fable 5.1, we used Opus 5 fallback for provider refusals” [1]. The footnote’s first sentence applies to all three Claude columns on that row: OpenAI, not Anthropic, ran the eval, and OpenAI’s own GPT-5.4 did the grading rather than whatever grader Anthropic uses on its own benchmark disclosures elsewhere. That much is a general methodology note, not a substitution — a legitimate thing for a competing lab to disclose once, in one place, about how it produced someone else’s numbers. The second sentence is narrower and does the actual swap: specifically for Fable 5.1, whenever the model’s own safeguards caused it to refuse a HealthBench Professional prompt — a live, expected outcome of Anthropic’s own biology and medical-question guardrails — OpenAI substituted an answer from Claude Opus 5 instead of scoring the refusal as a failure or leaving it out. Fable 5.1’s 58.1% is, by OpenAI’s own account four table-sections later, a blend: mostly Fable 5.1’s own answers, patched with Opus 5’s wherever Fable 5.1 declined to answer at all.

The two footnotes are not doing the same kind of substitution, and the difference matters more than it first appears. Footnote 17 swaps in Mythos — the identical weights with the safeguard removed entirely — for a benchmark built around the exact capability the safeguard exists to block. Footnote 11 is gentler: it patches Fable 5.1’s own refusals with a different production model’s actual answers, one question at a time, rather than replacing the whole run with an unsafeguarded twin. That second approach happens to track what Anthropic itself says its production safeguards do. Anthropic’s own launch page states that “in all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5” [2] — and HealthBench Professional, a medical-questions benchmark, is exactly the kind of biology-adjacent surface where Anthropic’s own safeguards would hand off to Opus 5 in ordinary production use. OpenAI’s footnote 11 fallback, in other words, approximates what Fable 5.1 running its actual production safeguards would have produced. Footnote 17’s Mythos swap does not have that same defense available for ExploitGym: Anthropic’s own disclosed intervention for cybersecurity tasks names Claude Opus 4.8, not an unsafeguarded Mythos run, so the number OpenAI printed under “Claude Fable 5.1” is not what Fable 5.1 in production would have scored either — it is what the same weights score with no safeguard standing in front of them at all, a materially larger capability than production Fable 5.1 ever exposes. Two footnotes, same superscript number, two different distances from what a reader would reasonably assume “Claude Fable 5.1” means on a benchmark table.

It is also worth being precise about which cells footnote 11 actually changes. All three Claude columns on the HealthBench Professional row — Fable 5.1, Fable 5, and Opus 5 — carry the “11” superscript, but the footnote’s two sentences do not apply evenly across them. The first sentence, on GPT-5.4 grading and length-adjusted, unclipped scoring, is a blanket methodology note that applies to all three: OpenAI re-ran and re-graded every Claude number on that row itself, rather than reusing whatever grading convention Anthropic used internally. Only the second sentence, the Opus-5-fallback substitution, is specific to the Fable 5.1 column. Fable 5 and Opus 5’s 60.9% and 56.4% are OpenAI’s own re-grading of those models’ actual answers; Fable 5.1’s 58.1% is OpenAI’s own re-grading of a mix of Fable 5.1’s actual answers and Opus 5’s stand-ins wherever Fable 5.1 declined to answer at all. One superscript, shared by three cells, disclosing two different things depending on which cell a reader happens to be looking at.

It is worth noting what OpenAI’s table does elsewhere, on a coding benchmark neither footnote touches. Terminal-Bench 4.0, in the earlier “Coding” section, prints Claude Fable 5.1 at 55.8% with no footnote mark at all [1] — the same 55.8% Anthropic’s own launch page reports as Fable 5.1’s genuine score on the same benchmark [2]. OpenAI plainly can print an unflagged, undisputed Fable 5.1 number when it has one. The two rows that got a footnote are the two rows where, on the evidence Anthropic has itself published, a clean Fable 5.1 number may not have existed to print. ExploitGym asks a model to develop working exploits from known vulnerabilities — precisely the capability Anthropic says Fable 5.1’s safeguards are built to decline, redirecting cybersecurity tasks that trigger an intervention to Claude Opus 4.8 rather than completing them as Fable 5.1 [2]. A benchmark built entirely around that declined capability would have little for Fable 5.1 itself to score on; Mythos, with the safeguard removed, had something to report instead. That is a plausible account of why OpenAI reached for Mythos’s number specifically there — it is this outlet’s own inference from Anthropic’s own disclosed safeguard behavior, not a claim OpenAI itself makes anywhere on the page, and it is offered as exactly that: a reasoned guess, not a fact traced to a primary source. Why ScreenSpot-Pro needed the same substitution is not something either company’s published material explains, and this piece does not guess at it further.

The Reach: How Far a Reader Travels to Find Out

None of the foregoing is a claim about concealment, and it should not be read as one. Footnote 17 exists. Footnote 11 exists. Both say, in plain sentences, exactly what happened to the numbers they are attached to. The question this piece asks is narrower and more checkable than “did OpenAI disclose this”: it is “how far does a reader have to travel, on the page as published, between the row and the sentence that explains it.” Call that distance a cell’s reach — a plain count, not a judgment, of how many full benchmark-category sections stand between a row and the footnotes block that follows all of them, plus the footnote’s own position in that block’s numbered sequence. It is a structural fact about the page’s layout, verifiable by anyone who opens it and counts, and it says nothing on its own about what any particular reader actually noticed, remembered, or was misled by — no test of that kind was run for this article, and none is claimed.

Run the count. ScreenSpot-Pro’s row sits in the first of OpenAI’s nine table-sections, Computer Use. Its footnote, 17, sits after the ninth. A reader moving from the row to the explanation crosses eight full sections — Professional, Coding, Academic, Science and Health, Cybersecurity, Alignment, Long Context, and Abstract reasoning — before the footnotes block even begins, and then has to find the seventeenth entry in a list of seventeen [1]. ExploitGym’s row sits in the sixth section, Cybersecurity; its reach to the same footnote 17 is shorter but still real — three full sections crossed (Alignment, Long Context, Abstract reasoning) before the footnotes block, then the same walk to the last entry in the list. HealthBench Professional’s row sits in the fifth section, Science and Health, its reach to footnote 11 crossing four sections (Cybersecurity, Alignment, Long Context, Abstract reasoning) before landing on the eleventh of seventeen footnotes.

A desk-mounted magnifying loupe on an articulating arm swinging into position over a printed table row, a small superscript numeral beside a percentage just entering sharp focus at the lens edge

Figure 1. Under magnification the superscript reads plainly enough; what it actually points to sits nowhere near this row [@openai-gpt-6-astra-launch]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Set the three OpenAI cases beside Anthropic’s Terminal-Bench 4.0 disclosure and the reach collapses to a single row of zeroes:

However many actual screens of scrolling that adds up to on a given monitor, at a given zoom level, is not something this piece measures — that would vary by device in a way the section count does not. What does not vary is the count itself: whichever of OpenAI’s own two cross-vendor substitutions a reader is checking, the sentence that explains it sits behind most or all of the rest of the comparison, collected with every other footnote on the page, rather than anywhere near the number it corrects.

A steel tape measure hooked to one printed row on a long fan-folded printout stretched across a desk, its blade extended most of the way toward a footnotes block at the far end but not yet reaching it

Figure 2. Nine benchmark-category tables sit between the row and the sentence that explains it; the seventeenth of seventeen footnotes is the one that does [@openai-gpt-6-astra-launch]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Anthropic Ran the Same Substitution Inside the Same Cell

The reach is not an abstract standard invented for this piece; OpenAI’s own table sits next to a same-week counterexample from the company on the other side of the row. This outlet’s own Reading the Claude Fable 5.1 Scorecard already reported, on Anthropic’s launch day, that “Anthropic reports Terminal-Bench 4.0 is a separate lead for Claude Mythos 5.1 specifically, at 60.9%, above Fable 5.1’s 55.8%” [3] — the identical kind of cross-safeguard substitution OpenAI’s footnote 17 discloses, on the identical benchmark, involving the identical two model names. That finding is not new here and this piece claims no credit for it; it is cited as the settled baseline against which OpenAI’s own convention can be measured.

What is worth looking at directly is the shape of Anthropic’s own disclosure, because it is not merely earlier in the document — it is not a footnote at all. Anthropic’s benchmark table lists Terminal-Bench 4.0 with a single cell under “Fable 5.1,” and that cell reads, in full: “55.8%60.9% (Mythos 5.1)” [2]. Both numbers are printed inside the same cell, on the same row, with the substitute model’s name spelled out in parentheses immediately next to the second figure. There is no footnote to open, no numbered marker to trace, no separate section to scroll past. The qualification and the number it qualifies occupy the same three square centimetres of the page.

Two printed proof sheets side by side under a lightbox, the left sheet with a thread running from a pin at one row out toward a distant pinned margin note, the right sheet with a red pencil circle drawn directly around a parenthetical note inside the same cell as its number

Figure 3. Anthropic's own table put its equivalent substitution inside the cell that needed it; OpenAI's sends the reader looking for a pin at the far edge of the page [@anthropic-fable-mythos-launch] [@adp-fable-scorecard]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Anthropic’s own page is not uniformly different from OpenAI’s here, and the difference is worth locating precisely rather than overstating. Two other rows carry their own bracketed markers — Terminal-Bench-Science 0.1 with a “[1]” and OSWorld 2.0 with a “[2]” — and both are resolved the same way OpenAI resolves its own: in a numbered “Footnotes” section collected near the end of the page, not in prose beside the table. Footnote 1 gives the score’s own uncertainty and a reproduction check: “The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise” [2]. Footnote 2 turns out to answer a different question than a reader tracking safeguard substitutions might expect: it discloses that OSWorld 2.0’s scores run on “the benchmark authors’ August 2026 task release,” not directly comparable to earlier published results, which is why the row carries no competitor figure [2] — a data-versioning caveat, not a safeguard one. The safeguard interventions that zeroed out Fable 5.1 and Fable 5 on OSWorld 2.0, and Fable 5 on AutomationBench, are disclosed separately, in a caption printed directly beneath the whole table rather than tied to either bracket [2]. So two of Anthropic’s three marked rows lean on the identical end-of-document footnote convention this piece is measuring OpenAI against, not some uniformly closer alternative. What is not equivocal is the row this piece is actually about. Terminal-Bench 4.0’s Mythos substitution carries no bracket and no footnote of any kind — alone among the substitutions examined on either company’s page, it needed no cross-reference at all, because the qualifying figure was set directly beside the number it qualifies. That is a narrower claim than “Anthropic doesn’t use footnotes,” and it is the only claim this piece makes about Anthropic’s own page.

I read the in-cell choice as the harder one to make, not the easier one. Printing “60.9% (Mythos 5.1)” directly beside Fable 5.1’s own 55.8% hands a reader, in the same glance, the fact that Fable 5.1’s real number is the smaller of the two — there is nowhere for that comparison to hide once both figures share a cell. A footnote at the bottom of a different page does not force that same glance; a reader has to go looking for it on purpose. Neither company is under any obligation to make its own headline numbers look worse at a glance, so the fact that one of them did, on this specific row, is worth noting on its own terms rather than treated as a default any publisher would reach for.

A Footnote Is Not, By Itself, a Failure

OpenAI has a straightforward answer available, and it deserves to be stated plainly rather than waved off. A numbered footnote, collected with its neighbors at the end of a document, is a completely standard convention — it is how academic papers annotate claims, how financial filings annotate figures, and how OpenAI’s own system cards and Preparedness Framework posts routinely handle methodology notes that would otherwise clutter a claim sentence [9]. Nothing about using a footnote, on its own, is evidence that OpenAI intended a reader to miss anything. Most of OpenAI’s seventeen footnotes on this page carry real, substantive methodology detail that would genuinely clutter a table if printed inline. Footnote 1 discloses that ARC-AGI-3’s 99.9% figure was run with “our responses API harness, which changes two settings to better match real-world performance,” and adds, unprompted, that “the changes do not specifically target ARC-AGI-3” — a caveat that complicates OpenAI’s own headline number and could easily have been left out entirely [1] [7]. Footnote 14 discloses that GPT-5.6 Sol’s 5.5% score on an internal exploit benchmark “is an artifact of the 300-turn limit in the benchmark, which is not a limit that real customers using max would have,” and that the same model scores 11.5% once that artificial ceiling is lifted — a note that argues against OpenAI’s own competitive interest by flagging a number that makes a rival product look artificially worse than a fairer setup would. Footnotes doing that kind of work are not filler, and collecting seventeen of them in one place, rather than interrupting nine different tables with inline asides, is a defensible editorial choice on a page already carrying dozens of numbers.

The claim this piece is making is narrower than “OpenAI did something wrong,” and it should not be read as broader than that. It is that inline-in-the-cell and footnoted-at-the-document’s-end are not the same disclosure, even when both are honest and both are true, because they hand a reader different amounts of work to do before the honest, true thing is visible at all. Whether that difference is a meaningful failure or an unremarkable editorial preference is, itself, a judgment call — one shaped by this outlet’s own house standard of checking every number against the page that reports it, not a neutral fact either company’s launch page settles on its own. A reader, or a competing lab, who considers a footnote fully sufficient regardless of where it sits has a coherent position; this piece’s position is that the sufficiency of a disclosure and its distance from the number it qualifies are two different things worth measuring separately, and that OpenAI’s own page, this week, is a clean case where the two came apart by nine table-sections and a full page of footnotes for one of its two cross-vendor substitutions.

A fanned row of small numbered footnote index cards on a desk in ascending order, one card sitting at a slight raised angle rather than flush with its flat neighbours

Figure 4. A footnote is one card in a long, ordered fan — legitimate, standard, and still the eleventh or seventeenth thing a reader has to open [@openai-gpt-6-astra-launch]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

What a Reader Who Trusts Only the Row Would Conclude

This outlet’s own standing practice — already applied to Anthropic’s launch scorecard on the day it published — is to check every vendor figure against the page that reports it before repeating it as fact. Applied here, to OpenAI’s own table, that practice produces a specific, bounded finding: three of the five cross-vendor Claude figures examined in this piece are not what their column header says, and two of those three sit behind the single longest reach on the entire page to find out why. That is not a claim that Astra’s real advantage over the genuine Fable 5.1 evaporates once the substitutions are corrected — on Terminal-Bench 4.0, the one benchmark where OpenAI used Fable 5.1’s own real, unflagged 55.8% against Astra’s 57.9%, Astra’s lead stands exactly as printed, undisputed by anything in this piece. The finding is smaller and more specific than a verdict on who is ahead: a reader who stops at the row, the way a row invites a reader to stop, walks away from ScreenSpot-Pro and ExploitGym having compared GPT-6 Astra to a Claude configuration that was never actually running under the name printed above it — not because OpenAI hid that fact, but because OpenAI put the correction as far from the number as this one page’s own footnote convention allows it to go, while the company standing next to it in the same benchmark cycle chose, on its own table, to put the identical kind of correction in the one place a reader is already looking.

The practical upshot, for this outlet and for anyone else treating either company’s launch-day table as a source rather than a headline, is narrow and mechanical rather than sweeping: a cross-vendor cell is not verified by reading the cell. Every one of the three OpenAI figures examined here required opening a document section most readers never reach before it was clear what model had actually produced the number. That is a cost worth paying once, in the reporting, so that it does not have to be paid by every reader who encounters the table afterward — which is the same reason this piece exists as a short, separate reading of one page rather than a footnote of its own, buried at the bottom of a longer piece about something else.

Sources

  1. OpenAI. GPT-6 Astra: A new generation of intelligence. OpenAI (2026).
  2. Anthropic. Introducing Claude Fable 5.1 and Claude Mythos 5.1. Anthropic (2026).
  3. Brecht Corbeel. Reading the Claude Fable 5.1 Scorecard. Absolute Digital Publishers (2026).
  4. Russell Brandom. Anthropic's new Fable release is cheaper, less restrictive. TechCrunch (2026).
  5. Ina Fried. "Welcome to the AGI era," OpenAI says as GPT-6 Astra debuts. Axios (2026).
  6. Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua. ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use. arXiv (2025).
  7. Greg Kamradt. OpenAI's GPT-6 Astra on ARC-AGI-3. ARC Prize (2026).
  8. Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, Rahul K. Arora, Foivos Tsimpourlas, Preston Bowman, Michael Sharman, Chi Tong, Kavin Karthik, Arnav Dugar, Akshay Jagadeesh, Khaled Saab, Johannes Heidecke, Ashley Alexander, Nate Gross, and Karan Singhal. HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats. arXiv (2026).
  9. OpenAI. Preparedness Framework. OpenAI (2025).

Originally published at https://absolutedigitalpublishers.com/articles/openai-buries-its-own-asterisk-two-screens-below-the-table.