OpenAI Says Astra Costs 63% Less Than a Model Priced Exactly Like It

Here is OpenAI’s sentence, in full, exactly as it appears next to the Terminal-Bench 4.0 row of its own GPT-6 Astra benchmark table: “Terminal-Bench 4.0 tests agents on complex terminal-based tasks, including software engineering, system configuration, and data analysis. GPT-6 Astra reaches a new high at 57.9%, compared with 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task, respectively” [1]. Astra beats Fable 5.1 on the benchmark by 2.1 points and, OpenAI says, does it for 63 cents on the dollar. Two paragraphs later in the same launch document, a second benchmark carries the identical shape of claim — Terminal-Bench Science 0.1, Astra at 64.6% against Fable 5.1’s 52.6%, “at approximately 31% lower estimated API cost” — and a third — BenchCAD, an industrial CAD-code-generation benchmark built from more than 17,900 execution-verified CadQuery programs spanning 106 real part families, from bevel gears to twist drills [6] — puts the gap at 86% [1]. Three benchmarks, three cost-per-task percentages, all pointing the same direction, all OpenAI’s own arithmetic on a rival’s model.

The claim would be unremarkable if Astra were simply the cheaper model per token — cheaper models finishing tasks more cheaply is not news. It isn’t. Astra launched September 3, 2026; Fable 5.1 launched two days earlier, September 1 [1] [3]. Read OpenAI’s own pricing page and Anthropic’s own pricing page side by side and the two models released in that same week charge the identical sticker price: $10 per million input tokens, $50 per million output tokens, on both [2] [3]. Whatever produces a 63% cost gap on Terminal-Bench 4.0 is not “Astra is priced lower.” It has to be built entirely from how the two models actually use those identically-priced tokens — and, as it turns out, from a mechanism where Astra’s own published rate is the more expensive one, not the cheaper one. That mechanism is the subject of this article, and it complicates OpenAI’s claim in a direction that runs opposite to what a reader skimming the launch page would assume.

Nothing here disputes that Astra could genuinely be cheaper to run on these benchmarks — that is plausible, and OpenAI’s own materials give independent reason to believe token-count efficiency is real for this model. What follows is narrower: a derivation, built entirely from numbers both companies chose to publish themselves, of how much work “63% lower cost” actually has to do before it can be taken at face value, and what specific number neither company has yet supplied that would let a buyer check it directly.

ADVERTISEMENT

Input and Output Match to the Dollar; Only the Cache Line Splits

Set the two companies’ own numbers next to each other and there is exactly one line item that differs.

Line item GPT-6 Astra Claude Fable 5.1 Ratio
Base input $10.00 / MTok $10.00 / MTok 1.0x
Output $50.00 / MTok $50.00 / MTok 1.0x
Cache write $12.50 / MTok (1.25x input) $12.50 / MTok, 5-min cache (1.25x input) 1.0x
Cache read (hit) $1.00 / MTok (10% of input) $0.25 / MTok (2.5% of input) 4.0x

Astra’s numbers come directly from OpenAI’s own model page: “Input $10.00, Cached input $1.00, Cache writes $12.50, Output $50.00” per million tokens, with cache writes “billed at 1.25x the uncached input token rate” [2]. Fable 5.1’s come from Anthropic’s own two primary pages. The launch announcement states the base rate directly — “Fable 5.1’s pricing is otherwise the same as Fable 5’s: $10 per million input tokens and $50 per million output tokens” — and separately, that “cache reads now cost 75% less, or $0.25 per million tokens” [3]. Anthropic’s pricing documentation gives the same figure in multiplier form: a cache read on Fable 5.1 and Mythos 5.1 runs “0.025x the base input price,” against “0.1x base input price” for every other current Claude model [4].

Line them up and three of four numbers tie exactly. The fourth does not: a cache hit on Astra costs four times what a cache hit on Fable 5.1 costs, in absolute terms ($1.00 versus $0.25) and in relative terms (10% of base input versus 2.5%). Nothing in either company’s own pricing page disputes this; it is simply the number each company chose to publish, sitting one column apart on two separate pricing tables that most readers of a benchmark-comparison table never open side by side.

Cache writes tell a smaller, consistent story. Astra bills cache writes “at 1.25x the uncached input token rate” — $12.50 per million tokens — with a single rate and no stated duration split [2]. Fable 5.1 offers the identical $12.50, 1.25x rate for a five-minute cache, plus a second option Astra’s page does not list at all: a one-hour cache write at $20, or 2x base input, for sessions that need a cached prefix to survive longer between turns [4]. Where Astra’s cache-write pricing matches the Claude family’s own short-duration tier exactly, Fable 5.1 offers a second, longer-duration option Astra’s published pricing has no equivalent for — a genuine structural difference, though a minor one next to what the read-side rate is doing to the bill on any workload that revisits a cached prefix more than once.

Two small mechanical timers set beside the two receipts, one wound for five minutes with its dial nearly run down, the other wound for one hour and barely moved off its start position
Figure 4. Astra's cache-write pricing offers a single 1.25x-rate option with no stated duration; Fable 5.1 offers the identical short rate plus a second, longer option Astra's page has no equivalent for — a one-hour cache write at 2x base input [@openai-astra-pricing] [@anthropic-pricing-docs].Image prompt and art direction by Brecht Corbeel; generation pending.

Astra’s Cache Rate Is the Claude Family’s Own Default — Fable 5.1 Is the Outlier

The instinct here might be to read Astra’s 10% cache rate as unusually expensive, engineered specifically to disadvantage it against Fable 5.1. Anthropic’s own documentation says the opposite: 10% of base input is what every Claude model charges for a cache hit except Fable 5.1 and Mythos 5.1 specifically. “Cache read (hit): 0.1x base input price (0.025x on Claude Fable 5.1 and Claude Mythos 5.1)… A cache hit costs 10% of the standard input price, which means caching pays off after one cache read for the 5-minute duration… On Claude Fable 5.1 and Claude Mythos 5.1, a cache hit costs 2.5% of the standard input price” [4]. Opus 5, Opus 4.8, Sonnet 5, Haiku 4.5 — every other current model in Anthropic’s own lineup — bills cache reads at the same 10% multiplier Astra uses. Fable 5.1 and Mythos 5.1 are the exception Anthropic carved out, on purpose, as the entire mechanism behind the cost claim in its own launch materials: “we’re reducing our pricing on cache reads… For highly agentic work, the savings will often be much larger — up to approximately 45%” [3]. This outlet’s own practitioner’s guide to Fable 5.1 already walks through what that discount does to a buyer’s bill — the 25%-typical, up-to-45%-agentic range, and the arithmetic behind it — and takes it as settled background here rather than re-deriving it: Building with Claude Fable 5.1: A Practitioner’s Guide.

ADVERTISEMENT

What that reframes is not whether the 4x gap is real — it is, on both companies’ own numbers — but which company should be described as departing from the industry’s own default. Astra’s cache economics are the unremarkable, standard-multiplier case. Fable 5.1’s are the one Anthropic chose to cut below its own family’s baseline, specifically for this model pair, specifically now. A comparison that treats Astra’s rate as the anomaly has the direction backward.

Anthropic states directly why it made that cut, in the same announcement that introduces Fable 5.1: “Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards” [3]. Read against the pricing table, “feedback… on price” resolves to one specific, targeted change — the cache-read multiplier, and nothing else in the rate card — rather than a general price cut across the model. Anthropic did not lower Fable 5.1’s input or output rate, and did not extend the discount to any other model in its current lineup. It cut exactly the one line item that determines how cheap a long, context-heavy agentic session gets, on exactly the two models — Fable 5.1 and Mythos 5.1 — being positioned this week as its answer to sustained, tool-using work. That is a deliberate competitive move on Anthropic’s part, not an incidental side effect of a broader repricing, and it is the move Astra’s identical-on-paper sticker price has to be read against.

A steel ruler laid diagonally across the cached-input line of two receipts at once, aligning one small figure against a visibly larger one printed directly beside it
Figure 1. The one line item that does not match: a cache read prices at 10% of base input on one receipt and 2.5% on the other — a four-to-one gap hiding under an identical sticker price [@anthropic-pricing-docs].Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Run the Blend and Astra Never Wins on Cached Tokens Alone

None of that settles what a 4x cache-rate gap does to a real task’s bill, because a real task mixes cached and fresh input tokens in proportions neither company discloses for these specific benchmarks. What can be stated without guessing the mix is the direction and shape of the effect, holding one deliberately conservative assumption fixed: that both models process a comparable total volume of input tokens on the same task, varying only in what share of those tokens are cache hits versus fresh reads. That assumption almost certainly understates Astra’s real handicap rather than overstating it, for reasons the next section takes up — but it is the cleanest way to isolate what the cache-rate gap alone contributes, before any claim about raw token-count efficiency enters the picture.

Call the cache-hit share ss — the fraction of input tokens that are cache reads rather than fresh reads. The blended input-token cost per model, in dollars per million tokens, is (1s)×10+s×r(1-s) \times 10 + s \times r, where rr is each model’s cache-read rate: $1.00 for Astra, $0.25 for Fable 5.1. The ratio of Astra’s blended input cost to Fable 5.1’s, at a given cache-hit share, is:

(1s)(10)+s(1.00)(1s)(10)+s(0.25)=109s109.75s \frac{(1-s)(10) + s(1.00)}{(1-s)(10) + s(0.25)} = \frac{10 - 9s}{10 - 9.75s}

Run that ratio across a stated bracket rather than picking one number and presenting it as fact:

Cache-hit share (ss) Astra blended input rate Fable 5.1 blended input rate Astra ÷ Fable 5.1
0% $10.00 $10.00 1.00x (parity)
30% $7.30 $7.075 1.03x
60% $4.60 $4.15 1.11x
90% $1.90 $1.225 1.55x
100% (all cache) $1.00 $0.25 4.00x

At zero cache reuse, the two models cost the same on input tokens — unsurprising, since the base rate ties. From there, every step toward a higher cache-hit share moves the ratio in exactly one direction: against Astra. There is no cache-hit share above 0% at which Astra’s own published rate produces a cheaper input-token bill than Fable 5.1’s, under the equal-token-count assumption. This is not a sensitivity finding that happens to favor one reading over another; it is the direct, unavoidable consequence of a 4x rate gap combined with identical base pricing, and it holds regardless of which specific share the real benchmark run actually used, because the function is monotonic across the entire range.

ADVERTISEMENT

Make one point on that table concrete with an illustrative turn — chosen only to show the mechanism at a realistic scale, not offered as either company’s actual Terminal-Bench 4.0 token count, which neither has disclosed. An agent re-sends a 100,000-token cached prefix (accumulated system instructions, tool schemas, command history) alongside 5,000 fresh input tokens, and produces 3,000 output tokens on that turn — a 95%-cache-hit-share turn, toward the high end of the bracket above.

Line item Astra cost Fable 5.1 cost
100,000 cached tokens 100,000 × $1.00 / 1,000,000 = $0.100 100,000 × $0.25 / 1,000,000 = $0.025
5,000 fresh input tokens 5,000 × $10 / 1,000,000 = $0.050 5,000 × $10 / 1,000,000 = $0.050
3,000 output tokens 3,000 × $50 / 1,000,000 = $0.150 3,000 × $50 / 1,000,000 = $0.150
Turn total $0.300 $0.225

Holding every token count identical between the two models, this single turn costs 33% more to run on Astra than on Fable 5.1 — purely because of the cache-read rate, before either model’s actual output length, turn count, or task-completion efficiency enters the calculation at all. Scale that turn across the many-turn trajectories a Terminal-Bench 4.0 run typically involves and the gap compounds with every additional cache read, not just once.

A narrow notepad beside the two receipts with three trial totals written in a column and a pencil bracket drawn partway around them, the bracket's closing stroke not yet complete
Figure 2. Run the same two rate sheets at a few different assumed cache-hit shares rather than pretending one blended number is disclosed, because neither company states the share its own claim depends on.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

A 63% Total Win Has to Be Paying Off a Debt First

Put the cache-rate table beside OpenAI’s own claim and the two make an odd pair. On Terminal-Bench 4.0, OpenAI reports Astra finishing at “approximately… 63% lower estimated API cost” than Fable 5.1 [1] — a total-cost figure that necessarily blends input, cache, and output tokens together, not an input-only number. Output tokens carry no rate asymmetry at all: both models charge $50 per million, so whatever output-token savings Astra achieves land at full value, undiluted by any handicap. Fresh input tokens are likewise a level field — same $10 rate on both. The only place a handicap exists is exactly the cache-read line just derived, and on that line, at any plausible cache-hit share, Astra is paying more per token than Fable 5.1, not less.

That means the 63% figure is not simply “Astra used fewer total tokens, priced at the same rate.” It is “Astra used enough fewer total tokens — concentrated, given where the rate asymmetry bites, in output and fresh-input volume — to first absorb a real cache-side cost penalty and then still come out 63% ahead overall.” The same shape holds for the other two benchmarks carrying a cost claim in OpenAI’s launch materials:

Benchmark Astra score Fable 5.1 score OpenAI’s cost claim
Terminal-Bench 4.0 57.9% 55.8% “approximately… 63% lower estimated API cost per task” [1]
Terminal-Bench Science 0.1 64.6% 52.6% “approximately 31% lower estimated API cost” [1]
BenchCAD (with tools) 95.9% 84.3% (reported) “approximately… 86% lower than Fable 5.1” [1]

All three carry the identical structural problem this article derives for Terminal-Bench 4.0: an input-side rate that runs against Astra at any nonzero cache-hit share, an output rate that ties exactly, and a headline percentage that has to be net of both. Whatever raw-token efficiency advantage sits underneath these three headline percentages is larger than the percentages state on their own, because part of what that advantage is doing — silently, with no line in OpenAI’s table acknowledging it — is closing a gap the cache-rate asymmetry opens up first.

There is independent, if indirect, support elsewhere in OpenAI’s own launch materials for exactly this kind of token-count lever being real and large. On a different benchmark entirely — Agents’ Last Exam, a thousand-plus-task evaluation of long-horizon, economically valuable work built by a consortium of more than 250 industry experts and academic researchers, where even the best current agents clear less than 1% of its hardest tier in full [7] — measured against Claude Opus 5 rather than Fable 5.1 — OpenAI states plainly that “Astra also uses approximately 65% fewer output tokens than Opus 5” at its highest-scoring settings [1]. That figure says nothing directly about the Fable 5.1 comparison; Opus 5 is a different Claude model, priced differently ($5 input / $25 output, not $10/$50) [4], and Agents’ Last Exam is not Terminal-Bench 4.0. But it establishes that “generates a similar or better answer using dramatically fewer output tokens” is a real, OpenAI-documented pattern for this model family, not a hypothetical this article is inventing to rescue its own arithmetic. If a comparable output-token gap holds on Terminal-Bench 4.0 against Fable 5.1 specifically — untested here, and OpenAI’s launch page does not say either way — it would be large enough on its own to explain most of a 63% total-cost gap even after conceding the cache-side handicap in full.

Terminal Agents Are Exactly the Workload Where This Debt Compounds

A fair objection to all of the above: maybe cache-read pricing barely matters for a task like Terminal-Bench 4.0, where fresh input, tool-call overhead, and output tokens could dominate the bill and cache hits could be a rounding error. If that were true, the entire derivation above would be a real but practically irrelevant asymmetry — accurate on paper, meaningless in a real invoice.

It is not the shape of workload where that objection holds most comfortably. Terminal-Bench 4.0 evaluates agents across “complex terminal-based tasks, including software engineering, system configuration, and data analysis” [1] — a benchmark built and maintained independently of either lab, hosted jointly by Stanford, the Harbor framework project, and the Laude Institute, whose fourth version recalibrated task timeouts and retired eight tasks that earlier models had already saturated [5] — multi-turn, tool-heavy trajectories where an agent repeatedly re-sends a large, mostly static prefix (system instructions, tool schemas, accumulated command history) on every turn of a long session. That is precisely the shape of workload prompt caching exists to serve, and precisely the shape where cache-hit share climbs highest as a proportion of total input tokens. Anthropic’s own guidance on Fable 5.1’s pricing makes the same point from the seller’s side: the up-to-45% savings figure applies specifically to “highly agentic work,” described as workloads where “cache reads make up most of the cost,” in contrast to a “typical” workload where the 25% figure applies instead [3]. If Terminal-Bench 4.0’s actual trajectories run anywhere near that “highly agentic” end of the spectrum — plausible given the benchmark’s own description, though neither company states the specific cache-hit share their own cost estimate used — the cache-hit share sits well above the low end of the bracket in the table above, not near zero. That pushes the ratio Astra has to overcome toward the higher end of the range shown — 1.1x to 1.55x or beyond on the cache-affected portion of the bill alone — which raises rather than lowers how large the underlying token-efficiency advantage has to be for the 63% headline to hold.

How large a share of a real agentic bill cache reads can reach is not a theoretical question on Fable 5.1’s own numbers. This outlet’s practitioner guide to the model works a twenty-turn agentic session where cached tokens make up the large majority of each turn’s input, and finds the cache-rate cut alone worth a 50% reduction against Fable 5’s own prior pricing on that specific, cache-heavy shape of session — above even the “up to 45%” ceiling Anthropic’s own announcement quotes for agentic work in general [3]. That example was built to isolate the cache-read line specifically, not to represent a blended real-world average, but it demonstrates the range is real at the high end, not a theoretical extreme nobody’s workload ever reaches. A Terminal-Bench 4.0 trajectory sitting anywhere near that shape is exactly where this article’s derived cache-rate disadvantage for Astra stops being a rounding error and starts being a meaningful fraction of the bill.

Neither Company Publishes the One Number That Would Settle It

The number that would resolve this cleanly — what cache-hit share OpenAI’s own cost estimate assumed for Astra and for Fable 5.1 on Terminal-Bench 4.0, or equivalently, the actual token counts behind the 63% figure — appears in neither company’s published materials. OpenAI’s launch page carries seventeen numbered footnotes attached to its benchmark table, covering everything from harness configuration on ARC-AGI-3 to grading methodology on HealthBench Professional to which Claude model actually produced a given score [1]; none of them addresses how the “estimated API cost per task” figures were computed, what token volumes they assumed, or what share of input tokens were treated as cache hits. Anthropic, for its part, never makes this specific comparison at all — the 63%, 31%, and 86% figures are OpenAI’s own estimates about a rival’s model, run in OpenAI’s own environment, not a number Anthropic published or verified independently on either side.

The nearest thing to independent verification comes from neither lab. Artificial Analysis runs its own Terminal-Bench 4.0 evaluations, outside either company’s environment, and scores GPT-6 Astra at 59.6% against Claude Fable 5.1’s 55.1% — a tighter gap than OpenAI’s self-reported 57.9%-to-55.8% pair, on different absolute numbers entirely, gathered under a different harness and a different set of runs [8]. That is independent confirmation that Astra wins this specific matchup; it is not independent confirmation of why, or of what cache-hit share either model’s cost figure assumed. Artificial Analysis publishes its own cost-per-task figures for both models on the same evaluation, and that dataset, too, carries no disclosed cache-hit share — the one number this section has been circling stays missing even in the one dataset built by neither vendor.

That silence is worth stating precisely rather than treating as either damning or irrelevant. It does not mean the 63% figure is wrong — nothing here contradicts it, and the Agents’ Last Exam output-token figure gives independent reason to think large token-count gaps are real for this model family. It means a reader cannot currently separate how much of Astra’s apparent cost edge comes from genuinely using fewer tokens to do the same work, and how much would shrink or vanish under a different, equally plausible assumption about how heavily either evaluation run relied on cached context. The honest position is a bounded one: at zero cache reuse, the cache-rate asymmetry contributes nothing and the 63% figure would need no adjustment; at high cache reuse — the more plausible end for this task class — the underlying token-efficiency gap responsible for the 63% headline has to be meaningfully larger than 63% would suggest in isolation, in order to first clear the cache-side handicap this article derives directly from both companies’ own published rates. If either company later discloses the actual cache-hit share behind this specific claim, and that share turns out to produce no meaningful difference from a same-price calculation, this article’s central inference — that the true efficiency gap runs larger than the headline number — would not survive intact, and should be retracted rather than defended.

A close view of a receipt's header row printing a "token count" field label with no figure filled in beside it, a pencil hovering just above without touching the paper
Figure 3. The one number that would settle the question — how many tokens, and what share of them cached, either company's own estimate actually used — appears in neither company's published cost claim.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The Sticker Price Was Never Where the Comparison Lived

Go back to the sentence this article opened on: Astra at 57.9% against Fable 5.1’s 55.8% on Terminal-Bench 4.0, “at approximately… 63% lower estimated API cost per task” [1]. Read next to a flat, identical $10/$50 sticker price, that sentence reads like a clean efficiency win, full stop. Read next to Anthropic’s own pricing table — where Fable 5.1’s cache-read rate sits at a quarter of Astra’s, an exception Anthropic itself carved out below its own model family’s standard multiplier — the same sentence reads differently: not simply “Astra is more efficient,” but “Astra is more efficient by an amount large enough to make a real, quantifiable pricing disadvantage disappear entirely and still post a headline number.” Both readings can be true at once. Only one of them is visible from the sticker price alone, and the sticker price is the only part of this comparison most coverage of the launch actually quotes.