A rate card is not a common ruler

OpenAI, Anthropic, and Google each publish a number denominated the same way: dollars per million tokens, input and output, for every model they sell through an API. Because the unit matches, it is tempting to drop three rate cards into one spreadsheet and read down the input column for the cheapest number. That comparison is not wrong so much as premature. It treats “dollars per million tokens” as though it names one measurable thing, when it actually names five separable choices a vendor has made — how a tier is priced, how a cache discount is structured, how a batch discount is structured, how context length is priced, and what a token even is on that vendor’s own tokenizer — three of which turn out to agree closely across all three companies, and two of which do not agree at all.

This article works through both halves honestly, as verified against each vendor’s own documentation on 12 August 2026, plus independent price-tracking analysis for the details no single vendor’s page states plainly. It is not a ranking. The series this piece belongs to gives OpenAI and Anthropic primary, detailed coverage and Google a smaller, proportionate mention, because that is where the public evidence and the two companies’ own competitive rate-cutting is thickest; nothing below should be read as declaring one vendor’s lineup superior to another’s, because a benchmark score and a price are different kinds of claim, and stacking a table of prices into a winner is exactly the kind of cross-vendor ranking this piece will not build.

Three ladders, priced and dated

As verified on 12 August 2026, OpenAI’s current flagship lineup is the GPT-5.6 family, three models sharing a context window above one million tokens: gpt-5.6-sol at 5 dollars input and 30 output per million tokens, cached input at 0.50; gpt-5.6-terra at 2 and 12, cached at 0.20; gpt-5.6-luna at 0.20 and 1.20, cached at 0.02 [1]. Anthropic’s current lineup, verified the same day, prices Claude Opus 5 at 5 dollars input and 25 output, with a 5-minute cache write at 6.25, a 1-hour write at 10, and a cache read at 0.50; Claude Sonnet 5 at 2 and 10, cache read at 0.20; Claude Haiku 4.5 at 1 and 5, cache read at 0.10 [3]. Google’s Gemini lineup is priced with the same structure but a smaller footprint in this article, consistent with the series-wide rule above: Gemini 3.1 Pro Preview runs 2 dollars input and 12 output for prompts at or under 200,000 tokens, and Gemini 3.6 Flash runs 1.50 and 7.50 with no stated context tiering at all [5].

ADVERTISEMENT

Two structural facts sit underneath those numbers and matter more than the numbers themselves. First, each vendor’s own ratio of output price to input price is a near-constant across its own lineup rather than a per-model tuning choice: OpenAI holds close to 6 to 1 across Sol, Terra, and Luna; Anthropic holds exactly 5 to 1 across every currently supported tier, Haiku 4.5 through Opus 5 [1, 3]. Second, retirement is itself a pricing signal on Anthropic’s card in a way worth flagging as a caution against reading “newer” as “more expensive”: Sonnet 4.6, an older model by Anthropic’s own numbering, is still billed at 3 dollars input and 15 output, above Sonnet 5’s 2 and 10, because Sonnet 5’s launch pricing — originally scheduled to rise to 3 and 15 on 1 September 2026 — was confirmed as permanent rather than temporary, a scheduled increase Anthropic states will not occur [3]. A rate card read on one day and trusted for a quarter is already a stale instrument; this is not a hypothetical caution, it is what happened to this exact tier inside the three weeks before this sentence was checked.

What actually converges: cache reads, batch jobs, and priority processing

A paired witness-dial instrument with two dial faces, one small pointer caught mid-settle a hair short of one tenth of its primary needle's reading, the other already steady at exactly half its own primary needle
Figure 1. A cache read costs roughly a tenth of a fresh token and a batch job roughly half the standard rate on every current rate card checked here — a structural ratio repeated across independently operated infrastructure, not a coincidence unique to one vendor.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Set the base rates aside and look instead at the multipliers vendors apply on top of them, and a genuine, checkable convergence appears. Anthropic’s documentation states a cache read costs exactly 0.1 times the base input price, uniformly across every current model [3, 4]. OpenAI’s cached-input price divided by its standard input price lands at exactly one-tenth for all three GPT-5.6 tiers: 0.50 against 5.00 for Sol, 0.20 against 2.00 for Terra, 0.02 against 0.20 for Luna [1]. Google’s cached-token price for Gemini 3.1 Pro Preview is 0.20 against a 2.00 standard rate, and for Gemini 3.6 Flash is 0.15 against 1.50 — one-tenth in both cases [5]. Three companies, none coordinating with the others, publishing the identical ratio on every current tier checked for this article.

The batch-processing discount converges just as tightly, at one-half rather than one-tenth. Anthropic’s Batch API processes requests asynchronously at exactly half the standard rate on both input and output tokens, across the entire current lineup [3]. OpenAI’s own documentation states its Batch API carries “a 50% cost discount compared to synchronous APIs,” with each batch completing within a 24-hour window “and often more quickly” [2]. Google’s Batch API documentation states the same figure for Gemini models, with a comparable 24-hour target turnaround [6]. A third ratio echoes the pattern for the two vendors that publish one: OpenAI’s fast-processing tier and Anthropic’s fast mode (a research preview available on Opus 5 and Opus 4.8) both price priority throughput at exactly double the standard rate, and both explicitly forbid combining that tier with the batch discount [1, 3]. Google’s public pricing does not currently publish an equivalent named priority multiplier, which is itself the kind of gap the series-wide rule asks this article to note rather than paper over with an assumed number.

None of this proves a shared serving architecture or a shared internal cost model; no vendor publishes the cost data that would confirm that, and this article is analysis of a repeated pattern, not a claim about why it repeats. What the convergence does establish, carefully, is that the shape of the tariff — a cache read near a tenth, a batch job near a half, priority near double — has settled onto common ground across independently operated infrastructure, even while the absolute level of the underlying rate differs by more than an order of magnitude across tiers within a single vendor’s own lineup.

Prices are dated claims, not constants

A rate-tag embossing press caught mid-strike over a small brass plate with its raised marks only half-formed, an earlier finished plate already fitted into a console visible behind it
Figure 2. A published price is a dated claim, not a constant; the rate one vendor cancelled and the rate another cut both moved inside the same three-week window this piece was checked against.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The Sonnet 5 example above is not an isolated case; it is a symptom of how often these numbers move. As verified 12 August 2026, independent tracking of OpenAI’s own change log records a price cut to the GPT-5.6 family on 30 July 2026: Terra’s rate fell by 20 percent and Luna’s by 80 percent, while Sol held at its original launch price, changes OpenAI itself attributed to rewritten production GPU kernels and a redesigned speculative-decoding draft model rather than to a change in the underlying weights [9]. That is a vendor explanation reported through an independent tracker, not a figure this article verified against OpenAI’s own change-log page directly, and it should be read with that provenance attached. Anthropic’s cancellation of its own scheduled Sonnet 5 increase, by contrast, is stated directly on Anthropic’s pricing page as verified above [3]. Google’s preview-tier models on the Gemini card carry their own instability, flagged directly on the pricing page: “Preview models may change before becoming stable and have more restrictive rate limits” [5].

ADVERTISEMENT

Put together, these are three different mechanisms for the same underlying fact — a serving-cost improvement passed through as a price cut, a scheduled increase cancelled before it took effect, and a preview tier whose price is provisional by the vendor’s own description — and every one of them happened inside a window measured in weeks rather than years. A comparison built from cached numbers, a training corpus, or a spreadsheet copied six months ago is not comparing today’s vendors; it is comparing whichever prices happened to be current on the day the copy was made. Every dollar figure in this article carries the verification date attached for exactly this reason, and any reader carrying one of these numbers forward should treat that date as an expiry warning, not a footnote.

What does not converge: how context length is priced

A bank of three stepped rotary attenuators mounted side by side, one selector knob caught mid-turn between two detents while the knob to its left sits fixed and the knob to its right already sits advanced
Figure 3. One vendor's rate card holds flat across its whole context window while two others step to a higher rate past a threshold — a genuine structural disagreement, not a rounding difference between rate cards.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Long-context pricing is where the three vendors make a genuinely different design choice, not a rounding difference between otherwise-similar rate cards. Anthropic’s Claude 4.6-and-later models bill the full context window at a single flat rate: a 900,000-token request costs the same per token as a 9,000-token one, with caching and batch discounts applying at that flat rate across the whole window [3]. OpenAI and Google instead step to a higher rate past a threshold. Independent tracking of OpenAI’s own documentation quotes the exact mechanism: “Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request,” applied to all three GPT-5.6 tiers — Sol rises from 5 and 30 dollars to 10 and 45, Terra from 2 and 12 to 4 and 18, Luna from 0.20 and 1.20 to 0.40 and 1.80 [9]. Google’s own pricing page states the same step for Gemini 3.1 Pro Preview at a different threshold: 2.00 dollars input and 12.00 output for prompts at or under 200,000 tokens, rising to 4.00 and 18.00 above it — again exactly double the input rate and one-and-a-half times the output rate [5].

That specific pairing, a 2x step on input and a 1.5x step on output, appearing identically at two different companies and two different thresholds, is worth stating precisely because it is a genuine structural echo sitting inside a genuine structural disagreement. Nothing published by either company explains why the ratio matches while the threshold does not; this article treats that match as an observation, not as evidence of a shared cause, since a decode-bound output stream and a prefill-bound input stream plausibly scale differently under long context for reasons specific to each company’s own serving stack, and neither company discloses the internal figures that would settle it. What is fully settled is the disagreement with Anthropic: one vendor’s rate card holds flat where the other two step, which means a workload’s typical context length changes its relative standing across vendors in a way no single “price per million tokens” figure captures. A comparison run only at short context and extrapolated to long-context workloads, or the reverse, is comparing under conditions where at least one of the three vendors is charged differently than the headline number states. Google’s Gemini 3.6 Flash tier, notably, carries no stated tiering at all — a reminder that this structural choice is not even uniform within one vendor’s own current lineup.

What does not converge: what counts as a token

A single paper ribbon threaded through three perforating wheelheads of different tooth pitch, the centre wheelhead's tooth caught mid-stroke while the ribbon trailing behind each head already shows a different spacing of finished perforations
Figure 4. The same passage of text is not the same number of tokens on every vendor's own count, and not even on one vendor's own count across generations; a raw per-token price compares different units before it compares anything else.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

A price quoted per million tokens implicitly assumes a token is a stable unit across the comparison. It is not, in two separate ways.

Within a single vendor, Anthropic states plainly that Claude 4.7 and later models, along with Claude Mythos Preview, “use a newer tokenizer that contributes to their improved performance on a wide range of tasks,” and that “this tokenizer produces approximately 30% more tokens for the same text” than Sonnet 4.6 and earlier models used [3]. A dollar figure per million tokens compared across two Claude generations is already comparing different denominators before a second vendor enters the picture.

Across vendors, the gap is documented but harder to pin to one clean multiplier, because it depends heavily on content type. An independent technical comparison that queried each vendor’s own token-counting endpoint directly — OpenAI’s tiktoken library locally, Anthropic’s count_tokens endpoint on the Messages API, and Google’s equivalent counting endpoint for Gemini, rather than approximating one vendor’s tokenizer with another’s — reported that English prose shows only “single-digit-percent differences” across vendors, that code blocks show OpenAI’s o200k-based tokenizer running “ten to twenty percent” more efficient than the older cl100k encoding on typical files, and that non-Latin scripts are where the real divergence sits: the older cl100k tokenizer runs close to one token per character on Japanese, Chinese, and Korean text, roughly four times worse than its own performance on English, with Google’s larger SentencePiece vocabulary narrowing that specific gap [10]. The methodological point that source makes is as important as any single figure it reports: “never trust a count from a tokenizer you don’t actually call in production” [10], because a local approximation of a vendor’s tokenizer — the only option for Anthropic’s, which is not published — will silently mismeasure the comparison it is meant to support.

ADVERTISEMENT

The practical consequence compounds with everything already established. A raw per-token price comparison assumes two vendors need the same number of tokens to represent the same request. Even holding content type fixed, that assumption already fails from one Claude generation to the next; across three vendors’ own tokenizer designs and three different serving philosophies, treating “tokens” as a portable, vendor-neutral unit is the single most common way an otherwise careful price comparison quietly stops being one.

The multiplier no rate card publishes

A mechanical totalizer with two coupled counting wheels behind a glass face, the outer visible wheel caught mid-click advancing by a single position while the smaller inner wheel behind it is a blurred streak from turning many times faster
Figure 5. A visible answer of a few hundred tokens can sit on top of many thousands billed as output before it; the hidden wheel turns for all of it, and the outer dial only for what the caller ever sees.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Reasoning models add a third source of divergence that sits inside the output-token column rather than beside it, and every vendor examined here handles it the same way in substance if not in the exact controls exposed to the caller: internal deliberation is generated as tokens, billed at the full output rate, and largely invisible to the party paying for it. An independent analysis of this mechanism across all three vendors states it plainly: providers “bill internal ‘thinking’ tokens at the full output rate” while keeping most of that stream out of the visible response, so “a hard question can generate tens of thousands of reasoning tokens behind a two-hundred-token visible answer,” at output rates running 25 to 50 dollars per million tokens on frontier tiers [11]. The same analysis notes the detail that matters for cross-vendor comparison specifically: OpenAI’s reasoning tokens are not returned via the API but are billed as output; Anthropic charges for the full thinking tokens generated rather than only a summary; and Gemini’s output pricing folds thinking tokens into the same billed total behind a visible summary [11]. The magnitude is controllable — OpenAI exposes a reasoning_effort parameter, Anthropic exposes a thinking budget, and both let a caller trade deliberation for cost directly — but the multiplier is not zero by default on any of the three, and none of the three publishes, on its rate card, what the typical ratio of hidden to visible tokens actually is for a given task.

This is the least comparable line item on any of the three cards, structurally, because it is the one where the published price is furthest from the price actually paid. Two models with identical published output rates can differ by a large factor in what a matched task costs, purely as a function of how much unseen deliberation each one’s default reasoning setting elects to generate, and that factor is invisible until a caller runs the task and reads the token count back from the response.

Why “cheapest model” is usually the wrong question

Collecting every distinction above into one expression makes explicit what a bare price comparison assumes without saying so. For vendor ii, let piinp^{\text{in}}_i and pioutp^{\text{out}}_i be the published input and output rates; let κi\kappa_i be a tokenizer expansion factor, the number of tokens vendor ii’s own tokenizer needs to encode one fixed reference passage, normalised so κi=1\kappa_i = 1 for whichever vendor is used as the baseline; let hih_i be the fraction of input served from cache and γi\gamma_i the cache-read multiplier; and let ρi\rho_i be the ratio of total output tokens generated, visible answer plus hidden reasoning, to the visible answer alone. A workload’s realised cost per completed task is then approximately

Ci  =  κi[(1hi)piin+hipiinγi]  +  ρipiout. C_i \;=\; \kappa_i\Big[(1-h_i)\,p^{\text{in}}_i + h_i\, p^{\text{in}}_i\,\gamma_i\Big] \;+\; \rho_i\, p^{\text{out}}_i .

Reading piinp^{\text{in}}_i and pioutp^{\text{out}}_i off a rate card and comparing them directly across vendors is equivalent to assuming κi\kappa_i, γi\gamma_i, and ρi\rho_i are all equal to each other and to 1 across every vendor in the comparison. This article has just shown that one of those three, γi\gamma_i, genuinely does converge close to 0.1 across all three companies’ published cards — the one term a naive comparison happens to get right by accident. The other two do not converge, are not published as clean multipliers by any vendor, and vary by content type and by task in ways that only measurement on the caller’s own workload can pin down. A price comparison that reports piinp^{\text{in}}_i and pioutp^{\text{out}}_i alone has silently set κi=ρi=1\kappa_i = \rho_i = 1 for every vendor, an assumption none of the sourcing above supports.

Even a comparison that correctly measures CiC_i is not yet a capability comparison, which is the deeper reason a “cheapest model” headline usually asks the wrong question. Artificial Analysis, an independent benchmarking organisation, maintains a live Intelligence Index alongside pricing across several hundred models from every major vendor, and explicitly plots “Intelligence Index vs. Cost per Task” as a quadrant chart rather than a single ranked list, with a stated methodology that derives “weighted average cost per Intelligence Index task” from “input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight” [8] — which is to say, a capability-adjusted cost figure of exactly the kind the equation above gestures toward, built by someone who actually ran the tasks rather than read the rate card. Epoch AI’s separate analysis of inference price trends makes the same point from the other direction: the price of reaching a fixed capability bar — GPT-4-level performance on PhD-level science questions, specifically — fell by a documented factor of roughly 40 per year, with the overall range across different capability bars running from roughly 9x to 900x per year depending on which bar is chosen, while the authors caution explicitly that “the fastest price drops in that range have occurred in the past year, so it’s less clear that those will persist” [7]. A decline that large and that uneven overwhelmingly reflects newer, differently designed models reaching a fixed bar more cheaply — substitution between architectures, not a like-for-like discount on one model’s own price.

Put the two halves of this article together and the honest instrument is neither “compare the sticker price” nor “trust the benchmark score” in isolation. It is: measure CiC_i on your own workload, at your own cache-hit rate and your own reasoning-effort setting, against a capability bar you actually need cleared — and only then ask which vendor cleared it for less, on the date you checked.

Predictions, with the observations that would falsify them

These are forecasts, clearly separated from the sourced analysis above. Horizon: 12 August 2028.

One. The cache-read ratio and the batch-discount ratio identified here will still be published at approximately 0.1 and 0.5 respectively by at least two of the three vendors, because both figures plausibly track a real, shared serving-economics constraint (a cache hit skips computation nearly entirely; a batched job trades latency for scheduling efficiency) rather than an arbitrary marketing choice. Disconfirmed if two or more of the three vendors have moved either published ratio by more than 50 percent of its current value by the horizon date.

Two. At least one of OpenAI or Google will publish a stated numeric ratio, typical or worst-case, between hidden reasoning tokens and visible output tokens for at least one model tier, because the gap documented in this article is large enough that sophisticated buyers will begin pricing it into procurement decisions and will ask for the number directly. Disconfirmed if, by the horizon date, no major vendor discloses any such ratio for any model.

Three. Independent capability-adjusted cost tracking, in the style of Artificial Analysis’s Intelligence Index or Epoch AI’s fixed-capability-bar price series, will become a more common citation in procurement discussions than raw per-token rate-card comparisons, because the gap this article documents between the two kinds of comparison is now well enough understood to be a documented methodological error rather than an oversight. Disconfirmed if raw per-token rate-card tables remain the dominant form of public vendor comparison with no capability adjustment by the horizon date.

What to take away

Three companies publish prices in the same unit, and that similarity is doing more work in most casual comparisons than it has earned. Two of the ratios sitting underneath the headline numbers — what a cache read costs against a fresh token, what a batched job costs against an interactive one — turn out to be genuinely, checkably the same across all three vendors’ current rate cards, a real structural fact worth knowing and trusting. Two other things that look like they should be comparable are not: how a vendor prices a request once it crosses a length threshold, and what a token even is for the text being priced, both of which differ by vendor, by content type, and in Anthropic’s case by generation within one vendor’s own lineup. Layered on top of all of it sits a reasoning-token multiplier no rate card states as a number, capable of moving the real cost of a matched task by a large factor that the published price alone cannot reveal.

None of this argues that price comparison across vendors is impossible or not worth doing. It argues that the comparison worth trusting is the one that measures a real workload’s cache-hit rate, tokenizer expansion, and reasoning overhead directly, against a capability bar the buyer actually needs cleared, rather than the one built from three columns copied off three rate cards on a single afternoon. The rate cards are real, dated, and — on the three ratios examined here — more alike than a skeptic would expect. The task is not.