Before the token: paying by the hour

Long before anyone billed a language-model API by the token, the commercial logic of renting someone else’s computer was already settled, and it was settled on a very different unit: the hour. On 25 August 2006, Amazon Web Services opened Amazon EC2 in limited public beta with a single instance type, the m1.small, in a single region, US East, billed at 10 cents per clock hour regardless of what ran on the machine [1]. The meter did not know or care whether the hour was spent training a model, serving a web page, or sitting idle; it counted elapsed time on a reserved machine, because elapsed time was the only thing the provider could cheaply and unambiguously verify.

That hourly-instance logic carried forward largely unchanged as cloud providers began building machine-learning-specific products on top of general-purpose compute. When Amazon Web Services introduced SageMaker at its re:Invent conference in November 2017, offering a managed way to build, train and deploy models, the underlying commercial pattern did not change even though the product did: SageMaker was announced as a new service for machine learning workflows, but its hosted inference endpoints were billed the same way general compute already was, by the hour an endpoint stayed provisioned, at a rate set by the size of the instance behind it [2]. A model answering one request an hour and a model answering ten thousand cost the operator the same amount, because the bill was for the machine, not for the work the machine did.

This is the baseline every later development in AI inference pricing departed from. Under instance-hour billing, the buyer carries the utilization risk: an idle accelerator still accrues its hourly charge, so cost discipline meant keeping utilization high or turning machines off, not shaping any individual request. That changed once providers began selling access to a specific model over an API rather than access to a machine, because it became possible, for the first time at consumer-software scale, to meter the actual unit of work a language model performs: the token. What follows is the history of how that unit got priced, mispriced, re-priced, and eventually split into the layered rate cards that govern the industry today.

ADVERTISEMENT

An API opens, and a price follows months later

OpenAI opened a commercial GPT-3 API in private beta in June 2020, but for its first stretch the beta ran without any published, general price list — access was granted by application, and the terms a given developer saw were not standardized [4]. Pricing itself was debated in public before OpenAI made anything official: as late as 4 September 2020, one of the more detailed contemporary write-ups of the still-unofficial pricing scheme was sourced not from an OpenAI announcement but from a Reddit post by outside researcher Gwern Branwen, because, as the piece put it at the time, “OpenAI hasn’t yet officially announced the GPT-3 pricing scheme” [4].

When the price list did arrive, effective 1 October 2020, it did not resemble the flat per-token meter the industry now takes for granted. OpenAI published four tiers rather than a single rate: Explore, a free tier good for 100,000 tokens or a three-month trial, whichever came first; Create, 100 dollars a month for 2 million tokens, plus 8 cents for every additional thousand tokens; Build, 400 dollars a month for 10 million tokens, plus 6 cents for every additional thousand; and Scale, priced by individual negotiation with the company [3, 4]. This was a subscription-plus-overage structure lifted directly from enterprise software, not a metering scheme built around what a transformer actually costs to run per request: a monthly access fee bought a token allowance, and only usage above that allowance was billed per token at all. The overage rate on the Build tier nonetheless previewed a number that would matter later — 6 cents per thousand tokens for the most capable engine was, within rounding, close to the rate at which OpenAI would eventually price plain per-token access to that same tier of model, once the subscription wrapper came off.

Reading this stretch of the record in order makes something easy to miss on its own: OpenAI shipped the product before it had a monetization answer, debated the answer semi-publicly for months, and only then converted an ad hoc access grant into a priced plan modeled on software subscriptions it already understood, rather than on any first-principles theory of transformer serving costs. The token was the unit from day one. The pricing structure wrapped around that unit took much longer to settle, and its first stable form was not the one that lasted.

An API gateway rack with a single status lamp just switching on, beside a pinboard where the rate-card slot next to it still sits empty with one pin held ready
Figure 1. The GPT-3 API opened to developers in June 2020 with no published price list at all; a tiered rate card followed only months later.

Four tiers become a token

By November 2021, when OpenAI lifted the application waitlist and let developers in supported countries sign up directly, access had moved from something granted case by case to something available on request: “those wishing to download and use the GPT-3 API have only to submit their email address” as of 18 November 2021 [5]. What the next fully documented price change shows is that, by then or soon after, the 2020 subscription-plus-quota structure was already gone. On 23 August 2022, OpenAI announced a price cut effective 1 September 2022 that reduced every model’s per-thousand-token rate by roughly two-thirds, and it was published as a straightforward per-token rate table across four named engines with no monthly minimum or bundled quota attached: Davinci fell from 6 cents to 2 cents per 1,000 tokens, Curie from 0.6 cents to 0.2 cents, Babbage from 0.12 cents to 0.05 cents, and Ada from 0.08 cents to 0.04 cents [6]. Contemporaneous reporting attributed the cut to growing competition from open-source alternatives, and quoted the company crediting its own serving-efficiency gains for making the lower rate possible — “our teams have made incredible progress in making our models more efficient to run” [6].

Two separate facts sit inside that one announcement, and they get conflated in almost every retelling of this period. One is a change in billing structure: a subscription-plus-quota product became a flat, no-minimum, pay-per-token product at some point in the roughly twenty-two months between the October 2020 tiered plan and the August 2022 price cut — the surviving public record fixes both endpoints of that change but not, from what is independently verifiable, the exact date the flat schedule replaced the tiered one. The other is a change in the price level within that already-flat structure, which is precisely dated: a two-thirds cut announced on 23 August 2022 and effective the following week. Together they set a pattern that has held ever since in this market: named model tiers, priced strictly per token, with headline rate cuts announced as discrete, dated events and tied explicitly to either competitive pressure, serving efficiency, or both at once.

ADVERTISEMENT
A four-card tiered rate sheet being taken down from a pinboard, one card caught mid-fall, while a single new printed strip is pinned in its place
Figure 2. By late 2022 the subscription-tiered plan of 2020 had given way to a flat per-token rate across four named model engines.

The market gets competitors, and competitors get prices

For its first two and a half years, “the API” effectively meant one company’s API. That changed across a compressed stretch of 2023. Microsoft brought Azure OpenAI Service to general availability on 16 January 2023, giving enterprise buyers a second procurement channel, through Azure rather than directly from OpenAI, to four versions of GPT-3 that traded performance for cost, alongside Codex and DALL-E 2, wrapped in Azure’s own billing, uptime and content-filtering commitments [8]. This did not introduce a new pricing philosophy so much as a second distribution channel for an existing one, but it mattered commercially: token pricing now had to be legible inside enterprise procurement processes built around cloud-consumption contracts, not only inside direct developer sign-up.

A more structurally new entrant arrived two months later. Anthropic introduced Claude on 14 March 2023, offering it through both a chat interface and an API, and, notably for pricing history, launched with two tiers from day one rather than four: “Claude, a state-of-the-art high-performance model,” and “Claude Instant, a lighter, less expensive, and much faster option” [7]. Two tiers instead of four collapsed the entire pricing decision to a single deliberate trade-off, capability against cost and speed, rather than OpenAI’s approach of naming four historical checkpoints within one model family. Anthropic’s entry mattered less for any specific number on its rate card at launch than for the fact that a rate card now had to be read against another company’s rate card at all; every pricing decision the incumbent made from that point on was implicitly a competitive one, whether or not it was announced that way.

The third disruption of 2023 did not come from a company selling API access at all. On 18 July 2023, Meta released Llama 2, stating plainly that it was “free for research and commercial use,” with model weights distributed through Azure, Amazon Web Services, Hugging Face and other providers rather than gated behind an API Meta itself operated [9]. This created a pricing comparison that had not existed before in this market: a buyer could now weigh a vendor’s per-token API rate against the fully loaded cost of renting raw compute, by the hour, in the older EC2 and SageMaker sense, to run the same open weights independently, or against a fast-growing set of third-party hosts that re-sold inference over those open weights at their own per-token rates. Token pricing was no longer only a question of which proprietary vendor to buy from; for a meaningful share of workloads it became a question of whether to buy tokens from anyone at all, or return to renting the machine by the hour and running the model oneself.

Three separate vendor rate-card clipboards laid side by side on a bench, one being flipped open to a fresh page while the other two lie settled
Figure 3. 2023 brought a second API vendor, a cloud-marketplace channel, and a freely downloadable open-weight model onto the same comparison bench.

The cache enters the price list

The next structural change to the rate card was not a new vendor or a lower headline rate; it was a new column. Google shipped it first at scale: on 27 June 2024, developers gained context caching in the Gemini API for both Gemini 1.5 Pro and Gemini 1.5 Flash, explicitly framed as a cost tool for “tasks that use the same tokens across multiple prompts” — send a large block of content once, cache it, and pay a reduced rate to refer back to it on later calls rather than repaying full price every time [10].

Anthropic followed on 14 August 2024 with prompt caching for Claude 3.5 Sonnet, Claude 3 Opus and Claude 3 Haiku, reaching general availability on 17 December 2024, and quantified the effect directly: cost reductions of up to 90 percent and response-time improvements reported at up to two times, or as much as 85 percent latency reduction on the longest prompts, with a worked example showing a 100,000-token book conversation running 79 percent faster and 90 percent cheaper once cached [11, 12]. OpenAI shipped its own version at its October 2024 developer event, announcing on 1 October 2024 that prompt caching would apply automatically, “no action needed” from the developer, to GPT-4o, GPT-4o mini and the newly released o1-preview and o1 models, at a 50 percent discount on the cached portion of any prompt over 1,024 tokens [13].

The pattern across all three launches is the same, and it is worth naming because it recurs at every later caching announcement across the industry: caching did not lower the price of a token so much as it split “input token” into two different goods carrying two different prices — a fresh token and a repeated one — and repetition, previously free to the seller and invisible to the buyer, became a line item with its own economics. It also introduced, for the first time in this history, a discount a buyer had to actively earn through how a request was structured, rather than one applied uniformly across the board; providers that made caching automatic, Google and OpenAI, versus opt-in, Anthropic’s original explicit-breakpoint mode alongside a later automatic mode, were making different bets about whether buyers could be trusted to structure their own requests well enough to capture the saving on their own.

ADVERTISEMENT
A memory cache module held just above its slot on an accelerator tray, its contacts not yet seated, with a stack of the same module already fitted in the row beside it
Figure 4. Prompt caching split "input token" into two different goods with two different prices: a fresh token and a repeated one.

Reasoning arrives, and part of the bill goes dark

Every pricing mechanism up to this point billed for something the buyer could see: an input they wrote, an output they received, a cached block they knew they had sent before. That stopped being universally true on 12 September 2024, when OpenAI shipped a preview of o1, a model built to spend extended computation deliberating before answering [16]. o1-preview launched priced at 15 dollars per million input tokens and 60 dollars per million output tokens, several times GPT-4o’s rate at the time [17]. The deliberation itself, the “reasoning tokens,” was billed at that 60-dollar output rate but never shown to the caller; OpenAI’s own later documentation states plainly that reasoning tokens “are not visible via the API” yet “still occupy space in the model’s context window and are billed as output tokens” [18], and a model-pricing reference for o1-preview notes the same fact even more bluntly: “reasoning tokens are priced identically to output tokens” [17].

The reaction was immediate, and it is itself part of this pricing history, because it shaped how every subsequent reasoning-model launch was received. A widely discussed Hacker News thread from the days after launch captured the objection precisely: one commenter noted that “you can’t verify that you’re paying what you should be if you can’t see the hidden tokens,” and others flagged the practical consequence that a request failing partway through reasoning could still be billed in full for computation that produced no visible output at all [19]. OpenAI’s stated rationale, competitive protection of its reasoning method together with safety monitoring of an unfiltered chain of thought, explained the design choice without resolving the billing objection it created; Wikipedia’s account of the model notes the restriction on revealing o1’s reasoning “has been described as a loss of transparency by developers who work with large language models” [16].

OpenAI moved o1 from preview to full release on 5 December 2024, and kept extending the same reasoning-token mechanism upward: a separate o1-pro tier, released in March 2025, priced deliberation and answer alike at 150 dollars per million input tokens and 600 dollars per million output tokens, ten times o1-preview’s own already-elevated rate [16]. Each step up that ladder repeated the same structural choice made at launch — more billed reasoning, still none of it shown — rather than introducing a new one, which is part of why the mechanism’s critics kept returning to the same objection rather than raising a new one each time.

The clearest illustration that reasoning-token billing was a genuinely new cost lever, not a marginal one, arrived from a direction OpenAI did not control. On 20 January 2025, the Chinese lab DeepSeek released R1, a reasoning model it positioned as matching OpenAI’s o1 on several published benchmarks, priced at 14 cents per million input tokens on a cache hit, 55 cents per million input tokens on a cache miss, and 2.19 dollars per million output tokens, a fraction of o1’s published rate on every line of the card [20]. The market reaction arrived within a week: on 27 January 2025, Nvidia’s stock fell nearly 17 percent in a single session, erasing roughly 593 billion dollars of market value, described as the largest one-day market-capitalization loss for any company in Wall Street history, as investors reassessed whether frontier-model economics justified the scale of accelerator spending the market had been pricing in [21]. Whatever the eventual full technical explanation for DeepSeek’s cost structure, the episode is a clean illustration of a mechanism this history has shown building since 2022: a published token rate is a competitive signal read well beyond the developers who actually call the API, all the way out to equity markets pricing the hardware underneath it.

A twin-tape billing printer where a narrow unlabeled tape spools out behind the visible itemised receipt, only the visible tape torn off and taken
Figure 5. A reasoning model's deliberation is billed at the output rate and never shown to the buyer who paid for it.

Batch queues, fast lanes, and the tiered present

The last structural addition to the rate card arrived alongside caching, and for a related reason: not every request needs to be served at the same latency, and a provider willing to defer a request can serve it more cheaply. OpenAI introduced a Batch API on 15 April 2024, offering a 50 percent discount against standard per-token rates in exchange for asynchronous processing with results guaranteed within 24 hours [14]. Anthropic followed on 8 October 2024 with a Message Batches API at the same 50 percent discount on both input and output tokens, reaching general availability on 17 December 2024, the identical date its prompt-caching feature also left beta, folding two of the period’s major pricing innovations into general availability within a single release window [15].

The rate card that results, as documented on each provider’s current pricing pages, is no longer one number per model but a small grid: a standard rate, a batch rate at roughly half of standard, a cached-read rate at roughly a tenth of standard, and, on at least one major provider, a premium “fast” rate above standard for buyers willing to pay for lower latency rather than accept a discount for tolerating more of it [23, 24]. Anthropic’s current documentation states this multi-axis structure explicitly: batch, caching and long-context pricing multipliers are all designed to stack, so a single request’s effective per-token cost can differ from the headline rate by an order of magnitude depending on how it was scheduled, how much of it was cached, and where it was permitted to run geographically [24].

Zoom out across the roughly two decades this history covers and the direction of travel is unambiguous even where the underlying mechanism kept changing. Epoch AI’s tracking of inference prices at fixed capability thresholds found the cost of reaching GPT-3-level general-knowledge performance fell from 60 dollars per million tokens in November 2021 to 18 cents per million tokens by February 2025, and the organization cautioned in the same analysis that the fastest of those declines were concentrated in the most recent year measured, making the trend’s persistence genuinely uncertain rather than a settled law [22]. Every mechanism this article has traced, hourly instances, subscription tiers, flat per-token rates, competitive entry, prompt caching, batching, and billed reasoning tokens, is a different technical and commercial answer to the same underlying pressure: an operator trying to charge for a unit of computation whose true marginal cost the buyer can never directly observe, worked out one dated rate-card change at a time.

A sorting rack with two labelled trays, a small paper ticket caught in mid-drop toward the slower tray while the faster tray already holds a settled stack
Figure 6. The current rate card is a grid, not a single number: a standard rate, a batch discount, a cache discount and a latency premium can all apply to the same request.

What this history does, and does not, explain

This article has deliberately stopped at the level of the rate card: what was announced, on what date, at what price, and in what structure. It has not tried to derive those prices from the underlying compute costs, model architectures or serving systems that made them possible — that is a different, analytical question, better answered by examining prefill and decode costs, batching mathematics and memory bandwidth directly than by reading pricing announcements in sequence. Epoch AI’s aggregate price-decline data shows the trend has been sharply downward, but a trend line does not explain any single provider’s decision to cut a specific rate on a specific date, any more than knowing that hardware improves over time explains why OpenAI cut prices in September 2022 specifically rather than in August or November of that year [22].

What the dated record does support is a narrower and more useful claim: every major structural feature of today’s inference pricing, the token as the billing unit, the tiered model family, the cache discount, the batch discount, the hidden-but-billed reasoning token, was introduced as a specific, attributable, datable event, usually with a stated commercial or competitive rationale attached at the time it happened. None of it was inevitable in the way that falling hardware costs are loosely inevitable. Each layer was a choice, made by a named company on a named date, most often in direct response to a competitor’s or a critic’s prior choice. The rate card in front of any buyer today is not a physical law describing the cost of computation. It is the accumulated, still-legible record of about two decades of those choices, still being added to.