Two instruments, one word

“Context window” names one thing on a specification page and a different thing on a benchmark leaderboard, and this publication’s companion piece on Claude’s long-context behaviour has already documented how wide that gap can run for a single vendor’s models [3]. This article does not repeat that ground. It asks the question the series was built to answer: when OpenAI, Anthropic, and Google each publish a token count for how much a frontier model can hold, and independent researchers separately measure how much of it a model actually uses correctly, how do the three vendors’ documented numbers compare, and how do their independently measured numbers compare — kept as two separate questions, because the two numbers come from two different instruments and answering with only one of them misleads.

That last clause matters for how this piece is built. The series rule governing it is explicit: OpenAI and Anthropic receive primary, detailed coverage; Google receives smaller, evidence-proportionate coverage; and nothing here should be read as a single ranking assembled across benchmarks that were not designed to be compared to one another. Where the evidence lets three providers be placed side by side on one benchmark, they will be. Where it does not, the gap itself is reported, because an unfilled cell in a comparison table is information too.

What each provider currently documents

Start with what is verifiable without running anything: the token counts each vendor states for its own current models, as published in their own developer-facing documentation.

ADVERTISEMENT

OpenAI’s current API documentation lists GPT-5.6 Sol, the frontier model in its GPT-5.6 family, with a context window of 1,050,000 tokens, of which up to 922,000 may be supplied as input, and up to 128,000 tokens may be generated as output in a single request; the same entry records a knowledge cutoff of 16 February 2026 [1]. That figure did not appear from nowhere. OpenAI’s April 2025 introduction of GPT-4.1 raised the model family’s context window to one million tokens, an eightfold increase over the prior 128,000-token ceiling, and the announcement paired that number with a specific internal claim: on OpenAI’s own needle-in-a-haystack evaluation, “GPT-4.1 consistently retrieves the needle accurately at all positions and all context lengths, all the way up to 1 million tokens” [2]. Read that sentence as what it is — a vendor’s report of its own internal test, not an independently reproduced result — and then read the sentence that follows it in the same announcement, because OpenAI states its own limitation plainly: “few real-world tasks are as straightforward as retrieving a single, obvious needle,” and users typically need a model to retrieve and relate multiple pieces of information rather than one. OpenAI’s own response to that gap was to publish a second, harder benchmark alongside the first: MRCR, short for Multi-Round Co-reference Resolution, which hides two, four, or eight near-identical requests throughout a long synthetic conversation and asks the model to retrieve the one matching a specific instance rather than any of the decoys [2]. A vendor naming its own headline test’s limitation and then building a harder one is worth recording as a fact distinct from the headline number itself.

Close view of one test rig's supply reel with a section of waxed paper tape lifted clear to show how much more remains wound on the reel than has passed through the reader head
Figure 1. A documented window size is a property of the reel; how much of it a model actually uses correctly is read off the head further down the bench, and the two figures are rarely the same number.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Anthropic’s current documentation states that Claude Opus 5 and several recent Opus and Sonnet models, along with Claude Fable 5 and Claude Mythos 5, carry a one-million-token context window on the Claude API at standard pricing, with no beta header required; Claude Sonnet 4.5 and Claude Haiku 4.5 remain at two hundred thousand tokens [3]. That figure has a documented history rather than having simply appeared at its current size: Anthropic first opened a one-million-token window to Claude Sonnet 4 in August 2025, a fivefold increase over the prior two-hundred-thousand-token limit, priced at a premium above the standard rate for any request drawing on more than two hundred thousand tokens — six dollars per million input tokens and twenty-two dollars fifty cents per million output tokens, against three dollars and fifteen dollars respectively below that threshold [4]. Current documentation confirms that premium has since been removed; the larger window is now the default at standard pricing on every model that supports it [3]. What is notably absent from Anthropic’s current context-window documentation is a headline recall percentage of the kind OpenAI and Google both lead with. In its place, the same page states a limitation directly, in its own words, as an established fact rather than a caveat appended after a strong claim: “As token count grows, accuracy and recall degrade, a phenomenon known as context rot” [3]. That documentation posture — naming the failure mode rather than a headline retrieval percentage — is itself a data point about how the three vendors currently choose to describe the same underlying gap, and it is reported here as exactly that: a difference in what gets stated first, not evidence that any one vendor’s underlying model recalls better or worse than the others.

Google’s documentation states that its Gemini models “come with large context windows of 1 million tokens or more,” with Gemini 1.5 Pro supporting up to two million tokens through the API [5]. Consistent with the series rule for this comparison, Google’s coverage here is proportionate to what independent evidence supports rather than expanded to match OpenAI’s and Anthropic’s sections — but the documentation is worth quoting precisely because it states its own caveat as plainly as Anthropic’s does. Google’s long-context guidance reports “up to 99% accuracy in many cases” on needle-in-a-haystack retrieval, then immediately qualifies it: “These tests consider the most basic setup, where you have a single needle you are looking for. In cases where you might have multiple needles… the model does not perform with the same accuracy. Performance can vary to a wide degree” [5]. All three vendors, in other words, currently publish some version of the same warning in their own documentation, in three different registers. That convergence is itself informative before a single independent benchmark is examined: none of the three companies whose own commercial interest runs toward a bigger, more impressive context-window number is currently claiming that the number alone settles the question of usable recall.

Google’s own earlier self-testing supplies the clearest illustration of that qualifier’s size, and it is worth reporting precisely because Google published it about its own model. The Gemini 1.5 technical report states near-perfect retrieval, above 99 percent, “up to at least 10M tokens, a generational leap over existing models such as Claude 3.0 (200k) and GPT-4 Turbo (128k)” [6] — a direct vendor comparison against two competitors’ models, both since superseded, and reported here as a claim made in Google’s own paper rather than as an independently confirmed ranking. A companion post from Google Cloud, published the same year, breaks that headline figure apart by task type on Gemini 1.5 Pro’s internal testing: better than 99.7 percent recall up to one million tokens and 99.2 percent recall at ten million tokens on the single-needle version of the test, but only 75 percent accuracy on a multi-round conversational-recall task spanning one million tokens, and only 60 percent recall on a version of the test with multiple needles hidden in one million tokens of context [7]. That same post also reports its own comparison figure for a competitor of the era, stating that GPT-4 Turbo averaged roughly 50 percent recall at its maximum 128,000-token context length with performance oscillating substantially [7] — a claim this article repeats only as an attributed statement from Google’s own comparative marketing material about a model OpenAI has since replaced twice over, not as an independently verified figure, and not as evidence about any model discussed elsewhere in this piece. The instructive result is the first half: even in a vendor’s own internal testing of its own model, the score on the easy single-needle version and the score on the harder multi-needle version differ by close to forty percentage points. That gap belongs to Google’s own published numbers, about Google’s own model, before any outside benchmark is involved at all.

The single test everyone still runs first, and what running it more only tells you

The design under all three vendors’ self-reported figures traces back to one open-source test an engineer named Greg Kamradt released in 2023: take an ordinary document long enough to fill a context window, insert one out-of-place sentence at a chosen depth, and ask the model to find it, sweeping both the depth and the total length to produce a grid of pass or fail results [12]. Every major provider adopted some version of it because it is cheap to run at any length a lab wants to advertise and it produces one clean, marketable number.

ADVERTISEMENT
A slim vertical pin-gauge mounted over one rig's paper tape, caught mid-drop toward a marked depth on the tape, not yet seated
Figure 2. The oldest long-context test buries one marked fact at a chosen depth and asks whether it comes back; every provider's rigs still run some version of it first.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

It is also, on its own, a comparatively easy test, and that fact is not a criticism invented by a critic outside the labs — OpenAI’s own GPT-4.1 documentation says as much when it describes the standard needle test as less demanding than real retrieval tasks, which is precisely why MRCR exists [2]. Google’s own long-context guidance makes the identical point about the same test in its own words: strong performance on the single-needle version does not carry over to the multiple-needle version [5]. When two competing vendors independently publish the same limitation about the same test design, in each case immediately after reporting a strong result on that test, the limitation deserves more weight than either vendor’s headline number does on its own. The question this leaves open is how large the gap actually runs once a benchmark is built specifically to find it — and that is a question the vendors’ own self-administered tests cannot answer, because a lab grading its own model on a test it also designed is not the same kind of evidence as an outside benchmark grading several vendors’ models on one shared, published methodology.

RULER: the first shared scoreboard, and who is missing from it

Researchers at NVIDIA built RULER specifically to move past the vanilla needle test, adding multi-needle retrieval, multi-hop variable tracing across the document, and aggregation tasks that require using many relevant pieces of the context rather than filtering everything else out [8]. The published cohort tested seventeen long-context language models across thirteen such tasks — fifteen open-weight models, and exactly two closed, commercially served models: GPT-4 and Gemini-1.5-Pro. No Claude model appears in RULER’s original published cohort, a gap this article states directly rather than filling with a number that was never measured.

On the two closed models that were tested, RULER’s published results show a clear divergence in shape rather than in a single blended score. GPT-4 scored 96.6 percent at four thousand tokens, still 96.3 percent at eight thousand, then fell steadily through 95.2 percent at sixteen thousand, 93.2 percent at thirty-two thousand, 87.0 percent at sixty-four thousand, and 81.2 percent at the study’s maximum tested length of 128,000 tokens; the paper’s own effective-length measure — the longest length at which a model’s score stays within a defined tolerance of a fixed reference standard — places GPT-4 at sixty-four thousand tokens, and its length-weighted average score across the tested range, which the authors compute to penalise degradation at longer lengths more heavily than a simple mean would, comes to 89.0, second among the seventeen models tested [8]. Gemini-1.5-Pro held closer to a flat line, scoring in the mid-nineties consistently across every tested length up to 128,000 tokens, with an effective length exceeding the study’s own maximum tested point and a length-weighted average of 95.5, the highest of the seventeen [8]. Read carefully, that ranking is a statement about seventeen specific models on thirteen specific synthetic tasks at lengths up to 128,000 tokens in early 2024 — not a general verdict on OpenAI versus Google, and not a result that says anything at all about Anthropic’s models, which this particular study never ran. RULER’s own headline finding matters independently of which vendor comes out ahead on it: despite near-perfect scores on the vanilla single-needle version of the test, “almost all models exhibit large performance drops as the context length increases” on the harder task set, and of the models in the study claiming to support 32,000 tokens or more, only half sustained what the authors call satisfactory performance at that length [8].

NoLiMa: what happens once the question stops sharing the needle’s words

A second, more recent benchmark attacks a different weakness in the original design. Kamradt’s test and RULER’s variants generally allow a needle that shares obvious vocabulary with the question asked about it, which lets a model succeed through something closer to string matching than genuine reasoning over content. NoLiMa, built by researchers publishing through ICML, replaces those needles with facts stated so that no wording in the question repeats wording in the needle, forcing the model to infer an association rather than locate a matching string [9]. Thirteen widely used models, each advertised to support at least 128,000 tokens of context, were tested at length steps from roughly one thousand tokens up to the maximum each model’s own documentation claimed.

A row of small blank mechanical flip-counters at the end of a test lane, most standing on their hit face, one caught mid-flip toward its miss face as the sweep reaches a longer length
Figure 3. Independent benchmarks that keep testing after the easy version passes find scores falling well before the documented window closes, and the fall is not the same shape for every model.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

This is the one benchmark in the current evidence base that reports comparable, matched-length numbers across all three vendors on one shared methodology, and it is worth presenting in full rather than summarised. GPT-4o scored 99.3 percent at NoLiMa’s short reference length, still 98.1 percent at one thousand tokens and 95.7 percent at four thousand, then declined to 89.2 percent at eight thousand, 81.6 percent at sixteen thousand, and 69.7 percent at thirty-two thousand tokens; an extended table in the same study carries GPT-4o further, to 62.4 percent at sixty-four thousand tokens and 56.0 percent at 128,000 [9]. Gemini 1.5 Pro started from a lower short-length baseline of 92.6 percent, falling to 82.7 percent at two thousand tokens, 75.4 percent at four thousand, 63.9 percent at eight thousand, and 48.2 percent at thirty-two thousand tokens. Gemini 2.0 Flash started at 89.4 percent, held closer to its baseline through the early lengths — 87.5 percent at two thousand tokens — before falling to 41.0 percent at thirty-two thousand, 33.0 percent at sixty-four thousand, and 16.4 percent at 128,000 tokens, the steepest late-length decline the extended table records for any of these four models [9]. Claude 3.5 Sonnet’s reported baseline was lower still, at 87.5 percent, falling to 61.7 percent at eight thousand tokens and 29.8 percent at thirty-two thousand, the study’s steepest decline among the four models at that specific length; the paper’s extended sixty-four- and 128,000-token figures were reported for GPT-4o and Gemini 2.0 Flash but not, in the portion of the results available here, for Claude 3.5 Sonnet or Gemini 1.5 Pro, itself a small illustration of how unevenly complete even one shared benchmark’s coverage can be across vendors [9].

Three things follow from that table, and none of them is “provider X beats provider Y.” First, the four models did not start from the same short-context baseline, so a raw score at any given length conflates two different quantities — how accurate the model is to begin with, and how much of that accuracy it retains as length grows — and a fair reading has to track both rather than either alone. Second, by the retention measure, GPT-4o kept roughly seventy percent of its own short-length score at thirty-two thousand tokens, Gemini 1.5 Pro kept roughly fifty-two percent, and Claude 3.5 Sonnet kept roughly thirty-four percent, an ordering that is not identical to the ordering of raw scores at that length. Third, and most important for dating this evidence honestly: GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Gemini 2.0 Flash are, as of this article’s publication, prior-generation models. NoLiMa is a real, peer-reviewed, methodologically careful result, and it is also already a historical snapshot of a model generation the vendors themselves have since superseded with GPT-5.6, Claude Opus 5 and Sonnet 5, and newer Gemini models. The paper’s finding that degrades reliably survives the passage of a model generation; the specific percentages attached to specific named models do not automatically transfer to their successors, and this article makes no claim that they do.

ADVERTISEMENT

The same underlying failure, three different shapes

Not every long-context weakness shows up as a declining accuracy percentage. A large industry study from Chroma, a company that builds retrieval infrastructure, tested eighteen current models — spanning OpenAI, Google, and Anthropic families — across several extended needle-retrieval tasks, a conversational-recall benchmark, and a task requiring exact verbatim replication of repeated text, and it found that degradation with length is real across the board but takes visibly different forms in different model families rather than one uniform curve [10].

Within OpenAI’s family, the study’s most striking finding was about confidence rather than raw accuracy: GPT-4.1 produced the highest hallucination rate under distraction of the models tested, generating confident, fluent, and incorrect answers rather than declining to answer, and it refused the exact-replication task outright in 2.55 percent of attempts, with refusals typically beginning once the text to be replicated passed roughly 2,500 words [10]. An older model in the same family, GPT-4 Turbo, showed a different and oddly specific pattern on that replication task: a local peak in performance around 500 words, preceded by mild overgeneration between 50 and 250 words and followed by undergeneration past the 500-word mark, a shape with no obvious single cause and no equivalent reported for the other vendors’ models [10]. Two smaller models in the same family showed still other idiosyncrasies: GPT-4.1 mini occasionally produced words absent from the source text entirely for certain repeated phrase pairs, and GPT-4.1 nano produced unexpected casing changes on specific tasks [10].

A light table where a rig's output tape is laid over its source tape for comparison, most sections aligning exactly, one small spliced-in loop of extra tape still curling and not yet trimmed flush
Figure 4. Different providers' failures under distraction take different physical shapes in the record, not one uniform decline, which is why a single ranking number hides more than it shows.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Google’s Gemini models showed a related but distinct failure signature on the same replication task: a tendency to generate words that were not present anywhere in the source input, typically beginning once the passage reached somewhere between 500 and 750 words, with Gemini 2.5 Pro showing the greatest variability in this specific behaviour among the Gemini variants tested [10]. That is a different failure from confident hallucination under distraction and a different one again from a sharp accuracy cliff at a specific token count — it is a fabrication pattern tied to a specific task type rather than to length in the abstract. Claude models were tested in the same study and behaved differently again, a finding this publication’s companion Claude piece has already reported in detail and will not be repeated here [10]. The point worth carrying forward is structural rather than comparative: three vendors’ models, tested on the same tasks by the same independent team, did not fail in the same way or even along the same axis, which is precisely why compressing this evidence into a single cross-vendor leaderboard number would misrepresent what was actually found.

Independent benchmarks do not agree with each other either

One further complication belongs in this account before any synthesis, because it cuts against a natural assumption: that “independent” and “in agreement” mean the same thing. A separate benchmark called Loong, built by a team publishing through EMNLP, argues that RULER’s style of test and others like it still rely on filler text that is largely irrelevant to the question being asked, so a model can score well by locating the relevant fragment and ignoring the rest — which is closer to retrieval than to the genuine multi-document reasoning real long-context tasks often require [11]. Loong is built so that “ignoring any document will lead to the failure of the answer,” and on that stricter standard the paper reports that contemporary long-context models still have substantial room for improvement, and that feeding a model retrieved excerpts underperforms giving it the full document set when every document genuinely matters [11]. Loong’s methodology and RULER’s are not measuring the same construct, so a model’s relative standing on one is not evidence about its standing on the other, and neither is more “correct” than the other — they simply probe different failure modes. This is exactly the trap the series-wide editorial rule for this comparison exists to prevent: stitching together scores from benchmarks that were never designed to be stitched together produces a ranking that looks precise and is not.

Why the most confident current claims are the least independently checked

Put the documentation and the independent benchmarks side by side and a pattern emerges that is more informative than any single score. Formalise the distinction the whole comparison rests on: let WW be the context window a vendor documents for a given model, a fixed engineering figure set by attention implementation, position encoding, and what the serving stack supports. Let S(n)S(n) be some benchmark’s measured accuracy at input length nWn \le W, and S0S_0 the same benchmark’s accuracy at a short reference length. For a chosen retention threshold τ\tau, define the effective context length as

Leff(τ)=max{nW : S(n)τS0} L_{\mathrm{eff}}(\tau) = \max\left\{\, n \le W \ :\ S(n) \ge \tau \cdot S_0 \,\right\}

WW is a documented constant, published on day one of a model’s release, identical no matter who asks. Leff(τ)L_{\mathrm{eff}}(\tau) is a measured variable — specific to a task, a threshold, a benchmark design, and a date — and every result surveyed above shows it running well below WW well before WW is reached. The two quantities answer different questions, and only one of them is available the moment a model ships.

A steel comparison rail running across three test lanes with three sliding stops, two already set to the same tick mark and the third caught still sliding toward it
Figure 5. Comparing three providers honestly means holding all three to the same tick mark on the same rail, not to whichever length each one's own documentation chose to headline.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

That timing gap is the article’s central finding, stated plainly: the newest, most-advertised context windows are, at the moment of writing, verified almost entirely by the vendors that built them. GPT-5.6 Sol’s 1,050,000-token window comes from OpenAI’s own documentation [1]. Claude Opus 5’s one-million-token window comes from Anthropic’s own documentation [3]. Current Gemini models’ million-token-plus windows come from Google’s own documentation [5]. None of RULER, NoLiMa, or the Chroma study evaluated any of these specific current models, because each was published before they existed. Independent, peer-reviewed, cross-vendor measurement of usable recall takes months to design, run, and publish after a model ships, and it typically arrives once the model it describes has already been superseded by the vendor’s next release — a genuine structural lag, not a flaw in any one study, and not evidence that current models perform worse than their predecessors did. It is evidence that the confidence a reader should place in any headline context-window number should be lower, not higher, the more recently that model shipped, simply because independent verification has not yet had time to catch up to it.

What this means for choosing a model against a real task

None of the above supports picking a smaller context window on principle, and none of it supports trusting a stated window as a guarantee. It supports a specific, checkable habit: test the actual task, at the actual length it requires, against the actual success criterion that matters, rather than trusting either a vendor’s headline retrieval percentage or an independent benchmark’s score on a task that resembles but is not the job at hand. A single-needle recall figure, self-reported or independently measured, describes single-needle recall; it does not describe multi-hop reasoning across scattered facts, aggregation across many relevant passages, or verbatim replication under distraction, and the evidence above shows those are measurably different skills that degrade at different rates and in different shapes even within one vendor’s own model family. Where a task genuinely requires holding an entire document coherent — a codebase, a contract, an agent’s own accumulated tool history — a documented window is the right ceiling to check first. Where the material is large and heterogeneous and most of it is irrelevant to any given query, the evidence here argues for treating the full window as a last resort rather than a default, and for building in some form of curation or retrieval ahead of it, regardless of which of the three vendors is supplying the model.

Predictions, with the observations that would falsify them

These are forecasts, clearly separated from the sourced findings above. Horizon: 12 August 2028.

One. Independent, peer-reviewed long-context benchmarks that include all three of OpenAI’s, Anthropic’s, and Google’s current flagship models on one shared methodology will remain rare relative to the pace of model releases, so that at any given time, the newest flagship from at least one of the three vendors will lack independent multi-vendor verification. Disconfirmed if, by 2028, a maintained, regularly updated cross-vendor long-context leaderboard exists that adds each new flagship model from all three vendors within one month of its release.

Two. Vendors will continue to supplement, rather than replace, single-needle retrieval claims with harder self-administered benchmarks of the kind OpenAI built with MRCR, because the single-needle number alone will have become recognised as insufficient marketing evidence on its own. Disconfirmed if a 2028-era frontier model launch from any of the three vendors leads with only a single-needle recall percentage and no harder multi-fact or non-literal benchmark alongside it.

Three. The gap between a raw accuracy score at a given length and that score’s retention relative to a short-context baseline will be reported more often in independent benchmarks, because the evidence above shows the two rank models differently and only one of them isolates degradation from starting ability. Disconfirmed if 2028-era long-context papers still report only raw accuracy at length with no baseline-relative retention figure.

Four. No single vendor will hold a durable, task-general lead in effective long-context recall across every benchmark design, because the evidence already shows the ordering changes with what is being tested — RULER, NoLiMa, and Loong do not agree with each other on which model family degrades least. Disconfirmed if, by 2028, one vendor’s models lead on effective context length across RULER-style, NoLiMa-style, and Loong-style tasks simultaneously and by a wide, consistent margin.

What to take away

A context window in tokens is a real, documented, and independently checkable engineering figure, and all three vendors currently publish one honestly as far as it goes. It is not, on its own, a measurement of what a model will actually recall correctly from material placed inside it, and every piece of independent evidence surveyed here — RULER’s divergent curves for the only two closed models it tested, NoLiMa’s four-way table showing different baselines and different retention rates across three vendors’ models, and Chroma’s finding that degradation takes a different shape in each family — supports treating the two numbers as separate claims rather than one. The most important structural fact in this comparison is not which vendor’s model recalls best; the evidence available does not support a single answer to that question, and building one anyway would be the exact error this series was built to avoid. It is that the newest, most confidently advertised numbers are the ones with the least independent verification behind them, because verification takes time a launch announcement does not wait for. Ask which instrument produced the number in front of you — the vendor’s own specification, or an outside benchmark built to test it — before treating either one as settled.