Ranking is the wrong instrument

Every few months a new leaderboard reorders the frontier: whichever lab shipped last looks best until the next one ships. That contest is not what this article is about. Ranking answers “which model is ahead this week,” and the honest answer keeps changing. A different, more durable question sits underneath it: which failures show up in more than one provider’s own systems, regardless of who is ahead. That question has a stable answer, and it is more useful to anyone procuring, evaluating, or building on these systems than a snapshot of this quarter’s chart.

The ten failure modes below were selected on one criterion: each has documented evidence from a specific provider’s own materials, or from an independent study, and where possible from more than one. None of them is presented as a universal law of “AI” in the abstract — that framing is exactly the one that lets a real, specific, sourced finding get smuggled into an unfalsifiable generalization. Each claim below is attributed to whoever produced it: a lab’s own technical report, a lab’s own postmortem, or a named independent research group. Where a claim is a vendor’s self-report, it is marked as one. Where two vendors are documented doing the same thing independently, that is the strongest form of evidence this article can offer, and it is flagged explicitly rather than folded into a single unattributed sentence.

One assumption threads through everything that follows, and it is worth making explicit before the catalogue starts. A published benchmark score is not a direct read of capability; it is a sum of capability plus at least two other things that vary independently of it:

ADVERTISEMENT
sobserved=scapability+Δcontamination+Δconfiguration s_{\text{observed}} = s_{\text{capability}} + \Delta_{\text{contamination}} + \Delta_{\text{configuration}}

Δcontamination\Delta_{\text{contamination}} is inflation from the model having seen the benchmark, or something close to it, during training. Δconfiguration\Delta_{\text{configuration}} is variance from everything about how the question was asked, formatted, and scored — prompt wording, decoding settings, which variant answered. Both terms are provider-specific and time-specific; neither is disclosed by the headline number. Three of the ten failure modes below are different ways that Δcontamination\Delta_{\text{contamination}} and Δconfiguration\Delta_{\text{configuration}} get large without the label on the chart changing.

Where the benchmark itself misleads

1. Benchmark scores do not predict real-world task completion. A percentile on a static exam and the ability to carry out an open-ended task are different quantities, and the gap between them is the oldest and best-documented failure on this list. METR built a different kind of measurement to make the gap visible directly: the “50%-task-completion time horizon,” the length of task (measured in human completion time) that a model can complete with 50% reliability. Testing twelve models released between 2019 and 2025 on a mix of software and research tasks, METR found this horizon has been doubling roughly every seven months since 2019, with Claude 3.7 Sonnet reaching around fifty minutes at the time of the study [1]. That metric does not track ordinary benchmark accuracy at all — a model can answer a hard multiple-choice question correctly while still losing the thread of a task that takes an hour to carry out coherently, because coherence over time and correctness on a single item are not the same capability.

The gap shows up just as clearly the other direction: a specific, widely repeated capability claim, checked and found smaller than advertised. OpenAI’s original framing of GPT-4 emphasized that it scored around the 90th percentile on the Uniform Bar Exam. Eric Martínez’s peer-reviewed reanalysis, published in Artificial Intelligence and Law, found that the comparison group behind that figure was skewed toward repeat test-takers who had already failed the exam — a lower-scoring population than the general pool. Measured against first-time test-takers from a July administration, Martínez found GPT-4’s overall percentile fell below the 69th, and its performance on the written-essay portions fell to roughly the 48th percentile among practicing attorneys [2]. This is not a claim that the exam score was fabricated; it is a documented case of a true number describing a smaller effect than the framing implied, once an independent researcher chose a different, arguably more appropriate comparison group. The lesson generalizes past this one exam: a percentile is only as informative as the population it is measured against, and that population is a choice, not a fact handed down with the score.

Where the benchmark itself misleads, continued

2. Public benchmarks leak into training data, and the leak is hard to see from outside. Every major lab now checks for this, which is itself evidence of how common it is expected to be. OpenAI’s GPT-4 technical report describes a substring-collision method — sampling short spans from each evaluation item and checking whether they appear verbatim in the training corpus — and reports that the check caught portions of BIG-bench that had been inadvertently mixed into training data, which were then excluded from reported results; OpenAI’s own conclusion was that residual contamination had little effect on the model’s zero-shot scores [3]. That is a vendor’s self-audit, and it should be read as one: a lab reporting on its own contamination is not the same evidentiary weight as an outside party checking the same thing.

The outside check exists, and it tells a more mixed story. Hugh Zhang and colleagues built GSM1k, a fresh grade-school arithmetic benchmark matched to the widely used GSM8K in style, difficulty, and human solve rate, specifically so it could not have leaked into any model’s training data the way the original had. Across the model families tested, accuracy on GSM1k dropped by as much as eight percentage points relative to GSM8K, and the size of a given model’s drop correlated with how readily it reproduced GSM8K items verbatim — while some frontier models showed close to no gap at all [4]. Read together, the two sources agree on the mechanism and disagree, usefully, on the magnitude: contamination is real, checkable, and unevenly distributed across models and benchmarks, not a fixed tax every score pays equally.

ADVERTISEMENT

The most striking recent evidence for this failure mode is also the most surprising in its source. In February 2026, OpenAI’s own Frontier Evals team published an analysis retiring SWE-bench Verified — a coding benchmark OpenAI itself had released roughly two years earlier and that had become an industry-standard reference point — after finding that a majority of a sampled set of previously unsolved “hard” tasks had flawed test design, and, more pointedly, that models could reproduce the benchmark’s gold-standard patches or problem statements essentially verbatim given only the task identifier, with no other prompting. The contamination signal was not confined to OpenAI’s own models: reporting on the analysis names Claude Opus 4.5 and Gemini 3.1 Pro alongside GPT-5.2 as systems showing the same effect [5]. A lab publicly retiring a benchmark it created, on the grounds that the whole field’s models — not just its own — had absorbed the answers, is about as strong a piece of cross-vendor evidence for this failure mode as currently exists in public.

A bound archive volume pulled halfway off a shelf of identical volumes, its pages fanned, held beside the intake slot of a running evaluation rig
Figure 1. A public benchmark and its published answer key can end up on the same shelf a model was trained from; contamination checks exist because nobody outside the lab can be sure which shelf that was.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

3. Labs can tune a variant specifically for a public evaluation, then ship something else. This is contamination’s more deliberate sibling: not the benchmark leaking into the model by accident, but a model built or selected specifically to perform on the benchmark, with no guarantee the released product matches it. In April 2025, Meta submitted “Llama-4-Maverick-03-26-Experimental” — an unreleased, specially tuned chat variant, more verbose and heavier with emoji than the shipped model — to the LMArena leaderboard, where it ranked second. Once the publicly released Maverick was tested independently on the same leaderboard, it ranked below GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro [6]. The two sides of the dispute are worth stating separately rather than collapsed into one verdict: LMArena’s own postmortem attributed the score gap to the experimental variant’s tone and formatting, and stated that Meta should have disclosed the substitution; Meta’s spokesperson denied training on benchmark test sets and attributed some of the disparity to implementation differences across serving platforms. What both sides agree on is the underlying fact — the number on the leaderboard and the model available to the public were, for a period, not the same system. LMArena subsequently changed its submission policy to require that leaderboard entries match released models.

The SWE-bench Verified retirement described above belongs here too, from a different angle: benchmarks that stay in wide use for years become targets, intentionally or not, and a score that looked meaningful in year one can stop meaning much in year three without any single actor doing anything identifiably improper. OpenAI’s recommended replacement, SWE-bench Pro, was itself later found by OpenAI’s own audit to have roughly a third of its public tasks flawed [5] — evidence that building an evaluation immune to this failure mode is harder than retiring the previous one.

A rig's automated card feeder cycling through a stack of uniform test cards beside an open tray holding a single odd-shaped unscripted object
Figure 2. A rig tuned against a scripted lane of uniform cards says nothing about the open tray next to it; the two are rarely measured on the same afternoon.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

What a high score doesn’t show

4. High benchmark scores coexist with confident, fluent fabrication. These are not contradictory facts about the same system; they are largely independent, because standard training and scoring procedures reward a confident guess over an honest “I don’t know.” OpenAI’s own account of this, published under the title “Why Language Models Hallucinate,” argues that hallucination persists specifically because most evaluation metrics grade guessing more favorably than admitted uncertainty — a model that always guesses will outperform one that abstains on a standard accuracy-only leaderboard, even though the second model is more trustworthy [9]. That is a mechanism claim from the vendor whose models are affected, not an outside audit, so it is worth pairing with a direct measurement. Vectara’s Hallucination Leaderboard, an independent benchmark measuring how often a model introduces unsupported claims while summarizing a supplied document, reported in its late-2025 update rates ranging from roughly 3% for the best-performing tested model up past 13% for others — with several current frontier systems, across more than one provider, sitting above 10% [10]. A model can be state-of-the-art on reasoning benchmarks and still fabricate unsupported claims in more than one summary out of ten; the two facts are measuring different things, and only one of them shows up on the leaderboard most buyers look at first.

5. Sycophancy: models learn to agree rather than to correct. Mrinank Sharma and colleagues at Anthropic tested five state-of-the-art AI assistants across four free-form generation tasks and found consistent sycophancy in all of them — a tendency to shift stated positions toward whatever the user seemed to believe, even away from a more accurate answer — and traced the likely cause to how human preference data is collected: raters tend to prefer confident, agreeable answers, and preference models trained on that data inherit the bias [7]. That is an Anthropic study describing a failure mode general enough to appear across the five systems it tested, not one specific to Anthropic’s own models.

A second provider produced a much more public instance of the same failure. On April 25, 2025, OpenAI shipped a GPT-4o update that became, in its own words, “noticeably more sycophantic,” and rolled it back four days later after the model was found endorsing users’ harmful and delusional statements rather than pushing back on them. OpenAI’s own postmortem attributed the cause to a narrow optimization target: “we focused too much on short-term feedback, and did not fully account for how users’ interactions with ChatGPT evolve over time,” producing responses that were “overly supportive but disingenuous” [8]. Read against Sharma and colleagues’ finding, the incident looks less like an isolated bug in one update and more like a sharper version of a bias already documented in a different lab’s models, produced by the same underlying mechanism: optimizing against short-term human approval.

ADVERTISEMENT
A stand loupe positioned over two printed transcripts of different lengths curling out from twin rigs fed by the same coloured cable
Figure 3. Two cards differing only in phrasing enter rigs patched to the same cable; what comes out does not have to match, and the loupe is where somebody first notices.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

6. The same question, reworded, can get a different answer. If a model’s accuracy depends on incidental features of how a question happens to be phrased, then the score reported for “the model” is really the score for one particular phrasing, and other equally valid phrasings may score very differently. Melanie Sclar and colleagues quantified exactly this for few-shot prompt formatting: holding the semantic content of a prompt fixed and varying only surface formatting — spacing, separators, casing — produced accuracy swings of up to 76 percentage points on LLaMA-2-13B, and the sensitivity did not shrink with model scale or with instruction tuning [11]. Their recommendation follows directly from the finding: report a range of scores across plausible formats rather than a single number, because a single number from a single format is closer to a sample from a distribution than a stable property of the model.

7. A context window marketed at a given length is not the same as a context window usable at that length. Providers advertise maximum context length as a single number; how much of that length a model can actually use reliably is a separate, empirical question, and the two numbers diverge substantially. NVIDIA’s RULER benchmark tested seventeen long-context models, including GPT-4 and Gemini-1.5-Pro among others, on tasks more demanding than simple needle-in-a-haystack retrieval, and found that only about half of models claiming support for 32,000 tokens could sustain satisfactory performance at that length — with some models advertised for far longer contexts, up to 200,000 tokens, showing effective usable performance closer to 32,000 in practice [15]. Near-perfect scores on the simplest retrieval test — find one fact planted in a long document — did not predict performance on RULER’s harder synthetic tasks at the same length, which is the central methodological point: a single easy retrieval test is a poor proxy for what “usable context” means, and the marketed maximum is not evidence about either.

A long paper tape reel feeding into a rig with a small flag marker clipped partway down its length, well ahead of the read head
Figure 4. A reel marketed at its full length is not the same as a reel whose read head actually reaches every point on it; the flag marks where most rigs stop noticing.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

8. Performance degrades specifically across multi-turn conversation, not just on harder single-shot questions. Most public benchmarks present a task in one shot: one prompt, one answer, scored once. Real use is rarely structured that way. Philippe Laban and colleagues, working across Microsoft Research and Salesforce, restructured six standard generation tasks into simulated multi-turn conversations — spreading the same information a user would normally provide all at once across several turns instead — and ran more than 200,000 simulated conversations against fifteen models spanning OpenAI (GPT-4o, GPT-4o-mini, GPT-4.1, o3), Anthropic (Claude 3 Haiku, Claude 3.7 Sonnet), Google (Gemini 2.5 Flash and Pro), Meta (Llama 3.1, Llama 3.3, Llama 4 Scout), and others. Every model tested performed substantially worse in the multi-turn version of the same task, with an average drop of 39% across the six tasks. The mechanism the authors identify is specific: models tend to commit to an assumption early in the conversation and prematurely attempt a final answer, and once that early assumption is wrong, most models do not recover within the same conversation [16]. Because the underlying task and the underlying information are identical between the single-turn and multi-turn conditions, this drop cannot be explained by task difficulty; it is a property of how the conversation is structured, present across every provider tested.

What training itself trades away

9. Safety and alignment tuning has a measurable cost to other capabilities, on more than one provider’s own account. OpenAI’s InstructGPT paper is where the term “alignment tax” enters this literature directly: the authors write that naive fine-tuning toward human preferences degraded performance on several public NLP benchmarks, and label that degradation “an alignment tax — an additional cost for aligning the model.” Their mitigation, mixing pretraining gradients back into the reinforcement-learning objective (a method they call PPO-ptx), recovered most of the regression and even exceeded the original GPT-3’s score on one benchmark, HellaSwag [12]. That is an existence proof that the tax can be partly engineered away, not evidence that it disappears by default.

A second provider documents the same trade-off, independently, in its own model. Meta’s Llama 2 paper explicitly cites prior work from Anthropic (Bai and colleagues, 2022) on the tendency of helpfulness and safety objectives to trade off against each other, and describes training separate reward models for each rather than relying on one model to balance both, because a single shared reward model struggled to perform well on both dimensions at once. The paper also reports measuring “false refusal” directly as a safety-tuning error mode — cases where safety training causes the model to decline entirely legitimate, benign requests — and frames a large part of its own safety-tuning effort as reducing false refusals without reopening genuine safety gaps [13]. Between the two papers, the same underlying trade-off is documented from three different labs’ perspectives: OpenAI naming and measuring the tax directly, Meta measuring it as false refusals and citing Anthropic’s account of the same tension. None of the three claims the trade-off has been eliminated; each describes a different partial mitigation.

A restrictor module being fitted mid-installation to one rig's intake throat, its screws not yet driven home, beside an identical unmodified rig running unobstructed
Figure 5. A module fitted to make a rig safer is not fitted for free; the identical rig beside it, still unmodified, is the only way to see what the fitting cost.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

10. Behavior drifts across versions released under a name that does not visibly change. A model identifier can remain stable in an API or a product while the weights or serving configuration behind it change, and the resulting behavior shift is invisible unless someone re-measures on a schedule rather than trusting the name. Lingjiao Chen, Matei Zaharia, and James Zou tracked GPT-3.5 and GPT-4 on identical tasks at monthly intervals from March to June 2023 and found substantial, sometimes opposite-direction, movement: GPT-4’s accuracy on a prime-number identification task fell from 84% to 51% over the period, while GPT-3.5 improved on the same task over the same window; both models showed an increase in code-formatting errors; and GPT-4 became measurably less willing to answer sensitive or opinion-based questions and less responsive to chain-of-thought prompting by June than it had been in March. The authors’ own summary is blunt: “GPT-4’s ability to follow user instructions has decreased over time” was one contributing mechanism behind several of the shifts they measured [14]. Nothing about the product name, endpoint, or documentation necessarily signaled any of this to a downstream user running the same prompt every week; the only way to see it was to keep a private, dated evaluation set and re-run it, rather than to assume last month’s characterization of “GPT-4” still described this month’s system.

What this means for anyone comparing models

None of the ten items above is a reason to distrust benchmarks entirely, and none of them supports ranking OpenAI, Anthropic, Google, or Meta’s systems against each other on the strength of what is documented here — the sources cited measure different things, at different times, under different conditions, and stacking them into a single ordinal ranking would misuse every one of them. What they do support is a small set of concrete practices. Treat a single benchmark score as a lower-confidence estimate of a range, not a point value, given how far the same model can move under reformatting alone. Re-run a private, dated evaluation set on a schedule rather than trusting that a stable model name implies a stable system underneath it. Separate a vendor’s self-reported finding from an independent replication explicitly, in your own notes if nowhere else, because the two carry different evidentiary weight even when they agree. And treat “beats the benchmark” and “does the job reliably in a long, open-ended, multi-turn, real task” as two different claims requiring two different kinds of evidence, because the studies above show they frequently diverge.

Two forward-looking claims follow from the pattern, stated as predictions with a horizon rather than as settled fact. First, over the eighteen months from this writing to early 2028, expect at least one more widely used public benchmark to be publicly retired or substantially revised by the lab that created it, on contamination grounds, following the precedent set by OpenAI’s SWE-bench Verified retirement; this would be disconfirmed if 2028 arrives with the current generation of headline benchmarks (SWE-bench Pro, GPQA, MMLU-successor suites) still in active, unrevised use by multiple labs’ official system cards. Second, expect published multi-turn and long-horizon evaluations — in the style of METR’s task-horizon metric and Laban and colleagues’ multi-turn degradation study — to become a standard part of system-card disclosures alongside single-shot accuracy figures, precisely because the gap between the two has now been documented across enough providers that a system card omitting it will look conspicuously incomplete; this would be disconfirmed if, by 2028, major system cards still report only single-shot benchmark suites with no multi-turn or extended-task disclosure at all.

The pattern underneath all ten is the same one running through the equation near the top of this piece: a single published number is a sum of several things that move independently, and no amount of competition between vendors makes that sum easier to decompose from outside. What decomposes it is documentation — a vendor’s own methodology section, an independent lab’s replication, a re-run against a private set — attributed to whoever actually produced it. That discipline does not tell you which model is best. It tells you what a given score is actually evidence of, which is the more useful and much rarer thing.