A number is not evidence until someone has rerun it
xAI has published a striking claim about a Grok model roughly every few months since early 2025, and outside of xAI, the pattern of what happens next is now well documented enough to work through methodically. This article does that: it takes three specific, dated rounds of claim and outside scrutiny — the Grok 3 launch benchmark chart, the Artificial Analysis Intelligence Index tracking across five Grok generations, and the ARC Prize Foundation’s testing and verification program — and uses them to build an explicit method for weighing any future xAI capability claim against the evidence that actually exists for it.
The scope here is deliberately narrow. This is not an assessment of whether Grok is a good or bad model, and it is not a cross-vendor ranking; incomparable benchmarks cannot support one, and where independent testers disagree about a comparison, the honest move is to describe the disagreement rather than resolve it. It is also not a piece about Colossus, xAI’s compute cluster in Memphis, Tennessee. Colossus is real infrastructure with its own separately documented facts — xAI describes it as a 200,000-GPU cluster used to train Grok 4 with an order of magnitude more reinforcement-learning compute than the company had previously applied [2] — and datacenter scale is not the same claim as model capability. A cluster being large is a fact about hardware procurement. A model scoring well on a named benchmark is a separate, differently-verified claim, checked through different channels, and this article stays on the second kind.
What an xAI capability claim actually consists of
A published score is never just a number. It is a specific checkpoint, queried under a specific method, scored against a specific set, on a specific date — and a claim is only as strong as the parts of that chain a reader can actually check. xAI’s own launch materials are unusually explicit about naming the pieces, which is precisely what makes it possible to audit them.
Take the Grok 4 announcement’s headline claim on Humanity’s Last Exam, a roughly 2,500-question expert benchmark: xAI reports 50.7% on the text-only subset, but only for a configuration called Grok 4 Heavy that uses “parallel test-time compute, which allows Grok to consider multiple hypotheses” before answering, and the announcement separately reports a lower figure for tool-free Grok 4 answering alone [2]. Those are not the same claim. A model that reaches 50.7% by reconciling several parallel attempts is a different computational product, at a different cost per answer, than a model that reaches a lower number on one attempt with no tools. xAI discloses which configuration produced which number; a reader who drops the qualifier and repeats “Grok 4 scored 50.7% on Humanity’s Last Exam” has already discarded the information needed to compare it to anything else.
The same announcement reports Grok 4 setting “state-of-the-art for closed models on ARC-AGI V2 with 15.9%,” a figure it describes as nearly double the roughly 8.6% then held by a leading closed competitor [2]. That specific number is worth holding onto, because it reappears independently later in this piece — and because ARC-AGI-2 is maintained by an outside organization with its own testing process, it is one of the few claims in this article that can be checked against a party other than xAI.
The Grok 3 dispute: a documented case of contested method
The clearest documented case of an xAI benchmark claim being publicly challenged is the Grok 3 launch, on 19 February 2025. xAI’s own release post reports Grok 3 in “Think” mode reaching 93.3% on AIME 2025, a competition mathematics benchmark, alongside strong scores on GPQA and LiveCodeBench, and describes training on the Colossus cluster with, in the company’s words, ten times the compute of its prior state-of-the-art model [1].
Within days, engineers at OpenAI publicly disputed the comparison chart built around that AIME figure. According to TechCrunch’s reporting, the dispute was specifically about method: xAI’s chart displayed Grok 3’s AIME score against a rival model’s score at “@1” — a single attempt per problem — while the 93.3% figure attributed to Grok 3 itself was reported at “consensus@64,” meaning the model was given 64 attempts per problem and the most frequent final answer across those 64 attempts was scored as the result [4]. Those are different measurements of different things. A single-attempt score describes what a user talking to the model once should expect; a 64-sample consensus score describes what the model’s most common answer looks like after a considerably larger compute spend than a normal chat turn, and consensus sampling raises the reported figure without proportionally telling a reader what any one response would have looked like. TechCrunch reported that when the rival model’s own consensus-sampled score was used for the comparison instead of its single-attempt score, it came out ahead of the number xAI had charted [4].
xAI co-founder Igor Babuschkin responded on the social platform X that OpenAI had itself published comparison charts favorable to its own models in the past, an argument about precedent rather than a rebuttal of the specific method used [4]. TechCrunch’s own conclusion, after examining both sides, was that the truth sat “somewhere in between”: xAI had not invented a number, but the chart’s framing was selectively favorable, and the deeper problem — that most public benchmark charts, from every vendor, omit the computational cost behind each bar — was bigger than any one company’s release [4]. That is analysis this article adopts rather than a verdict it is in a position to independently confirm; it is presented here as a specific journalistic finding, attributed to its source, not as this publication’s own re-derivation of the underlying scores.
The general lesson generalizes past this one chart. Stella Biderman and thirty co-authors, writing from three years of building the open-source Language Model Evaluation Harness, document the same failure mode across the field: language models are highly sensitive to prompt format and sampling configuration, comparisons across papers and press releases routinely differ in exactly these hidden variables, and reproducibility suffers because the configuration is so often left out of the headline number [12]. The Grok 3 chart is one dated instance of a documented, general problem, not an isolated one.
Where an outside record lines up with the claim
Contested cases are not the only kind. ARC-AGI-2, the benchmark behind Grok 4’s 15.9% claim, is maintained by the nonprofit ARC Prize Foundation, and its scoring is not purely self-reported: the foundation states that results on the ARC-AGI-1 and ARC-AGI-2 semi-private evaluation sets — data withheld from public release specifically so a model cannot have trained on the exact questions — “are run and verified by ARC Prize,” while broader claims on other benchmarks are separately flagged as “scored on a public set and self-reported,” with the foundation explicit that it cannot vouch for the authenticity of that second category [7]. The distinction matters because it tells a reader which of a vendor’s numbers passed through an outside gate and which did not. Grok 4’s 15.9% ARC-AGI-2 figure falls into the first, checked category, and independent technical coverage published in the weeks after launch reported the same figure without contradiction [2].
An earlier ARC Prize post is worth reading alongside that number, because it shows the foundation testing xAI’s models directly rather than only cataloguing xAI’s own claims. In a June 2025 post titled “We tested every major AI reasoning system. There is no clear winner,” the foundation ran its own open-source testing harness against Grok 3 and Grok 3 Mini on ARC-AGI-1 and reported Grok 3 Mini (low) at 16.5% accuracy for roughly one cent per task, against plain Grok 3 at 5.5% accuracy for a comparable cost — a result the post used to argue that a smaller, cheaper configuration was the more useful one for that task, and to state plainly that “naked accuracy scores are marketing, not science” without a cost figure attached [5]. That framing is now explicit in this article’s own method section below: a capability number without its cost, sampling method, and verification status is an incomplete claim regardless of which vendor issued it.
From one-off checks to standing verification
A single blog post confirming a number is weaker evidence than a standing institutional process built to keep confirming or disconfirming future numbers, and that second kind of infrastructure now exists for at least one Grok-relevant benchmark. In November 2025, the ARC Prize Foundation announced “ARC Prize Verified,” a certification program the foundation built specifically because, in its own words, “self-reported or third-party figures often vary in dataset curation, prompting methods, and many other factors, which prevents an apples-to-apples comparison of results” [6]. The program formalizes hidden-set testing, adds an outside academic panel — drawn, per the announcement, from NYU, UCLA, Columbia, and the Santa Fe Institute — to audit the testing process itself, and issues an official leaderboard badge only to scores that pass through it [6]. xAI is named in that same announcement as one of five sponsoring AI labs [6], which is a notable fact in its own right: it is xAI funding, at least in part, the infrastructure that can contradict xAI’s own future marketing claims, not merely a case of an outside party checking a rival’s homework.
This is also where a composite measure earns its own paragraph, because a composite index is not a repeat measurement of a single-benchmark headline claim — it is a different object. Artificial Analysis, an independent evaluation firm that runs every model it tracks through one fixed harness, reports an “Intelligence Index” built from nine separate evaluations, including Humanity’s Last Exam and GPQA Diamond among others [8]. A composite index of this kind is, formally, a weighted sum
where each
When two independent measures disagree with each other
The harder methodological case is not vendor versus outside evidence — it is outside evidence versus other outside evidence. LMArena, the human-preference leaderboard that grew out of UC Berkeley’s Chatbot Arena project and rebranded to Arena in 2026, ranks models by aggregating blind pairwise votes from real users into an Elo-style score, with results bracketed by a statistical confidence interval [13]. On that leaderboard, Grok 4.1 Thinking reached the top overall position at an Elo of 1483 as of the source’s reporting date, described as roughly 31 points ahead of the next non-xAI model, a sharp rise from Grok 4’s earlier rank of 33rd on the same board [13].
Read next to Artificial Analysis’s Intelligence Index — which places Grok models well behind the frontier on several task-accuracy components even as later Grok generations climb [8, 9] — the two independent measures do not tell a consistent story about where Grok models rank, because they are not measuring the same thing. A pairwise human-preference vote captures which response a person reading two answers side by side liked better: tone, formatting, apparent confidence, and perceived helpfulness all load onto that vote alongside correctness. A task-accuracy composite scores whether an answer was actually right against a held answer key. The source covering Grok 4.1’s Arena result makes exactly this point explicit, cautioning that leaderboard dominance on human preference does not establish superior reliability under scrutiny, and that a model can, in the piece’s own words, be “better at sounding like it is” trustworthy without that translating into a task-accuracy edge [13]. Neither leaderboard is wrong. A reader who wants to know whether Grok 4.1 will draft a pleasant-sounding email and a reader who wants to know whether it will get a graduate-level physics problem right are asking two different questions, and each of the two independent leaderboards above answers one of them.
What is confirmed about parameter counts, and what is not
Now the part this article is under the strictest instruction to get right: how large a given Grok model actually is. Exactly one figure in this lineage is confirmed in the strong sense — tied to something a third party can independently inspect rather than to a reported remark. In March 2024, xAI open-sourced the weights of Grok-1 itself, describing it in its own release post as a 314-billion-parameter mixture-of-experts model with roughly a quarter of those weights active on any given token [3]. Because the weights themselves were published under an open license, that parameter count is not a claim resting on trust in xAI’s disclosure — anyone with the hardware to load the checkpoint can verify it directly, which is the strongest form of confirmation a parameter count can have.
No later Grok model has been accompanied by a comparable technical disclosure. The figures that circulate for Grok 3, Grok 4, and Grok 5 — most commonly “3 trillion” for Grok 3 and Grok 4, and “6 trillion” for Grok 5 — do not originate in a system card, a research paper, or released weights. The independent AI model tracker LifeArchitect.ai, maintained by researcher Alan D. Thompson, traces the 6-trillion Grok 5 figure specifically to a reported remark by xAI’s chief executive in a conversation with an investor in November 2025, in which he is quoted describing Grok 5 as a 6-trillion-parameter model against 3 trillion for Grok 3 and Grok 4 [11]. That is a real, attributed statement from xAI’s own leadership, and this article does not dispute that it was said. But it is categorically different from a technical disclosure: it carries no accompanying architecture description, no active-parameter figure for what is very likely a mixture-of-experts design, and no way for an outside party to check it the way Grok-1’s open weights could be checked. This article treats it, and every other trillion-scale figure attached to a closed Grok model, as a reported claim rather than a confirmed one, and states no specific parameter count for Grok 3, Grok 4, Grok 4.5, Grok 4.6, or Grok 5 as settled fact.
The existence of a specialized discipline for estimating exactly these unstated figures is itself evidence of how unresolved they are. Epoch AI, an independent research organization, maintains a documented methodology for estimating undisclosed parameter counts and training compute from whatever architectural or hardware detail is available, and explicitly grades every estimate by confidence: “confident” for an estimate accurate within roughly a factor of three, “likely” within roughly a factor of ten, and “speculative” within roughly a factor of thirty [10]. That a serious research organization needs a three-tier uncertainty scale, with the widest tier spanning a thirty-fold range, to talk about frontier model sizes at all is a direct measure of how little the field’s leading labs — xAI very much included — actually disclose. A number produced by this kind of estimation is useful and worth citing as an estimate; it is not equivalent to a vendor-confirmed figure, and this article has not used any such estimate for a current Grok model as a stand-in for a confirmed count.
A method for weighing the next xAI claim
Every case above resolves into the same short checklist, and it is offered here as a method rather than a scorecard for any single number.
Identify the exact configuration, not the model name. “Grok 4” is not one measurement; Grok 4, tool-free Grok 4, and Grok 4 Heavy produced different numbers in xAI’s own materials [2]. A claim that drops the configuration has already lost the information needed to compare it to anything.
Identify the sampling method behind the number. A single-attempt score and a consensus-of-many-attempts score answer different questions at different costs, and the Grok 3 dispute shows how much a chart can imply by leaving that distinction off the page [4].
Check whether an outside party reran it, or only repeated it. ARC Prize’s semi-private-set testing is a rerun; most benchmark citations in press coverage are a repetition of the vendor’s own number [7]. Only the first counts as independent confirmation.
If the figure is a composite index, check whose weights built it. An index score is not a repeat measurement of a headline claim; it is a different construct, and disagreement in scale between an index and a vendor’s own number is expected, not suspicious [8].
Expect independent measures to disagree with each other, and ask what each one actually measures before treating disagreement as evidence of error. Arena’s human-preference Elo and Artificial Analysis’s task-accuracy index rank Grok’s recent generations differently because a preference vote and a graded answer key are not the same instrument [13, 8].
Treat any parameter or compute figure for a closed model as unconfirmed unless it is tied to released weights or a technical paper. Grok-1’s 314 billion parameters are confirmed because the weights were published [3]. Every later Grok figure currently in circulation is a reported remark or an outside estimate, explicitly labeled with an uncertainty range by the researchers who produce it [11, 10].
Keep model claims and datacenter claims in separate columns. Colossus’s GPU count is a hardware procurement fact, disclosed on xAI’s own terms; a Grok model’s benchmark score is a claim about a trained system’s behavior, checked, when it is checked at all, by a different party through a different process [2].
Predictions, and what would falsify them
These are forecasts, kept separate from the documented cases above. Horizon: 12 August 2028.
One. More vendors will follow xAI’s Grok 4.6 pattern of citing an independent index’s own construct directly in launch marketing when it is favorable, rather than publishing only an in-house comparison chart. Disconfirmed if major frontier-lab launch posts in 2028 have reverted to citing only self-produced benchmark charts with no reference to an outside index.
Two. At least one additional benchmark maintainer will introduce a formal verification tier comparable to ARC Prize Verified — hidden test data plus an outside academic audit panel — with a major lab, plausibly including xAI, listed as a sponsor. Disconfirmed if, by the horizon date, ARC Prize Verified remains the only program of this kind covering a frontier reasoning benchmark.
Three. No current-generation closed Grok model’s exact parameter count will be confirmed by a primary xAI technical disclosure by the horizon date, unless xAI again open-sources weights as it did once with Grok-1. Disconfirmed if xAI publishes a system card, paper, or equivalent technical document stating an explicit parameter count for any Grok 4-series or Grok 5-series model.
What the record actually supports
Strip away the marketing framing on both sides of every dispute above and a narrower, more defensible picture remains. Grok 4’s ARC-AGI-2 figure sits inside a category ARC Prize itself verifies, and independent coverage repeated it without contradiction. Grok 3’s launch-day AIME comparison rested on a sampling-method mismatch that TechCrunch’s own reporting called selectively framed rather than fabricated. Artificial Analysis’s composite index tracks real movement across Grok’s generations on a scale that was never going to match xAI’s own headline numbers, because it is measuring a different, independently weighted construct. Arena’s human-preference leaderboard and that same composite index rank recent Grok models differently, and both are right about what they each measure. And exactly one parameter count in this entire lineage — Grok-1’s 314 billion, confirmed by published weights — is settled; everything larger currently circulating for Grok 3 through Grok 5 is a reported remark or a labeled estimate, not a disclosure, and belongs in the tray of clippings rather than the folder of confirmed record until xAI closes that gap itself.