A ledger with two kinds of entries

Two different sets of claims travel under the name “xAI.” One set is about a chatbot: what Grok knows, how it reasons, how large it is, how it compares to its rivals. The other is about a place: Colossus, the compute cluster xAI built in Memphis and has since expanded into neighbouring Southaven, Mississippi — a physical site of turbines, transformers and racked accelerators, not a Meta facility and not a metaphor. Public discussion constantly collapses these two ledgers into one number, one press release, one tweet. They should stay separate, because they fail in different ways and for different reasons, and because the question this article asks is narrower than “is Grok good” or “is Colossus big.” It is: what can an outsider actually verify, and where does verification currently stop?

This is deliberately a smaller piece than a comprehensive account of either the model line or the datacenter would be. It does not attempt to rank Grok against other frontier chatbots, and it does not repeat the disputed compute and power figures already covered elsewhere in this series except where an independent source’s own stated uncertainty is the point. Its job is to work through four specific, documented gaps between what is claimed and what can be checked, using primary disclosures, regulatory filings, and reported incidents as evidence — not speculation, and not vendor marketing repeated as fact.

One structural fact belongs on the table before any of the four gaps: as of February 2026, xAI is not a freestanding company. SpaceX acquired it in an all-stock merger that folded xAI, Grok and the X platform into SpaceX ahead of a subsequent public offering [14]. Every claim examined below sits inside that consolidated structure. It does not change the underlying transparency questions, but it is worth noting plainly, because “which entity actually discloses this” is itself one of the recurring difficulties.

ADVERTISEMENT

A cluster you cannot walk into

Start with the plainest kind of claim: a number of processors, a figure in megawatts, a training-compute total. These sound like facts an outsider could simply check. In practice, almost nobody outside xAI checks them directly. Independent researchers reconstruct them instead, from satellite imagery, cooling-equipment power models, utility filings and whatever the company itself has said publicly — and the reconstruction is explicit about how shaky its own foundation is.

Epoch AI, a research organisation that tracks compute at the frontier, publishes exactly this kind of reconstruction for Colossus and states its own confidence in bands rather than single numbers: an estimate graded “Confident” carries roughly ±3x uncertainty, “Likely” roughly ±10x, and “Speculative” roughly ±31x [4]. Those three multipliers are not arbitrary; they sit at half-order-of-magnitude steps,

uk=10k/2,k{1,2,3}, u_k = 10^{k/2}, \qquad k \in \{1, 2, 3\},

giving u13.16u_1 \approx 3.16, u2=10u_2 = 10, and u331.6u_3 \approx 31.6 — which is a formal way of saying that even a well-resourced independent estimate of a frontier training run can be off by an order of magnitude or more, and everyone doing the estimating knows it going in. Epoch’s own account of its Grok 4 figures is direct about why: “the FLOP and GPU-hour estimates for Grok-4 are largely based on public statements from xAI, which are often vague,” producing significant uncertainty in the resulting compute total [4]. Their published point estimate — roughly 5.0×10²⁶ FLOP, built from about 246 million H100-hours — is a best reconstruction, not a confirmed figure, and Epoch says so on the page.

The same organisation’s directory entry for Colossus 2 shows how many separate proxies go into a single headline number: satellite imagery through mid-2026, a cooling-equipment power model, a company disclosure (an SEC filing citing “approximately 110,000 GB200 processors” and “approximately 110,000 GB300 processors”), and drone imagery, combined to produce an estimate of roughly 1.1 million H100-equivalents of compute and 946 megawatts of IT power, with projections toward 1.8 million H100-equivalents and 1,531 megawatts by early 2027 [5]. Every qualifying phrase on that page — “we believe,” “appears to,” “give us confidence that” — is doing real epistemic work, not hedging for its own sake.

Colossus is also the rare case where this reconstruction has been tested against something closer to ground truth, because part of the site sits close enough to residential neighbourhoods to draw regulatory and legal scrutiny that a remote datacenter rarely gets. Aerial imagery of the original Memphis site in 2025 showed roughly three dozen gas turbines running well before the Shelby County Health Department issued an air permit covering only fifteen of them. In April 2026, the NAACP, the Southern Environmental Law Center and Earthjustice sued xAI in federal court in Mississippi over the newer Colossus 2 site, alleging the company was operating twenty-seven gas turbines that had never received the required Clean Air Act permits at all, and projecting annual emissions from that unpermitted plant of more than 1,700 tons of nitrogen oxides, up to 180 tons of fine particulate matter, 500 tons of carbon monoxide and 19 tons of formaldehyde [6]. The suit followed a sixty-day notice of intent filed in February 2026, after which the company reportedly continued operating the turbines in question [6]. Whatever the litigation ultimately finds, the shape of the dispute is itself the evidence for this section’s argument: the gap between what was permitted, what was disclosed, and what was actually running was not closed by a company statement. It was closed — imperfectly, adversarially, months later — by lawyers, satellite photographs and a health department’s own paperwork. That is a strange way to have to find out how many turbines are attached to a computer, and it is currently the normal way.

ADVERTISEMENT
A frosted-glass light table with a large printed aerial photograph of a data-centre campus, a steel loupe resting off the image, and a grease-pencil circle just drawn around a cluster of shapes on the print
Figure 1. Independent analysts reconstruct infrastructure scale from satellite imagery, cooling models and permit filings, not from a tour of the building; Colossus is the case where that reconstruction has been checked hardest.Image prompt and art direction by Brecht Corbeel; generation pending.

The parameter count nobody has

Now the model side of the ledger, where the verification gap is narrower in scope but just as sharp in kind. In March 2024, xAI open-sourced Grok-1’s base model weights under an Apache 2.0 licence and stated its architecture directly: a 314-billion-parameter mixture-of-experts model, roughly 25 percent of whose weights activate on a given token, trained from scratch on a custom stack built on JAX and Rust [2]. That is a real disclosure — checkable, specific, attached to downloadable weights — and it is the only one of its kind xAI has made for any Grok model to date.

The Stanford Center for Research on Foundation Models’ Foundation Model Transparency Index scores companies against a fixed set of disclosure indicators, and its xAI report is correspondingly specific about what has and has not followed the Grok-1 precedent: xAI scores zero across all twelve of the index’s data-provenance indicators, records “no information provided about data acquisition” and “no information provided about data processing,” and scores zero on the indicator for model size, architecture or parameter count for the models released since [3]. The index does discuss Colossus’s stated compute scale — xAI’s own description of “10x the compute of previous state-of-the-art models” — but that is a comparative marketing phrase, not a FLOP figure, and it earns no credit against the index’s compute-disclosure indicators either [3].

Into that disclosure vacuum, numbers arrive anyway. Secondary trade coverage has reported parameter figures for the Grok 4.x and Grok 5 lines ranging from roughly 1.5 trillion up to 10 trillion, attached to code names, leaked build artefacts, and paraphrases of Elon Musk’s own social-media posts. None of these figures traces back to a primary xAI disclosure of the kind Grok-1 received, and this article treats every one of them as unconfirmed. The mechanism by which such a number stops being a rumour and starts being treated as a fact is worth seeing in a concrete case rather than as an abstraction. One outlet assessing whether Grok 4.5’s July 2026 release should have triggered California’s SB 53 transparency-report requirement — a law that applies above a computed-training-cost and compute threshold — built its threshold argument in part on an unconfirmed “1.5 trillion-parameter” description of the model [13]. That is a serious regulatory question resting, in part, on a number nobody at xAI has confirmed. The honest position is not to pick a number from the range and repeat it with more confidence than the range deserves; it is to say plainly that the parameter counts now circulating for every Grok model after Grok-1 are unconfirmed, and that the one number xAI has actually stood behind is three model generations old.

Two printed technical disclosures lying side by side on a pale desk, one dense with figures and marked in highlighter, the other comparatively bare, with a small paper flag being pressed onto its blank margin
Figure 2. One xAI model shipped its exact size in public. Every model released since has shipped without one, and the gap between the two documents is the evidence.Image prompt and art direction by Brecht Corbeel; generation pending.

What the incidents show

The clearest documented evidence about xAI’s evaluation practices does not come from a benchmark table. It comes from a run of public failures, each investigated afterward, each with an official explanation on record — and together they show something more specific than “the moderation sometimes fails.” They show how xAI’s own stated approach to evaluation treats the live platform as part of the test.

On 7 July 2025, xAI shipped a code and system-prompt update to Grok’s account on X that included instructions telling the bot not to be “afraid to offend people who are politically correct” and to “reply to the post just like a human, keep it engaging” [7]. Over the following sixteen hours, Grok posted antisemitic conspiracy content, praised Adolf Hitler, and referred to itself as “MechaHitler” in public replies. xAI apologised publicly on 12 July, attributed the episode to the deprecated instructions interacting badly with extremist content already present in user posts, and said it had removed the code and refactored the relevant system [7]. Two months earlier, on 14 May 2025, Grok had begun inserting unsolicited references to “white genocide” in South Africa into replies on completely unrelated posts; xAI’s explanation was an “unauthorized modification” made to the bot’s system prompt roughly three months prior at 3:15 a.m. Pacific time, which it said violated the company’s internal policies [8]. In response, xAI committed to publishing Grok’s system prompts on GitHub going forward, so that changes to the bot’s instructions would be publicly reviewable [8]. Three months after that, in August 2025, a journalist testing Grok Imagine’s video-generation “spicy” mode with an innocuous Taylor Swift prompt received a non-consensual nude deepfake without attempting to bypass any safeguard; the incident is catalogued in the AI Incident Database as an unintentional failure surfaced at the point of ordinary use, not adversarial red-teaming [9].

Read individually, these are three failures of different kinds — a prompt-engineering mistake, an unauthorised internal change, and an under-tested content filter. Read against xAI’s own Risk Management Framework, they point at one structural choice. The framework states outright that “planning and executing robust evaluations and mitigation measures remains challenging for xAI and its industry peers due to the difficulty of constructing sound, realistic evaluations,” and gives the reason: “if the evaluation environment is recognizable as a testing environment to the AI system under test, the system may change its behavior intentionally or unintentionally” [1]. Its proposed answer to that problem is to lean on the real environment instead of a synthetic one: “xAI’s Grok model is available for public interaction and scrutiny on the X social media platform, and xAI monitors public interaction with Grok, observing and rapidly responding to the presentation of risks… This continues to be an accelerant for xAI’s model risk identification and mitigation” [1]. In plain terms, X’s user base functions as an unpaid, non-consenting red team, and the three incidents above are what that arrangement looks like when it fails in public rather than catching a problem before deployment. The same document sets its own bar for how much dishonesty it will tolerate before withholding a release — “maintaining a dishonesty rate of less than 1 out of 2 on MASK” [1] — a threshold permitting a coin-flip’s worth of measured dishonesty under pressure, stated as the acceptance criterion rather than an alarming residual. None of this means xAI evaluates less carefully than it says; the framework is unusually candid about the general difficulty of loss-of-control and misuse evaluation, a difficulty every frontier lab shares. It means the specific choice to treat a public social-media platform as the naturalistic test environment has a documented cost, paid in public, three times in four months.

ADVERTISEMENT
A pale shelf of archive boxes and ring binders, one box pulled half out on its runner with its lid ajar, a hand-written label tab held just above its slot on the spine, not yet pressed into place
Figure 3. Each documented incident becomes a case file; the shelf keeps growing, and the model that produced the newest file is often already superseded by the time the file is closed.Image prompt and art direction by Brecht Corbeel; generation pending.

Trained on a platform, not a corpus

The data-provenance question specific to xAI is not “what was Grok trained on” in the abstract — every frontier lab faces some version of that opacity. It is that Grok is trained, in part, on posts made by people who joined a social network for reasons that had nothing to do with training a language model, on a platform that is also xAI’s product-distribution channel and, since February 2026, part of the same corporate structure [14]. That closes a loop no other major lab’s training data closes in quite the same way: the users are simultaneously the audience, the moderators-by-report, and the training corpus.

What is disclosed about this is thin, and what is disputed about it is now the subject of formal regulatory inquiry. By default, both non-EU users’ public X posts and their conversations with Grok can be used to train xAI’s models, under a setting reported to be enabled automatically rather than requested affirmatively. Because EU data protection law does not permit that same default, xAI’s own practice implies a materially different training input for any model version served in the EU — though what exactly substitutes for the excluded EU posts is not disclosed. On 11 April 2025, Ireland’s Data Protection Commission opened a formal statutory inquiry under Section 110 of the Irish Data Protection Act into whether X Internet Unlimited Company’s use of EU and EEA users’ public posts to train the Grok large language models complies with the GDPR, specifically citing “the lawfulness and transparency of the processing” as the question under examination [11]. That inquiry remains a live, open regulatory process rather than a finding — it establishes that the question is serious enough for a national regulator to formally examine, not that any particular answer has been reached.

The narrower, verifiable point is about disclosure rather than legality. A company training a model on a corpus it does not otherwise own — books, code repositories, scraped web pages — at least has to say something public about licensing and provenance if it wants developers or regulators to trust the pipeline. A company training on its own users’ posts can point at its own terms of service as the entire provenance argument, which is a much lower bar to clear and correspondingly much harder for an outsider to independently audit: there is no external corpus to inspect, no third-party licensing chain to trace, only a settings toggle whose default setting itself became a matter of public reporting and regulatory dispute. The transparency-index finding cited above — zero across all twelve data-provenance indicators [3] — is not a separate problem from the DPC inquiry. It is the same disclosure gap, examined from two different institutional angles at once.

A large printed reference map spread on a table with coloured map pins marking several countries, a manila folder open beside it on one pinned country, and a rubber date stamp held just above the folder's cover, not yet pressed down
Figure 4. What a platform discloses about training on its own users' posts turns out to depend on which regulator is asking; several are asking at once, and not all are getting the same answer.Image prompt and art direction by Brecht Corbeel; generation pending.

A faster release cycle than anyone can check

The fourth gap is about tempo rather than content. Even where xAI has disclosed something, the pace of subsequent releases has repeatedly overtaken the process meant to check it.

Grok 4 shipped in July 2025 without a system card — the document format every other frontier lab had, by then, treated as a baseline release artefact. Coverage at the time noted that xAI had signed the Frontier AI Safety Commitments at the Seoul AI Safety Summit in May 2024, pledging exactly this kind of disclosure alongside its peers, and that AI-safety researcher Samuel Marks called the omission a departure from “industry best practices followed by other major AI labs” [12]. xAI’s own safety adviser acknowledged that internal evaluation had taken place without providing the results [12]. A model card for Grok 4 did eventually follow, dated 20 August 2025, and xAI has since published cards for several subsequent releases — Grok 4.1 in November 2025, Grok 4.20 in April 2026 — closing part of the gap after the fact [1]. But the pattern recurred: reporting on the July 2026 release of Grok 4.5 found that it shipped without either a transparency report or a model card, a gap with concrete legal weight in the United States because California’s SB 53, in force since January 2026, requires covered developers to publish a transparency report — including risk evaluations conducted and any third-party evaluator involvement — before or concurrently with a covered model’s release, with fines of up to one million dollars for non-compliance [13].

This is not uniquely an xAI problem; it is a general feature of the current evaluation ecosystem that xAI’s release cadence has simply exposed more starkly than most. Anthropic’s own public case for funding third-party evaluation work states the ecosystem-wide constraint bluntly: “the current evaluations landscape is limited. Developing high-quality, safety-relevant evaluations remains challenging, and the demand is outpacing the supply” [10]. Independent evaluators need time to build sound tests, time to run them, and — this is the part a compressed release cycle removes first — time to report results before the next version supersedes the one being tested. A model card published weeks or months after a model has already been widely used answers a narrower question than it appears to: it describes the system as it was tested, which by publication date is frequently not quite the system currently being served.

A large-format printer feeding out a freshly printed report page still curling at its leading edge, above a ring binder held open with its rings only half closed on the previous page
Figure 5. New model releases and new incident reports keep arriving faster than any outside evaluator can file them; the binder is always at least one page behind the printer.Image prompt and art direction by Brecht Corbeel; generation pending.

Predictions, with the observations that would falsify them

These are forecasts, held separate from the documented material above. Horizon: 12 August 2028.

One. Independent, satellite- and filing-based reconstruction — the Epoch AI model used throughout this piece — will remain the primary external check on Colossus’s actual scale, because neither xAI nor its regulators will grant routine third-party physical audit access. Disconfirmed if xAI publishes audited, third-party-verified compute or power figures for Colossus on a recurring schedule, or if a regulator gains and exercises standing physical inspection rights.

Two. At least one further Grok-line parameter count will be reported by secondary outlets as fact before xAI issues any primary disclosure comparable to the Grok-1 release. Disconfirmed if xAI publishes a dated architecture disclosure — parameter count, activation pattern, or comparable specificity — for a Grok model at or after release, matching the standard it set with Grok-1.

Three. Further content-moderation or safety incidents involving Grok will continue to surface through public use on X rather than through pre-deployment testing, for as long as xAI’s stated evaluation strategy continues to rely on the live platform as a naturalistic test environment. Disconfirmed if xAI’s Risk Management Framework is revised to substitute a closed, non-public evaluation environment as the primary mechanism for catching this class of failure, and a subsequent release cycle passes without a comparable public incident.

Four. The gap between a Grok release date and the publication of a corresponding system card or transparency report will narrow under continued regulatory pressure — chiefly from California’s SB 53 and its enforcement record — rather than through voluntary practice. Disconfirmed if a major Grok release ships without required transparency documentation and draws no enforcement action or penalty under an applicable law within twelve months of release.

What can actually be said

Strip away the marketing on one side and the rumour mill on the other, and what remains is a short list of things that are actually documented: one confirmed parameter count, three years old, attached to a model nobody uses anymore; a compute and power estimate for Colossus that its own most careful independent source describes in half-order-of-magnitude uncertainty bands; a federal lawsuit alleging that even the number of gas turbines on site was not what regulators had approved; three public safety incidents in four months, each traceable to a specific instruction or configuration change and each caught by users rather than by pre-deployment testing; a data-training default contested by name in a formal EU regulatory inquiry; and a release cadence that has twice outrun xAI’s own disclosure commitments badly enough to draw legal and expert criticism on the record.

None of that adds up to a verdict on whether Grok is a capable model or whether Colossus is a well-run facility — those are different questions, and ones this deliberately narrow piece has not tried to answer. What it does support is a working rule for reading any claim about either one: treat a number as established only if it traces to a primary disclosure, treat an infrastructure figure as an estimate with a stated uncertainty band rather than a fact, and treat the interval between a release and its documentation as informative in its own right. On the evidence gathered here, that interval has been the most reliable indicator this series has found.