Two different questions wearing one name
“xAI” answers to two different kinds of claim, and this article is only about one of them.
The first kind is about Colossus: how many GPUs are installed at the Memphis site, how many days it took to bring them online, what powers them, and who else now rents time on the finished cluster. These are, for the most part, documented, checked in places by independent reporters and research groups against satellite imagery, permitting filings, and utility disclosures. Where the record is thin or contested — and on power sourcing, it is both — this piece says so directly.
The second kind of claim is about Grok: how large the model trained on that cluster actually is, measured in parameters. That number is, for every Grok model released after March 2024, not documented anywhere xAI has put its name to. What circulates instead is a chain of social-media statements and outlet-to-outlet repetition that drifts from 1.5 trillion to 2 trillion to figures as high as 15 trillion depending on which week’s post is being summarized.
The premise of this article is that these two kinds of claim should never lend each other credibility. A well-documented cluster does not make an undocumented parameter count more likely to be true. This piece keeps them apart throughout, comparing xAI’s infrastructure build — the part that is actually written down — against the equally documented, differently shaped approaches of Microsoft and OpenAI, Google, and Amazon.
What Colossus actually is, on xAI’s own account
xAI’s most specific infrastructure claims are about speed. NVIDIA’s own account of the joint deployment, published as the cluster went live, states that xAI built the Memphis facility and populated it with 100,000 Nvidia Hopper GPUs in 122 days, with training beginning just 19 days after the first server rack arrived on the floor — a timeline the release frames as far shorter than comparable builds, which it says have historically taken measured in years rather than months [5]. xAI’s own Grok 3 announcement corroborates the direction, though not with a parameter count: it says Grok 3 was “trained on our Colossus supercluster with 10x the compute of previous state-of-the-art models,” a comparison stated entirely in compute, never in model size [2].
The cluster kept growing. By the time xAI announced Anthropic’s access to compute on the system in 2026, xAI’s own description put Colossus 1 at “over 220,000 NVIDIA GPUs, including dense deployments of H100, H200, and next-generation GB200 accelerators” [4]. xAI’s Series E announcement, four months earlier, described “the world’s largest AI supercomputers at Colossus I and II, ending the year with over one million H100 GPU equivalents” and said the newly raised $20 billion would accelerate “our world-leading infrastructure buildout” [3]. Those are xAI’s own figures about xAI’s own hardware — a primary disclosure about infrastructure, not a rumor about a model, and this piece repeats it on that basis.
Worth sitting with: a rival lab renting capacity on Colossus 1 is a form of external validation a parameter-count rumor could never provide. Anthropic’s engineers now have direct, paying access to whatever xAI actually built, and a cluster that could not do what xAI claims would be a commercial problem in a way a wrong number on a fan blog never is.
Measuring the speed claim against everyone else’s timeline
A single company’s own account of its own speed is worth checking against a method that does not depend on any company’s cooperation. Epoch AI, an independent research group, built exactly that: using satellite imagery, local permitting records, and company disclosures, it tracked the interval between construction start and the point a facility could be documented as drawing at least one gigawatt of power, across seven major AI data center campuses that broke ground within the past three years [7].
The gap is real and it is large. By Epoch AI’s tracking, xAI’s Colossus 2 reached the 1-gigawatt mark roughly 12 months after construction began. Every other tracked campus took longer: Amazon and Anthropic’s New Carlisle, Indiana site and Microsoft’s Fayetteville campus each took about 23 months; OpenAI’s Stargate site in Abilene took about 25 months; Microsoft’s Fairwater campus took about 32 months; Amazon’s Ridgeland site took about 43 months; and Meta’s Prometheus campus, the slowest tracked, took roughly 101 months [7]. Across all seven, the researchers describe a 1-to-3.6-year range from groundbreaking to gigawatt-scale power, with xAI at the fast end by a wide margin.
Epoch AI is explicit about what its method cannot pin down: satellite-imagery timing is imprecise, cooling-load estimates that back out a power figure can be wrong by roughly a factor of two, permitting practice is not uniform across jurisdictions, and companies do not use the word “operational” the same way [7]. None of that erases the gap between 12 months and 101 months, but it does mean the month counts are estimates with real error bars, not a scoreboard accurate to the week — a distinction the source insists on and this article preserves.
Two ways to spread a training run across ground
“Fast” and “slow,” on their own, describe only one axis. A second, separate choice each company has made and documented is where to put the capacity: concentrated on one site, or spread deliberately across several.
Colossus is the concentrated case almost by design — one former appliance factory in Memphis, expanded on the same footprint and later a neighboring site, rather than federated across cities. Microsoft has taken the opposite approach with Fairwater. Its own account of the flagship Wisconsin site describes “the largest and most sophisticated AI factory we’ve built yet,” a 315-acre, three-building, 1.2-million-square-foot campus using “a single flat networking” fabric to bind hundreds of thousands of GPUs into what Microsoft calls one supercomputer rather than independent servers [9]. Fairwater is explicitly one node of several: Microsoft pairs the Wisconsin site with a second campus in Georgia, describing the pair as connected pieces of a planet-scale “AI superfactory” designed from the outset to span more than one location rather than to maximize what a single site can hold.
OpenAI’s Stargate program, built with Oracle and SoftBank, documents the multi-site pattern even more explicitly, since each expansion has been announced as a named list of separate sites rather than a single campus growing larger. SoftBank’s release on the September 2025 expansion states that five new sites — in Texas, New Mexico, Ohio, and an undisclosed Midwest location — brought Stargate’s planned capacity to “nearly 7 gigawatts” backed by “over $400 billion in investment over the next three years,” on top of the flagship Abilene, Texas campus already under construction, with a stated ambition of $500 billion and 10 gigawatts of committed capacity [12]. Three of the five new sites are developed with Oracle; two with SoftBank and its SB Energy affiliate — spreading geography, ownership, and power procurement across multiple corporate partners, with no analogue in xAI’s wholly controlled Memphis buildout.
Amazon’s Project Rainier sits between the two patterns. Its primary campus is concentrated — Amazon’s announcement describes “nearly half a million” custom Trainium2 chips brought online at a campus in St. Joseph County, Indiana, built for a single anchor tenant, Anthropic, in under a year from the project’s initial announcement [11]. But Amazon also runs a second, smaller Rainier campus in Mississippi, and has disclosed plans to scale Anthropic’s access toward more than a million Trainium2 chips across “direct usage and Amazon Bedrock” by the end of 2025 — Bedrock being Amazon’s shared cloud service, which spreads capacity across many customers rather than one [11]. Rainier is concentrated at the site level but distributed at the ownership level in a way neither Colossus nor Stargate is.
Custom silicon changes what “infrastructure” even means
Google’s documented approach differs from all of the above in a more basic way: its primary infrastructure decision is not only about sites and power, but about the chip itself. Google’s account of its seventh-generation TPU, Ironwood, describes a purpose-built accelerator delivering “42.5 Exaflops” at full pod scale, a single pod made of “9,216 liquid cooled chips” tied together over Google’s own inter-chip interconnect, with a doubling of performance-per-watt over the prior TPU generation and a roughly 30-fold efficiency gain over Google’s first Cloud TPU from 2018 [10]. None of that is Nvidia hardware, and none of it is described in terms comparable to a GPU count, because it is not a fleet of merchant GPUs assembled quickly — it is a chip Google designs, fabricates through partners, and deploys inside datacenters engineered around it.
That difference matters for the comparison itself. A GPU-based buildout like Colossus or Rainier can, in principle, be measured on a shared yardstick — GPU count, generation, interconnect — because the chips are the same ones available to any buyer with capital and allocation. A TPU pod cannot: the relevant unit of “infrastructure speed” for Google is design-and-fabrication lead time on custom silicon, which does not reduce to a days-to-gigawatt question at all. This is one of the places where a single ranking across companies would misrepresent what is being compared, because the axis of comparison itself changes.
Amazon’s variant: silicon designed with one tenant in mind
Amazon’s approach is a third pattern, related to but different from Google’s. Like Google, Amazon has moved toward chips it designs itself — Trainium2 — but unlike Google’s TPU fleet, which serves Google’s own products and a broad set of Cloud customers, Project Rainier’s headline deployment was built around and substantially dedicated to a single primary tenant, Anthropic, from the outset [11].
Amazon’s physical description of the deployment is itself a statement about architecture, not just scale: an “UltraServer” groups four physical machines and 64 Trainium2 chips over Amazon’s own NeuronLink interconnect, and tens of thousands of UltraServers are joined into an “UltraCluster” over its Elastic Fabric Adapter networking [11]. That is a training fabric engineered jointly with Rainier’s anchor customer — closer in spirit to Google’s custom-silicon model than to xAI’s open-market GPU assembly, but organized, unlike Google’s shared fleet, primarily around one partner’s needs.
The power question, and where the field genuinely disagrees
Nothing in the speed and siting comparisons above explains why xAI could build faster than everyone else it has been measured against. Power does. A gigawatt of capacity is worthless without a gigawatt of electricity to run it, and the four companies compared here have made visibly different, and in places contested, choices about where that electricity comes from.
xAI’s documented approach is to generate a substantial share of its own power on site rather than wait for grid interconnection. Global Energy Monitor’s tracking of the Colossus 2 power station identifies the installation as built around Solar Turbines Titan-350 gas turbine units, each rated at roughly 35 to 38 megawatts, owned through a joint venture in which Solaris Energy Infrastructure holds 50.1 percent and xAI holds 49.9 percent; the tracker records Mississippi regulators granting xAI’s operating entity temporary approval to run the turbines for up to twelve months without the permit that would ordinarily be required first, with a buildout plan reaching more than 1.1 gigawatts of turbine capacity by the second quarter of 2027 [6]. Independent energy-sector reporting frames this as a deliberate trade: Latitude Media’s account of 2025 describes xAI adding roughly a gigawatt “incredibly quickly” using more than twenty additional gas turbines through the year, a choice that “sparked immediate environmental backlash over smog and pollution” locally, while Amazon added a larger total volume of new capacity, 3.8 gigawatts, largely through conventional grid and power-purchase channels, and Microsoft and Google instead emphasized carbon-matching and long-dated clean-power contracts [8].
Those long-dated contracts are themselves a documented, if slower, infrastructure strategy. Microsoft’s agreement with Constellation Energy is a 20-year power-purchase agreement for the full 835-megawatt output of a restarted Three Mile Island Unit 1, now renamed the Crane Clean Energy Center, with Constellation committing roughly $1.6 billion to the restart and targeting commercial operation in 2028 pending Nuclear Regulatory Commission review [13]. Amazon’s parallel nuclear strategy through Talen Energy at the Susquehanna plant has already met a regulatory limit the turbine approach has not: in November 2024 the Federal Energy Regulatory Commission voted 2–1 to reject an amendment that would have expanded the Amazon-Talen arrangement from 300 to 480 megawatts, which Talen said it believed was mistaken; the same reporting notes Meta separately abandoned a planned nuclear partnership after a rare bee species was found on the proposed site [14].
None of this supports a claim that one power strategy is simply better. The documented record shows three different exposures. xAI’s on-site generation converts a permitting delay into an environmental and regulatory dispute it is fighting in public, in exchange for capacity that does not wait on the grid or on FERC. Microsoft’s and Amazon’s nuclear contracts convert that delay into years of regulatory review and construction risk, in exchange for power not subject to a comparable local-emissions fight. Where those exposures net out is a judgment call this article declines to make, because the parties disagree about it and the dispute is ongoing, not settled.
What the speed premium is actually buying
The comparisons above describe what differs. It is worth being explicit about why a company would accept the risk xAI has accepted, because running turbines ahead of a permit is otherwise just recklessness with extra steps.
Model the choice as a simple trade. A firm can reach a target power capacity
whenever it takes the faster, costlier, more legally exposed route over the slower, cheaper, better-permitted one. The model does not say what
Decoupling Colossus’s size from Grok’s
Everything above is about the cluster. None of it is a statement about how large any Grok model is, and the two should not be allowed to blur into each other, which is the mistake this article is built to avoid.
xAI has disclosed a specific parameter count for exactly one Grok model. Its March 2024 release of Grok-1’s weights states plainly: “Grok-1 is a 314 billion parameter Mixture-of-Experts model trained from scratch by xAI,” with roughly a quarter of those weights active on any given token [1]. That is a primary disclosure, directly attributable, and it is the only Grok parameter count this article treats as an established fact.
Every model since has been handled differently. Grok-2’s weights were later released publicly too, but xAI’s release materials do not state a parameter count anywhere in them; independent analysts who examined the checkpoint files calculated the architecture themselves, reverse-engineering the mixture-of-experts structure under the model’s eight-way tensor parallelism to arrive at roughly 270 billion total parameters with about 115 billion active per token [15]. That estimate is worth more than a rumor — derived from the actual published weights, not a social-media claim — but it is still not an xAI statement, and this article reports it as a calculation, attributed to whoever did the calculating.
From Grok 3 onward, even that trail disappears. xAI’s Grok 3 announcement describes training compute relative to earlier models and says nothing about size [2], and no subsequent xAI model card examined for this piece states a parameter count for Grok 4, 4.1, 4.5, or 4.6. What fills the gap is a pattern of numbers attributed to Elon Musk’s own social-media statements: reporting on a post dated July 18, 2026 describes him calling the then-upcoming Grok 4.6 a “2 trillion parameter” model entering final training, while noting explicitly that this came from Musk’s account on X rather than “an official xAI model card or technical documentation,” and that xAI “has not yet announced” the model’s launch date or technical metrics [16]. Other figures circulating for the same and later models — 1.5 trillion, 2.1 trillion, 6 trillion, 10 trillion — trace through the same kind of chain: a founder statement or an unnamed leak, repeated and sometimes altered across outlets, never converging because no single document anchors any of them.
This is the article’s central disclosure claim: xAI has been comparatively forthcoming about Colossus, publishing GPU counts, timelines, and funding figures under its own name, while becoming steadily less forthcoming about Grok’s architecture with each release. Those are two trend lines moving in opposite directions inside one company. Treating a well-documented cluster as evidence for an undocumented parameter figure — reasoning that “the world’s largest AI supercomputer” must imply a proportionately enormous model — is exactly the inference this article declines to make, because xAI’s own compute-comparison language is compatible with many combinations of model size and training duration, and xAI has not said which one it chose.
What is actually different, without a ranking
Four defensible, independently documented approaches sit next to one another. xAI concentrates compute on one site, brought online faster than any comparably sized facility Epoch AI has tracked, paid for partly with on-site gas generation that has drawn open regulatory and environmental objection. Microsoft and OpenAI spread capacity across multiple named sites tied together by wide-area network fabric and multiple corporate partners, trading single-site speed for diversification, and lean more on long-dated clean-power contracts carrying their own multi-year construction and regulatory risk. Google routes around the GPU market’s timeline entirely by designing and fabricating its own accelerator, making “days to a gigawatt” a category error for its infrastructure rather than a metric it loses on. Amazon splits the difference: custom silicon, deployed at a concentrated flagship site, but built jointly with and substantially dedicated to a single primary tenant rather than run as a general-purpose fleet.
None of the four is a strictly better version of another; each trades away something the others kept. Speed against permitting and emissions scrutiny. Site concentration against exposure to a single location, utility, and set of regulators. Custom silicon against a chip’s own multi-year design cycle. Shared-fleet flexibility against a training partner locked in from the start. Where credible people disagree about how those trades will look in five years — and reporting shows they do, from FERC commissioners split 2–1 to Amazon’s own account of adding more raw capacity than xAI while doing it more slowly — this article records the disagreement rather than resolving it.
What would change this assessment
These are dated forecasts, kept separate from the sourced comparison above. Horizon: August 2028.
One. xAI’s speed advantage will narrow rather than disappear, as competitors adopt faster permitting workarounds and modular power deployment of their own. Disconfirmed if a hyperscaler-scale campus not built by xAI reaches 1 gigawatt in under nine months, matching or beating Colossus 2’s documented pace on Epoch AI’s methodology.
Two. xAI will publish at least one further parameter count as specific as Grok-1’s, most plausibly tied to an open-weights release rather than a flagship product launch, because that is the only channel through which xAI has disclosed such a figure to date. Disconfirmed if, by the horizon date, no xAI model card or weights release for any Grok generation after 4.6 states a specific parameter count.
Three. At least one of xAI’s currently temporary or unpermitted gas-turbine installations will be forced into a materially different operating status — permitted, relocated, or curtailed — by regulatory or legal action. Disconfirmed if the Colossus-area turbine fleet continues operating under the same temporary or provisional authority it holds today, unchanged, through the horizon date.
Four. The rumored parameter figures for Grok 4-series and Grok 5 models will continue to disagree with one another across outlets by a factor of two or more, because no primary disclosure exists to converge them. Disconfirmed if independent outlets’ reported figures for a single named Grok model converge to within twenty percent of one another absent a new xAI disclosure.
What to take away
Set side by side, the documented infrastructure records of xAI, Microsoft and OpenAI, Google, and Amazon describe four genuinely different bets about what matters most when building the physical plant behind a frontier AI lab — raw speed, geographic and partner diversification, custom silicon, or a tenant-anchored fleet — each with a cost the other three avoided. That comparison can be made honestly because each company put real numbers on the record, and because independent researchers and reporters have, in places, checked those numbers against satellite imagery, permits, and utility filings.
The Grok parameter count cannot be compared this way, because it has not been put on the record in the same manner since Grok-1. A reader who wants to know how big the model running on Colossus actually is should notice that the cluster’s size is not evidence for it, that xAI’s own most recent word on the subject is a compute ratio rather than a parameter count, and that every larger number now circulating traces back to a statement xAI itself did not make in any document bearing its name.