A room before it was a datacenter
Every account of AI infrastructure that starts with the accelerator is starting in the middle. The engineering problem is older than the transistor: a dense concentration of electrical work produces heat faster than an ordinary room removes it, and somebody has to decide how the room gives that heat back to the outside world. That problem has been solved and re-solved at rising scale for eight decades, and each solution left a specific, dated, attributable record. This article follows that record from the first purpose-built machine room through the liquid-cooled AI racks now shipping, and closes on a new version of the same collision — this time over power delivery rather than heat removal.
Three patterns recur, and naming them first makes the chronology easier to read. First, a new class of dense electronics repeatedly outgrows whatever cooling method the previous generation settled on, and someone responds either by conditioning the room around the machine or by putting the machine directly into a liquid. Second, once a technique proves out, it gets codified — turned into a standard, a guideline, or an industry-wide layout convention — so that equipment from different manufacturers can share infrastructure without each buyer re-deriving the physics. Third, the constraint itself eventually moves: once heat rejection is solved at a given scale, the next bottleneck appears somewhere upstream, currently at the point where power reaches the site at all.
1947: heat becomes the first constraint a computer room has to meet
ENIAC is usually remembered for arithmetic. Its own record-keepers remembered it, just as urgently, for heat. Martin Weik’s contemporary account, written by an engineer who worked on the machine at the Ballistic Research Laboratory and published in Ordnance in January-February 1961, describes roughly 19,000 vacuum tubes, 1,500 relays, and hundreds of thousands of passive components consuming almost 200 kilowatts of electrical power [1]. After ENIAC’s 1947 move to Aberdeen Proving Ground, Weik records that “the substantial quantity of heat which had to be dissipated into the warm, humid Aberdeen atmosphere created a heat-removal problem of major proportions” [1]. The fix was not a new cooling technology; it was dedicated, heavy-duty forced-air conditioning built around the machine, because no ordinary room ventilation of the period could carry the load. Weik also records that ENIAC’s continuous operation depended on solving a second, related problem years later: power-line fluctuations made running the tubes directly off transformer mains unworkable, a problem addressed in early 1952 by installing an independent motor-generator set to condition the supply before it reached the machine [1]. Even in the first computer room on record, power quality and heat removal were already two faces of one plant-engineering problem, not two separate concerns.
That is the shape everything after it repeats: a machine’s electrical load outruns the room built around it, and the response is purpose-built environmental engineering rather than a change to the computation itself.
1966-1985: when air stopped being enough, engineers put the circuits in liquid
Air conditioning bought time, but it did not scale indefinitely with circuit density. IBM’s own engineering literature traces what happened next in unusually direct terms. A 2008 paper by Michael Ellsworth and colleagues at IBM, presented at the IEEE Intersociety Conference on Thermal and Thermomechanical Phenomena in Electronic Systems, reviews water cooling across five generations of IBM large systems, beginning with the System/360 Model 91 in 1966 and running forward to the Power 575 system IBM had just announced at the time of writing [2]. The title the authors chose for that retrospective, “Back to the future,” is itself a dated data point: by 2008, water cooling in IBM’s largest systems was understood internally as a return, not a debut. The System/360 Model 91 needed it because, as a machine built for the highest sustained computational throughput of its era — scientific workloads including space-exploration, astronomy, and physics calculations — its central processor’s heat load exceeded what conditioned room air could remove from a cabinet of that density, and IBM built a dedicated coolant distribution unit into the machine’s own footprint to manage it directly [2].
Cray Research reached the same conclusion by a different route less than two decades later. The Cray-2, which Cray Research introduced in 1985 as the fastest machine in the world at the time, abandoned indirect liquid cooling entirely in favor of full immersion: dense stacks of circuit boards were submerged directly in Fluorinert, an electrically inert fluorocarbon liquid, which was itself circulated through a chilled-water heat exchanger and cooling tower built into the machine’s housing [3]. The Computer History Museum’s technical description of the surviving unit notes that the design earned the machine the nickname “Bubbles,” a direct consequence of choosing to solve the heat-removal problem at the board level rather than the room level [3]. Both machines made the same engineering judgment for the same reason: past a certain density, air handling anywhere in the room is too indirect a path for the heat to travel, and the liquid has to touch the electronics, or something very close to them, directly.
The interval when air cooling won
What is easy to miss, looking only at the landmark machines, is that liquid cooling did not become the default even after Cray and IBM had proven it worked. The generations of computing that followed — minicomputers, then the microprocessor-based servers that came to dominate general-purpose computing from the 1980s onward — mostly ran on air, and the industry spent a long interval getting very good at it. ASHRAE’s own technical committee later described exactly how good: its 2021 white paper on liquid cooling records that the server industry had driven the fraction of total server power consumed by cooling fans down “from levels as high as 20% down to as low as 2% in some cases” over the preceding decades of air-cooled design refinement [7]. A fan-power fraction that low is not an incidental detail; it is the signature of an industry that had spent years optimizing airflow paths, fan curves, and server chassis geometry specifically so that air could keep doing the job. For most of the 1990s and into the 2000s, air was winning, and it was winning because enormous, sustained engineering effort had been spent making it win.
1999-2004: the industry writes down a common answer
Winning with air at the server level did not mean every room using air was well designed, and this is where the datacenter as an engineered discipline, rather than a collection of individually tuned machine rooms, begins to take shape. According to Electronics Cooling’s account of the committee’s own history, thermal engineers from roughly fifteen information-technology manufacturers began meeting informally in 1999, prompted by contacts made at ASME conferences, to address a shared problem: there was no common language between the people who built IT equipment and the people who built the buildings and HVAC systems around it [4]. That consortium published its first joint analysis of IT equipment power trends through the Uptime Institute in 1999, then in 2002 moved its work under ASHRAE to form Technical Committee 9.9 and reach a wider audience [4].
The committee’s first published deliverable, Thermal Guidelines for Data Processing Environments, appeared in 2004. Before that document, Electronics Cooling notes, there was no single industry source specifying the temperature and humidity envelope equipment should be designed to tolerate; the 2004 first edition established a single recommended class with a maximum equipment inlet temperature of 25 degrees Celsius [4]. That single number did something the room-by-room engineering of the previous decades had not: it let a server designed by one manufacturer and a cooling system designed by another share a specification, rather than each buyer or integrator re-deriving safe operating limits from scratch. The guideline was revised as the industry gained confidence in wider tolerances — Electronics Cooling records the equipment temperature ceiling rising to 27 degrees Celsius in the 2008 edition, and the 2011 revision adding expanded allowable classes reaching as high as 40 and 45 degrees Celsius for equipment able to tolerate them [4].
The layout that made air cooling reliable at scale: hot aisle, cold aisle
Standardizing the temperature envelope solved one half of the problem; the other half was getting conditioned air to the equipment intake without letting it mix with the equipment’s own exhaust first. According to Upsite Technologies’ account of the practice’s origin, Robert “Dr. Bob” Sullivan implemented the first intentional hot-aisle/cold-aisle arrangement in the mid-1990s, breaking with the then-common practice of orienting every rack in a room to face the same direction [6]. Turning alternating rows to face each other so that cold supply air reached rack fronts from one aisle while hot exhaust collected in the alternating aisle behind them, rather than everywhere at once, meant a room’s air handlers no longer had to fight recirculating hot exhaust to keep intake temperatures down. Once the arrangement existed, Upsite notes, “a science of best practices quickly emerged to optimize the benefits of that separation” — restricting perforated floor tiles to cold aisles only, controlling underfloor obstructions that disrupted airflow, and using blanking panels to seal unused rack space so cold air could not simply bypass the equipment it was meant to reach [6]. None of this required a new cooling technology. It required treating the room’s airflow as a designed system with a hot side and a cold side that were not supposed to touch, and it became, in the years following, one of the standard reference layouts for air-cooled facilities generally.
Physical containment of that separation — solid or curtain barriers placed at the ends and tops of aisles so hot and cold air genuinely could not mix rather than merely tending not to — followed as the logical extension once the aisle geometry itself was established practice, turning a beneficial airflow pattern into something closer to two physically distinct air systems sharing one room.
The guidelines already anticipated a return to liquid
It would be a mistake to read ASHRAE’s committee as purely an air-cooling standards body that liquid cooling later disrupted. The same committee published dedicated liquid-cooling guidance well before AI accelerators created urgent demand for it. Electronics Cooling’s 2008 review of ASHRAE’s Liquid Cooling Guidelines for Datacom Equipment Centers, published in 2006 as part of ASHRAE’s Datacom series, describes a roughly hundred-page document covering the interface between facility chilled-water systems and liquid cooling loops attached directly to IT racks, motivated by density trends already visible at the time: “air-cooled systems are now struggling to provide the needed level of thermal performance for many installations,” the review states, summarizing the book’s own framing [5]. That 2006 guidance predates the GPU-accelerated computing era by roughly a decade. The committee was not caught off guard by liquid cooling’s return; it had already written the interface specification for facilities to use whenever density crossed the threshold that made air impractical again, and simply waited for the industry to need it at scale.
2022-2024: accelerator power crosses the line a second time
The line got crossed. NVIDIA’s own product documentation for the H100 GPU marks a useful anchor point on the way there: the H100 PCIe card, in a product brief NVIDIA published in November 2022, operates up to a maximum thermal design power of 350 watts using a passive heat sink that “requires system airflow to operate the card properly within its thermal limits” — a fully air-cooled part, dependent on the surrounding server’s fans rather than any liquid loop [8]. Two years later, NVIDIA’s rack-scale Blackwell platform made the opposite design choice mandatory rather than optional. The GB200 NVL72, as NVIDIA’s own product page describes it, integrates 72 Blackwell GPUs and 36 Grace CPUs into a single liquid-cooled rack presented as “an exascale computer in a single rack,” built around a 72-GPU NVLink domain moving 130 terabytes per second of GPU-to-GPU traffic [9]. NVIDIA states the design was built specifically to deliver dramatically more performance within a comparable power envelope relative to air-cooled H100 infrastructure — the company’s own framing is “25x more performance at the same power” [9]. Reaching that density at that power draw is what made liquid cooling non-optional again: the same ASHRAE white paper that documented the 2% fan-power floor of the air-cooling era also renamed the committee’s facility water classes to carry their upper temperature limits directly in the class name — W17 through W45 — and stated plainly that required facility water temperatures are expected to fall over time as chip heat flux rises, the opposite of what would make an operator’s life easier [7]. Between a 350-watt air-cooled card in 2022 and a rack mandating liquid contact at the chip two years later sits the same collision ENIAC’s engineers met in 1947 and IBM’s engineers met in 1966, replayed at a scale neither could have specified hardware for.
2024-2026: the constraint moves upstream, to the meter
Solving heat rejection at rack scale did not remove the next bottleneck; it exposed one. Once a facility can absorb whatever heat its densest racks produce, the binding limit on how much AI compute a site can run becomes how much power reaches the site in the first place, and grid interconnection has not kept pace with cluster buildout schedules. The clearest visible response has been operators generating power on-site, behind their own meter, rather than waiting for a utility interconnection. xAI’s Colossus facility in Memphis is the most heavily documented case: reporting by TechCrunch, based on aerial photography and thermal imaging commissioned by the Southern Environmental Law Center, found at least 35 natural-gas combustion turbines installed around the site with a peak generating capacity of roughly 421 megawatts, with the Southern Environmental Law Center alleging in a notice of intent to sue that the turbines were installed and operated without the preconstruction or operating air permits Tennessee law generally requires for new pollution sources [10]. Whatever the outcome of that specific regulatory dispute, the underlying engineering decision it documents — build generation on-site rather than wait for a grid connection — is the same reflex ENIAC’s engineers showed when they installed their own motor-generator set in 1952 rather than accept mains fluctuations, applied at a scale roughly a thousand times larger.
The scale of the underlying demand growth is what makes that reflex increasingly common rather than exceptional. The International Energy Agency’s April 2026 update reports that data centre electricity demand grew 17% in 2025 against overall global electricity demand growth of just 3%, and identifies tightening supply chains for gas turbines and transformers as an active constraint on new capacity, alongside strained planning and regulatory systems struggling to keep pace with the pipeline of proposed projects [11]. That is a facility-wide figure covering all data centres, not an AI-specific measurement, and it should not be quoted as though it were the latter. But it establishes the scale of the pressure pushing operators toward behind-the-meter generation: when interconnection queues and equipment supply chains cannot deliver power on a cluster’s construction timeline, on-site generation becomes the substitute, exactly as it was the substitute for unreliable mains power at Aberdeen Proving Ground in 1952 — except that the modern version measures its self-generated capacity in gigawatts rather than kilowatts, and burns gas rather than turning a motor-generator set.
What this history actually shows
Read end to end, the record does not show steady, incremental improvement. It shows the same specific collision recurring at each new order of magnitude in density, met each time by one of exactly two responses: condition the room more aggressively around the machine, or put a liquid directly against the machine. ENIAC forced dedicated air conditioning because ordinary room ventilation could not remove 200 kilowatts of tube heat. The System/360 Model 91 and the Cray-2 forced direct liquid contact because even conditioned room air could not remove the heat their densest components produced. The industry then spent roughly two decades proving that air, properly engineered — driven down to a two-percent fan-power overhead, delivered through a room laid out with the intake and exhaust streams kept physically apart — could handle almost everything short of the most extreme systems, and it wrote that consensus down as a shared standard so the whole industry could build against one specification instead of many. Accelerator power density broke that consensus a second time, on a schedule visible in NVIDIA’s own product documentation across just two years. And now the same collision is repeating one level further upstream, where the limit is no longer how fast heat leaves a rack but how fast power reaches a site, and where the answer emerging is the direct-connection instinct of 1952 restaged with combustion turbines instead of a motor-generator set.
None of this required a conceptual breakthrough at any step. It required engineers to notice, each time, that the previous generation’s solved problem had been solved for a lower density than the next generation would present, and to reach for direct contact — liquid on the chip, generation on the site — once the indirect path through the room or through the grid stopped being fast enough.
Predictions, with the observations that would falsify them
These are forecasts, clearly separated from the sourced history above. Horizon: 15 August 2029.
One. ASHRAE’s facility water classes will be revised again to accommodate colder required supply temperatures, continuing the direction already stated in the 2021 white paper, rather than stabilizing at W17-W45. Disconfirmed if the 2029 edition of the liquid cooling guidelines widens rather than tightens the coldest required class.
Two. Behind-the-meter generation at AI clusters will be retrofitted with emissions controls or replaced by grid-connected or lower-emissions on-site generation, rather than remaining predominantly uncontrolled combustion turbines, as regulatory exposure of the kind documented at Colossus becomes a recognized cost of the approach. Disconfirmed if new large AI clusters in 2029 are still predominantly powered at commissioning by uncontrolled, unpermitted on-site combustion turbines.
Three. A specification for the interface between behind-the-meter generation and datacenter electrical plant will be codified by a standards body, the way ASHRAE codified the air-cooling and liquid-cooling interfaces before them. Disconfirmed if by 2029 no major standards body has published guidance specific to on-site or behind-the-meter generation for datacenters.
What to take away
Cooling and power engineering for computing has never advanced on a smooth curve. It has advanced by hitting the same ceiling — a density of electrical work that the existing path for removing heat or delivering power cannot handle — at successively larger scales, roughly once per generation of hardware, and by responding with one of two moves each time: engineer the indirect path harder, or go direct. ENIAC’s forced-air room, the System/360 Model 91’s water jackets, the Cray-2’s Fluorinert tank, ASHRAE’s 2004 temperature envelope, Sullivan’s hot-aisle/cold-aisle layout, the GB200 NVL72’s mandatory cold plates, and the gas turbines now standing behind AI clusters are not seven different stories. They are the same story, dated and re-measured seven times. The current chapter, in which the binding constraint has moved from the rack to the meter, is not the end of that pattern. It is simply the most recent place the pattern has reached.