A bench with no winner declared

Ask a datacenter engineer which cooling technology is “best,” or which power-sourcing strategy is “right,” and a careful one will answer with a question of their own: best for what site, at what existing rack density, in what climate, at what point in an interconnection queue. That deflection is not evasion. It is the correct response to a genuinely multidimensional problem that trade press coverage routinely flattens into a single leaderboard.

This article makes the comparison structured instead of flattened. It treats AI datacenter engineering as two largely separable decisions. The first is how to reject heat from a rack: traditional air cooling, direct-to-chip cold plates, rear-door heat exchangers, or immersion. The second is how to source the power that becomes that heat in the first place: grid interconnection, behind-the-meter generation, or on-site renewables paired with storage. Both decisions get evaluated against four concrete axes — efficiency, retrofit cost, water consumption, and deployment speed — because a technology that wins on one axis routinely loses on another, and the industry argument that looks like a technical dispute is often just two people weighting the same four axes differently.

The governing rule, stated by this publication’s editorial standard and worth stating again here because this article is organised entirely around it: never build a cross-vendor ranking from incomparable data, and where engineers or regulators genuinely disagree, characterise the disagreement rather than picking a side. Four cooling methods and three power-sourcing strategies appear below. None of them is declared the winner, because the evidence does not support declaring one.

ADVERTISEMENT

Four ways to reject heat from a rack

Every cooling method on the bench solves the same physical problem — move heat from a semiconductor junction to the outside world — using a different combination of proximity to the heat source, working fluid, and mechanical complexity. ASHRAE’s Technical Committee 9.9, the standards body whose guidance underlies almost all commercial data center thermal design, frames the four options as points on a single spectrum of how close the working fluid gets to the die before the heat is captured [5]. That spectrum is the organising idea for this section.

Air cooling: the reference case

Air cooling remains the installed base against which every liquid alternative is measured, and it is not obsolete — it is bounded. ASHRAE’s own historical data shows the server industry drove fan power down “from levels as high as 20% down to as low as 2% in some cases” over roughly a decade, a genuine efficiency achievement built entirely on moving more air with less energy [5]. The bound is architectural rather than a failure of engineering effort: a rack aperture can only accept so much airflow before acoustic limits, floor-tile delivery capacity, and fan power as a share of total draw all move in the wrong direction simultaneously. ASHRAE’s guidance notes that products requiring more than 100 cubic feet per minute per rack unit are already on the market, which — relying solely on raised-floor delivery through a single tile rated near 1,900 cfm — would occupy only nineteen rack units of usable height before airflow, not power, becomes the binding constraint [5]. Air cooling’s genuine advantages are retrofit simplicity (nothing new touches the IT hardware), zero water at the rack, and a supply chain every facilities team already knows. Its genuine limit is a density ceiling that GPU-class racks now regularly exceed.

A finned aluminium heat sink and frame fan on the comparison bench with the fan caught part-seated in its plastic shroud
Figure 1. Air cooling remains the reference case on the bench, the baseline every liquid alternative is still measured against.Image prompt and art direction by Brecht Corbeel; generation pending.

Direct-to-chip: heat capture at the package

Direct-to-chip cold plates capture heat exactly where the temperature is highest and the thermal resistance to the transport fluid is lowest — at the processor package itself, sometimes extended to memory modules mounted on both sides of a DIMM card [5]. This proximity is the source of its efficiency advantage: a cold plate typically removes the majority of a rack’s heat load directly into a liquid loop, leaving a residual air load for memory, drives, and power supplies that a standard room-level air system can still handle. The retrofit calculus, according to ASHRAE’s synthesis of first-cost comparisons across cooling solutions, is density-dependent rather than fixed: “as rack density increases, the installed first cost per MW decreases,” to the point that “there is a rack density point, which varies per cooling solution, at which the first cost to build and deploy liquid cooling is actually lower than air-cooling the same IT equipment load, making the payback period instantaneous” [5]. Below that density point, direct-to-chip is a capital-intensive retrofit requiring a new facility water loop, coolant distribution units, and blind-mate plumbing on every server tray. Above it, air cooling is the expensive option. The crossover point is a site-specific engineering calculation, not a universal constant — which is precisely why comparing “direct-to-chip cost” against “air cooling cost” as if each had one number is a category error.

A machined copper cold-plate module on the comparison bench with one bronze quick-disconnect coolant fitting caught mid-connection
Figure 2. Direct-to-chip cooling captures heat at the package, where the temperature is highest and the fluid path is shortest.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Rear-door heat exchangers: the low-disruption retrofit

A rear-door heat exchanger intercepts heat at the rack’s exhaust boundary rather than at the package, replacing (or supplementing) the rear door of a standard rack with a liquid-cooled coil. ASHRAE frames it explicitly as the retrofit path for facilities running into raised-floor airflow limits: rather than rebuild room-level delivery, “one could consider replacing or supplementing raised-floor cooling with a closely coupled coil such as a rear door heat exchanger” [5]. The appeal is architectural conservatism — the rack itself, the servers inside it, and the room’s existing air-handling philosophy all stay largely intact; only the rear door and a facility water connection change. A real deployment illustrates the ceiling on that conservatism: the Frontera supercomputer at the Texas Advanced Computing Center, at the time the ninth most powerful system in the world at 39 petaflops across 8,008 nodes, ran a hybrid design with CPUs direct-liquid-cooled and the rest of each node air-cooled through a rear-door heat exchanger in the rack [5]. That hybrid pattern is common precisely because a rear door alone rarely captures enough heat for the densest modern racks — it is usually paired with, not a substitute for, package-level cooling once density climbs far enough.

A rear-door heat exchanger panel propped on the comparison bench with its aluminium fin core exposed and a bronze supply valve handle caught mid-turn
Figure 3. A rear-door heat exchanger asks the least of the room it retrofits into, and the smallest thermodynamic ambition on the bench is also the fastest one to install.Image prompt and art direction by Brecht Corbeel; generation pending.

Immersion: the ceiling, and the cost of reaching it

Immersion cooling submerges IT hardware directly in a dielectric fluid, eliminating server fans and the air path entirely. ASHRAE credits it with “the benefits of broad temperature support, high heat capture, high density, and flexible hardware and deployment options,” and notes intense commercial interest, “with a large number of startup and established players enabling large integrated solutions and proof of concepts in all key market segments” [5]. The same guidance is equally direct about the cost of that ceiling: immersion “can create additional challenges and complexities, specifically with regard to deployment and service. For tank applications, a crane or two-man lift is often required to remove IT equipment hardware for service.” For rack-based tank designs, “sealing the immersion fluid can be challenging,” and “interoperability with an immersion fluid with IT equipment hardware can also be an issue that may impact warranty” — serious enough that ASHRAE recommends “a materials compatibility assessment and warranty impact evaluation” before any deployment decision [5]. Even the thermal ceiling has a ceiling: as component heat density rises, single-phase natural-convection immersion can be exceeded, forcing “a move to combine other cooling technologies or a wholesale shift to forced convection or two-phase immersion cooling altogether” [5]. Immersion is the cooling method most likely to be the correct answer for the highest-density future racks, and the one whose service model, floor-loading requirements, and fluid-compatibility risk are least like anything a conventional data center operations team already runs.

ADVERTISEMENT
A clear-walled immersion cooling tank on the comparison bench with a server blade caught partway into the dielectric fluid on its lifting bail
Figure 4. Immersion cooling submerges the whole assembly, and the highest efficiency on the bench comes with the hardest handling.Image prompt and art direction by Brecht Corbeel; generation pending.

Water is not where the liquid-versus-air framing predicts it

A common shorthand treats “liquid cooling” and “water-hungry” as synonyms, and “air cooling” as the water-frugal default. Lawrence Berkeley National Laboratory’s 2024 national data center energy model shows this shorthand is wrong on its own terms, because water consumption tracks the heat-rejection route chosen at the facility level — not the rack-level choice between air and liquid.

Water Usage Effectiveness, the metric LBNL uses throughout its modeling, is defined analogously to the more familiar Power Usage Effectiveness:

WUEsite=VwaterEIT \mathrm{WUE_{site}} = \frac{V_{\mathrm{water}}}{E_{\mathrm{IT}}}

where VwaterV_{\mathrm{water}} is total on-site water consumption in litres and EITE_{\mathrm{IT}} is the electricity delivered to IT equipment, giving WUE units of litres per kilowatt-hour rather than PUE’s dimensionless ratio [4]. What LBNL’s simulations across roughly 965 US weather stations and multiple facility archetypes show is that this number is driven overwhelmingly by which heat-rejection route a facility uses at the plant level — an economizer decision layered on top of, and largely independent from, the cooling method at the rack. Facilities running “water-cooled chiller systems without economizers exhibit the highest WUE, largely attributed to substantial cooling tower water usage,” and this pattern holds for both air-cooled and liquid-cooled IT equipment alike when both sit behind a waterside economizer [4]. Conversely, facilities using airside economizers “allow for the shutdown of chilled water systems during favorable weather conditions, resulting in substantial water conservation,” regardless of whether the racks behind them are air- or liquid-cooled [4].

The direct-to-chip and rear-door sections above already contain the mechanism that reconciles this with rack-level physics: liquid-cooled IT equipment “can often be operated at higher water/refrigerant temperatures than air-cooled IT systems,” which lets a waterside economizer run in “free cooling” mode for more of the year and can — when paired with a dry cooler rather than an evaporative tower — reduce water use even as it improves PUE [4]. But that outcome depends on which heat-rejection route sits behind the rack, not on the liquid-versus-air choice by itself. A facility can be liquid-cooled and water-intensive (waterside economizer with a cooling tower) or liquid-cooled and nearly water-free (dry cooler with no adiabatic assist), and the same bifurcation exists on the air-cooled side. LBNL’s own caution deserves repeating here rather than paraphrasing: “there are tradeoffs between low PUEs and low site WUEs. For example, water-cooled chillers and other evaporation-based cooling systems are generally more energy efficient than an air-cooled chiller or other waterless systems. While air-cooled chillers use no water, they use more energy” [4]. Treating water use as a property of the rack-level cooling technology, rather than of the facility-level heat-rejection route chosen independently, is the single most common error in public discussion of this trade-off.

Three ways to get power to a datacenter

The power-sourcing decision is nominally simpler — there are fewer named options — but the axes pull in sharper opposition than they do on the cooling side, because deployment speed and grid impact are now regulated concerns rather than purely engineering ones.

Grid interconnection: still the default, no longer a passive one

Connecting to the existing transmission grid remains the default path, and it remains the slowest one at scale. The IEA’s 2025 assessment states plainly that transmission line construction “requires four to eight years in advanced economies,” that wait times for transformers and cables “have doubled in the past three years,” and that gas turbine deliveries — relevant even to grid-side peaking capacity, not only behind-the-meter generation — now face “lead times of several years” [1]. Its headline risk estimate is specific: “unless these risks are addressed, around 20% of planned data centre projects could be at risk of delays” [1]. Tyler Norris’s congressional testimony, drawing on his Duke University research, adds a harder edge to the same picture: transformer order lead times have “grown to two to five years, up from less than one year in 2020, while costs have surged by 80%,” and “some utilities have quoted interconnection delays for new large loads ranging up to 7 to 10 years” [3].

ADVERTISEMENT

What has changed more recently than the queue itself is what grid operators now demand from a connected data center once it is through the queue. NERC’s May 2026 Reliability Guideline for large loads reclassifies AI datacenters from passive consumers to active participants whose electrical behavior must be characterised, monitored, and constrained. Because large training loads “operate cyclically” and can cycle “in the electromechanical range (i.e., 0.1–2 Hz for inter-area and local modes),” they risk “interacting with natural low frequencies and causing widespread forced oscillations” — a risk NERC states is not hypothetical, citing “one cryptocurrency mining facility in ERCOT” and “one data center in Dominion” where “complex dynamic control interactions have been demonstrated to cause oscillations” [2]. On harmonics, the guideline directs transmission providers to require harmonic current spectra “up to the 100th harmonic order” against IEEE Standard 519-2022 limits, and specifies that “if design changes result in 10% or more increase in projected harmonic emissions, interconnection studies should be re-performed” [2] — a provision that turns even a mid-project GPU vendor swap into a potential re-study trigger. The guideline’s own framing of urgency is stark: customer-initiated load reduction and oscillation events “can transpire in a matter of seconds, leaving real-time operators little to no time to respond” [2]. Grid interconnection therefore trades a genuinely lower marginal cost per delivered megawatt for a queue measured in years and, increasingly, an ongoing compliance relationship that did not exist for a conventional commercial load.

The same body of research complicates the “grid is simply slow” framing, however. Norris and colleagues’ analysis of 22 of the largest US balancing authorities — covering roughly 95 percent of the country’s peak load — found an average system-wide load factor of just 53 percent, meaning “nearly half of US electrical infrastructure remains unused at any given time, on average” [3]. Their concept of curtailment-enabled headroom quantifies what that slack is worth to a flexible new load: 76 gigawatts of new load, equivalent to about 10 percent of national peak demand, could be integrated system-wide with an average annual curtailment rate of just 0.25 percent; 98 gigawatts at a 0.5 percent curtailment rate; 126 gigawatts at 1.0 percent [3]. The curtailment burden this implies is modest in duration — “the average curtailment event lasts about two hours, and nearly 90% of hours during which load reduction is required retain at least half of the new load” [3]. This does not shrink the interconnection queue’s paperwork, but it changes what “grid-connected” can mean: a datacenter willing to accept occasional, short curtailment can plausibly interconnect faster and at lower system cost than one insisting on firm, uninterruptible service from day one.

Behind-the-meter generation: buying speed, inheriting regulatory risk

Behind-the-meter generation — typically natural gas turbines sited with, or physically co-located with, the datacenter — exists explicitly to route around the interconnection queue rather than to be cheaper in isolation. Norris’s testimony captures the trade candidly: “lead times for gas turbines have reportedly reached four years,” with NextEra’s chief executive stating on a recent earnings call that new gas projects “won’t be available at scale until 2030, and then only in certain pockets of the US” [3]. Four years is still faster than a seven-to-ten-year interconnection queue in a constrained market, which is the entire commercial logic behind the approach — but it is a four-year wait for a turbine, not an instant bypass, and the turbine supply chain is now itself a shared, congested resource across every developer making the same bet simultaneously.

The regulatory risk behind-the-meter power carries is not hypothetical either; a specific case demonstrates it in unusual detail. In November 2024, the Federal Energy Regulatory Commission rejected, on a 2–1 vote, an amended interconnection service agreement that would have let the Susquehanna nuclear plant in Pennsylvania send an increased, co-located load directly to an adjacent Amazon Web Services data center, raising the behind-the-meter supply from 300 megawatts to 480 megawatts [6]. The majority’s stated reasoning was narrowly procedural — that grid operator PJM “has not met its burden to show that these provisions are necessary for any interest unique to the interconnection of the Susquehanna” plant — but the dissent framed the stakes as a genuine reliability question, with Chairman Willie Phillips warning that “in failing to accept the agreement, we are rejecting protections that the interconnected transmission owner says will enhance reliability” [6]. That is a live regulatory disagreement inside a single federal order, not a settled question this article can adjudicate. The practical resolution took seven months: Talen and Amazon restructured the arrangement in June 2025 into an eighteen-billion-dollar, seventeen-year power purchase agreement for up to 1,920 megawatts, moving the deal from a behind-the-meter, co-located structure to a grid-connected, front-of-the-meter retail structure, with the energy flowing through the grid rather than directly to the data center [6]. The episode is the clearest available evidence that behind-the-meter speed is not simply an engineering property of a gas turbine; it is contingent on a regulatory treatment that, in this instance, did not hold, and it can cost a developer months of restructuring to discover which treatment actually applies to their project.

On-site renewables plus storage: fast, partial, and accounted for differently

On-site solar or wind paired with battery storage is generally the fastest power-sourcing option to permit and construct at moderate scale, and it is the only one of the three approaches that changes what a facility can honestly claim about the carbon content of its electricity — a point developed fully in the next section. Its engineering limitation is capacity, not construction speed: intermittent generation backed by storage sized for hours, not days, cannot by itself deliver the continuous, high-load-factor power a training cluster needs, which is why on-site renewables function in essentially every deployed case as a supplement to grid or behind-the-meter firm power rather than a replacement for it.

Sizing that supplement correctly is itself an active area of engineering research, not a solved, off-the-shelf calculation. A 2025 framework presented at the SC25 Sustainable Supercomputing Workshop extends an existing computing-and-energy co-simulator with NREL’s System Advisor Model to study exactly this problem: how the sizing and composition of on-site generation and storage jointly affects a data center’s long-term sustainability and power reliability, capturing both the operational emissions from grid draw and the embodied emissions of the hardware itself [9]. That a purpose-built simulation framework is still being developed to answer “how much solar and how much storage” for a given workload and reliability target is itself informative: unlike a gas turbine’s dispatchable output or a grid interconnection’s contracted capacity, an on-site renewables-plus-storage system’s adequacy depends on the joint statistics of weather, workload, and storage state of charge, and getting the sizing wrong shows up as either wasted capital or unmet load rather than as a single obvious failure mode.

A mobile instrumentation cart beside the comparison bench with flow and power meters wired to the sample trays, one bronze-handled probe clamp caught mid-attachment
Figure 5. Comparing these methods honestly means metering every one of them the same way, and that instrument is still being clipped on.Image prompt and art direction by Brecht Corbeel; generation pending.

Carbon accounting depends on where the meter sits

The choice among these three power-sourcing strategies is inseparable from a question that sounds like bookkeeping but is actually substantive: what counts as “carbon-free” power, and on what time resolution.

The Greenhouse Gas Protocol’s Scope 2 Guidance, the dominant global standard for corporate electricity-emissions accounting since its original 2015 publication, defines two accounting methods side by side rather than one [8]. The location-based method assigns a facility the average emissions intensity of its local grid, regardless of what the facility contracted to buy. The market-based method lets a facility claim the emissions profile of electricity it has contractually acquired — through a power purchase agreement, a renewable energy certificate, or on-site generation — provided that instrument meets the standard’s quality criteria, evaluated against annual totals. That annual-total framing is precisely what a 2024 systems-level study by Riepin and Brown identifies as the accounting gap that 24/7 carbon-free energy procurement is designed to close: hourly matching “overcome[s] the limitations of established procurement schemes, such as the temporal mismatch between clean electricity supply and buyers’ demand that is inherent to ‘volumetric’ matching,” and the authors find that hourly-matched procurement commitments have “consistent beneficial effects on participants and the electricity system,” including that “even as grids become cleaner over time, the hourly matching strategy contributes significantly to system-level emissions reduction” in a way annual matching does not automatically deliver [7].

This is where the power-sourcing choice and the carbon-accounting choice become the same decision in practice, and where reasonable analysts weight the trade-offs differently. A datacenter drawing from the grid and holding annual renewable energy certificates can report a low market-based emissions figure under the current Scope 2 Guidance even while drawing fossil generation at the specific hours it runs its heaviest training jobs — a gap the annual accounting window does not see. A datacenter with genuine on-site renewables plus storage, sized and dispatched to actually match its consumption hour by hour, cannot hide behind that gap, but pays for the privilege in the capacity and reliability limitations described above. The accounting standard itself is in motion on exactly this question: the GHG Protocol’s own page records a public consultation period from October 2025 to January 2026 on updates to the Scope 2 Guidance [8], a live signal that the annual-versus-hourly matching question is unresolved at the standards level, not merely in industry commentary. Riepin and Brown’s finding that hourly matching drives real system-level benefit is evidence for tightening the standard; the operational cost and complexity documented in the on-site renewables-plus-storage discussion above is the evidence on the other side. Both are legitimate inputs to a decision this article does not resolve.

Where the two decisions meet

The cooling comparison and the power-sourcing comparison are not independent in practice, because the axis that dominates one often constrains the other at the same site. A greenfield campus with an unconstrained interconnection date and abundant water can reasonably choose immersion cooling paired with grid power and accept both the highest service complexity and the longest queue, because neither constraint binds hard enough to force a compromise. A retrofit into an existing colocation facility with legacy air-handling infrastructure and a tight schedule is a far more natural fit for rear-door heat exchangers — the least disruptive cooling upgrade — paired with behind-the-meter generation or curtailment-enabled grid capacity, because the retrofit cost of immersion or the multi-year wait for firm interconnection would each independently break the schedule. A site in a water-stressed region gains more from choosing a dry-cooler-backed liquid loop than from choosing any particular power source, because — as the WUE discussion above shows — the water decision is made at the facility heat-rejection layer, largely independent of which rack-level cooling technology or which power contract sits upstream of it.

None of these pairings is a rule; each is a description of which constraint happens to bind hardest for a particular combination of site, climate, existing infrastructure, and schedule. A comparison honest about that variation cannot produce a single ranked list, and should not try to.

Where the experts actually disagree

Four disagreements recur across the sources above and deserve to be named rather than resolved. First, on behind-the-meter co-location: FERC’s own commissioners split on whether the Susquehanna arrangement threatened or protected grid reliability, an unsettled tension between letting large loads self-supply quickly and ensuring the arrangement does not shift cost or risk onto other ratepayers [6]. Second, on load flexibility as a grid resource: Norris’s curtailment-enabled headroom framework argues existing infrastructure can absorb far more load than conventional planning assumes, provided new loads accept modest curtailment [3]; NERC’s guideline, drafted by the reliability body responsible when that assumption fails, treats the same flexible, cyclical load profile primarily as a new source of oscillation and ramping risk requiring active mitigation [2]. Both positions come from credentialed experts working from real operating data; they differ on how much confidence to place in forecasting and controlling large-load behavior in real time. Third, on immersion cooling’s practical readiness: ASHRAE simultaneously documents strong commercial momentum and a list of unresolved risks — fluid sealing, warranty interoperability, materials compatibility — serious enough to warrant formal evaluation before deployment [5]. Fourth, on carbon accounting granularity: annual market-based matching and hourly 24/7 matching are both defensible standards under active development, and the GHG Protocol’s own open consultation is the clearest sign the standards body has not yet decided between them [8] [7].

Predictions, with the observations that would falsify them

These are forecasts, explicitly separated from the sourced analysis above. Horizon: 15 August 2029.

One. Rear-door heat exchangers and direct-to-chip cold plates will increasingly appear together in the same rack rather than as competing choices, because neither alone captures enough heat from the highest-density GPU generations. Disconfirmed if single-method deployments (pure air, pure direct-to-chip, or pure immersion, each without a secondary rack-level method) remain the dominant new-build pattern in 2029.

Two. Interconnection agreements that specify a contracted curtailment allowance, priced explicitly against faster queue placement, will become a standard commercial product offered by grid operators, rather than remaining a research finding. Disconfirmed if major grid operators in 2029 still offer only firm-service interconnection with no curtailment-linked fast-track option.

Three. Behind-the-meter co-location structures will be used disproportionately for new-build, greenfield sites rather than for adding load to existing regulated generation assets, because the Talen-Amazon precedent raised the regulatory cost of the latter path specifically. Disconfirmed if co-location amendments to existing utility-owned generation interconnection agreements become routine and uncontested by 2029.

Four. Hourly-matched (24/7) carbon accounting will be formally incorporated as an optional, clearly labeled reporting tier within the GHG Protocol Scope 2 Guidance, existing alongside rather than replacing annual market-based matching. Disconfirmed if the standard is finalized with only one matching window, whether annual or hourly, and no dual-tier option.

None of these predictions requires a technological breakthrough. Each follows from tensions already visible in the guidance, testimony, and regulatory record cited above.

What to take away

Seven approaches, four axes, no single winner. Air cooling is cheap to retrofit and bounded by airflow physics. Direct-to-chip captures heat where it is hottest but needs a facility water loop and enough rack density to justify one. Rear-door heat exchangers are the least disruptive upgrade and the least thermodynamically ambitious. Immersion reaches the highest density and efficiency and demands a service model most operations teams do not yet have. Grid interconnection is the cheapest power per delivered megawatt and the slowest to obtain, though curtailment-enabled headroom research suggests that slowness is partly a design choice. Behind-the-meter generation buys years of queue time at the cost of turbine lead times and a regulatory treatment a live FERC case shows is not yet settled. On-site renewables plus storage deploy fastest and change what a facility can honestly claim about its carbon accounting, but cannot alone carry a training cluster’s continuous load.

Comparing these options honestly means keeping all four axes on the table at once, attributing every claim about performance or savings to the specific study or standard that produced it, and resisting the pull toward a single ranked list. The bench holds four cooling samples and, upstream of it, three ways to power them. It does not hold a verdict.