Equation 19 · The Physical Plant: Power, Cooling, and Networks in an AI Datacenter
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol T_ring
ing is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol p
p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol N
N occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol B
B occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
The cost structure of a collective is arithmetic, not policy. For a ring AllReduce over p ranks reducing N bytes at per-link bandwidth B with per-hop latency , each rank moves . so bandwidth cost saturates near 2N/B while the latency term grows linearly in p . Two consequences follow. The slowest link sets the pace for every rank, because the operation does not complete until all ranks have contributed. And large jobs avoid large collectives: Meta reports that multi-dimensional parallelism keeps “the number of GPUs in the largest collective to hundreds of GPUs even when running a job that is tens of thousands of GPUs,” which is why their analysis focuses on…
Read the full surrounding passage
The cost structure of a collective is arithmetic, not policy. For a ring AllReduce over p ranks reducing N bytes at per-link bandwidth B with per-hop latency , each rank moves . so bandwidth cost saturates near 2N/B while the latency term grows linearly in p . Two consequences follow. The slowest link sets the pace for every rank, because the operation does not complete until all ranks have contributed. And large jobs avoid large collectives: Meta reports that multi-dimensional parallelism keeps “the number of GPUs in the largest collective to hundreds of GPUs even when running a job that is tens of thousands of GPUs,” which is why their analysis focuses on collectives spanning 16 to 128 GPUs [ 6 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to The Physical Plant: Power, Cooling, and Networks in an AI Datacenter