Equation 6 · How AI Datacenter Interconnects Actually Work
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol V
V is part of the quantity the equation computes from the expression on the right.
Symbol S
S occurs above the fraction bar. The numerator is divided by the entire denominator below it.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The ring algorithm is the one worth understanding mechanically, because it is bandwidth-optimal and because its cost model exposes exactly why topology matters. In a ring all-reduce across p workers, each worker is logically placed on a ring; a gradient buffer of size S is split into p chunks, and each worker simultaneously sends one chunk to its ring-neighbor while receiving a different chunk from its other neighbor, reducing (summing) as chunks arrive. A complete all-reduce takes 2(p-1) such steps, each moving S/p bytes [ 5 ] . The total data volume any one worker sends over the whole operation is: . which is the reason ring all-reduce is called bandwidth-optimal: in the…
Read the full surrounding passage
The ring algorithm is the one worth understanding mechanically, because it is bandwidth-optimal and because its cost model exposes exactly why topology matters. In a ring all-reduce across p workers, each worker is logically placed on a ring; a gradient buffer of size S is split into p chunks, and each worker simultaneously sends one chunk to its ring-neighbor while receiving a different chunk from its other neighbor, reducing (summing) as chunks arrive. A complete all-reduce takes 2(p-1) such steps, each moving S/p bytes [ 5 ] . The total data volume any one worker sends over the whole operation is: . which is the reason ring all-reduce is called bandwidth-optimal: in the limit, total traffic per worker approaches twice the buffer size regardless of how many workers participate, rather than growing with p . What does grow with p is the number of sequential steps, 2(p-1) , and each step’s latency is bounded below by the slowest link and the slowest worker in the ring — which is precisely why NCCL switches to a tree algorithm for smaller messages and larger worker counts, trading some bandwidth efficiency for a p step count instead of a linear one [ 5 ] . This is an analytical property of the algorithm, not a vendor claim; the tree-versus-ring choice NCCL actually makes at runtime, however, is implementation-specific tuning NVIDIA has published as measured behavior on its own systems, and should be read as a vendor-characterized default rather than a universal law [ 5 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How AI Datacenter Interconnects Actually Work