← Back to article

Equation 6 · How AI Datacenter Interconnects Actually Work

What does this equation mean?

V=2(p−1)⋅Sp≈2Sfor large pV = 2(p-1)\cdot\frac{S}{p} \approx 2S \quad \text{for large } p

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withS
Divide byp
This relates toV
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

VV

Symbol V

V is part of the quantity the equation computes from the expression on the right.

Understand this part →

pp

Symbol p

the rather than growing with.

Understand this part →

SS

Symbol S

S occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The ring algorithm is the one worth understanding mechanically, because it is bandwidth-optimal and because its cost model exposes exactly why topology matters. In a ring all-reduce across p workers, each worker is logically placed on a ring; a gradient buffer of size S is split into p chunks, and each worker simultaneously sends one chunk to its ring-neighbor while receiving a different chunk from its other neighbor, reducing (summing) as chunks arrive. A complete all-reduce takes 2(p-1) such steps, each moving S/p bytes [ 5 ] . The total data volume any one worker sends over the whole operation is: V=2(p−1)⋅Sp≈2Sfor large pV = 2(p-1)\cdot\frac{S}{p} \approx 2S \quad \text{for large } p. which is the reason ring all-reduce is called bandwidth-optimal: in the…
Read the full surrounding passage
The ring algorithm is the one worth understanding mechanically, because it is bandwidth-optimal and because its cost model exposes exactly why topology matters. In a ring all-reduce across p workers, each worker is logically placed on a ring; a gradient buffer of size S is split into p chunks, and each worker simultaneously sends one chunk to its ring-neighbor while receiving a different chunk from its other neighbor, reducing (summing) as chunks arrive. A complete all-reduce takes 2(p-1) such steps, each moving S/p bytes [ 5 ] . The total data volume any one worker sends over the whole operation is: V=2(p−1)⋅Sp≈2Sfor large pV = 2(p-1)\cdot\frac{S}{p} \approx 2S \quad \text{for large } p. which is the reason ring all-reduce is called bandwidth-optimal: in the limit, total traffic per worker approaches twice the buffer size regardless of how many workers participate, rather than growing with p . What does grow with p is the number of sequential steps, 2(p-1) , and each step’s latency is bounded below by the slowest link and the slowest worker in the ring — which is precisely why NCCL switches to a tree algorithm for smaller messages and larger worker counts, trading some bandwidth efficiency for a log⁡\log p step count instead of a linear one [ 5 ] . This is an analytical property of the algorithm, not a vendor claim; the tree-versus-ring choice NCCL actually makes at runtime, however, is implementation-specific tuning NVIDIA has published as measured behavior on its own systems, and should be read as a vendor-characterized default rather than a universal law [ 5 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How AI Datacenter Interconnects Actually Work

See this formula across 1 published context →

Browse the mathematical compendium →