← Back to article

Equation 7 · How AI Datacenter Interconnects Actually Work

What does this equation mean?

pp

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the rather than growing with. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

pp

Symbol p

the rather than growing with.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

which is the reason ring all-reduce is called bandwidth-optimal: in the limit, total traffic per worker approaches twice the buffer size regardless of how many workers participate, rather than growing with p . What does grow with p is the number of sequential steps, 2(p-1) , and each step’s latency is bounded below by the slowest link and the slowest worker in the ring — which is precisely why NCCL switches to a tree algorithm for smaller messages and larger worker counts, trading some bandwidth efficiency for a log⁡\log p step count instead of a linear one [ 5 ] . This is an analytical property of the algorithm, not a vendor claim; the tree-versus-ring choice NCCL actually makes at runtime,…
Read the full surrounding passage
which is the reason ring all-reduce is called bandwidth-optimal: in the limit, total traffic per worker approaches twice the buffer size regardless of how many workers participate, rather than growing with p . What does grow with p is the number of sequential steps, 2(p-1) , and each step’s latency is bounded below by the slowest link and the slowest worker in the ring — which is precisely why NCCL switches to a tree algorithm for smaller messages and larger worker counts, trading some bandwidth efficiency for a log⁡\log p step count instead of a linear one [ 5 ] . This is an analytical property of the algorithm, not a vendor claim; the tree-versus-ring choice NCCL actually makes at runtime, however, is implementation-specific tuning NVIDIA has published as measured behavior on its own systems, and should be read as a vendor-characterized default rather than a universal law [ 5 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How AI Datacenter Interconnects Actually Work

Browse the mathematical compendium →