← Back to article

Equation 7 · How AI Datacenter Interconnects Actually Work

What does this equation mean?

p−1p-1

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

pp

Symbol p

p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Take the canonical case: a ring all-reduce across p accelerators, each holding a shard of a gradient vector. Mechanically, at the level this article has been building toward, each of the 2(p-1) steps in a ring reduce-scatter-then-all-gather is nothing more exotic than one RDMA WRITE of the kind described above. A collective library such as NCCL posts one work request per hop describing a source range inside the sending GPU’s memory and a destination range inside the next rank’s GPU memory, reachable because both ends registered that memory for GPUDirect RDMA access; the sending NIC’s DMA engine reads the chunk straight out of HBM, frames it as one or more RoCEv2 or native InfiniBand packets,…
Read the full surrounding passage
Take the canonical case: a ring all-reduce across p accelerators, each holding a shard of a gradient vector. Mechanically, at the level this article has been building toward, each of the 2(p-1) steps in a ring reduce-scatter-then-all-gather is nothing more exotic than one RDMA WRITE of the kind described above. A collective library such as NCCL posts one work request per hop describing a source range inside the sending GPU’s memory and a destination range inside the next rank’s GPU memory, reachable because both ends registered that memory for GPUDirect RDMA access; the sending NIC’s DMA engine reads the chunk straight out of HBM, frames it as one or more RoCEv2 or native InfiniBand packets, and the receiving NIC’s DMA engine writes the payload straight into the destination GPU’s memory when it arrives — with the destination GPU never having posted a matching receive, because a WRITE does not require one. Repeated p-1 times to scatter partial sums around the ring and p-1 more times to circulate the fully reduced result, that is the entire mechanism; nothing about it requires any step to be more sophisticated than the doorbell-and-DMA pattern already described.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to How AI Datacenter Interconnects Actually Work

See this formula across 4 published contexts →

Browse the mathematical compendium →