Symbol T_ring
ing is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Consider the two canonical all-reduce algorithms. A ring arranges participants in a cycle and performs a reduce-scatter followed by an all-gather, each in p-1 steps carrying n/p bytes: . The bandwidth term converges to 2n as p grows — it stops depending on the number of participants — which is why the ring is the bandwidth-optimal choice for large messages. NVIDIA’s own performance documentation encodes exactly this factor, defining the bus bandwidth of an all-reduce by applying a correction of 2(p-1)/p to the naive size-over-time figure, on the reasoning that an all-reduce requires 2(p-1) data transfers across p links; all-gather, reduce-scatter and all-to-all get a…
ing is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →n is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →p is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →β is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →gamma is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 15 · AI Hardware & Semiconductors
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Consider the two canonical all-reduce algorithms. A ring arranges participants in a cycle and performs a reduce-scatter followed by an all-gather, each in p-1 steps carrying n/p bytes: . The bandwidth term converges to 2n as p grows — it stops depending on the number of participants — which is why the ring is the bandwidth-optimal choice for large messages. NVIDIA’s own performance documentation encodes exactly this factor, defining the bus bandwidth of an all-reduce by applying a correction of 2(p-1)/p to the naive size-over-time figure, on the reasoning that an all-reduce requires 2(p-1) data transfers across p links; all-gather, reduce-scatter and all-to-all get a…