← Mathematical compendium

Published equation contexts

Tallreduce≈2 (p−1) α+2 p−1p⋅NβlinkT_{\mathrm{allreduce}} \approx 2\,(p-1)\,\alpha + 2\,\frac{p-1}{p}\cdot\frac{N}{\beta_{\mathrm{link}}}

Why this formula appears here

Their cost has a shape worth internalising. For the standard ring formulation of an all-reduce over p devices and N bytes, the operation decomposes into a reduce-scatter followed by an all-gather, each of p - 1 steps in which every device sends N/p bytes. Total time is approximately Tallreduce≈2 (p−1) α+2 p−1p⋅NβlinkT_{\mathrm{allreduce}} \approx 2\,(p-1)\,\alpha + 2\,\frac{p-1}{p}\cdot\frac{N}{\beta_{\mathrm{link}}}. with α\alpha the per-step latency and βlink\beta_{\mathrm{link}} the per-link bandwidth. The bandwidth term approaches 2N/βlink\beta_{\mathrm{link}} and stops growing with p ; the latency term grows linearly in p . Small, frequent collectives are therefore latency bound and scale badly, while large ones are bandwidth bound and scale well — which is precisely why gradient bucketing and overlapping…

Read the full article-specific guide →

Read the representative guide

TallreduceT_{\mathrm{allreduce}}

Symbol T_allreduce

TaT_allreduce is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Tallreduce≈2 (p−1) α+2 p−1p⋅Nβlink,T_{\mathrm{allreduce}} \approx 2\,(p-1)\,\alpha + 2\,\frac{p-1}{p}\cdot\frac{N}{\beta_{\mathrm{link}}},

Equation 30 · AI Hardware & Semiconductors

What an AI Accelerator Actually Is: Silicon, Packaging, and the Memory It Can Reach

This equation gives an approximation: it relates the quantities while allowing an approximation.

Their cost has a shape worth internalising. For the standard ring formulation of an all-reduce over p devices and N bytes, the operation decomposes into a reduce-scatter followed by an all-gather, each of p - 1 steps in which every device sends N/p bytes. Total time is approximately Tallreduce≈2 (p−1) α+2 p−1p⋅NβlinkT_{\mathrm{allreduce}} \approx 2\,(p-1)\,\alpha + 2\,\frac{p-1}{p}\cdot\frac{N}{\beta_{\mathrm{link}}}. with α\alpha the per-step latency and βlink\beta_{\mathrm{link}} the per-link bandwidth. The bandwidth term approaches 2N/βlink\beta_{\mathrm{link}} and stops growing with p ; the latency term grows linearly in p . Small, frequent collectives are therefore latency bound and scale badly, while large ones are bandwidth bound and scale well — which is precisely why gradient bucketing and overlapping…

Meanings in this article

Equation guide → · Article →