← All parts of this equation

Equation 30 · Part 12 · What an AI Accelerator Actually Is: Silicon, Packaging, and the Memory It Can Reach

Numerator: p-1

Tallreduce≈2 (p−1) α+2 p−1p⋅Nβlink,T_{\mathrm{allreduce}} \approx 2\,(p-1)\,\alpha + 2\,\frac{p-1}{p}\cdot\frac{N}{\beta_{\mathrm{link}}},
p−1p-1

What this part means

The complete quantity above the fraction bar.

Its job in the formula

p-1 occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

Their cost has a shape worth internalising. For the standard ring formulation of an all-reduce over p devices and N bytes, the operation decomposes into a reduce-scatter followed by an all-gather, each of p - 1 steps in which every device sends N/p bytes. Total time is approximately Tallreduce≈2 (p−1) α+2 p−1p⋅NβlinkT_{\mathrm{allreduce}} \approx 2\,(p-1)\,\alpha + 2\,\frac{p-1}{p}\cdot\frac{N}{\beta_{\mathrm{link}}}. with α\alpha the per-step latency and βlink\beta_{\mathrm{link}} the per-link bandwidth. The bandwidth term approaches 2N/βlink\beta_{\mathrm{link}} and stops growing with p ; the latency term grows linearly in p . Small, frequent collectives are therefore latency bound and scale badly, while large ones are bandwidth bound and scale well — which is precisely why gradient bucketing and overlapping…

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.