Equation 30 · Part 1 · What an AI Accelerator Actually Is: Silicon, Packaging, and the Memory It Can Reach
Symbol T_allreduce
What this part means
llreduce is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
llreduce is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol T_allreduce→Article meaning
The passage around this formula
Their cost has a shape worth internalising. For the standard ring formulation of an all-reduce over p devices and N bytes, the operation decomposes into a reduce-scatter followed by an all-gather, each of p - 1 steps in which every device sends N/p bytes. Total time is approximately . with the per-step latency and the per-link bandwidth. The bandwidth term approaches 2N/ and stops growing with p ; the latency term grows linearly in p . Small, frequent collectives are therefore latency bound and scale badly, while large ones are bandwidth bound and scale well — which is precisely why gradient bucketing and overlapping…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [11] NVIDIA Collective Communication Library (NCCL) Documentation ↗
- [9] NVIDIA Hopper Architecture In-Depth ↗
- [3] TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings ↗
These citations provide research context; check each source for the exact claim it supports.