Equation 3 · The Network Is the Computer Again
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol p
p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol m
m occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
Denominator: m+p-1
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction.
What the article says around this equation
Pipeline parallelism splits the model by depth, giving each device a contiguous run of layers. The traffic is point-to-point rather than collective — activations forward, gradients backward, between neighbouring stages only — and the volume is modest. Its cost is not bandwidth but idleness. GPipe’s contribution was to split each batch into micro-batches so that stages could work concurrently instead of waiting for a whole batch to traverse the pipeline [ 6 ] . With p stages and m micro-batches, the standard accounting gives a bubble fraction of . which says the only cure for pipeline idleness is more micro-batches in flight, and micro-batches cost memory. Narayanan and…
Read the full surrounding passage
Pipeline parallelism splits the model by depth, giving each device a contiguous run of layers. The traffic is point-to-point rather than collective — activations forward, gradients backward, between neighbouring stages only — and the volume is modest. Its cost is not bandwidth but idleness. GPipe’s contribution was to split each batch into micro-batches so that stages could work concurrently instead of waiting for a whole batch to traverse the pipeline [ 6 ] . With p stages and m micro-batches, the standard accounting gives a bubble fraction of . which says the only cure for pipeline idleness is more micro-batches in flight, and micro-batches cost memory. Narayanan and colleagues composed tensor, pipeline and data parallelism together, reported 502 petaFLOP/s on 3,072 GPUs at 52 percent of theoretical peak, and introduced an interleaved pipeline schedule they report improves throughput by over ten percent at comparable memory [ 5 ] .
Sources cited in the surrounding passage
- [6] GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism ↗
- [5] Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Network Is the Computer Again