Equation 3 · How AI Inference Serving Actually Works
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →Numerator: FLOPs performed
The complete quantity above the fraction bar. FLOPs count floating-point arithmetic operations.
Denominator: bytes read and written
The complete quantity below the fraction bar; it must be nonzero for this division. Bytes measure the data moved or stored.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The standard way to make this precise is a piece of hardware-performance reasoning called the roofline model, and stating it once explains why every mitigation below targets one phase and not the other. Every accelerator has two ceilings: a peak arithmetic rate, in floating-point operations per second, and a peak memory-bandwidth rate, in bytes per second. Which one binds a given piece of work depends on that work’s arithmetic intensity — the ratio of arithmetic performed to bytes moved: . A computation is compute-bound when I exceeds the hardware’s own ratio of FLOPs to bandwidth, and memory-bound when it falls below it. Prefill, processing every position of a prompt…
Read the full surrounding passage
The standard way to make this precise is a piece of hardware-performance reasoning called the roofline model, and stating it once explains why every mitigation below targets one phase and not the other. Every accelerator has two ceilings: a peak arithmetic rate, in floating-point operations per second, and a peak memory-bandwidth rate, in bytes per second. Which one binds a given piece of work depends on that work’s arithmetic intensity — the ratio of arithmetic performed to bytes moved: . A computation is compute-bound when I exceeds the hardware’s own ratio of FLOPs to bandwidth, and memory-bound when it falls below it. Prefill, processing every position of a prompt against the same weight matrices in one pass, reuses each loaded weight across many tokens at once, so its arithmetic intensity is high and it sits on the compute-bound side of the roofline — which is what the 76 percent utilization figure above looks like in practice [ 1 ] . Decode loads the same weights, plus the accumulated key-value cache, to produce a single new token; each byte fetched from memory is used for one position’s worth of arithmetic and then it is gone. Its arithmetic intensity is low by construction, not by implementation quality, so decode sits on the memory-bound side of the roofline no matter how efficient the kernel is. That is the structural reason the 29-millisecond decode figure cannot simply be optimized away the way a compute-bound kernel can: no amount of faster arithmetic helps a step that is waiting on bytes, not FLOPs [ 1 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How AI Inference Serving Actually Works