Equation 4 · How AI Inference Serving Actually Works
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the ratio of arithmetic performed to bytes moved. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A computation is compute-bound when I exceeds the hardware’s own ratio of FLOPs to bandwidth, and memory-bound when it falls below it. Prefill, processing every position of a prompt against the same weight matrices in one pass, reuses each loaded weight across many tokens at once, so its arithmetic intensity is high and it sits on the compute-bound side of the roofline — which is what the 76 percent utilization figure above looks like in practice [ 1 ] . Decode loads the same weights, plus the accumulated key-value cache, to produce a single new token; each byte fetched from memory is used for one position’s worth of arithmetic and then it is gone. Its arithmetic intensity is low by…
Read the full surrounding passage
A computation is compute-bound when I exceeds the hardware’s own ratio of FLOPs to bandwidth, and memory-bound when it falls below it. Prefill, processing every position of a prompt against the same weight matrices in one pass, reuses each loaded weight across many tokens at once, so its arithmetic intensity is high and it sits on the compute-bound side of the roofline — which is what the 76 percent utilization figure above looks like in practice [ 1 ] . Decode loads the same weights, plus the accumulated key-value cache, to produce a single new token; each byte fetched from memory is used for one position’s worth of arithmetic and then it is gone. Its arithmetic intensity is low by construction, not by implementation quality, so decode sits on the memory-bound side of the roofline no matter how efficient the kernel is. That is the structural reason the 29-millisecond decode figure cannot simply be optimized away the way a compute-bound kernel can: no amount of faster arithmetic helps a step that is waiting on bytes, not FLOPs [ 1 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.