Published equation contexts
Why this formula appears here
That fact is that a request has two phases with different physics, and the split is not incidental to the architecture — it follows directly from what autoregressive generation requires. Prefill consumes the entire prompt in one pass. Every position in the prompt is known in advance, so the accelerator computes attention and feed-forward outputs for all of them at once, as one large, dense matrix multiplication. Decode then produces the answer one token at a time, and each step depends on the one before it: the model cannot compute position t+1 until it has sampled position t , so there is no way to parallelize across output positions the way prefill parallelizes across input positions. The…
Read the representative guide
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Inference Economics
How AI Inference Serving Actually Works
One position after t in the sequence described by the article.
That fact is that a request has two phases with different physics, and the split is not incidental to the architecture — it follows directly from what autoregressive generation requires. Prefill consumes the entire prompt in one pass. Every position in the prompt is known in advance, so the accelerator computes attention and feed-forward outputs for all of them at once, as one large, dense matrix multiplication. Decode then produces the answer one token at a time, and each step depends on the one before it: the model cannot compute position t+1 until it has sampled position t , so there is no way to parallelize across output positions the way prefill parallelizes across input positions. The…