← Back to article

Equation 1 · How AI Inference Serving Actually Works

What does this equation mean?

t+1t+1

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

One position after t in the sequence described by the article. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

tt

Symbol t

a position index.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

That fact is that a request has two phases with different physics, and the split is not incidental to the architecture — it follows directly from what autoregressive generation requires. Prefill consumes the entire prompt in one pass. Every position in the prompt is known in advance, so the accelerator computes attention and feed-forward outputs for all of them at once, as one large, dense matrix multiplication. Decode then produces the answer one token at a time, and each step depends on the one before it: the model cannot compute position t+1 until it has sampled position t , so there is no way to parallelize across output positions the way prefill parallelizes across input positions. The…
Read the full surrounding passage
That fact is that a request has two phases with different physics, and the split is not incidental to the architecture — it follows directly from what autoregressive generation requires. Prefill consumes the entire prompt in one pass. Every position in the prompt is known in advance, so the accelerator computes attention and feed-forward outputs for all of them at once, as one large, dense matrix multiplication. Decode then produces the answer one token at a time, and each step depends on the one before it: the model cannot compute position t+1 until it has sampled position t , so there is no way to parallelize across output positions the way prefill parallelizes across input positions. The later positions simply do not exist yet.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to How AI Inference Serving Actually Works

See this formula across 1 published context →

Browse the mathematical compendium →