Equation 14 · How AI Inference Serving Actually Works
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Its footprint follows directly from the shape of the cache. For a model with L layers, key-value attention heads, head dimension , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is . with the factor of two accounting for storing both keys and values. Two things follow immediately from this equation, and both matter more than the equation’s arithmetic itself. It scales linearly with context length, so a conversation twice as long holds twice the cache. And it scales linearly with the batch b — the exact quantity the previous section identified as the only lever available to make a…
Read the full surrounding passage
Its footprint follows directly from the shape of the cache. For a model with L layers, key-value attention heads, head dimension , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is . with the factor of two accounting for storing both keys and values. Two things follow immediately from this equation, and both matter more than the equation’s arithmetic itself. It scales linearly with context length, so a conversation twice as long holds twice the cache. And it scales linearly with the batch b — the exact quantity the previous section identified as the only lever available to make a memory-bound decode step efficient. The key-value cache therefore competes directly, in the same pool of device memory, with the batch size continuous batching is trying to grow. It is not a side cost of serving; it is the resource that sets the ceiling on the technique the previous section depends on.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to How AI Inference Serving Actually Works