Published equation contexts
Why this formula appears here
Its footprint follows directly from the shape of the cache. For a model with L layers, key-value attention heads, head dimension , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is . with the factor of two accounting for storing both keys and values. Two things follow immediately from this equation, and both matter more than the equation’s arithmetic itself. It scales linearly with context length, so a conversation twice as long holds twice the cache. And it scales linearly with the batch b — the exact quantity the previous section identified as the only lever available to make a…
Read the representative guide
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 14 · Inference Economics
How AI Inference Serving Actually Works
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Its footprint follows directly from the shape of the cache. For a model with L layers, key-value attention heads, head dimension , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is . with the factor of two accounting for storing both keys and values. Two things follow immediately from this equation, and both matter more than the equation’s arithmetic itself. It scales linearly with context length, so a conversation twice as long holds twice the cache. And it scales linearly with the batch b — the exact quantity the previous section identified as the only lever available to make a…
Meanings in this article
- : the total resident size.
- : the number of layers.
- : the number of key-value attention heads.
- : the head dimension.
- : the number of positions in each sequence.
- : the number of concurrently served sequences.
- : the bytes stored per element.