Equation 13 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol b_max
ax is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol M_device
evice occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol M_weights
eights occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol M_kv
v occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol s
s is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Denominator: M_kv(s)
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.
Sources cited in the article section
- [2] Efficiently Scaling Transformer Inference ↗
- [3] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations give research context. Read each source to check which claims it supports.
Return to Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs