Equation 13 · Part 1 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs
Symbol b_max
What this part means
ax is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
ax is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol b_max→Article meaning
The passage around this formula
What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [2] Efficiently Scaling Transformer Inference ↗
- [3] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations provide research context; check each source for the exact claim it supports.