Equation 13 · Part 2 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs
Symbol M_device
What this part means
evice occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Its job in the formula
evice occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Full expression→Symbol M_device→Article meaning
The passage around this formula
What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [2] Efficiently Scaling Transformer Inference ↗
- [3] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations provide research context; check each source for the exact claim it supports.