Equation 13 · Part 1 · Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall
Symbol M_kv
What this part means
v is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
v is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol M_kv→Article meaning
The passage around this formula
None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as . with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [11] Efficiently Scaling Transformer Inference ↗
- [12] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations provide research context; check each source for the exact claim it supports.