Equation 13 · Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol M_kv
v is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol H
H is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol d_h
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol b
b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as . with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…
Read the full surrounding passage
None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as . with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such as multi-query attention, can extend the usable context length by up to 32 times for a fixed memory budget, precisely because they attack the capacity term of that formula without touching the bandwidth term [ 11 ] . Kwon and colleagues showed the complementary move on the software side: treating the KV cache as pageable memory rather than a contiguous reservation removed fragmentation waste and improved serving throughput by 2 to 4 times at matched latency, entirely by using existing capacity more efficiently rather than adding any [ 12 ] .
Sources cited in the surrounding passage
- [11] Efficiently Scaling Transformer Inference ↗
- [12] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations give research context. Read each source to check which claims it supports.
Return to Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall