Equation 16 · Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol d_h
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such as multi-query attention, can extend the usable context length by up to 32 times for a fixed memory budget, precisely because they attack the capacity term of that formula without touching the bandwidth term [ 11 ] . Kwon and colleagues showed the complementary move on the software side: treating the KV cache as…
Read the full surrounding passage
with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such as multi-query attention, can extend the usable context length by up to 32 times for a fixed memory budget, precisely because they attack the capacity term of that formula without touching the bandwidth term [ 11 ] . Kwon and colleagues showed the complementary move on the software side: treating the KV cache as pageable memory rather than a contiguous reservation removed fragmentation waste and improved serving throughput by 2 to 4 times at matched latency, entirely by using existing capacity more efficiently rather than adding any [ 12 ] .
Sources cited in the surrounding passage
- [11] Efficiently Scaling Transformer Inference ↗
- [12] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
These citations give research context. Read each source to check which claims it supports.
Return to Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall