Equation 1 · AI Memory Systems and the Bandwidth Wall in Practice: An Advanced Technical Guide
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol M_KV
V is part of the quantity the equation computes from the expression on the right.
Symbol L
L is one factor in the product that computes the quantity on the left.
Symbol H
H is one factor in the product that computes the quantity on the left.
Symbol d_h
is one factor in the product that computes the quantity on the left.
Symbol b
b is one factor in the product that computes the quantity on the left.
Symbol s
s is one factor in the product that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The key-value cache is the part of a serving system’s memory footprint that a purchasing decision cannot fix, because it does not exist until a request is running. It grows with every generated token, in proportion to context length, batch size, layer count and head dimension, and its lifetime is exactly the lifetime of the request that owns it — arriving and departing unpredictably, at whatever size that request’s conversation happens to reach. A capacity plan sized for the average context length will be wrong for the tail, and a naive contiguous allocator sized for the tail will waste most of its memory on requests that never reach it. with L transformer layers, H key-value heads, head…
Read the full surrounding passage
The key-value cache is the part of a serving system’s memory footprint that a purchasing decision cannot fix, because it does not exist until a request is running. It grows with every generated token, in proportion to context length, batch size, layer count and head dimension, and its lifetime is exactly the lifetime of the request that owns it — arriving and departing unpredictably, at whatever size that request’s conversation happens to reach. A capacity plan sized for the average context length will be wrong for the tail, and a naive contiguous allocator sized for the tail will waste most of its memory on requests that never reach it. with L transformer layers, H key-value heads, head dimension, b batch size, s sequence length and p bytes per stored element, the factor of two counting keys and values separately. Every strategy below is, mechanically, an attempt to shrink one term of this product without shrinking the model’s effective context.
Sources cited in the article section
- [4] Efficient Memory Management for Large Language Model Serving with PagedAttention ↗
- [5] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache ↗
These citations give research context. Read each source to check which claims it supports.
Return to AI Memory Systems and the Bandwidth Wall in Practice: An Advanced Technical Guide