← Back to article

Equation 13 · Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall

What does this equation mean?

Mkv  ≈  2 L H dh b T,M_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

MkvM_{\mathrm{kv}}

Symbol M_kv

MkM_kv is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

LL

Symbol L

the number of layers.

Understand this part →

HH

Symbol H

H is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

dhd_h

Symbol d_h

dhd_h is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

bb

Symbol b

b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

TT

Symbol T

the number of tokens.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as Mkv  ≈  2 L H dh b TM_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T. with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…
Read the full surrounding passage
None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as Mkv  ≈  2 L H dh b TM_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T. with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such as multi-query attention, can extend the usable context length by up to 32 times for a fixed memory budget, precisely because they attack the capacity term of that formula without touching the bandwidth term [ 11 ] . Kwon and colleagues showed the complementary move on the software side: treating the KV cache as pageable memory rather than a contiguous reservation removed fragmentation waste and improved serving throughput by 2 to 4 times at matched latency, entirely by using existing capacity more efficiently rather than adding any [ 12 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall

See this formula across 1 published context →

Browse the mathematical compendium →