← Back to article

Equation 18 · Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall

What does this equation mean?

TT

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the number of tokens. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

TT

Symbol T

the number of tokens.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such as multi-query attention, can extend the usable context length by up to 32 times for a fixed memory budget, precisely because they attack the capacity term of that formula without touching the bandwidth term [ 11 ] . Kwon and colleagues showed the complementary move on the software side: treating the KV cache as…
Read the full surrounding passage
with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such as multi-query attention, can extend the usable context length by up to 32 times for a fixed memory budget, precisely because they attack the capacity term of that formula without touching the bandwidth term [ 11 ] . Kwon and colleagues showed the complementary move on the software side: treating the KV cache as pageable memory rather than a contiguous reservation removed fragmentation waste and improved serving throughput by 2 to 4 times at matched latency, entirely by using existing capacity more efficiently rather than adding any [ 12 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall

Browse the mathematical compendium →