Equation 8 · AI Memory Systems and the Bandwidth Wall in 2035: Scenarios, Signals, and Falsifiable Predictions
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol H_kv
v is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, key-value heads, head dimension , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately
Sources cited in the article section
- [15] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model ↗
- [14] Efficient Streaming Language Models with Attention Sinks ↗
- [16] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models ↗
These citations give research context. Read each source to check which claims it supports.