Equation 17 · Part 2 · AI Memory Systems and the Bandwidth Wall in 2035: Scenarios, Signals, and Falsifiable Predictions
Symbol d_h
What this part means
the head dimension.
Its job in the formula
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol d_h→Article meaning
Where the article explains it
For a model with L layers, key-value heads, head dimension , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately = 2 \, L \, \, \, S \, B \, p bytes, the leading factor of two accounting for keys and values together.
The passage around this formula
Peer-reviewed evidence, three different levers. DeepSeek-V2 replaces full multi-head keys and values with a low-rank latent projection — reducing to a much smaller compressed dimension — and reports a 93.3 percent reduction in KV-cache size relative to the company’s own prior 67-billion-parameter dense model, while extending supported context to 128,000 tokens [ 15 ] . That shrinks the per-token cost of the equation above. StreamingLLM…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [15] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model ↗
- [14] Efficient Streaming Language Models with Attention Sinks ↗
- [16] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models ↗
These citations provide research context; check each source for the exact claim it supports.