← Mathematical compendium

Published equation contexts

Hkv⋅dhH_{\mathrm{kv}} \cdot d_h

Why this formula appears here

Peer-reviewed evidence, three different levers. DeepSeek-V2 replaces full multi-head keys and values with a low-rank latent projection — reducing HkvH_{\mathrm{kv}} ⋅\cdot dhd_h to a much smaller compressed dimension — and reports a 93.3 percent reduction in KV-cache size relative to the company’s own prior 67-billion-parameter dense model, while extending supported context to 128,000 tokens [ 15 ] . That shrinks the per-token cost of the equation above. StreamingLLM instead shrinks effective S : it keeps only a small window of recent tokens plus a handful of initial “attention sink” tokens, and reports enabling stable generation over sequences up to four million tokens with up to a 22.2-times…

Read the full article-specific guide →

Read the representative guide

HkvH_{\mathrm{kv}}

Symbol H_kv

HkH_kv is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Hkv⋅dhH_{\mathrm{kv}} \cdot d_h

Equation 17 · Semiconductors

AI Memory Systems and the Bandwidth Wall in 2035: Scenarios, Signals, and Falsifiable Predictions

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Peer-reviewed evidence, three different levers. DeepSeek-V2 replaces full multi-head keys and values with a low-rank latent projection — reducing HkvH_{\mathrm{kv}} ⋅\cdot dhd_h to a much smaller compressed dimension — and reports a 93.3 percent reduction in KV-cache size relative to the company’s own prior 67-billion-parameter dense model, while extending supported context to 128,000 tokens [ 15 ] . That shrinks the per-token cost of the equation above. StreamingLLM instead shrinks effective S : it keeps only a small window of recent tokens plus a handful of initial “attention sink” tokens, and reports enabling stable generation over sequences up to four million tokens with up to a 22.2-times…

Meanings in this article

Equation guide → · Article →