← All parts of this equation

Equation 13 · Part 1 · AI Memory Systems and the Bandwidth Wall in 2035: Scenarios, Signals, and Falsifiable Predictions

Symbol M_KV

MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p
MKVM_{\mathrm{KV}}

What this part means

MKM_KV is part of the quantity the equation computes from the expression on the right.

Its job in the formula

MKM_KV is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, HkvH_{\mathrm{kv}} key-value heads, head dimension dhd_h , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p. bytes, the leading factor of two accounting for keys and values together. That equation is worth writing out because it exposes the one real lever every mitigation below actually pulls: none of them escapes linear growth in S and B . Each instead shrinks one of the other factors — most consequentially…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.