← All parts of this equation

Equation 13 · Part 2 · AI Memory Systems and the Bandwidth Wall in 2035: Scenarios, Signals, and Falsifiable Predictions

Symbol L

MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p
LL

What this part means

the number of layers.

Its job in the formula

L is an input to the expression that computes the quantity on the left.

Where the article explains it

For a model with L layers, HkvH_{\mathrm{kv}} key-value heads, head dimension dhd_h , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p.

The passage around this formula

Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, HkvH_{\mathrm{kv}} key-value heads, head dimension dhd_h , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p. bytes, the leading factor of two accounting for keys and values together. That equation is worth writing out…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.