← Mathematical compendium

Published equation contexts

MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p

Why this formula appears here

Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, HkvH_{\mathrm{kv}} key-value heads, head dimension dhd_h , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p. bytes, the leading factor of two accounting for keys and values together. That equation is worth writing out because it exposes the one real lever every mitigation below actually pulls: none of them escapes linear growth in S and B . Each instead shrinks one of the other factors — most consequentially…

Read the full article-specific guide →

Read the representative guide

MKVM_{\mathrm{KV}}

Symbol M_KV

MKM_KV is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p

Equation 13 · Semiconductors

AI Memory Systems and the Bandwidth Wall in 2035: Scenarios, Signals, and Falsifiable Predictions

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, HkvH_{\mathrm{kv}} key-value heads, head dimension dhd_h , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately MKV=2 L Hkv dh S B pM_{\mathrm{KV}} = 2 \, L \, H_{\mathrm{kv}} \, d_h \, S \, B \, p. bytes, the leading factor of two accounting for keys and values together. That equation is worth writing out because it exposes the one real lever every mitigation below actually pulls: none of them escapes linear growth in S and B . Each instead shrinks one of the other factors — most consequentially…

Meanings in this article

  • LL: the number of layers.
  • dhd_h: the head dimension.
  • SS: the sequence length.
  • BB: the batch size.
  • pp: the number of bytes.
Equation guide → · Article →