← All parts of this equation

Equation 1 · Part 7 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

Symbol p

Mkv=2 L nkv dhead s b p,M_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p ,
pp

What this part means

p is an input to the expression that computes the quantity on the left.

Its job in the formula

p is an input to the expression that computes the quantity on the left.

The passage around this formula

A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: Mkv=2 L nkv dhead s b pM_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p . for L layers, nkvn_{\mathrm{kv}} key–value heads, head dimension dheadd_{\mathrm{head}} , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.