← All parts of this equation

Equation 1 · Part 4 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

Symbol d_head

Mkv=2 L nkv dhead s b p,M_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p ,
dheadd_{\mathrm{head}}

What this part means

the head dimension.

Its job in the formula

dhd_head is an input to the expression that computes the quantity on the left.

Where the article explains it

for L layers, nkvn_{\mathrm{kv}} key–value heads, head dimension dheadd_{\mathrm{head}} , sequence length s , batch size b , and p bytes per element.

The passage around this formula

A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: Mkv=2 L nkv dhead s b pM_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p . for L layers, nkvn_{\mathrm{kv}} key–value heads, head dimension dheadd_{\mathrm{head}} , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.