← Mathematical compendium

Published equation contexts

Mkv=2 L nkv dhead s b pM_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p

Why this formula appears here

A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: Mkv=2 L nkv dhead s b pM_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p . for L layers, nkvn_{\mathrm{kv}} key–value heads, head dimension dheadd_{\mathrm{head}} , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.

Read the full article-specific guide →

Read the representative guide

MkvM_{\mathrm{kv}}

Symbol M_kv

MkM_kv is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Mkv=2 L nkv dhead s b p,M_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p ,

Equation 1 · Foundation Models

Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: Mkv=2 L nkv dhead s b pM_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p . for L layers, nkvn_{\mathrm{kv}} key–value heads, head dimension dheadd_{\mathrm{head}} , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.

Meanings in this article

Equation guide → · Article →