← Back to article

Equation 1 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

What does this equation mean?

Mkv=2 L nkv dhead s b p,M_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p ,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operations2 L n_kv d_head s b p
Result or conditionM_kv
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

MkvM_{\mathrm{kv}}

Symbol M_kv

MkM_kv is part of the quantity the equation computes from the expression on the right.

Understand this part →

LL

Symbol L

the number of layers.

Understand this part →

nkvn_{\mathrm{kv}}

Symbol n_kv

nkn_kv is an input to the expression that computes the quantity on the left.

Understand this part →

dheadd_{\mathrm{head}}

Symbol d_head

the head dimension.

Understand this part →

ss

Symbol s

the sequence length.

Understand this part →

bb

Symbol b

the batch size.

Understand this part →

pp

Symbol p

p is an input to the expression that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: Mkv=2 L nkv dhead s b pM_{\mathrm{kv}} = 2\, L \, n_{\mathrm{kv}} \, d_{\mathrm{head}} \, s \, b \, p . for L layers, nkvn_{\mathrm{kv}} key–value heads, head dimension dheadd_{\mathrm{head}} , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

See this formula across 1 published context →

Browse the mathematical compendium →