← All parts of this equation

Equation 14 · Part 2 · How AI Inference Serving Actually Works

Symbol L

Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p
LL

What this part means

the number of layers.

Its job in the formula

L is an input to the expression that computes the quantity on the left.

Where the article explains it

For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p.

The passage around this formula

Its footprint follows directly from the shape of the cache. For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p. with the factor of two accounting for storing both keys and values. Two things…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.