← All parts of this equation

Equation 14 · Part 1 · How AI Inference Serving Actually Works

Symbol M_kv

Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p
MkvM_{\mathrm{kv}}

What this part means

the total resident size.

Its job in the formula

MkM_kv is part of the quantity the equation computes from the expression on the right.

Where the article explains it

For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is

The passage around this formula

Its footprint follows directly from the shape of the cache. For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p. with the factor of two accounting for storing both keys and values. Two things follow immediately from this equation, and both matter more than the equation’s arithmetic itself. It scales linearly with context length, so a conversation twice as long holds twice the cache. And it scales linearly with the batch b — the exact quantity the previous section identified as the only lever available to make a…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.