← All parts of this equation

Equation 14 · Part 3 · How AI Inference Serving Actually Works

Symbol h_kv

Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p
hkvh_{kv}

What this part means

the number of key-value attention heads.

Its job in the formula

hkh_kv is an input to the expression that computes the quantity on the left.

Where the article explains it

For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p.

The passage around this formula

Its footprint follows directly from the shape of the cache. For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p. with the factor of two accounting for storing both keys and values. Two things follow…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.