← All parts of this equation

Equation 14 · Part 4 · How AI Inference Serving Actually Works

Symbol d_h

Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p
dhd_h

What this part means

the head dimension.

Its job in the formula

dhd_h is an input to the expression that computes the quantity on the left.

Where the article explains it

For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p.

The passage around this formula

Its footprint follows directly from the shape of the cache. For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is Mkv=2 L hkv dh s b pM_{\mathrm{kv}} = 2\,L\,h_{kv}\,d_h\,s\,b\,p. with the factor of two accounting for storing both keys and values. Two things follow immediately from this equation, and both matter more than…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.