← All parts of this equation

Equation 13 · Part 4 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

Symbol M_kv

bmax⁡≈Mdevice−MweightsMkv(s),b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)} ,
MkvM_{\mathrm{kv}}

What this part means

MkM_kv occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

MkM_kv occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)} . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.