← All parts of this equation

Equation 13 · Part 7 · Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

≈

bmax⁡≈Mdevice−MweightsMkv(s),b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)} ,
≈

What this part means

Approximately equal to; the equality is not exact.

Its job in the formula

Approximately equal to; the equality is not exact.

The passage around this formula

What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)} . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.