← All parts of this equation

Equation 26 · Part 7 · Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

≈

bmax⁡≈Mdevice−MweightsMkv(s),b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)},
≈

What this part means

Approximately equal to; the equality is not exact.

Its job in the formula

Approximately equal to; the equality is not exact.

The passage around this formula

Trace one design decision through that structure to see why the tiers behave as they do. Suppose the middle tier’s key–value cache is halved by moving from multi-head to grouped-query attention. Achievable batch size is roughly bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)}. so halving the denominator roughly doubles concurrency at a given context length, which roughly halves cost per request at fixed hardware. The quality cost of that change, per the uptraining results, is small and concentrated on tasks that stress long-range attention [ 8 ] . A buyer sees a price cut. What actually happened is a targeted quality trade whose effects appear on a specific class of input.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.