← Mathematical compendium

Published equation contexts

bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)}

Why this formula appears here

Trace one design decision through that structure to see why the tiers behave as they do. Suppose the middle tier’s key–value cache is halved by moving from multi-head to grouped-query attention. Achievable batch size is roughly bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)}. so halving the denominator roughly doubles concurrency at a given context length, which roughly halves cost per request at fixed hardware. The quality cost of that change, per the uptraining results, is small and concentrated on tasks that stress long-range attention [ 8 ] . A buyer sees a price cut. What actually happened is a targeted quality trade whose effects appear on a specific class of input.

Read the full article-specific guide →

Read the representative guide

bmax⁡b_{\max}

Symbol b_max

bmb_max is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
MdeviceM_{\mathrm{device}}

Symbol M_device

MdM_device occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →
MweightsM_{\mathrm{weights}}

Symbol M_weights

MwM_weights occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →
MkvM_{\mathrm{kv}}

Symbol M_kv

MkM_kv occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
ss

Symbol s

s is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
Mdevice−MweightsM_{\mathrm{device}} - M_{\mathrm{weights}}

Numerator: M_device - M_weights

The complete quantity above the fraction bar.

Read this term in its guide →
Mkv(s)M_{\mathrm{kv}}(s)

Denominator: M_kv(s)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (2)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

bmax⁡≈Mdevice−MweightsMkv(s),b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)},

Equation 26 · Foundation Models

Dense, Sparse, and Distilled: Comparing Approaches to Frontier Model Capacity

This equation gives an approximation: it relates the quantities while allowing an approximation.

Trace one design decision through that structure to see why the tiers behave as they do. Suppose the middle tier’s key–value cache is halved by moving from multi-head to grouped-query attention. Achievable batch size is roughly bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)}. so halving the denominator roughly doubles concurrency at a given context length, which roughly halves cost per request at fixed hardware. The quality cost of that change, per the uptraining results, is small and concentrated on tasks that stress long-range attention [ 8 ] . A buyer sees a price cut. What actually happened is a targeted quality trade whose effects appear on a specific class of input.

Equation guide → · Article →
bmax⁡≈Mdevice−MweightsMkv(s),b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)} ,

Equation 13 · Foundation Models

Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs

This equation gives an approximation: it relates the quantities while allowing an approximation.

What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore bmax⁡≈Mdevice−MweightsMkv(s)b_{\max} \approx \frac{M_{\mathrm{device}} - M_{\mathrm{weights}}}{M_{\mathrm{kv}}(s)} . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.

Equation guide → · Article →