Symbol b_max
ax is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
Trace one design decision through that structure to see why the tiers behave as they do. Suppose the middle tier’s key–value cache is halved by moving from multi-head to grouped-query attention. Achievable batch size is roughly . so halving the denominator roughly doubles concurrency at a given context length, which roughly halves cost per request at fixed hardware. The quality cost of that change, per the uptraining results, is small and concentrated on tasks that stress long-range attention [ 8 ] . A buyer sees a price cut. What actually happened is a targeted quality trade whose effects appear on a specific class of input.
ax is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →evice occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →eights occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →v occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →s is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →The complete quantity above the fraction bar.
Read this term in its guide →The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 26 · Foundation Models
This equation gives an approximation: it relates the quantities while allowing an approximation.
Trace one design decision through that structure to see why the tiers behave as they do. Suppose the middle tier’s key–value cache is halved by moving from multi-head to grouped-query attention. Achievable batch size is roughly . so halving the denominator roughly doubles concurrency at a given context length, which roughly halves cost per request at fixed hardware. The quality cost of that change, per the uptraining results, is small and concentrated on tasks that stress long-range attention [ 8 ] . A buyer sees a price cut. What actually happened is a targeted quality trade whose effects appear on a specific class of input.
Equation guide → · Article →Equation 13 · Foundation Models
This equation gives an approximation: it relates the quantities while allowing an approximation.
What binds is memory, and specifically the cache. Weights are shared across the batch; the key–value cache is not. Each concurrent request carries its own, and each grows with every token it generates. Achievable batch size is therefore . which falls as contexts lengthen. This is the mechanism behind an effect users notice without explanation: long-context workloads cost disproportionately more, because they crowd out the concurrency that made short-context serving cheap.
Equation guide → · Article →