Published equation contexts
Why this formula appears here
This is the reason batching and caching matter for price rather than only for latency: a serving system that keeps a GPU at 80% average utilization across a day divides its fixed hourly cost across roughly twice as many tokens as one running at 40%, and can profitably charge roughly half as much per token for the same margin. Independent benchmarking gives some visibility into what utilization current hardware and software combinations can actually achieve under realistic load. MLCommons’ MLPerf Inference benchmark suite tests submitted systems under both a latency-bounded “server” scenario and an unconstrained “offline” scenario designed to maximize throughput through batching, and recent…
Read the representative guide
Symbol U
U occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →Symbol T_max
the chip’s maximum achievable token throughput under ideal batching.
Read this term in its guide →Numerator: C_fixed + C_energy
The complete quantity above the fraction bar.
Read this term in its guide →Denominator: U × T_max
The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · AI Industry
How AI Inference Economics Actually Work
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
This is the reason batching and caching matter for price rather than only for latency: a serving system that keeps a GPU at 80% average utilization across a day divides its fixed hourly cost across roughly twice as many tokens as one running at 40%, and can profitably charge roughly half as much per token for the same margin. Independent benchmarking gives some visibility into what utilization current hardware and software combinations can actually achieve under realistic load. MLCommons’ MLPerf Inference benchmark suite tests submitted systems under both a latency-bounded “server” scenario and an unconstrained “offline” scenario designed to maximize throughput through batching, and recent…
Meanings in this article
- : the amortized hardware and facility cost per unit time.
- : the energy cost per unit time.
- : the chip’s maximum achievable token throughput under ideal batching.
- : the margin term.