← Back to article

Equation 1 · How AI Inference Economics Actually Work

What does this equation mean?

price per token≳Cfixed+CenergyU⋅Tmax⁡+m\text{price per token} \gtrsim \frac{C_{\text{fixed}} + C_{\text{energy}}}{U \cdot T_{\max}} + m

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

CfixedC_{\text{fixed}}

Symbol C_fixed

the amortized hardware and facility cost per unit time.

Understand this part →

CenergyC_{\text{energy}}

Symbol C_energy

the energy cost per unit time.

Understand this part →

UU

Symbol U

U occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

Tmax⁡T_{\max}

Symbol T_max

the chip’s maximum achievable token throughput under ideal batching.

Understand this part →

mm

Symbol m

the margin term.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

Cfixed+CenergyC_{\text{fixed}} + C_{\text{energy}}

Numerator: C_fixed + C_energy

The complete quantity above the fraction bar.

Understand this part →

U⋅Tmax⁡U \cdot T_{\max}

Denominator: U × T_max

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

What the article says around this equation

This is the reason batching and caching matter for price rather than only for latency: a serving system that keeps a GPU at 80% average utilization across a day divides its fixed hourly cost across roughly twice as many tokens as one running at 40%, and can profitably charge roughly half as much per token for the same margin. Independent benchmarking gives some visibility into what utilization current hardware and software combinations can actually achieve under realistic load. MLCommons’ MLPerf Inference benchmark suite tests submitted systems under both a latency-bounded “server” scenario and an unconstrained “offline” scenario designed to maximize throughput through batching, and recent…
Read the full surrounding passage
This is the reason batching and caching matter for price rather than only for latency: a serving system that keeps a GPU at 80% average utilization across a day divides its fixed hourly cost across roughly twice as many tokens as one running at 40%, and can profitably charge roughly half as much per token for the same margin. Independent benchmarking gives some visibility into what utilization current hardware and software combinations can actually achieve under realistic load. MLCommons’ MLPerf Inference benchmark suite tests submitted systems under both a latency-bounded “server” scenario and an unconstrained “offline” scenario designed to maximize throughput through batching, and recent rounds have shown substantial throughput gains attributable specifically to newer accelerator generations and to serving-software improvements like better batching and KV-cache management, rather than to raw chip count alone [ 9 ] . (Fact, attributed to the cited benchmark reporting; MLPerf figures describe controlled benchmark conditions and are not a direct measurement of any specific commercial provider’s live-traffic utilization, which providers do not publish.) price per token≳Cfixed+CenergyU⋅Tmax⁡+m\text{price per token} \gtrsim \frac{C_{\text{fixed}} + C_{\text{energy}}}{U \cdot T_{\max}} + m. Here CfixedC_{\text{fixed}} is the amortized hardware and facility cost per unit time, CenergyC_{\text{energy}} the energy cost per unit time, Tmax⁡T_{\max} the chip’s maximum achievable token throughput under ideal batching, U ∈\in (0,1] the realized utilization fraction, and m a margin term. This is not a model any provider discloses or that this article claims to have measured; it is a bookkeeping identity stated to make one point precisely: U appears in the denominator, so it multiplies with, rather than merely adds to, every other lever in this article. Halving fixed cost through better batching and halving it again through utilization gains compound rather than add. (Analysis: an accounting identity offered to expose the assumption, not a fitted or disclosed cost model.)

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How AI Inference Economics Actually Work

See this formula across 1 published context →

Browse the mathematical compendium →