← Back to article

Equation 21 · AI Inference Economics in Practice: An Advanced Technical Guide

What does this equation mean?

Qmax⁡Q_{\max}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the maximum throughput it sustains per hour at full utilization. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Qmax⁡Q_{\max}

Symbol Q_max

the maximum throughput it sustains per hour at full utilization.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Self-hosting is the limiting case, where the commitment term is effectively indefinite and the fixed cost is capital and operating expense rather than an hourly rate — and it inherits the identical discipline. The serving-stack parameters covered earlier are the mechanism by which a fixed hardware footprint’s effective Qmax⁡Q_{\max} is set: vLLM’s guidance that gpumu_memoryuy_utilization , maxnx_numbm_batchedtd_tokens , and maxnx_numsm_seqs are the highest-impact throughput knobs [ 17 ] , and Mooncake’s demonstration that disaggregating the KV-cache tier can shift throughput by a reported multiple in specific scenarios [ 15 ] , are the difference between a fleet’s real utilization ceiling and the number on a…
Read the full surrounding passage
Self-hosting is the limiting case, where the commitment term is effectively indefinite and the fixed cost is capital and operating expense rather than an hourly rate — and it inherits the identical discipline. The serving-stack parameters covered earlier are the mechanism by which a fixed hardware footprint’s effective Qmax⁡Q_{\max} is set: vLLM’s guidance that gpumu_memoryuy_utilization , maxnx_numbm_batchedtd_tokens , and maxnx_numsm_seqs are the highest-impact throughput knobs [ 17 ] , and Mooncake’s demonstration that disaggregating the KV-cache tier can shift throughput by a reported multiple in specific scenarios [ 15 ] , are the difference between a fleet’s real utilization ceiling and the number on a spec sheet. Evaluating self-hosting against a vendor’s on-demand rate without first establishing what the hardware can do compares a vendor’s optimized number against an unoptimized one, and will reliably conclude self-hosting costs more than it needs to.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to AI Inference Economics in Practice: An Advanced Technical Guide

Browse the mathematical compendium →