Equation 10 · The Deployment Envelope: Small Models Where the Power Is Not
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol E_q
is part of the quantity the equation computes from the expression on the right.
Symbol barP
barP is one factor in the product that computes the quantity on the left.
Symbol r
set by the thermal wall, and r falls as the device heats — so a long deliberation is charged at a worsening exchange rate the longer it runs, which is exactly the behaviour the sustained-load measurements show [ 6 , 5 ].
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Deliberation. This is the sharpest inversion in the whole subject. The most reliable modern route to harder reasoning is to spend more computation at inference time — more tokens, more attempts, more search. In a datacentre that compute is available on demand and is charged for. On a battery-powered, thermally capped device it is the single most expensive thing you can do, because energy per query scales with tokens generated: . with mean power , tokens generated and sustained token rate r . Both and r are set by the thermal wall, and r falls as the device heats — so a long deliberation is charged at a worsening exchange rate the longer it…
Read the full surrounding passage
Deliberation. This is the sharpest inversion in the whole subject. The most reliable modern route to harder reasoning is to spend more computation at inference time — more tokens, more attempts, more search. In a datacentre that compute is available on demand and is charged for. On a battery-powered, thermally capped device it is the single most expensive thing you can do, because energy per query scales with tokens generated: . with mean power , tokens generated and sustained token rate r . Both and r are set by the thermal wall, and r falls as the device heats — so a long deliberation is charged at a worsening exchange rate the longer it runs, which is exactly the behaviour the sustained-load measurements show [ 6 , 5 ] . The remedy that works best in the cloud is the one the device can least afford. Any claim that a small model closes a reasoning gap by thinking longer needs to be checked against the device’s sustained rate, not its first-token rate.
Sources cited in the article section
- [19] Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws ↗
- [15] Gemma 3 Technical Report ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Deployment Envelope: Small Models Where the Power Is Not