Equation 31 · How a Model Actually Gets Small Enough to Run on a Phone
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol E
The expected value operator: the probability-weighted average of the quantity inside its brackets.
Symbol x
x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →Denominator: 12
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
for a signed integer of b bits. An 8-bit integer has 2^8=256 representable levels; a 4-bit integer has 2^4=16 . That sixteen-fold reduction in available codes is the entire cost of quantization in one number. Treating the rounding error as approximately uniform over one quantization step s , its expected squared magnitude is the classical result . so halving the number of bits, which roughly doubles s at fixed range, roughly quadruples the expected squared error per weight. That is why INT8 is usually described as close to free and INT4 is not: the error a network has to absorb does not grow gently as bits are removed, it grows quadratically in the step size, and every bit…
Read the full surrounding passage
for a signed integer of b bits. An 8-bit integer has 2^8=256 representable levels; a 4-bit integer has 2^4=16 . That sixteen-fold reduction in available codes is the entire cost of quantization in one number. Treating the rounding error as approximately uniform over one quantization step s , its expected squared magnitude is the classical result . so halving the number of bits, which roughly doubles s at fixed range, roughly quadruples the expected squared error per weight. That is why INT8 is usually described as close to free and INT4 is not: the error a network has to absorb does not grow gently as bits are removed, it grows quadratically in the step size, and every bit below eight is removed from an already-narrow budget.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to How a Model Actually Gets Small Enough to Run on a Phone