Equation field guide
Quantization: from a value to an integer code
Quantization maps many possible numerical values onto a smaller set of codes. It helps a small device store and process a neural network using compact integer arithmetic.
Two directions
Encoding goes from an original value r to a code: approximately q = round(r/S) + Z, followed by clipping to the allowed integer range. Decoding goes back from the code to the represented value: r̂ = S(q − Z). The hat on r̂ reminds us this can be an approximation to the original r.
The equation in the article writes the decoded value as r. It is best read as “the real value represented by this code”, not “every possible original value can be recovered exactly”.
Why information can be lost
Values between adjacent grid points are rounded to one of them. If S = 0.25, both 0.88 and 0.94 may map to the same code and decode to 1.0. Values outside the chosen range can also be clipped.
Smaller steps improve precision but cover a smaller total range with the same number of codes. Larger steps cover more range but lose finer differences.
Where this is used
In edge AI, weights and activations are often quantized so inference can use small integer storage and efficient integer operations. A network contains many values, and a scale and zero-point may be set for an entire tensor or for separate channels.