← Back to the equation

Equation field guide

Quantization: from a value to an integer code

Quantization maps many possible numerical values onto a smaller set of codes. It helps a small device store and process a neural network using compact integer arithmetic.

Two directions

Encoding goes from an original value r to a code: approximately q = round(r/S) + Z, followed by clipping to the allowed integer range. Decoding goes back from the code to the represented value: r̂ = S(q − Z). The hat on r̂ reminds us this can be an approximation to the original r.

The equation in the article writes the decoded value as r. It is best read as “the real value represented by this code”, not “every possible original value can be recovered exactly”.

Why information can be lost

Values between adjacent grid points are rounded to one of them. If S = 0.25, both 0.88 and 0.94 may map to the same code and decode to 1.0. Values outside the chosen range can also be clipped.

Smaller steps improve precision but cover a smaller total range with the same number of codes. Larger steps cover more range but lose finer differences.

Where this is used

In edge AI, weights and activations are often quantized so inference can use small integer storage and efficient integer operations. A network contains many values, and a scale and zero-point may be set for an entire tensor or for separate channels.

Sources and further reading