Equation 1 · Part 6 · Edge AI Electronics and Sensor Systems in Practice: Quantization, Power Budgets, and Sensor Front Ends
Decoded value
What this part means
Subtract Z from the stored code, then multiply by S. For q = 14, Z = 10, and S = 0.25, the code represents r = 0.25 × (14 − 10) = 1.00.
Its job in the formula
This part belongs to the expression shown above. Read it together with the other parts of the formula.
Full expression→Decoded value→Article meaning
The passage around this formula
A model trained in 32-bit floating point does not run on a microcontroller with a few hundred kilobytes of RAM and no floating-point unit worth using for anything but the occasional scalar. The standard fix is quantization: representing weights and activations as low-bit integers, most commonly 8-bit, and executing the forward pass using integer arithmetic end to end. The scheme that made this practical for commodity hardware is an affine mapping between a real value and its stored integer, . where r is the real-valued number, q is the stored integer, S > 0 is a scale factor, and Z is an integer zero-point chosen so that real zero maps exactly onto a representable integer…
Learn the underlying idea
Quantization maps many possible numerical values onto a smaller set of codes. It helps a small device store and process a neural network using compact integer arithmetic.
Open the illustrated quantization: from a value to an integer code guide →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.