← All parts of this equation

Equation 1 · Part 6 · Edge AI Electronics and Sensor Systems in Practice: Quantization, Power Budgets, and Sensor Front Ends

Decoded value

r=S(q−Z),r = S \left( q - Z \right),
r=S(q−Z)r=S(q-Z)

What this part means

Subtract Z from the stored code, then multiply by S. For q = 14, Z = 10, and S = 0.25, the code represents r = 0.25 × (14 − 10) = 1.00.

Its job in the formula

This part belongs to the expression shown above. Read it together with the other parts of the formula.

The passage around this formula

A model trained in 32-bit floating point does not run on a microcontroller with a few hundred kilobytes of RAM and no floating-point unit worth using for anything but the occasional scalar. The standard fix is quantization: representing weights and activations as low-bit integers, most commonly 8-bit, and executing the forward pass using integer arithmetic end to end. The scheme that made this practical for commodity hardware is an affine mapping between a real value and its stored integer, r=S(q−Z)r = S \left( q - Z \right). where r is the real-valued number, q is the stored integer, S > 0 is a scale factor, and Z is an integer zero-point chosen so that real zero maps exactly onto a representable integer…

Read this part in the article →

Learn the underlying idea

Quantization maps many possible numerical values onto a smaller set of codes. It helps a small device store and process a neural network using compact integer arithmetic.

Open the illustrated quantization: from a value to an integer code guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.

Further reading for this equation