Glossary

Quantization

Beginner

Rounding numbers onto a small set of allowed values so they take fewer bits to store and are cheaper to compute with.

Novice

Converting a tensor from a wide format (usually FP32 or BF16) to a narrow one such as INT8, FP8 or FP4, with a scale factor that maps the tensor’s range onto the format’s range. It introduces rounding error and, if values exceed the range, clipping error.

Expert

q=clamp(round(x/s)+z)q = \mathrm{clamp}(\mathrm{round}(x/s) + z) and x^=s(q−z)\hat{x} = s(q - z) for integers, or a scaled cast for small floats. Choices: format, granularity (tensor, channel, block), calibration of ss (max, percentile, MSE), and whether it happens after training (PTQ) or is simulated during training (QAT).