Quantization
Beginner
Rounding numbers onto a small set of allowed values so they take fewer bits to store and are cheaper to compute with.
Novice
Converting a tensor from a wide format (usually FP32 or BF16) to a narrow one such as INT8, FP8 or FP4, with a scale factor that maps the tensor’s range onto the format’s range. It introduces rounding error and, if values exceed the range, clipping error.
Expert
and for integers, or a scaled cast for small floats. Choices: format, granularity (tensor, channel, block), calibration of (max, percentile, MSE), and whether it happens after training (PTQ) or is simulated during training (QAT).
Explained in Number formats (Architectures).
See also: Scale factor, Integer quantization (INT8, INT4).