Scale factor
Beginner
A shared multiplier stored alongside a group of small numbers, telling you how big they really are.
Novice
A number stored once per tensor, row or block that the quantized values are multiplied by to get back the real values. It lets a narrow format cover whatever range the data actually has.
Expert
Real-valued (often FP32) per-tensor or per-channel scales for INT8 and FP8; power-of-two E8M0 scales per 32-element block in MX; E4M3 scales per 16 elements plus an FP32 tensor scale in NVFP4. In a dot product the scales factor out and are applied once per block or per output.
Explained in Number formats (Architectures).
See also: Quantization, Microscaling (MX).