Loss scaling
Beginner
Multiplying numbers up before a calculation so that tiny ones don’t round to zero, then dividing back afterward.
Novice
A trick for FP16 training: multiply the loss by a constant (say 8 or 1,024) before backpropagation so every gradient is scaled up by the same amount and stays above FP16’s smallest value. The gradients are divided back down before the weight update.
Expert
Shifts gradient distributions into FP16’s representable window. Dynamic loss scaling raises the scale until an overflow (Inf/NaN) appears, then skips the step and backs off. BF16 rarely needs it; FP8 uses per-tensor scale factors instead, because skipping steps on overflow would happen too often.
Explained in Number formats (Architectures).
See also: FP16 (half precision), Scale factor.