Outliers (in activations)
Beginner
A few numbers in a block of data that are much bigger than all the others.
Novice
Values far larger than the typical value in a tensor. Large language models develop a few activation channels with values up to about twenty times larger than the rest, and they matter for accuracy, so they can’t simply be clipped.
Expert
Systematic large-magnitude features that appear in specific hidden dimensions of transformers (about 0.1% of features, up to ~20× larger, in all layers beyond a few billion parameters). With a per-tensor scale they force a coarse grid on everything else; fixes include mixed-precision decomposition, finer scaling granularity and rotation transforms.
Explained in Number formats (Architectures).
See also: Integer quantization (INT8, INT4), Microscaling (MX).