Glossary

Outliers (in activations)

Beginner

A few numbers in a block of data that are much bigger than all the others.

Novice

Values far larger than the typical value in a tensor. Large language models develop a few activation channels with values up to about twenty times larger than the rest, and they matter for accuracy, so they can’t simply be clipped.

Expert

Systematic large-magnitude features that appear in specific hidden dimensions of transformers (about 0.1% of features, up to ~20× larger, in all layers beyond a few billion parameters). With a per-tensor scale they force a coarse grid on everything else; fixes include mixed-precision decomposition, finer scaling granularity and rotation transforms.