Glossary

Mixed-precision training

Beginner

Training a model with small, fast number formats for most of the math and a bigger format where accuracy really matters.

Novice

A training recipe where matrix multiplications use 16-bit (or 8-bit) inputs, but results are added up in 32-bit and a 32-bit “master” copy of the weights receives the updates.

Expert

Low-precision GEMM inputs, FP32 accumulation, an FP32 master copy of weights and optimizer state, plus loss scaling for FP16 and per-tensor scaling for FP8. Precision is chosen per operation: reductions, normalizations and softmax typically stay in FP32.

Explained in Number formats (Architectures).

See also: Loss scaling, Accumulator.

All 896 terms →