Glossary

Optimizer state

Beginner

Extra notes the learning method keeps about each number in the model, like which way it has been moving lately.

Novice

Extra values the optimizer keeps per weight. The common Adam optimizer stores two running averages per weight plus a high-precision copy of the weight, so it uses several times more memory than the weights.

Expert

For mixed-precision Adam: fp32 master weights, momentum and variance, 12 bytes per parameter (K=12K = 12 in the ZeRO paper), on top of 2-byte weights and 2-byte gradients: 16 bytes per parameter in total.