Optimizer state
Beginner
Extra notes the learning method keeps about each number in the model, like which way it has been moving lately.
Novice
Extra values the optimizer keeps per weight. The common Adam optimizer stores two running averages per weight plus a high-precision copy of the weight, so it uses several times more memory than the weights.
Expert
For mixed-precision Adam: fp32 master weights, momentum and variance, 12 bytes per parameter ( in the ZeRO paper), on top of 2-byte weights and 2-byte gradients: 16 bytes per parameter in total.
Explained in Why the network looks this way (Systems).
See also: ZeRO / FSDP (sharded data parallelism), Weights (parameters).