Activations (training memory)
Beginner
The in-between results a model produces while working on an example, which it has to keep until it has learned from that example.
Novice
The intermediate outputs of every layer during the forward pass. Training keeps them until the backward pass uses them, so their memory grows with batch size, sequence length and layer count.
Expert
Per transformer layer about bytes in 16-bit training (Korthikanti et al.); the attention term grows with . Reduced by tensor and sequence parallelism, and traded for compute by recomputation.
Explained in Why the network looks this way (Systems).
See also: Activation recomputation (checkpointing).