Glossary

Activations (training memory)

Beginner

The in-between results a model produces while working on an example, which it has to keep until it has learned from that example.

Novice

The intermediate outputs of every layer during the forward pass. Training keeps them until the backward pass uses them, so their memory grows with batch size, sequence length and layer count.

Expert

Per transformer layer about sbh (34+5as/h)sbh\,(34 + 5as/h) bytes in 16-bit training (Korthikanti et al.); the attention term grows with s2s^2. Reduced by tensor and sequence parallelism, and traded for compute by recomputation.