Glossary

Activation recomputation (checkpointing)

Beginner

Throwing away some in-between results to save memory, then working them out again when they are needed.

Novice

A memory-saving trick: keep only some layer inputs during the forward pass and recompute the rest just before the backward pass needs them. It costs extra computation.

Expert

Full recomputation stores only each layer’s input (≈2sbh\approx 2sbh bytes) and costs one extra forward pass (about 33% more FLOPs). Selective recomputation re-does only the cheap attention core, removing the 5as/h5as/h term for a few percent of compute.