Activation recomputation (checkpointing)
Beginner
Throwing away some in-between results to save memory, then working them out again when they are needed.
Novice
A memory-saving trick: keep only some layer inputs during the forward pass and recompute the rest just before the backward pass needs them. It costs extra computation.
Expert
Full recomputation stores only each layer’s input ( bytes) and costs one extra forward pass (about 33% more FLOPs). Selective recomputation re-does only the cheap attention core, removing the term for a few percent of compute.
Explained in Why the network looks this way (Systems).
See also: Activations (training memory).