Glossary

ZeRO / FSDP (sharded data parallelism)

Beginner

A way for chips holding copies of a model to split their notes among themselves, instead of each chip keeping all of them.

Novice

A way to save memory in data parallelism: instead of every replica storing the full optimizer state, gradients and weights, each stores only its share and fetches the rest when needed.

Expert

ZeRO stage 1 shards optimizer state, stage 2 adds gradients, stage 3 (PyTorch FSDP) adds weights. Stages 1–2 move the same 2Ψ2\Psi per step as plain DP; stage 3 moves 3Ψ3\Psi, because weights are all-gathered before use.