ZeRO / FSDP (sharded data parallelism)
Beginner
A way for chips holding copies of a model to split their notes among themselves, instead of each chip keeping all of them.
Novice
A way to save memory in data parallelism: instead of every replica storing the full optimizer state, gradients and weights, each stores only its share and fetches the rest when needed.
Expert
ZeRO stage 1 shards optimizer state, stage 2 adds gradients, stage 3 (PyTorch FSDP) adds weights. Stages 1–2 move the same per step as plain DP; stage 3 moves , because weights are all-gathered before use.
Explained in Why the network looks this way (Systems).
See also: Data parallelism (DP), Optimizer state.