Data parallelism (DP)
Beginner
Giving every chip its own copy of the model and a different slice of the examples, then having them share what they learned.
Novice
Each accelerator, or group of them, holds a full copy of the model and trains on a different part of each batch. At the end of every step the copies average their gradients so they stay identical.
Expert
Replicas of degree exchange gradients once per step with an all-reduce, of the gradient buffer per rank. Communication can overlap with the backward pass; memory is not reduced unless state is sharded (ZeRO/FSDP).
Explained in Why the network looks this way (Systems).
See also: All-reduce, ZeRO / FSDP (sharded data parallelism).