Glossary

Data parallelism (DP)

Beginner

Giving every chip its own copy of the model and a different slice of the examples, then having them share what they learned.

Novice

Each accelerator, or group of them, holds a full copy of the model and trains on a different part of each batch. At the end of every step the copies average their gradients so they stay identical.

Expert

Replicas of degree dd exchange gradients once per step with an all-reduce, 2(d−1)/d2(d-1)/d of the gradient buffer per rank. Communication can overlap with the backward pass; memory is not reduced unless state is sharded (ZeRO/FSDP).