Context parallelism
Beginner
Splitting a very long input, such as a whole book, across chips so no single chip has to hold all of it.
Novice
Dividing the sequence of a very long input across accelerators for every layer, including attention. Each accelerator still needs keys and values from the others, which it receives by passing them around a ring or gathering them.
Expert
Partitions the full sequence across CP ranks; attention needs remote and , exchanged by ring passing overlapped with blockwise compute (Ring Attention) or by all-gather (Llama 3). Used for 100K-token-scale contexts.
Explained in Why the network looks this way (Systems).
See also: Sequence parallelism.