Sequence parallelism
Beginner
Splitting a long piece of text into chunks handled by different chips.
Novice
Splitting work along the sequence (the list of tokens) instead of along the weights. In the Megatron sense it splits the layer parts that tensor parallelism leaves whole, such as normalization, so their memory is shared too.
Expert
Megatron sequence parallelism partitions LayerNorm and dropout along inside the TP group, replacing each all-reduce with a reduce-scatter plus all-gather at equal bandwidth and dividing all activations by .
Explained in Why the network looks this way (Systems).
See also: Tensor parallelism, Context parallelism.