Glossary

Sequence parallelism

Beginner

Splitting a long piece of text into chunks handled by different chips.

Novice

Splitting work along the sequence (the list of tokens) instead of along the weights. In the Megatron sense it splits the layer parts that tensor parallelism leaves whole, such as normalization, so their memory is shared too.

Expert

Megatron sequence parallelism partitions LayerNorm and dropout along ss inside the TP group, replacing each all-reduce with a reduce-scatter plus all-gather at equal bandwidth and dividing all activations by tt.