Glossary

Tensor parallelism

Beginner

Splitting each step of an AI model’s math across several chips. They must swap part-answers many times per step.

Novice

Dividing the weight matrices of each model layer among several accelerators. Each computes part of every layer, then they combine partial results, typically with all-reduce operations, several times per layer.

Expert

Intra-layer model parallelism (Megatron-style column/row splits). Communication is on the critical path, per layer, with message sizes proportional to batch × sequence × hidden size, so it is normally confined to the scale-up domain.