Expert parallelism (EP)
Beginner
Putting different specialists of a model on different chips and sending each word to the chip that holds its specialist.
Novice
Placing the experts of a mixture-of-experts layer on different accelerators. Before each expert layer, every accelerator sends its tokens to wherever their chosen experts live, and the results are sent back afterwards.
Expert
Experts sharded over an EP group, usually carved out of the data-parallel dimension. Each MoE layer needs a dispatch and a combine all-to-all forward and again backward, moving about elements per rank each time.
Explained in Why the network looks this way (Systems).
See also: Mixture of experts (MoE), All-to-all.