Glossary

Expert parallelism (EP)

Beginner

Putting different specialists of a model on different chips and sending each word to the chip that holds its specialist.

Novice

Placing the experts of a mixture-of-experts layer on different accelerators. Before each expert layer, every accelerator sends its tokens to wherever their chosen experts live, and the results are sent back afterwards.

Expert

Experts sharded over an EP group, usually carved out of the data-parallel dimension. Each MoE layer needs a dispatch and a combine all-to-all forward and again backward, moving about k bsh (EP−1)/EPk\,bsh\,(\mathrm{EP}-1)/\mathrm{EP} elements per rank each time.