Mixture of experts (MoE)
Beginner
A model built from many specialist parts, where each word is handled by only a few of the specialists.
Novice
A model whose feed-forward layers are replaced by many parallel ‘experts’. A small router sends each token to a few of them, so the model can have many more parameters without doing more work per token.
Expert
Sparse layers with experts and top- routing; total parameters grow with while FLOPs per token track . Introduces load balancing, token dropping or capacity limits, and all-to-all traffic when experts live on different devices.
Explained in Why the network looks this way (Systems).
See also: Expert parallelism (EP), All-to-all.