Glossary

Mixture of experts (MoE)

Beginner

A model built from many specialist parts, where each word is handled by only a few of the specialists.

Novice

A model whose feed-forward layers are replaced by many parallel ‘experts’. A small router sends each token to a few of them, so the model can have many more parameters without doing more work per token.

Expert

Sparse layers with EE experts and top-kk routing; total parameters grow with EE while FLOPs per token track kk. Introduces load balancing, token dropping or capacity limits, and all-to-all traffic when experts live on different devices.