Transformer
Beginner
The design behind today’s chatbots. It reads a whole passage at once and works out which words matter to which.
Novice
A neural network built from repeated blocks, each with an attention layer (which mixes information between positions in the input) and a feed-forward layer (two matrix multiplications applied to each position).
Expert
Stacks of multi-head attention and position-wise MLP blocks with residual connections and normalization. Almost all of its FLOPs are GEMMs: the Q/K/V/output projections, the MLP, and attention’s and products.
Explained in What the workload needs (Architectures).
See also: Attention, Large language model (LLM).