Attention
Beginner
The step where a model compares every word with every other word to decide which ones to pay attention to.
Novice
A layer that scores how relevant each position is to each other position (a matrix multiply of queries by keys), turns the scores into weights with softmax, and takes a weighted mix of the values (another matrix multiply).
Expert
per head. Its FLOPs and score matrix grow as the square of sequence length; during decode it reads every cached key and value per new token, which makes it memory-bound.
Explained in What the workload needs (Architectures).
See also: Transformer, KV cache.