Glossary

Attention

Beginner

The step where a model compares every word with every other word to decide which ones to pay attention to.

Novice

A layer that scores how relevant each position is to each other position (a matrix multiply of queries by keys), turns the scores into weights with softmax, and takes a weighted mix of the values (another matrix multiply).

Expert

softmax(QKT/d) V\mathrm{softmax}(QK^T/\sqrt{d})\,V per head. Its FLOPs and score matrix grow as the square of sequence length; during decode it reads every cached key and value per new token, which makes it memory-bound.

Explained in What the workload needs (Architectures).

See also: Transformer, KV cache.

All 896 terms →