Glossary

Grouped-query attention (GQA)

Beginner

A way to build an AI model so the notes it keeps about the chat are much smaller.

Novice

An attention design in which groups of query heads share one key/value head, so the KV cache shrinks by the group size. Multi-query attention (MQA) is the extreme case with a single shared key/value head.

Expert

With hh query heads and gg KV heads the KV cache shrinks h/gh/g times, and the arithmetic intensity of attention in decode rises to about 2h/g2h/g ÷ bytes per element FLOPs per byte. Llama 3 uses 8 KV heads at every size.

Explained in The memory wall (Architectures).

See also: KV cache.

All 896 terms →