Grouped-query attention (GQA)
Beginner
A way to build an AI model so the notes it keeps about the chat are much smaller.
Novice
An attention design in which groups of query heads share one key/value head, so the KV cache shrinks by the group size. Multi-query attention (MQA) is the extreme case with a single shared key/value head.
Expert
With query heads and KV heads the KV cache shrinks times, and the arithmetic intensity of attention in decode rises to about ÷ bytes per element FLOPs per byte. Llama 3 uses 8 KV heads at every size.