KV cache
Beginner
The model’s notes on everything it has read so far, kept so it doesn’t have to reread the whole conversation for every new word.
Novice
The keys and values that attention computed for every earlier token, stored in memory so each new token only needs its own. It grows with the length of the conversation and the number of users served at once.
Expert
2 × layers × KV heads × head dimension × bytes per token per sequence. It is read in full on every decode step, so at long contexts or large batches it, not the weights, dominates memory traffic.
Explained in What the workload needs (Architectures).
See also: Attention, Decode (generation).