Glossary

Decode (generation)

Beginner

The second part of answering: the model writes its reply one word at a time, and each word depends on the ones before.

Novice

The inference phase that produces output tokens one after another. Each step must read all the model’s weights to produce just one token per user, so it is limited by memory speed.

Expert

A sequential loop of forward passes, one token per sequence each. Weight GEMMs shrink to MM = batch size, so intensity is about batch size (in FLOP/byte at 16-bit) until KV-cache reads take over.