Prefill
Beginner
The first part of answering: the model reads your whole question at once.
Novice
The inference phase that runs the prompt through the model. All prompt tokens are processed together, so the matrix multiplications are large and the chip’s math units stay busy.
Expert
One forward pass over tokens that also fills the KV cache. Weight GEMMs have = tokens in the batch, so it is usually compute-bound; it sets the time to first token.
Explained in What the workload needs (Architectures).
See also: Decode (generation), KV cache.