Glossary

Prefill

Beginner

The first part of answering: the model reads your whole question at once.

Novice

The inference phase that runs the prompt through the model. All prompt tokens are processed together, so the matrix multiplications are large and the chip’s math units stay busy.

Expert

One forward pass over B×LinputB \times L_{\mathrm{input}} tokens that also fills the KV cache. Weight GEMMs have MM = tokens in the batch, so it is usually compute-bound; it sets the time to first token.