Token
Beginner
A small chunk of text, often a word or part of a word. Language models read and write text one token at a time.
Novice
The unit of text a language model processes: a word, part of a word or a punctuation mark, mapped to a number. How text is split into tokens depends on the model’s tokenizer.
Expert
Throughput and cost are quoted per token. Prefill processes all prompt tokens in parallel; decode emits one token per sequence per forward pass.
Explained in What the workload needs (Architectures).
See also: Prefill, Decode (generation).