Speculative decoding
Beginner
A small, quick model guesses the next few words, and the big model checks all the guesses at once.
Novice
An inference trick in which a small draft model proposes several tokens and the large model verifies them in one pass, keeping the ones it agrees with. The output is the same as the large model’s.
Expert
It spends spare compute in a memory-bound regime: verifying k tokens costs about one weight read, so accepted tokens per weight read rise. Gains depend on the draft’s acceptance rate and shrink as batching already uses the compute.
Explained in The memory wall (Architectures).
See also: Decode (generation), Memory-bound.