Batching
Beginner
Serving many requests together so each number fetched from memory gets used for all of them at once.
Novice
Running several inputs through a model in one pass. The weights are fetched once and used for every input in the batch, which raises the work done per byte.
Expert
Raises weight reuse linearly with batch size, at the cost of per-request latency and KV-cache capacity. It does nothing for per-sequence traffic such as the KV cache.
Explained in What the workload needs (Architectures).
See also: Decode (generation), Arithmetic (operational) intensity.