Glossary

Batching

Beginner

Serving many requests together so each number fetched from memory gets used for all of them at once.

Novice

Running several inputs through a model in one pass. The weights are fetched once and used for every input in the batch, which raises the work done per byte.

Expert

Raises weight reuse linearly with batch size, at the cost of per-request latency and KV-cache capacity. It does nothing for per-sequence traffic such as the KV cache.