Glossary

Continuous batching

Beginner

Letting users get on and off the shared bus at every step, so nobody waits for the slowest answer to finish.

Novice

A serving method that rebuilds the batch at every decode step: finished requests return immediately and waiting requests join, which keeps the batch full.

Expert

Iteration-level scheduling, introduced in Orca. It needs a KV-cache allocator that tolerates sequences of different, growing lengths, which is what paged KV-cache managers provide.

Explained in The memory wall (Architectures).

See also: Batching, KV cache.

All 896 terms →