Glossary

Inference scenario

Beginner

The pattern in which questions arrive during an MLPerf test: one at a time, in a steady stream from many users, or all at once in a big pile.

Novice

MLPerf Inference has four: Single stream (one query at a time, latency measured), Multistream (groups of 8), Server (random arrivals, highest rate that still meets a latency limit) and Offline (everything at once, raw throughput).

Expert

Server arrivals follow a Poisson process and the metric is the highest arrival rate at which the tail latency still meets the benchmark’s bound (for LLMs, time to first token and time per output token). Offline needs at least 24,576 samples.

Explained in Comparing chips (Architectures).

See also: Tail latency.

All 896 terms →