Glossary

Tail latency

Beginner

How long the slowest few answers take. An app feels slow if even one answer in a hundred is slow.

Novice

A high percentile of response time, such as the 99th percentile: 99% of requests finish faster than this. Services set limits on it, not on the average.

Expert

Batching and deep queues raise throughput but stretch the tail, so a latency bound caps usable batch size. MLPerf Server reports throughput only at a load where the tail bound holds, estimated with an early-stopping statistical test.

Explained in Comparing chips (Architectures).

All 896 terms →