Glossary

Batch (batch size)

Beginner

How many requests a chip works on together. Bigger batches keep the hardware busier but make each request wait longer.

Novice

The number of inputs processed together by the same weights. In a matrix multiplication it adds rows to the activation matrix, which gives a weight-stationary array more work per weight load.

Expert

Raises weight reuse and utilization but adds queueing latency; inference services cap it with response-time limits (7 ms at the 99th percentile for one TPU v1 workload).

Explained in Systolic arrays (Architectures).

See also: Array utilization.

All 896 terms →