Batch (batch size)
Beginner
How many requests a chip works on together. Bigger batches keep the hardware busier but make each request wait longer.
Novice
The number of inputs processed together by the same weights. In a matrix multiplication it adds rows to the activation matrix, which gives a weight-stationary array more work per weight load.
Expert
Raises weight reuse and utilization but adds queueing latency; inference services cap it with response-time limits (7 ms at the 99th percentile for one TPU v1 workload).