Glossary

Occupancy

Beginner

How many teams of workers a GPU core is holding at once, compared with how many it could hold.

Novice

The number of warps resident on a GPU core divided by the maximum it supports. More resident warps give the scheduler more choices while others wait on memory.

Expert

Limited by warp slots, registers per thread, shared memory per block and block-count limits. Necessary for latency hiding only up to the Little’s-law requirement; tuned GEMM kernels often run at low occupancy and hide latency with ILP instead.

Explained in SIMT and GPUs (Architectures).

See also: Latency hiding, Register file.

All 896 terms →