Occupancy
Beginner
How many teams of workers a GPU core is holding at once, compared with how many it could hold.
Novice
The number of warps resident on a GPU core divided by the maximum it supports. More resident warps give the scheduler more choices while others wait on memory.
Expert
Limited by warp slots, registers per thread, shared memory per block and block-count limits. Necessary for latency hiding only up to the Little’s-law requirement; tuned GEMM kernels often run at low occupancy and hide latency with ILP instead.