Memory coalescing
Beginner
Fetching data for a whole team in one trip instead of one trip per worker.
Novice
When the 32 threads of a warp read neighboring addresses, the hardware combines their requests into a few large memory transactions. Scattered addresses need many transactions and waste bandwidth.
Expert
Global loads are served in 32-byte sectors; a warp touching 32 consecutive 4-byte words needs 4 sectors, while 32 scattered words can need 32, cutting effective bandwidth up to 8×.