Gradient
Beginner
A list of nudges, one for every number in the model, saying which way to change it and by how much.
Novice
For each weight, the rate at which the training error would change if that weight changed. The backward pass computes one gradient per weight, so gradients take as much memory as the weights themselves.
Expert
for every parameter, produced layer by layer in reverse during the backward pass. In data parallelism, replicas average their gradients with an all-reduce (or reduce-scatter) before the optimizer step.
Explained in Why the network looks this way (Systems).
See also: All-reduce, Data parallelism (DP).