Glossary

Gradient

Beginner

A list of nudges, one for every number in the model, saying which way to change it and by how much.

Novice

For each weight, the rate at which the training error would change if that weight changed. The backward pass computes one gradient per weight, so gradients take as much memory as the weights themselves.

Expert

∂loss/∂w\partial \mathrm{loss} / \partial w for every parameter, produced layer by layer in reverse during the backward pass. In data parallelism, replicas average their gradients with an all-reduce (or reduce-scatter) before the optimizer step.