Training step
Beginner
One round of learning: the model tries a batch of examples, checks how wrong it was, and nudges all its numbers a little.
Novice
One iteration of training: a forward pass over a batch of examples, a backward pass that computes gradients, and an optimizer update of every weight. Large models take hundreds of thousands of steps.
Expert
Forward, backward and optimizer phases over a global batch, possibly split into microbatches whose gradients are accumulated. Synchronous training ends every step with all replicas holding identical weights.
Explained in Why the network looks this way (Systems).
See also: Gradient, Microbatch.