Training
Beginner
Teaching a neural network by showing it examples and nudging its numbers after each mistake.
Novice
Running examples forward through the model, measuring the error, then running backward to compute how to change every weight, and updating them. It is done once per model, on large clusters, for weeks.
Expert
Forward plus backward is about 3× the forward FLOPs ( per token). It needs memory for weights, gradients, optimizer states and saved activations; large token batches keep the GEMMs compute-bound.