Inference
Beginner
Using a trained neural network to answer a question, recognize a picture or write text.
Novice
Running a trained model forward to get an output. It happens every time someone uses the model, so its total cost over a model’s life can exceed training’s, and it usually has a response-time limit.
Expert
Forward pass only ( FLOPs per token), with latency targets per request. Small batches and the sequential decode loop make it far more often memory-bound than training.