MFU (model FLOPs utilization)
Beginner
A fair score for how well a training job uses its chips. Count only the math the model truly needs, and compare it with the chips’ top speed.
Novice
Observed tokens per second × the operations each token needs (about 6 per model parameter for training) ÷ the system’s peak operations per second.
Expert
Introduced with PaLM. Unlike hardware FLOPs utilization (HFU), it excludes recomputation (rematerialization) and other implementation-specific work, so it can’t be inflated by doing extra arithmetic. PaLM 540B reported 46.2% MFU and 57.8% HFU.