Every new AI chip comes with a big number on the box: so many trillion math steps per second. That number is real, but it is a , like the highest number on a car’s speedometer. A chip almost never reaches it on a real job.
Comparing AI chips works the same way. This chapter asks three questions you can bring to any chip:
- How fast is it on real work? Not the box number, but what it really gets done.
- How much power does it use? Count the whole computer and the cooling, not just the chip.
- What does each answer cost? Add the price and the power bill. Then divide by all the work it does.
You will also meet , a set of shared tests that lets companies compare machines under the same rules.
The earlier chapters described the main ways AI chips are built. This one is about judging them. A chip’s spec sheet leads with its , the maximum number of operations per second it could perform in theory. On real models, chips deliver a fraction of that. Google’s first TPU, for example, had a peak of 92 trillion operations per second but averaged 21.4 trillion across the six production workloads its designers measured.1
Fair comparison takes three steps, and this chapter covers each:
- Measure delivered performance on a defined task under shared rules. The industry standard is , run by the MLCommons consortium.1011
- Measure power at the right boundary: chip, server, rack or whole building. This gives .
- Add up cost: the purchase price spread over the hardware’s life plus electricity and cooling. This is the , and dividing it by the work done gives cost per unit of work, such as dollars per million generated words.
A final section shows how to read vendor claims critically, because the same chip can be made to look very different by choosing what to count.
You know a neural network is mostly matrix multiplication, and the earlier chapters covered how GPUs, systolic arrays, spatial meshes and wafer-scale designs execute it. This chapter is about the measurement problem: turning “chip X has N TFLOPS” into a defensible statement about delivered work per dollar.
The pieces, each made precise below:
- Peak versus delivered. and its training-specific form, , with published values from 21% to 58%, and why inference often sits lower.2
- Benchmark methodology. MLPerf’s divisions, availability categories, inference scenarios, time-to-train scoring and the statistical rules that stop cherry-picking.67
- Power boundaries. Chip TDP, measured chip power, node wall power, rack and facility (), and why “incremental” and “total” perf/W can differ by more than 2×.1
- TCO models. Capex amortization, energy opex, duty cycle and idle power, reduced to dollars per million .
- Claim hygiene. Sparsity-doubled peaks, precision mismatches, batch-size and latency choices, system-size choices, and what MLCommons’ own messaging rules forbid.9
Peak throughput: A’s spec sheet says 2,000 TFLOP/s, B’s 1,000. A looks twice as good.
A chip’s top speed isn’t measured. It is worked out on paper: the number of math parts, times how fast each one works.
Real AI jobs fall short of it. Often the math parts sit idle, waiting for numbers to arrive from memory, like cooks with no ingredients. Sometimes there just isn’t enough work to keep them all busy.
The share of top speed a chip really reaches is called its . Google tested its first AI chip on six real jobs. On average it reached about a quarter of its top speed, and one job reached only 3%.1 Big, carefully tuned training jobs reach about half.2
Peak throughput is arithmetic, not measurement. For a chip built from many identical cores, each with matrix units:
For example, NVIDIA’s A100 has 108 streaming multiprocessors, each doing 1,024 fused multiply-adds per clock in 16-bit formats. A multiply-add counts as 2 , and the boost clock is 1.41 GHz, so trillion FLOP/s (312 TFLOP/s).14 AMD’s MI300X whitepaper gives the same kind of recipe: 304 compute units × 2,048 FLOPs per clock × 2.1 GHz ≈ 1,307 TFLOP/s.15
Delivered throughput is what you measure running a real model, and = delivered ÷ peak. It falls short for three main reasons:
- Memory bandwidth. If a computation does few operations per byte it loads (low ), the chip finishes the math before the next data arrives. The workload chapter’s roofline shows this limit, and generating text one at a time is the classic case (see The memory wall).
- Latency limits. many requests raises utilization, but a service that must answer within a few milliseconds can’t wait to fill a big batch. In Google’s first TPU paper, a 7-millisecond limit held a GPU to 37% of the throughput it could reach with bigger batches.1
- Everything else. Communication between chips, operations that aren’t matrix multiplies, and software overheads all take time during which the matrix units idle.
Published utilization figures span a wide range:
| System and job | Delivered ÷ peak |
|---|---|
| First-generation TPU, average of six inference workloads (2015 hardware) | 21.4 ÷ 92 TOPS ≈ 23% |
| Same chip, best and worst workload | 93% and 3% |
| GPT-3 training on V100 GPUs (as computed by the PaLM authors) | 21.3% |
| Megatron-LM, 1-trillion-parameter model on 3,072 GPUs | 52% |
| PaLM 540B training on 6,144 TPU v4 chips | 46.2% (model FLOPs) |
Sources: Jouppi et al.; Chowdhery et al.; Narayanan et al.123
Peak is at a stated precision and clock. Two vendor whitepapers show the arithmetic: A100 = 108 SMs × 1,024 FP16 FMA/clock × 2 × 1.41 GHz = 311.9 TFLOP/s;14 MI300X = 304 CUs × 2,048 FP16 FLOPs/clock × 2.1 GHz = 1,307.4 TFLOP/s.15 Note what this assumes: every MAC issues every cycle at the boost clock, with no stall, no non-matrix instruction and no power or thermal throttling.
Utilization: be precise about the ratio
“Utilization” is ambiguous until you name the numerator. PaLM’s authors separated two:2
- Hardware FLOPs utilization (HFU): FLOPs the hardware actually executed ÷ peak. It counts rematerialization (recomputing activations to save memory), so a run can raise HFU by doing more work without training any faster. It also depends on how FLOPs are counted (analytically or with performance counters).
- Model FLOPs utilization (): observed tokens/s × the FLOPs the model mathematically needs per token ÷ peak. For a dense decoder that is about 6N per training token (2N forward, 4N backward) plus attention. Its numerator is a pure throughput measurement.
PaLM 540B ran at 46.2% MFU and 57.8% HFU on 6,144 TPU v4 chips. Recomputed on the same basis, GPT-3 was 21.3%, Gopher 32.5% and Megatron-Turing NLG 530B 30.2%.2 Megatron-LM reported 52% of theoretical peak per GPU at 3,072 GPUs, counted with its own analytical FLOP formula.3 Treat cross-paper comparisons with care: the peak in the denominator, the attention FLOPs and the recomputation accounting all vary.
Why inference often sits lower
Hardware counters on the first TPU make the causes concrete. Across six production workloads it averaged 21.4 of 92 TOPS. Its two convolutional networks reached 86.0 and 14.1 TOPS; its two LSTMs reached 3.7 and 2.8; for the four memory-bound MLP and LSTM workloads, 44–62% of cycles were weight-load stalls.1 Latency adds a second ceiling: under a 7 ms 99th-percentile bound for one MLP, a CPU ran at 42% and a GPU at 37% of the throughput they reached with relaxed latency, because the bound forced batch 16 instead of 64.1 The same two forces (low arithmetic intensity in decode, and limits on ) govern LLM serving today; see The memory wall.
Peak assumes every math unit works every clock tick. Switch to delivered to see real jobs.
If every company tests its chip on its own favorite job, the numbers can’t be compared. So companies and universities agreed on a shared set of tests called . The first round of its tests for running trained models, in 2019, got over 600 results from 14 groups.4
The tests work like a sports league with a rule book:
- Same tasks. Everyone runs the same AI models on the same data.
- Same quality. A machine can’t go faster by giving sloppier answers.
- Rivals check each other. Before results come out, the other companies that entered review them.8
- Clear labels. Each result says if the machine is for sale today, coming soon, or only a lab test.
In some results, everyone must run exactly the same AI model, so only the machines differ. In others, people may change the model too. That shows off clever ideas, but makes fair comparisons harder.10
is a family of benchmark suites maintained by MLCommons, a consortium of companies and universities. The two most watched are MLPerf Training and MLPerf Inference (datacenter and edge). Each round, organizations submit results; the other submitters review them; then everything is published together, with code and logs.8
What gets measured
- Training measures : the wall-clock time for a system to train a reference model until it reaches a fixed quality on held-out data. Measuring time to a quality target, rather than raw samples per second, stops a system from looking fast by learning badly.5
- Inference measures throughput or latency in four that mimic how inference is used: one query at a time (Single stream), groups of 8 (Multistream), random arrivals from many users with a latency limit (Server), and a big batch job (Offline). The results as a whole must reach a quality target set as a percentage of the reference model’s accuracy, such as 99% or 99.9%.6
Divisions and categories
Every result carries two labels. Its says what was allowed to change:
- Closed: the model must be equivalent to the reference. For inference, numbers may be converted to a lower-precision format using a supplied calibration set, but the model can’t be retrained. This is the apples-to-apples hardware comparison.6
- Open: a different model or retraining is allowed, as long as the task is the same. It shows what new ideas can achieve, not which hardware is faster.
Its says whether you can get the system: Available (can be bought or rented now), Preview (must become available about one round later, with results that hold up), or RDI (research, development or internal).8
The rules are public on GitHub and are the real specification; the summaries below follow the versions checked in October 2026.
Inference
- Architecture. A fixed load generator (LoadGen) issues queries to the submitter’s system under test and timestamps the answers; data set and accuracy script are also fixed. Everything behind the interface (framework, kernels, batching, scheduling) is the submitter’s choice.4
- Scenarios and metrics. Single stream: 90th-percentile latency. Multistream: 99th-percentile latency of 8-sample queries. Server: Poisson arrivals; the metric is the highest arrival rate at which the benchmark’s tail-latency bound holds. Offline: one query holding at least 24,576 samples; the metric is throughput. Minimum duration 600 s.6
- LLM latency bounds. For Llama2-70b, the Server bound is a time to first token (TTFT) of 2,000 ms and a time per output token (TPOT) of 200 ms; a separate Interactive category tightens them to 450 ms and 40 ms. Its accuracy target is 99.9% of the FP32 reference’s ROUGE scores, with output length at least 90% of the reference.6
- Closed-division quantization. Any “purely mathematical, reproducible” quantization to any format is allowed if it meets the quality target, uses only the supplied calibration set and is publicly described. No retraining. Datacenter systems must run with ECC enabled on DRAM and HBM.6
- Divisions. Closed, Open, and a Network division for systems whose queries arrive over a network.6
Training
- Metric. Time to train to a quality target, timed from first touching training data. Larger batches can raise throughput yet need more epochs: in the v0.5 round ResNet-50 took about 64 epochs at a 4K minibatch but over 80 at 16K, about 30% more computation.5
- Scoring. Each result is a set of runs (10 for most current benchmarks, 3 for the largest LLMs); the fastest and slowest are dropped and the rest averaged. Running many sets to pick the best is prohibited, and runs that converge faster than the reference convergence points allow are rejected or rescaled.7
- Precision. Closed training allows a pre-approved list of formats (FP32 down to FP8, MXFP4/MXFP6, NVFP4, INT8/INT4 and others) and only “numerically safe” compensations such as scaling. From v6.1, every log must record the lowest precision used in linear layers, attention and communication.7
Availability and review
Available hardware must have pricing (public or on request) and have shipped to at least one third party. A Preview system must be resubmitted as Available in the round after 140 days (or the next round, whichever is later) with performance within 2%, or the Preview result is marked invalid. RDI components can’t reappear as Available until the round after next or 221 days, whichever is longer. Review is by a committee that includes a representative of every submitter, which “operates on balance of interests rather than by avoiding conflict of interest”: competitors audit each other.8
What MLPerf does not fix: system size. Vendors choose how many chips to submit, and as the TPU v4 authors noted, ideally systems would be compared at equal size, cost or power, “but that is not required.”16 Price is not part of any result.
All labels match: a like-for-like comparison of two systems. Now try changing Y’s labels.
Speed is only half the story. A chip that is twice as fast but uses three times the electricity can cost more to run. So engineers also ask how much work a chip does for each bit of electricity. This is its . A watt measures how fast something uses electricity; a phone charger uses about 20.
The tricky part is what to count. Picture rings around the chip:
- The chip alone. The number on the spec sheet.
- The whole computer. The AI chips share a box with other processors, memory and fans. In one well-known AI computer, the AI chips were given only about half of its top power.14
- The building. A datacenter is a building full of computers. In the U.S., cooling and other equipment there add about 40% more, on average.17
The wider the ring, the worse every chip looks, and the more honest the number. A fair comparison uses the same ring for both chips.
is throughput ÷ power, for example tokens per second per watt, which is the same as tokens per joule. Both the top and the bottom of that fraction need care.
Which power?
- TDP is not power use. A chip’s is the heat its cooling must handle, a design limit. MLPerf’s power group lists relying on TDP or power-supply ratings as a myth: they “often grossly overestimate actual power consumption.”11 Measured server power in AI training averages about three-quarters of the rated maximum.17
- The chip is not the system. NVIDIA’s DGX A100 server holds 8 GPUs rated at 400 W each (3,200 W) inside a system rated at 6,500 W maximum.14 Two server processors, a terabyte of memory, network adapters, storage, fans and power-supply losses make up the rest.
- The building multiplies it. is total facility power ÷ IT power. A 2024 Lawrence Berkeley National Laboratory study estimates the U.S. average fell from 1.6 in 2014 to 1.4 in 2023, with large modern facilities lower.17
MLPerf’s power measurements take the middle ring: the whole system at the wall, including compute nodes and the network between them, but not yet the building’s cooling. They report efficiency as samples (or queries) per joule.11
Why the ring matters
Google’s first TPU paper showed how much the boundary changes the answer. Counting the whole server, the TPU server had 17 to 34 times the performance per watt of a CPU server. Subtracting the host server’s power to count only the accelerators (“incremental” performance per watt) raised that to 41 to 83 times.1 Same hardware, same measurements, more than twice the ratio, just by moving the boundary.
Define the boundary before the ratio. From the inside out:
| Boundary | Denominator | Typical pitfall |
|---|---|---|
| Chip, rated | or max board power | A cooling limit; not the draw on your workload |
| Chip, measured | Package (+HBM) power from telemetry | Ignores host, NICs, fans, conversion losses |
| Node, wall | AC power into the server | Ignores switches and storage shared across nodes |
| Rack / cluster | Nodes + fabric switches + storage share | Switch power is hard to attribute to one job |
| Facility | IT power × | PUE is a site average, not a chip property |
Rated versus measured. The TPU v4 paper lists TDP as “N.A.” and instead reports measured minimum/mean/maximum chip-plus-HBM power on production applications (121/170/192 W).16 LBNL’s 2024 report cites measurements of 8-GPU servers rated at 10.2 kW that drew about 7.9 kW (≈78% of rated) in computationally saturated training, idled near 18% of rated, and models annual operation at 70% of rated.17
MLPerf Power’s choices. Datacenter submissions measure full-system power over the timed execution phase with a SPEC-certified analyzer or, at scale, hardware counters plus estimates for switches. Cooling and storage nodes are not yet included. The power group explicitly rejects three shortcuts: isolating accelerator power (“Myth #1”), using TDP or PSU ratings (“Myth #2”), and folding in PUE (“Myth #3”), because PUE describes the building, not the ML system, and can’t be verified across submitters.11
Total versus incremental. For an accelerator added to a host, define and . Jouppi et al. report both: the TPU server at 17–34× a CPU server on total, 41–83× on incremental; the GPU server at 1.2–2.1× and 1.7–2.9×.1 Incremental is the right metric only if the host would exist anyway; in a purpose-built AI server it flatters the accelerator.
Accuracy target interacts with energy. On BERT in MLPerf Inference v3.1 and v4.0 datacenter Offline results, moving from the 99% to the 99.9% accuracy target cut samples per joule by about half on average, though newer quantization methods narrowed the gap.11
Chip: 400 W thermal design power (TDP), a cooling limit rather than measured use. Work per watt looks best here.
In the end, a company wants to know what each answer costs. Two kinds of money go into it:
- Buying the chip. A $25,000 chip used for four years costs about $17 a day, busy or not.
- Running it. Mostly the electricity bill, including the cooling.
Add them up and divide by all the work done. The total is called the . The team behind Google’s first AI chip wrote that when you buy thousands of computers, cost matters more than raw speed.1
This leads to some surprises:
- A slower chip can win if it is much cheaper.
- The same two chips can swap places in a country where electricity costs more.
adds up everything a system costs over its life. For comparing accelerators, a simple model is enough to see the trade-offs. Per hour:
- = purchase price ÷ (useful life in hours). Use the chip’s price plus its share of the server, network and installation.
- Energy opex = (chip power + its share of server power) × PUE × electricity price.
- Cost per unit of work = (capex + opex per hour) ÷ (work done per hour).
A worked example, with made-up but plausible numbers:
| Item | Value |
|---|---|
| Price with server share, 4-year life | $25,000 ÷ 35,040 h = $0.71/h |
| Power: 700 W chip + 500 W server share, PUE 1.3 | 1.56 kW |
| Electricity at $0.10/kWh | $0.16/h |
| Delivered: 500 TFLOP/s on a 70-billion-parameter model | ≈ 3,571 tokens/s ≈ 12.9 million/h |
| Cost per million tokens | ($0.71 + $0.16) ÷ 12.9 ≈ $0.068 |
Here the purchase is about 82% of the cost. That is typical of expensive accelerators at ordinary electricity prices; for reference, U.S. commercial customers paid 14.53 ¢/kWh and industrial customers 9.77 ¢/kWh on average in July 2026.18 Real deployments add networking, buildings, staff and software, and real hardware prices are negotiated and seldom published.1
The example assumes the chip is busy all the time. If it is busy half the time, capex stays the same while the tokens halve, so cost per token nearly doubles. Idle servers still draw power, too: a measured 8-GPU server idled at about 18% of its rated power.17
A per-accelerator hourly model, which is what the simulation below implements:
Here is the depreciation life in years, the busy fraction (duty cycle), the idle draw as a fraction of loaded draw, and utilization. for decoding a dense -parameter model. LBNL cites a measured 8-GPU node idling at about 18% of rated power and puts loaded draw at roughly 0.7–0.8 of rated, which makes about 0.25; the simulation uses 0.2.17
Three structural consequences:
- Utilization and duty cycle divide capex directly. for the capex term. Halving utilization through a software regression costs as much as doubling the price.
- Energy share . With $25–40k capex, ~1.2–1.6 kW per accelerator, PUE 1.3 and $0.10/kWh it is roughly 12–23% (15% for chip A and 18% for chip B in the “Faster on paper” preset). It approaches half only when capex per watt-year is small: cheap chips, long lives, or electricity well above $0.30/kWh.
- The ranking can flip with site parameters. A chip that is faster but power-hungry wins where power is cheap and loses where it is dear; try the “Power bill” preset and then lower the electricity price.
What the simple model omits matters at scale: network fabric and switches per accelerator (see Scale-out networking), facility capex per kW of critical power, failure and repair, software engineering to reach the assumed U, and financing. Power-limited sites change the objective entirely: when megawatts rather than dollars are the binding constraint, work per facility watt (the system-level perf/W) becomes the figure of merit.
24 busy hours: $17.12 capex + $3.74 energy = $20.87 for 309 M tokens → $0.068 per million. Energy is 18%.
Companies want their chips to look good. There are many honest-sounding ways to pick a flattering number. A few questions catch most of them:
- “With sparsity”? Some top speeds assume half the numbers are zero and can be skipped. That doubles the box number, but most AI models aren’t built that way.14
- Same kind of numbers? Math on small, rough numbers runs faster than math on big, exact ones. Compare like with like.
- How many chips? A result from 72 chips will beat one from 8. Check the count.
- Can you buy it? A lab test is a promise, not a product.
Five checks catch most misleading comparisons:
- Dense or sparse peak? Many spec sheets list a second, doubled peak “with sparsity”. It assumes : in every group of four weights, at least two are zero. Both NVIDIA’s A100 whitepaper (312 dense / 624 sparse TFLOPS in BF16) and AMD’s MI300X whitepaper (1,307 / 2,615) list both.1415 A dense model runs at the dense rate. The sparsity chapter covers when the pattern pays off.
- Which precision? Halving the bits roughly doubles peak throughput on most matrix units, so an FP8 number is not comparable with a BF16 number.15 (See Number formats.) The fair question is which precision your model can use without losing accuracy. MLPerf handles this by fixing the accuracy target, not the format.6
- Which batch size and latency? Throughput quoted at a huge batch may be impossible under the response-time limit your service needs. Prefer results that state a latency bound, like MLPerf’s Server scenario.
- How many chips, and what scale? MLPerf lets submitters choose system size, and its guidelines require comparisons to state chip count. Dividing a result by the chip count gives a derived number, not an official MLPerf metric.9
- Whose power, and whose price? Perf/W against TDP, or a cost comparison using list prices, is the claimant’s calculation. MLCommons forbids submitters from deriving perf/W from TDP or power-supply ratings.9
A sixth habit: look for the same benchmark from several submitters using the same chip. In MLPerf the spread between them, which comes from servers and software, can be as large as the gap between rival chips (see “By the numbers”).
A checklist, with what each trick does to the number:
| Claim pattern | Effect | Check |
|---|---|---|
| Peak “with sparsity” | 2× the dense rate (2:4 structured) | Is the model pruned to 2:4 at equal accuracy? Memory traffic doesn’t halve. |
| Lowest precision listed | FP8 = 2× BF16 on MI300X; INT4 = 4× BF16 on A100 | Same accuracy target? Accumulation precision? Which layers really run there? |
| Throughput at max batch | Raises U toward compute-bound | Stated TTFT/TPOT or p99 bound? Compare Server, not Offline. |
| “Per chip” from a big system | Hides scaling losses, or gains from more aggregate memory | Compare at equal system size; per-chip is a derived metric. |
| Perf/W on TDP | Mixes a thermal limit with a measured result, and leaves out the host | Measured wall power with a stated boundary. |
| Geometric mean across workloads | Weights every workload equally | Does the mix match yours? |
Two concrete headline numbers show the pattern. The DGX A100 table lists “5 (GPU Tensor PFLOP)” for 8 GPUs: 8 × 624 TFLOPS, the sparse FP16 rate; the dense figure is 2.5 PFLOPS.14 The MI300X table’s largest entry is 5,229.8 TFLOPS: FP8 with sparsity, 4× its dense BF16 rate.15 Both are accurately labeled in the whitepapers; the distortion happens when one vendor’s sparse FP8 figure is set beside another’s dense BF16.
On means: the first TPU paper reported the TPU die 14.5× a contemporary CPU die by geometric mean over six applications, but 29.2× when weighted by the actual production mix; for the GPU the two means were 1.1× and 1.9×.1 Neither is wrong; they answer different questions. The same paper rebutted, in its fallacies section, the claim that its competitors would have looked better if run differently, which is a reminder that vendor studies, this one included, are written by interested parties.
MLCommons’ messaging guidelines make several of these rules binding for MLPerf results: comparisons must identify any difference in version, division, category, verified status, scenario or chip count; MLPerf results may not be compared with non-MLPerf results; derived metrics (cost, perf/W from anything but measured MLPerf power) must be marked as such.9
The headline as printed, correctly labeled in the whitepaper. Apply each check to get a number you can compare.
Below are two made-up chips, A and B. Chip A has the bigger number on the box. In the bars, the outline is each chip’s top speed and the filled part is its real speed. Then look at the cost per million tokens (a token is a piece of a word). Try the buttons at the top. Then drag the sliders and see if you can make chip A win.
The calculator compares two hypothetical accelerators serving a 70-billion-parameter language model. Pick a chip to edit, then set its peak, utilization, power, server overhead and price; the datacenter settings apply to both. Things to try:
- In “Faster on paper”, how much utilization does A need before it beats B on cost?
- In “Power bill”, lower the electricity price and PUE. At what price does A win?
- In “Pricey chip”, change the lifetime. Does a longer life rescue A?
The model is the hourly TCO equation from the previous section, with and idle draw at 20% of loaded. Expert view adds the duty cycle, energy per token, lifetime capex and energy, and the worked equation for the chip being edited. Try:
- The “Idle fleet” preset (30% busy). How does the capex share change, and does the ranking?
- Find the break-even electricity price in “Power bill”, and compare it with the EIA averages below.
- Make A and B equal except utilization and price. What price ratio offsets a 1.5× utilization advantage?
Everything is illustrative: no real chip’s price or utilization is implied, and the model omits network, facility capex, staff and software.
- First TPU: average delivered of 92 TOPS peak
- 21.4 TOPS
- PaLM 540B training, model FLOPs utilization
- 46.2%
- U.S. average datacenter PUE, 2023
- ≈ 1.4
- 8-GPU server, measured ÷ rated power in training
- ≈ 78%
Sources: Jouppi et al. 2017; Chowdhery et al. 2022; LBNL 2024.1217
What these numbers mean:
- 21.4 of 92: Google’s first AI chip did, on average, less than a quarter of the work its top speed promised.
- 46.2%: one of the best-tuned big training runs ever reported used a bit under half of its chips’ top speed. That counted as excellent.
- 1.4: for every 10 watts the computers use, a typical U.S. datacenter uses about 4 more for cooling and power equipment.
- 78%: even a busy AI computer uses less than the most power its label allows. So the label makes the bill look bigger than it is.
Below are real MLPerf results. Machines built around chips from different companies can land surprisingly close together.
How to read the table: Server ÷ Offline shows what the latency limit costs. The two large 8-accelerator systems lose 2–3% when queries arrive at random and must start within 2 s and stream at 200 ms per token; the smaller 4-card system loses 29%. A system with more headroom can hold large batches without breaking the latency bound.6
Even a careful comparison gives something up:
- Shared tests versus your job. The tests use a few AI models. Yours might behave differently, so the test winner might not be your winner.
- More users versus faster replies. A chip can serve more people per second if each one waits a bit longer.
- Cheapest versus fastest. The cheapest chip per answer might be too slow for a live chat.
- Today versus tomorrow. Better software can make a chip much faster next year. Today’s winner may not last.
- Standard benchmark versus representative workload. MLPerf’s Closed division fixes the model so results compare hardware, but that means it lags behind the newest models, and vendors can tune heavily for the few benchmarks that count. Your own workload is the best benchmark, if you can run it.
- Throughput versus latency. raises tokens per second per chip and lowers cost, but stretches each user’s wait. A cost per token quoted without a latency limit is half a number.
- Cost per token versus capacity. A chip with a lower cost per token but less memory may need more chips to hold a large model, adding network cost and complexity.
- Efficiency versus flexibility. Specialized designs can win on perf/W for the models they were built for, while general-purpose ones keep working when models change. A four-year depreciation life assumes the hardware stays useful that long.
- Measured versus modeled. Power and price are often estimated rather than measured, and a TCO model is only as good as its utilization and duty-cycle assumptions.
- Goodhart’s law on benchmarks. MLPerf’s defenses (reference convergence points, run minimums with fastest and slowest dropped, compliance tests, and peer review by competitors) limit overfitting, but tuning effort concentrates on scored models.78 The same-chip spreads in the v6.1 table are a direct measure of how much the software and system stack moves the result.13
- Scale choice. Single-node results hide interconnect behavior; rack-scale results fold in the fabric, which is often where the real differences lie (see Scale-up fabrics). Compare at the scale you will deploy.
- Accuracy target as a knob. Lower targets permit more aggressive quantization; MLPerf Power’s BERT data show the 99.9% target roughly halving average samples per joule versus 99%.11 Make sure two claims use the same quality bar.
- Power capacity versus capex. When the site’s power budget is the binding constraint, optimize delivered work per facility watt; when capital is binding, optimize $/token. These can select different chips from the same data.
- Failure modes in TCO models. Assuming peak instead of delivered throughput, 100% duty cycle, TDP instead of measured power, and list price instead of system price each bias the answer in a predictable direction; a model that assumes all four can be off by several times.
Batch 8: 323 tokens/s, 24.8 ms per output token. Cost per token is 4.6× the best that meets the limit (batch 256).
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
1. MFU from first principles
For a dense decoder-only transformer with parameters, layers, heads, head dimension and sequence length , PaLM’s Appendix B gives training FLOPs per token as .2 With a system peak of FLOP/s, the theoretical maximum is
and
Worked example from the paper: Megatron-Turing NLG 530B trained at 65,430 tokens/s on 2,240 A100s with 312 peak TFLOP/s each. Ignoring attention:
or 30.2% with attention.2 For inference decode, replace with about per generated token; the same structure applies, but decode is usually bandwidth-bound, so a low MFU there is expected rather than a failure.
2. Early stopping in the Server scenario
A latency percentile estimated from finitely many queries is uncertain. MLPerf Inference treats each query as a Bernoulli trial that meets the latency bound with unknown probability . After over-latency queries, it finds the smallest number of under-latency queries such that a “minimally failing” system (one that meets the bound with probability exactly the target percentile minus a tolerance ) would produce a result this good with probability at most :
equivalently
where is the binomial cumulative distribution and the regularized incomplete beta function, solved by binary search. MLPerf uses and . A system comfortably inside its bound stops early; one close to it must run more queries; one outside it never passes.6 For a 90th-percentile target with 99% confidence and 0.5% margin, the suggested minimum is 23,886 inferences, rounded to 24,576 ().6
p = 0.9, t = 20: stop once h ≥ 305 under-latency queries (n = 325). Observed late share 6.15% vs 10% allowed.
3. Olympic scoring and convergence checks in training
Time to train is a random variable: initialization and data order change how many epochs a run needs. MLPerf Training therefore runs independent seeds ( for most current benchmarks), discards the fastest and slowest, and averages the rest; one non-converging run may be dropped as the slowest, two invalidate the result. Submitters may run runs only if they pick the consecutive runs (by an objective order such as timestamp) whose score is closest to the median, which prevents choosing the luckiest set. Reference convergence points, measured from the reference implementation at several batch sizes, set a floor on epochs to converge; submissions that converge statistically faster are rejected or have their scores scaled.7
4. Why Server throughput falls below Offline
Offline gives the system every sample at once, so it can run the largest batch that fits in memory and keep the pipeline full. Server sends Poisson arrivals at rate ; to keep the 99th-percentile time to first token under its bound, the scheduler must start new requests promptly and cap the batch so that time per output token stays under its bound too. Queueing raises tail latency sharply as utilization approaches 1, so the reported sits where the tail, not the mean, hits the limit; the first TPU paper makes the same point about input queues raising throughput but stretching response time.1 In the v6.1 Llama2-70b data, the two 8-accelerator vendor systems in the table kept 97–98% of Offline throughput (other 8-accelerator servers in the round kept roughly 92–100%), while a 4-card system kept 71%.13
5. A complete TCO comparison, step by step
Using the simulation’s “Faster on paper” preset (all values hypothetical):
| Step | Chip A | Chip B |
|---|---|---|
| Peak × U = delivered (TFLOP/s) | 2,000 × 0.22 = 440 | 1,000 × 0.50 = 500 |
| Tokens/s at 140 GFLOP/token | 3,143 | 3,571 |
| Facility power: (chip + server) × 1.3 | (1,000 + 600) × 1.3 = 2,080 W | (700 + 500) × 1.3 = 1,560 W |
| Capex/h over 4 years | $40,000 ÷ 35,040 = $1.142 | $25,000 ÷ 35,040 = $0.713 |
| Energy/h at $0.10/kWh | $0.208 | $0.156 |
| $ per million tokens | 1.350 ÷ 11.31 = $0.119 | 0.869 ÷ 12.86 = $0.068 |
| System perf/W (delivered TFLOP/s per facility W) | 0.21 | 0.32 |
A has twice B’s peak and 1.4× B’s peak perf/W per chip watt (2.0 vs 1.43 TFLOP/s/W), yet B delivers 14% more tokens per chip, 51% more per facility watt, and costs 43% less per token. Break-even for A on cost requires , or a price near $19,500 at 22% utilization. The arithmetic is trivial; the hard part in practice is measuring and power honestly on your workload.
That’s the end of the AI chapters. You have seen what AI needs from a chip, the main ways chips are built for it, and how to judge them. Next, the FPGA chapters look at chips you can rewire. But no AI chip works alone. Thousands of them are packed into computers, wired together and cooled in giant buildings. The Systems guide picks up from here.
The last group in this guide, FPGAs, trades efficiency for chips you can rewire. Beyond the chip, this chapter kept running into server power, networks between chips, cooling and the facility. Those are the subject of the Systems guide. It starts with packaging and chiplets, moves through the server and the networks that connect thousands of accelerators, explains why the network looks the way it does, and ends with power and cooling, where the PUE in this chapter comes from.
Most of the terms this chapter left as inputs are outputs of system design: server overhead watts, the switch power MLPerf Power finds hard to attribute, the scaling efficiency behind rack-scale results, and the PUE of a liquid-cooled hall. The Systems guide covers them from the package up: Packaging and chiplets, Board and server, Scale-up fabrics, Scale-out networking, Optics, parallelism and Power and cooling.
Accelerator die — This guide: how the die computes, and how to judge it. (Architectures guide)
Q1A chip’s peak is 1,000 TFLOP/s and it delivers 350 TFLOP/s on a model. What is its utilization?
Q2Two MLPerf Inference results are both for the same model. One is in the Closed division, one in the Open division. Why be careful comparing them?
Q3A datacenter has a PUE of 1.5. A server draws 10 kW. How much power does the building draw for it?
Q4A spec sheet lists 624 TFLOPS “with sparsity” and 312 TFLOPS dense. Which should you use to estimate an ordinary dense model?
Sources
Show Hide 18 sources
- In-Datacenter Performance Analysis of a Tensor Processing Unit92 TOPS peak but 21.4 TOPS average across six production workloads (Table 3); 7 ms 99th-percentile limit forces small batches (Table 4); total vs incremental performance/Watt; TCO is the best cost metric and power correlates with it; geometric vs weighted means.
- PaLM: Scaling Language Modeling with PathwaysDefines model FLOPs utilization (MFU) vs hardware FLOPs utilization; PaLM 540B at 46.2% MFU and 57.8% HFU; GPT-3 21.3%, Gopher 32.5%, MT-NLG 30.2%; Appendix B formula and worked example.
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM502 petaFLOP/s on 3072 GPUs, 52% of theoretical peak per GPU.
- MLPerf Inference BenchmarkWhy the four scenarios exist; the LoadGen / system-under-test split; first round in October 2019 with over 600 submissions from 14 organizations.
- MLPerf Training BenchmarkTime-to-train to a quality target as the metric; optimizations that raise throughput can lengthen time to solution; ResNet-50 at a 16K minibatch needs over 80 epochs vs about 64 at 4K.
- MLPerf Inference RulesClosed, Network and Open divisions; quantization and calibration rules; the four scenarios, their metrics and 600 s minimum duration; accuracy targets; Llama2-70b latency limits; early-stopping criterion.
- MLPerf Training RulesClosed vs Open; pre-approved numerical formats; required logging of lowest precision; minimum number of runs; dropping fastest and slowest runs; reference convergence points.
- General MLPerf Submission RulesReview committee of submitters; Available, Preview and RDI categories; Preview systems must become Available after 140 days or by the next round and match within 2%.
- MLPerf Results Messaging GuidelinesComparisons must state differences in version, division, category, scenario and chip count; only MLPerf-measured system power may be used for perf/W; derived metrics such as cost must be labeled as such.
- MLPerf Inference: DatacenterClosed division for apples-to-apples comparisons, Open division for innovation; Available, Preview and RDI; power reported as system power (Server, Offline) or energy per stream.
- MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AIFull-system power, including compute nodes and interconnect but not yet cooling or storage; myths about component power, TDP and PUE; Samples/Joule; 99.9% accuracy averaged half the efficiency of 99% on BERT.
- MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark ResultsPublished 16 September 2026; 30 submitting organizations; new end-to-end RAG and edge agentic benchmarks.
- MLPerf Inference v6.1 results (summary.csv and per-system logs)Llama2-70b (99.9%) Closed, Available, datacenter: tokens/s for each system and scenario, chip counts and weight precisions.
- NVIDIA A100 Tensor Core GPU Architecture (whitepaper)Dense and 2:4-sparse peak rates (312 / 624 TFLOPS BF16); 108 SMs × 1,024 FMA per clock; 400 W TDP; DGX A100 with 8 GPUs, 5 PFLOPS headline and 6,500 W maximum power.
- Introducing AMD CDNA 3 Architecture (white paper)MI300X peak rates dense and with sparsity (1,307.4 / 2,614.9 TFLOPS BF16); 304 CUs at 2,100 MHz and the FLOPs-per-clock formula; 750 W maximum power.
- TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings275 peak TFLOPS (bf16 or int8); measured min/mean/max power instead of TDP; MLPerf lets vendors pick system size; peak FLOPS/s do not predict real performance.
- 2024 United States Data Center Energy Usage ReportPUE definition; U.S. average PUE from 1.6 (2014) to 1.4 (2023); 8-GPU servers rated 10.2 kW measured near 7.9 kW in saturated training; idle about 18% of rated; datacenters used 4.4% of U.S. electricity in 2023.
- Electric Power Monthly, Table 5.6.A: Average Price of Electricity to Ultimate Customers by End-Use SectorU.S. average for July 2026: 14.53 ¢/kWh commercial, 9.77 ¢/kWh industrial.