A neural network is billions of numbers, multiplied and added over and over. A computer stores each number as a row of bits. A bit is a single 0 or 1, like a light switch that is off or on.
How many bits each number gets matters a lot. Fewer bits means less to store, less to move around, and smaller, cheaper math circuits on the chip.
Ordinary programs use 32 bits for a number like 3.14. AI chips now use 16, 8 or even 4. With 4 bits there are only 16 possible patterns of 0s and 1s. That sounds far too few to work, and on its own it is. This chapter is about the tricks that make it work.
What the workload needs showed that AI is dominated by matrix multiplication: long chains of operations. This chapter is about the numbers those operations work on. How a value is encoded decides three costs at once: how much memory a model needs, how many bytes move per operation, and how large and power-hungry each multiplier is.
For decades the default for real numbers was 32-bit (FP32). Around 2017 training moved to 16-bit formats, with FP32 kept for the sensitive parts.1 In 2022 two 8-bit floating-point formats were proposed for deep learning and later standardized.23 In 2023 a group of chip and cloud companies published an open specification for block-scaled formats down to 4 bits, called microscaling (MX).4 Each step roughly halves memory and data movement, and lets a chip of the same size do more math.
The rest of the chapter covers:
- how floating point trades range against precision, and why that split is a design choice;
- BF16 versus FP16, and the mixed-precision recipe that made 16-bit training work;
- FP8, integer formats, and the role of the scale factor;
- microscaling and FP4, where every 32 values share a scale;
- what fewer bits buy in hardware, and why the running sum still needs many bits.
Given that a network is mostly matrix multiplication (What the workload needs), the number format sets three things at once: bytes per operand, which moves the of every kernel; multiplier area and energy, which grow roughly with the square of significand width56; and the numerical error the model has to absorb. The history is a steady walk down that ladder:
- FP32 → 16-bit with mixed precision: 16-bit operands, FP32 accumulation, FP32 master weights, and loss scaling for FP16.1 BF16 removed the need for loss scaling by keeping FP32’s exponent.7
- 16 → 8-bit floats: E4M3 and E5M2, standardized by OCP, used with per-tensor scale factors.23
- Per-tensor → per-block scaling: the OCP MX specification (MXFP8, MXFP6, MXFP4, MXINT8), because a single scale is not enough below 8 bits.48
In parallel, inference has long used integer quantization with INT8 operands and INT32 accumulators.9 This chapter goes through the encodings bit by bit, the scaling schemes that make narrow formats usable, the error each introduces, accumulation precision, and the hardware cost model that motivates all of it. The algorithms section derives the error models.
FP32: 1 sign, 8 exponent, 23 mantissa bits. 2^23 steps between 1 and 2; the reference for size and cost.
Scientific notation, with bits
Scientists write big numbers in a short way. Three million becomes 3 × 1,000,000. The 3 gives the digits, and the “× 1,000,000” gives the size.
Computers do the same thing with bits. It’s called , because the decimal point can float left or right. One bit says plus or minus. A few bits give the size, called the . The rest hold the digits.
This lets the same few bits describe a speck of dust or a galaxy. The catch is that each number keeps only a few exact digits. Small numbers get small rounding errors, and big numbers get big ones.
Range against detail
With a set number of bits, every bit given to the size is taken from the digits. Picking that split is the main choice behind every format in this chapter.
Three fields
A binary number has a sign bit , an field of bits, and a field of bits. For an ordinary (“normal”) number, the value is
The exponent is stored with a bias so the field can be an unsigned number: in FP32, means . The leading “1 +” is not stored at all. Every normal number starts with a 1 in binary, so formats leave it out, a free extra bit called the hidden bit.10 FP32 has 1 sign bit, 8 exponent bits and 23 mantissa bits, so its significand (the hidden 1 plus the mantissa) has 24 bits of precision.10
A worked example
Take the number 6.5. In binary it is 110.1, which is . So , the exponent is 2, and the mantissa bits are 101 followed by zeros. In FP32 the stored exponent is (binary 10000001). In a tiny format with 3 mantissa bits the mantissa would be 101, and 6.5 fits exactly. The number 6.6 would not: the nearest values with 3 mantissa bits at that size are 6.5 and 7.0, so 6.6 rounds to 6.5.
Precision is relative
Between one power of two and the next (a binade, say 4 to 8), the mantissa bits split the interval into equal steps. With 3 mantissa bits, the values between 4 and 8 are 4, 4.5, 5, 5.5, … 7.5; between 8 and 16 they are 8, 9, 10, …. The step doubles every binade, so the relative error stays roughly constant: rounding to the nearest value is wrong by at most about of the number, whatever its size.10 For FP32 that is about , roughly seven significant decimal digits.
Range
The exponent field sets how many binades the format spans, its . FP32’s 8 exponent bits reach from about to .7 Values beyond the top either become infinity or are clamped to the largest value, depending on the format and mode. Values below the smallest normal number become numbers, which give up precision gradually instead of jumping straight to zero.10
Rounding
When a value falls between two representable numbers, the standard rule is : pick the closer one, and on an exact tie pick the one whose last bit is 0. The OCP 8-bit and MX specifications both require this mode for conversions.34
Encoding
An ExMy format has 1 sign bit, exponent bits and mantissa (trailing significand) bits.4 With bias , a normal encoding () has value ; a subnormal () has .3 The hidden bit gives precision .10 Rounding to nearest bounds the relative error of any in-range normal value by (Goldberg’s with ).10
Range, precision and special values
The exponent field buys binades; the mantissa buys resolution within a binade. Two other choices change the count of usable values:
- Reserved codes. IEEE formats spend the all-ones exponent on and NaN. That is a whole binade, which matters little at 32 bits and a lot at 8. OCP E4M3 keeps only S.1111.111 as NaN, which lifts its maximum from 240 to 448 and its range from 17 to 18 binades.2 The MX FP6 and FP4 types reserve nothing at all.4
- Subnormals. They add binades of gradually decreasing precision below the smallest normal. In E4M3 they extend the bottom from to .3 They cost hardware (a leading-zero count and normalization shift), which is why some datapaths flush them to zero.
Two consequences shape the rest of the chapter. First, a float’s error is relative: the absolute step at magnitude is about , so small values are represented as precisely (in relative terms) as large ones, until they fall out of the range. Second, the bias only positions the window; it doesn’t widen it. A tensor whose values span more binades than the format has will lose either its smallest or its largest values, whatever scale is applied. That is why range (exponent bits) and granularity of scaling are the two levers for low-precision formats.
| Format | S/E/M | Bias | Max normal | Min normal | Relative step |
|---|---|---|---|---|---|
| FP32 | 1/8/23 | 127 | |||
| BF16 | 1/8/7 | 127 | |||
| FP16 | 1/5/10 | 15 | 65,504 | ||
| FP8 E4M3 | 1/4/3 | 7 | 448 | 0.125 | |
| FP8 E5M2 | 1/5/2 | 15 | 57,344 | 0.25 | |
| FP4 E2M1 | 1/2/1 | 1 | 6 | 1 | 0.5 |
Values from the BF16 study’s comparison table, the OCP FP8 specification and the OCP MX specification.734
E4M3: normal values from 0.016 to 240 (14 binades), 8 steps per binade. 6.6 rounds to 6.5 (1.5% off).
When AI training moved from 32 bits to 16, there were two ways to cut the bits in half. Both are still used.
, the older one, keeps more digits but can’t reach very big or very tiny numbers. Anything smaller than about 0.00000006 turns into zero.
, invented at Google, keeps the full size range of the 32-bit format and gives up digits instead.5 It’s simply the first half of a 32-bit number.
That matters for training. As a network learns, it makes tiny fixes to its numbers. Many fixes are so small that FP16 turns them into zero, and the learning stalls. BF16 can hold them, which is a big reason it became the favorite for training.
Training still keeps a few things in 32 bits, like the main copy of the model’s numbers. Mixing sizes like this cuts the memory a model needs about in half, without hurting the result.1
FP16: precision without range
IEEE half precision, , has 1 sign, 5 exponent and 10 mantissa bits. Its largest value is 65,504, its smallest normal value about , and its smallest subnormal about .7 That is plenty for weights and activations, but not for the gradients computed during training (the small corrections that tell each weight which way to move). In one published analysis about 5% of weight-gradient values were smaller than and would become zero in FP16; in a detection model, many activation-gradient values fell below FP16’s smallest number, and the network failed to train in FP16 without a fix.1
The 2018 mixed-precision recipe fixed this with three techniques:1
- An FP32 master copy of the weights. Updates are often too small to register on an FP16 weight, so they are applied to an FP32 copy, which is rounded to FP16 for the next step.
- . Multiply the loss by a constant before backpropagation, so every gradient is scaled up by the same factor and lands inside FP16’s range; divide it back out before the update. For one detection model, a factor of just 8 was enough to match FP32.
- FP32 accumulation. Multiply in FP16 but add the products into FP32 sums, converting to FP16 only when writing results to memory.
The result matched FP32 accuracy across many networks and nearly halved memory use.
BF16: range without precision
(“brain floating point”) takes the opposite trade: 1 sign, 8 exponent and 7 mantissa bits. It is simply FP32 with the bottom 16 mantissa bits dropped, so it has FP32’s full range and converting between them is trivial.7 It started as a storage format for shrinking the data exchanged between machines in distributed training, then became a compute format.7 A 2019 study trained image, speech, language and recommendation models in BF16 and matched FP32 in the same number of iterations with no hyperparameter changes, because no loss scaling is needed.7
BF16 is coarse: with 7 mantissa bits, the step between neighboring values is of the number, so rounds back to 1. Two to three significant decimal digits is enough for the inputs of a matrix multiplication, but not for the sum, so BF16 is also paired with FP32 accumulation. Google’s TPUs, for example, multiply in BF16 and accumulate in FP32.5
Two splits of 15 bits
FP16 (E5M10, bias 15) and BF16 (E8M7, bias 127) spend the same 15 non-sign bits differently. FP16 has and spans to 65,504 (about 40 binades including subnormals); BF16 has and spans the whole FP32 normal range, about to .7 FP16 has 8× finer resolution within a binade; BF16 has 254 binades of normal numbers against FP16’s 30. Which one wins depends on whether a tensor’s problem is resolution or range.
- Gradients are range-limited. Micikevicius et al. found weight gradients whose exponents fell below −24 (about 5% of values in one network), and an SSD detector, many of whose activation gradients (recorded in FP32 training) fell below FP16’s smallest subnormal, that diverged in FP16 without scaling; a loss scale of 8 (three binades) restored FP32 accuracy.1 Automatic mixed-precision tools adjust the loss scale at run time by provoking and detecting overflows.2
- Weight updates are resolution-limited. An update (learning rate × gradient) smaller than half an ulp of the weight vanishes under round-to-nearest, which is why the master weights stay in FP32 even with BF16.1
- GEMM inputs tolerate either, as long as the products are summed at higher precision. Both the FP16 recipe and BF16 practice accumulate in FP32.17
BF16’s hardware argument is the multiplier: an 8-bit significand multiplier instead of 11 or 24 bits.7 Google states that a BF16 multiplier is about half the area of an FP16 one and one eighth of an FP32 one, because multiplier size scales with the square of the mantissa width.5 The and ratios agree with that to first order. The conversion from FP32 is a rounding of the low 16 bits; the BF16 study used round-to-nearest-even.7
Gradient 0.0000012. FP16: stored as 0.0000013; BF16: 0.0000012. FP16 has more digits but reaches only down to about 0.00000006.
Eight bits
With 8 bits there are only 256 possible values. AI uses two kinds. FP8 is a tiny floating-point number, with size bits and digit bits.2 INT8 holds only whole numbers, from −127 to 127, evenly spaced like marks on a ruler.
The shared multiplier
Neither one can cover the real numbers in a model by itself. So a whole block of numbers gets one extra number stored with it, called a . Multiply each stored value by it to get the real value. It’s usually picked so the biggest number in the block just fits.
That works well until one number is much bigger than the rest. Chatbot models have a few of these giants, up to about 20 times bigger than the others, and they matter.11 If the scale stretches to fit a giant, the normal numbers get squeezed onto a few ruler marks. Many of them round to zero.
FP8: two variants
In 2022 engineers from NVIDIA, Arm and Intel proposed two 8-bit floating-point encodings, and in 2023 they were published as an Open Compute Project standard, OFP8.23
- E4M3: 4 exponent bits, 3 mantissa bits. Largest value 448, smallest , 18 binades. Recommended for weights and activations.
- E5M2: 5 exponent bits, 2 mantissa bits. Largest value 57,344, smallest , 32 binades. Recommended for gradients, which need range more than precision.23
Even 32 binades isn’t enough to hold every tensor in a network with a fixed bias, so each tensor gets a stored in higher precision. Values are multiplied by the scale before conversion so that the tensor’s largest magnitude lands near the format’s maximum; anything that still overflows is saturated to the maximum rather than turned into infinity.2 With this recipe, FP8 training matched 16-bit results on image, speech and language models up to 175 billion parameters, with no hyperparameter changes.2
Integers and quantization
For running trained models (inference), the long-standing alternative is . Each real value is represented by an integer through a scale and a zero-point : .9 With symmetric quantization , and is typically chosen as for INT8. A small example: if a tensor’s largest magnitude is 2.54, then , the value 0.73 is stored as (a tie, rounded to even), and it comes back as 0.72.
Integers have two advantages. The products of two INT8 values are exact integers that can be summed exactly in a 32-bit integer ,9 and integer multipliers are smaller than floating-point ones (more on that below). An 8-bit quantization workflow published in 2020 kept every network studied within 1% of its floating-point accuracy.12
Rounding error against clipping error
Choosing the scale is a trade-off. A larger scale covers more of the tensor, so fewer values are clipped at the ends, but each step is wider, so every value suffers more rounding error. A smaller scale does the reverse.13 Practical tools choose the range from sample data, using the maximum, a high percentile, or a measure that minimizes information loss.12
Outliers
Large language models make this trade-off much harder. Beyond a few billion parameters, a small number of hidden dimensions carry values up to about 20 times larger than the rest, in every layer. They are only about 0.1% of the features, but setting them to zero destroys the model’s accuracy.11 With one scale per tensor, those outliers set the step size, and the ordinary values are crushed onto a few levels. Two responses followed: keep the outlier dimensions in 16-bit and the remaining 99.9% in 8-bit,11 or use a separate scale for smaller pieces of the tensor (per row, per channel, or per block).13 The second idea, taken to small blocks, is microscaling.
OFP8 in detail
E4M3 has bias 7, no infinities and a single NaN mantissa pattern (S.1111.111), which extends it to 448 and 18 binades; E5M2 has bias 15 and IEEE-style specials (S.11111.00 is ), reaching 57,344 and 32 binades.3 E5M2 is effectively FP16 with 8 fewer mantissa bits, which makes conversion trivial.2 Conversion from wider types rounds to nearest even and then applies one of two overflow modes: saturating (clamp to ) or non-saturating (E4M3 → NaN, E5M2 → ).3 Training recipes use saturation, because the loss-scaling trick of skipping a step on overflow would skip too many steps with FP8’s narrow range.2
The designers kept the IEEE-like biases and left scaling to software: a per-tensor scale in higher precision can take any real value, while a programmable exponent bias is equivalent to power-of-two scales only.2 A scale chosen so that amax maps near the format’s maximum is applied before the cast, and its inverse is applied once per dot product (or folded into the next operation), so the cost is amortized over many multiply-accumulates.2
Integer quantization in detail
The affine scheme makes real zero exactly representable () and expands an integer GEMM into an int8 × int8 core plus correction terms that cost for an multiplication.9 Symmetric () quantization avoids the corrections, and an NVIDIA study found it sufficient for INT8 on all networks it examined.12 The quantizer is
with two error sources: rounding error, uniform within for in-range values, and clipping error for values outside . Raising trades the second for the first.13 Calibration picks from activation statistics (max, entropy, or a percentile such as 99.99%).12 Granularity is the other knob: one scale per tensor, per output channel (common for weights), or per group of elements; finer groups generally improve accuracy at some cost in overhead.13
Float versus integer at 8 bits
For the same 8 bits, the difference between INT8 and FP8 is only where the grid points sit: uniform for INT8, denser near zero and sparser far out for FP8. So floats do better on distributions with heavy tails and outliers, and integers on uniform ones; on a Gaussian, the best low-exponent float formats and INT8 come out close.14 The simulator reproduces this: on its near-Gaussian “Weights” tensor (only two mild outliers), INT8 with a per-tensor scale has lower error than E4M3, and on the outlier-heavy “Activations” tensor the order flips.
Emergent outliers
Dettmers et al. found that in transformers, large-magnitude features (up to ~20× the other dimensions) appear first in about a quarter of layers and, around 6.7B parameters, in all of them. They occupy about 0.1% of feature dimensions; zeroing them degrades perplexity by 600–1000%.11 They are highly systematic: at 6.7B parameters about 150,000 outliers per sequence fall in only 6 feature dimensions. LLM.int8() exploits that, sending those dimensions through a 16-bit GEMM and quantizing the rest to 8 bits with a separate scale per row and column.11 Fine-grained block scaling attacks the same problem by limiting how many values each outlier can hurt.
INT8, one scale: s = 20 / 127 = 0.1575. Ordinary values snap to multiples of s; 2 of 7 round to 0, RMS error 5% of their typical size.
A multiplier for every small group
The fix for giants is simple: give every small group of numbers its own multiplier. Then a giant only spoils its own group. This is called . In 2023, a group of chip and cloud companies agreed on a shared standard for it.4
In the standard, every 32 numbers share one multiplier. The numbers themselves can be as small as 4 bits each.4
Four bits
The 4-bit format is called . It can only hold 0, 0.5, 1, 1.5, 2, 3, 4 and 6, plus their negatives.4 That sounds hopeless. But with a fresh multiplier for every 32 numbers, it’s enough to store a big chatbot model.
One openly released model stores most of its numbers this way. That shrinks it enough to fit 120 billion numbers in the memory of a single AI chip.15
The MX specification
The Open Compute Project’s Microscaling Formats specification (version 1.0, September 2023), written by engineers from AMD, Arm, Intel, Meta, Microsoft, NVIDIA and Qualcomm, defines a family of block-scaled formats.4 An MX block has three parts:
- elements, all in the same narrow format. For every concrete MX format .
- One shared scale , in a format called E8M0: 8 bits holding just an exponent, so is a power of two between and .
- The value of element is simply .
| Format | Element | Element bits | Block | Bits per value |
|---|---|---|---|---|
| MXFP8 | FP8 E4M3 or E5M2 | 8 | 32 | 8.25 |
| MXFP6 | FP6 E2M3 or E3M2 | 6 | 32 | 6.25 |
| MXFP4 | FP4 E2M1 | 4 | 32 | 4.25 |
| MXINT8 | INT8 | 8 | 32 | 8.25 |
Formats and block size from the specification’s Table 1; bits per value = element bits + 8/32.4
Choosing the scale
The specification’s conversion rule: set to the largest power of two not exceeding the block’s largest magnitude, divided by the largest power of two the element format can represent. Then divide every value by and round it to the element format, clamping anything that still exceeds the format’s maximum.4 For FP4, whose largest power of two is 4, a block whose biggest value is 20 gets , and its elements are stored as multiples of .
Why blocks of 32
A scale per tensor works for 8-bit formats but has been shown to be insufficient below 8 bits, because the formats have too little range to cover a whole tensor.8 A scale per 32 values follows the data closely while costing only a quarter of a bit per value. In the specification’s accompanying study, 8-bit MX formats ran inference directly on FP32-trained models with minimal accuracy loss and no calibration; 6-bit MX formats trained large transformers to FP32 accuracy; and training with 4-bit MX weights cost only a minor accuracy drop.8
FP4 and INT4
FP4 (E2M1) has 1 sign, 2 exponent and 1 mantissa bit, no infinities or NaN, and the magnitudes 0, 0.5, 1, 1.5, 2, 3, 4 and 6.4 INT4 has the evenly spaced integers −7 to 7 (or −8 to 7). Four-bit weights are attractive because, for a fixed memory budget, a bigger model with 4-bit weights tends to beat a smaller model with 8-bit weights; a study of more than 35,000 experiments found 4 bits almost universally optimal for total model bits, with small block sizes and the choice of data type being the main ways to improve further.16
Variations on the MX idea exist. NVIDIA’s NVFP4 uses the same E2M1 elements but blocks of 16 values, a scale stored in E4M3 (so it is not limited to powers of two), and an extra FP32 scale per tensor, at 4.5 bits per value.17
Semantics
An MX format is a triple: scale type ( bits), element type ( bits) and block size ; a block costs bits, and its memory layout is not prescribed.4 The concrete formats all use and an E8M0 scale: unsigned, bias 127, exponents −127 to 127, a single NaN code (0xFF) that poisons the whole block, no infinities and no zero.4 E8M0 covers a superset of FP32’s exponent range.8 FP6 (E2M3, max 7.5; E3M2, max 28) and FP4 (E2M1, max 6) reserve no codes; conversion to them must support round-to-nearest-even and saturation to . MXINT8 elements are two’s complement with an implicit scale, i.e. a sign, one integer bit and six fraction bits.4
Conversion
emax_elem = exponent of the element format's largest normal
(E4M3: 8, E5M2: 15, E2M1: 2, E3M2: 4, E2M3: 2, INT8: 0)
shared_exp = floor(log2(max_i |V_i|)) - emax_elem
X = 2^shared_exp # stored as E8M0: shared_exp + 127
P_i = round_to_element(V_i / X) # ties to even; clamp normals
# above the element max to +/-max
value_i = X * P_i- 1L3The block’s largest magnitude lands in the element format’s top binade, so the full exponent range of the element is used.
- 2L5The scale is a power of two: applying it is an exponent add, not a multiply.
- 3L6Because the floor rounds the block maximum’s exponent down, V/X can exceed the element max (e.g. 6 < V/X < 8 in FP4) and is clamped.
The rule follows the specification’s minimal semantics; implementations may use other algorithms.48 The scale is shared along one axis, usually the reduction dimension of the matrix multiplication, so converting and transposing do not commute: a matrix that is reduced along different axes in the forward and backward passes has to be converted separately for each.8
The dot product
For two MX blocks, : the element products are reduced first and the two power-of-two scales applied once. A longer dot product is a sum of block dot products and should produce an FP32 result; the internal precision and order of operations are implementation-defined.4 In hardware, the inner sum is a small fixed-point or narrow-float adder tree over 32 products, and the scale is an exponent addition before the result joins a wider accumulator.
The power-of-two penalty and NVFP4
Because rounds down, lies in . For E2M1 that is , and everything in saturates to 6: the block maximum itself can lose up to 25%. Rounding the exponent up instead avoids clipping but leaves up to one binade of the tiny element range unused. NVIDIA’s NVFP4 paper describes MXFP4 as potentially losing up to one binade of dynamic range because of power-of-two scale rounding, and changes three things: 16-element blocks, E4M3 block scales with fractional precision, and a second, per-tensor FP32 scale that keeps the E4M3 scales in range. At least one value per block (its amax) is then represented at near-FP8 precision.17 The cost is bits per value instead of 4.25.
What 4 bits has been shown to do
- Inference: Dettmers and Zettlemoyer’s bit-level scaling laws (19M–176B parameters) found 4-bit weights almost universally optimal for accuracy per total model bit; below that, 3-bit reverses the trend.16
- Deployment: the gpt-oss models store their mixture-of-experts weights, over 90% of parameters, in MXFP4 at 4.25 bits per parameter.15
- Training: a 12B-parameter model pretrained on 10 trillion tokens with NVFP4 GEMMs reached 62.58% on MMLU-pro against 62.62% for an FP8 baseline. It needed Random Hadamard transforms to spread out block outliers, two-dimensional block scaling, stochastic rounding for gradients, and a few layers kept in higher precision.17
One scale for all 16: X = 4. FP4 steps are multiples of X/2, so 10 of 15 ordinary values round to 0 (5 in block 1, 5 in block 2).
Smaller multipliers
A chip multiplies the way you do on paper: every digit of one number times every digit of the other, then add up the rows.
So fewer digits shrink the multiplier fast. Going from 32-bit numbers to BF16 makes it about eight or nine times smaller.5 The same chip can then fit many more of them.
Less to move
Moving numbers costs even more energy than multiplying them. Fetching one number from memory can use about 200 times the energy of the math it feeds.6 Half the bits means half as much to move.
The total stays big
There’s one place chips don’t skimp. When thousands of results are added up, the running total gets many more bits, often 32. A small total would lose the tiny additions once it got big. It would be like weighing a feather on a truck scale.18
Multiplier cost grows with the square of the width
Multiplying two -bit significands takes roughly one-bit partial products, which are then added together in an array of adders. So the area and energy of a multiplier grow roughly with , while an adder grows roughly with .6 Google explains BF16 this way: its multipliers are about half the size of FP16’s and eight times smaller than FP32’s.5 Measured figures quoted in a widely used survey put numbers on it: an 8-bit integer multiply uses 18.5 times less energy and 27.5 times less area than a 32-bit floating-point multiply, and an 8-bit add 30 times less energy than a 32-bit floating-point add.6
Integer versus floating point at the same width
A floating-point multiply needs more than the significand multiplier: it adds the exponents, and before products can be summed they have to be shifted to line up. An integer multiply needs neither. One gate-count estimate found an FP8 (E4M3) multiply-accumulate about 53% more expensive than INT8 with the same kind of accumulator, and 183% more if it has to accumulate in FP32.14 That is the case for integers in inference chips, and the price paid for FP8’s wider range.
Throughput and bandwidth
Chips turn the smaller multipliers into throughput. On one GPU generation, published figures give FP16 8 times and INT8 16 times the FP32 matrix throughput, with 2× and 4× less data to move per value.12 Data movement is the bigger energy item: a DRAM access costs orders of magnitude more energy than an arithmetic operation,6 so in memory-bound work (where is low), the bandwidth saving is what makes narrower formats faster.
Accumulation precision
Every output of a matrix multiplication is a dot product: thousands of products added into one . The inputs can be narrow; the sum cannot.
- Integers accumulate exactly if the accumulator is wide enough: INT8 inputs use a 32-bit integer accumulator.9
- Floats lose information in long sums. Adding a small number to a large one shifts the small one right to align the binary points, and bits that fall off the end are gone. Once the running sum is more than about times bigger than the next term, the term disappears entirely, an effect called .18
That is why every recipe in this chapter pairs narrow inputs with FP32 (or wider) accumulation: FP16 and BF16 training,17 FP8 matrix units, which offer FP16 or FP32 accumulators,19 and the MX specification, whose general dot product should produce an FP32 result.4 The cost is modest: there is one accumulator per output, but hundreds or thousands of multiplications feeding it.
A first-order cost model
For a floating-point multiply-add the main blocks are: a significand multiplier (), an exponent adder (), an alignment shifter and adder for the accumulation (, with log-depth shifting), and normalization and rounding. For integers only the multiplier and a fixed-point adder remain. The survey by Sze et al., citing Horowitz’s ISSCC 2014 figures, states the scaling explicitly: multiplier energy and area scale roughly quadratically with bits, adder and memory cost roughly linearly.6
| Format | Significand | (multiplier, relative) | vs FP32 |
|---|---|---|---|
| FP32 | 24 | 576 | 1 |
| FP16 | 11 | 121 | 0.21 |
| BF16 | 8 | 64 | 0.11 |
| INT8 | 8 | 64 | 0.11 |
| FP8 E4M3 | 4 | 16 | 0.028 |
| FP8 E5M2 | 3 | 9 | 0.016 |
| FP4 E2M1 | 2 | 4 | 0.007 |
This is the multiplier array only. It explains why BF16 multipliers are about half of FP16’s and an eighth of FP32’s,5 but it overstates the savings for tiny floats, whose cost is dominated by the shifters, the exponent logic and above all the accumulation. The van Baalen et al. gate-count study makes the point: FP8-E4 with a fixed-point accumulator is 53% more expensive than INT8, and with an FP32 accumulator 183% more. Their INT8 design needs a 27-bit fixed-point accumulator (15 bits of product plus 12 bits of headroom for 4,096 products); FP8-E4, whose products span many more binades, would need 37 bits.14 That is consistent with published peak figures, where halving the input width roughly doubles matrix throughput rather than quadrupling it.1219
Accumulation error
Floating-point accumulation is inexact and order-dependent, and the error grows with the reduction length .14 The failure mode is : when the running sum exceeds an addend by more than , the addend is shifted out entirely, and the problem worsens with long sums and non-zero-mean data.18 Wang et al. showed that splitting a dot product into chunks summed separately and then combined (chunk-based accumulation), plus , let FP8 training use 16-bit additions.18
A real-world example: DeepSeek-V3’s authors measured that FP8 GEMMs on the GPU they used retain only about 14 bits of accumulation precision inside the matrix unit, giving a maximum relative error of nearly 2% for . Their fix copies partial sums to FP32 registers every 128 elements of , which is chunked accumulation implemented in software.20 The general lesson is that the narrow-format story is really two stories: input precision, which sets multiplier cost and bandwidth, and accumulation precision, which sets numerical quality and is much cheaper to keep high.
BF16: p = 8, 64 partial products, about one ninth of FP32 (Google says about eight times smaller).
BF16 running sum: stuck at 256 after 1,024 additions of 1. With 7 mantissa bits, a sum of 256 has a step of 2, so +1 rounds away.
The simulator holds 256 made-up numbers, shaped like the real ones in a model. Most are small, and a few are giants (the red triangles). Pick a format along the top. The gray bars are the original numbers, and the pink bars are the rounded copies.
Try INT4 on “Big outliers”: almost everything turns into zero. Then try MXFP4, which gives each group of 32 numbers its own multiplier. Which groups turn red?
Choose a format and a sample tensor. The histogram shows the input values (gray) and the quantized values (pink); blue ticks under the axis are the representable values at the current scale. The bit panel decodes one value field by field. On “Activations + outliers”, compare INT8 with one scale for the tensor (RMS error 4.8% of , 17 values flushed to zero) with MXFP8 (2.7%), then INT4 (78%, 214 of 256 values zeroed) with MXFP4 (32%). The colored strip shows each scale group’s error: with MX formats only the blocks holding an outlier turn red. On “Gradients (tiny)”, compare FP16 (six values flushed to zero) with BF16.
Encodings follow OFP8 and MX v1.0 exactly: round-to-nearest-even, subnormals, saturation, E8M0 scales computed as . FP8 and integer formats use an FP32 per-tensor scale mapping amax to the format maximum. Click bits to decode arbitrary patterns (try S.1111.111 in E4M3, or the all-ones exponent in E5M2). The block-size slider sweeps MX from to 256: MXFP4 on the activations goes from 18% to 32% to 52% RMS error of as grows from 8 to 32 to 256, while the scale overhead falls from 1 to 0.03 bits per value. Watch the clipped count too: a block maximum just below the next power of two saturates in E2M1 and E4M3 alike. The readouts include SQNR and a multiplier-cost estimate.
- Bits per number, classic format
- 32
- Bits per number, common for training today
- 16 or 8
- Values an FP4 number can take
- 15
- Numbers sharing one multiplier
- 32
What these numbers mean:
- The classic format uses 32 bits per number. AI training moved to 16, but kept 32 bits for the main copy of the model.1
- Tests show that 8-bit numbers can train models as well as 16-bit ones, even very large chatbot models.2
- FP4 has 16 patterns of bits. But plus zero and minus zero are the same number, which leaves 15 values, from −6 to 6.4
- In the shared standard, 32 numbers share one multiplier. That adds only a quarter of a bit to each one.4
- Largest FP16 / E4M3 / E5M2 value
- 65,504 / 448 / 57,344
- MXFP4 bits per value
- 4.25
- INT8 vs FP32 multiply energy
- 18.5× less
- FP8-E4 MAC cost vs INT8 (gate estimate)
- +53%
| Format | Bits | Sign / exponent / mantissa | Range (largest) | Typical use |
|---|---|---|---|---|
| FP32 | 32 | 1 / 8 / 23 | Master weights, accumulators | |
| BF16 | 16 | 1 / 8 / 7 | Training | |
| FP16 | 16 | 1 / 5 / 10 | 65,504 | Training with loss scaling; inference |
| FP8 E4M3 | 8 | 1 / 4 / 3 | 448 | Weights and activations |
| FP8 E5M2 | 8 | 1 / 5 / 2 | 57,344 | Gradients |
| INT8 | 8 | two’s complement | 127 × scale | Inference |
| MXFP4 | 1 / 2 / 1, shared E8M0 | Weights for inference; research training |
Bit layouts and ranges from the OCP specifications and the BF16 study; uses as recommended by the FP8 and MX papers.73428
Energy. An 8-bit integer multiply uses 18.5× less energy and 27.5× less area than a 32-bit floating-point multiply; an 8-bit add uses 30× less energy than a 32-bit float add.6 Throughput. On one GPU generation, INT8 matrix math ran at 16× the FP32 rate and INT4 at 32×.12
- E4M3 / E5M2 dynamic range
- 18 / 32 binades
- FP8 GEMM accumulation measured on one GPU
- ≈ 14 bits
- NVFP4 vs FP8 pretraining, MMLU-pro (12B, 10T tokens)
- 62.58% vs 62.62%
- MXFP4 share of gpt-oss parameters
- over 90%
Format parameters
| Format | Bias | Max | Min subnormal | Specials |
|---|---|---|---|---|
| FP16 | 15 | 65,504 | IEEE , NaN | |
| E4M3 | 7 | 448 | NaN only (S.1111.111) | |
| E5M2 | 15 | 57,344 | IEEE , NaN | |
| E3M2 (FP6) | 3 | 28 | 0.0625 | none |
| E2M3 (FP6) | 1 | 7.5 | 0.125 | none |
| E2M1 (FP4) | 1 | 6 | 0.5 | none |
| E8M0 (scale) | 127 | (no zero) | one NaN (0xFF) |
From the FP16 column of the BF16 study, the OFP8 specification and the MX specification.734
Published results worth knowing
- FP8 (E4M3 forward, E5M2 gradients, per-tensor scales) matched 16-bit training on CNNs, RNNs and transformers up to 175B parameters with unchanged hyperparameters.2
- DeepSeek-V3 (671B total parameters) trained with FP8 GEMMs, E4M3 on all tensors, 1 × 128 activation tiles and 128 × 128 weight blocks; in tests at two smaller scales the FP8 model’s loss stayed within 0.25% of a BF16 baseline.20
- 8-bit MX direct-cast inference on FP32-trained models with minimal loss; 6-bit MX training at FP32 quality.8
- Emergent outliers in LLMs: about 0.1% of features, up to 20× larger, in all layers from about 6.7B parameters.11
Every bit you remove makes the chip faster and the model smaller. It also makes each number a little less exact. Models can handle a lot of rounding, but not endless amounts. Four bits only works with careful tricks, and fewer than that is much harder.16
The sneaky part is how things go wrong. Tiny numbers quietly become zero. A few giants can blur everything else. Nothing crashes; the model just gets a bit worse. That’s why teams test a model carefully after changing its number format.
What you give up for what you get
| Choice | You get | You give up |
|---|---|---|
| Fewer bits | Less memory and bandwidth, smaller multipliers | Precision; more engineering to keep accuracy |
| More exponent bits | Range: fewer values lost to underflow or overflow | Precision within each power of two |
| Integer instead of float | Smaller, exact multiply-accumulate hardware | Uniform grid that handles outliers and tiny values poorly |
| Finer scaling (per block) | Outliers only hurt their own block | Scale storage ( bits per value) and per-block hardware |
| Wide accumulator | Accurate long dot products | Some area and energy per output (small next to the multipliers) |
How it goes wrong
- Underflow. Values below the smallest representable number become zero, as FP16 gradients do without loss scaling.1
- Overflow. Values above the maximum become infinity or NaN unless the conversion saturates; one NaN then spreads through every result it touches. FP8 recipes saturate for this reason.2
- Outliers. With one scale per tensor, a few huge values set the step size for everything and ordinary values lose most of their precision.11
- Accumulation loss. Narrow running sums drop small terms; errors grow with the length of the dot product.1820
- Calibration mismatch. A scale chosen from sample data can be wrong for real inputs, clipping values the calibration set never showed.12
Each failure shows up as a quiet drop in accuracy rather than a crash, so format changes are validated against the full-precision model on real evaluation tasks. Peak speeds quoted for a new format only count if the model still works in it.
The design space
- Element width and type. At fixed bits, exponent versus mantissa trades range for SQNR (about 6 dB per mantissa bit). Integers maximize resolution for a known range; floats tolerate unknown or heavy-tailed distributions.1422
- Scale granularity and type. Per-tensor FP32 scales are free in storage but expose every value to the tensor’s worst outlier. Per-block scales cost bits per value; power-of-two scales cost only an exponent add but can waste up to a binade; E4M3 scales are more precise but need a second-level tensor scale for range.417
- Accumulator width. Exact fixed-point accumulation is cheapest for integers; float inputs with many exponent bits need very wide fixed-point accumulators or a floating accumulator, which is where FP8 loses much of its hardware advantage over INT8.14
- Software burden. Every step down the ladder has added machinery: master weights and loss scaling for FP16, per-tensor scaling for FP8, tile and block scaling plus promoted accumulation in FP8 training at scale, and Hadamard rotations, stochastic rounding and high-precision layers for FP4 training.12017
Failure modes in detail
- Range failures are binary and data-dependent: a tensor that fits during calibration may overflow on an unusual input. Saturating conversion bounds the damage; non-saturating conversion turns it into NaN or .3
- Resolution failures are statistical: SQNR drops gradually, and the effect on accuracy depends on the layer. Operations other than matrix multiplications, such as nonlinearities and normalizations, typically run in higher precision anyway,2 and the NVFP4 recipe also kept a few linear layers in higher precision for stability.17
- Systematic bias: round-to-nearest of updates smaller than half an ulp is always zero, so many small updates are lost rather than averaged; stochastic rounding removes that bias at the cost of noise and a random-number source.18
- Hardware-specific numerics: the MX spec leaves dot-product internal precision implementation-defined, so the same model in the same format can produce slightly different results on different chips.4 DeepSeek-V3’s accumulation measurement is an example of such a detail mattering at scale.20
Tap a failure mode. Each one shows up as a quiet drop in accuracy, not a crash, which is why format changes are validated on real tasks.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Rounding to a small float
Every format in this chapter can be produced by one routine: find the binade, round the significand to bits with ties to even, saturate, and handle the bottom of the range with subnormals. This is what the simulator runs, per the OFP8 and MX conversion rules.34
emin = 1 - bias # exponent of the smallest normal
a = |x|
e = max(floor(log2(a)), emin) # below emin: subnormal spacing
ulp = 2^(e - m) # step size in this binade
r = round_half_even(a / ulp) * ulp
if r > max_normal: r = max_normal # SAT mode; NONSAT gives Inf or NaN
return sign(x) * r # r == 0 means the value underflowed- 1L3Clamping e at emin is what makes subnormals: the step stops shrinking at 2^(emin − m).
- 2L4The step doubles each binade, so the relative error is bounded by 2^−(m+1), i.e. 2^−p.
- 3L5Ties go to the even neighbor; rounding up can carry into the next binade, which is still a valid value.
- 4L6OFP8 requires both saturating and non-saturating modes; MX FP6/FP4 must support saturation.
Noise model for a uniform grid
If values fall at random positions between grid points spaced apart, the rounding error is uniform on and its mean square is . For a symmetric -bit integer quantizer with a per-tensor scale set by the tensor’s largest magnitude, with . For values with RMS , the signal-to-quantization-noise ratio () is
Three readings of this formula:
- Each extra bit doubles and adds dB.
- Each doubling of the crest factor costs the same 6 dB, a full bit. An outlier 20× the RMS costs about 26 dB, more than four bits.
- For INT8 with , as in the simulator’s activation tensor, the formula predicts dB on the ordinary values; the simulator measures an RMS error of 4.8% of , which is 26.3 dB.
Widening the range past is never useful, but narrowing it below trades this rounding noise for clipping error on the tail: the total error is roughly plus the mean of over values beyond the clipping point . Calibration methods choose to balance the two.1312
Noise model for a float
Within a binade the step is , so the error is still uniform, with mean square . Relative to , and averaging over a value spread evenly through the binade (the mean of for uniform on is ), the relative noise power is . That gives, independent of scale while values stay in range,
The linear dependence on mantissa bits matches the analysis of block formats by Rouhani et al.22 The prediction is 31.9 dB for E4M3 (), 25.8 dB for E5M2 and 55.9 dB for BF16; the simulator measures 32–34, 26–28 and 54–62 dB on its three tensors. The model breaks down at the edges: values pushed into the subnormal range lose relative precision, values below the smallest subnormal are lost, and values above the maximum clip. That is why a float’s SQNR is flat across scales only within its dynamic range, and why the scale factor for FP8 matters even though the float grid is “self-scaling.”
Block scaling
Block scaling replaces the tensor’s crest factor with each block’s. For an MX block with scale , the block maximum lands in of the element format: on average the top of the element range is under-used by about half a binade, and in FP4 a maximum in saturates. Shared-exponent analyses show the resulting SQNR growing linearly with mantissa bits and logarithmically with finer block granularity.22 The cost side is exact: bits per value for an E8M0 scale. In the simulator, MXFP4 on the activation tensor falls from 18% to 32% to 52% RMS error of as goes from 8 to 32 to 256, for 1, 0.25 and 0.03 bits of overhead.
Accumulation error
Summing terms one after another in a float with unit roundoff gives an error bound that grows in proportion to . Splitting the sum into chunks of length , summing each, and then summing the partial sums reduces the bound to order ,18 which is smallest when . For that is ; DeepSeek-V3’s promotion interval of 128 sits in the same range, and it was their answer to errors of nearly 2% measured at .20
For integers the question is width, not rounding: a -bit × -bit product needs bits, and summing of them needs more. The INT8 example in van Baalen et al. is bits for 4,096 products; a fixed-point accumulator that exactly covers FP8-E4 products would need 37 bits.14 That is why 32-bit integer accumulators are standard for INT8.9
Stochastic rounding
Round-to-nearest maps every value in to , so an update smaller than half an ulp of the weight is lost every time, however often it is applied. Stochastic rounding instead rounds up with probability , so and small updates survive on average. Wang et al. used it to make 16-bit weight updates work in FP8 training,18 and the NVFP4 recipe uses it for gradients.17
After 150 updates of 0.001: exact 1.15, nearest-even 1 (every update below ulp/2 = 0.0039 is lost), stochastic 1.1641 (unbiased: E[round(x)] = x).
Q1BF16 and FP16 are both 16 bits. What is the key difference?
Q2A tensor is quantized to INT8 with one scale for the whole tensor, . Typical values are about 1, but one outlier is 25. What happens to the typical values?
Q3How many bits per value does MXFP4 cost, counting the shared scale?
Q4Why do 8-bit matrix units usually add up their products in 32 bits?
Sources
Show Hide 22 sources
- Mixed Precision TrainingFP16 storage with an FP32 master copy of weights, loss scaling and FP32 accumulation; nearly halves memory; values below 2^−24 become zero in FP16; about 5% of weight gradients fall below that; an SSD model diverges without loss scaling and trains with a scale of 8.
- FP8 Formats for Deep LearningDefines E4M3 and E5M2; E4M3 drops infinities to reach 448 (240 otherwise); E4M3 for weights and activations, E5M2 for gradients; per-tensor scale factors with saturation; FP8 training matches 16-bit on models up to 175B parameters.
- OCP 8-bit Floating Point Specification (OFP8), Revision 1.0Bit layouts, biases (7 and 15), special values and ranges of E4M3 (18 binades, max 448) and E5M2 (32 binades, max 57,344); round-to-nearest-even and saturating or non-saturating conversion.
- OCP Microscaling Formats (MX) Specification, Version 1.0A block of k = 32 elements sharing one E8M0 power-of-two scale; MXFP8, MXFP6, MXFP4 and MXINT8; FP4 (E2M1) and FP6 encodings; the conversion rule for the shared scale; dot-product semantics with implementation-defined internal precision.
- BFloat16: The secret to high performance on Cloud TPUsMultiplier size scales with the square of the mantissa width; a BF16 multiplier is about half the size of an FP16 one and eight times smaller than FP32; the name comes from Google Brain; TPU v2/v3 multiply in BF16 and accumulate in FP32 in a 128 × 128 systolic array.
- Efficient Processing of Deep Neural Networks: A Tutorial and SurveyCiting Horowitz (ISSCC 2014): an 8-bit fixed-point add uses 30× less energy than a 32-bit float add; an 8-bit fixed-point multiply 18.5× less energy and 27.5× less area than a 32-bit float multiply; multiplier cost scales about quadratically with bits, adder and memory cost about linearly; DRAM access costs orders of magnitude more than arithmetic.
- A Study of BFLOAT16 for Deep Learning TrainingTable of FP32, FP16 and BF16 bit layouts and ranges; BF16 keeps FP32’s range so training needs no hyperparameter changes; FMA units built from 8-bit multipliers; BF16 began as a storage format in DistBelief and TensorFlow; round-to-nearest-even conversion; FP32 accumulation.
- Microscaling Data Formats for Deep LearningPer-tensor scaling is insufficient below 8 bits; MX conversion algorithm; 8-bit MX direct-cast inference with minimal loss; 6-bit MX training matches FP32; 4-bit MX weights train with a minor accuracy drop.
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only InferenceThe affine scheme r = S(q − Z) with a real scale S and an integer zero-point Z; 8-bit integer operands with a 32-bit integer accumulator.
- What Every Computer Scientist Should Know About Floating-Point ArithmeticSignificand, base and precision p; relative rounding error bounded by machine epsilon (β/2)β^−p; IEEE single precision with 8 exponent bits and a hidden bit giving p = 24; denormals and gradual underflow.
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleOutlier features up to 20× larger than other dimensions emerge in all transformer layers around 6.7B parameters; they are about 0.1% of features but zeroing them wrecks perplexity; outlier dimensions kept in 16-bit, the other 99.9% in 8-bit.
- Integer Quantization for Deep Learning Inference: Principles and Empirical EvaluationRelative math throughput and bandwidth of FP16, INT8 and INT4 versus FP32 on a 2018 GPU architecture; quantization granularity; calibration by max, entropy or percentile; an 8-bit workflow within 1% of floating point on all networks studied.
- A White Paper on Neural Network QuantizationThe quantization function with scale and zero-point; widening the range cuts clipping error but raises rounding error; per-channel and per-group quantizers.
- FP8 versus INT8 for efficient deep learning inferenceGate-count estimate: an FP8-E4 multiply-accumulate costs about 53% more than INT8 with fixed-point accumulation, 183% more with an FP32 accumulator; integer (Kulisch) accumulation is exact; floats suit outlier-heavy distributions, integers uniform ones.
- gpt-oss-120b & gpt-oss-20b Model CardMixture-of-experts weights (over 90% of parameters) post-trained to MXFP4 at 4.25 bits per parameter, letting the 120B model fit on one 80 GB GPU.
- The case for 4-bit precision: k-bit Inference Scaling LawsMore than 35,000 experiments from 19M to 176B parameters: for a fixed total number of model bits, 4-bit weights are almost universally optimal; small block sizes and the choice of data type are what help.
- Pretraining Large Language Models with NVFP4NVFP4: 16-element blocks, E4M3 block scales and an FP32 tensor scale; MXFP4’s power-of-two scale can waste up to one binade; table of microscaling formats and speedups on one GPU generation; a 12B model trained on 10 trillion tokens in 4-bit matching FP8.
- Training Deep Neural Networks with 8-bit Floating Point NumbersSwamping: a small addend is lost when the running sum is larger by more than 2^(mantissa+1); chunk-based accumulation and stochastic rounding allow 16-bit additions.
- NVIDIA Hopper Architecture In-DepthFP8 tensor cores with E4M3 and E5M2 inputs and FP16 or FP32 accumulation; FP8 doubles throughput over FP16/BF16; software chooses between FP8 and 16-bit per layer and manages scaling.
- DeepSeek-V3 Technical ReportFP8 training of a 671B-parameter model; 1×128 and 128×128 scaling tiles; E4M3 on all tensors; FP8 GEMM accumulation on the GPU used retains about 14 bits, giving up to nearly 2% error at K = 4096, fixed by promoting partial sums to FP32 every 128 elements; loss within 0.25% of BF16.
- Introducing AMD CDNA 4 Architecture (white paper)CDNA 3 added both OCP FP8 variants; CDNA 4 adds hardware support for MXFP8, MXFP6 and MXFP4 and doubles matrix throughput for 16-bit and smaller types.
- With Shared Microexponents, A Little Shifting Goes a Long WayBlock data representations; quantization signal-to-noise ratio grows linearly with mantissa bits (about 6.02 dB per bit) and logarithmically with block granularity; power-of-two block scales are cheap in hardware.