Architectures · Chapter 7 of 16 · AI accelerators

Number formats

Computers can store numbers with many digits or just a few. AI often works fine with fewer, like rounding 3.14159 to 3.1, and that makes chips faster and cheaper to run.

Training moved from 32-bit floating point to 16-bit formats like BF16, and inference increasingly uses 8-bit and 4-bit formats. Fewer bits mean smaller multipliers and less data to move, at some cost in precision.

FP32, FP16 and BF16, the two FP8 variants, integer quantization, block-scaled microscaling (MX) formats and FP4, accumulation precision, and how a multiplier’s cost grows with mantissa width.

A neural network is billions of numbers, multiplied and added over and over. A computer stores each number as a row of bits. A bit is a single 0 or 1, like a light switch that is off or on.

How many bits each number gets matters a lot. Fewer bits means less to store, less to move around, and smaller, cheaper math circuits on the chip.

Ordinary programs use 32 bits for a number like 3.14. AI chips now use 16, 8 or even 4. With 4 bits there are only 16 possible patterns of 0s and 1s. That sounds far too few to work, and on its own it is. This chapter is about the tricks that make it work.

What the workload needs showed that AI is dominated by matrix multiplication: long chains of operations. This chapter is about the numbers those operations work on. How a value is encoded decides three costs at once: how much memory a model needs, how many bytes move per operation, and how large and power-hungry each multiplier is.

For decades the default for real numbers was 32-bit (FP32). Around 2017 training moved to 16-bit formats, with FP32 kept for the sensitive parts. In 2022 two 8-bit floating-point formats were proposed for deep learning and later standardized. In 2023 a group of chip and cloud companies published an open specification for block-scaled formats down to 4 bits, called microscaling (MX). Each step roughly halves memory and data movement, and lets a chip of the same size do more math.

The rest of the chapter covers:

  • how floating point trades range against precision, and why that split is a design choice;
  • BF16 versus FP16, and the mixed-precision recipe that made 16-bit training work;
  • FP8, integer formats, and the role of the scale factor;
  • microscaling and FP4, where every 32 values share a scale;
  • what fewer bits buy in hardware, and why the running sum still needs many bits.

Given that a network is mostly matrix multiplication (What the workload needs), the number format sets three things at once: bytes per operand, which moves the of every kernel; multiplier area and energy, which grow roughly with the square of significand width; and the numerical error the model has to absorb. The history is a steady walk down that ladder:

  • FP32 → 16-bit with mixed precision: 16-bit operands, FP32 accumulation, FP32 master weights, and loss scaling for FP16. BF16 removed the need for loss scaling by keeping FP32’s exponent.
  • 16 → 8-bit floats: E4M3 and E5M2, standardized by OCP, used with per-tensor scale factors.
  • Per-tensor → per-block scaling: the OCP MX specification (MXFP8, MXFP6, MXFP4, MXINT8), because a single scale is not enough below 8 bits.

In parallel, inference has long used integer quantization with INT8 operands and INT32 accumulators. This chapter goes through the encodings bit by bit, the scaling schemes that make narrow formats usable, the error each introduces, accumulation precision, and the hardware cost model that motivates all of it. The algorithms section derives the error models.

FP3232 bits · 4,294,967,296 patternsSexponentmantissavalues between 1 and 28,388,608 steps: too fine to draw12memory, data moved1multiplier size1relative to FP32 (dashed outline)
Format

FP32: 1 sign, 8 exponent, 23 mantissa bits. 2^23 steps between 1 and 2; the reference for size and cost.

One number in four formats. Memory and data moved shrink with the bit count; the multiplier shrinks roughly with the square of the significand width; precision drops with the mantissa bits.Share freely with credit: ‘Figure from chipfieldguide.com’

Scientific notation, with bits

Scientists write big numbers in a short way. Three million becomes 3 × 1,000,000. The 3 gives the digits, and the “× 1,000,000” gives the size.

Computers do the same thing with bits. It’s called , because the decimal point can float left or right. One bit says plus or minus. A few bits give the size, called the . The rest hold the digits.

This lets the same few bits describe a speck of dust or a galaxy. The catch is that each number keeps only a few exact digits. Small numbers get small rounding errors, and big numbers get big ones.

Range against detail

With a set number of bits, every bit given to the size is taken from the digits. Picking that split is the main choice behind every format in this chapter.

Three fields

A binary number has a sign bit SS, an field EE of ee bits, and a field MM of mm bits. For an ordinary (“normal”) number, the value is

value=(−1)S×2E−bias×(1+M2m)\text{value} = (-1)^S \times 2^{E - \text{bias}} \times \left(1 + \frac{M}{2^m}\right)

The exponent is stored with a bias so the field can be an unsigned number: in FP32, E=127E = 127 means 202^0. The leading “1 +” is not stored at all. Every normal number starts with a 1 in binary, so formats leave it out, a free extra bit called the hidden bit. FP32 has 1 sign bit, 8 exponent bits and 23 mantissa bits, so its significand (the hidden 1 plus the mantissa) has 24 bits of precision.

A worked example

Take the number 6.5. In binary it is 110.1, which is 1.101×221.101 \times 2^2. So S=0S = 0, the exponent is 2, and the mantissa bits are 101 followed by zeros. In FP32 the stored exponent is 2+127=1292 + 127 = 129 (binary 10000001). In a tiny format with 3 mantissa bits the mantissa would be 101, and 6.5 fits exactly. The number 6.6 would not: the nearest values with 3 mantissa bits at that size are 6.5 and 7.0, so 6.6 rounds to 6.5.

Precision is relative

Between one power of two and the next (a binade, say 4 to 8), the mm mantissa bits split the interval into 2m2^m equal steps. With 3 mantissa bits, the values between 4 and 8 are 4, 4.5, 5, 5.5, … 7.5; between 8 and 16 they are 8, 9, 10, …. The step doubles every binade, so the relative error stays roughly constant: rounding to the nearest value is wrong by at most about 2−(m+1)2^{-(m+1)} of the number, whatever its size. For FP32 that is about 6×10−86 \times 10^{-8}, roughly seven significant decimal digits.

Range

The exponent field sets how many binades the format spans, its . FP32’s 8 exponent bits reach from about 1.2×10−381.2 \times 10^{-38} to 3.4×10383.4 \times 10^{38}. Values beyond the top either become infinity or are clamped to the largest value, depending on the format and mode. Values below the smallest normal number become numbers, which give up precision gradually instead of jumping straight to zero.

Rounding

When a value falls between two representable numbers, the standard rule is : pick the closer one, and on an exact tie pick the one whose last bit is 0. The OCP 8-bit and MX specifications both require this mode for conversions.

Encoding

An ExMy format has 1 sign bit, xx exponent bits and yy mantissa (trailing significand) bits. With bias bb, a normal encoding (E>0E > 0) has value (−1)S×2E−b×(1+2−mM)(-1)^S \times 2^{E-b} \times (1 + 2^{-m} M); a subnormal (E=0E = 0) has (−1)S×21−b×(0+2−mM)(-1)^S \times 2^{1-b} \times (0 + 2^{-m} M). The hidden bit gives precision p=m+1p = m + 1. Rounding to nearest bounds the relative error of any in-range normal value by ε=2−p\varepsilon = 2^{-p} (Goldberg’s (β/2) β−p(\beta/2)\,\beta^{-p} with β=2\beta = 2).

Range, precision and special values

The exponent field buys binades; the mantissa buys resolution within a binade. Two other choices change the count of usable values:

  • Reserved codes. IEEE formats spend the all-ones exponent on ±∞\pm\infty and NaN. That is a whole binade, which matters little at 32 bits and a lot at 8. OCP E4M3 keeps only S.1111.111 as NaN, which lifts its maximum from 240 to 448 and its range from 17 to 18 binades. The MX FP6 and FP4 types reserve nothing at all.
  • Subnormals. They add mm binades of gradually decreasing precision below the smallest normal. In E4M3 they extend the bottom from 2−62^{-6} to 2−92^{-9}. They cost hardware (a leading-zero count and normalization shift), which is why some datapaths flush them to zero.

Two consequences shape the rest of the chapter. First, a float’s error is relative: the absolute step at magnitude ∣x∣|x| is about 2⌊log⁡2∣x∣⌋−m2^{\lfloor \log_2 |x| \rfloor - m}, so small values are represented as precisely (in relative terms) as large ones, until they fall out of the range. Second, the bias only positions the window; it doesn’t widen it. A tensor whose values span more binades than the format has will lose either its smallest or its largest values, whatever scale is applied. That is why range (exponent bits) and granularity of scaling are the two levers for low-precision formats.

FormatS/E/MBiasMax normalMin normalRelative step 2−m2^{-m}
FP321/8/231273.40×10383.40 \times 10^{38}1.17×10−381.17 \times 10^{-38}1.2×10−71.2 \times 10^{-7}
BF161/8/71273.39×10383.39 \times 10^{38}1.17×10−381.17 \times 10^{-38}7.8×10−37.8 \times 10^{-3}
FP161/5/101565,5046.10×10−56.10 \times 10^{-5}9.8×10−49.8 \times 10^{-4}
FP8 E4M31/4/374482−62^{-6}0.125
FP8 E5M21/5/21557,3442−142^{-14}0.25
FP4 E2M11/2/11610.5

Values from the BF16 study’s comparison table, the OCP FP8 specification and the OCP MX specification.

Sexponentmantissaevery positive normal value (log scale)0.00010.01110010,000range: 14 binadeszoom: the binade from 4 to 86.6→ 6.5488 steps of 0.5: 6.6 → 6.5

E4M3: normal values from 0.016 to 240 (14 binades), 8 steps per binade. 6.6 rounds to 6.5 (1.5% off).

One 8-bit float with a movable split. Top: every positive normal value on a log scale (each mark is one value). Bottom: the steps between 4 and 8, where 6.6 has to round.Share freely with credit: ‘Figure from chipfieldguide.com’

When AI training moved from 32 bits to 16, there were two ways to cut the bits in half. Both are still used.

, the older one, keeps more digits but can’t reach very big or very tiny numbers. Anything smaller than about 0.00000006 turns into zero.

, invented at Google, keeps the full size range of the 32-bit format and gives up digits instead. It’s simply the first half of a 32-bit number.

That matters for training. As a network learns, it makes tiny fixes to its numbers. Many fixes are so small that FP16 turns them into zero, and the learning stalls. BF16 can hold them, which is a big reason it became the favorite for training.

Training still keeps a few things in 32 bits, like the main copy of the model’s numbers. Mixing sizes like this cuts the memory a model needs about in half, without hurting the result.

FP16: precision without range

IEEE half precision, , has 1 sign, 5 exponent and 10 mantissa bits. Its largest value is 65,504, its smallest normal value about 6.1×10−56.1 \times 10^{-5}, and its smallest subnormal about 6×10−86 \times 10^{-8}. That is plenty for weights and activations, but not for the gradients computed during training (the small corrections that tell each weight which way to move). In one published analysis about 5% of weight-gradient values were smaller than 2−242^{-24} and would become zero in FP16; in a detection model, many activation-gradient values fell below FP16’s smallest number, and the network failed to train in FP16 without a fix.

The 2018 mixed-precision recipe fixed this with three techniques:

  1. An FP32 master copy of the weights. Updates are often too small to register on an FP16 weight, so they are applied to an FP32 copy, which is rounded to FP16 for the next step.
  2. . Multiply the loss by a constant before backpropagation, so every gradient is scaled up by the same factor and lands inside FP16’s range; divide it back out before the update. For one detection model, a factor of just 8 was enough to match FP32.
  3. FP32 accumulation. Multiply in FP16 but add the products into FP32 sums, converting to FP16 only when writing results to memory.

The result matched FP32 accuracy across many networks and nearly halved memory use.

BF16: range without precision

(“brain floating point”) takes the opposite trade: 1 sign, 8 exponent and 7 mantissa bits. It is simply FP32 with the bottom 16 mantissa bits dropped, so it has FP32’s full range and converting between them is trivial. It started as a storage format for shrinking the data exchanged between machines in distributed training, then became a compute format. A 2019 study trained image, speech, language and recommendation models in BF16 and matched FP32 in the same number of iterations with no hyperparameter changes, because no loss scaling is needed.

BF16 is coarse: with 7 mantissa bits, the step between neighboring values is 1/1281/128 of the number, so 1+1/2561 + 1/256 rounds back to 1. Two to three significant decimal digits is enough for the inputs of a matrix multiplication, but not for the sum, so BF16 is also paired with FP32 accumulation. Google’s TPUs, for example, multiply in BF16 and accumulate in FP32.

Two splits of 15 bits

FP16 (E5M10, bias 15) and BF16 (E8M7, bias 127) spend the same 15 non-sign bits differently. FP16 has p=11p = 11 and spans 2−242^{-24} to 65,504 (about 40 binades including subnormals); BF16 has p=8p = 8 and spans the whole FP32 normal range, about 2−1262^{-126} to 3.4×10383.4 \times 10^{38}. FP16 has 8× finer resolution within a binade; BF16 has 254 binades of normal numbers against FP16’s 30. Which one wins depends on whether a tensor’s problem is resolution or range.

  • Gradients are range-limited. Micikevicius et al. found weight gradients whose exponents fell below −24 (about 5% of values in one network), and an SSD detector, many of whose activation gradients (recorded in FP32 training) fell below FP16’s smallest subnormal, that diverged in FP16 without scaling; a loss scale of 8 (three binades) restored FP32 accuracy. Automatic mixed-precision tools adjust the loss scale at run time by provoking and detecting overflows.
  • Weight updates are resolution-limited. An update (learning rate × gradient) smaller than half an ulp of the weight vanishes under round-to-nearest, which is why the master weights stay in FP32 even with BF16.
  • GEMM inputs tolerate either, as long as the products are summed at higher precision. Both the FP16 recipe and BF16 practice accumulate in FP32.

BF16’s hardware argument is the multiplier: an 8-bit significand multiplier instead of 11 or 24 bits. Google states that a BF16 multiplier is about half the area of an FP16 one and one eighth of an FP32 one, because multiplier size scales with the square of the mantissa width. The (pBF16/pFP16)2=(8/11)2≈0.53(p_{\mathrm{BF16}}/p_{\mathrm{FP16}})^2 = (8/11)^2 \approx 0.53 and (8/24)2≈0.11(8/24)^2 \approx 0.11 ratios agree with that to first order. The conversion from FP32 is a rounding of the low 16 bits; the BF16 study used round-to-nearest-even.

FP16subnormal: 1.0% offBF16kept, 0.24% off10⁻⁹10⁻⁶10⁻³110³gradientlog scale; pale = subnormal (fewer digits)
Loss scaling

Gradient 0.0000012. FP16: stored as 0.0000013; BF16: 0.0000012. FP16 has more digits but reaches only down to about 0.00000006.

Same 16 bits, different split: FP16 (5 exponent, 10 mantissa) against BF16 (8 exponent, 7 mantissa). Bars show each format’s range on a log scale; the pale part is FP16’s subnormals. Loss scaling multiplies by 1,024 before rounding.Share freely with credit: ‘Figure from chipfieldguide.com’

Eight bits

With 8 bits there are only 256 possible values. AI uses two kinds. FP8 is a tiny floating-point number, with size bits and digit bits. INT8 holds only whole numbers, from −127 to 127, evenly spaced like marks on a ruler.

The shared multiplier

Neither one can cover the real numbers in a model by itself. So a whole block of numbers gets one extra number stored with it, called a . Multiply each stored value by it to get the real value. It’s usually picked so the biggest number in the block just fits.

That works well until one number is much bigger than the rest. Chatbot models have a few of these giants, up to about 20 times bigger than the others, and they matter. If the scale stretches to fit a giant, the normal numbers get squeezed onto a few ruler marks. Many of them round to zero.

FP8: two variants

In 2022 engineers from NVIDIA, Arm and Intel proposed two 8-bit floating-point encodings, and in 2023 they were published as an Open Compute Project standard, OFP8.

  • E4M3: 4 exponent bits, 3 mantissa bits. Largest value 448, smallest 2−9≈0.0022^{-9} \approx 0.002, 18 binades. Recommended for weights and activations.
  • E5M2: 5 exponent bits, 2 mantissa bits. Largest value 57,344, smallest 2−162^{-16}, 32 binades. Recommended for gradients, which need range more than precision.

Even 32 binades isn’t enough to hold every tensor in a network with a fixed bias, so each tensor gets a stored in higher precision. Values are multiplied by the scale before conversion so that the tensor’s largest magnitude lands near the format’s maximum; anything that still overflows is saturated to the maximum rather than turned into infinity. With this recipe, FP8 training matched 16-bit results on image, speech and language models up to 175 billion parameters, with no hyperparameter changes.

Integers and quantization

For running trained models (inference), the long-standing alternative is . Each real value rr is represented by an integer qq through a scale SS and a zero-point ZZ: r=S×(q−Z)r = S \times (q - Z). With symmetric quantization Z=0Z = 0, and SS is typically chosen as max⁡∣r∣/127\max|r| / 127 for INT8. A small example: if a tensor’s largest magnitude is 2.54, then S=0.02S = 0.02, the value 0.73 is stored as round(0.73/0.02)=round(36.5)=36\mathrm{round}(0.73 / 0.02) = \mathrm{round}(36.5) = 36 (a tie, rounded to even), and it comes back as 0.72.

Integers have two advantages. The products of two INT8 values are exact integers that can be summed exactly in a 32-bit integer , and integer multipliers are smaller than floating-point ones (more on that below). An 8-bit quantization workflow published in 2020 kept every network studied within 1% of its floating-point accuracy.

Rounding error against clipping error

Choosing the scale is a trade-off. A larger scale covers more of the tensor, so fewer values are clipped at the ends, but each step is wider, so every value suffers more rounding error. A smaller scale does the reverse. Practical tools choose the range from sample data, using the maximum, a high percentile, or a measure that minimizes information loss.

Outliers

Large language models make this trade-off much harder. Beyond a few billion parameters, a small number of hidden dimensions carry values up to about 20 times larger than the rest, in every layer. They are only about 0.1% of the features, but setting them to zero destroys the model’s accuracy. With one scale per tensor, those outliers set the step size, and the ordinary values are crushed onto a few levels. Two responses followed: keep the outlier dimensions in 16-bit and the remaining 99.9% in 8-bit, or use a separate scale for smaller pieces of the tensor (per row, per channel, or per block). The second idea, taken to small blocks, is microscaling.

OFP8 in detail

E4M3 has bias 7, no infinities and a single NaN mantissa pattern (S.1111.111), which extends it to 448 and 18 binades; E5M2 has bias 15 and IEEE-style specials (S.11111.00 is ±∞\pm\infty), reaching 57,344 and 32 binades. E5M2 is effectively FP16 with 8 fewer mantissa bits, which makes conversion trivial. Conversion from wider types rounds to nearest even and then applies one of two overflow modes: saturating (clamp to ±max⁡\pm\max) or non-saturating (E4M3 → NaN, E5M2 → ±∞\pm\infty). Training recipes use saturation, because the loss-scaling trick of skipping a step on overflow would skip too many steps with FP8’s narrow range.

The designers kept the IEEE-like biases and left scaling to software: a per-tensor scale in higher precision can take any real value, while a programmable exponent bias is equivalent to power-of-two scales only. A scale chosen so that amax maps near the format’s maximum is applied before the cast, and its inverse is applied once per dot product (or folded into the next operation), so the cost is amortized over many multiply-accumulates.

Integer quantization in detail

The affine scheme r=S(q−Z)r = S(q - Z) makes real zero exactly representable (q=Zq = Z) and expands an integer GEMM into an int8 × int8 core plus correction terms that cost O(N2)O(N^2) for an N3N^3 multiplication. Symmetric (Z=0Z = 0) quantization avoids the corrections, and an NVIDIA study found it sufficient for INT8 on all networks it examined. The quantizer is

q=clamp(round(x/s)+z, qmin⁡, qmax⁡),x^=s(q−z)\begin{aligned} q &= \mathrm{clamp}\big(\mathrm{round}(x/s) + z,\ q_{\min},\ q_{\max}\big), \\ \hat{x} &= s(q - z) \end{aligned}

with two error sources: rounding error, uniform within ±s/2\pm s/2 for in-range values, and clipping error for values outside [s(qmin⁡−z), s(qmax⁡−z)][s(q_{\min} - z),\ s(q_{\max} - z)]. Raising ss trades the second for the first. Calibration picks ss from activation statistics (max, entropy, or a percentile such as 99.99%). Granularity is the other knob: one scale per tensor, per output channel (common for weights), or per group of elements; finer groups generally improve accuracy at some cost in overhead.

Float versus integer at 8 bits

For the same 8 bits, the difference between INT8 and FP8 is only where the grid points sit: uniform for INT8, denser near zero and sparser far out for FP8. So floats do better on distributions with heavy tails and outliers, and integers on uniform ones; on a Gaussian, the best low-exponent float formats and INT8 come out close. The simulator reproduces this: on its near-Gaussian “Weights” tensor (only two mild outliers), INT8 with a per-tensor scale has lower error than E4M3, and on the outlier-heavy “Activations” tensor the order flips.

Emergent outliers

Dettmers et al. found that in transformers, large-magnitude features (up to ~20× the other dimensions) appear first in about a quarter of layers and, around 6.7B parameters, in all of them. They occupy about 0.1% of feature dimensions; zeroing them degrades perplexity by 600–1000%. They are highly systematic: at 6.7B parameters about 150,000 outliers per sequence fall in only 6 feature dimensions. LLM.int8() exploits that, sending those dimensions through a 16-bit GEMM and quantizing the rest to 8 bits with a separate scale per row and column. Fine-grained block scaling attacks the same problem by limiting how many values each outlier can hurt.

zoom: −1.4 to 1.4−101originalafter roundingoutlier20step s = 0.15752 of 7 flushed to 0
Format

INT8, one scale: s = 20 / 127 = 0.1575. Ordinary values snap to multiples of s; 2 of 7 round to 0, RMS error 5% of their typical size.

A per-tensor scale is set by the largest magnitude. Ticks are the representable values near zero; each ordinary value (top) snaps to the nearest one (bottom). Illustrative values.Share freely with credit: ‘Figure from chipfieldguide.com’

A multiplier for every small group

The fix for giants is simple: give every small group of numbers its own multiplier. Then a giant only spoils its own group. This is called . In 2023, a group of chip and cloud companies agreed on a shared standard for it.

In the standard, every 32 numbers share one multiplier. The numbers themselves can be as small as 4 bits each.

Four bits

The 4-bit format is called . It can only hold 0, 0.5, 1, 1.5, 2, 3, 4 and 6, plus their negatives. That sounds hopeless. But with a fresh multiplier for every 32 numbers, it’s enough to store a big chatbot model.

One openly released model stores most of its numbers this way. That shrinks it enough to fit 120 billion numbers in the memory of a single AI chip.

The MX specification

The Open Compute Project’s Microscaling Formats specification (version 1.0, September 2023), written by engineers from AMD, Arm, Intel, Meta, Microsoft, NVIDIA and Qualcomm, defines a family of block-scaled formats. An MX block has three parts:

  • kk elements, all in the same narrow format. For every concrete MX format k=32k = 32.
  • One shared scale XX, in a format called E8M0: 8 bits holding just an exponent, so XX is a power of two between 2−1272^{-127} and 21272^{127}.
  • The value of element ii is simply X×PiX \times P_i.
FormatElementElement bitsBlockBits per value
MXFP8FP8 E4M3 or E5M28328.25
MXFP6FP6 E2M3 or E3M26326.25
MXFP4FP4 E2M14324.25
MXINT8INT88328.25

Formats and block size from the specification’s Table 1; bits per value = element bits + 8/32.

Choosing the scale

The specification’s conversion rule: set XX to the largest power of two not exceeding the block’s largest magnitude, divided by the largest power of two the element format can represent. Then divide every value by XX and round it to the element format, clamping anything that still exceeds the format’s maximum. For FP4, whose largest power of two is 4, a block whose biggest value is 20 gets X=16/4=4X = 16 / 4 = 4, and its elements are stored as multiples of 4×{0,0.5,1,1.5,2,3,4,6}4 \times \{0, 0.5, 1, 1.5, 2, 3, 4, 6\}.

Why blocks of 32

A scale per tensor works for 8-bit formats but has been shown to be insufficient below 8 bits, because the formats have too little range to cover a whole tensor. A scale per 32 values follows the data closely while costing only a quarter of a bit per value. In the specification’s accompanying study, 8-bit MX formats ran inference directly on FP32-trained models with minimal accuracy loss and no calibration; 6-bit MX formats trained large transformers to FP32 accuracy; and training with 4-bit MX weights cost only a minor accuracy drop.

FP4 and INT4

FP4 (E2M1) has 1 sign, 2 exponent and 1 mantissa bit, no infinities or NaN, and the magnitudes 0, 0.5, 1, 1.5, 2, 3, 4 and 6. INT4 has the evenly spaced integers −7 to 7 (or −8 to 7). Four-bit weights are attractive because, for a fixed memory budget, a bigger model with 4-bit weights tends to beat a smaller model with 8-bit weights; a study of more than 35,000 experiments found 4 bits almost universally optimal for total model bits, with small block sizes and the choice of data type being the main ways to improve further.

Variations on the MX idea exist. NVIDIA’s NVFP4 uses the same E2M1 elements but blocks of 16 values, a scale stored in E4M3 (so it is not limited to powers of two), and an extra FP32 scale per tensor, at 4.5 bits per value.

Semantics

An MX format is a triple: scale type (ww bits), element type (dd bits) and block size kk; a block costs w+kdw + kd bits, and its memory layout is not prescribed. The concrete formats all use k=32k = 32 and an E8M0 scale: unsigned, bias 127, exponents −127 to 127, a single NaN code (0xFF) that poisons the whole block, no infinities and no zero. E8M0 covers a superset of FP32’s exponent range. FP6 (E2M3, max 7.5; E3M2, max 28) and FP4 (E2M1, max 6) reserve no codes; conversion to them must support round-to-nearest-even and saturation to ±max⁡\pm\max. MXINT8 elements are two’s complement with an implicit 2−62^{-6} scale, i.e. a sign, one integer bit and six fraction bits.

Conversion

OCP MX conversion of one block (spec §6.3, as Algorithm 1 in Rouhani et al.)text
emax_elem  = exponent of the element format's largest normal
             (E4M3: 8, E5M2: 15, E2M1: 2, E3M2: 4, E2M3: 2, INT8: 0)
shared_exp = floor(log2(max_i |V_i|)) - emax_elem
X          = 2^shared_exp                   # stored as E8M0: shared_exp + 127
P_i        = round_to_element(V_i / X)      # ties to even; clamp normals
                                            # above the element max to +/-max
value_i    = X * P_i
  1. 1L3The block’s largest magnitude lands in the element format’s top binade, so the full exponent range of the element is used.
  2. 2L5The scale is a power of two: applying it is an exponent add, not a multiply.
  3. 3L6Because the floor rounds the block maximum’s exponent down, V/X can exceed the element max (e.g. 6 < V/X < 8 in FP4) and is clamped.

The rule follows the specification’s minimal semantics; implementations may use other algorithms. The scale is shared along one axis, usually the reduction dimension of the matrix multiplication, so converting and transposing do not commute: a matrix that is reduced along different axes in the forward and backward passes has to be converted separately for each.

The dot product

For two MX blocks, Dot(A,B)=XAXB∑iPA,iPB,i\mathrm{Dot}(A, B) = X_A X_B \sum_i P_{A,i} P_{B,i}: the element products are reduced first and the two power-of-two scales applied once. A longer dot product is a sum of block dot products and should produce an FP32 result; the internal precision and order of operations are implementation-defined. In hardware, the inner sum is a small fixed-point or narrow-float adder tree over 32 products, and the scale is an exponent addition before the result joins a wider accumulator.

The power-of-two penalty and NVFP4

Because ⌊log⁡2amax⌋\lfloor \log_2 \mathrm{amax} \rfloor rounds down, amax/X\mathrm{amax}/X lies in [2emax⁡,2emax⁡+1)[2^{e_{\max}}, 2^{e_{\max}+1}). For E2M1 that is [4,8)[4, 8), and everything in (6,8)(6, 8) saturates to 6: the block maximum itself can lose up to 25%. Rounding the exponent up instead avoids clipping but leaves up to one binade of the tiny element range unused. NVIDIA’s NVFP4 paper describes MXFP4 as potentially losing up to one binade of dynamic range because of power-of-two scale rounding, and changes three things: 16-element blocks, E4M3 block scales with fractional precision, and a second, per-tensor FP32 scale that keeps the E4M3 scales in range. At least one value per block (its amax) is then represented at near-FP8 precision. The cost is 4+8/16=4.54 + 8/16 = 4.5 bits per value instead of 4.25.

What 4 bits has been shown to do

  • Inference: Dettmers and Zettlemoyer’s bit-level scaling laws (19M–176B parameters) found 4-bit weights almost universally optimal for accuracy per total model bit; below that, 3-bit reverses the trend.
  • Deployment: the gpt-oss models store their mixture-of-experts weights, over 90% of parameters, in MXFP4 at 4.25 bits per parameter.
  • Training: a 12B-parameter model pretrained on 10 trillion tokens with NVFP4 GEMMs reached 62.58% on MMLU-pro against 62.62% for an FP8 baseline. It needed Random Hadamard transforms to spread out block outliers, two-dimensional block scaling, stochastic rounding for gradients, and a few layers kept in higher precision.
Block 1×40.80−1.3−20.402.12−0.60201610−0.20Block 2×40.50−0.901.62−0.300.70−1.2−20.202.42top: original · bottom: X × FP4 valueE2M1 magnitudes: 0, 0.5, 1, 1.5, 2, 3, 4, 610 of 15 ordinary values → 0
Scale granularity

One scale for all 16: X = 4. FP4 steps are multiples of X/2, so 10 of 15 ordinary values round to 0 (5 in block 1, 5 in block 2).

FP4 (E2M1) values with power-of-two scales chosen by the MX rule. Groups of 8 here to fit; the MX standard uses 32, at 4 + 8/32 = 4.25 bits per value. Illustrative values.Share freely with credit: ‘Figure from chipfieldguide.com’

Smaller multipliers

A chip multiplies the way you do on paper: every digit of one number times every digit of the other, then add up the rows.

So fewer digits shrink the multiplier fast. Going from 32-bit numbers to BF16 makes it about eight or nine times smaller. The same chip can then fit many more of them.

Less to move

Moving numbers costs even more energy than multiplying them. Fetching one number from memory can use about 200 times the energy of the math it feeds. Half the bits means half as much to move.

The total stays big

There’s one place chips don’t skimp. When thousands of results are added up, the running total gets many more bits, often 32. A small total would lose the tiny additions once it got big. It would be like weighing a feather on a truck scale.

Multiplier cost grows with the square of the width

Multiplying two pp-bit significands takes roughly p×pp \times p one-bit partial products, which are then added together in an array of adders. So the area and energy of a multiplier grow roughly with p2p^2, while an adder grows roughly with pp. Google explains BF16 this way: its multipliers are about half the size of FP16’s and eight times smaller than FP32’s. Measured figures quoted in a widely used survey put numbers on it: an 8-bit integer multiply uses 18.5 times less energy and 27.5 times less area than a 32-bit floating-point multiply, and an 8-bit add 30 times less energy than a 32-bit floating-point add.

Integer versus floating point at the same width

A floating-point multiply needs more than the significand multiplier: it adds the exponents, and before products can be summed they have to be shifted to line up. An integer multiply needs neither. One gate-count estimate found an FP8 (E4M3) multiply-accumulate about 53% more expensive than INT8 with the same kind of accumulator, and 183% more if it has to accumulate in FP32. That is the case for integers in inference chips, and the price paid for FP8’s wider range.

Throughput and bandwidth

Chips turn the smaller multipliers into throughput. On one GPU generation, published figures give FP16 8 times and INT8 16 times the FP32 matrix throughput, with 2× and 4× less data to move per value. Data movement is the bigger energy item: a DRAM access costs orders of magnitude more energy than an arithmetic operation, so in memory-bound work (where is low), the bandwidth saving is what makes narrower formats faster.

Accumulation precision

Every output of a matrix multiplication is a dot product: thousands of products added into one . The inputs can be narrow; the sum cannot.

  • Integers accumulate exactly if the accumulator is wide enough: INT8 inputs use a 32-bit integer accumulator.
  • Floats lose information in long sums. Adding a small number to a large one shifts the small one right to align the binary points, and bits that fall off the end are gone. Once the running sum is more than about 2m+12^{m+1} times bigger than the next term, the term disappears entirely, an effect called .

That is why every recipe in this chapter pairs narrow inputs with FP32 (or wider) accumulation: FP16 and BF16 training, FP8 matrix units, which offer FP16 or FP32 accumulators, and the MX specification, whose general dot product should produce an FP32 result. The cost is modest: there is one accumulator per output, but hundreds or thousands of multiplications feeding it.

A first-order cost model

For a floating-point multiply-add the main blocks are: a p×pp \times p significand multiplier (area∝p2\text{area} \propto p^2), an exponent adder (∝e\propto e), an alignment shifter and adder for the accumulation (∝accumulator width\propto \text{accumulator width}, with log-depth shifting), and normalization and rounding. For integers only the multiplier and a fixed-point adder remain. The survey by Sze et al., citing Horowitz’s ISSCC 2014 figures, states the scaling explicitly: multiplier energy and area scale roughly quadratically with bits, adder and memory cost roughly linearly.

FormatSignificand ppp2p^2 (multiplier, relative)vs FP32
FP32245761
FP16111210.21
BF168640.11
INT88640.11
FP8 E4M34160.028
FP8 E5M2390.016
FP4 E2M1240.007

This is the multiplier array only. It explains why BF16 multipliers are about half of FP16’s and an eighth of FP32’s, but it overstates the savings for tiny floats, whose cost is dominated by the shifters, the exponent logic and above all the accumulation. The van Baalen et al. gate-count study makes the point: FP8-E4 with a fixed-point accumulator is 53% more expensive than INT8, and with an FP32 accumulator 183% more. Their INT8 design needs a 27-bit fixed-point accumulator (15 bits of product plus 12 bits of headroom for 4,096 products); FP8-E4, whose products span many more binades, would need 37 bits. That is consistent with published peak figures, where halving the input width roughly doubles matrix throughput rather than quadrupling it.

Accumulation error

Floating-point accumulation is inexact and order-dependent, and the error grows with the reduction length KK. The failure mode is : when the running sum exceeds an addend by more than 2m+12^{m+1}, the addend is shifted out entirely, and the problem worsens with long sums and non-zero-mean data. Wang et al. showed that splitting a dot product into chunks summed separately and then combined (chunk-based accumulation), plus , let FP8 training use 16-bit additions.

A real-world example: DeepSeek-V3’s authors measured that FP8 GEMMs on the GPU they used retain only about 14 bits of accumulation precision inside the matrix unit, giving a maximum relative error of nearly 2% for K=4096K = 4096. Their fix copies partial sums to FP32 registers every 128 elements of KK, which is chunked accumulation implemented in software. The general lesson is that the narrow-format story is really two stories: input precision, which sets multiplier cost and bandwidth, and accumulation precision, which sets numerical quality and is much cheaper to keep high.

p = 8: 8 × 8 partial productsdashed: FP32 (24 × 24)64 partial productsmultiplier area ∝ p²0.11 of FP32bytes moved ∝ bits1/2
Format

BF16: p = 8, 64 partial products, about one ninth of FP32 (Google says about eight times smaller).

Each dot is a one-bit partial product of a p × p significand multiplier. The faint grid is FP32’s, for scale. Area and energy grow roughly with p², data moved with the bit count.Share freely with credit: ‘Figure from chipfieldguide.com’
002,0482,0484,0964,096N →running sumexact sum1,024BF16 accumulator256768 addends loststalls at 2^8
Accumulator
Promote to FP32 every 128

BF16 running sum: stuck at 256 after 1,024 additions of 1. With 7 mantissa bits, a sum of 256 has a step of 2, so +1 rounds away.

The running sum of N ones, rounded to the accumulator format after every addition (round to nearest even). The dashed line is the exact sum. Chunks: sum 128 terms, then add the chunk into FP32.Share freely with credit: ‘Figure from chipfieldguide.com’

The simulator holds 256 made-up numbers, shaped like the real ones in a model. Most are small, and a few are giants (the red triangles). Pick a format along the top. The gray bars are the original numbers, and the pink bars are the rounded copies.

Try INT4 on “Big outliers”: almost everything turns into zero. Then try MXFP4, which gives each group of 32 numbers its own multiplier. Which groups turn red?

Choose a format and a sample tensor. The histogram shows the input values (gray) and the quantized values (pink); blue ticks under the axis are the representable values at the current scale. The bit panel decodes one value field by field. On “Activations + outliers”, compare INT8 with one scale for the tensor (RMS error 4.8% of σ\sigma, 17 values flushed to zero) with MXFP8 (2.7%), then INT4 (78%, 214 of 256 values zeroed) with MXFP4 (32%). The colored strip shows each scale group’s error: with MX formats only the blocks holding an outlier turn red. On “Gradients (tiny)”, compare FP16 (six values flushed to zero) with BF16.

Encodings follow OFP8 and MX v1.0 exactly: round-to-nearest-even, subnormals, saturation, E8M0 scales computed as 2⌊log⁡2amax⌋−emax⁡,elem2^{\lfloor \log_2 \mathrm{amax} \rfloor - e_{\max,\mathrm{elem}}}. FP8 and integer formats use an FP32 per-tensor scale mapping amax to the format maximum. Click bits to decode arbitrary patterns (try S.1111.111 in E4M3, or the all-ones exponent in E5M2). The block-size slider sweeps MX from k=8k = 8 to 256: MXFP4 on the activations goes from 18% to 32% to 52% RMS error of σ\sigma as kk grows from 8 to 32 to 256, while the scale overhead falls from 1 to 0.03 bits per value. Watch the clipped count too: a block maximum just below the next power of two saturates in E2M1 and E4M3 alike. The readouts include SQNR and a p2p^2 multiplier-cost estimate.

Loading simulation…
Bits per number, classic format
32
Bits per number, common for training today
16 or 8
Values an FP4 number can take
15
Numbers sharing one multiplier
32

What these numbers mean:

  • The classic format uses 32 bits per number. AI training moved to 16, but kept 32 bits for the main copy of the model.
  • Tests show that 8-bit numbers can train models as well as 16-bit ones, even very large chatbot models.
  • FP4 has 16 patterns of bits. But plus zero and minus zero are the same number, which leaves 15 values, from −6 to 6.
  • In the shared standard, 32 numbers share one multiplier. That adds only a quarter of a bit to each one.
Largest FP16 / E4M3 / E5M2 value
65,504 / 448 / 57,344
MXFP4 bits per value
4.25
INT8 vs FP32 multiply energy
18.5× less
FP8-E4 MAC cost vs INT8 (gate estimate)
+53%
FormatBitsSign / exponent / mantissaRange (largest)Typical use
FP32321 / 8 / 233.4×10383.4 \times 10^{38}Master weights, accumulators
BF16161 / 8 / 73.4×10383.4 \times 10^{38}Training
FP16161 / 5 / 1065,504Training with loss scaling; inference
FP8 E4M381 / 4 / 3448Weights and activations
FP8 E5M281 / 5 / 257,344Gradients
INT88two’s complement127 × scaleInference
MXFP44+8/324 + 8/321 / 2 / 1, shared E8M06×21276 \times 2^{127}Weights for inference; research training

Bit layouts and ranges from the OCP specifications and the BF16 study; uses as recommended by the FP8 and MX papers.

Energy. An 8-bit integer multiply uses 18.5× less energy and 27.5× less area than a 32-bit floating-point multiply; an 8-bit add uses 30× less energy than a 32-bit float add. Throughput. On one GPU generation, INT8 matrix math ran at 16× the FP32 rate and INT4 at 32×.

E4M3 / E5M2 dynamic range
18 / 32 binades
FP8 GEMM accumulation measured on one GPU
≈ 14 bits
NVFP4 vs FP8 pretraining, MMLU-pro (12B, 10T tokens)
62.58% vs 62.62%
MXFP4 share of gpt-oss parameters
over 90%

Format parameters

FormatBiasMaxMin subnormalSpecials
FP161565,5042−24≈5.96×10−82^{-24} \approx 5.96 \times 10^{-8}IEEE ±∞\pm\infty, NaN
E4M374482−92^{-9}NaN only (S.1111.111)
E5M21557,3442−162^{-16}IEEE ±∞\pm\infty, NaN
E3M2 (FP6)3280.0625none
E2M3 (FP6)17.50.125none
E2M1 (FP4)160.5none
E8M0 (scale)12721272^{127}2−1272^{-127} (no zero)one NaN (0xFF)

From the FP16 column of the BF16 study, the OFP8 specification and the MX specification.

Published results worth knowing

  • FP8 (E4M3 forward, E5M2 gradients, per-tensor scales) matched 16-bit training on CNNs, RNNs and transformers up to 175B parameters with unchanged hyperparameters.
  • DeepSeek-V3 (671B total parameters) trained with FP8 GEMMs, E4M3 on all tensors, 1 × 128 activation tiles and 128 × 128 weight blocks; in tests at two smaller scales the FP8 model’s loss stayed within 0.25% of a BF16 baseline.
  • 8-bit MX direct-cast inference on FP32-trained models with minimal loss; 6-bit MX training at FP32 quality.
  • Emergent outliers in LLMs: about 0.1% of features, up to 20× larger, in all layers from about 6.7B parameters.

Every bit you remove makes the chip faster and the model smaller. It also makes each number a little less exact. Models can handle a lot of rounding, but not endless amounts. Four bits only works with careful tricks, and fewer than that is much harder.

The sneaky part is how things go wrong. Tiny numbers quietly become zero. A few giants can blur everything else. Nothing crashes; the model just gets a bit worse. That’s why teams test a model carefully after changing its number format.

What you give up for what you get

ChoiceYou getYou give up
Fewer bitsLess memory and bandwidth, smaller multipliersPrecision; more engineering to keep accuracy
More exponent bitsRange: fewer values lost to underflow or overflowPrecision within each power of two
Integer instead of floatSmaller, exact multiply-accumulate hardwareUniform grid that handles outliers and tiny values poorly
Finer scaling (per block)Outliers only hurt their own blockScale storage (8/k8/k bits per value) and per-block hardware
Wide accumulatorAccurate long dot productsSome area and energy per output (small next to the multipliers)

How it goes wrong

  • Underflow. Values below the smallest representable number become zero, as FP16 gradients do without loss scaling.
  • Overflow. Values above the maximum become infinity or NaN unless the conversion saturates; one NaN then spreads through every result it touches. FP8 recipes saturate for this reason.
  • Outliers. With one scale per tensor, a few huge values set the step size for everything and ordinary values lose most of their precision.
  • Accumulation loss. Narrow running sums drop small terms; errors grow with the length of the dot product.
  • Calibration mismatch. A scale chosen from sample data can be wrong for real inputs, clipping values the calibration set never showed.

Each failure shows up as a quiet drop in accuracy rather than a crash, so format changes are validated against the full-precision model on real evaluation tasks. Peak speeds quoted for a new format only count if the model still works in it.

The design space

  • Element width and type. At fixed bits, exponent versus mantissa trades range for SQNR (about 6 dB per mantissa bit). Integers maximize resolution for a known range; floats tolerate unknown or heavy-tailed distributions.
  • Scale granularity and type. Per-tensor FP32 scales are free in storage but expose every value to the tensor’s worst outlier. Per-block scales cost 8/k8/k bits per value; power-of-two scales cost only an exponent add but can waste up to a binade; E4M3 scales are more precise but need a second-level tensor scale for range.
  • Accumulator width. Exact fixed-point accumulation is cheapest for integers; float inputs with many exponent bits need very wide fixed-point accumulators or a floating accumulator, which is where FP8 loses much of its hardware advantage over INT8.
  • Software burden. Every step down the ladder has added machinery: master weights and loss scaling for FP16, per-tensor scaling for FP8, tile and block scaling plus promoted accumulation in FP8 training at scale, and Hadamard rotations, stochastic rounding and high-precision layers for FP4 training.

Failure modes in detail

  • Range failures are binary and data-dependent: a tensor that fits during calibration may overflow on an unusual input. Saturating conversion bounds the damage; non-saturating conversion turns it into NaN or ∞\infty.
  • Resolution failures are statistical: SQNR drops gradually, and the effect on accuracy depends on the layer. Operations other than matrix multiplications, such as nonlinearities and normalizations, typically run in higher precision anyway, and the NVFP4 recipe also kept a few linear layers in higher precision for stability.
  • Systematic bias: round-to-nearest of updates smaller than half an ulp is always zero, so many small updates are lost rather than averaged; stochastic rounding removes that bias at the cost of noise and a random-number source.
  • Hardware-specific numerics: the MX spec leaves dot-product internal precision implementation-defined, so the same model in the same format can produce slightly different results on different chips. DeepSeek-V3’s accumulation measurement is an example of such a detail mattering at scale.
0minUnderflowmax→ max or ∞Overflowcoarse stepOutliersswampedAccumulationcalibrated rangeCalibration

Tap a failure mode. Each one shows up as a quiet drop in accuracy, not a crash, which is why format changes are validated on real tasks.

The five failure modes of narrow formats, with the usual fix for each. Schematic pictures, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

Rounding to a small float

Every format in this chapter can be produced by one routine: find the binade, round the significand to mm bits with ties to even, saturate, and handle the bottom of the range with subnormals. This is what the simulator runs, per the OFP8 and MX conversion rules.

quantize(x) to an ExMy format, saturating (illustrative pseudocode)text
emin = 1 - bias                      # exponent of the smallest normal
a    = |x|
e    = max(floor(log2(a)), emin)      # below emin: subnormal spacing
ulp  = 2^(e - m)                      # step size in this binade
r    = round_half_even(a / ulp) * ulp
if r > max_normal: r = max_normal     # SAT mode; NONSAT gives Inf or NaN
return sign(x) * r                    # r == 0 means the value underflowed
  1. 1L3Clamping e at emin is what makes subnormals: the step stops shrinking at 2^(emin − m).
  2. 2L4The step doubles each binade, so the relative error is bounded by 2^−(m+1), i.e. 2^−p.
  3. 3L5Ties go to the even neighbor; rounding up can carry into the next binade, which is still a valid value.
  4. 4L6OFP8 requires both saturating and non-saturating modes; MX FP6/FP4 must support saturation.

Noise model for a uniform grid

If values fall at random positions between grid points spaced Δ\Delta apart, the rounding error is uniform on [−Δ/2,Δ/2][-\Delta/2, \Delta/2] and its mean square is Δ2/12\Delta^2/12. For a symmetric bb-bit integer quantizer with a per-tensor scale set by the tensor’s largest magnitude, Δ=amax/qmax⁡\Delta = \mathrm{amax} / q_{\max} with qmax⁡=2b−1−1q_{\max} = 2^{b-1} - 1. For values with RMS σ\sigma, the signal-to-quantization-noise ratio () is

SQNR=10log⁡10 ⁣(12 σ2qmax⁡2amax2)≈10.8+20log⁡10qmax⁡−20log⁡10amaxσ  dB\begin{aligned} \mathrm{SQNR} &= 10 \log_{10}\!\left(\frac{12\,\sigma^2 q_{\max}^2}{\mathrm{amax}^2}\right) \\ &\approx 10.8 + 20 \log_{10} q_{\max} \\ &\quad - 20 \log_{10}\frac{\mathrm{amax}}{\sigma}\ \ \mathrm{dB} \end{aligned}

Three readings of this formula:

  • Each extra bit doubles qmax⁡q_{\max} and adds 20log⁡102≈6.0220 \log_{10} 2 \approx 6.02 dB.
  • Each doubling of the crest factor amax/σ\mathrm{amax}/\sigma costs the same 6 dB, a full bit. An outlier 20× the RMS costs about 26 dB, more than four bits.
  • For INT8 with amax/σ≈21\mathrm{amax}/\sigma \approx 21, as in the simulator’s activation tensor, the formula predicts 52.9−26.4≈26.552.9 - 26.4 \approx 26.5 dB on the ordinary values; the simulator measures an RMS error of 4.8% of σ\sigma, which is 26.3 dB.

Widening the range past amax\mathrm{amax} is never useful, but narrowing it below amax\mathrm{amax} trades this rounding noise for clipping error on the tail: the total error is roughly Δ(α)2/12\Delta(\alpha)^2/12 plus the mean of (∣x∣−α)2(|x| - \alpha)^2 over values beyond the clipping point α\alpha. Calibration methods choose α\alpha to balance the two.

Noise model for a float

Within a binade [2e,2e+1)[2^e, 2^{e+1}) the step is 2e−m2^{e-m}, so the error is still uniform, with mean square 22(e−m)/122^{2(e-m)}/12. Relative to x2x^2, and averaging 1/x21/x^2 over a value spread evenly through the binade (the mean of 1/u21/u^2 for uu uniform on [1,2)[1, 2) is 1/21/2), the relative noise power is 2−2m/242^{-2m}/24. That gives, independent of scale while values stay in range,

SQNR≈6.02 m+13.8 dB\mathrm{SQNR} \approx 6.02\,m + 13.8\ \mathrm{dB}

The linear dependence on mantissa bits matches the analysis of block formats by Rouhani et al. The prediction is 31.9 dB for E4M3 (m=3m = 3), 25.8 dB for E5M2 and 55.9 dB for BF16; the simulator measures 32–34, 26–28 and 54–62 dB on its three tensors. The model breaks down at the edges: values pushed into the subnormal range lose relative precision, values below the smallest subnormal are lost, and values above the maximum clip. That is why a float’s SQNR is flat across scales only within its dynamic range, and why the scale factor for FP8 matters even though the float grid is “self-scaling.”

Block scaling

Block scaling replaces the tensor’s crest factor with each block’s. For an MX block with scale X=2⌊log⁡2amax⌋−emax⁡X = 2^{\lfloor \log_2 \mathrm{amax} \rfloor - e_{\max}}, the block maximum lands in [2emax⁡,2emax⁡+1)[2^{e_{\max}}, 2^{e_{\max}+1}) of the element format: on average the top of the element range is under-used by about half a binade, and in FP4 a maximum in (6,8)⋅X(6, 8) \cdot X saturates. Shared-exponent analyses show the resulting SQNR growing linearly with mantissa bits and logarithmically with finer block granularity. The cost side is exact: 8/k8/k bits per value for an E8M0 scale. In the simulator, MXFP4 on the activation tensor falls from 18% to 32% to 52% RMS error of σ\sigma as kk goes from 8 to 32 to 256, for 1, 0.25 and 0.03 bits of overhead.

Accumulation error

Summing NN terms one after another in a float with unit roundoff u=2−pu = 2^{-p} gives an error bound that grows in proportion to NN. Splitting the sum into chunks of length CL\mathit{CL}, summing each, and then summing the partial sums reduces the bound to order N/CL+CLN/\mathit{CL} + \mathit{CL}, which is smallest when CL≈N\mathit{CL} \approx \sqrt{N}. For N=4096N = 4096 that is CL=64\mathit{CL} = 64; DeepSeek-V3’s promotion interval of 128 sits in the same range, and it was their answer to errors of nearly 2% measured at K=4096K = 4096.

For integers the question is width, not rounding: a bb-bit × bb-bit product needs 2b−12b - 1 bits, and summing 2g2^g of them needs gg more. The INT8 example in van Baalen et al. is 15+12=2715 + 12 = 27 bits for 4,096 products; a fixed-point accumulator that exactly covers FP8-E4 products would need 37 bits. That is why 32-bit integer accumulators are standard for INT8.

Stochastic rounding

Round-to-nearest maps every value in (xk−ulp/2, xk+ulp/2)(x_k - \mathrm{ulp}/2,\ x_k + \mathrm{ulp}/2) to xkx_k, so an update smaller than half an ulp of the weight is lost every time, however often it is applied. Stochastic rounding instead rounds up with probability (x−⌊x⌋)/ulp(x - \lfloor x \rfloor)/\mathrm{ulp}, so E[round(x)]=xE[\mathrm{round}(x)] = x and small updates survive on average. Wang et al. used it to make 16-bit weight updates work in FP8 training, and the NVFP4 recipe uses it for gradients.

1.01.21.40100200300updates →weightexact1.15SR1.164RNE1
Update size

After 150 updates of 0.001: exact 1.15, nearest-even 1 (every update below ulp/2 = 0.0039 is lost), stochastic 1.1641 (unbiased: E[round(x)] = x).

A BF16 weight at 1.0 (ulp 2^−7) receives N equal updates smaller than half an ulp. Round-to-nearest-even drops each one; stochastic rounding keeps them on average. One fixed random sequence.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 4 correct
  1. Q1BF16 and FP16 are both 16 bits. What is the key difference?

  2. Q2A tensor is quantized to INT8 with one scale for the whole tensor, s=max⁡∣x∣/127s = \max|x| / 127. Typical values are about 1, but one outlier is 25. What happens to the typical values?

  3. Q3How many bits per value does MXFP4 cost, counting the shared scale?

  4. Q4Why do 8-bit matrix units usually add up their products in 32 bits?

Sources

Show Hide 22 sources
  1. Mixed Precision TrainingPaulius Micikevicius, Sharan Narang et al. (Baidu and NVIDIA) · ICLR 2018, arXiv:1710.03740 · 2018FP16 storage with an FP32 master copy of weights, loss scaling and FP32 accumulation; nearly halves memory; values below 2^−24 become zero in FP16; about 5% of weight gradients fall below that; an SSD model diverges without loss scaling and trains with a scale of 8.
  2. FP8 Formats for Deep LearningPaulius Micikevicius et al. (NVIDIA, Arm and Intel) · arXiv:2209.05433 · 2022Defines E4M3 and E5M2; E4M3 drops infinities to reach 448 (240 otherwise); E4M3 for weights and activations, E5M2 for gradients; per-tensor scale factors with saturation; FP8 training matches 16-bit on models up to 175B parameters.
  3. OCP 8-bit Floating Point Specification (OFP8), Revision 1.0AMD, Arm, Google, Intel, Meta and NVIDIA · Open Compute Project · 2023Bit layouts, biases (7 and 15), special values and ranges of E4M3 (18 binades, max 448) and E5M2 (32 binades, max 57,344); round-to-nearest-even and saturating or non-saturating conversion.
  4. OCP Microscaling Formats (MX) Specification, Version 1.0Bita Darvish Rouhani et al. (AMD, Arm, Intel, Meta, Microsoft, NVIDIA, Qualcomm) · Open Compute Project · 2023A block of k = 32 elements sharing one E8M0 power-of-two scale; MXFP8, MXFP6, MXFP4 and MXINT8; FP4 (E2M1) and FP6 encodings; the conversion rule for the shared scale; dot-product semantics with implementation-defined internal precision.
  5. BFloat16: The secret to high performance on Cloud TPUsShibo Wang and Pankaj Kanwar · Google Cloud Blog · 2019Multiplier size scales with the square of the mantissa width; a BF16 multiplier is about half the size of an FP16 one and eight times smaller than FP32; the name comes from Google Brain; TPU v2/v3 multiply in BF16 and accumulate in FP32 in a 128 × 128 systolic array.
  6. Efficient Processing of Deep Neural Networks: A Tutorial and SurveyVivienne Sze, Yu-Hsin Chen, Tien-Ju Yang and Joel Emer · Proceedings of the IEEE, arXiv:1703.09039 · 2017Citing Horowitz (ISSCC 2014): an 8-bit fixed-point add uses 30× less energy than a 32-bit float add; an 8-bit fixed-point multiply 18.5× less energy and 27.5× less area than a 32-bit float multiply; multiplier cost scales about quadratically with bits, adder and memory cost about linearly; DRAM access costs orders of magnitude more than arithmetic.
  7. A Study of BFLOAT16 for Deep Learning TrainingDhiraj Kalamkar et al. (Intel and Facebook) · arXiv:1905.12322 · 2019Table of FP32, FP16 and BF16 bit layouts and ranges; BF16 keeps FP32’s range so training needs no hyperparameter changes; FMA units built from 8-bit multipliers; BF16 began as a storage format in DistBelief and TensorFlow; round-to-nearest-even conversion; FP32 accumulation.
  8. Microscaling Data Formats for Deep LearningBita Darvish Rouhani et al. · arXiv:2310.10537 · 2023Per-tensor scaling is insufficient below 8 bits; MX conversion algorithm; 8-bit MX direct-cast inference with minimal loss; 6-bit MX training matches FP32; 4-bit MX weights train with a minor accuracy drop.
  9. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only InferenceBenoit Jacob et al. (Google) · CVPR 2018, arXiv:1712.05877 · 2017The affine scheme r = S(q − Z) with a real scale S and an integer zero-point Z; 8-bit integer operands with a 32-bit integer accumulator.
  10. What Every Computer Scientist Should Know About Floating-Point ArithmeticDavid Goldberg · ACM Computing Surveys 23(1) (free in the ACM Digital Library’s open 1951–2000 archive) · 1991Significand, base and precision p; relative rounding error bounded by machine epsilon (β/2)β^−p; IEEE single precision with 8 exponent bits and a hidden bit giving p = 24; denormals and gradual underflow.
  11. LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleTim Dettmers, Mike Lewis, Younes Belkada and Luke Zettlemoyer · NeurIPS 2022, arXiv:2208.07339 · 2022Outlier features up to 20× larger than other dimensions emerge in all transformer layers around 6.7B parameters; they are about 0.1% of features but zeroing them wrecks perplexity; outlier dimensions kept in 16-bit, the other 99.9% in 8-bit.
  12. Integer Quantization for Deep Learning Inference: Principles and Empirical EvaluationHao Wu, Patrick Judd, Xiaojie Zhang, Mikhail Isaev and Paulius Micikevicius · arXiv:2004.09602 · 2020Relative math throughput and bandwidth of FP16, INT8 and INT4 versus FP32 on a 2018 GPU architecture; quantization granularity; calibration by max, entropy or percentile; an 8-bit workflow within 1% of floating point on all networks studied.
  13. A White Paper on Neural Network QuantizationMarkus Nagel et al. (Qualcomm AI Research) · arXiv:2106.08295 · 2021The quantization function with scale and zero-point; widening the range cuts clipping error but raises rounding error; per-channel and per-group quantizers.
  14. FP8 versus INT8 for efficient deep learning inferenceMart van Baalen et al. (Qualcomm AI Research) · arXiv:2303.17951 · 2023Gate-count estimate: an FP8-E4 multiply-accumulate costs about 53% more than INT8 with fixed-point accumulation, 183% more with an FP32 accumulator; integer (Kulisch) accumulation is exact; floats suit outlier-heavy distributions, integers uniform ones.
  15. gpt-oss-120b & gpt-oss-20b Model CardOpenAI · arXiv:2508.10925 · 2025Mixture-of-experts weights (over 90% of parameters) post-trained to MXFP4 at 4.25 bits per parameter, letting the 120B model fit on one 80 GB GPU.
  16. The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers and Luke Zettlemoyer · ICML 2023, arXiv:2212.09720 · 2022More than 35,000 experiments from 19M to 176B parameters: for a fixed total number of model bits, 4-bit weights are almost universally optimal; small block sizes and the choice of data type are what help.
  17. Pretraining Large Language Models with NVFP4NVIDIA · arXiv:2509.25149 · 2025NVFP4: 16-element blocks, E4M3 block scales and an FP32 tensor scale; MXFP4’s power-of-two scale can waste up to one binade; table of microscaling formats and speedups on one GPU generation; a 12B model trained on 10 trillion tokens in 4-bit matching FP8.
  18. Training Deep Neural Networks with 8-bit Floating Point NumbersNaigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen and Kailash Gopalakrishnan (IBM) · NeurIPS 2018, arXiv:1812.08011 · 2018Swamping: a small addend is lost when the running sum is larger by more than 2^(mantissa+1); chunk-based accumulation and stochastic rounding allow 16-bit additions.
  19. NVIDIA Hopper Architecture In-DepthMichael Andersch et al. · NVIDIA Technical Blog · 2022FP8 tensor cores with E4M3 and E5M2 inputs and FP16 or FP32 accumulation; FP8 doubles throughput over FP16/BF16; software chooses between FP8 and 16-bit per layer and manages scaling.
  20. DeepSeek-V3 Technical ReportDeepSeek-AI · arXiv:2412.19437 · 2024FP8 training of a 671B-parameter model; 1×128 and 128×128 scaling tiles; E4M3 on all tensors; FP8 GEMM accumulation on the GPU used retains about 14 bits, giving up to nearly 2% error at K = 4096, fixed by promoting partial sums to FP32 every 128 elements; loss within 0.25% of BF16.
  21. Introducing AMD CDNA 4 Architecture (white paper)AMD · AMD Instinct technical documentation · 2025CDNA 3 added both OCP FP8 variants; CDNA 4 adds hardware support for MXFP8, MXFP6 and MXFP4 and doubles matrix throughput for 16-bit and smaller types.
  22. With Shared Microexponents, A Little Shifting Goes a Long WayBita Rouhani et al. · ISCA 2023, arXiv:2302.08007 · 2023Block data representations; quantization signal-to-noise ratio grows linearly with mantissa bits (about 6.02 dB per bit) and logarithmically with block granularity; power-of-two block scales are cheap in hardware.