Architectures · Chapter 12 of 16 · AI accelerators

Sparsity, in-memory and analog compute

Some designs skip multiplying by zero, do math inside the memory itself, or use smooth signals instead of on-off switches. Each idea saves energy and brings new problems.

Sparsity hardware skips multiplications by zero. In-memory computing does multiply-accumulate inside memory arrays to avoid moving data. Analog compute adds and multiplies with currents or charges, trading precision for energy.

Structured and unstructured sparsity and the hardware that exploits each, SRAM and resistive compute-in-memory, converter overheads, noise and accuracy limits in analog compute, and what published prototypes report.

Most AI chips work the same way. They fetch numbers, multiply them, and add them up. This chapter looks at three ideas that try to do it differently.

  • Skip the zeros. Many numbers inside an AI model can be set to zero with little harm. Anything times zero is zero, so a chip can skip that work. This is called .
  • Do the math inside the memory. Carrying numbers from memory to the math parts takes far more energy than the math itself. does the math where the numbers are kept, so they never travel.
  • Use analog electricity. Normal chips use signals that are just on or off. lets flows of electricity add themselves up in a wire instead. It is very cheap, but a little fuzzy.

All three save energy on paper. This chapter is really about the catch: what each one gives up in return.

Neural networks spend most of their time multiplying matrices. The AI chips in this guide make that faster by adding more multipliers (GPUs), reusing data inside an array (systolic arrays), or keeping data on-chip (wafer-scale and SRAM-heavy designs). This chapter covers three ideas that attack the cost from other directions:

  • : if a weight is zero, its multiply-accumulate () can be skipped. Trained networks can often have half or more of their weights removed by with little accuracy loss.
  • (CIM): fetching a value from off-chip DRAM can cost around 200 times the energy of the arithmetic done on it, so CIM does the MAC inside the memory array that holds the weights.
  • : most CIM designs multiply and add with currents or charges. That is very efficient at low precision but introduces noise, and it needs converters to get back to digital.

The common thread is a trade: each idea removes one cost (wasted multiplies, data movement, digital logic) and adds another (indexing overhead, rigidity, error). Whether the trade pays depends on details that this chapter walks through.

Digital accelerators have converged on dense matrix engines fed by a memory hierarchy, and their energy is dominated by data movement: in a classic breakdown, a DRAM access costs about 200× a MAC, a global-buffer access 6×, and a register-file access about 1×. Even with good reuse, Murmann estimates that digital processing elements struggle to go below about 150–200 fJ per 8-bit MAC once local buffer traffic is counted, against about 75 fJ for a 16-bit MAC’s arithmetic alone in 28 nm. That gap is what the three ideas here target:

  • reduces the number of MACs and bytes. The hard part is turning a zero count into wall-clock and energy savings on hardware that wants regular work.
  • makes the weight-stationary dataflow total: weights never leave the array, so weight-fetch energy goes away and only activations and results move.
  • replaces adder trees with Kirchhoff summation or charge sharing. It wins only in the low-SNR regime, roughly below 8 bits, because thermal noise makes each extra bit cost about 4× the energy.

The chapter covers the patterns of sparsity and what hardware exploits, SRAM and resistive CIM, converter overheads, noise and accuracy, and what open prototypes actually measure. The recurring lesson is that the headline mechanism (skip a zero, sum a current) is cheap; the surrounding machinery (indices, load balance, ADCs, calibration) sets the real cost.

memorymultiplysumno weight fetchfetch×x1×x2×x3×x40.8#10.1#3−0.5#30.05#3skippedskippedcurrents add on the wire+y+ Baseline: nothing skipped, nothing kept in place.− Each weight fetched from DRAM costs ≈ 200× a MAC.
Scheme

Ordinary digital: weights are fetched from memory, multiplied by the inputs, and summed by an adder tree.

One four-term dot product, the ordinary way and with each of the chapter’s three ideas. Faded parts are the cost removed; amber parts are the cost added. Weight values illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

After an AI model is trained, many of its numbers are tiny and hardly matter. You can set them to zero, let the model practice a bit more, and it often works just as well. In some models that name what is in a photo, researchers removed nine out of ten numbers this way. The models did just as well.

Here is the twist: zeros scattered at random are hard for a chip to use. Its math parts work in lockstep, like rowers in a boat. If one rower still has lots to do, everyone waits.

So chip designers ask for zeros in a pattern. A popular rule is “in every group of four numbers, keep only two.” Every rower then gets the same amount of work, and the multiplying runs twice as fast.

There are two sources of zeros. Weight sparsity is created on purpose by : the classic recipe trains a network, removes the smallest weights, and retrains the rest. It cut AlexNet’s weights by 9× and VGG-16’s by 13× without accuracy loss. Activation sparsity happens by itself: the ReLU function sets every negative value to zero, so many inputs to the next layer are zero.

Where the zeros are matters

Zeros can be unstructured (anywhere) or follow a pattern, called . Unstructured pruning keeps the most accuracy for a given number of zeros, but it is the hardest to speed up:

  • Bookkeeping. Each kept value needs an index saying where it belongs, stored in a . With 8-bit weights and 16-bit indices, the index can be twice the size of the data.
  • Irregular access. Values are fetched from scattered addresses, which wastes memory bandwidth.
  • Load imbalance. Some rows keep more weights than others, and parallel hardware waits for the busiest one.

Two structured patterns dominate in practice. In 2:4 sparsity, every group of four consecutive weights keeps at most two. Half the weights are always zero, every group looks the same, and each kept value needs only a 2-bit position. GPUs with sparse matrix units run these at twice the dense math rate. In block sparsity, whole tiles (say 8×8 or 32×32) are zero, so an ordinary matrix engine can skip a tile at a time with almost no overhead. Larger blocks are easier to exploit but cost more accuracy.

The rule of thumb that falls out: hardware speeds up sparsity it can count on in advance. Unstructured sparsity needs very high zero fractions, or special hardware, before it pays.

Zeros come from two places: static weight sparsity from , and dynamic activation sparsity (ReLU zeros, which an accelerator can exploit by gating the MAC for about 45% energy savings, or by skipping the cycle for a 1.37× throughput gain in the designs Sze et al. survey). Each can be unstructured or , and the pattern, more than the zero count, decides whether hardware benefits.

Unstructured: most accurate, least exploitable

Magnitude pruning plus retraining reaches 9–13× parameter reduction on older CNNs, and SparseGPT prunes 175-billion-parameter language models to 60% unstructured sparsity in one shot with negligible perplexity increase. But Mishra et al. note that the performance benefit of such patterns on matrix pipelines is “negligible and at times negative,” even at 95% sparsity. The obstacles:

  1. Metadata. CSR/CSC store an index per nonzero; with 8-bit values and 16-bit column indices the overhead is up to 200%.
  2. Data-dependent access. Indirection adds latency to matrix reads and wastes cache lines.
  3. Imbalance. Per-row nonzero counts vary, so SIMD lanes or array rows idle while the busiest finishes.

In software on GPUs, the best sparse kernels reached about 27% of a V100’s single-precision peak and gave 1.2–2.1× end-to-end speedups on sparse Transformers and MobileNets, at the moderate sparsity real models tolerate. Purpose-built hardware does better. EIE stored pruned fully connected layers in compressed sparse column form in on-chip SRAM, skipped zero activations, and attributed 10× of its energy advantage to sparsity. Dataflow designs with distributed SRAM can trigger a MAC only for each nonzero weight that arrives, so zeros are never streamed or stored.

N:M structured sparsity

2:4 keeps at most two nonzeros per group of four along the reduction dimension. The compressed matrix is half the width, plus 2 bits per kept value saying which of the four positions it came from: 12.5% overhead for 16-bit values and 25% for 8-bit. The sparse matrix unit uses that metadata to select the matching half of the dense operand and runs a GEMM of half the depth, for twice the peak math rate. Because every group has the same count, there is no imbalance, and a value’s position follows from the compression rate with no pointer chasing. The recipe is to train dense, prune each group to its two largest magnitudes, and retrain with the original schedule, which matched dense accuracy across the vision, translation and language tasks reported.

Block and coarse structures

Block sparsity zeroes whole b×bb \times b tiles. A dense engine skips tiles with only a per-block index, and OpenAI’s kernels showed speedups near 1/(1−s)1/(1-s) over cuBLAS at high sparsity with 32×32 blocks, while generic cuSPARSE was slower than simply multiplying the zeros. Coarser still, pruning whole channels or filters just yields a smaller dense layer, but larger groups lose more accuracy.

weights (empty = zero)time: multiplies, then idlelane 16lane 23lane 35lane 42step endsSpeedup 1.33× (1.07× with index decoding)storagedense 256 bvalues 128 b+ index 256 b
Sparsity pattern

Half the weights are zero, but the busiest lane still has 6 of 8: at best 1.33×. The 16-bit indexes make storage 150% of dense.

Four lanes in lockstep, one row of eight weights each. Zeros are skipped, but the step ends only when the busiest lane is done. Storage counts 8-bit values plus their indexes. Zero positions illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

In most chips, memory and math sit apart. Every number is read out of memory and sent down a wire to be multiplied. Then the answer is sent back. For AI, that back-and-forth uses much more energy than the math.

builds the math into the memory. The memory is a grid of tiny cells, and each cell holds one of the model’s numbers. Each cell does its own small multiply right where it sits, and the results add up down each column.

Some designs use , the fast memory found on every chip, with a few extra parts in each cell. Others use newer kinds of memory, such as . It stores a number as how easily electricity flows through a tiny cell.

A weight-stationary systolic array keeps each weight in one place while inputs flow past it (see Systolic arrays). takes that to the limit: the memory cells that store the weights also do the multiplying, so weights are never read out at all.

Resistive crossbars

Picture a : horizontal and vertical wires with a programmable resistor at each crossing. Store each weight as the cell’s conductance GG (how easily current flows). Put each input on its row as a voltage VV. By Ohm’s law each cell passes a current I=V×GI = V \times G, a multiplication. By Kirchhoff’s current law, the currents entering a column wire add up. So each column current is a full dot product, and all the columns together compute a matrix-vector product in one step.

The resistors can be , (PCM) or cells. All are non-volatile and dense, so a chip can hold millions of weights without any external memory. They have drawbacks: limited precision, wire resistance that distorts large arrays, costly writes, and cell-to-cell variation.

SRAM-based CIM

, the fast memory on every logic chip, can also compute. In analog SRAM CIM, each cell (with a few extra transistors or a small capacitor) multiplies its stored weight bit by an input and contributes a current or charge to a shared wire; capacitors are popular because their values can be matched closely. A 65 nm research processor with a 590-kilobit charge-domain SRAM array ran CIFAR-10 image classification at the same accuracy as a digital implementation. SRAM cells hold only one bit and lose data when power is off, so multi-bit weights are spread over several cells and the arrays are less dense than resistive ones.

Digital CIM keeps the SRAM array but replaces the analog summing with tiny logic gates and adder trees beside the cells. It gives exact answers and more flexible mapping, at some cost in energy efficiency.

is the extreme of weight stationarity: weights are written once and every MVM reads them in place. Designs split along two axes: the storage device and whether accumulation is analog or digital.

Resistive crossbars (RRAM, PCM, flash)

In a , row inputs (voltage levels, pulse widths or serial bits) meet a conductance matrix and each column current is Ij=∑iViGijI_j = \sum_i V_i G_{ij}. Signed weights use a pair of devices, w∝G+−G−w \propto G^+ - G^-; IBM’s PCM chip reports that using two devices per sign improves SNR enough to move from roughly 3-bit-equivalent to 3–4-bit-equivalent weight precision relative to an 8-bit digital engine. The practical limits of large arrays are IR drop along wires (wire energy dominates at around 1k×1k), write energy and multiple programming pulses, and device-to-device and cycle-to-cycle variation with nonlinear conductance. arrays usually put a series selector transistor in each cell (one-transistor-one-resistor, 1T1R); NeuRRAM integrated 3 million such devices with 130 nm CMOS. drifts after programming (next section). In flash, each transistor acts as a programmable variable resistor holding a weight; Mythic presented a 40 nm design of this kind with 8-bit analog compute at Hot Chips 2018.

SRAM CIM: analog and digital

Early analog SRAM CIM drove word lines with a 5-bit DAC and let bit-cell currents discharge the bit line, saving 12× in energy compared with reading binary weights out and computing separately. The bit-line discharge is nonlinear and cells vary, so that design needed a DAC that pre-distorts the input and boosting to combine weak classifiers. Charge-domain designs sum on capacitors instead: each bit cell drives a local capacitor, and charge sharing across the column computes the sum. With about 1% capacitor matching, the cumulative error stays below one input step until NN exceeds about 10,000. Jia et al. integrated a 590 kb charge-domain array with a CPU and used bit-parallel weights and bit-serial inputs to scale precision, reaching software-level CIFAR-10 accuracy at 1 and 4 bits.

Digital CIM (DIMC) puts bitwise multipliers and adder trees next to the bit cells. It is noise-free and maps more flexibly, and in the Houshmand et al. survey its efficiency depends strongly on technology node and precision. For analog CIM (AIMC), by contrast, a newer node raises compute density but only marginally improves energy efficiency.

The system problem

A CIM array computes only matrix-vector products against weights already written into it. Activation functions, products between two activation tensors, and data routing between arrays still need digital logic and buffers; IBM’s speech demonstration ran its vector operations and activations on a host. Retaining the array’s efficiency for a complete accelerator is the open challenge. SRAM CIM capacity is small, so large models need off-chip weight buffers, which brings back the data movement CIM was meant to remove; dense non-volatile devices are the proposed fix.

conductance G (µS) = weightV1 0.20 VV2 0.10 VV3 0.00 V6.0 µA11.0 µA8.0 µAcolumn current= dot product105030401020203050

Column currents 6.0, 11.0, 8.0 µA: each is a full dot product, summed by Kirchhoff’s current law. Tap a cell to see its multiply.

A 3 × 3 resistive crossbar. Conductance G stores the weight, row voltage V is the input; each cell multiplies by Ohm’s law and each column sums by Kirchhoff’s law. Ideal devices, illustrative values.Share freely with credit: ‘Figure from chipfieldguide.com’

Inside an analog grid, numbers are electricity. A bigger input is a stronger push. A bigger stored number lets more electricity through. Where wires meet, the flows just add together. That adding is free; nature does it.

The rest of the chip still works with normal numbers, though. So at the edge of each grid sit translators. One turns each number into electricity on the way in. Another, called an , measures what comes out and turns it back into a number.

The measuring is the costly part. Telling 256 levels apart takes real work, and each step finer costs a lot more. In one design study, measuring used more than half of all the power.

needs converters at both ends of the array:

  • A on each row turns a digital input into a voltage or a pulse of a certain length. Many designs avoid multi-bit DACs by feeding inputs one bit at a time over several cycles, which only needs a simple on/off driver.
  • An on each column (or shared by several columns) measures the summed current and outputs a digital number.

The ADC sets the cost. A column summing many products produces a wide range of possible values, and resolving all of them takes many bits. Each additional bit doubles the number of levels, and once the converter is limited by electrical noise, each extra bit roughly quadruples the energy. In ISAAC, a well-known design study for resistive crossbars, the ADCs took 58% of the power and 31% of the area of each tile.

Two tricks help. Sharing one ADC conversion across many rows spreads its cost over many MACs, so arrays should sum at least a few hundred rows per conversion. And not every bit of the sum is needed: the network only needs enough precision that the ADC’s error is small compared with errors the model already tolerates.

The result is a clear sweet spot. Below about 8 bits of precision, analog MACs can beat digital ones; above that, the converters make digital the better choice. That is one reason CIM pairs naturally with the low-precision formats in Number formats.

Precision has to be assembled from low-precision pieces, because a 16-bit multiply in one cell would need a 16-bit , 2162^{16} conductance levels and an of more than 16 bits. ISAAC’s scheme, typical of the field:

  • Inputs bit-serial: v=1v = 1 bit per cycle, so the DAC is an inverter and a 16-bit input takes 16 cycles, combined by shift-and-add. A 2-bit DAC would have added 63% chip area and 7% power for no throughput gain.
  • Weights bit-sliced: w=2w = 2 bits per cell, so a 16-bit weight spans 8 cells whose column results are shifted and added digitally.
  • ADC resolution follows from the column sum: A=log⁡2R+v+wA = \log_2 R + v + w, minus 1 if vv or ww is 1. For R=128R = 128 rows, v=1v = 1, w=2w = 2 this gives 9 bits; storing a column’s weights inverted when that keeps the MSB zero saves one bit, allowing an 8-bit ADC.

Even so, ISAAC’s 96 ADCs per tile took 58% of tile power and 31% of area, and the authors found a 9-bit ADC never worth its cost. Murmann’s amortized model makes the scaling explicit: EMAC=EADC/N+Ecell+ElogicE_{\mathrm{MAC}} = E_{\mathrm{ADC}}/N + E_{\mathrm{cell}} + E_{\mathrm{logic}}, with EADCE_{\mathrm{ADC}} fitted to published converters as 100 fJ⋅ENOB+1 aJ⋅4ENOB100\,\mathrm{fJ} \cdot \mathrm{ENOB} + 1\,\mathrm{aJ} \cdot 4^{\mathrm{ENOB}}. For 4-bit operands over N=1152N = 1152 rows the predicted MAC energy is 3.8 fJ; beyond 8 bits it exceeds 200 fJ, which digital designs already reach, and above 6–7 bits competing with digital is “challenging.” Published macros agree: the AIMC macro with the best peak efficiency, about 1,800 TOP/s/W, got there by optimizing its converters and using a large array so that the array, not the ADCs, dominated power.

Alternatives to the classic column ADC trade differently: NeuRRAM integrates voltage-mode neurons that double as the ADC and activation function, avoiding large current-sinking amplifiers; IBM’s 64-core PCM chip gives each of a core’s 256 outputs its own time-based current ADC, trimmed once at calibration, for fully parallel MVMs; and IBM’s speech chip encodes inputs as pulse durations and uses ramp-based conversion.

digitalindigitaloutDACarrayADCADC levels: 2^3 = 8truereadserror 0.047ADC energy (log)300 fJISAAC tile powerADCs 58%

3-bit ADC: 8 levels. True 0.618 reads as 0.571 (error 0.047). Energy per conversion 300 fJ.

The converters around an analog array. The ADC rounds the column sum to one of 2^b levels; energy per conversion from Murmann’s fit to published ADCs. Column values illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Normal chips give the exact same answer every time. Analog chips don’t quite. No two memory cells are perfectly alike, some slowly change over time, and electricity itself jiggles a little. The answers come out close, but not exact.

AI models can live with a bit of fuzz, the way you can still read a slightly blurry sign. The trick is to train the model to expect it. While it learns, researchers add random errors on purpose. So it learns to give good answers anyway. This is called .

It makes a big difference. In one test, a model trained the normal way sorted pictures correctly only about 25% of the time once chip-like noise was added. Trained with practice noise, it got about 86% right.

Analog errors come from several sources, which the NeuRRAM team listed after measuring their chip:

  • Limited precision when programming each cell, and differences from cell to cell.
  • Cells that slowly change after being written. In PCM this is called .
  • Voltage lost along long wires (IR drop), and electrical coupling between neighboring wires.
  • The ADC’s limited resolution and range.

Simulations that model only some of these can overestimate accuracy; NeuRRAM found a 2.3-point gap on CIFAR-10 between a simulation that modeled only two error sources and the real chip.

Designers fight back at every level:

  • Program carefully. Write each cell, read it back, and adjust until it is close enough (write-verify).
  • Calibrate. Measure each array’s output range on real data and set the ADC to match.
  • Train for it. raised CIFAR-10 accuracy under NeuRRAM’s device noise from 25.34% to 85.99%.
  • Use two cells per weight. Averaging two devices reduces random error.

The best published results come within about a percentage point of software on a small image task, and further from it on a large speech model (see By the numbers). IBM’s 64-core PCM chip measured a precision equivalent to roughly 3- to 4-bit weights.

For a column summing NN products with independent per-device error of standard deviation σ\sigma (as a fraction of the largest weight, the maximum conductance), the error variance is σ2∑xi2\sigma^2 \sum x_i^2 while the signal variance is E[w2]∑xi2E[w^2] \sum x_i^2, so the relative error ≈σ/rms⁡(w)\approx \sigma / \operatorname{rms}(w) independent of NN. Programming error therefore caps SNR no matter how many ADC bits you buy; the simulation below shows the plateau. The sources NeuRRAM lists are IR drop on input wires, drivers and crossbar wires; limited programming resolution; conductance relaxation; capacitive coupling; and ADC resolution and range. Modeling only relaxation and ADC effects overestimated CIFAR-10 accuracy by 2.32 points versus measurement.

Drift

follows G(t)=G(t0) (t/t0)−νG(t) = G(t_0)\,(t/t_0)^{-\nu}. A global per-core rescale of the MVM output corrects the mean, but ν\nu depends on conductance state and device, so the residual error grows with time. RRAM instead relaxes abruptly within about a second of programming. NeuRRAM measured a relaxation spread near 10% of the maximum conductance, far larger than its write-verify tolerance; three rounds of re-programming cut it by 29%, and hardware-aware training largely mitigates the rest.

Mitigations, by layer of the stack

  • Device: iterative write-verify; two devices per signed weight to average error.
  • Circuit: per-ADC trimming and on-chip calibration so converters share one linear transfer curve; per-layer selection of input amplitude to fill the ADC range.
  • Training: using measured device statistics (25.34% → 85.99% CIFAR-10 in NeuRRAM’s evaluation), then chip-in-the-loop progressive fine-tuning (a further 1.99 points), where measured outputs of programmed layers train the layers not yet programmed.
  • System: keep precision-sensitive operations digital. IBM’s speech demonstration ran vector-vector products, biases, activations and the joint layer in software.

Murmann notes that inference tolerates errors (and can be desensitized by noise injection) down to about 4 bits, but training does not; at least 8-bit floats tend to be required. That makes analog CIM an inference technology. For reference, IBM’s 64-core chip measured precision equivalent to 3–4-bit weights in an 8-bit-I/O engine.

trained boundaryas programmedmisclassifiedaccuracy per chip190%290%3100%470%595%Mean 89%
Training

Normally trained: thin margin. Five chips, each programmed with σ = 15% error: 89% correct on average.

A toy two-input classifier stored in noisy analog cells: each programmed chip draws a slightly different boundary. Illustrative; NeuRRAM’s measured CIFAR-10 accuracy under device noise rose from 25.34% to 85.99% with noise-aware training.Share freely with credit: ‘Figure from chipfieldguide.com’

The simulation has two tabs. On Sparsity, pick where the zeros go and how many numbers to remove. An ordinary chip gets no faster, but a chip built for the pattern does. On Analog MAC, make the memory cells sloppier or the measuring finer. Watch how often the chip gets the wrong answer. Then see how much energy goes into measuring.

The Sparsity tab prunes a 16×16 weight matrix by magnitude in one of three patterns and compares the multiply speedup on hardware without sparsity support and on hardware built for that pattern. The Analog MAC tab runs 160 fixed trials of a 64-row crossbar with 4 columns and compares the analog result, after the ADC, with exact math. Things to try:

  • Set unstructured sparsity to 50%. Is the specialized hardware actually faster? Now try 2:4.
  • Compare “weight energy kept” for unstructured, 2:4 and block at 50%. Which loses the most?
  • In Analog MAC, set the noise to 5% and sweep the ADC from 2 to 12 bits. Where does the error stop improving, and what happens to energy after that?

Expert view adds the busiest-row count (load imbalance), a software sparse-kernel bar for unstructured sparsity, a rows-per-column slider NN, output SNR (shown as equivalent bits), Murmann’s B+log⁡2NB + \log_2 \sqrt{N} rule, and an ADC energy-per-conversion curve. All constants are illustrative; the ADC energy fit is Murmann’s.

  • Find the unstructured sparsity where the index-driven engine first beats 1×, and the one where the software kernel breaks even. Compare with the 1/3 storage break-even.
  • With σ=0\sigma = 0, compare the ADC bits you set with the equivalent bits shown at N=64N = 64 and N=256N = 256. Why is there a gap (the outputs don’t fill the ADC’s full scale), and how does it compare with the rule of thumb, which is the ADC’s own required ENOB?
  • Raise NN from 16 to 256 at 10 bits and watch the ADC share of energy per MAC.
Loading simulation…
Peak math of 2:4 sparse vs. dense, on supporting hardware
2×
Metadata per kept 2:4 value
2 bits
ADC share of tile power, ISAAC (simulated)
58%
Predicted energy per 4-bit analog MAC (model)
3.8 fJ

Sources: Mishra et al. and the A100 whitepaper for sparsity; ISAAC for the ADC share; Murmann’s model for the MAC energy, which assumes a large, well-amortized array.

What these numbers mean:

  • Twice the math is what a chip built for the “two of every four” pattern gets from the same parts. The model has to follow the pattern.
  • 2 bits is the tiny note the chip keeps for each kept number. It says which of the four spots it came from.
  • 58% means that in one well-known analog design, more than half the power went to measuring the answers, not to working them out.
  • 3.8 femtojoules is a best-case guess for one small analog multiply-and-add. A joule is about the energy it takes to lift an apple one meter. A femtojoule is a millionth of a billionth of that. A digital chip uses around 200 femtojoules for a bigger, more exact (8-bit) multiply-and-add, so on paper the small analog one looks tens of times cheaper. At the same precision the gap shrinks, and real analog chips do much worse than the best case.

Reading the tables:

  • The sparsity numbers are peak math rates. A 2× peak only becomes 2× real speed when the matrix multiply is large and math-bound; the MI300A study shows how little can be left in small or isolated runs.
  • The CIM chips are research prototypes. The IBM speech chip ran parts of the network on a host computer, and its word error rate rose from 7.45% to 9.26%.
  • figures are not comparable across papers without knowing the precision, the assumed sparsity, and whether the number covers the array, the chip or the whole system.

Calibrating the figures:

  • 2:4 cuts weight storage to 62.5% of dense at 8 bits and 56% at 16 bits (including metadata); measured INT8 GEMM speedups on A100 approach 2× as GEMM-K grows at M=N=10,240M = N = 10{,}240.
  • The Jia et al. figures are 1b-, normalized to 1-bit operations; throughput scales down linearly with the weight and input precisions. Houshmand et al. avoid bit-normalized metrics because they distort the landscape, and compare designs only at a stated 50% operand sparsity.
  • The HERMES chip’s 9.76 TOPS/W is for 1-phase MVMs at full utilization; the accuracy experiments used a higher-precision 4-phase read mode.
  • The speech system’s figures are estimates for the five analog chips plus modeled digital work: 6.94 TOPS/W with conventional weight mapping, where the analog-to-digital operation ratio was 325:1 (vector ops and activations in software), and 6.7 TOPS/W with the weight expansion that gave the 9.258% word error rate. Its 14× is a projection, relative to the best MLPerf energy-efficiency submission for that task.

Skipping zeros

Removing numbers from anywhere keeps the model smartest, but chips can barely use it. A strict pattern makes the chip faster, but costs the model a little.

Computing in memory

The memory grid is great at one job: multiplying inputs by a fixed table of numbers. Everything else still needs ordinary circuits. And a big model may not fit, so it has to be loaded in pieces, which brings back the trips.

Analog

Analog math is cheap only when you don’t need much detail. Finer measuring gets expensive fast, and the fuzz never fully goes away. Analog chips are also built for using a trained model, not for training one.

ApproachYou gainYou give up
Unstructured sparsityMost accuracy for a given number of zeros; big memory savings at high sparsityIndex overhead, irregular memory access, idle lanes; little speedup on ordinary matrix hardware
2:4 structured sparsityA reliable 2× in math on supporting hardware; small metadataFixed at 50%; needs retraining or careful pruning; no gain where memory, not math, is the limit
Block sparsityWorks on ordinary matrix engines; speedup tracks sparsityMore accuracy loss, often offset by a bigger model
SRAM CIMUses standard chip processes; fast; digital versions are exactLow density, volatile; large models need weights reloaded from outside
Resistive CIM (RRAM, PCM, flash)Dense, non-volatile weight storage; a whole model can stay on chipNoise, drift, slow and energy-hungry writes; extra process steps
Analog MACs in generalVery low energy per MAC at 4–8 bitsConverters dominate; accuracy loss; inference only; harder to program

Ways designs go wrong

  • Sparsity that stays on paper. A model pruned to 60% unstructured sparsity runs no faster on a dense matrix engine, and storing 16-bit indices for 8-bit weights makes it bigger, not smaller.
  • Peak, not delivered. 2× sparse math does nothing for a layer that is waiting on memory, which is common during token-by-token generation (see The memory wall).
  • Simulated accuracy. Analog results estimated in software from a few device measurements tend to be too optimistic.
  • Forgetting the rest of the chip. An array that is 100× more efficient helps little if data movement and digital operations around it dominate.

Sparsity: accuracy, pattern, speed

At equal sparsity, accuracy loss rises from unstructured to N:M to block to channel, while exploitability rises in the same order. 2:4 sits at a sweet spot because its constraint is local enough for retraining to absorb, which the A100 table shows: dense and 2:4 accuracy match within noise across ResNet, Mask R-CNN, GNMT, Transformer-XL and BERT. Its limits: it is fixed at 50%, its gain is in math only, and it applies to one operand (the second matrix and the output stay dense). Block sparsity typically recovers accuracy by enlarging the model, which is useful only when the larger dense model would have been impractical. Activation sparsity is known only at run time, so hardware must detect it on the fly, as the zero-gating and zero-skipping designs do, and its benefit varies with the input.

CIM: density, flexibility, utilization

  • Array utilization. Large arrays amortize converters but are underused by depthwise and pointwise layers; small arrays map well but pay peripheral overhead, and many small macros raise feature-map traffic.
  • Capacity. Current SRAM CIM designs cannot hold more than about 10 MB of weights at reasonable area, so larger networks need off-chip weight buffers.
  • Write cost. Resistive devices need multiple programming pulses and have finite endurance, which rules out frequently rewritten weights.
  • Programmability. Few analog CIM designs can be reconfigured for diverse models; NeuRRAM’s transposable array is one attempt.

Analog: where the efficiency goes

Every analog design meets the same three-way tension. More rows per conversion amortize the ADC but grow the required ENOB by log⁡2N\log_2 \sqrt{N} and worsen IR drop. More bits per cell or per input pulse cut cycles but exponentially shrink allowable rows (ISAAC’s A=log⁡2R+v+wA = \log_2 R + v + w). Higher precision costs 4× per bit in the thermal-noise regime. Murmann’s caution about binarized arrays is instructive: mixed-signal demonstrators reach about 4 fJ/MAC, but digital designs with similar architecture and dataflow get close, so mixed-signal is not necessarily a game changer for binarized arrays, and the case for it rests on multi-bit operands.

Unstructured✗ no gain2:4✗ no gainBlock✓ speedupChannel✓ speedupeasier for hardware to exploitless accuracy lost
Hardware

Ordinary matrix engine: only block and channel sparsity speed it up (2 of 4). Tap a pattern for its trade.

Four sparsity patterns, each 50% zeros. Accuracy loss rises left to right, and so does how easily hardware turns the zeros into speed. Tap a pattern; change the hardware.Share freely with credit: ‘Figure from chipfieldguide.com’
16-bit input: 16 cycles16-bit weight: 8 cellsrows summed per conversion: 128IR dropADC resolution A9 bitsADC energy per 16-bit MAC (log)1.2 pJ

16 cycles per input, 8 cells per weight, 128 rows per conversion → the ADC needs 9 bits; ADC energy ≈ 1.2 pJ per 16-bit MAC.

Every analog design trades rows per conversion, bits per cell and bits per input cycle against ADC resolution, using ISAAC’s formula. ADC energy from Murmann’s fit; combining them this way is illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

Five models, each simple enough to check by hand; the simulation implements most of them.

1. Sparse speedup on parallel hardware

Let dd be the density (1 − sparsity). For a dense engine, multiply time is unchanged: T=TdenseT = T_{\mathrm{dense}}. A software sparse kernel that reaches a fraction η\eta of dense peak gives speedup S=η/dS = \eta/d, so it breaks even only at d=ηd = \eta. Gale et al.’s kernels reached 27% of single-precision peak on their best problems; taking η≈0.27\eta \approx 0.27 as a best case, this simple model needs more than about 73% sparsity. On hardware with LL lanes each handling one row, the step time is set by the busiest row:

S=Cmax⁡rnnzr⋅(1+o)S = \frac{C}{\max_r \mathrm{nnz}_r \cdot (1 + o)}

where CC is the row length and oo the per-nonzero index-decoding overhead. Imbalance (max vs. mean nnz) and oo both eat into 1/d1/d. The simulation uses C=16C = 16 and o=0.25o = 0.25 (illustrative). For 2:4, nnzr=C/2\mathrm{nnz}_r = C/2 for every row and oo is absorbed into the multiplexer, so S=2S = 2 exactly on the math pipeline. For blocks, S=nblocks/(nnonzero (1+ob))S = n_{\mathrm{blocks}} / \bigl(n_{\mathrm{nonzero}}\,(1 + o_{\mathrm{b}})\bigr), approaching 1/d1/d. In every case the end-to-end gain obeys Amdahl’s law over the non-GEMM and memory-bound parts.

2. Storage arithmetic

With bb-bit values and an ii-bit index per nonzero, storage relative to dense is d (b+i)/bd\,(b + i)/b.

  • Unstructured, b=8b = 8, i=16i = 16: 3d3d, so it saves memory only for d<1/3d < 1/3.
  • 2:4, b=16b = 16: (2⋅16+2⋅2)/(4⋅16)=36/64(2 \cdot 16 + 2 \cdot 2)/(4 \cdot 16) = 36/64, about 44% saved; b=8b = 8: 20/3220/32, about 38% saved.
  • Block bs×bsb_{\mathrm{s}} \times b_{\mathrm{s}} with one ii-bit index per kept block: d (1+i/(b bs2))d\,\bigl(1 + i/(b\, b_{\mathrm{s}}^2)\bigr).

3. Pruning recipes

  1. Train–prune–retrain (Han et al.): train dense, remove weights below a magnitude threshold, retrain the survivors; iterate for higher sparsity.
  2. 2:4 workflow (Mishra et al.): train dense; in each group of four, zero the two smallest magnitudes; retrain with the same optimizer and learning-rate schedule as the original run, trading one extra training run for no hyperparameter search.
  3. One-shot, second-order (SparseGPT): prune layer by layer, solving a reconstruction problem so the remaining weights compensate, with no retraining; scales to 175B parameters and supports 2:4 and 4:8.

4. Crossbar MVM, converters and SNR

Ideal crossbar: Ij=∑iVi(Gij+−Gij−)I_j = \sum_i V_i \left(G^+_{ij} - G^-_{ij}\right). With bit-serial vv-bit inputs and ww-bit cells across RR rows, the column needs A=log⁡2R+v+wA = \log_2 R + v + w (−1 when v or w=1)(-1 \text{ when } v \text{ or } w = 1) bits to be lossless. In practice designs choose fewer, following an SNR argument. Murmann assumes uniform activations quantized to BB bits; the quantization noise of NN accumulated products sets an SNRsig=22B\mathrm{SNR}_{\mathrm{sig}} = 2^{2B}, and the ADC’s input-referred noise must be k2k^2 times smaller. For a full-scale range FS\mathrm{FS}, that gives

ENOB=B+log⁡2 ⁣(k⋅FS⋅N)\mathrm{ENOB} = B + \log_2\!\left(k \cdot \mathrm{FS} \cdot \sqrt{N}\right)

With k=2k = 2 and FS=0.5\mathrm{FS} = 0.5 this is B+log⁡2NB + \log_2 \sqrt{N}, about 5 extra bits for N=1152N = 1152. Energy per MAC is then

EMAC=EADCN+Ecell+Elogic,EADC=k1⋅ENOB+k2⋅4ENOBE_{\mathrm{MAC}} = \frac{E_{\mathrm{ADC}}}{N} + E_{\mathrm{cell}} + E_{\mathrm{logic}}, \qquad E_{\mathrm{ADC}} = k_1 \cdot \mathrm{ENOB} + k_2 \cdot 4^{\mathrm{ENOB}}

with k1=100 fJk_1 = 100\,\mathrm{fJ} and k2=1 aJk_2 = 1\,\mathrm{aJ}. Murmann fits this to published ADC measurements; the linear term covers the part of converter energy that grows with each added bit, and the 4ENOB4^{\mathrm{ENOB}} term is the thermal noise limit: noise power is kT/CkT/C, so halving the step (one more bit) needs 4× the capacitance and 4× the energy. At ENOB=8\mathrm{ENOB} = 8 the two terms are 800 fJ and 66 fJ; at 12 bits, 1.2 pJ and 16.8 pJ. The simulation uses this fit with b=ENOBb = \mathrm{ENOB}, adds an illustrative DAC term (44 fJ/bit × 4 bits shared by 64 columns, a per-bit figure Houshmand et al. fitted to published macros) and 1 fJ per cell.

Device error adds independently: for per-device error σ\sigma, relative output error ≈σ/rms⁡(w)\approx \sigma / \operatorname{rms}(w), so the output SNR is bounded by roughly 20log⁡10 ⁣(rms⁡(w)/σ) dB20 \log_{10}\!\bigl(\operatorname{rms}(w)/\sigma\bigr)\,\mathrm{dB} however fine the ADC. Effective bits follow from SNR=6.02⋅ENOB+1.76 dB\mathrm{SNR} = 6.02 \cdot \mathrm{ENOB} + 1.76\,\mathrm{dB}. Combining the two error sources, total error2≈σdev2+σADC2\text{total error}^2 \approx \sigma_{\mathrm{dev}}^2 + \sigma_{\mathrm{ADC}}^2; the ADC should be sized so its term is a fraction of the device term, not driven to zero. That is the knee you can find in the simulation.

5. Drift compensation

For PCM, G(t)=G(t0) (t/t0)−νG(t) = G(t_0)\,(t/t_0)^{-\nu}. Global drift compensation rescales each core’s MVM results by an appropriate factor, removing the mean drift. The residual comes from the spread of ν\nu across devices and states, and it grows over time.

G / G_maxν=.045−25ν=.060−11ν=.050−19ν=.065−7ν=.040−19ν=.055−15ν=.062−9ν=.048−25programmednowerror, % of G_maxt = 24.0 h
Global drift compensation

t = 24.0 h: mean conductance down 30%; rms error 17.5% of G_max. Turn on global compensation.

PCM drift, G(t) = G(t₀)(t/t₀)^−ν, for eight devices with different ν. A global rescale fixes the mean but not the spread. t₀ = 60 s and ν = 0.040–0.065 are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 4 correct
  1. Q1A weight matrix is pruned to 50% sparsity with zeros placed anywhere. On a matrix engine with no sparsity support, how much faster does the multiply run?

  2. Q2Why does 2:4 structured sparsity need so little metadata?

  3. Q3In an analog compute-in-memory array, where does most of the energy typically go?

  4. Q4Why do analog AI chips usually run models with only about 4 to 8 bits of precision?

Sources

Show Hide 19 sources
  1. Efficient Processing of Deep Neural Networks: A Tutorial and SurveyVivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, Joel Emer · arXiv (Proceedings of the IEEE, 2017) · 2017Relative energy of DRAM, buffer, PE and register-file accesses vs. a MAC; ReLU activation sparsity and zero-skipping hardware; CSR/CSC formats; structured vs. fine-grained pruning; SRAM and resistive in-memory computing and their ADC/DAC, IR-drop, write-energy and variation drawbacks.
  2. Learning both Weights and Connections for Efficient Neural NetworksSong Han, Jeff Pool, John Tran, William J. Dally · arXiv (NeurIPS 2015) · 2015Train, prune small weights, retrain: AlexNet parameters cut 9× (61 M to 6.7 M) and VGG-16 13× (138 M to 10.3 M) with no accuracy loss.
  3. NVIDIA A100 Tensor Core GPU Architecture (whitepaper)NVIDIA · NVIDIA technical whitepaper · 2020Fine-grained 2:4 structured sparsity; sparse MMA instructions double Tensor Core throughput; compressed weights halve footprint; train dense, prune 2:4, fine-tune; Table 11 dense vs. 2:4 accuracy; fine- vs. coarse-grained sparsity primer.
  4. Accelerating Sparse Deep Neural NetworksAsit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, Paulius Micikevicius · arXiv · 20212:4 format with 2-bit metadata per kept value (12.5%/25% overhead for 16/8-bit), CSR overhead up to 200%; dense vs. sparse TOPS table; measured INT8 sparse GEMM speedup approaching 2× for large GEMMs; unstructured sparsity rarely speeds up matrix pipelines.
  5. Sparse GPU Kernels for Deep LearningTrevor Gale, Matei Zaharia, Cliff Young, Erich Elsen · arXiv (SC 2020) · 2020Practical speedups are hard at moderate sparsity; kernels reach 27% of V100 single-precision peak on the top-performing problems; 1.2–2.1× end-to-end speedups for sparse Transformer and MobileNet models.
  6. GPU Kernels for Block-Sparse WeightsScott Gray, Alec Radford, Diederik P. Kingma · OpenAI technical report · 2017Block-sparse kernels efficient with blocks as small as 8×8; speedup over cuBLAS close to 1/(1 − s) at high sparsity; cuSPARSE and per-block cuBLAS baselines slower than dense.
  7. EIE: Efficient Inference Engine on Compressed Deep Neural NetworkSong Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, William J. Dally · arXiv (ISCA 2016) · 2016Accelerator for compressed, unstructured-sparse fully connected layers: weight sparsity, weight sharing, skipping zero activations (3×); 600 mW; 120× energy saving from SRAM over DRAM, 10× from sparsity, 8× from weight sharing.
  8. Cerebras Architecture Deep Dive: First Look Inside the HW/SW Co-Design for Deep LearningSean Lie · Hot Chips 34 · 2022Fully distributed single-cycle SRAM; dataflow triggers that filter out zeros for unstructured sparsity; sparse GEMM executed as one AXPY per nonzero weight.
  9. Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300AAaron Jarmusch, Connor Vitz, Sunita Chandrasekaran · arXiv · 2026CDNA3 supports 2:4 structured sparsity; isolated FP8 sparse GEMMs (rocSPARSE) measured 0.97–1.02× vs. dense because a constant 3.5–5.8 µs encoding overhead offsets the halved math; about 1.3× per stream under concurrency.
  10. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-ShotElias Frantar, Dan Alistarh · arXiv (ICML 2023) · 2023One-shot pruning of OPT-175B and BLOOM-176B to 60% unstructured sparsity with negligible perplexity increase, no retraining; extends to 2:4 and 4:8.
  11. Mixed-Signal Computing for Deep Neural Network InferenceBoris Murmann · IEEE Transactions on VLSI Systems, vol. 29, no. 1 (open copy on NSF PAR) · 2021Analog wins only at low SNR; kT/C means each extra bit quadruples energy; 16-bit MAC ≈ 75 fJ (28 nm), 1 kB SRAM ≈ 1 pJ/byte, digital PEs ≈ 150–200 fJ/MAC at 8 bits; E_MAC = E_ADC/N + …; E_ADC = 100 fJ·ENOB + 1 aJ·4^ENOB; ENOB = B + log2(k·FS·√N); 3.8 fJ/MAC predicted at 4 bits.
  12. ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in CrossbarsAli Shafiee, Anirban Nag, Naveen Muralimanohar, Rajeev Balasubramonian, John Paul Strachan, Miao Hu, R. Stanley Williams, Vivek Srikumar · ISCA 2016 (author copy, University of Utah) · 2016Bit-serial inputs with 1-bit DACs, 2-bit cells, ADC resolution = log R + v + w (− 1); weight-flipping encoding saves an ADC bit; ADCs are 58% of tile power and 31% of tile area; a 2-bit DAC adds 63% area.
  13. A Microprocessor implemented in 65nm CMOS with Configurable and Bit-scalable Accelerator for Programmable In-memory ComputingHongyang Jia, Yinqi Tang, Hossein Valavi, Jintao Zhang, Naveen Verma · arXiv (IEEE JSSC 2020) · 2018590 kb charge-domain SRAM in-memory accelerator in 65 nm; 152/297 1b-TOPS/W at 1.2/0.85 V; bit-parallel/bit-serial precision scaling; CIFAR-10 at 89.3%/92.4% (1-b/4-b), matching digital.
  14. Benchmarking and modeling of analog and digital SRAM in-memory computing architecturesPouya Houshmand, Jiacong Sun, Marian Verhelst · arXiv · 2023Survey and cost model of analog vs. digital SRAM IMC; AIMC best peak ≈ 1,800 TOP/s/W when ADC/DAC costs are amortized by a large array; node barely affects AIMC efficiency but strongly affects DIMC; comparisons assume 50% operand sparsity.
  15. Edge AI without Compromise: Efficient, Versatile and Accurate Neurocomputing in Resistive Random-Access Memory (NeuRRAM)Weier Wan et al. · arXiv (Nature 2022) · 202148-core, 3-million-device RRAM CIM chip in 130 nm; 5–8× lower energy-delay product than prior RRAM CIM; 4-bit-weight software-comparable accuracy; list of non-idealities; noise-injection training (25.34% → 85.99% CIFAR-10).
  16. A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inferenceManuel Le Gallo, Riduan Khaddam-Aljameh, et al. · arXiv (Nature Electronics 2023) · 2022IBM HERMES Project Chip: 14 nm, 64 cores of 256×256 PCM; 63.1 TOPS at 9.76 TOPS/W for 8-bit I/O MVMs; per-row ADCs; CIFAR-10 92.81% vs. 93.67% software; precision equivalent to 3–4-bit weights; PCM drift G(t) = G(t0)(t/t0)^−ν.
  17. An analog-AI chip for energy-efficient speech recognition and transcriptionS. Ambrogio, P. Narayanan, A. Okazaki, et al. · Nature 620, 768–775 (open access, PubMed Central) · 202314 nm chip with 35 million PCM devices across 34 tiles; up to 12.4 TOPS/W chip-sustained; estimated five-chip system 6.94 TOPS/W (conventional mapping, 9.475% WER) or 6.7 TOPS/W (weight expansion, 9.258% WER) vs. 7.452% software baseline; projected 14× over the best MLPerf efficiency; vector ops, activations and the joint layer run in software.
  18. Analog Computation in Flash Memory for Datacenter-scale AI Inference in a Small ChipDave Fick, Mike Henry (Mythic) · Hot Chips 30 · 2018Flash transistors modeled as variable resistors holding the weights, with ADCs converting column currents to digital codes (slide 15); 40 nm process (slide 12); 8-bit analog compute is about half the 0.5 pJ/MAC total (slide 23); initial product holds 50M weights (slide 20).
  19. ADC Performance Survey 1997–presentBoris Murmann · GitHub (bmurmann/ADC-survey)Open dataset of ADC performance from ISSCC and the VLSI Symposium, the kind of data the E_ADC fit is drawn from.