Chips are made on a thin, round slice of silicon called a , about the size of a dinner plate. Hundreds of copies of a chip are printed on it side by side. Then the wafer is sawn apart, and each chip ends up about the size of a fingernail.1
A few AI chip designers asked a bold question. What if we didn’t cut the wafer at all, and made the whole thing one giant chip?
The reason is distance. An AI model is a huge pile of numbers that the chip reads again and again. Reading them from memory on the same chip is fast. Sending them between separate chips is slow and wastes energy. One giant chip keeps everything close.
The idea isn’t new. People tried it in the 1980s and gave up, because every wafer has a few tiny flaws.13 This chapter shows how today’s giant chips get around that.
A chip is normally one : a rectangle cut from a 300 mm that carries many copies of the design. Two things cap a die’s size. The lithography scanner prints one field of 26 × 33 mm at a time, which sets the of about 858 mm².14 And the bigger a die is, the more likely it is to catch a killer defect, so falls as area grows.15
AI workloads push against both limits. A neural network’s learned numbers, its , have to be read for every input, and moving them between chips costs far more time and energy than moving them a millimeter across one chip. This chapter covers two related answers:
- Wafer-scale integration: making one part from most of a wafer, using redundancy so that defects don’t kill it. Cerebras’s first Wafer Scale Engine, shown in 2019, covered 46,225 mm² with 400,000 cores and 18 GB of on-chip memory.1
- SRAM-heavy designs: chips of ordinary size that spend much of their area on fast on-chip memory () and little or none on external memory, with execution planned in advance by a compiler.67
Both trade capacity for speed. On-chip memory is fast but small, so a large model has to be spread across many chips or wafers. How defects are handled is covered in depth in the Transistors guide’s Making them chapter; here we use its yield models and add redundancy.
You know a neural network runs mostly on matrix multiplication. The chips in this chapter start from the observation that for many of those multiplications, especially in inference, the limit is moving weights and activations rather than doing arithmetic, and that moving data across a package or a board costs, by one vendor’s estimates, one to two orders of magnitude more energy per bit than moving it across a die.23 Two design families follow:
- stitches dozens of reticle fields into one part with wiring across the , and survives defects with fine-grained redundancy. Cerebras’s WSE-3 has 4 trillion transistors, 900,000 cores and 44 GB of SRAM on one wafer.3 Tesla’s Dojo training tile (as presented at Hot Chips 2022) and a UCLA/UIUC prototype took the chiplet route instead, bonding known-good dies onto a wafer-sized substrate.1011
- SRAM-heavy, reticle-sized chips, such as Groq’s tensor streaming processor (220 MB SRAM) and Graphcore’s IPU (896 MiB), keep the working set on-die and hand scheduling to the compiler.67
The chapter works through the engineering problems in order: the reticle and the cost of off-die links, yield with redundancy, cross-reticle wiring, power delivery and cooling, keeping weights in SRAM, , and the capacity ceiling that pushes all of these designs back to multi-chip systems. Yield statistics are derived in the Transistors fabrication chapter; Under the hood extends them to repairable parts.
Separate chips: data leaves through a few pins and crosses the board over SerDes links, about 5–10 pJ per bit by Cerebras’s estimate.
The machine that prints chips works a bit like a slide projector. It shines light through a pattern and prints one patch of the wafer at a time. Each patch is about the size of a postage stamp, and a normal chip has to fit inside one.14
So big AI computers are usually built from many stamp-sized chips wired together. Inside a chip, wires are tiny and packed by the thousand. Between chips, data has to squeeze out through far fewer connections and cross a circuit board. That trip is slower and uses much more energy.1
A wafer-scale chip still prints patch by patch. Then tiny wires join the patches into one chip.1
The comes from the scanner: it projects the mask pattern onto a field of 26 × 33 mm, then steps to the next field. High-NA EUV scanners halve the field to 26 × 16.5 mm, so the ceiling is getting lower, not higher.14 A die must fit in one field. Everything bigger is a system of several dies.
Connecting separate dies costs in three ways:
- Bandwidth. On a die, wires are spread across the whole area. Off a die, signals can only leave through its edge or its bumps, so the number of connections is far smaller.
- Energy. Long chip-to-chip links use circuits that send data fast over a few wires. Cerebras estimated about 5–10 picojoules per bit for such links, against 0.05–0.15 pJ/bit for short wires between neighbors on its wafer.23
- Latency and complexity. Traffic between chips goes through switches and protocol layers; splitting a model across devices means the software has to decide what goes where.23
Normally the dies on a wafer are electrically independent. The between them is where the saw cuts, and it also holds the fab’s test structures. Cerebras worked with its foundry to run wires across the scribe lines, extending the cores’ grid network across die boundaries so the whole wafer looks like one uniform array.1 On the second-generation part those bridges are under 1 mm long.2
The scanner field is 26 × 33 mm (858 mm²) today and 26 × 16.5 mm with high-NA anamorphic optics.14 A 46,225 mm² wafer-scale part is about 54 full fields. Each WSE-2 die is 17 × 30 mm, within one field, and is repeated in a 12 × 7 grid of 84 dies, each a 66 × 154 array of cores (the slide lists 10,156 cores per die).2
The motivation is the gap between on-die and off-die interconnect. Cerebras’s own comparison, on an equal 826 mm² of silicon (their estimate of a GPU’s links, not a measurement), is:2
| 826 mm² of silicon | Bandwidth | Energy | Power |
|---|---|---|---|
| Off-chip SerDes links (vendor estimate) | 0.6 TB/s | 10 pJ/bit | 60 W |
| Wafer fabric sub-region | 4.3 TB/s | 0.15 pJ/bit | 6 W |
The 2024 version of the comparison puts each WSE-3 die’s edge links at 2,880 GB/s counting both directions (480 parallel 24 Gb/s lanes), at 0.05 pJ/bit per end, against 5 pJ/bit per end for 100 Gb/s SerDes.3 The reason is physics, not cleverness: a sub-millimeter parallel wire needs no equalization, clock recovery or long-reach driver, so it costs about what any on-chip wire costs. The cross-scribe interface is source-synchronous and parallel, with redundancy, link training and an auto-correction state machine.2
The alternatives trade differently. Chiplets on a silicon interconnect substrate keep links short (the UCLA/UIUC system uses 10 µm copper-pillar pitch and 200–300 µm links at 0.063 pJ/bit) and let every die be tested before assembly, but add bonding yield to worry about.11 Tesla’s Dojo tile, as presented at Hot Chips 2022, combined 25 known-good compute dies and 40 I/O dies on a system-on-wafer substrate.10
No fields exposed yet. Normally the scribe line is the saw lane and holds test structures; the two dies are electrically separate.
Every wafer picks up a few tiny flaws while it is made. A speck of dust lands, or a wire doesn’t quite form. If a flaw lands on a normal chip, that chip is usually thrown away. A whole-wafer chip is sure to catch several flaws.16
In the 1980s, a famous computer designer named Gene Amdahl tried to build wafer-sized chips. He found they would only work if almost every part came out perfect, so he stopped.13
Today’s answer is to stop needing perfection. The giant chip is built from hundreds of thousands of cores. A core is a small processor, here smaller than a grain of sand. Some are . After testing, any core with a flaw is switched off and a spare takes its place. A flaw now costs one speck instead of the whole chip.14
The yield of a die with area in a process with (killer defects per cm²) is, in the simplest model, .15 At :
- a 1 cm² die yields ;
- a 6.25 cm² die (near the reticle limit) yields ;
- a 400 cm² wafer-sized die with no redundancy yields , about one in .
The Making them chapter explains where this formula comes from, covers its more realistic relatives, and lets you try it on a wafer map. For a wafer-scale part, the formula says the part never works, so the design must tolerate defects.
Redundancy does that in three steps:
- Make the unit of failure tiny. If a defect kills one core, the loss is one core’s area. Cerebras puts a WSE-3 core at about 0.05 mm².4
- Add spares everywhere. The design includes extra cores and extra links in its on-chip network. After testing, defective cores are switched off, spares replace them, and the extra links reconnect the grid.1
- Hide the repair from software. Programs always see a complete, uniform grid of cores, whatever was repaired underneath.3
Cerebras reports 970,000 physical cores on WSE-3, of which 900,000 are enabled, and says it enables about 93% of its silicon area. With an assumed 0.001 defects per mm² (0.1 per cm²), it expects about 46 defects per wafer, which cost only about 2.2 mm² of cores.4 These are the vendor’s own figures.
The same idea works on ordinary-size chips. Graphcore’s 823 mm² IPU uses “23/24 tile redundancy”: one spare tile for every 24.7 Groq’s chip is built from 21 identical horizontal slices of its datapath (“superlanes”), one of which is a spare, hidden from the compiler, so a die with one bad superlane can still be used as a fully working part.6 And it works between chiplets: Dojo assembled each tile from dies that had already passed test, and added harvesting and configurable routing for yield.10
Without redundancy, a 400 cm² monolithic part at has an expected 40 defects and Poisson yield . Negative binomial clustering helps , but that is still not a product.15 The requirement becomes: every defect must land on something that can be replaced, and there must be a replacement within reach.
Three properties make that work:
- Granularity. Area lost per defect the replaceable unit’s area. WSE-2’s core is 228 × 170 µm (0.039 mm²) with 48 kB SRAM and half its area in memory.2 WSE-3’s is about 0.05 mm², so Cerebras’s 46 expected defects remove about 2.2 mm².4
- Locality. A spare can only replace a core it can be wired in for, and the repaired network must still look like a complete rectangular mesh to software. The sim models this as independent repair regions, each with its own spares. That is why spare counts are set by the tail of the per-region defect distribution, not by the wafer-wide mean (see Under the hood).
- Everything is redundant. Cerebras provisions redundant fabric links as well as cores.1 Any non-repairable area multiplies yield by : just 1% of a 400 cm² part (4 cm²) left unprotected costs a third of all wafers at .
Cerebras’s numbers for WSE-3: 970,000 physical cores, 900,000 enabled (about 7% spare; Cerebras puts silicon utilization at 93%), against an expected ~46 defects at 0.001/mm².4 The ratio of spares to expected defects (about 1,500:1) is a reminder that locality and clustering, not the mean, set the budget; the blog does not say how the spares are placed.
The chiplet route moves the problem. In the UCLA/UIUC prototype, each chiplet is tested before bonding, so die defects are screened out, but each of its ~2,000 copper-pillar I/Os can fail to bond. With one pillar per pad (≥ 99.99% each) a chiplet bonds correctly only 81.46% of the time, which across 2,048 chiplets means about 380 bad ones. Two pillars per pad raise that to 99.998% and about one bad chiplet per wafer.11 The network then has to survive the few that remain: a single dimension-ordered mesh lost over 12% of source–destination paths with five dead chiplets, and two independent meshes (X-then-Y and Y-then-X) cut that below 2%. The clock is forwarded tile to tile from any edge tile, so it reaches every working tile unless all four of its neighbors have failed.11
2 defects, all repaired: each row skips its dead cores and shifts along to a spare, so software still sees a full 8 × 5 grid.
Making the circuits is only half the job. A dinner-plate chip also needs power, cooling and a safe way to hold it.
Power. The first wafer-scale computer used about as much power as a dozen electric kettles running at once.5 Ordinary circuit boards can’t carry that much electricity. So it is fed straight down into the face of the wafer.1
Cooling. Air can’t carry away that much heat, so water flows through a metal pressed against the wafer.1
Growing and shrinking. Silicon and circuit boards grow by different amounts when they warm up. Across a whole wafer, that would tear the connections apart. So the builders designed a special connector that can stretch a little.1
Cerebras listed five problems it had to solve for its first wafer-scale engine: connecting across dies, yield, , package assembly, and power and cooling.1 The first two are covered above. The other three are about the package.
Power delivery
A modern chip runs from a supply of around a volt, so even modest power means very large currents. On an ordinary chip, the circuit board’s copper planes carry that current to the package. For a wafer, Cerebras found the current density too high for board planes and instead sent current perpendicular to the wafer, through its face, so each patch of silicon is fed from directly behind it.1 Tesla’s 2022 Dojo tile did the same: data moved horizontally across the tile, and power arrived vertically, up to 15 kW per tile.9
The UCLA/UIUC research prototype, a smaller 725 W system, shows what happens when power comes from the edge instead. Feeding 2.5 V at the edge, the supply sags to about 1.4 V at the center from resistance in the wafer’s thin wiring, so every chiplet carries its own regulator (an ) to produce a steady 1.1 V.11 The same voltage-drop problem appears on every chip as ; see Power planning.
Cooling
Heat is removed by water flowing through a on the wafer.1 Dojo’s tile integrated its cooling with the power delivery in one module, so tiles could be added without new cooling designs.10 The Power and cooling chapter covers the rack and facility side.
Packaging
No standard package fits a wafer. Cerebras designed a connector between wafer and board that absorbs the difference in thermal expansion, and built custom tools to align the board, connector, wafer and cold plate precisely.1
Cross-reticle wiring
Scribe lines are normally keep-out: they are destroyed at dicing and carry the fab’s test structures. A monolithic wafer-scale part needs the foundry to allow product wiring through them; Cerebras did this with TSMC at 16 nm and extended it to 7 nm and 5 nm.123 The link is designed like a very short parallel bus, not a SerDes: under 1 mm, source-synchronous, with redundancy, link training and an auto-correction state machine.2 The result, per Cerebras, is the same bandwidth between dies as within one, so the compiler sees one uniform mesh.2
Power delivery
The numbers that matter are current and distance. At a supply around 0.8 V (an illustrative assumption, not a published figure), Dojo’s 15 kW tile would draw on the order of 19,000 A. Distributing that laterally through thin on-wafer metal is hopeless, because for a uniformly loaded region fed at its edge the IR drop grows with the square of its size (derived in Under the hood). Cerebras delivers, and the 2022 Dojo tile delivered, power vertically through the face of the part, so each region’s current path is short regardless of the wafer’s size.19
The UCLA/UIUC design quantifies the edge-fed alternative at much lower power. Its 1,024 tiles draw up to 350 mW each at 1.21 V, about 290 A in total, through two 2 µm-thick metal layers of the interconnect wafer. Feeding 2.5 V at the edge leaves about 1.4 V at the center under full load, so each chiplet has an LDO tracking a 1.4–2.5 V input to deliver 1.0–1.2 V.11 The costs are visible: LDO efficiency is roughly , so edge tiles waste more than half their input power; and because off-chip capacitors can only sit at the wafer’s edge, up to 70 mm away, about 35% of each tile’s area went to on-chip decoupling capacitance (about 20 nF per tile).11 Vias through the full wafer thickness were the alternative, but were not yet available for the prototype.
Clocks
A single passive clock net across 15,100 mm² and 1,024 sinks has more than 450 pF and 120 nH of parasitics, limiting it to sub-MHz. The UCLA/UIUC design instead generates a fast clock in an edge tile and forwards it tile to tile, inverting it at each hop so duty-cycle distortion alternates instead of accumulating.11 Graphcore runs its 823 mm² die from a mesochronous clock with about three cycles of global drift.7
Heat and mechanics
Heat flux, not total power, sets the cooling problem. Dojo’s 15 kW over 25 D1 dies of 645 mm² each is about 93 W/cm² of compute silicon.9 CS-1 drew 20 kW as a whole system;5 even if all of it went through the 462 cm² wafer, that would be about 43 W/cm². Cerebras described its wafer’s heat density as too high for direct air cooling and moved the heat into water through a .1 Strain from grows with distance from the part’s center; Cerebras judged that a traditional package would put too much mechanical stress on a wafer and developed a custom connector between wafer and board that absorbs the movement.1
Pick data, power or heat to follow, or warm it up. Tap a layer to name it.
Chips use two main kinds of memory. On-chip memory sits right next to the math and answers very fast. Separate memory chips sit beside the processor. They hold far more, but each trip to them is slower and uses more energy.
For AI this matters a lot. A chatbot holds billions of numbers it learned in training, called . It reads all of them once for every word it writes. The faster it reads them, the faster the words come out.3
So some chips fill much of their space with on-chip memory. They keep the whole model there, right beside the math. The dinner-plate wafer chip holds 44 gigabytes this way, about as much as ten movies.3 The catch is that on-chip memory takes up a lot of room.
stores each bit in a small circuit of about six transistors built in the same process as the logic (see Memory cells). It can be placed anywhere on the die, in many small banks, each next to the logic that uses it. stacks DRAM dies beside the processor in the same package: far denser, but it is reached through a limited number of wires at the edge of the processor.
The SRAM-heavy designs spread memory evenly across the die:
- Cerebras gives every core its own 48 kB of SRAM in eight banks, readable in one clock cycle, with about half of each core’s area taken by memory.2
- Groq’s tensor streaming processor has 220 MB of SRAM in 88 slices, delivering up to 80 TB/s, with no caches: the compiler addresses the memory banks directly.6
- Graphcore’s second IPU gives each of 1,472 tiles 624 KiB, 896 MiB in all, at 47 TB/s.7
Why it helps: when a language model generates text one token at a time, each token requires reading every weight once, and there is little arithmetic per byte read. Speed is then set by memory bandwidth, not by math.3 The memory wall chapter explains this in detail. Cerebras compares 21 PB/s of SRAM bandwidth on its wafer with about 3.35 TB/s of HBM bandwidth on one GPU.3
The price is capacity per area. Taking published figures, the wafer holds 44 GB in 46,225 mm² and Graphcore 896 MiB in 823 mm², roughly 1 MB of SRAM per mm² of chip, logic included.37 At that density a reticle-sized chip holds well under 1 GB, while one HBM-equipped accelerator package holds tens of GB: Graphcore’s own comparison lists 40 or 80 GiB of off-die memory for a GPU of the time, next to 40 MiB of SRAM on that GPU’s own die.7
For a matrix–vector product, as in single-stream decode, each weight is used once per token, so arithmetic intensity is about one multiply-accumulate per weight read: the roofline puts it firmly on the bandwidth slope. Cerebras’s slides put the bytes-per-FLOP needed at about 0.005 for GEMM, 1 for GEMV and 2–3 for vector operations, and present distributed SRAM, next to every datapath, as the way to supply that ratio across a whole chip.2
Several of these designs still pair SRAM with DRAM for capacity. Dojo’s interface processors carried 32 GB of HBM each, 160 GB per tile edge, and a full training matrix had 1.3 TB of SRAM next to 13 TB of DRAM.10 Graphcore’s chips use host DDR as their off-die memory.7 SRAM is the working store; DRAM remains the warehouse.
SRAM in many small banks, each beside the logic that uses it: total bandwidth grows with the number of cores. Graphcore’s chip: 896 MiB at 47 TB/s.
Most processors work like a busy restaurant kitchen. Orders arrive at random times, and the kitchen decides on the fly what to cook next. That is flexible, but nobody can say exactly how long a dish will take.
Some of these chips work like an orchestra playing from sheet music instead. Before the music starts, a program called the compiler writes out every note for every player: which number moves where, and on which tick of the chip’s clock.6
Because everything is planned, the same job always takes exactly the same time. Chips like this are called . They answer quickly and predictably. The price is that they can’t adapt if something unexpected happens.
In a conventional processor, hardware makes timing decisions as the program runs: caches decide what to keep, arbiters decide which request wins a shared resource, and networks route packets around congestion. Each of these makes the time an operation takes depend on what else is happening.
removes them. Groq’s design rule is “no reactive components”: no arbiters, crossbars, replay mechanisms or caches. The compiler knows where every piece of data is on the chip at every moment and schedules every instruction to the clock cycle.6 Data moves along the chip through chains of registers, one hop per cycle, so the time to get from one unit to another is a simple subtraction.6
Graphcore uses a related scheme called execution: all tiles compute independently, then synchronize, then exchange data, then compute again. The compiler schedules every transfer to a precise cycle after the synchronization, so the on-chip exchange has no queues or arbiters.78
What it buys and what it costs:
- Predictable latency. A response takes the same time every run, which matters when a service promises fast replies.
- Less control hardware. Groq reports that instruction dispatch takes under 3% of the chip area.6
- Everything must be known in advance. Shapes and data placement must be fixed at compile time, and anything that varies at run time is harder to handle. Graphcore noted that whole-graph compilation grew slow as models grew.7
Not every wafer-scale design is deterministic. Cerebras cores start work when data arrives over the network (dataflow triggering), which lets them skip zeros in sparse data, though the routes themselves are configured statically.2 The Dataflow chapter covers how compilers map layers onto meshes like these.
Groq’s tensor streaming processor is the cleanest published example. The chip is functionally sliced: columns of memory, vector, matrix and data-reshaping units, with instructions flowing vertically to 144 independent instruction queues and operands flowing horizontally on eastward and westward stream-register paths.6 The design rules:
- No arbiters, crossbars, replay or caches; a defined memory consistency model with no reordering.
- The compiler tracks every tensor’s location and schedules at cycle granularity.
- A stream register is one hop per cycle, so travel time is a subtraction.
- NOP, SYNC and NOTIFY instructions let the compiler align the 144 queues; the result can be seen as a fully pipelined 144-wide VLIW machine.
Peak is about 750 TOPS of INT8 at 900 MHz from four 320 × 320 matrix engines, and instruction dispatch costs under 3% of area.6
Determinism is harder across chips. Link latency varies, there is no shared clock, and clocks drift. Groq gives each chip a free-running hardware-aligned counter, deskews links at run time, paces links in software so they never overflow or underflow, and uses forward error correction, which corrects most transmission errors deterministically. All on-chip SRAM and datapaths carry SECDED error-correcting codes. The chip-to-chip network is then software-scheduled, with no hardware flow control.6
Graphcore’s model relaxes determinism to the superstep: tiles run asynchronously within a compute phase, then a hardware barrier (about 150 cycles on chip, 15 ns per hop between chips) precedes an exchange that the compiler schedules cycle by cycle, with addressing “by time and select state” over a 1,600-way receive multiplexer per tile.7 A superstep costs , so load imbalance between tiles is paid every step. Graphcore also noted that bulk synchrony makes it harder to tune out supply-voltage margin for supply transients (all tiles switch phase together, so their current steps coincide).7
Cerebras sits between the two: 24 statically configured routes (“colors”) per link, each with its own buffering, so the network is non-blocking and its paths fixed at compile time, while cores fire handlers when data arrives and skip zero operands.2 The common thread is that the compiler, not run-time hardware, owns placement and routing. That is a natural fit for SRAM-only machines, where every access latency is fixed.
Finished at tick 18. Red stalls are cache misses and lost arbitration; run it again.
Here’s the catch with keeping everything in on-chip memory: there isn’t much of it. A large AI model can have 70 billion numbers or more. That takes about 140 gigabytes to store.3
Even the dinner-plate wafer chip, with 44 gigabytes, needs four of them for that model. A normal-size chip built the same way would need hundreds.3
So these designs end up as teams of chips after all, working like an assembly line.3
The arithmetic is simple. Weight memory = parameters × bytes per parameter. A 70-billion-parameter model needs 140 GB in a 16-bit format, 70 GB in an 8-bit format (see Number formats). Chips needed = weight memory ÷ usable memory per chip, rounded up.
- Cerebras maps Llama 3.1-8B (16 GB at FP16) onto one wafer and Llama 3.1-70B (140 GB) onto four, using 176 GB of SRAM.3
- At 220 MB per chip, holding 140 GB of weights in SRAM takes at least 637 Groq-style chips, or 319 at 8 bits, before any room for anything else.6 (Our arithmetic, not a published deployment.)
And weights are not all. Serving a language model also needs a that grows with conversation length and number of users. Training needs far more: gradients and optimizer state on top of the weights.
Designs have taken three routes around the limit:
- Pipeline across chips or wafers. Each chip holds a few layers (); only activations cross between them. Cerebras reports that its four-wafer 70B mapping needs about 100 Gb/s of the 1.2 Tb/s of I/O each system has.3
- Stream the weights. For training, Cerebras keeps weights off the wafer in an external memory system and streams them in one layer at a time (), sending gradients back out for the update.2
- Add DRAM after all. Dojo attached HBM through separate interface processors; Graphcore uses the host’s DDR memory.107
The parallelism chapter in Systems explains the ways a model can be split, and what each costs in communication.
Weight footprint is bytes; usable memory per device is , with the share reserved for , activations, code and buffers. The device count is , and in a pipeline mapping it is also the minimum number of stages. Published data points:
- WSE-3 (44 GB): Llama 3.1-8B (16 GB at FP16) on one wafer; Llama 3.1-70B (140 GB) on four, 176 GB total, which leaves about 20% of the SRAM () for everything else. Layers are mapped to wafer regions sized by their memory and compute needs, and the model’s weights and KV cache stay in the region’s SRAM.3
- Between wafers, only activations cross: Cerebras reports needing about 100 Gb/s of 1.2 Tb/s of I/O per system, and says its four-wafer 70B mapping needs four network hops through I/O of under 5 µs, under 1% of latency.3
The same arithmetic for a 220 MB device gives devices for FP16 weights and 319 for FP8, before any ; a pipeline that deep has hundreds of chip-to-chip hops per token, which is why multi-chip determinism matters so much to that design.6
Training changes the ratio. With mixed-precision Adam, model state is about 16 bytes per parameter (FP16 weights and gradients plus 12 bytes of FP32 master weights and two moments),12 so a 3B-parameter model’s state alone (48 GB) would overflow a 44 GB wafer. Cerebras’s answer is : an external store (MemoryX) holds weights and runs the optimizer; weights are broadcast onto the wafer one layer at a time; gradients stream out and are reduced on the way. Scaling to many systems is then data-parallel only.2 This swaps the capacity limit for an external-bandwidth requirement, and only works because a training batch provides enough reuse per streamed weight.
Tesla’s system shows the same pressure from the other side: a training matrix with 1.3 TB of SRAM is paired with 13 TB of high-bandwidth DRAM, and the software keeps a single copy of parameters, replicating them just in time.10
A token enters the first stage. It must pass through every layer of the model, in order.
The top panel shows the same square of silicon twice, with the same flaws (the dots). On the left, it is cut into small chips, and any chip with a flaw is thrown away. On the right, it stays one giant chip with a few spare cores in each block.
Raise the flaws slider and watch what happens. Then set the spares to zero: the giant chip always fails. The bottom panel asks how many chips it takes to hold a whole AI model.
Panel 1 cuts 400 cm² of silicon two ways and puts the same defects on both. On the left, conventional dies fail if they contain any defect; on the right, a wafer-scale part is divided into 400 regions of 2,000 small cores, and a region is repaired as long as its defects don’t exceed its spare cores. Panel 2 computes how many chips it takes to hold a model’s weights. Things to try:
- Set spares to 0. Why does the wafer-scale part fail even at low defect density?
- Find the fewest spares that keep the part working at , then at 0.5 and 1.0.
- Compare the usable silicon of 10 mm and 25 mm dies as rises.
- Load a 70B model at BF16 and switch between a reticle-size SRAM chip and a wafer-scale one.
Panel 1 shows one deterministic sample wafer (defects placed by a fixed seed) and, alongside, the model’s expectation: Poisson or negative binomial () die yield for the conventional dies, and for the wafer-scale part. The Expert view adds per die and per region and the probability that every region is covered. Panel 2 adds an HBM capacity slider and a reserve for KV cache and activations, and converts the SRAM needed into silicon area at about 1 MB/mm². Try:
- At , how many spares per region give 99% part yield under Poisson? Switch to clustered: how many now? Why does clustering help the dies and hurt the wafer?
- At , compare the spares needed with the expected defects per region.
- Set a 405B model at FP8 with a 20% reserve. How many wafers of SRAM is that?
The model assumes each defect kills exactly one core and that links, routers and clocks never fail, which makes it optimistic; see Under the hood for what that leaves out.
- Scanner field (reticle)
- 26 × 33 mm
- WSE-3 silicon
- 46,225 mm²
- WSE-3 on-chip SRAM
- 44 GB
- Groq TSP on-chip SRAM
- 220 MB
Sources: field size from IEEE Spectrum; WSE-3 figures from Cerebras’s Hot Chips 2024 slides; TSP figure from Groq’s Hot Chips 34 slides.1436
What these numbers mean:
- 26 × 33 mm is the biggest patch the printing machine can make at once. That is a bit bigger than a postage stamp. A normal chip has to fit inside it.
- 46,225 mm² is a square about 21.5 cm on each side, cut from a round wafer. It covers about 54 of those stamp-sized patches.
- 44 GB of on-chip memory is enough for a mid-sized AI model. A large one still needs four wafers.3
- 220 MB is what a normal-size chip built around on-chip memory holds. That is about 200 times less than the wafer.
Today’s wafer chip catches about 46 flaws on every wafer, its maker says. It keeps working, because each flaw costs only one tiny core.4
Reading the table:
- Wafer-scale parts hold tens of GB of SRAM; reticle-sized SRAM-heavy chips hold hundreds of MB. Either way, that is far less than one GPU package with HBM.
- Every generation of the monolithic wafer kept the same 46,225 mm² and 84-die grid and gained from the process node: cores went from 400,000 to 900,000 and SRAM from 18 to 44 GB.
- The chiplet-based designs (the 2022 Dojo tile and the UCLA/UIUC prototype) build their wafer from dies tested before assembly.
Ratios worth computing from the table:
- SRAM per area: about 0.95 MB/mm² (WSE-3), 1.1 MB/mm² (GC200), 0.68 MB/mm² (D1), logic included.
- Core size: 46,225 mm² ÷ 970,000 physical cores ≈ 0.048 mm² per core including everything else on the wafer, consistent with the 0.05 mm² Cerebras quotes.4
- Generational scaling at fixed area: WSE → WSE-3 (16 nm → 5 nm) is 3.3× the transistors but 2.25× the cores and 2.4× the SRAM, so each core grew (about 1.5× the transistors per core) while its SRAM stayed at 48 kB.13 The area cost of SRAM at advanced nodes is covered in The memory wall.
- Bandwidth per capacity: WSE-3 offers about 0.48 PB/s per GB of SRAM, Groq about 0.36 PB/s per GB. A GPU with 3.35 TB/s of HBM bandwidth over tens of GB is at about 0.0001 PB/s per GB or less, three to four orders of magnitude lower.3
Defect arithmetic, from the vendor: ~46 defects per wafer at 0.001/mm², costing ~2.2 mm², against 70,000 spare cores.4 CS-1 drew 20 kW as a system;5 a Dojo tile was specified at 15 kW.9
Going big and keeping everything on the chip means making some bargains. The see-saws below show four of them.
None of these is right or wrong. These chips are great at answering one person’s question very quickly. They are worse at holding a gigantic model cheaply.
SRAM vs. HBM
SRAM gives enormous bandwidth and single-cycle access, but at roughly 1 MB per mm² it holds far less per chip than HBM.37 It wins when the model fits and every byte is read often, as in generating text for one user at a time. It loses when capacity dominates, as in very large models, long conversations (whose KV cache grows), or training.
Wafer-scale vs. many chips
A wafer replaces off-chip links with on-chip wires, cutting energy per bit by roughly two orders of magnitude by its maker’s estimate.3 In exchange it needs foundry cooperation to wire across scribe lines, custom power delivery and cooling, and a custom package, and it is a single, very expensive unit of failure if repair doesn’t cover a wafer.1
Monolithic wafer vs. chiplets on a wafer
Bonding tested chiplets lets each die be screened before assembly and mixes die types (Dojo combined compute and I/O dies), but adds bonding yield and makes every link a bonded connection.1011
What goes wrong
- Unprotected logic. Any part of the wafer without a spare is a single point of failure for the whole wafer. That is why links, not just cores, need redundancy.1
- Voltage sag. Feeding power from the edges leaves the center short; the UCLA/UIUC design saw 2.5 V fall to about 1.4 V.11
- The model outgrows the chip. A model that no longer fits forces a re-mapping across more chips or wafers, with new communication costs.3
- Compiler bottleneck. Static scheduling moves complexity into software; Graphcore found whole-graph compilation grew slow as models grew.7
Capacity, bandwidth and batch
SRAM-resident designs win on bandwidth per byte stored, which is what single-stream decode needs. At high batch, HBM devices reuse each weight across many sequences, raise arithmetic intensity and close the gap, while the SRAM design must spend capacity on each sequence’s KV cache. Cerebras’s pitch is the low-latency end of that curve; it argues that surplus bandwidth also lets several users share the pipeline.3 The Comparing chips chapter discusses how to compare such claims on equal terms.
Cost structure
A wafer-scale part converts die-yield risk into spare area plus a small probability of losing a whole wafer. Wafer-scale integration has long been pitched as saving the cost of packaging and testing many separate chips,13 but it adds a custom package, connector and cooling of its own.1 Cost per good die (wafer cost ÷ good dies) for conventional parts grows much faster than area;16 for a repairable wafer it is roughly wafer cost ÷ P(part works), with the spare fraction as an additional tax on delivered compute.
Failure modes
- Non-repairable critical area. Yield is multiplied by ; 4 cm² unprotected on a 400 cm² part costs a third of wafers at 0.1/cm².
- Clustered defects exhaust local spares even when the mean is low; spare budgets are set by the tail.15
- Assembly yield. For chiplet-on-wafer, bonding defects at thousands of I/Os per die can dominate unless pads are duplicated.11
- Supply and clock distribution. Edge-fed supplies droop quadratically with size; passive clock nets become too slow; both push toward local regulation and forwarded clocks.11
- Synchronized power transients. Bulk-synchronous execution moves all tiles between phases together, complicating voltage-margin reduction.7
- Cross-chip determinism. Link latency variation and clock drift must be compensated, and link errors corrected in a way that doesn’t disturb the schedule.6
Each see-saw leans toward the side of the bargain this design takes. Tap one for the details.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Four models behind the chapter and the simulation: yield with regional repair, the cost of non-repairable area, the IR drop of an edge-fed wafer, and capacity-driven device counts.
1. Die yield, briefly
With defects falling independently at density , the number in area is Poisson with mean , so
and the zero-defect yield is . If itself varies with a gamma distribution, the count is negative binomial with cluster parameter and
is close to Poisson, close to Seeds’ model.15 Derivations and the Murphy model are in the fabrication chapter.
2. Yield with regional repair
Split the part into independent repair regions, each with mean defects and spare units that can replace any failed unit in the region. Assume each defect kills one unit and units are small enough that two defects rarely share one. Then:
For , (1 cm² regions at ), Poisson:
| Spares per region | P(region ok) | P(part ok) |
|---|---|---|
| 0 | 0.90484 | |
| 1 | 0.99532 | 0.153 |
| 2 | 0.999845 | 0.940 |
| 3 | 0.9999962 | 0.998 |
Three observations. First, the step from to is enormous: each extra spare cuts the per-region failure probability by roughly a factor , so part yield climbs very steeply once the tail is covered. Second, the spare tax here is tiny (2 of 2,000 cores is 0.1%). Third, the required grows with : at , gives 93% and gives 99.4%; at , gives 99.6%.
Clustering changes the tail. For the negative binomial is geometric,
At that is : gives (versus 0.94 Poisson), gives 0.9975. The same clustering raises conventional die yield (a 6.25 cm² die goes from 54% to 62%), because defects that bunch together spare other dies. Repairable parts care about the worst region; conventional dies care about the typical one.
Delivered compute is for units per region. Maximizing it over gives the spare budget; in practice the budget is set well above this simple optimum to cover what the model ignores: defects that kill several adjacent units, parametric failures (slow or leaky cores), test escapes, and failures in the repair logic itself. Cerebras’s ~7% spare cores against ~46 expected defects suggests real budgets are far more conservative than this model.4
3. Non-repairable area
Any area whose failure cannot be repaired multiplies part yield by . With and :
- 1% unprotected (4 cm²): .
- 0.1% unprotected (0.4 cm²): .
That is why wafer-scale designs make links, routers and clock paths redundant too.111 The UCLA/UIUC dual-mesh result is the network analogue: with one dimension-ordered mesh, a single dead chiplet cuts every path that routes through it, and five dead chiplets disconnected over 12% of pairs; a second mesh routed Y-then-X gives most pairs a second path and brought that under 2%.11
4. IR drop on an edge-fed wafer
Model the power grid as a disc of radius with sheet resistance (Ω per square), drawing a uniform current density (A/m²), fed at its rim. The current crossing a circle of radius is everything drawn inside it, . Spread over the circumference , the sheet current density is , so the field is . Integrating from the center to the rim:
At fixed power density, drop grows as : doubling the part’s radius quadruples the droop. Total current grows as too, so raising the edge voltage only helps until regulators and losses take over, which is the compromise the UCLA/UIUC prototype made (2.5 V at the edge, about 1.4 V at the center, LDOs to 1.1 V).11 Vertical delivery replaces with the spacing between feed points, which is independent of the part’s size. That is the geometric argument for feeding power through the face, as Cerebras does and the 2022 Dojo tile did.19
Edge-fed, R = 150 mm: center droop ρs·j·R²/4 = 36× the 25 mm reference. The UCLA/UIUC prototype saw 2.5 V at the edge fall to about 1.4 V at the center.
5. Device count and pipeline depth
with parameters, bytes per parameter, memory per device and the reserve for KV cache, activations and buffers. In a pipeline mapping, , and each token crosses boundaries. Two consequences:
- Per-token latency includes , so deep SRAM-only pipelines need very fast, predictable links. Cerebras reports that its four-wafer 70B mapping needs four network hops through I/O of under 5 µs, under 1% of latency.3
- Inter-device bandwidth per token is one activation vector per boundary (hidden size × bytes), tiny next to the weights each stage reads locally, which is why pipeline splits suit capacity-limited SRAM designs.3
For training, is replaced by the full optimizer state (about 16 bytes per parameter with mixed-precision Adam), the problem Cerebras’s weight streaming addresses by keeping weights and the optimizer in an external store.2
6. What the simulation leaves out
- Defects that hit links, routers, clocks or I/O (it treats all area as repairable cores).
- Defects large enough to kill several cores, and parametric failures.
- Scribe lanes, edge exclusion and partial dies: both options use the same 200 × 200 mm square.
- Correlation between regions on the same wafer: the clustered case uses an independent negative binomial per region for the expectation, and a heuristic cluster pattern for the sample map.
- Memory for anything but weights, unless you set the reserve.
Q1With a defect density of 0.1 per cm², roughly what fraction of 6.25 cm² dies have no defect, using ?
Q2Why do wafer-scale designs use very small cores?
Q3A 70-billion-parameter model is stored in a 16-bit format. About how much memory do its weights need?
Q4What makes an SRAM-only, deterministically scheduled chip easier for its compiler to plan?
Sources
Show Hide 16 sources
- Wafer-Scale Deep Learning (Hot Chips 31 slides)WSE: 46,225 mm², 1.2 trillion transistors, 400,000 cores, 18 GB on-chip memory, 9 PB/s memory and 100 Pb/s fabric bandwidth, 16 nm. Challenges: cross-die connectivity, yield, thermal expansion, package assembly, power and cooling. Wires added across the scribe line with the foundry; redundant cores and fabric links; a custom wafer-to-PCB connector; power delivered perpendicular to the wafer; water-cooled cold plate.
- Cerebras Architecture Deep Dive: First Look Inside the HW/SW Co-Design for Deep Learning (Hot Chips 34 slides)WSE-2: 850,000 cores, 40 GB SRAM, 20 PB/s, 7 nm; 84 dies of 17 × 30 mm in a 12 × 7 grid; 10,156 cores per die; 228 × 170 µm core with 48 kB SRAM in 8 banks, 50:50 logic to SRAM area, 1.1 GHz, 30 mW; 2D mesh with single-cycle hops and 24 static routes; bridges under 1 mm across scribe lines with redundancy and training; fabric 0.15 pJ/bit vs a 10 pJ/bit SerDes estimate; weight streaming from off-wafer MemoryX.
- Wafer-Scale AI: GPU Impossible Performance (Hot Chips 2024 slides)WSE-3: 4 trillion transistors, 900,000 cores, 44 GB SRAM, 21 PB/s; 48 kB SRAM per core; 84 dies; die-to-die links crossing reticle boundaries extended to 5 nm, with built-in redundancy so software sees a uniform mesh; 0.05 pJ/bit on-wafer vs 5 pJ/bit SerDes; layers mapped to wafer regions as a pipeline; Llama 3.1-70B (140 GB at FP16) on 4 wafers, passing only activations between them.
- 100x Defect Tolerance: How Cerebras Solved the Yield ProblemWSE-3 core about 0.05 mm²; 970,000 physical cores with 900,000 active (93% of the silicon area enabled); about 0.001 defects per mm² assumed; about 46 defects on a 46,225 mm² wafer costing about 2.2 mm² of cores; redundant routing around failed cores.
- Fast Stencil-Code Computation on a Wafer-Scale ProcessorCS-1: a 462 cm² wafer with 380,000 cores, each with 48 KB of SRAM (18 GB total), one-cycle load-to-use latency, a 7 × 12 array of 84 identical dies, and 20 kW total system power.
- The Groq Software-defined Scale-out Tensor Streaming Multiprocessor (Hot Chips 34 slides)Design for determinism: no arbiters, crossbars, replay or caches; the compiler knows every tensor’s location and schedules instructions cycle by cycle; 220 MB on-chip SRAM in 88 slices at up to 80 TB/s; flat memory hierarchy; about 750 TOPS INT8 at 900 MHz; under 3% of area for instruction dispatch; multi-chip determinism with aligned counters, deskew and a software-scheduled network.
- Graphcore Colossus Mk2 IPU (Hot Chips 33 slides)GC200: 59.3 billion transistors, 7 nm, 823 mm², 1,472 tiles with 624 KiB each, 896 MiB SRAM at 47 TB/s; 23/24 tile redundancy; bulk-synchronous execution with on-chip sync in about 150 cycles; compiler-scheduled exchange with no queues or arbiters; GPU and TPU comparison of on-die and off-die memory; large SRAM moves DRAM power out of the logic die.
- Dissecting the Graphcore IPU Architecture via MicrobenchmarkingFirst-generation IPU: 1,216 tiles with 256 KiB each, 304 MiB per chip, 45 TB/s aggregate memory bandwidth at 6-cycle latency; the bulk synchronous parallel model of compute, exchange and barrier phases.
- The Microarchitecture of Tesla’s Exa-Scale Computer (Hot Chips 34 slides)D1 die: 7 nm, 645 mm², 354 nodes with 1.25 MB SRAM each, 440 MB total, 2 GHz; training tile of 5 × 5 known-good D1 dies; 15 kW per tile with vertical power delivery and cooling and a horizontal data plane; 11 GB of SRAM per tile.
- Super-Compute System Scaling for ML Training (Hot Chips 34 slides)System-on-wafer tile of 25 compute dies and 40 I/O dies; known-good dies, harvesting and configurable routing for yield; interface processors with 32 GB HBM each, 160 GB per tile edge; a training matrix with 1.3 TB SRAM and 13 TB DRAM; workloads run almost entirely out of SRAM.
- Designing a 2048-Chiplet, 14336-Core Waferscale Processor1,024 compute and 1,024 memory chiplets on a silicon interconnect wafer, 15,100 mm², 725 W peak; edge power delivery with supply falling from 2.5 V at the edge to about 1.4 V at the center, per-chiplet LDOs and about 35% of tile area spent on decoupling capacitance; clock forwarded from edge tiles; two pillars per pad lift chiplet bonding yield from 81.46% to 99.998%; two dimension-ordered networks cut disconnected paths from over 12% to under 2% with five faulty chiplets.
- ZeRO: Memory Optimizations Toward Training Trillion Parameter ModelsMixed-precision Adam keeps fp16 parameters and gradients (2 + 2 bytes) plus 12 bytes of fp32 optimizer state per parameter: 16 bytes in total.
- Wafer-scale integrationDefinition; Trilogy Systems (1980, Gene Amdahl) and its failure, with Amdahl’s remark that it would need 99.99% yield; other 1980s attempts; routing around damaged sub-circuits.
- This Machine Could Keep Moore’s Law on TrackToday’s EUV scanners expose a 26 × 33 mm field on the wafer; high-NA anamorphic optics halve it to 26 × 16.5 mm.
- Yield Modeling and Analysis (IEOR 130 course notes)Killer defects and die yield; the Poisson model; clustering makes Poisson pessimistic for large dies; the negative binomial model with cluster parameter α (α ≥ 10 ≈ Poisson, α = 1 ≈ Seeds).
- Cost (EEC 116 lecture handout)Dies per wafer; the negative binomial yield formula; die cost = wafer cost ÷ (dies per wafer × yield); a 43 mm² die at 71% yield against a 296 mm² die at 9%.