Pick up a modern AI processor and you might think you’re holding one chip. Sometimes you’re holding a dozen or more. Under the lid, several pieces of silicon sit side by side, and some are stacked on top of others.9 Each piece is called a . The case that holds them and wires them together is the .
The hard part is the walkways. Inside a chip, signals zip along billions of tiny wires. To get from one chip to another, they must pass through little metal dots on each chip’s face. The more dots you can squeeze in, the more data can flow. That’s what this chapter is about.
For most of computing history a processor was one die (a single rectangle of silicon) in its own , the carrier that connects it to a circuit board and carries heat away. The largest chips now break that rule. A design splits one product into several dies and joins them inside one package, side by side or stacked, with stacks of memory placed right beside the compute dies.19
This chapter covers:
- Why a design gets split at all: the size limit of chipmaking tools, yield, and mixing processes.
- The three ways to wire dies together: through an ordinary package board, through a slab or strip of silicon (2.5D), or by stacking them (3D).
- The single number that ties these together: how closely the connections can be spaced.
- HBM memory stacks, and UCIe, an open standard for links between chiplets.
- What it costs in heat, yield and money.
It builds on two other chapters. Yield and the reticle limit come from Making them, and bumps, wafer sort and assembly from After tapeout. Here the question is what the package does for the system.
You know an AI accelerator is a big chip with fast memory. This chapter is about the layer between the die and the board, where most of the interesting system decisions now happen: how a design is partitioned into , how those dies are joined, and what the joins cost.
UCIe’s own list of key metrics is a good map of the chapter: (per mm of die edge and per mm²), which comes from data rate and ; ; latency; channel reach; reliability; and the cost difference between standard and advanced packaging.4 We’ll derive the first two from bump geometry, follow them through 2.5D, 3D and , and finish with the yield and cost model that decides whether a split pays.
The yield statistics themselves are in Making them; bumping, sort and known-good-die test are in After tapeout. Why accelerators need this much memory bandwidth is the subject of The memory wall.
From outside: one package under a metal lid. Lift it to see the dies.
There are three good reasons to build one product from several pieces.
First, there’s a size limit. Chips are printed with light, and the machine can’t print anything bigger than about 26 by 33 millimeters. That’s a bit bigger than a postage stamp.2 If you want more chip than that, you need more than one piece.
Second, dust. One speck can ruin the chip it lands on. A big chip is more likely to get hit, and with small pieces each speck wastes less.8
Third, each part can come from the factory that suits it. The newest factories make the best switches, but they cost a fortune. Some parts don’t get much better there, so they’re made somewhere cheaper.1
Splitting isn’t free. Each piece needs extra circuits to talk to its neighbors, and joining them costs money.1 So it only pays off for chips that would be very big or very costly.
The reticle limit
Lithography scanners expose one field at a time, and today’s field is 26 × 33 mm, so no single die can exceed about 858 mm². High-NA EUV, the next generation of tools, halves the field to 26 × 16.5 mm.2 Server CPUs and GPUs have been growing toward that for years.1 Chiplets let the total silicon in one product exceed it: the four dies of AMD’s first-generation EPYC server processor already added up to more than one reticle.1
Yield
A die’s (the share that come out working) falls quickly with area, because a bigger die is more likely to catch a killer defect, and fewer large rectangles fit on a round wafer.1 Four small dies tested separately waste far less silicon per defect than one large one. AMD estimated that its first EPYC, built from four chiplets, cost 41% less to make than an equivalent monolithic (single-die) design would have.1 The yield arithmetic is worked through in Making them; the simulation below applies it to a split.
Mixing processes
Once the pieces are separate, each can use the that suits it. In AMD’s second-generation EPYC, the CPU cores sit on 7 nm dies, while the memory and PCIe interfaces sit on a 12 nm I/O die. Those interfaces are mostly analog circuits that barely shrink on a newer node, so the expensive process would buy almost nothing for them.1 This is .
Reuse
The same EPYC generation used one I/O die with two to eight core dies to make a whole product line from 16 to 64 cores, from just two designs instead of up to four separate chips.1
The bill
Each chiplet must carry its own clocking, power management, test and debug circuits, plus the interfaces that talk to its neighbors, so two dies’ worth of function takes more than two dies’ worth of silicon.1 Then come the package, the assembly steps, and the energy and delay of every signal that crosses between dies.
Four forces
- Field size. The 26 × 33 mm scanner field caps a monolithic die at about 858 mm², and high-NA EUV’s anamorphic optics halve it to 26 × 16.5 mm.2 Die sizes of server CPUs and GPUs have been approaching that cap.1
- Defect-limited yield. Reticle-scale dies are expensive because area raises the chance of one or more defects.1 Splitting turns a single large Bernoulli trial into several small ones whose failures can be discarded independently, provided each die is tested before assembly.
- Node economics. Analog PHYs and much of the I/O die don’t scale, so a 7 nm core die plus a 12 nm I/O die beats an all-7 nm part.1 The same logic now puts compute on N5 and cache/I/O on N6 in a single accelerator.8
- Reuse and NRE. One I/O die and one core die covered 16 to 64 cores; a monolithic line could have needed four tapeouts, each with its own mask set and physical design.1
What it costs
Per-chiplet overhead is real silicon: replicated clocking, power management, test and debug, plus the die-to-die PHYs.1 Feng and Ma’s cost model assumes about 10% of each chiplet’s area for die-to-die interfaces, following EPYC. With that, they find splitting can save up to half the die cost, but packaging took about 30% of the total in the 16-core case they modeled, and for mature nodes the D2D and packaging overheads ate most of the savings.3 Their summary rule: splitting pays when the cost of die defects exceeds the total cost the packaging adds, and two or three chiplets usually capture most of the benefit.3
Non-recurring cost pulls the other way. Every distinct die needs its own masks and sign-off, so for a single product at modest volume a monolithic die is often cheaper overall; chiplets win through reuse across many products.3 The opposite bet, a whole wafer as one die with redundancy instead of packaging, is in Wafer-scale.
Monolithic: 800 mm² each, within the 858 mm² field. Specks kill 1 of 1: 100% of the silicon scrapped.
4 core dies (7 nm) + 1 I/O die (12 nm) = 32 cores. Two designs cover the whole range, 16 to 64 cores.
A chip connects to the outside through tiny metal dots on its face, called bumps. The chip is flipped over, and its bumps are soldered onto a small board inside the package. Wires in that board can join two chips that sit on it.
That board is made a lot like the green circuit board inside a computer. Its wires are fairly thick, so the bumps must sit about a tenth of a millimeter apart. That’s about the width of a hair.4
The fix is a thin slab of silicon under the chips, or just a small strip under the edges where they meet. Its wires are much finer.6 So the bumps can sit two to five times closer, and many times more of them fit.4
Bumps and the package substrate
In a package, a grid of solder bumps covers the die’s face. The die is flipped onto the , a multilayer laminate (often called an organic substrate, because its insulator is a resin) that routes signals out to the board. The spacing of the bumps, the , is the number to watch: on an organic substrate the UCIe standard assumes 100–130 µm.4 Bumping itself is described in After tapeout.
Two dies on the same substrate can talk through its wiring. That’s cheap and reaches far (UCIe allows up to 25 mm), but the laminate’s lines and spaces are above 10 µm wide, so only a few signals fit between bumps.43
Interposers and bridges (2.5D)
A is a passive slab of silicon with chip-style wiring on top and (TSVs, vertical metal plugs) to pass signals and power down to the substrate. Its wiring can have lines and spaces below a micrometer.3 The dies attach to it with at 25–55 µm pitch, and the link between neighbors can be up to 2 mm long.4
A full interposer has a cost: every signal and power connection of every die must pass through it.6 A puts fine wiring only where it is needed. A small silicon die with several wiring layers is embedded in a cavity in the organic substrate, under the facing edges of two dies. Only that region needs fine-pitch microbumps; the rest of each die keeps ordinary coarse bumps.6 Intel’s EMIB, in volume production since 2017, is the best-known example.6
Shoreline
Side by side, only a strip of bumps behind each die’s facing edge can join the neighbor. The length of edge used is the (or beachfront), and side-by-side link bandwidth is quoted per millimeter of it.4 That figure, the , is the number of signal bumps behind each millimeter of edge times the data rate of each. UCIe’s targets:
| Standard package (organic) | Advanced package (interposer or bridge) | |
|---|---|---|
| Bump pitch | 100–130 µm | 25–55 µm |
| Reach | up to 25 mm | up to 2 mm |
| Bandwidth per mm of edge | 28–224 GB/s | 165–1,317 GB/s |
| Energy target | 0.5 pJ/bit | 0.25 pJ/bit |
Source: UCIe 1.1 targets; the bandwidth ranges span data rates of 4–32 Gb/s per lane.4 The advanced package gives about six times the bandwidth per millimeter at half the energy per bit, but its links must be short and the silicon costs more.
Three substrates, three wiring densities
Feng and Ma summarize the ladder: organic substrate (lines and spaces above 10 µm, highest per-lane rates), fan-out redistribution layers (above 2 µm) and silicon interposers (above 0.4 µm, lowest per-lane rates but the highest pin counts).3 Die-to-die bandwidth is pins × rate, and finer wiring lets you escape more pins from a given edge. UCIe accordingly defines two physical profiles: a standard-package profile at 100–130 µm bump pitch with up to 25 mm of reach, and an advanced-package profile at 25–55 µm with up to 2 mm.4
Organic-substrate links are not obsolete. AMD’s first EPYC ran high-bandwidth SerDes links across the package substrate between its four dies, avoiding the cost of a silicon interposer.1
Interposer versus bridge
A full interposer gives every die access to fine wiring, but all signal and power vias must traverse it, and its area grows with the whole die complex. An embedded bridge confines fine-pitch wiring to the die-to-die region and lets each link use its own bridge, while the rest of each die keeps coarse bumps straight to the substrate.6 The price is two bump pitches on one die: fine at the bridge, coarse everywhere else. Variants add TSVs to the bridge to feed vertical power to HBM.6
Bandwidth per millimeter, step by step
- Bumps on a square grid of pitch give bumps per mm². At 45 µm that is per mm².
- Not all are signals: power, ground, clocks, valid and sideband take a share. Call the signal share .
- Only a strip of depth behind the edge is used, so signals per mm of shoreline .
- Multiply by the lane rate and divide by 8 for bytes: shoreline density .
UCIe’s own bump maps pin this down. An x64 advanced module occupies 388.8 µm of shoreline at every pitch; its depth is about 388 µm at 25 µm pitch (16 bump columns), 1,043 µm at 45 µm (10 columns) and 1,585 µm at 55 µm (8 columns).4 At 32 GT/s a module moves 256 GB/s each way, and 512 GB/s over 0.3888 mm is the 1,317 GB/s/mm at the top of the advanced range; a standard x16 module moves 64 GB/s each way.4 In the simulation, and reproduce that figure at 45 µm.
Organic substrate: cheap and long-reaching, but lines and spaces above 10 µm let only a few signals out per mm of edge. 0.5 pJ/bit.
The next step is to stack chips, one face down on top of another.7 Now the whole area where they overlap can hold connections, not just one edge.
For the tightest packing, engineers skip the solder and fuse copper pads straight together. This is called .7 The pads can sit just 9 thousandths of a millimeter apart. That fits many times more connections than a silicon slab can.7
The problem is heat. The bottom chip’s heat has to pass through the top chip to escape. And two hot spots stacked together get hotter than either one alone.14
From edges to areas
In , one die sits on another and connections run vertically across their overlap. carry signals and power through a thinned die when a third layer or the package must be reached. Because the whole overlap can hold connections, the count scales with area divided by pitch squared, not with edge length.5
Hybrid bonding
Solder bottom out at pitches of tens of micrometers. removes the solder. Copper pads are built into each die’s top surface, surrounded by silicon oxide and recessed very slightly. After a chemical treatment of the oxide, the dies are pressed face to face so the oxide bonds; annealing then makes the copper expand across the gap and fuse. Production parts place bonds about 9 µm apart.7
It comes in two forms. Wafer-on-wafer bonds two whole wafers: alignment is easiest and pitch finest (400 nm in research), but every die on one wafer gets paired with whatever die lands on top of it, good or bad. Chip-on-wafer places individual dies, so each can be tested first and dies of different sizes can be mixed, at coarser pitch (2 µm in research).7 The surfaces must be extraordinarily flat and clean: a bowed wafer leaves whole regions unconnected.7
What it buys
UCIe’s 3D profile targets about 4 TB/s per square millimeter at 9 µm pitch, and under 0.05 pJ per bit, five times better than its interposer profile.5 AMD’s MI300 accelerator stacks its compute dies on I/O dies at the same 9 µm pitch and moves up to 17 TB/s vertically between them.8
What it costs
Heat. Stacking puts power-hungry circuits on top of each other. An Arm and UT Austin study folded a 7 nm CPU into two bonded tiers and found it ran up to 12 °C hotter than the flat version under a worst-case workload; putting logic on one tier and memory on the other halved the increase to about 6 °C.14 Heat and power delivery are covered further under Trade-offs.
Pitch scaling and what it does to the PHY
UCIe-3D is specified for bond pitches under 10 µm (and functional from 10 to 25 µm), with connection count scaling as over the entire chiplet area rather than its edges.5 Its targets by pitch:
| Bond pitch | Bandwidth density | Energy target |
|---|---|---|
| 9 µm | 4 TB/s/mm² | < 0.05 pJ/b |
| 5 µm | ≈ 12 TB/s/mm² | |
| 3 µm | ≈ 35 TB/s/mm² | < 0.02 pJ/b |
| 1 µm | ≈ 300 TB/s/mm² |
Source: UCIe 2.0 presentation.5 Each step follows : , , . The data rate is capped at 4 GT/s, chosen to equal the SoC’s logic frequency, with 80 lanes per cluster.5 With thousands of bonds per mm², bandwidth comes from width rather than speed. Running lanes near the core clock removes serializers, most of the clocking and the equalization a lateral link needs, which is where the energy saving comes from. The same presentation lists electrostatic-discharge protection falling from 30 V to 5 V CDM, and possibly to none for wafer-to-wafer bonds, which removes more capacitance from every lane.5
Bonding flows
Recessed Cu pads in oxide, oxide activation, face-to-face oxide bonding, then an anneal (typically around 300 °C) in which copper expansion closes the gap.7 Wafer-on-wafer reaches 400 nm pitch in research with near-perfect planarity from CMP; chip-on-wafer has reached about 2 µm because singulated dies can’t be planarized the same way, but it permits known-good-die selection.7 Wafer-to-wafer therefore suits stacks of high-yield, same-size dies; chip-to-wafer suits large or immature logic dies.
Thermal coupling
Stacking roughly doubles per footprint when both tiers are active, and heat from the lower tier must cross the upper tier and the bond layer to reach the heat sink. In Mathur et al.’s calibrated simulation of a two-tier 7 nm CPU, overlapping hotspots raised peak temperature by up to 12 °C over 2D; placing logic over memory, so hotspots don’t overlap, cut that to about 6 °C.14 That is why the 3D products in this chapter pair compute with cache or I/O dies rather than stacking compute on compute.
1. Prepare: Copper pads are built into each die’s surface, surrounded by silicon oxide and recessed very slightly. Both faces must be extraordinarily flat.
2D baseline: the whole CPU on one tier. Every hotspot has the heat sink directly above it.
AI chips need to read huge amounts of data every second. Normal memory chips on a circuit board can’t keep up. (high-bandwidth memory) fixes this with packaging.11
Several memory chips are stacked like pancakes, with wires running straight down through them. The stack sits right next to the processor, only a few millimeters away.11
Being so close lets each stack use over a thousand data wires at once. Moving data such a short way also takes much less energy.119 The biggest AI chips are ringed by eight of these stacks.910
What a stack is
An stack is several DRAM dies piled on a logic . Data travels across a DRAM die from its memory banks to a central strip, down to the base die, and a short way across the base die to the bumps that face the processor.11 The stack sits on the same interposer as the processor, beside it.
HBM3 runs up to 6.4 Gb/s per pin for 819 GB/s per stack, in stacks 4, 8 or 12 dies high, with 16 high planned.12 Each stack’s interface is 1,024 bits wide.9 HBM4 doubles that to 2,048 bits at up to 8 Gb/s, or up to 2 TB/s per stack, with stacks up to 16 high and 64 GB.13
Why it has to be close
Thousands of signals can only be wired between two dies over the fine wiring of silicon, and those thin wires lose signal quickly with length: HBM2’s interface could travel only about 5–7 mm across an interposer at 2 Gb/s.11 That is why HBM stacks ring the edges of the compute die, and why the compute die’s edge length limits how many stacks fit.
Energy
Reading a bit from HBM2 cost about 3.9 pJ, compared with about 14 pJ for GDDR5 memory on a circuit board. Most of that HBM energy is spent inside the stack; the interposer wires add only about 0.3 pJ/bit.11 Why AI chips need this much bandwidth in the first place is the subject of The memory wall.
The stack’s energy budget
O’Connor et al. break down an HBM2 access: about 1.21 pJ/bit of row activation, 2.24 pJ/bit to move data across the DRAM die, down the TSVs and across the base die (about 9.9 mm in all), and only 0.3 pJ/bit on the interposer I/O, for 3.92 pJ/bit including ECC.11 The package link is the cheap part. At that energy a 60 W DRAM power budget caps bandwidth near 2 TB/s, and 4 TB/s would cost over 120 W.11 The authors concluded that future high-bandwidth DRAM must cut this internal data-movement energy, not just add pins.11
Reach and floorplan
HBM’s I/O is unterminated, wide and slow per pin, which only works over a short, low-loss channel. On an interposer, a copper line 1.5 µm thick, 1 µm wide and 2.5 µm from its neighbors supports about 5 mm at 2 Gb/s; the HBM2 PHY could reach roughly 5–7 mm.11 The PHYs therefore sit on the compute die’s edge facing each stack, and HBM competes with die-to-die links and off-package I/O for the same . A larger compute die has more edge, and so room for more stacks.
Generations
| Width per stack | Rate per pin | Bandwidth per stack | Channels | Stack height | |
|---|---|---|---|---|---|
| HBM3 (JESD238, 2022) | 1,024 bits | up to 6.4 Gb/s | 819 GB/s | 16 | 4, 8, 12 (16 planned) |
| HBM4 (JESD270-4, 2025) | 2,048 bits | up to 8 Gb/s | up to 2 TB/s | 32 | 4 to 16 |
Sources: JEDEC announcements1213; HBM3 width from AMD’s bandwidth arithmetic (8,192 bits across eight stacks).9 Doubling width per stack doubles the microbumps and interposer routing per stack, so each generation tightens the packaging as well as the DRAM. Taller stacks push toward hybrid bonding between DRAM dies: production stacks are 8 to 12 dies high, and 16-high hybrid-bonded HBM has been demonstrated.7
The stack: DRAM dies on a base die, beside the processor on one interposer. Step a bit of data from a memory bank to the processor.
Chiplets from different companies need to agree on the details. Where does each connection go? How fast do you send? How do you catch mistakes? is an open set of rules for all of that. A group of chip and cloud companies started it in 2022.4
Think of it like USB, but for the inside of a package. If chiplets follow the same rules, ones from different makers can be mixed. The newest version covers stacked chips too.45
was founded in March 2022 to define a any company could build to.4 It has three layers, like a network:
- Physical layer. The electrical signaling and the exact bump map. Lanes are grouped into modules of 16 (standard package) or 64 (advanced package) data lanes each way, plus a forwarded clock and a “valid” signal. Rates run from 4 to 32 Gb/s per lane.4
- Die-to-die adapter. Packs data into fixed-size blocks called flits, adds an error check, and resends anything that arrives corrupted.4
- Protocol layer. What the data means. UCIe can carry PCIe and CXL, the standards used between chips on a server board, or any company’s own protocol.4 PCIe and CXL are explained in Board and server.
Compared with the high-speed serial links that cross a circuit board (), UCIe’s designers claim about 20 times the I/O performance at a twentieth of the power, because a link a few millimeters long needs far less signal conditioning.4
Bumps fail sometimes. The advanced-package profile includes spare lanes that can replace a broken one; the standard profile instead falls back to fewer lanes.4 Version 2.0 added the 3D profile for hybrid bonding and common features for managing, testing and debugging a package full of chiplets.5
Physical layer
A module has 16 (standard) or 64 (advanced) single-ended data lanes in each direction, plus one valid lane, a differential forwarded clock and a track lane for calibration; a link is one, two or four modules. Supported rates are 4, 8, 12, 16, 24 and 32 GT/s, and a device must support every rate up to its maximum. At 32 GT/s that is 64 GB/s per standard module and 256 GB/s per advanced module, each direction. An always-on sideband (two lanes each way at 800 MHz) handles training and management using otherwise depopulated bumps.4 The bump maps are fixed in the specification, with die rotation and mirroring allowed, so dies from different vendors line up.4
Adapter and protocols
The die-to-die adapter builds 68-byte or 256-byte flits with a 2-byte header and CRC covering each 128-byte payload chunk, and replays on CRC failure. Raw mode bypasses the adapter for protocols that bring their own.4 PCIe and CXL map on for plug-and-play attach; streaming mode carries a vendor’s on-die fabric protocol (AXI, CHI and others), so a CPU or switch can be built from identical chiplets without protocol conversion.4
Targets
Latency under 2 ns transmit plus receive including adapter and PHY; low-power entry and exit in 0.5–1 ns; and FIT far below 1 with flit-mode CRC and retry.4 Repair differs by profile: spare lanes in advanced packages, width degradation in standard ones. Version 1.1 added per-lane error counters for run-time link-health monitoring (for automotive use) and a 32-lane advanced option that cuts PHY area by about 40% for lower-cost packaging.4 Version 2.0 added UCIe-3D and a DFx architecture for discovery, test, debug and telemetry across all chiplets in a package, from die sort to field.5
Many shipping multi-die products still use proprietary links tuned to one package: AMD’s Infinity Fabric and NVIDIA’s NV-HBI are examples (see By the numbers). The standard’s value is the open chiplet market it enables, more than any one link’s performance.
Standard package: 16 data lanes each way. At 32 GT/s, 64 GB/s per module in each direction. Tap a layer.
Pick a way to join two chips: side by side (2D), on a silicon slab (2.5D) or stacked (3D). Then slide the gap between connections and watch the data per second climb. See how big the jump is when you switch to stacking.
Choose a packaging technology with the 2D / 2.5D / 3D buttons. The top drawing is a schematic side view; the bottom one is the bump field drawn to scale (gold bumps carry data, gray ones power and ground). Things to try:
- In 2.5D, change the pitch from 45 µm to 25 µm. By what factor do the connections per mm grow?
- Double the data rate instead. The sim holds energy per bit fixed, so power follows bandwidth either way. In real links, why does a faster lane usually cost more energy per bit?
- Switch to 3D. Why is its density quoted per mm² instead of per mm of edge?
The link model is for side-by-side links (signal share , bump-field depth , shoreline , pitch ) and for a bonded overlap . Bandwidth is summed over both directions, and power is bandwidth × energy per bit. The signal shares (0.38, 0.68, 0.65) are calibrated so the defaults land near UCIe’s published densities. The lower panel compares silicon cost per good system for a monolithic die and a split into chiplets, with negative binomial yield. Try:
- In 2.5D, set 32 Gb/s and 45 µm and compare the shoreline density with UCIe’s 1,317 GB/s/mm. Then shrink the depth: what does a shallower PHY cost you?
- Find the logic area at which four chiplets stop beating a monolithic die at . How does the crossover move with and with die-to-die overhead?
- Compare tested and untested chiplets at . Why does the gap explode?
It leaves out signal integrity (reach is shown, not enforced), package and interposer cost, bonding yield per connection, test escapes and NRE. Use it for scaling, not for quotes.
- Largest die one exposure can print
- 26 × 33 mm
- Bump pitch: organic / interposer (UCIe)
- 100–130 / 25–55 µm
- Hybrid-bond pitch in products
- ≈ 9 µm
- HBM4 bandwidth per stack
- up to 2 TB/s
Sources: IEEE Spectrum, UCIe 1.1 and JEDEC’s HBM4 announcement.24713
What these numbers mean:
- 26 × 33 mm is about the size of a large postage stamp. No single chip can be bigger than that.
- 100 to 130 µm is the gap between connections on an ordinary package board. A µm is a thousandth of a millimeter, so that’s about the width of a hair. A silicon slab cuts the gap to 25–55 µm.
- 9 µm is the gap when stacked chips are fused copper to copper. That’s over ten times closer than on a package board, so over a hundred times as many connections fit.
- 2 TB/s means 2 trillion bytes a second from one memory stack. That’s like reading hundreds of HD movies every second.
| Organic substrate (2D) | Interposer / bridge (2.5D) | Hybrid bonding (3D) | |
|---|---|---|---|
| Bump or bond pitch | 100–130 µm | 25–55 µm | < 10 µm |
| Lane rate | 4–32 GT/s | 4–32 GT/s | up to 4 GT/s |
| Reach | ≤ 25 mm | ≤ 2 mm | vertical |
| Bandwidth per mm of edge | 28–224 GB/s | 165–1,317 GB/s | n/a |
| Bandwidth per mm² | 22–125 GB/s | 188–1,350 GB/s | ≈ 4,000 GB/s at 9 µm |
| Energy target | 0.5 pJ/b | 0.25 pJ/b | < 0.05 pJ/b at 9 µm |
UCIe’s published targets for its three profiles; ranges span the supported data rates.45
Two cross-checks. MI300X’s HBM figure is plain arithmetic: .9 And NV-HBI’s 10 TB/s across one die boundary is in the range the simulation’s 2.5D model gives for a die edge of about 8–15 mm at 16–32 Gb/s and 45 µm pitch. NVIDIA doesn’t publish NV-HBI’s pitch or lane count, so that is a consistency check, not a reverse-engineering.
- Close or far. Tightly packed connections only work over short distances. A package board can link chips a couple of centimeters apart. A silicon slab reaches only a couple of millimeters.4
- Heat. Stacked chips trap heat. So designers often stack a hot chip with a cooler memory chip, not two hot chips together.14
- One bad piece ruins it all. If any chip in the package is broken, the whole package fails. So every piece is tested before it’s joined.1
- Cost. Fancy packaging is expensive. For a small chip, one plain chip is often cheaper.3
Bandwidth against reach and cost
Organic substrates are cheap and reach far but give the fewest connections per millimeter; interposers and bridges give about six times more at half the energy but only over a couple of millimeters; hybrid bonding gives far more again but only vertically.45 Advanced packaging also adds its own yield losses: the interposer itself must be good, and every bonding step can fail.3
Known-good die
A package works only if every die in it works, so dies are tested at wafer sort and only are assembled.1 When assembly fails, the good dies inside are lost too, and with expensive dies that waste becomes a large part of the cost.3 Testing each die alone, then each partial stack, is covered in After tapeout.
Heat
Stacking raises the power per unit of footprint and makes the lower die’s heat cross the upper one. Careful partitioning (memory under logic, so hot spots don’t overlap) can halve the temperature penalty.14 Cooling the whole package is the subject of Power and cooling.
Power delivery
Every die needs current from below. With a full interposer, all power passes through it; bridges leave power on short paths through the substrate.6 A package designed to accept many different chiplets must be provisioned for the worst combination, so it can end up over-built for most.1
Design and volume
Every distinct die needs its own masks and sign-off. For one product made in modest volume, a single die is often cheaper overall; chiplets pay back through reuse across many products.3
Where designs go wrong
- Over-splitting. Savings from finer granularity are marginal past two or three chiplets while D2D area and packaging overhead keep growing.3
- Advanced packaging on a mature node. At 14 nm-class defect densities, Feng and Ma find D2D and packaging overhead above 25% for organic multi-chip modules and above 50% for 2.5D, wiping out most of the yield gain; advanced packaging pays mainly with advanced logic.3
- Packaging yield and wasted known-good dies. With a large monolithic interposer, interposer yield and bonding defects scrap good dies; at 7 nm and 900 mm² in 2.5D, packaging cost was comparable to chip cost in their model.3
- Shoreline exhaustion. HBM PHYs, die-to-die links and off-package SerDes all compete for the same edge, and edge grows only as .411
- Generality against optimization. A common chiplet interface sized for the most demanding die makes every other die pay its area; a package provisioned for any chiplet combination is over-designed for di/dt, electromigration and decoupling on most of them.1
- Stacking hot on hot. Overlapping hotspots cost up to 12 °C in a two-tier CPU study; thermally aware partitioning halved it.14
- Hybrid-bond process windows. Bonding needs near-perfect flatness and cleanliness; a warped wafer leaves regions unconnected, and wafer-on-wafer pairs dies regardless of whether they are good.7
- Coherence and security. Mixing chiplets from several sources raises open questions about coherence protocols, memory consistency, denial-of-service between chiplets and side channels on the shared fabric.1
5 mm on an organic substrate, within its 25 mm reach: up to 224 GB/s per mm of edge at 0.5 pJ/bit.
40 fresh dies, 6 of them secretly bad. Nothing has been tested.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Three small models: connection count and bandwidth density from bump geometry, link power, and the silicon cost of a split versus a monolithic die. These are the equations behind the simulation.
1. Counting connections
For a square bump array of pitch (in mm) with a share of bumps carrying data:
A side-by-side link uses a strip of depth behind the shoreline ; a stacked link uses the whole overlap area :
Bandwidth, counting both directions, is in GB/s for a lane rate in Gb/s. Worked at the simulation’s 2.5D defaults (, , , , ):
- bumps/mm²; signals/mm².
- signals per mm of edge; signals.
- , or about 670 GB/s per mm of shoreline.
The calibration check is UCIe’s own bump map: a 64-lane module (128 data lanes counting both directions) in 388.8 µm × 1,043 µm at 45 µm pitch carries 512 GB/s at 32 GT/s, which is 1,317 GB/s per mm of shoreline.4 The model gives 1,343 GB/s/mm at the same settings. For 3D, at 9 µm and 4 GT/s gives , UCIe-3D’s stated target.5
Two consequences fall straight out. Edge links scale as , so halving pitch at fixed depth quadruples bandwidth; at fixed bump count it instead shrinks the PHY’s depth, which is how UCIe’s 388.8 µm module goes from 1,585 µm deep at 55 µm pitch to 388 µm at 25 µm.4 And 3D links scale with area, so a stacked die’s die-to-die bandwidth grows with its footprint while a side-by-side die’s grows only with its perimeter.
x64 module at 45 µm: 388.8 × 1,043 µm (10 columns), 512 GB/s at 32 GT/s → 1,317 GB/s per mm of shoreline at every pitch. PHY area 0.41 mm²: finer pitch buys a shallower PHY, not a wider link.
2. Power at bandwidth
One terabit per second at one picojoule per bit is one watt. The 5.4 TB/s link above at 0.25 pJ/b: . NV-HBI-class bandwidth of 10 TB/s at the same energy would be 20 W, under 3% of a 750 W accelerator such as MI300X.9 UCIe claims 20× the I/O performance at a twentieth of the power of off-package SerDes, which is why links this wide stay inside the package.4 The 3D numbers show why stacking wins on energy: 4 TB/s through 1 mm² at under 0.05 pJ/b is under 1.6 W.5
3. Silicon cost of a split
The simulation’s lower panel uses the standard textbook model:
- Dies per wafer for area on a wafer of diameter :15
- Negative binomial yield with defect density and cluster parameter :153
- Cost per good die:15
- Split into chiplets with die-to-die overhead : each chiplet has area
- With known-good-die test and a bond yield per chiplet:
- Without test, every package needs good dies:
Worked at the simulation defaults (, , , , , a $10,000 illustrative 300 mm wafer):
- Monolithic: , , so about $318 per good die.
- Chiplets: , , , so about $45 per good chiplet; per system.
- Untested: of packages work, so about $353, worse than monolithic.
So the tested split costs about 0.59× the monolithic silicon. Push the area down to 100 mm² and the ratio rises above 1: the overhead outweighs the yield gain. This is the curve behind Feng and Ma’s turning point.3
What the simple model leaves out, from Feng and Ma’s fuller version: the raw package and substrate cost; for 2.5D the interposer’s own cost and yield ; chip-to-interposer bond yield per chip and interposer-to-substrate bond yield ; and the good dies scrapped when packaging fails. Their packaging term is
and the last term, wasted known-good dies, is why chip-last assembly (build the interposer or redistribution layers first, then bond the dies) is preferred for expensive dies.3 NRE adds a per-die fixed cost (masks, sign-off) plus per-area terms, amortized over volume.3
4. Why stacking heats up
To first order, temperature rise is power times thermal resistance to the heat sink. Two tiers in one footprint double the power through the same sink area, and the lower tier’s heat adds the resistance of the upper die and the bond layer to its path. Overlapping hotspots compound both. Mathur et al. quantify it for a calibrated 7 nm two-tier CPU: up to 12 °C above the 2D design, about 6 °C with logic over memory.14 Every degree of headroom lost there comes out of clock frequency or sustained workload before throttling.
Q1You halve the bump pitch of a die-to-die link and keep everything else the same. Roughly what happens to the number of connections in the same area?
Q2A die-to-die link moves 4 TB/s at 0.25 pJ/bit. About how much power does the link use?
Q3Why is a silicon interposer able to wire dies together so much more densely than an organic package substrate?
Q4Why do designers often build the I/O part of a chiplet product on an older, cheaper process?
Sources
Show Hide 15 sources
- Understanding Chiplets Today to Anticipate Future Integration Opportunities and LimitsDie sizes approaching the reticle limit; reticle-scale dies are costly because of defects; per-chiplet overheads (clocking, power management, test, debug, interfaces) mean 2X functionality costs more than 2X area; known-good die; first-gen EPYC: four 14 nm chiplets over SerDes on the package substrate without an interposer, total silicon above the reticle limit, estimated 41% cheaper than a monolithic equivalent; second-gen EPYC: 7 nm core dies plus a 12 nm I/O die because analog PHYs scale poorly; one I/O die and up to eight core dies give 16–64-core parts; generality vs optimization, power provisioning, coherence and security challenges.
- This Machine Could Keep Moore’s Law on TrackToday’s scanner exposure field is 26 by 33 mm; high-NA EUV halves it to 26 by 16.5 mm.
- Chiplet Actuary: A Quantitative Cost Model and Multi-Chiplet Architecture ExplorationOrganic substrate line/space above 10 µm, fan-out RDL above 2 µm, silicon interposer above 0.4 µm; negative binomial yield (1 + DS/c)^−c; 10% die-to-die area overhead after EPYC; packaging cost terms including interposer yield, bonding yields and wasted known-good dies; up to 50% die-cost saving but packaging up to 30% of cost; split pays when die-defect cost exceeds packaging cost; two or three chiplets usually enough; per-chip NRE such as masks favors monolithic at low volume; advanced packaging pays only at advanced nodes.
- The UCIe 1.1 Specification: Future Applications of ChipletsMotivation (beyond reticle, mixed processes, smaller dies yield better; 20× I/O performance at 1/20th power vs off-package SerDes); layers (PHY, die-to-die adapter, protocol); 16-lane standard and 64-lane advanced modules, 4–32 GT/s, spare lanes vs width degradation; 68 B and 256 B flits with 2 B header and CRC; x64 bump maps with 388.8 µm shoreline and ~388/1043/1585 µm depth at 25/45/55 µm; KPI table: 100–130 vs 25–55 µm bump pitch, ≤25 vs ≤2 mm reach, 28–224 vs 165–1317 GB/s/mm, 22–125 vs 188–1350 GB/s/mm², 0.5 vs 0.25 pJ/b, under 2 ns latency; x32 option; founded March 2022.
- UCIe 2.0 Specification: Advancing an Open Ecosystem for On-Package Chiplet InnovationUCIe-3D: up to 4 GT/s (at SoC logic frequency), width 80, optimized for bond pitch under 10 µm (10–25 µm functional), connections scale inversely with pitch squared over the whole chiplet area; 4 TB/s/mm² at 9 µm, ~12 at 5 µm, ~35 at 3 µm, ~300 at 1 µm; under 0.05 pJ/b at 9 µm and under 0.02 at 3 µm; ESD dropping from 30 V to 5 V CDM, none possible for wafer-to-wafer; manageability and DFx.
- Embedded Multi-die Interconnect Bridge (EMIB) Technology BriefSmall silicon bridges with multiple routing layers embedded in a cavity in the organic substrate; tight microbump pitch only at the bridge; a full interposer requires all signal and power vias to pass through it; in volume production since 2017; TSVs added in one variant for vertical power to HBM.
- Hybrid Bonding Plays Starring Role in 3D ChipsSolder microbumps at tens of µm; hybrid bonds about 9 µm apart in products, 2 µm chip-on-wafer and 400 nm wafer-on-wafer in research; recessed copper pads in oxide, bonded oxide, anneal expands copper; wafers must be nearly perfectly flat; chip-on-wafer lets dies be tested first; HBM stacks 8 to 12 dies high, hybrid bonding for 16 layers and more.
- AMD’s Next GPU Is a 3D-Integrated SuperchipMI300 uses TSMC SoIC (hybrid bonding, no solder) to stack compute chiplets on I/O dies and CoWoS (a silicon interposer) to reach eight HBM stacks; compute in N5, I/O and cache in N6; 9 µm vertical pitch reused from V-cache; up to 17 TB/s vertically; smaller chips yield better.
- Introducing AMD CDNA 3 Architecture (white paper)Up to 8 vertically stacked accelerator dies (5 nm) on 4 I/O dies (6 nm), 8 HBM stacks; 2 of 40 compute units per die disabled for yield; 1024 bits × 8 stacks at 5.2 Gb/s = 5.3 TB/s for MI300X; previous generation joined two dies over an interposer bridge; MI300X rated at 750 W.
- Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory EraTwo reticle-sized dies joined by NV-HBI at 10 TB/s, programmed as one GPU; 208 billion transistors; TSMC 4NP; eight 12-high HBM3E stacks on 16 × 512-bit controllers, 288 GB, 8 TB/s.
- Fine-Grained DRAM: Energy-Efficient DRAM for Extreme Bandwidth SystemsHBM2 about 3.9 pJ/bit vs GDDR5 about 14 pJ/bit; inside the stack, data moves over the DRAM die, down TSVs and across the base die (2.24 pJ/bit), the interposer wires add only 0.3 pJ/bit; 4 TB/s would need 120 W of DRAM power at HBM2 energy; HBM2 I/O reaches roughly 5–7 mm on an interposer at 2 Gb/s.
- JEDEC Publishes HBM3 Update to High Bandwidth Memory (HBM) StandardJESD238 HBM3, announced January 27, 2022: up to 6.4 Gb/s per pin, 819 GB/s per device, 16 independent channels (up from 8), 4-, 8- and 12-high TSV stacks with provision for 16-high. May need a regular browser; the site blocks some automated clients.
- JEDEC and Industry Leaders Collaborate to Release JESD270-4 HBM4 Standard: Advancing Bandwidth, Efficiency, and Capacity for AI and HPCJESD270-4 HBM4, announced April 16, 2025: up to 8 Gb/s across a 2048-bit interface, up to 2 TB/s per stack, 32 channels (up from 16), 4- to 16-high stacks, up to 64 GB. May need a regular browser; the site blocks some automated clients.
- Thermal Analysis of a 3D Stacked High-Performance Commercial Microprocessor using Face-to-Face Wafer Bonding TechnologySub-10 µm face-to-face wafer bonding; overlapping hotspots raise power density; a 7 nm CPU folded into two tiers runs up to 12 °C hotter than 2D, logic-over-memory partitioning halves that to about 6 °C.
- Cost (EEC 116 lecture handout)Dies per wafer = π(d/2)²/A − πd/√(2A); die yield (1 + D·A/α)^−α with α ≈ 3; die cost = wafer cost ÷ (dies per wafer × die yield).