Extras · Industries · AI accelerators and datacenter

AI accelerators and datacenter

Dies as large as lithography can print, stitched together in the package and fed by stacked memory, all to keep matrix math busy.

Google TPU v1 matrix unit
65,536 8-bit MACs
TPU v1, start to deployment
15 months
HBM4 interface per stack
2,048 bits, up to 2 TB/s
Wafer-scale WSE-3 vs H100 die
46,225 mm² vs 814 mm²

At a glance

Where the flow bends

  1. 02Architecture

    Planners start from the math AI needs and how fast memory can feed it. They also decide early whether to split the design across several chips in one package.

    Before any circuit is written, architects run real neural networks through a software model of the proposed chip. The model tells them how much math hardware, on-chip memory and memory bandwidth to build, and whether to split the design into several smaller chips that share one package.

    Performance models of real workloads set the balance of multiply-accumulate units, on-chip SRAM and HBM stacks. If the design is split into chiplets, the bandwidth, latency and standard (UCIe or proprietary) of the links between dies are chosen here, and they are expensive to change once the package and interposer are designed.

  2. 04Verification

    Testing a design this big on ordinary computers is far too slow. So teams load it onto special machines that can be rewired to act like the chip. Real AI programs run on it months before the chip exists.

    Simulating the chip’s design in software is far too slow to run AI programs on it. Teams also load the design onto special reprogrammable hardware (emulators and FPGAs) that runs it much faster, so the compiler and other software can be tested against it before the chip is built.

    Emulator and FPGA-prototype capacity becomes a schedule resource, booked like any other. Multi-die designs also need the links between dies verified at package level, and the performance the RTL actually delivers is checked against the architecture model alongside the usual functional coverage.

  3. 06Design for test

    A chip this big will almost always have a few flaws. Designers add spare parts, plus tests that find the bad pieces and swap in the spares.

    With this much silicon, some part of almost every chip will be defective. Designers add spare cores, spare wires and repairable memory, plus test circuits that find the faults. Each chip is tested before packaging, so one bad chip doesn’t ruin a package full of good ones.

    Known-good-die test at wafer sort, lane repair on die-to-die links, built-in self-repair for SRAM, and harvesting (selling dies with a few units disabled) all drive DFT. A defect found after assembly scraps the interposer and the HBM stacks attached to it as well.

  4. 07Floorplanning

    The chip is as big as the factory can print. Its edges are crowded with links to memory and to other chips.

    The chip is close to the largest size the factory can print. Much of its layout is decided by its edges: the circuits that talk to the stacked memory and to neighboring chips must sit on particular edges, lined up with the wiring in the package beneath.

    Dies sit near the reticle limit and the floorplan is set by shoreline. HBM PHYs and die-to-die interfaces need die edge, and their bump maps are co-designed with the interposer. Bandwidth per millimeter of edge bounds how much shoreline each interface takes, so the I/O ring often fixes the die outline.

  5. 08Power planning

    These chips use huge amounts of power. Getting it to every part of the chip, and then getting the heat out, are among the hardest jobs.

    These chips draw very large currents at low voltage. Getting that current to every transistor without the voltage sagging takes thick power wiring that crowds out the wires carrying signals. Some new manufacturing processes move the power wiring to the back of the chip.

    Thousands of multiply units switching on the same clock edge cause sudden current steps, so dynamic IR drop, resonance in the package’s power network and heat density dominate. Backside power delivery frees front-side routing tracks for signals but changes the heat path and how failed chips are debugged.

  6. 12Signoff

    The final checks cover the whole package, because several chips and memory towers must work as one.

    The final checks of timing, power and signal quality cover the whole package, not just one chip: the links between chips, the silicon wiring layer they sit on, and how heat spreads through the stack.

    Static timing analysis spans dies, with a timing budget for each interface; signal and power integrity are analyzed at package level; IR-drop and electromigration checks use temperature maps; and physical verification covers the interposer and any hybrid-bonded interfaces.

  7. 13GDS & tapeout

    These chips use the newest and most costly way of making chips. A mistake can cost millions of dollars and months of time.

    These chips use the newest and most expensive manufacturing processes, and AI moves fast. Teams send designs to the factory on tight schedules and reuse finished chiplets across several products to spread the cost.

    Mask cost and wafer price at leading nodes push I/O and base dies onto older, cheaper processes so they can be reused across products. Tapeout also starts the package assembly flow, with its own lead time and test steps.

Today’s AI runs mostly on one kind of math. It multiplies huge grids of numbers, over and over. An AI chip is built to do just that one job, thousands of times at once.

Picture a kitchen with thousands of cooks. Hiring more cooks is easy. The hard part is getting food to all of them fast enough. AI chip designers spend most of their effort on that kind of delivery: bringing in numbers and power, and taking away heat.

These chips are often as big as a factory can print in one piece, about the size of a postage stamp. That ceiling is called the . To go bigger, designers put several chips and towers of memory side by side in one package, working as one.

Big tech companies now design their own AI chips. Google did a quick sum: if people used voice search for three minutes a day, it would need twice as many computers. Using ordinary chips, that cost too much.

A , the kind of program behind modern AI, spends almost all its time on one operation: matrix multiplication, which combines two large grids of numbers by multiplying and adding them pairwise. The small step repeated inside it is the : multiply two numbers and add the result to a running total. An AI accelerator is a chip that spends most of its area on MAC units and on the fast memory that keeps them supplied. Google’s first TPU had 65,536 of them working in parallel.

How those MAC units are arranged, and why feeding them is the hard part, is the subject of the AI accelerator chapters in the Architectures guide. This page is about something else: how building such a chip changes the design flow, the sequence of steps that turns a written specification into the file a factory manufactures from (the Design Flow guide walks through them one by one).

Four limits shape every step:

  • Size. A chip is printed onto a silicon wafer by a machine that can only expose a limited area at once, roughly 800 mm², or about the size of a postage stamp. One piece of silicon, a die, can’t be bigger than that. NVIDIA’s H100 GPU die is 814 mm², right at the limit.
  • Memory bandwidth. The model’s numbers don’t fit on the chip, so they live in separate memory chips. How many bytes per second can flow between memory and chip often decides real speed. The usual answer is , memory chips stacked into a tower beside the processor.
  • Power and heat. These chips draw very large currents at low voltage. Delivering that current without the voltage sagging, and carrying the heat away, get harder each generation.
  • Yield and cost. Every wafer has a few random defects. The bigger the die, the more likely it catches one, and the newest manufacturing processes are the most expensive.

An accelerator balances four things under a hard area ceiling: throughput, on-chip , off-chip bandwidth and the power budget. The workload side of that balance (arithmetic intensity, rooflines, why decode is memory-bound) is covered in What the workload needs and The memory wall. Here we follow what that balance does to each stage of the implementation flow.

Yield is the hidden fifth constraint, and it is worth working through once. Cost per good die is the wafer cost divided by dies per wafer times the fraction of dies that work (the yield):

cost per good die=wafer costdies per wafer×yield\text{cost per good die} = \frac{\text{wafer cost}}{\text{dies per wafer} \times \text{yield}}

Both terms in the denominator fall as die area grows: fewer rectangles fit on a round wafer, and each one is more likely to contain a random defect. In a textbook example with a 30 cm wafer and a fixed defect density:

Die areaDies per waferYieldCost per good die
60 mm²1,09284.8%1× (baseline)
120 mm²52872.1%2.43×
240 mm²25152.5%7.0×

Quadrupling the area raised cost per good die about sevenfold. An accelerator near the sits far out on that curve.

The industry’s answers are architectural and physical at once. Split the system into on a or in a 3D stack; put HBM beside the compute die; add redundancy so a defect costs a small fraction of the die; and move power delivery to the wafer backside. Each answer adds work that a single-die design never needed: designing the package alongside the chip, testing each die before assembly, closing timing across the links between dies, and signing off heat and power for the whole package. The packaging technology itself is covered in the Systems guide’s Packaging and chiplets chapter.

wafer · dots = defectsdies per wafer1,092yield84.8%cost per good die1×relative to 60 mm²
Die area

60 mm² die: 1,092 dies per wafer, 84.8% of them work, so each good die costs 1× the 60 mm² one.

A textbook example: 30 cm wafer, the same defect density for every size. Wafer drawing not to scale; the numbers are the example’s.Share freely with credit: ‘Figure from chipfieldguide.com’
  • Size limits. One chip can only be so big. So big AI systems are built from several chips packed close together.
  • Memory hunger. AI models hold billions of numbers. To feed the chip fast enough, memory chips are stacked like pancakes and set right beside it.
  • Power and heat. One of these chips can use about as much power as a microwave oven. All of that ends up as heat, in a space smaller than your palm.
  • Speed. AI changes fast. A chip that shows up a year late may be built for last year’s AI.

Memory. The main memory of a computer is , dense but relatively slow memory on separate chips. stacks several DRAM chips into a tower, connects them with metal-filled holes drilled through the silicon, and places the tower beside the processor on a shared base. The latest standard, HBM4, moves data at up to 8 billion bits per second on each of 2,048 wires, for up to 2 terabytes per second from one stack. An accelerator surrounds itself with several stacks, and each stack needs its own (the circuit that drives and receives those wires) on the edge of the processor die. Those PHYs take up most of the die’s edge, which matters later in the flow.

Packaging. The package is what a die is mounted in so it can connect to a circuit board. For AI chips it has become part of the design. In 2.5D packaging, dies sit side by side on a , a slice of silicon used only for wiring. In 3D packaging they are stacked. Stacked dies have traditionally been joined by tiny solder balls with pitches (center-to-center spacing) in the tens of micrometers. fuses copper pads directly, and researchers have made such connections every 400 nanometers. The standard defines how chiplets talk across a package, and its backers present chiplets as a way to build systems larger than one die can be.

Time to market. AI models change quickly, so a late chip may be tuned for last year’s workloads. Google designed, verified, built and deployed its first TPU in 15 months.

Power delivery. On a conventional chip, the power grid and the signal wires share the same stack of metal layers above the transistors, so every track the grid uses is a track routing can’t. imec reports that power interconnect takes at least 20% of routing resources. Its simulations of with buried power rails, which moves the grid under the transistors, cut by about 7×. Intel’s PowerVia test chip showed over 6% higher frequency and 30% less power lost in delivery. Intel also had to write new design rules to keep heat in check and invent new debug methods, because the transistors now sit between two stacks of wiring. At the extreme, a wafer-scale part can’t be fed from its edges at all: Cerebras delivers power vertically through more than 300 voltage regulator modules spread across the wafer.

Design for yield. If a defect lands in a block, that whole block has to be switched off, so the size of the smallest replaceable unit decides how much a defect costs. In Cerebras’s WSE-3, a single 46,225 mm² wafer-scale processor, each core is about 0.05 mm², against about 6 mm² for one streaming multiprocessor (the repeated compute block) of an 814 mm² H100. A reconfigurable on-chip network routes around bad cores. The smaller the repair unit, the less silicon each defect wastes, which is what makes a die far past the reticle limit viable. The wafer-scale chapter covers the architecture side.

Tapeout cost. Treat headline figures with care. A widely quoted 2018 estimate put a 5 nm design at $542.2 million; Semiconductor Engineering argues such numbers are inflated and fall as a process matures. Masks, licensed IP, verification and software are all real costs, and they push teams toward reusable chiplets and fewer respins (repeat tapeouts to fix bugs).

heat outinterposerHBMHBMHBMHBMcompute diereticle limit ≈ 800 mm²power in
Compute

An accelerator package from above. Tap or hover any part.

An accelerator package seen from above, not to scale. Split the compute into chiplets and run it to see data and power flow.Share freely with credit: ‘Figure from chipfieldguide.com’

An AI chip adds five extra jobs to the usual steps. Engineers plan with real AI programs, copy one small tile many times, test on stand-in machines, plan the package, and check each chip before it goes in. Step through them in the picture below.

Architecture and design code. Teams first write a fast software model of the chip and run real neural networks through it, to choose sizes before committing to circuits. The circuit is then described in a hardware description language as (register-transfer level code, which says what values are stored and how they change on each tick of the clock). For an accelerator much of that code is generated: one processing element is written once and repeated into a large grid, often a that passes numbers between neighbors so each value read from memory is used many times. How systolic arrays work is covered in their own chapter.

Verification. Verification means checking the design does what it should before it is built. Normally that is done by simulating the RTL in software, but software simulation of a large chip is far too slow to run a compiler and a full AI model on it. So teams add and prototyping: they load the design onto reprogrammable hardware that runs it much faster. The open-source FireSim project, for example, ran an exact model of a 1,024-computer cluster on rented cloud FPGAs at an effective 3.4 MHz: less than 1,000 times slower than real hardware, and several orders of magnitude faster than software simulators.

Physical design. Physical design turns the checked design into a layout: where each of millions of logic gates goes, and how the wires between them run. For an accelerator it is done in pieces: one tile is laid out and finished, then copied across the die. The outline of the chip is set by the memory and chip-to-chip connections on its edges, and delivering power is a major task in its own right.

Testing and tapeout. Design-for-test (DFT) adds circuits that let a factory tester check each manufactured chip for defects. In a multi-chip package every die must be a , tested before assembly. Spare lanes, cores and memory rows are added so test can switch around defects. Tapeout is the moment the finished layout is sent to the factory.

Floorplans built around memory blocks. A floorplan is the first rough layout: the die outline and where the big blocks go. In TPU v1 the 24 MiB unified buffer (the main on-chip memory) took almost a third of the die and the matrix unit a quarter, while control logic was just 2%. The buffer’s size was chosen partly to match the matrix unit’s pitch on the die. Expect floorplans in which the placement of SRAM and datapath is decided by hand, and the automatic tools mostly fill in the logic around them. That is the opposite of a control-heavy CPU, where most of the area is irregular logic the placer arranges.

Closing the links between dies. When a design spans several dies, each die is implemented and checked on its own, so the link between them needs an agreed timing budget: how much of the clock period each side, and the package wiring between them, may use. The link also gets its own signal- and power-integrity analysis and a way to repair failed lanes. UCIe’s advanced-package profile includes spare lanes and targets 165 to 1,317 GB/s per millimeter of die edge, depending on data rate. That number tells the floorplanner how much each link will consume. For example, at the low end, 1 TB/s between two dies would take about 6 mm of edge on each.

Schedule versus PPA. PPA (power, performance, area) is how a design’s quality is scored. Time-to-market pressure leaves some of it on the table. The TPU v1 authors estimate that more aggressive logic synthesis and block design could have raised the clock by 50% had they had more than 15 months. Teams choose which PPA to give up to hit the market window.

Signoff for the whole package. Signoff is the final set of checks before tapeout. For a multi-die product, heat, warping and power integrity are analyzed for the full stack of dies, interposer, HBM and package substrate, not die by die. With backside power, the methods used to find and analyze defects in failed chips must change as well.

ModelTileEmulatePackageKGDneural networkperformance modelMACsSRAMbandwidthsizeschosen
1 / 5

A fast software model runs real neural networks, to size MAC units, on-chip memory and bandwidth before any RTL exists.

What an accelerator adds to the design flow, step by step. Drawings are sketches, not layouts.Share freely with credit: ‘Figure from chipfieldguide.com’

Around 2013, Google saw that AI would soon need far more computing than it had. So it rushed to build its own chip, the Tensor Processing Unit, or TPU.

The team kept the design simple so it could move fast. Most of the chip was one giant grid of small multipliers, plus a big block of memory.

It went from start to work in Google’s datacenters in just 15 months. On Google’s AI jobs, it ran about 15 to 30 times faster than other chips of its day.

Google’s paper on its first TPU describes a chip designed for speed of delivery as much as for speed of computation. Several of its choices were made to cut schedule risk, not to maximize performance.

  • An add-in card, not a new computer. To reduce the chance of delaying deployment, the TPU plugged into existing servers over PCIe, the standard slot used by graphics cards. The server’s own processor sent it instructions, so the TPU never had to fetch its own, which kept its control logic small and simpler to check.
  • One big, regular block. A single grid of 256 × 256 MAC units, fed by 24 MiB of on-chip memory. A regular structure is easier to lay out and verify than many different blocks.
  • Result. About 15–30× faster than the graphics chips and processors of the time on Google’s AI workloads, and 30–80× more operations per watt.

Why the grid passes numbers between neighbors, and how well it was kept busy, is covered in the systolic arrays chapter.

Two lessons from the paper bear on the design flow. First, memory bandwidth, not the matrix unit, set delivered performance: four of the six production workloads were limited by memory bandwidth. The authors estimate that using a GPU’s GDDR5 memory in place of the DDR3 weight memory the TPU shipped with would have tripled achieved throughput and raised operations per watt to nearly 70× the GPU’s. That is a decision made at the architecture stage, before any RTL, and it is why later accelerators moved to HBM. The memory wall chapter explains the mechanism.

Second, risk reduction shaped the architecture. A coprocessor that receives complex, multi-cycle instructions from the host, a single large matrix unit and minimal control logic kept verification tractable on a 15-month schedule, and the authors note that a longer schedule could have bought a higher clock. Simplifying the design to fit the verification budget is a choice every accelerator team makes in some form.

Later generations moved toward system scale. TPU v4 adds SparseCores, small dataflow processors that speed up embedding-heavy models 5–7× while using about 5% of die area and power, and it reaches 2.1× the performance and 2.7× the performance per watt of TPU v3. The design target shifted from one chip to a 4,096-chip machine, which moves many of the hard problems out of the chip flow and into the system (see the Systems guide).

hostCPUPCIeTPU v1unified buffer24 MiB · ≈ ⅓256×256 · ≈ ¼2%otherDDR3weightsthroughput1×
Weight memory

As shipped, with DDR3: four of six workloads were limited by memory bandwidth. Tap any part.

TPU v1 block diagram. Block sizes follow the paper’s area shares; positions are a sketch, not the real floorplan.Share freely with credit: ‘Figure from chipfieldguide.com’

Sources

Show Hide 12 sources
  1. In-Datacenter Performance Analysis of a Tensor Processing UnitNorman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, et al. · arXiv (ISCA 2017) · 2017TPU v1: 65,536 MACs, 92 TOPS, 700 MHz, 28 MiB on-chip memory; the 2013 voice-search projection; 15-month schedule; PCIe coprocessor with host-sent CISC instructions; systolic execution to cut SRAM reads; die floorplan (unified buffer almost a third, matrix unit a quarter, control 2%); memory-bound apps and the GDDR5 what-if; 50% clock from more aggressive synthesis.
  2. A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial IntelligenceYudhishthira Kundu, Manroop Kaur, Tripty Wig, et al. · arXiv · 2025Reticle limit about 815 mm²; H100 die 814 mm² vs WSE-3 46,225 mm²; WSE-3 core 0.05 mm² vs H100 SM 6 mm²; reconfigurable fabric; over 300 VRMs delivering power vertically.
  3. IC Manufacturing, Cost, Power, and Dependability (COE 501 lecture slides)Muhamed Mudawar · King Fahd University of Petroleum and MineralsDie yield and die cost formulas; worked example on a 30 cm wafer: 60, 120 and 240 mm² dies give 1,092, 528 and 251 candidates, 84.8%, 72.1% and 52.5% yield, and cost per good die rising about 7.0× from smallest to largest (ratios recomputed from the slide’s own inputs; the slide’s 962 good dies for 60 mm² is a multiplication slip for 926, which gave its 2.52× and 7.28×).
  4. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsNorman P. Jouppi, George Kurian, Sheng Li, et al. · arXiv (ISCA 2023) · 20234,096-chip machines; optical circuit switches under 5% of system cost and under 3% of power; SparseCores 5–7× faster on embeddings using 5% of die area and power; 2.1× the performance and 2.7× the performance per watt of TPU v3.
  5. High Bandwidth MemoryWikipediaHBM stacks DRAM dies joined by TSVs and microbumps, connected to the processor through a substrate such as a silicon interposer.
  6. JEDEC and Industry Leaders Collaborate to Release JESD270-4 HBM4 Standard: Advancing Bandwidth, Efficiency, and Capacity for AI and HPCJEDEC · JEDEC Solid State Technology Association (press release) · 2025JEDEC’s own release, April 16, 2025. HBM4: up to 8 Gb/s across a 2,048-bit interface, up to 2 TB/s per stack; 32 channels per stack. May need a regular browser; the site blocks some automated clients.
  7. Hybrid Bonding Plays Starring Role in 3D ChipsSamuel K. Moore · IEEE Spectrum · 2024Solder microbumps have pitches in the tens of micrometers; imec’s wafer-on-wafer hybrid bonds at a 400 nm pitch; use in stacked cache and HBM.
  8. The UCIe 1.1 Specification: Future Applications of ChipletsDebendra Das Sharma · Universal Chiplet Interconnect Express (UCIe) Consortium · 2023Chiplets let SoCs exceed the reticle; standard (100–130 µm pitch, up to 25 mm) vs advanced package (25–55 µm, up to 2 mm); 165–1,317 GB/s per mm of shoreline in the advanced package; spare lanes in the advanced profile.
  9. Backside power deliveryNaoto Horiguchi and Eric Beyne · imec · 2022Power interconnect takes at least 20% of routing resources; backside delivery with buried power rails cut IR drop about 7× in simulation.
  10. Intel Is All-In on Backside Power DeliverySamuel K. Moore · IEEE Spectrum · 2023PowerVia test chip: over 6% frequency gain, 30% less power loss; new design rules for thermal issues and new debug methods.
  11. What Will That Chip Cost?Brian Bailey · Semiconductor Engineering · 2023A 2018 estimate put a 5 nm design at $542.2M; the article argues such headline figures are inflated and fall as a node matures.
  12. FireSim: FPGA-Accelerated Cycle-Exact Scale-Out System Simulation in the Public CloudSagar Karandikar, Howard Mao, Donggyu Kim, et al. · ISCA 2018 (author copy) · 2018FPGA-accelerated, cycle-exact simulation of a 1,024-node cluster at a 3.4 MHz simulated clock, under 1,000× slower than real time and several orders of magnitude faster than existing software simulators.