Guide 3 of 4 · 16 chapters

Architectures

Different chips are built for different jobs. This guide shows how a CPU, a GPU, an AI chip and an FPGA each work, and why each one is fast at some jobs and slow at others.

Organized by kind of chip, then by idea: how CPUs run one program fast, how GPUs run thousands of threads at once, the main ways chips are built for AI, and how FPGAs can be rewired after they are made.

Pipelines, speculation and coherence in CPUs; SIMT execution in GPUs; roofline analysis, dataflow taxonomies and the memory wall in AI accelerators; and LUT fabrics and the FPGA flow, using only openly published data.

The Expert level assumes you know binary numbers and logic gates, that a program is a list of instructions, and that a neural network runs mostly on matrix multiplication.

  1. CPUs · general-purpose cores

    01Instructions and pipelines
    A CPU follows a list of simple commands: load a number, add two numbers, jump to another step. It works like an assembly line, starting the next command before the last one is done.
    An instruction set is the contract between software and hardware. A pipeline splits each instruction into stages, fetch, decode, execute, memory and write back, so several are in flight at once. Hazards, where one instruction needs another’s result, stall the line unless forwarding gives it a shortcut.
    RISC and CISC instruction sets with RISC-V as the worked example, the classic five-stage pipeline, structural, data and control hazards, forwarding and the load-use stall, the branch penalty, and CPI and the iron law of performance.
    Full chapter
  2. 02Branch prediction and out-of-order execution
    A CPU guesses which way a program will go and starts work early. It also runs steps in whatever order is ready first, then puts the results back in order.
    Branch predictors guess the outcome and target of each branch before it is known, so the pipeline keeps filling. Out-of-order cores look ahead through a window of instructions, run whichever have their inputs ready, and retire them in program order so the result looks sequential.
    Branch predictors from two-bit counters through gshare to TAGE, target buffers and return stacks, the misprediction penalty, Tomasulo’s reservation stations, register renaming, the reorder buffer and precise exceptions, load/store queues, ILP limits and window sizing, and what Spectre and Meltdown showed about speculation’s side effects.
    Full chapter
  3. 03Caches and coherence
    Main memory is far away and slow, so a CPU keeps copies of recent data in small, fast memories close by. When several cores share data, they have to agree on which copy is current.
    Caches work because programs reuse recent data and touch nearby addresses. L1, L2 and L3 caches trade size for speed, and a TLB speeds up address translation. Coherence protocols keep the copies held by different cores consistent when one of them writes.
    Locality, cache organization (sets, ways and lines), the three Cs, replacement and write policies, average memory access time, prefetching, TLBs and virtually indexed caches, the MESI protocol, snooping versus directories, false sharing, and memory consistency models.
    Full chapter
  4. 04Multicore and vector units
    Clock speeds stopped rising in the mid-2000s because chips got too hot, so chips added more cores instead. Each core can also do the same math on a whole row of numbers at once.
    When power limits stopped clock rates climbing, CPUs went multicore. Amdahl’s and Gustafson’s laws say how much extra cores help, threads can share a core, the cores share a cache and memory over an on-chip network, and vector (SIMD) units apply one instruction to many values, the idea GPUs take much further.
    The end of Dennard scaling, the power and ILP walls, Amdahl’s and Gustafson’s laws under a power budget, simultaneous multithreading, rings, meshes, sliced last-level caches, NUMA and chiplets, fixed-width and length-agnostic SIMD (AVX-512, NEON, SVE, RISC-V V), and big and little cores.
    Full chapter
  5. GPUs · thousands of threads at once

    05SIMT and GPUs
    A GPU is a chip with thousands of small math units that all follow the same steps on different numbers. It was built to draw video game graphics and turned out to suit AI too.
    GPUs group threads into warps that run in lockstep. Each core keeps many warps ready, so while one waits on memory another runs. Dedicated matrix units now do most of the AI math.
    SIMT execution, warp scheduling and occupancy, the register file and shared-memory hierarchy, matrix-multiply instructions on tensor units, and the programmability that keeps GPUs general purpose.
    Full chapter
  6. AI accelerators · built for neural networks

    06What the workload needs
    AI mostly does one thing over and over: multiply big tables of numbers. A chip’s speed depends on how fast it does math and how fast it can fetch the numbers.
    Neural networks are dominated by matrix multiplication and attention. Arithmetic intensity, the number of operations per byte moved, tells you whether a job is limited by math or by memory. The roofline model plots both limits on one chart.
    Operation counts for matrix multiplication, convolution and attention; the arithmetic intensity of prefill versus decode and how batching and the KV cache move it; the roofline model and its ceilings; and why training and inference stress different resources.
    Full chapter
  7. 07Number formats
    Computers can store numbers with many digits or just a few. AI often works fine with fewer, like rounding 3.14159 to 3.1, and that makes chips faster and cheaper to run.
    Training moved from 32-bit floating point to 16-bit formats like BF16, and inference increasingly uses 8-bit and 4-bit formats. Fewer bits mean smaller multipliers and less data to move, at some cost in precision.
    FP32, FP16 and BF16, the two FP8 variants, integer quantization, block-scaled microscaling (MX) formats and FP4, accumulation precision, and how a multiplier’s cost grows with mantissa width.
    Full chapter
  8. 08Systolic arrays
    A systolic array is a grid of math units that pass numbers to their neighbors in a steady rhythm, like a bucket brigade. Each number gets used many times without another trip to memory.
    Data flows through a grid of multiply-accumulate units. A weight-stationary array keeps the weights in place and streams inputs past them; an output-stationary array keeps the running sums in place. Reuse inside the array cuts memory traffic.
    The dataflow taxonomy (weight-, output-, input- and row-stationary), utilization and fill-and-drain overhead, mapping matrix multiplication and convolution onto an array, and published TPU designs as case studies.
    Full chapter
  9. 09Dataflow and spatial meshes
    Some AI chips split the work across many small processors laid out in a grid. Each has its own memory, and neighbors pass data straight to each other.
    Spatial chips place many cores, each with local memory, on an on-chip network. A compiler maps each layer onto groups of cores and schedules the data moving between them, so data stays on the chip.
    Spatial and dataflow execution, tiling and placing layers, network-on-chip bandwidth and congestion, the compiler’s central role, and the trade-off between flexibility and efficiency.
    Full chapter
  10. 10Wafer-scale and SRAM-heavy designs
    Chips are usually cut from a big disc of silicon. Some designs keep the whole disc as one giant chip, or fill chips with fast memory, to keep the data right next to the math.
    Wafer-scale chips connect many regions across one wafer and route around manufacturing defects. SRAM-heavy designs keep a model’s weights in on-chip memory for very fast access and spread large models over many chips.
    Yield and redundancy at wafer scale, wiring across reticle boundaries, power delivery and cooling, deterministic scheduling in designs without off-chip memory, and the capacity limits that push them to many chips.
    Full chapter
  11. 11The memory wall
    AI chips can often do math faster than they can fetch the numbers. Keeping the math units busy is one of the hardest parts of the design.
    Compute speed has grown faster than memory bandwidth. High-bandwidth memory (HBM) stacks DRAM right beside the chip, and large on-chip SRAM reuses data. Generating text one token at a time is usually limited by memory bandwidth, not math.
    Bandwidth and capacity trends, HBM generations and the packaging constraints they bring, SRAM area cost at advanced nodes, KV-cache growth during decode, and techniques such as batching and compression that raise arithmetic intensity.
    Full chapter
  12. 12Sparsity, in-memory and analog compute
    Some designs skip multiplying by zero, do math inside the memory itself, or use smooth signals instead of on-off switches. Each idea saves energy and brings new problems.
    Sparsity hardware skips multiplications by zero. In-memory computing does multiply-accumulate inside memory arrays to avoid moving data. Analog compute adds and multiplies with currents or charges, trading precision for energy.
    Structured and unstructured sparsity and the hardware that exploits each, SRAM and resistive compute-in-memory, converter overheads, noise and accuracy limits in analog compute, and what published prototypes report.
    Full chapter
  13. 13Comparing chips
    Comparing AI chips takes more than reading the biggest number on the box. What matters is how fast they run real AI programs and what that costs.
    Peak numbers rarely match real workloads. Benchmarks like MLPerf measure complete training and inference tasks under shared rules. Performance per watt and total cost of ownership add power, cooling and price.
    MLPerf methodology and divisions, utilization versus peak throughput, performance per watt at chip, server and rack level, total-cost-of-ownership models, and how to read vendor claims critically.
    Full chapter
  14. FPGAs · hardware you can rewire

    14Lookup tables and the FPGA fabric
    An FPGA is a chip you can rewire after it is made. It is full of small logic blocks and switches, and loading a new file joins them into a different circuit.
    An FPGA is a grid of lookup tables, flip-flops, memory blocks and DSP blocks, linked by programmable routing that takes most of its area. A configuration file, the bitstream, sets every table and switch, so the same chip can become many different circuits.
    LUTs as small truth tables, logic blocks and carry chains, island-style routing and switch boxes, block RAM, DSP slices and hard IP, SRAM, flash and antifuse configuration, bitstreams and partial reconfiguration, and why routing dominates FPGA area and delay.
    Full chapter
  15. 15From hardware code to bitstream
    To use an FPGA, you describe a circuit in code. Software tools turn that code into the file that wires up the chip. The steps look like designing a chip, but the chip already exists, so a build takes minutes or hours, not months.
    The FPGA flow mirrors the chip design flow: synthesis maps the design onto lookup tables and hard blocks, packing and placement fit it onto fixed sites, routing picks existing wires and switches, and timing analysis checks the clock. The result is a bitstream that loads in seconds.
    LUT mapping with K-feasible cuts, FlowMap and area recovery; packing; annealing and analytic placement on typed sites; negotiated-congestion routing on a fixed graph; timing closure without gate sizing; the open bitstream databases behind Yosys and nextpnr; HLS and on-chip debug.
    Full chapter
  16. 16Where FPGAs win
    An FPGA is slower and uses more power than a chip built for one job, but it can be changed. That makes it good for testing new chips, for products made in small numbers, for jobs that change often, and for jobs where every microsecond counts.
    FPGAs trade efficiency for flexibility. They win where a custom chip would never pay back its design cost, where the design must change after it ships, when prototyping chips before tapeout, and in networking and trading jobs that need low, predictable latency. Hybrids mix FPGA fabric with fixed processors and engines.
    The FPGA–ASIC gap in area, speed and power and how hard blocks narrow it, NRE and break-even volume, field upgrades, prototyping economics, SmartNICs and low-latency trading, datacenter and AI-inference deployments, radiation and obsolescence, and adaptive SoCs, eFPGAs and structured ASICs.
    Full chapter