Before anyone designs a chip in detail, someone has to decide what big pieces it has. They also decide how those pieces talk to each other. That job is architecture. The architect turns the spec’s wish list into a drawing of boxes and arrows.
Architects test their plan on computer models first, because changing it later costs months. Every later step in building the chip has to live with these choices.
A chip that runs a phone or a server is not one big circuit. It is dozens of blocks, each with its own job, wired together. Architecture is the stage where engineers decide what those blocks are. It is stage 2 of the flow. It starts from the specification, the written description of what the product must do: which programs it must run (its workloads), how fast, within what power budget and cost, and which outside connections it needs. It ends with a block-level plan that the next stage turns into code. That code is called RTL (register-transfer level): a description of the hardware that later tools convert into logic gates.
The plan lists every block, what it does, how blocks connect and how fast each connection is, and how the chip is clocked and powered. Some teams separate this chip-level plan from the of each block: for a processor, how it breaks each instruction into steps and how big its caches are.
Most chips today are a (SoC): processor cores, accelerators, on-chip memory, the wiring between them (the interconnect) and controllers for external memory and I/O, all on one piece of silicon. A processor core runs software, a list of instructions, each a small step such as “add these two numbers” or “load this value from memory.” An is a block built for one kind of work, such as graphics, video or AI. The architect’s job is to balance (performance, power and area, where area means silicon and silicon means cost) across those blocks for the target workloads.
Specialization is the biggest lever. A general processor spends far more energy working out what each instruction means than doing the work itself: fetching and interpreting one instruction costs 10× to 4000× the energy of a simple operation such as an add. A block that skips that overhead does its narrow job far more cheaply.1 How AI accelerators in particular are organized is covered in the Architectures guide, for example in SIMT and GPUs and Systolic arrays.
Architects test their choices on models before anything is built: formulas and spreadsheets first, then simulators (programs that imitate the proposed chip running real software) such as gem5, and faster “transaction-level” models written in SystemC.34
Architecture defines the problem everyone downstream solves. It fixes:
- the partition: which blocks exist and how the design nests into sub-blocks;
- the and the bandwidth between its levels;
- the interconnect, and how copies of data in different caches are kept consistent;
- how many pipeline stages each datapath has, and so how much logic must fit between registers in one clock cycle;
- the clock and power domains.
Those choices bound what later stages can reach. Synthesis maps the RTL onto gates, the floorplan places the big blocks on the die, and timing closure makes every path meet the clock. None of them can rescue a pipeline stage that holds too many levels of gates for one cycle, or a wide bus that the architecture says must cross the whole die in a single cycle. Fixing either means changing the RTL, and so the architecture.
The work is quantitative, and it runs in rounds. First, simple formulas prune the space of possible designs: Amdahl’s law (how much speeding up part of a job helps), the roofline model (whether a task is limited by compute or by memory bandwidth), AMAT (average memory latency) and queueing models (how waiting time grows with load). Each is explained further down this page. Second, cycle-level simulators rank the surviving designs on recorded workloads. Third, power and area models such as McPAT turn the simulator’s activity counts (how often each cache, core or link is used) into watts and square millimeters.6
Hardware and software are designed together. For many accelerators the algorithm itself changes, trading extra arithmetic, which is cheap in custom hardware, for less memory traffic, which is expensive; even then, memory usually ends up taking most of the accelerator’s area and power.1 The Architectures guide follows that thread in The memory wall.
- Energy gap, 4-core processor chip vs. custom ASIC, HD video encoding
- 500×
- Energy to fetch and decode an instruction vs. one add
- 10–4000×
- Share of the TPU v1 die spent on control logic
- 2%
An ASIC (application-specific integrated circuit) is a chip custom-built for one task. The figures come from Hameed et al., Dally et al. and Google’s paper on the TPU v1, a custom chip for running neural networks.219
Tap or hover a block to read what it does, then choose where the video job runs.
Architecture starts from the spec: what the chip must do, how fast, on how much power, and at what cost.
It ends with a drawing of all the blocks and how they connect. Each block gets a budget: this much space and this much power.
| Direction | What | Form |
|---|---|---|
| In | Product spec: workloads, performance targets, power and heat budget, cost or area target, interfaces | Requirements document |
| In | Representative workloads | Benchmark programs, application code, recorded traces of instructions and memory addresses |
| In | Technology assumptions: gate delay, wire delay, SRAM density and energy, energy per operation | Spreadsheets, early data on the target process’s cells and memories |
| In | Blocks available to license or reuse (IP): processor cores, interconnect, memory controllers, PHYs | IP catalogs and datasheets |
| Out | Architecture spec: blocks, interfaces, memory map, clock, reset and power plan | Document and block diagram |
| Out | Performance model | Analytical scripts, simulator configurations (gem5), a SystemC/TLM virtual platform |
| Out | Per-block budgets: area, power, latency in clock cycles, clock frequency | Spreadsheet; later, first drafts of the timing constraints (SDC) and power intent (UPF) |
| Out | Make-or-buy list | IP list with versions and license terms |
A few of these need a word of explanation.
- Traces are recordings of a real program’s instructions and memory addresses, which a simulator can replay.
- PHYs (physical-layer circuits) are the mostly analog circuits that drive signals on and off the chip, for example to memory chips or over USB.
- The memory map gives every block and memory a range of addresses, so software knows where to read and write.
- The architecture spec is the contract with the RTL team: what each block must do and how it talks to the others.
- A virtual platform is a software model of the whole chip that can run real programs, so software teams can start long before the chip exists. SystemC TLM provides standard interfaces for exchanging such models across architecture analysis, software development and performance analysis.4
- The power plan later becomes a UPF file, written in the IEEE 1801 standard format, which says which regions can be switched off or run at other voltages and which special cells that requires.7
Two outputs are easy to underrate.
Interface definitions. For every connection the spec pins down the protocol (the rules for requesting and answering a transfer), the data width, how many requests may be outstanding at once, whether responses must come back in order, and the address map. These are what let blocks from different teams and suppliers work together. Get one wrong and it surfaces late, when the blocks are first connected.
Per-block budgets. Large chips are built hierarchically: each big block is synthesized, placed and routed on its own, then assembled. A block’s budget of area, power and cycles of latency becomes its timing budget (how much of each clock cycle its inputs and outputs may use) and the first input to the floorplan. Treat every budget as an estimate with an error bar and record the assumption behind it, so that when RTL numbers arrive you know which decisions to revisit.
Inputs on the left, outputs on the right. Tap or hover one to read what it holds.
Splitting the chip into blocks
- Processor cores can run any app. They are flexible, but they spend a lot of energy figuring out what each command means.
- do one kind of job, such as video or AI. They are like the pizza oven. In one study, ordinary processors used about 500 times more energy than a special video block to do the same video job.2
- Memory comes in layers. A tiny, fast memory sits right next to the processor, like the fridge at each station. The huge main memory is farther away and slower, like the big cooler down the hall.
- Wires between blocks are the hallways. If they are too narrow, fast blocks sit idle, waiting for data.
Assembly lines
A chip has a clock: a signal that ticks billions of times a second, like a drummer keeping the beat. On each tick, every part of the chip takes one step.
Inside a processor, each command is split into small steps, like an assembly line. This is called . Shorter steps let the clock tick faster. But sometimes the processor guesses wrong about what comes next. Then it throws away the half-finished work on the line. A longer line throws away more. So there is a sweet spot, and you can find it below.
Many hands, with a catch
More workers only help with the part of a job they can share. Some parts must be done in order, by one worker.17
Saving power
Running everything at full speed all the time would drain a phone’s battery fast. So architects split the chip into areas that can be slowed down or switched off on their own.
Partitioning
Start from the workloads. Profile them (measure where the time and energy go) and decide which parts run on general-purpose processor cores and which deserve hardware of their own. A typical SoC has:
- a cluster of CPU cores (central processing units, the general-purpose processors), each with small private and sharing a larger one; cache levels are named L1, L2 and L3, from fastest to largest;
- accelerators: a GPU for graphics, an NPU (neural processing unit) for AI, video encoders and decoders, a DSP (digital signal processor) for audio and radio, a crypto engine;
- on-chip SRAM, fast memory built from transistors on the die itself;
- the interconnect that links them all;
- controllers and PHYs for the external main memory;
- I/O controllers for USB, cameras, displays and networks;
- housekeeping blocks for clock generation, power management, security and debug.
Accelerators win in two ways that multiply. Specialization cuts the energy per operation, and the energy saved pays for doing many operations in parallel.1 A study by Hameed et al. shows how big the gap is. A four-core processor chip encoding HD video (720p, in the H.264 format) used 500× more energy than an ASIC, a chip custom-built for the job. Adding SIMD units, which apply one instruction to several numbers at once, improved energy 10× but left it 50× worse than the ASIC, because 90% of the energy still went to overhead. Only units built for the specific video algorithms brought it within 3× of the ASIC.2
Memory hierarchy and bandwidth
Memory comes in layers, the . Closest to the arithmetic are registers, a few dozen storage slots inside the processor. Next comes the cache: a small, fast memory that automatically keeps copies of recently used data in fixed-size chunks called lines, each with a tag recording which address it came from. If the data is there, that is a hit; if not, a miss, and the line must be fetched from the next level. Accelerators often use instead: on-chip memory that software fills and empties on purpose, so its timing is predictable. Google’s TPU v1, a chip for running neural networks, has a 24 MiB scratchpad called the Unified Buffer (a MiB is bytes, about 1.05 million).9 Off chip sits , the large, slow main memory.
Time on a chip is counted in clock cycles. The clock is a signal that ticks billions of times a second, and each tick is one cycle. Memory latency is summarized by , .10 A worked example: an L1 cache answers in 1 cycle and misses 5% of the time; an L2 answers in 10 cycles and misses 20% of the requests that reach it; DRAM takes 200 cycles. A request that misses L1 costs cycles on average, so cycles.
Bandwidth, how many bytes per second a memory can deliver, is the other half. Once arithmetic is cheap, accelerators often wait on memory rather than on math: four of the six production neural networks studied on the TPU v1 were limited by memory bandwidth.9 The AI architectures guide treats this problem in depth in The memory wall.
Interconnect and coherence
The interconnect is how blocks reach each other. A is one shared set of wires that blocks take turns using. A is a switch that connects any input directly to any output, so several transfers can happen at once. A (NoC) is a small packet network of routers and links, laid out as a ring, a grid (mesh) or a grid whose edges wrap around (torus); most on-chip networks put a router next to each block.11 Buses suit a few slow peripherals, crossbars a handful of fast blocks, and networks-on-chip dozens of blocks.
When several cores keep copies of the same address in their own caches, a write by one core would leave the others holding stale data. is the hardware that prevents this, by invalidating or updating the other copies. One approach, snooping, broadcasts every request to all caches; it is simple but hard to scale to many cores. The other keeps a directory that records which caches might hold each line and messages only those, which scales better at the cost of storage to keep those records.12 How the caches themselves are organized, and how protocols such as MESI work, is covered in Caches and coherence.
Pipelining and the clock
Logic on a chip is organized between flip-flops (also called registers): one-bit storage elements that capture their input on each clock tick. Each cycle, signals leave one set of flip-flops, pass through logic gates, and must reach the next set before the following tick. The slowest such path sets how fast the clock can run.
splits long logic into shorter stages with registers in between, like an assembly line. The clock period becomes roughly . A processor runs instructions this way, several in flight at once. The catch is : moments when the next instruction can’t proceed, most often after a branch (an instruction that decides what runs next, like an if-statement). A guesses the outcome; a wrong guess throws away the work already in the pipeline, and a deeper pipeline throws away more. So , the average number of cycles per instruction, rises with depth. Performance is , and it peaks at a finite depth. Hartstein and Puzak derived that optimum as a formula.19 How the pipeline itself runs instructions, and what stalls it, is the subject of Instructions and pipelines.
Parallelism and its limit
The other route to speed is doing more at once: more cores, wider SIMD units, larger arrays of multiply-accumulate (MAC) units. Its limit is : , where is the fraction of the original run time that gets faster and is how much faster that part gets.18 Accelerate 80% of the run time by 10× and you get ×. Even an infinitely fast accelerator caps out at 5×, because the other 20% still takes its full time.
Clock and power domains
Blocks that need different clock speeds, or that must stop independently, get their own . A signal passing between two unrelated clocks is a . The receiving flip-flop may catch it in the middle of changing and hang between 0 and 1 for a while, a state called metastability, so every crossing needs a synchronizer circuit. Crossing bugs are among the most common reasons for a re-spin, a costly second round of manufacturing to fix a chip that came back broken.13
A is a region that can be switched off when idle () or run at its own voltage, which the chip may lower together with the clock speed when there is little work (, dynamic voltage and frequency scaling). The plan is written down as power intent in a UPF file.7
Reuse, make or buy
Few teams design every block. Standard pieces such as CPU cores, the interconnect, memory and PCIe controllers (PCIe is the standard expansion connection in PCs and servers) and PHYs are usually licensed from a supplier or reused from earlier chips as , and the team’s own effort goes into the blocks that set the product apart. Open-source generators, programs that produce a block’s design from a list of parameters, show the idea. The Chipyard project collects generators for an in-order RISC-V core (Rocket), which runs instructions strictly in program order, an out-of-order core (BOOM), which lets later instructions go ahead while earlier ones wait, a matrix-multiply accelerator (Gemmini) and a network-on-chip (Constellation).14 RISC-V is an open instruction set.
When to split into chiplets
A design can also be split across several dies (separate pieces of silicon) in one package; each piece is a . Reasons include going beyond the reticle limit (the largest area the manufacturing equipment can pattern in one exposure), better yield (the share of chips that work, which is higher for small dies), using the best manufacturing process for each die, and reusing one die across products.15 The cost is die-to-die links that add latency, power and packaging expense. The Systems guide covers this in Packaging.
Partition by traffic
A good partition keeps heavy, latency-critical traffic inside a block and puts traffic that can tolerate delay on block boundaries. Count the boundaries before you commit, because each one costs twice later. Every boundary becomes a boundary in hierarchical implementation, where the block is laid out on its own and needs its own timing budget and pin plan. Every boundary between unrelated clocks also becomes a clock-domain-crossing structure (a synchronizer or an asynchronous FIFO, a small buffer written on one clock and read on the other) that must be verified.
Memory system
AMAT is recursive: the L1 miss penalty is the L2’s own AMAT, and so on down to DRAM. The tension at L1 is capacity against hit time: a bigger cache misses less but answers more slowly, and every instruction that touches memory waits on it. One rule of thumb is to use the biggest L1 that keeps hit time at 1–2 cycles, about 16–64 KB in modern technology.10
Accelerators often trade caches for scratchpads. The TPU v1 used a large software-managed memory and a deterministic execution model, with no caches, branch prediction or out-of-order execution, because predictable timing matched its 99th-percentile latency limits better than CPU and GPU features that improve only the average case.9
Bandwidth also needs concurrency. says the average number of requests in flight equals the request rate times how long each one takes.22 Turned around: to sustain bandwidth when each request takes latency , a requester must keep bytes outstanding. At 17 GB/s and 100 ns that is 1,700 bytes, about 27 cache lines of 64 bytes, and the hardware needs a tracking slot for each one: miss-handling registers in a cache, tags in a DMA engine, buffer entries in the interconnect.
Interconnect and coherence
Topology choice trades four things: hop count (routers a packet passes through, a proxy for latency and energy), bisection bandwidth (the total bandwidth across a cut that splits the network into two equal halves, a proxy for how much chip-wide traffic it can carry), router degree (ports per router, which sets router size) and how easily the links lay out on a flat die. Under uniform random traffic, average hop count runs ring > mesh > torus, and a mesh has half the bisection bandwidth of a torus of the same size.11
Coherence choices drive interconnect traffic. A directory avoids broadcasts but adds storage and, when another cache holds the data, an extra trip through the directory.12 Accelerators are often attached without coherence and move data with DMA (direct memory access, a block that copies data between memories without a processor’s help); the TPU v1 uses a programmable DMA controller between host memory and its Unified Buffer.9 Coherent attachment simplifies software but puts the accelerator’s traffic into the coherence protocol, right next to the CPU caches. The protocols, directory organizations and false sharing are worked through in Caches and coherence.
Specialization and its limits
Accelerator speedup is bounded three ways. The first is the offloadable fraction: Amdahl’s law, with the time to hand work over and synchronize counted in the part that isn’t sped up. The second is memory bandwidth, which the roofline model captures. The third is programmability: a block too narrow for next year’s algorithms becomes dead silicon. Dally et al. describe accelerator design as parallel programming guided by a cost model in which arithmetic is free and global memory is expensive.1 The TPU v1 paper estimates that replacing its DDR3 memory with the faster GDDR5 used on GPUs of the time would have roughly tripled achieved throughput, a reminder that the memory interface is part of the accelerator.9 The Architectures guide develops these trade-offs for AI chips, starting from What the workload needs.
Clock and power domain strategy
Each power domain multiplies downstream work. Signals leaving a domain that can be off need isolation cells, which hold them at a fixed value so they don’t feed garbage into live logic. Signals between domains at different voltages need level shifters. Registers that must keep their contents through power-off need retention versions. The domain also needs power switches, a set of power states to verify, and an area of its own in the floorplan.78
Unrelated (asynchronous) clock domains save power and let blocks run at their own rates, a style called GALS, globally asynchronous, locally synchronous. But every crossing needs structural and formal CDC checks, because neither RTL simulation nor static timing analysis (STA, which checks every path against the clock without simulating) catches metastability bugs.13 Fewer, well-chosen domains usually beat fine-grained ones whose savings disappear into overhead.
Make or buy, chiplets
Make-or-buy is a risk decision as much as a cost one. Protocol versions, verification quality, configurability, test hooks, power intent and license terms all land on the team that integrates the block. Generator-based IP makes reuse easier, because one generator produces many configurations.14
For chiplets, the partition must respect die-to-die bandwidth. The UCIe consortium, which publishes an open die-to-die standard, cites 20× the I/O performance at 1/20 the power of off-package SerDes (the high-speed serial links used between packages). Its targets are still not free: 0.25 to 0.5 pJ per bit and under 2 ns of added latency for transmit plus receive.15 Each crossing costs PHY area at the die edge, energy per bit and nanoseconds that an on-die wire avoids, so put chiplet cuts where traffic is thin.
AMAT = 1 + 0.05 × (10 + 0.2 × 200) = 3.5 cycles. Pick a level to see one request’s trip.
Three transfers want to happen at the same time. Pick an interconnect to see which can.
f = 0.80, s = 10: speedup = 1 / (0.20 + 0.080) = 3.57×. Upper limit 1 / 0.20 = 5.0×.
Busy: every domain on, the CPU at its top clock and voltage.
Architects can’t afford to build a chip just to see if an idea works. So they build models. A simulator is a program that pretends to be the chip. It runs real apps on the pretend chip, only much more slowly. gem5 is a popular free one.3
A predicts how fast a design will be before it exists. Models come in levels, from fast and rough to slow and precise.
- Analytical models. Formulas in a spreadsheet or short script: Amdahl’s law, AMAT, and the . Roofline says a task’s attainable performance is the smaller of two limits: the chip’s peak compute, or its memory bandwidth the task’s (operations done per byte fetched from memory).16 Example: a chip that can do 400 GOPS (billion operations per second), with 17 GB/s of memory bandwidth, running a task that does 2 operations per byte, reaches only GOPS. Memory, not math, is the limit. The Architectures guide works through rooflines for AI chips in What the workload needs.
- Architectural simulators. Programs that imitate the proposed chip, cycle by cycle, while it runs real software. gem5 offers processor models ranging from a simple one-instruction-per-cycle model to a detailed out-of-order core; a mode that runs a single program and imitates the operating system’s services, and a full-system mode that boots a real operating system; and memory models that include cache coherence.3
- . Written in SystemC, a C++ library for modeling hardware. Blocks exchange whole transactions, such as “read 64 bytes from this address,” through function calls instead of setting values on individual wires as RTL does. The loosely-timed style is fast enough to boot an operating system and develop software; the approximately-timed style adds timing detail for exploring the architecture and analyzing performance.5
- Power and area models. McPAT models cores, caches, networks-on-chip and memory controllers. It estimates power from activity counts, how often each part is used, that the performance simulator passes in.6
Models are driven by workloads: benchmark suites (standard test programs), the product’s own applications, or traces recorded from them. closes the loop: when a model shows a bottleneck, sometimes the cheapest fix is in the software. Dally et al. note that the algorithms often have to change for an accelerator, trading extra computation for less memory bandwidth.1
Treat the models as a ladder trading speed for fidelity. Analytical models evaluate thousands of design points per second and catch first-order mistakes, such as a bandwidth shortfall or an Amdahl cap. gem5’s CPU models sit at deliberately different points on the speed-versus-accuracy spectrum; the simplest is suited to fast-forwarding, so one study can skip quickly through a program’s start-up and then measure the region of interest with the detailed model.3 RTL simulation and FPGA emulation (running the RTL on a reprogrammable chip) come later; they are the reference you calibrate the faster models against.
Roofline is most useful with ceilings: lower lines under the roof that show the limit when an optimization is missing, for example code that doesn’t use SIMD instructions or reads memory in a scattered order. The roofline authors measured sustainable memory bandwidth with their own microbenchmarks rather than trusting the datasheet peak.16 For the TPU v1, the ridge point (where the slope meets the flat roof, the intensity needed to reach peak) sat at 1,350 operations per byte of weight memory, against 13 for a contemporary server CPU and 9 for a GPU. Its measured MLP and LSTM layers, network types that use each weight only a few times, sat under the bandwidth slope.9
Early power models run off the same activity counts. McPAT reads per-component access counts from the performance simulator through an XML file and turns them into dynamic power (from switching) and leakage power (which flows even when nothing switches).6 Before RTL exists their absolute numbers carry wide error bars; use them to compare candidates, and re-anchor them to synthesis results as blocks mature.
Start with a large design space: core counts, cache sizes, accelerators, interconnects. Each dot is one candidate.
Below is an assembly line for commands. It starts with 5 steps. Add steps and watch the speed climb. Work done climbs too, then levels off and starts to fall. That’s because a longer line wastes more each time the processor guesses wrong.
Now make the processor guess right more often. The sweet spot moves to a longer line.
To see how a processor makes these guesses, read Branch prediction and out-of-order execution.
The simulation models a processor pipeline. You set three things: the pipeline depth , the number of stages (1–30); the register overhead per stage, the fixed time each stage’s registers add (30–150 ps, where a picosecond is a trillionth of a second); and the branch predictor accuracy, how often it guesses right (80–99%).
The model splits 2400 ps of logic evenly across the stages, so the clock period is . : every extra stage adds to the average cycles per instruction, where covers wrong guesses and other hazards. , in billions of instructions per second. At the defaults (, 90 ps, 90%) you get a 1.75 GHz clock (1.75 billion ticks a second), CPI 1.32 and performance 1.33. Find the peak: it is at , with a 4.48 GHz clock, CPI 2.36 and performance 1.90. Depths from 15 to 20 all land within 1% of the peak, so the curve is flat near the top.
How real predictors reach such accuracies, and how a core keeps working past a stall, is the subject of Branch prediction and out-of-order execution.
The sim also shows the formula and the analytic optimum. Time per instruction is the clock period times CPI, , where is the register overhead. Setting its derivative in to zero gives . That has the same square-root form as Hartstein and Puzak’s : total logic delay over latch overhead in the numerator and a hazard term in the denominator (the symbols are explained under “Under the hood”).19 Sweep the inputs: 99% prediction accuracy moves the peak to , 80% moves it to (exactly tied with 16), and 150 ps of overhead moves it to .
The flat top is the practical lesson. A few stages either side of the optimum cost under 1%, so real teams pick depth for power, verification effort and timing margin. Weighting power in the metric pushes the optimum shallower still.19
Real cores also shrink itself, with TAGE-class predictors and out-of-order windows that overlap stalls; see Branch prediction and out-of-order execution.
At an architecture review, the team walks through the drawing and asks three questions.
- Does every job have a home? Each feature in the spec needs a block that does it, or a piece of software.
- Are the busiest hallways wide and short? Blocks that pass lots of data should sit close together, with fast memory nearby.
- Does it fit the budget? Add up the power and space of every block. The total must fit the spec, with some left over.
Below is an illustrative architecture summary for a small smart-camera chip (not a real product), written the way a block-level spec might list it. Each line names a block and its key numbers; the numbered notes under the code explain the abbreviations. A few that recur:
- KB, MB: thousands and millions of bytes; mm2: square millimeters of silicon.
- GHz: billions of clock ticks per second; W: watts of power.
- 128-bit, 256-bit: how many wires carry data side by side on a connection.
- Coherent / non-coherent: whether the hardware keeps that block’s view of memory consistent with the processors’ caches, or software has to.
Below are an illustrative block-level summary, a back-of-the-envelope model run against it, and a trade-off table of the kind an architecture review works through. The numbers are invented but internally consistent; check that the model’s conclusions follow from the summary.
# Illustrative smart-camera SoC, one die (not a real product)
BLOCK cpu 4x in-order RISC-V, 32 KB L1I + 32 KB L1D each, 1 MB shared L2
BLOCK npu 16x16 MAC array, 8-bit, 0.8 GHz, 2 MB scratchpad, DMA engine
BLOCK media image signal processor + video encoder, fixed function
BLOCK sram 4 MB shared on-chip SRAM, 8 banks
BLOCK ddr LPDDR controller + PHY, 32-bit, 4266 MT/s, 17 GB/s peak
BLOCK io camera input, USB, SPI, UART, GPIO
BLOCK sys PLLs, power manager, boot ROM, security, debug
NOC 2D mesh 3x3, 128-bit links, 1.0 GHz
LINK cpu -> noc coherent, L2 is the point of coherence
LINK npu -> noc 256-bit, non-coherent DMA, driver flushes caches
LINK media -> noc non-coherent, streaming, needs QoS priority
CLOCK cpu 0.4-1.6 GHz (DVFS) | npu 0.8 | noc 1.0 | ddr from PHY | aon 32 kHz
POWER PD_CPU (gated, DVFS) | PD_NPU (gated) | PD_MEDIA (gated) | PD_AON (always on)
BUDGET 2.5 W typical: cpu 0.9, npu 0.8, media 0.4, ddr+io 0.4
AREA 25 mm2 target; SRAM is about 40% of it- 1L2Four small RISC-V processor cores that run instructions in program order, each with 32 KB level-1 caches for instructions (L1I) and data (L1D), sharing a 1 MB L2. The L2 size trades AMAT against area; it becomes one or more SRAM macros in the floorplan.
- 2L3The AI accelerator (NPU): a 16×16 grid of multiply-accumulate (MAC) units on 8-bit numbers. 256 MACs × 2 ops × 0.8 GHz = 409.6 GOPS peak. The scratchpad holds weights and activations so layers can reuse them; the DMA engine copies data in and out.
- 3L4The image signal processor turns raw sensor data into pictures; the encoder compresses video. Fixed function means hardwired, not programmable.
- 4L6LPDDR is low-power DRAM. Peak bandwidth = 4 bytes (32 bits) × 4266 million transfers per second (MT/s) ≈ 17 GB/s. Plan on sustaining well below peak.
- 5L7Simple I/O: SPI and UART are slow serial links to other chips; GPIO pins are general-purpose on/off signals.
- 6L8Housekeeping. PLLs generate the clocks; the boot ROM holds the first code run at power-on.
- 7L10A network-on-chip: a 3×3 grid of routers joined by 128-bit links. A small mesh gives every block a nearby port without one long shared bus.
- 8L11The L2 is the point of coherence, where the four cores’ copies of data are kept consistent.
- 9L12Non-coherent DMA keeps NPU traffic out of the coherence protocol, at the cost of explicit cache maintenance in software (the driver flushes caches before handing data over).
- 10L13Camera data arrives at a fixed rate and can’t wait, so the interconnect needs a priority scheme (QoS, quality of service).
- 11L15Five clocks; aon is the always-on domain, ticking at 32 kHz while the rest sleeps. Several crossings are between unrelated clocks; each needs a synchronizer or async FIFO, and CDC checks.
- 12L16Four power domains (PD). Each gated domain becomes a voltage area with isolation cells and a power switch network.
- 13L17The power budget, split by block. The sum must fit the product’s battery and heat limits with margin.
- 14L18SRAM-heavy designs are common: memory often dominates accelerator area.
The script below is an illustrative first-pass model of that chip, the kind of thing an architect writes before any simulator is set up. It checks four things: a roofline for the NPU (is it limited by compute or by memory?), AMAT for the processor caches, Amdahl’s law for one camera frame (how much does the NPU really help?), and Little’s law for the number of memory requests that must be in flight to use the full DRAM bandwidth.
# arch_model.py: back-of-the-envelope model of the SoC above
peak_ops = 2 * 16 * 16 * 0.8e9 # NPU: 409.6e9 ops/s
dram_bw = 17.0e9 # bytes/s, peak
ridge = peak_ops / dram_bw # ~24 ops/byte
def roofline(intensity): # ops per DRAM byte
return min(peak_ops, dram_bw * intensity)
conv = roofline(200.0) # conv layer, weights reused: 409.6e9 (compute-bound)
fc = roofline(2.0) # fully connected, batch 1: 34.0e9 (memory-bound)
# CPU: average memory access time, in cycles
l1_hit, l1_miss = 1, 0.05
l2_hit, l2_miss = 12, 0.20 # L2 local miss rate
dram_lat = 150
amat = l1_hit + l1_miss * (l2_hit + l2_miss * dram_lat) # 3.1 cycles
# Amdahl: the NPU speeds up 85% of the frame time by 20x
f, s = 0.85, 20.0
speedup = 1 / ((1 - f) + f / s) # 5.2x, capped at 6.7x by the other 15%
# Little's law: 64-byte requests in flight to sustain peak DRAM bandwidth
latency = 100e-9 # seconds
in_flight = dram_bw * latency / 64 # ~27 outstanding lines- 1L2Peak compute: MACs × 2 operations (a multiply and an add) × clock rate.
- 2L4The ridge point. Tasks doing fewer than about 24 operations per byte fetched can’t reach peak, no matter how many MACs you add.
- 3L9A convolution layer (common in image networks), split into tiles well, reuses each weight many times, so it sits right of the ridge.
- 4L10A fully connected layer run on one input at a time (batch 1) uses each weight once: about 8% of peak. This is the case that sizes the scratchpad and the DRAM interface.
- 5L14“Local” miss rate: the share of requests reaching L2 that also miss there.
- 6L16L1 miss penalty = 12 + 0.2 × 150 = 42 cycles, so AMAT = 1 + 0.05 × 42 = 3.1 cycles.
- 7L201 ÷ (0.15 + 0.0425) ≈ 5.2×. With an infinitely fast NPU it would still be 1 ÷ 0.15 ≈ 6.7×. Speeding up the CPU’s 15% now pays more than a bigger NPU.
- 8L2417 GB/s × 100 ns = 1,700 bytes ≈ 27 lines in flight. If the NPU’s DMA engine can track only 8 outstanding requests, it can’t reach peak bandwidth.
Architecture reviews often summarize choices in a table like the one below. The last column is the one the physical design team, the engineers who lay out the chip, cares about.
| Decision (illustrative) | Option A | Option B | What it changes later |
|---|---|---|---|
| NPU scratchpad size | 2 MB: smaller die, more DRAM traffic on large layers | 8 MB: fewer DRAM trips, about 4× the scratchpad SRAM area | SRAM blocks dominate the floorplan; more DRAM traffic means more power in the memory interface |
| Interconnect | Shared bus: simplest, bandwidth shared by all | 2D mesh NoC: scalable, adds router area and hop latency | A bus is one long bundle of wires driven by many blocks, hard to make fast; a mesh needs regularly placed routers and short links |
| CPU pipeline | 5 stages: lower clock, simpler hazards | 10 stages: higher clock, more flip-flops and a bigger cost per wrong branch guess | Logic per stage sets how hard timing closure is; more flip-flops add load on the clock network and power |
| NPU coherence | Non-coherent DMA | Coherent port into the L2 | Coherent traffic adds snoop or directory load near the CPU cluster |
| Power domains | 2 (always-on, everything else) | 5 (one per major block) | Each domain needs a voltage area, isolation, level shifters, switches and its own verification states |
| Packaging | One die | Compute chiplet plus I/O chiplet | Die-to-die PHYs on facing edges; the cut must carry the traffic that crosses it |
Each spec feature maps to a block (or to software on the CPU). An unmapped feature is a missing block or a missing driver.
A house’s blueprint fixes where the walls and pipes go before the builders arrive. In the same way, the architect’s choices become fixed limits for everyone who builds the chip later.
Later stages turn the plan into a physical layout, and each architectural choice becomes a fixed constraint there.
- Memories become , pre-built blocks placed as a unit. Each SRAM in the architecture is a fixed rectangle in the , the map of where big blocks sit on the die. The TPU v1’s 24 MiB Unified Buffer takes almost a third of the die, and its size was picked partly so it would line up with the pitch (row spacing) of the matrix unit beside it.9
- Power domains become voltage areas. Each domain in the UPF file becomes a , a region of the layout with its own supply and its isolation and level-shifting cells.8
- Clock domains become a list of crossings that need synchronizers.13 Each domain also gets its own clock tree later, the network of wires and buffers that delivers the clock to every flip-flop.
- Wide links become wiring demand. A 256-bit link across the die is hundreds of long wires, often needing extra registers along the way so signals can make the trip within a clock cycle.
- Pipeline depth becomes , the number of gates a signal passes through between registers. Fewer stages means more gates per stage, and a harder (the slowest path, which sets the clock) in synthesis and placement.
The architecture’s cycle budgets are physical claims. A link that must cross a 10 mm die in one cycle, a cache sized for a 1-cycle hit, or a stage holding a certain number of delays all assume wire and gate delays that the floorplan must then deliver. (One FO4 is the delay of an inverter driving four copies of itself, a unit that stays roughly constant across manufacturing processes.) Hrishikesh et al. put the best clock period for integer code at about 8 FO4: 6 of useful logic and about 2 of latch and clock overhead. They also found that further pipelining could at best double integer performance over the designs of the day.20
At that depth, overhead is a quarter of the cycle, so the clock tree’s skew (the clock arriving at different flip-flops at slightly different times) and jitter (variation from one tick to the next) matter as much as the logic. Budget global links in cycles from the start. Adding pipeline registers to a link whose protocol assumed one-cycle latency is an RTL change, which physical design can’t make on its own.
Block budgets turn into and pin plans for . Accelerator datapaths with regular structure (MAC arrays, banked SRAM) floorplan well when the architecture sizes them to tile cleanly, as the TPU’s buffer-to-array pitch match shows.9
The architecture as drawn: blocks, a wide link, two clocks and a power domain. Switch views, or tap a part.
- Planning for the wrong job. If the example programs don’t match what people really run, the chip is fast at the wrong things.
- Forgetting the hallways. Lots of number-crunching power is useless if data can’t reach it fast enough.
- Speeding up only part of the work. A brilliant special block helps little if most of the time is spent somewhere else.
- Hopeful guesses. Early guesses about power and size tend to grow as the real design fills in. Teams keep some room spare for that.
- Unrepresentative workloads. Small benchmark programs whose data fits entirely in the caches never exercise main memory, so they hide memory problems. Use the product’s own applications and traces as early as possible.
- Peak instead of sustained bandwidth. A roofline drawn with the memory datasheet’s peak number is optimistic. The roofline authors measured the bandwidth a real program can sustain with small test programs; do the same, or model it.16
- Ignoring Amdahl. An accelerator that covers 60% of the run time caps the speedup at 2.5×, however fast it is. Count the time spent handing work to it and collecting results as part of the time that isn’t sped up.
- Memory-bound accelerators. Once arithmetic is cheap, memory dominates. The TPU v1 left performance on the table because its memory interface was too slow for most of its workloads.9
- Too many clock and power domains. Each one adds clock domain crossings, isolation cells and power states to verify. Crossing bugs are a common reason for re-spins.13
- Late surprises from IP. A licensed block that expects a different version of the connection protocol, resets differently or has a different power plan turns into weeks of integration work. Check interfaces at architecture time.
- Uncalibrated models. Simulator results without a calibration point are hypotheses. Calibrate against RTL or silicon of a similar block, and track the error as blocks arrive.
- Averages hiding tails. Average utilization looks fine while queues explode in bursts. Size for the tail. In the simplest queueing model (M/M/1: random arrivals, random service times, one server), the mean time in system is , so at 90% utilization it is already 10× the service time.21
- Too little concurrency. Bandwidth targets that ignore Little’s law fail because the requesters can’t keep enough transactions in flight to cover the latency.22
- Coherence and ordering found late. Mixing coherent and non-coherent agents without a clear contract for ordering and cache maintenance (who flushes what, and when) produces bugs that only appear under specific traffic.
- Optimizing a flat optimum. Near the best pipeline depth, performance in the model above changes by under a percent per stage, while latch count and power keep climbing with depth.19
- Chiplet cuts in the wrong place. Cutting through a high-bandwidth path forces wide die-to-die interfaces, with their energy per bit, PHY area and latency.15
Memory-bound: 17.0 GB/s × 2 ops/byte = 34.0 GOPS, 8% of the 410 GOPS peak.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Optimal pipeline depth
Hartstein and Puzak model the time to run instructions as a function of six quantities: the total logic delay of the processor, the latch overhead added by each stage, the number of stages , the number of hazards , the average superscalar degree (instructions issued per cycle), and , the average fraction of the pipeline stalled by each hazard. More stages shorten the logic part of each cycle () but add overhead per cycle and lengthen each stall. Differentiating with respect to and setting the result to zero gives:19
The formula shows the opposing pulls. A higher superscalar degree shortens the optimum, and fewer hazards lengthen it. A larger ratio of logic delay to latch overhead () leaves more room for pipelining. Performance-only studies of several machines put the optimum at 8–10 FO4 per stage, logic plus overhead.19
Hrishikesh et al. reached a similar answer by simulating a high-performance core in 100 nm technology: about 6 FO4 of useful logic plus about 2 FO4 of overhead for integer code, 6 FO4 in total for floating point, and insensitive to latch and clock-skew overhead.20 Power changes the answer. With the metric (billions of instructions per second, cubed, per watt), Hartstein and Puzak’s optimum moved to about 22.5 FO4, roughly 7 stages, and it depends strongly on how the number of latches grows with depth. More dynamic power shortens the optimum; clock gating (stopping the clock to idle registers) lengthens it.19 The simulation above uses the performance-only form.
Roofline
, with intensity measured in operations per byte moved to or from DRAM. On log-log axes the bandwidth term is a line of slope 1 and the compute term is flat. The ridge point’s x-coordinate is the minimum intensity needed for peak; if it sits far right, only kernels with very high intensity reach peak. Ceilings add lower roofs for missing optimizations: in-core ones (instruction-level parallelism, SIMD, a balanced mix of floating-point adds and multiplies) and memory ones (unit-stride access, memory affinity, software prefetching).16
For accelerators, draw one roofline per level of the hierarchy. Raising intensity through tiling or a larger scratchpad moves a kernel right; raising bandwidth lifts the slope. The TPU v1 analysis is the standard worked example.9
Amdahl, Gustafson, and multicore
Amdahl’s 1967 argument was a limit: if a fraction of the work must run serially, even an infinite number of processors speeds the job up by at most .17 In formula form, with serial fraction and parallel fraction of the time on one processor (), speedup on processors . Gustafson observed that problem size usually grows with the machine. Measuring and on the parallel machine instead gives , and he reported speedups of 1016–1021 on a 1024-processor machine for applications with a 0.4–0.8% serial fraction.17 Use Amdahl when the spec fixes the job (latency per frame or per query) and Gustafson when it fixes the time and lets the work grow (throughput, batch size).
Hill and Marty add a hardware cost model. A chip has room for base-core equivalents (BCEs), each the area of one simple core, and a core built from BCEs delivers , so doubling single-thread performance costs four BCEs. Asymmetric chips, with one big core and many small ones, can beat symmetric ones by a wide margin and are never worse: for and the best asymmetric speedup is 125.0 against 51.2 symmetric. Methods that look locally inefficient can be globally efficient because they shorten the serial phase while the other cores sit idle.18 The same reasoning justifies keeping a strong CPU next to a large accelerator. Multicore and vector units adds a power budget to the same models.
Cache misses: the 3Cs and AMAT
The classify misses by simulation. Compulsory misses occur even in an infinite cache. Capacity misses occur even with full associativity and perfect replacement. Conflict misses are the remainder, caused by where blocks are allowed to sit. Larger capacity cuts capacity misses, higher associativity cuts conflict misses, and larger lines cut compulsory misses when nearby data is used together, each at a cost in hit or miss latency.10
AMAT nests: , where is a level’s hit time and its miss rate (local: counted over the requests that reach that level). It is a latency model. Out-of-order cores overlap misses, so the effective stall time also depends on how many misses are in flight at once (memory-level parallelism), which Little’s law ties to bandwidth.
NoC topologies
Topology fixes hop count, bisection bandwidth, node degree and layout cost. A bus is a single shared channel and a natural broadcast medium. A crossbar connects inputs to outputs in one hop and sits inside every router. Among direct networks (a router at every node), a ring is a -ary 1-cube; a mesh drops a torus’s wraparound links and loses its edge symmetry, which concentrates traffic on the center channels. Under uniform random traffic, average hop count runs ring > mesh > torus, and a mesh has half the bisection bandwidth of the equivalent torus.11
A torus’s long wraparound wires are tamed by folding, which interleaves the nodes to equalize link lengths. Concentration attaches several cores to one router to save area and power, at the risk of bottlenecks during bursts. Packet latency is head latency, roughly proportional to hop count, plus serialization latency for a packet of length on a channel of bandwidth .11 Chipyard’s Constellation is one open-source generator for exploring these choices.14
Queueing intuition for contention
Model a congested link or memory controller as an M/M/1 queue: requests arrive at random at rate , a single server handles them at rate , and service times are also random. Utilization must stay below 1 for the queue to be stable, and the mean time in system is , which is the service time divided by .21 At 50% load the latency doubles; at 90% it is 10×; near saturation it grows without bound. Real on-chip traffic often comes in bursts that this model doesn’t capture, so treat its curve as a best case.
Little’s law, , is the companion tool. It gives the occupancy (buffer entries, outstanding requests) needed to sustain a throughput at a given latency, and it holds whatever the arrival pattern or service order.22 Together they explain why architects plan interconnect and memory controllers for well under 100% average utilization, and why a NoC’s latency curve stays flat until it suddenly isn’t.
ρ = 0.50: W = S/(1 − ρ) = 2.0 × S = 20 ns; L = ρ/(1 − ρ) = 1.0 requests in the system on average.
Q1A cache answers in 1 cycle when it has the data, misses 10% of the time, and a miss costs 20 more cycles. What is the average memory access time (AMAT)?
Q2An accelerator can do at most 1,000 billion operations per second (GOPS) and fetch 100 GB/s from main memory. A task does 2 operations for every byte it fetches. What does the roofline model predict?
Q3An AI accelerator speeds up 60% of a program’s run time by 10×. What is the overall speedup?
Q4When many processor cores keep copies of the same data in their own caches, why do large chips track the copies with a directory instead of snooping?
Q5In the pipeline simulation with default settings, why does performance fall when you go well past 18 stages?
Sources
Show Hide 22 sources
- Domain-Specific Hardware AcceleratorsInstruction overhead 10× to 4000× a simple ADD; efficiency from specialization, performance from parallelism; memory dominates DSA area and power; algorithms must be codesigned.
- Understanding Sources of Inefficiency in General-Purpose ChipsH.264 encoder: a four-core CMP was 500× less energy efficient than an ASIC; SIMD gave 10× energy, still 50× worse; specialized units reached within 3× of the ASIC.
- The gem5 SimulatorCPU models from AtomicSimple (suited to fast-forwarding) to out-of-order O3 on a speed-versus-accuracy spectrum; syscall-emulation and full-system modes; Classic and Ruby memory systems with coherence protocols.
- SystemC Transaction Level Modeling (TLM)TLM standard interfaces for model exchange across architecture analysis, software development, performance analysis and hardware verification; loosely-timed and approximately-timed styles; part of IEEE Std 1666.
- OSCI TLM-2.0 Language Reference ManualTLMs communicate by function calls rather than pin-level events; loosely-timed style suits software development on a virtual platform and can boot an OS; approximately-timed style suits architectural exploration and performance analysis; generic payload and sockets for interoperability.
- McPAT: An Integrated Power, Area, and Timing Modeling Framework for Multicore and Manycore ArchitecturesModels cores, NoCs, shared caches, memory controllers and multi-domain clocking; consumes activity statistics from a performance simulator via XML; dynamic, short-circuit and leakage power.
- Low Power Design, Verification, and Implementation with IEEE 1801 UPF (tutorial)UPF (IEEE 1801) captures power intent: power domains, supply rails, power switches, isolation cells, level shifters, retention registers and power states; power gating, multi-voltage and DVFS; RTL + UPF is verified, then drives synthesis, DFT and place and route.
- Unified Power Format (upf)Reads UPF into the physical design flow; set_domain_area gives each power domain an area in the floorplan; isolation, level-shifter and power-switch strategies.
- In-Datacenter Performance Analysis of a Tensor Processing Unit24 MiB Unified Buffer, almost a third of the die, sized partly to match the matrix unit’s pitch; control is 2% of the die; 4 of 6 apps memory-bound; rooflines with ridge points at 1350, 13 and 9 ops/byte; GDDR5 would triple achieved TOPS.
- Caches (continued), MIT 6.5900 Lecture L03AMAT = hit time + miss rate × miss penalty; compulsory, capacity and conflict misses; biggest L1 that keeps hit time at 1–2 cycles.
- ECE 1749H: Interconnection Networks for Parallel Computer Architectures: TopologyBus, crossbar, ring, mesh, torus; most on-chip networks use direct topologies; hop count ring > mesh > torus; mesh has half the bisection bandwidth of a torus; folding and concentration; latency = head + serialization.
- Directory-Based Cache Coherence, MIT 6.5900 Lecture L13Snoopy schemes broadcast and are hard to scale; directories message only possible sharers and need extra storage; data held in another cache passes through the directory, adding latency.
- Pragmatic Formal Verification Methodology for Clock Domain Crossing (CDC)GALS designs create asynchronous domains; CDC paths are prone to metastability and need synchronizers; CDC bugs are a common cause of re-spins; RTL simulation and STA are not enough; structural CDC analysis.
- Chipyard ComponentsOpen-source generator-based SoC IP: Rocket (in-order) and BOOM (out-of-order) RISC-V cores, Gemmini matrix-multiply accelerator, Constellation NoC generator.
- The UCIe 1.1 Specification: Future Applications of ChipletsChiplets exceed the reticle limit, enable die reuse, smaller dies for yield, process choice per die; 20× I/O performance at 1/20 the power of off-package SerDes; targets of 0.5 / 0.25 pJ/b and under 2 ns Tx+Rx latency.
- Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore ArchitecturesAttainable performance = min(peak compute, peak bandwidth × operational intensity); ridge point; computational and memory ceilings; sustainable DRAM bandwidth measured with microbenchmarks.
- Reevaluating Amdahl’s LawAmdahl’s 1967 argument that a serial fraction s caps speedup at 1/s; states Amdahl’s formula; scaled speedup s′ + p′N; speedups of 1016–1021 on 1024 processors with 0.4–0.8% serial fraction.
- Amdahl’s Law in the Multicore EraSpeedup = 1/((1 − f) + f/S); base core equivalents; perf(r) = √r; asymmetric chips beat symmetric ones (125.0 vs 51.2 at f = 0.975, n = 256).
- The Optimum Pipeline Depth Considering Both Power and PerformanceJournal version of the authors’ MICRO-36 paper. Summarizes the authors’ 2002 theoretical and simulated study of performance-optimal depth and restates its optimum p_opt² = N_I·t_p ÷ (α·γ·N_H·t_o); performance-only studies find 8–10 FO4; weighting power shortens the optimum; latch count grows with depth.
- The Optimal Logic Depth Per Pipeline Stage is 6 to 8 FO4 Inverter DelaysOptimal clock period about 8 FO4 for integer code (6 useful + about 2 overhead), 6 FO4 for floating point; deeper pipelining at best 2× for integer.
- M/M/1 Queue (2.854 lecture notes)One server with service rate µ, arrival rate λ < µ; utilization ρ = λ/µ; mean number in system ρ/(1 − ρ); by Little’s law, mean time in system Ws = 1/(µ − λ).
- Little’s Law as Viewed on Its 50th AnniversaryL = λW; holds without assuming a stationary arrival process and independent of queue discipline; applications in computer architecture.