Architectures · Chapter 4 of 16 · CPUs

Multicore and vector units

Clock speeds stopped rising in the mid-2000s because chips got too hot, so chips added more cores instead. Each core can also do the same math on a whole row of numbers at once.

When power limits stopped clock rates climbing, CPUs went multicore. Amdahl’s and Gustafson’s laws say how much extra cores help, threads can share a core, the cores share a cache and memory over an on-chip network, and vector (SIMD) units apply one instruction to many values, the idea GPUs take much further.

The end of Dennard scaling, the power and ILP walls, Amdahl’s and Gustafson’s laws under a power budget, simultaneous multithreading, rings, meshes, sliced last-level caches, NUMA and chiplets, fixed-width and length-agnostic SIMD (AVX-512, NEON, SVE, RISC-V V), and big and little cores.

For about thirty years, computer chips got faster in a simple way. Their clocks ticked faster every year. A clock tick is the beat that tells every part of the chip when to take its next step.

Around 2005, that stopped. Faster clocks made chips too hot to cool. So chip makers did something else. They put several processors on one chip. Each one is called a core, and a chip with several is a chip. The biggest chips for data centers now have nearly two hundred cores.

Each core has one more trick. It can do the same math on a whole row of numbers in one step. GPUs, the next chapter, take that idea much further.

Until the mid-2000s, a new processor was faster mainly because its clock was faster. Then clock rates stopped climbing, and chips started adding cores instead: complete processors, side by side on one piece of silicon, each running its own stream of instructions. A chip built this way is a processor.

The reason was power. Transistors kept shrinking, but after about 2005 they stopped getting more energy-efficient as fast as they got more numerous, the end of (the device physics is in Speed and power). Running more transistors faster would have made chips too hot to cool. This chapter, the last of the CPU chapters, covers what designers did instead:

  • Two walls: the power wall that stopped clock rates, and the ILP wall that stopped single cores from getting much wider.
  • More cores, and the two laws (Amdahl’s and Gustafson’s) that say how much they help.
  • Threads sharing a core, to fill the gaps one program leaves.
  • Connecting the cores: rings, meshes, a shared cache, memory that is nearer to some cores than others, and chips built from several smaller dies.
  • Vector units, which apply one instruction to many numbers.
  • Big and little cores on the same chip.

It builds on the earlier CPU chapters: Instructions and pipelines, Branch prediction and out-of-order execution and Caches and coherence.

Treat the chip as a power-limited system. You have a die full of transistors and a thermal budget, the . Switching power is P≈αCV2fP \approx \alpha C V^2 f, and in the usual operating range the supply voltage must rise roughly in proportion to the clock, so one core’s power grows roughly as f3f^3. Once supply voltage stopped scaling with feature size, the cheapest way to turn transistors into throughput was more cores at moderate clocks, not one core at a higher one.

The rest of the chapter is about how that budget gets spent:

  • Cores versus clock. With P∝Nf3P \propto N f^3, a fixed budget gives f∝N−1/3f \propto N^{-1/3} and peak throughput Nf∝N2/3Nf \propto N^{2/3}, until the voltage floor, leakage and the uncore stop the gain. Amdahl’s serial fraction decides whether that throughput is usable.
  • Threads per core. SMT trades per-thread latency for utilization of an expensive core.
  • Data movement. Rings give way to meshes; the sliced and memory controllers make latency depend on placement (); chiplets add die-to-die hops.
  • Width per instruction. SIMD amortizes fetch, decode and scheduling over 4–16 lanes, from fixed-width (SSE/AVX/AVX-512, NEON) to length-agnostic (SVE, RVV) designs.
  • Asymmetry. Big and little cores in one package, scheduled by capacity and energy.

Products appear only as case studies, x86 (AMD, Intel), Arm and RISC-V side by side, with dated figures in an As-of box.

dieslice 0LLCslice 1LLCslice 2LLCslice 3LLCcoreL2DRAMDRAMoff-chip →

A generic eight-core processor. Tap or hover a part to see what it is and where this chapter covers it.

A generic eight-core processor (not a real product). Each core has its own threads, vector unit and private caches; the shared cache, interconnect and memory are where cores meet.Share freely with credit: ‘Figure from chipfieldguide.com’

In 1993 the fastest chip’s clock ticked about 200 million times a second. By 2004 it ticked 3.6 billion times a second. Since then, the record has barely moved.

Why? Every tick flips millions of tiny switches. Each flip uses a little energy and makes a little heat. A faster clock also needs a stronger push of electricity, called voltage. So the heat climbs much faster than the speed does.

By the early 2000s, a single chip gave off about 100 watts. That’s as much heat as a bright, old-style light bulb, from a piece of silicon smaller than a postage stamp. Fans and metal fins can only carry so much heat away. This limit is called the .

But the switches kept getting smaller. So every few years, chips had room for twice as many. Chip makers spent that room on more cores, each running at a sensible speed. Pick a line on the chart below to see it happen.

Each time a logic signal switches, it charges or drains a little capacitance. The power that takes is the ,

Pdyn=α C V2f,P_{\mathrm{dyn}} = \alpha\, C\, V^2 f,

where α\alpha is how often signals switch, CC the capacitance, VV the supply voltage and ff the clock. For decades each manufacturing generation shrank transistors and lowered the voltage together, so a chip with twice the transistors at a higher clock used about the same power per square millimeter. Around 2005 the voltage stopped falling, because lowering it further would have made transistors leak too much current (the full story is in The end of Dennard scaling). From then on, a faster clock meant much more heat.

The trend data shows the turn:

  • Clock. The fastest clock rose from about 200 MHz in 1993 to 3.65 GHz in 2004 and has never passed 4.7 GHz since.
  • Single-thread performance. It grew 52% a year from 1986 to 2002 and less than 20% a year after. A straight-line fit to the trend data gives about 45% a year before 2004 and about 12% a year since.
  • Power. Typical processor power rose from a few watts to around 100 W by the early 2000s, and its growth slowed sharply after that.
  • Cores. The first multicore processors appear in the data in 2004–05. In 2005 Intel followed IBM’s Power4 and Sun’s Niagara in moving its high-performance processors to multiple cores. A widely read article that year told programmers that “the free lunch is over”: software would have to run in parallel to keep getting faster.

The Berkeley parallel-computing report summed it up as a new rule: “power is expensive, but transistors are ‘free’. That is, we can put more transistors on a chip than we have the power to turn on.” Researchers later estimated that a growing share of a chip would have to be switched off at any moment, which they called .

In the range where a chip is designed to run, the supply needed for a clock ff rises roughly linearly with it, V≈V0+kfV \approx V_0 + k f, and switching power per core goes as CV2fC V^2 f. Take the sim’s illustrative curve (0.6 V up to 1 GHz, then +0.15 V per GHz): one core at 4 GHz needs 1.05 V and burns 1.052×4≈4.41.05^2 \times 4 \approx 4.4 units, while two cores at 2 GHz need 0.75 V and burn 2×0.752×2≈2.32 \times 0.75^2 \times 2 \approx 2.3 units, the same peak throughput for about half the power. That ratio is the economic argument for multicore, provided the work can be split.

The architecture-level evidence came before the device-level crisis. Agarwal et al. modeled an aggressive out-of-order core from 250 nm to 35 nm and found that, with clock gains shrinking and wire delay growing, no scaling strategy delivered more than 12.5% a year, against the 50–60% the industry was used to. The Berkeley report’s “brick wall” combined three limits: power, memory latency and ILP. It also noted that static (leakage) power had reached as much as 40% of the total in desktops and servers, which matters for the multicore trade-off because idle cores still leak unless they are power-gated.

Esmaeilzadeh et al. pushed the argument one step further. Combining device projections with multicore performance models, they estimated that 21% of a fixed-size chip would have to be dark at 22 nm and more than 50% at 8 nm, and that multicore scaling alone would deliver only 7.9× average speedup through 2024. Their exact percentages depend on the projections; the lasting result is that performance per watt, not transistors per die, sets the design. The responses all spend transistors to save energy: more and simpler cores, wider vector units that amortize instruction overhead, heterogeneous cores and fixed-function accelerators such as those in the AI chapters.

transistors (×1000)single-thread perf.clock (MHz)power (W)logical corespower wall1101001k10k100k1M10M197019801990200020102020data: K. Rupp et al. (CC BY 4.0)
Series

Fifty years of microprocessors on one logarithmic axis: each gridline is 10× the one below. Pick a series to isolate it.

Microprocessor trends, 1971–2021. Clock and power flatten around 2005 while transistors keep growing; the extra transistors became cores. Data: K. Rupp et al., CC BY 4.0.Share freely with credit: ‘Figure from chipfieldguide.com’

There was a second wall. Engineers had already made each core clever. It could start several steps of a program in the same tick, as long as those steps didn’t depend on each other.

But programs are full of steps that need the answer from the step before. You can’t ice a cake before you bake it, however many bakers you hire. So making one core wider helped less and less.

More cores get around this, but only if the job is split into separate parts. Each part runs on its own core. Splitting the job is the programmer’s task.

Instruction-level parallelism (ILP) is the number of instructions from a single program that can run at the same time. Through the 1990s, cores found more of it by issuing several instructions per cycle and by running them out of program order, as Branch prediction and out-of-order execution explains. The measure of success is , instructions completed per clock cycle.

The limit is the program itself. Instructions depend on each other’s results, branches must be guessed, and loads that miss the cache take hundreds of cycles. In a 1995 simulation study, even an 8-issue core sustained under 1.5 instructions per cycle, using only 19% of its issue slots. The figure below shows the effect on a small piece of code: past about four instructions per cycle, a wider core barely finishes sooner.

Wider cores also cost more. Every extra issue slot needs more register-file ports, more comparisons to find ready instructions and more wires to pass results between units, and the delay of that logic grows faster than the width does. The Berkeley report called this the : diminishing returns on finding more parallelism inside one instruction stream.

Multicore changes where the parallelism comes from. Instead of hardware discovering it within one program, software supplies it as separate threads, which can run on separate cores. That gets around the ILP wall, at the price of making parallel programming everyone’s problem.

Even with perfect branch prediction and an unbounded window, is bounded by the dataflow graph: total instructions divided by the critical-path length in cycles. Real cores sit well below that bound because the window is finite, mispredictions flush it and misses stall it. Tullsen et al. broke the 81% of wasted issue slots in their 8-issue model into many causes with no dominant one, which is why no single fix (a bigger predictor, a bigger cache) recovered it.

The cost side grows superlinearly. Palacharla, Jouppi and Smith modeled rename, the wakeup and select logic that finds ready instructions, and the bypass network that forwards results: rename and wakeup delays have components quadratic in issue width, a fully bypassed design needs 2⋅IW2⋅S2 \cdot IW^2 \cdot S bypass paths for issue width IWIW and SS pipeline stages after the first result-producing stage, and the bypass wire delay grows quadratically with issue width. These sit on the critical path that sets the clock. Agarwal et al.’s wire-delay model is the quantitative version: as features shrink, structures that must be reached in one cycle cover a shrinking fraction of the die, so wide, deep cores lose clock rate or IPC, and annual gains from a conventional core fall to 12.5% or less.

Two responses followed. Inside the core, simultaneous multithreading fills slots that one thread can’t (next sections). Across the chip, thread-level parallelism replaces instruction-level parallelism: the Berkeley report’s “new conventional wisdom” is that increasing parallelism, not clock frequency, is the primary way to raise processor performance.

cycle →11 cycles12345678instruction 0instruction 11instruction 12instruction 22instruction 1instruction 2instruction 13instruction 14instruction 15instruction 16instruction 17instruction 24instruction 25instruction 27instruction 31instruction 3instruction 4instruction 5instruction 8instruction 18instruction 20instruction 6instruction 7instruction 19instruction 21instruction 9instruction 23instruction 29instruction 10instruction 26instruction 28instruction 30■ 1-cycle op ■ load, 3 cycles □ empty slotIPC by issue width1.011.922.732.943.263.28width:
Issue width

Issue width 4: 11 cycles, IPC 2.91, 73% of issue slots filled. No width can beat 10 cycles: the longest chain of dependent instructions.

The same 32 instructions on an ideal core of growing issue width (illustrative program; loads take 3 cycles, other instructions 1). IPC flattens because instructions wait on each other’s results.Share freely with credit: ‘Figure from chipfieldguide.com’

More cores only help if the job can be shared out. Most jobs have some steps that can be shared and some that can’t.

Say 90% of a job can be split among cores, and 10% must be done by one core alone. Even with endless cores, that 10% still takes its full time. So the job can never get more than 10 times faster. This rule is called .

There’s a hopeful side, too. When people get more cores, they usually give the computer bigger jobs: a sharper picture, a longer weather forecast. The part that can be shared grows, and the part that can’t stays small. Then more cores keep paying off. That idea is .

Suppose a fraction pp of a program’s running time on one core can be spread perfectly over NN cores, and the rest must run on one. gives the speedup:

SAmdahl(N)=1(1−p)+p/N  ≤  11−p.S_{\mathrm{Amdahl}}(N) = \frac{1}{(1 - p) + p/N} \;\le\; \frac{1}{1 - p}.

At p=0.95p = 0.95, 64 cores give about 15× and no number of cores can beat 20×. The Architecture stage of the Design Flow guide uses the same law to size accelerators; here it sizes core counts.

Gustafson pointed out in 1988 that people rarely run a fixed-size problem on a bigger machine. They grow the problem (finer grids, more time steps) to fill the time they have. If the parallel part grows with NN and the serial part doesn’t, the useful measure is the :

SGustafson(N)=(1−p′)+p′N,S_{\mathrm{Gustafson}}(N) = (1 - p') + p' N,

where p′p' is the parallel fraction measured on the NN-core run. His group at Sandia reported speedups of 1016–1021 on a 1024-processor machine, for applications whose serial time was 0.4–0.8% of the run.

Neither law is “right”. Amdahl fits a job of fixed size with a deadline, such as one web request or one video frame. Gustafson fits throughput work that grows with the hardware. Most server and scientific work is closer to Gustafson; much interactive desktop and phone work is closer to Amdahl, which is one reason single-thread speed still matters.

The two formulas measure pp on different machines, so they don’t contradict each other. Writing ss for the serial fraction on one core, the same program’s serial fraction measured on NN cores is s′=s/(s+(1−s)/N)s' = s/(s + (1-s)/N); substituting it into Gustafson’s formula gives exactly Amdahl’s. What differs is the assumption about what is held fixed: work (Amdahl) or time (Gustafson).

Hill and Marty made Amdahl’s law a design tool for multicore chips. Give the chip nn base-core equivalents (BCEs) of area and let a core built from rr BCEs run at perf(r)=r\mathrm{perf}(r) = \sqrt{r}. A symmetric chip of n/rn/r such cores gives

Ssym=(1−fperf(r)+f rperf(r) n)−1,S_{\mathrm{sym}} = \left( \frac{1-f}{\mathrm{perf}(r)} + \frac{f\,r}{\mathrm{perf}(r)\, n} \right)^{-1},

and an asymmetric chip with one rr-BCE core plus n−rn - r single-BCE cores runs the serial part on the big core and the parallel part on all of them. For f=0.975f = 0.975 and n=256n = 256 the best asymmetric design reaches 125.0 against 51.2 for the best symmetric one. That is the area argument for heterogeneous cores later in this chapter.

Power changes the axis. Hill and Marty count area; a power-limited chip also pays for frequency. Per-core power rising roughly as f3f^3 means a budget split across NN cores forces each to slow down, which costs the serial phase directly. Two hardware responses follow: boost clocks, which give one busy core the power that idle, switched-off cores aren’t using, and big cores dedicated to serial work. The sim below and Under the hood put numbers on both.

1×14×416×1664×64256×2561024×1024cores (log scale) →speedupideal (N×)limit 1/(1−p) = 100×Amdahl: 39×
Model

Fixed job, 99% parallel: 64 cores give 39×; the limit is 100× however many cores you add.

Speedup against core count on log scales. Amdahl’s law assumes a fixed job; Gustafson’s assumes the parallel work grows with the machine, as it often does in practice.Share freely with credit: ‘Figure from chipfieldguide.com’

A core is often stuck waiting. It asks the main memory for a number, and the number takes hundreds of ticks to arrive. Meanwhile, most of the core sits idle.

So many cores can work on two jobs at once. Each job keeps its own notes inside the core. In every tick, the core takes steps from both jobs. When one job is stuck, the other keeps the core busy. This is called .

Each job a core keeps track of is a . The computer sees a core with two of them as two processors. Each job runs a little slower than if it had the core alone. But together they get more done.

A is the state a core needs to run one instruction stream: its registers and program counter. Intel’s manual describes a core with Hyper-Threading as two or more “logical processors”, each with its own architectural state, sharing the core’s execution engine. RISC-V specifications call a hardware thread a hart.

There are three ways to share a core among threads. Coarse-grained multithreading switches when one thread stalls on a long miss. Fine-grained multithreading switches every cycle. (SMT) issues from several threads in the same cycle. Tullsen, Eggers and Levy showed that fine-grained switching used only about 40% of a wide core, because it can only hide whole empty cycles, while SMT could reach up to 4× the throughput of a single-threaded and 2× that of fine-grained multithreading in their model. The figure shows why: one thread leaves both whole cycles and parts of cycles empty, and SMT fills both.

Real gains are smaller and depend on the workload. AMD, for example, reports that two threads per core gave 34% more throughput than one on its Zen 4 cores, up from 25% on Zen 3. Not every design uses SMT: Intel’s E-core Xeons keep each core single-threaded to give each thread predictable performance.

What gets duplicated is small: architectural registers, PC, some control state and the interrupt controller per logical processor. Everything expensive, the execution engine, the caches and the bus interface, is shared between the threads. That is why SMT is cheap in area and why its benefit is largest on wide out-of-order cores with low single-thread utilization; the wider Zen 4 back end is AMD’s explanation for its higher SMT uplift.

In Tullsen et al.’s terms, a miss creates vertical waste (a fully empty cycle) and limited ILP creates horizontal waste (a part-filled cycle). Fine-grained multithreading removes only vertical waste; SMT attacks both, and is “only limited by the issue bandwidth of the processor”. The costs are per-thread latency (a thread runs slower than alone), contention for shared caches and queues, which can make SMT a net loss for cache-sensitive code, and isolation: threads on one core share microarchitectural state, so some designs give each thread its own core. Intel cites “performance isolation” as a reason its E-cores are single-threaded.

SMT also interacts with power. Two threads on one core use the same clock and voltage, so the extra throughput comes at nearly the cost of the extra switching it causes, without paying for a second core’s leakage and uncore share. That makes it one of the cheaper ways to raise throughput per watt when there are enough threads.

cycle →25% of slots usedslot 1slot 2slot 3slot 4Ashaded = thread stalled on a cache miss
Threads

One thread: 25% of issue slots used. 7 of 16 cycles are completely empty (a cache miss) and most others are part-empty (not enough independent instructions).

Issue slots of a 4-wide core over 16 cycles (illustrative threads). Each color is a thread; shaded bars below mark when a thread is stalled on memory.Share freely with credit: ‘Figure from chipfieldguide.com’

Cores need to share data. So a chip has roads between the cores, the shared memory and the connections to main memory.

With a few cores, one loop road past every core works fine. With dozens of cores, a loop gets too long. So big chips use a grid of roads instead, like city blocks.

All the cores share one big on-chip memory, the . It is split into pieces, one beside each core.

On a big chip, some main memory sits closer to some cores than others. Reaching far memory takes longer. This is called . The computer tries to keep each program’s data near the core running it.

Some of the biggest chips are really several small chips in one package. Each small chip is a .

Every core needs to reach every other core’s data, the shared cache and the memory controllers. The wiring that does this is a :

  • Ring. One stop per core, data moving around in either direction. Simple and fast for a handful of cores, but the average trip grows in proportion to the number of cores. Intel’s description of its move away from rings: as cores were added, “the access latency increased and available bandwidth per core diminished.”
  • Mesh. A grid of routers, each linked to its neighbors; traffic travels along a column and then a row by a shortest path. Trips grow only with the square root of the number of cores. Large x86 and Arm server chips use meshes: Arm’s CMN-700 mesh supports up to 256 cores per die, and NVIDIA’s Grace CPU uses a mesh with over 3.2 TB/s of bisection bandwidth. (See and Dataflow and spatial meshes for the general idea.)

The shared cache. Behind each core’s private caches sits the (LLC), shared by all cores and split into slices, one per core tile. In Intel’s recent Xeon layout a core tile holds a core with its private L2, an LLC slice with the logic that tracks which cores hold copies of each line, and a mesh router. Keeping those copies consistent is , covered with the protocols in Caches and coherence.

Near and far memory. On a large chip, and always between sockets, memory controllers are closer to some cores than others. The Linux kernel’s description: “memory access time and effective memory bandwidth varies depending on how far away the cell containing the CPU … is from the cell containing the target memory.” The kernel therefore allocates memory from the node local to the core that asks for it. Chips can expose their own halves or quarters as NUMA nodes: Intel calls this sub-NUMA clustering. On the 64-core SG2042 RISC-V processor, which has four NUMA regions, how threads are placed across regions and four-core clusters measurably changes performance.

Chiplets. Very large CPUs are now built from several dies in one package. AMD’s first EPYC used four 8-core , which it estimated was 41% cheaper than one big die; the second generation put up to eight 8-core compute dies around a separate I/O die that holds the memory and I/O controllers. Intel’s recent Xeons similarly separate compute and I/O chiplets. The price is another level of distance: cores on the same die share a cache and talk quickly; cores on different dies go through the die-to-die links. Packaging, die-to-die links and UCIe are covered in Packaging and chiplets.

For a bidirectional ring of NN stops the average distance is about N/4N/4 and the diameter N/2N/2; for a k×kk \times k mesh with XY routing the average is about 2k/32k/3 and the diameter 2(k−1)2(k-1). At 64 cores that is 16 versus about 5 hops on average, which is why rings stopped at a couple of dozen stops. Intel’s Skylake-SP move to a mesh came with a cache rebalance: L2 grew from 256 KB to 1 MB per core and the shared LLC shrank from 2.5 MB (inclusive) to 1.375 MB (non-inclusive) per core, trading shared capacity for fewer trips across a larger network.

Slicing the LLC distributes both capacity and coherence work. Each address belongs to one home slice, and that slice’s snoop filter (a directory of sorts) tracks which private caches may hold the line, so a miss in one core’s L2 costs a mesh trip to the home slice and, if another core holds it modified, a forward from that core. The latency of a shared-cache hit therefore depends on where the home slice is, which makes even a single die somewhat non-uniform; sub-NUMA clustering restricts each half of the die to its own memory controllers and LLC slices to shorten those trips. The coherence protocol itself (-style states, snooping versus directories) and the ways software causes coherence traffic, such as two cores writing different variables in the same , are in Caches and coherence.

Chiplet CPUs add a third tier of distance. With eight 8-core dies, only 7 of a core’s 63 peers share its die and cache; every other exchange crosses two die-to-die links and the I/O die. Putting the memory controllers on the central I/O die makes DRAM roughly equally far from every compute die, while core-to-core and cache-to-cache latency depends on which dies are involved. Operating systems handle this with the same node abstraction and local-allocation policy they use for sockets.

MCMC2D mesh, 8×4; orange: worst case (XY routing), 10 hopsavg 4.0 hops · worst 10
Topology
Cores

Mesh, 32 cores: average 4.0 hops, worst 10. Mesh distance grows with √N, which is why big chips use meshes.

Three generic on-chip layouts: a ring, a 2D mesh and chiplets around an I/O die. Hop counts are exact for these idealized drawings.Share freely with credit: ‘Figure from chipfieldguide.com’

Lots of computer work does the same thing to many numbers. Brighten every pixel in a photo. Add up long lists of prices. Mix two sounds together.

So each core has a . One instruction tells it to do the same math on a whole row of numbers at once: 4, 8 or 16 of them, depending on the chip. Giving one order for 16 numbers takes far less work than giving 16 orders.

Older designs fix the row length inside the program. A program written for rows of 8 always uses rows of 8. Newer designs let the same program use whatever row length the chip has. Try both kinds in the figure, and watch what happens to the numbers left over at the end.

A (single instruction, multiple data) unit applies one instruction to a short vector of numbers held in a wide register. Each major has its own :

  • x86. SSE, introduced with the Pentium III, and SSE2 with the Pentium 4, use 128-bit registers; AVX widened them to 256 bits; AVX-512 has thirty-two 512-bit registers plus eight mask registers that switch individual lanes on or off. Intel’s newer AVX10 defines one converged version of these instructions for both its performance and efficiency cores. AMD’s current server chips also support AVX-512.
  • Arm. Advanced SIMD, better known as NEON, has thirty-two 128-bit registers. The Scalable Vector Extension (SVE) lets each chip choose a vector length from 128 to 2,048 bits, in steps of 128.
  • RISC-V. The V extension leaves the register length (VLEN) to each implementation, as any power of two up to 65,536 bits; the version for application processors requires at least 128.

The split that matters is fixed-width versus . SSE, AVX, AVX-512 and NEON code is compiled for one register width: to use wider registers you recompile, and the elements left over at the end of a loop need extra code. SVE and RISC-V V programs instead ask the hardware how many elements fit and loop accordingly. In SVE’s designers’ words, the model “allows code to run and scale automatically across all vector lengths without recompilation.” In RISC-V V, the program passes the number of elements still to do, and an instruction called vsetvli answers with how many the hardware will handle this time.

The width of the instruction set and the width of the hardware can differ. AMD’s Zen 4 runs 512-bit AVX-512 instructions on a 256-bit datapath over two cycles; even so, needing up to 50% fewer instructions than 256-bit code saved enough power to raise the effective clock in power-limited systems. NVIDIA’s Grace uses Arm cores with four 128-bit SVE2 units each. The 64-core SG2042, a RISC-V server chip, has 128-bit vectors but implements a pre-standard version of the extension (0.7.1) that mainline compilers didn’t support, a reminder that software support matters as much as the hardware.

A vector instruction amortizes the per-instruction overhead (fetch, decode, rename, scheduling, retirement) over 4–16 lanes, the same argument that drives matrix instructions in GPUs. AMD’s Zen 4 numbers make the power side concrete: AVX-512 cut the instruction count by up to half relative to AVX2, reduced front-end and out-of-order tracking power, and gave a 17% higher effective frequency on HPL in power-limited systems, even though most 512-bit operations take two passes through a 256-bit datapath. The flip side is that a wider unit switches more capacitance whenever it runs, so a chip sized for full-width vector code on every core must budget power for it, which is what the sim’s vector-width control models.

Masks and predicates

Conditionals and loop tails need per-lane control. AVX-512 adds eight opmask registers (k0–k7); SVE adds sixteen predicate registers and drives loop control from them, with whilelt building a predicate of the lanes still in range; it also adds first-fault loads, which let a loop read speculatively past the end of an array of unknown length. RISC-V V uses register v0 as a mask and handles tails by shortening vl itself.

Fixed width versus length-agnostic

  • Binary portability. A length-agnostic binary runs on any implementation; a fixed-width one must be rebuilt (or multiversioned) to use wider units. That is why the x86 family has accumulated SSE, AVX, AVX-512 and AVX10 code paths.
  • Tails and alignment. Predication or a shorter vl handles the last partial vector without a scalar remainder loop.
  • What gets harder. The vector length is no longer a compile-time constant. Compilers can’t vectorize by unrolling a loop a fixed number of times, need instructions that use the current length implicitly (SVE’s index and inc), and must place spilled vector registers in stack regions whose offsets are multiples of the length. RISC-V’s register grouping (LMUL) lets one instruction operate on 2, 4 or 8 registers at once, mainly so mixed-width data can be processed at the same element count.

Under the hood walks through the same loop in SVE and RISC-V V assembly.

pass8 lanes (256-bit) →5 passes · 48% lanes busy1012345672891011121314153 scalar164 scalar175 scalar18■ active lane ■ scalar remainder □ inactive lane
Code
Vector width

Fixed 256-bit width: 2 full vector passes of 8 lanes, then a scalar loop for the 3 leftover elements: 5 passes. This binary only ever uses 256-bit registers.

One loop, y[i] = a·x[i] + y[i], on 32-bit numbers. Each row is one pass of the loop. Fixed-width code needs a scalar tail; length-agnostic code masks off unused lanes.Share freely with credit: ‘Figure from chipfieldguide.com’

Not every job needs a big, fast core. Checking for new messages can be done slowly. Opening an app should be quick, because someone is waiting.

So many chips, especially in phones, mix two kinds of core. Big cores are fast but use lots of energy. Little cores are slower but thrifty. The computer sends each job to the kind that suits it. These are .

Arm calls this idea big.LITTLE. Intel calls its two kinds P-cores (performance cores) and E-cores (efficient cores). Both kinds understand the same instructions, so any program can run on either.

run the same instruction set but sit at different points of speed and power. The idea was tested in a 2003 study: giving the operating system a choice of cores cut energy by 39% on average for a 3% loss in performance, more than lowering the voltage and clock of one kind of core could achieve.

The Linux kernel describes such systems in terms of capacity, a core’s performance relative to the fastest one: capacity = work done per clock tick × maximum clock. Its example is Arm’s big.LITTLE, where the big cores have “more pipeline stages, bigger caches, smarter predictors” and reach higher clocks. Its energy-aware scheduler, which runs only on such mixed systems, picks for each task the core that uses the least energy “with a minimal impact on throughput.”

The same trade shows up in servers, usually as whole chips of one kind. AMD’s compact Zen 4c core has the same instructions and the same work per clock as Zen 4, a lower maximum clock, and about 35% less area, so twice as many fit on a compute die. Intel sells Xeons built from its E-cores, in modules of two or four cores sharing an L2 cache, alongside its P-core Xeons. Many simpler cores give more throughput per watt and per square millimeter, which is the parallel case of Amdahl’s and Gustafson’s laws.

Hill and Marty’s asymmetric chip is the area argument: run the serial phase on a big core, worth r\sqrt{r} small cores of performance for rr of their area, and the parallel phase on everything. Kumar et al. made the power argument with real core designs at one ISA, picking the core per program phase. Production systems need three things the models assume away:

  • A scheduler that knows the cores. Linux’s capacity-aware scheduling uses capacity = work per Hz × max frequency to keep tasks that need more than a little core can give off the little cores, and energy-aware scheduling uses a per-core energy model to place each task.
  • ISA parity. A thread can migrate at any time, so every core must accept every instruction the thread might issue. Intel’s AVX10 specification makes the converged vector ISA a feature of both P-cores and E-cores. The RISC-V V specification notes that a thread with live vector state cannot, in general, migrate between harts whose VLEN differs, so mixed designs keep one architectural vector length even when their datapaths differ.
  • Shared structures. Big and little cores share the cache hierarchy and memory, so their coherence and interconnect latencies differ, and a little core’s cache misses can still occupy shared bandwidth.

Brand names for the same structure: Arm’s big.LITTLE, Intel’s P-cores and E-cores; in servers, dense-core variants such as AMD’s Zen 4c and Intel’s E-core Xeons apply the same idea one chip at a time.

2 bigbigbigtime2 senergy8 J8 littletime4 senergy4 J1 big + 4 littlebigtime2 senergy8 Jfilled = busy core · dashed = power-gated
Task

Urgent single thread: the big core finishes in 2 s; the all-little chip needs 4 s, since extra cores can’t speed up one thread. The mixed chip matches the big chip.

Three chips of equal area. A big core has the area of four little ones, runs twice as fast and draws four times the power (illustrative). Pick a task.Share freely with credit: ‘Figure from chipfieldguide.com’

You are designing a chip that may use at most 120 watts. Pick how many cores it has, how fast its clock ticks, and how many numbers its vector unit handles at once. Then pick a job: one that goes step by step, one that is mixed, or one that splits well.

Add cores and watch the power bar. To stay under the limit, you’ll have to slow the clock. Then press “Show me the best chip” for each job. Is the best chip the same every time?

The sim is a power-budget explorer. Every core runs at the same clock and voltage; the voltage rises with the clock above 1 GHz; and the budget is checked with every core busy. You set the core count, the clock, the vector width and the job’s parallel and vectorizable fractions. Speedup is measured against one core at 3 GHz running everything as scalar code. Things to try:

  • Pick “Serial job” (5% parallel). Jump to the best design: a few cores at the full 5 GHz. Now set 64 cores and fit the clock. Why is the result slower than the reference core?
  • Pick “Mixed job” and jump to the best design. Look at the run-time bar: what share of the time goes to the 10% of work that is serial?
  • Pick “Parallel job” (99% parallel, 80% vectorizable). Why does the best design now use many slow cores and the widest vectors, when the mixed job preferred 128-bit ones?

The model, also in the sim spec: V=0.6 VV = 0.6\,\mathrm{V} up to 1 GHz, then +0.15 V/GHz+0.15\,\mathrm{V/GHz}; per busy core (1+0.25 w/128) nF⋅V2f(1 + 0.25\,w/128)\,\mathrm{nF} \cdot V^2 f (the vector term only if the job has vector code), plus 0.3 A⋅V0.3\,\mathrm{A} \cdot V of leakage per core and 15 W of uncore. Run time is Amdahl’s, split again into scalar and vector work, with vector work w/32w/32 times faster. The Expert view adds the budget, a boost toggle (serial code runs on one core with the others power-gated, at the fastest clock that fits) and a Gustafson toggle (parallel work grows with NN). Try:

  • Mixed job: the best design is 12 cores at 4.8 GHz. Turn boost on and the best design jumps to 64 cores at 1.9 GHz, 15.7× instead of 11.8×. This is why server chips advertise both a base and a boost clock.
  • With boost off, set the mixed job to 512-bit vectors at 12 cores and fit the clock: the vector units’ power pulls the whole chip down to 3.8 GHz and the speedup falls to about 10×. Wide vectors pay only when most of the work is vector code.
  • Halve and double the budget (60 W, 240 W) for the mixed job. The best core count moves 6 → 12 → 24 while the clock barely changes. Extra power buys cores, not frequency.

Deliberately left out: memory bandwidth and caches, coherence and synchronization costs, SMT, per-core voltage domains and temperature-dependent leakage. Parallel work scales perfectly, so the curves are upper bounds.

Loading simulation…
Fastest clock in the trend data: by 2004 vs since
3.65 vs ≤ 4.7 GHz
Single-thread growth per year: 1986–2002 vs after
52% vs < 20%
SMT throughput gain on AMD cores: Zen 3 → Zen 4
25% → 34%
Arm Neoverse V2 core + 2 MB L2 at 2.8 GHz (5 nm)
1.4 W · 2.5 mm²

What these numbers mean:

  • 3.65 versus 4.7: by 2004 the fastest chip ticked 3.65 billion times a second. In all the years since, no chip in the data has gone past 4.7 billion.
  • 52% versus less than 20%: one core used to get about half again as fast every year. After 2002, it got less than a fifth faster each year.
  • 25% → 34%: on one company’s chips, letting each core juggle two jobs got about a third more work done.
  • 1.4 watts: one modern server core, with its own memory, uses about as much power as a small flashlight bulb. That is why a chip can hold dozens of them.

The two EPYC parts are this chapter’s trade-off in one product line. In about the same power, one chip has three times the cores at two-thirds of the base clock: 192 cores at 2.25 GHz in 500 W, or 64 at 3.3 GHz in 400 W. The first suits throughput work with many independent threads; the second suits work with serial bottlenecks, and AMD pitches it as a host for GPU servers.

The Neoverse V2 figure shows why core counts could grow so far: at 2.8 GHz and 0.75 V, a big core with its L2 needs about 1.4 W and 2.5 mm². At that rate the cores of a 72-core chip need around 100 W; the rest of the budget goes to the shared cache, the mesh, memory and I/O.

Per-core power from the TDPs: about 500/192 ≈ 2.6 W per core for the 192-core part, 400/64 ≈ 6.3 W for the 64-core part, both including uncore. The ratio, about 2.4× power per core for about 1.5× base clock (and 1.35× boost), is consistent with power rising much faster than linearly in frequency once voltage must rise with it. The denser part’s Zen 5c cores follow AMD’s compact-core approach: the same IPC and features at lower maximum clocks in less area, which the Zen 4c generation put at about 35% less core-plus-L2 area.

Vector datapaths differ more than their ISAs suggest: full 512-bit AVX-512 on Zen 5, 512-bit operations on a 256-bit datapath on Zen 4, and 4 × 128-bit SVE2 on Neoverse V2, where Arm’s manual gives the SVE vector length as 128 bits, so vector-length-agnostic code sees 128-bit vectors and gets width from issuing to four units. Measured, rather than peak, comparisons belong to Comparing chips.

More cores are not free speed. They come with catches.

Someone has to split the job into pieces. That is hard work for programmers, and some jobs barely split at all.

The cores also share things. They share the big on-chip memory, and they share the pipes to main memory. Add more cores, and each one gets a smaller share. Past a point, extra cores just wait for data. Slide the number of cores in the figure to find that point.

And every core must run a bit slower, so that all of them fit in the power limit.

You getYou give up
More throughput per watt from many moderate coresSingle-thread speed, unless the chip can boost one core; software must supply threads
Higher core utilization with SMTSlower individual threads, shared-cache contention, weaker isolation between threads
A large shared last-level cacheLatency that depends on which slice holds the data; coherence traffic across the chip
Cheaper, larger chips from chipletsSlower communication between dies; another level of non-uniform access
Wide vector unitsPower and clock rate when vector code runs; code that must be vectorized to benefit
Big and little coresA scheduler that must place every task well, and identical instruction sets on both

Common ways multicore performance goes wrong

  • The serial part. A section of code only one thread can run, or a lock every thread waits for, caps the speedup by Amdahl’s law.
  • Memory bandwidth. All cores share the same memory channels, so bandwidth per core falls as cores are added; streaming code stops scaling once the channels are full.
  • Coherence traffic. Cores that write the same data, or different data that happen to share a , bounce that line between their caches (see Caches and coherence).
  • Placement. A thread far from its memory, or threads spread badly across clusters and NUMA regions, can lose a large share of performance.
  • Power versus serial speed. A fixed TDP split across more cores lowers the all-core clock; boost clocks win back serial performance only while the other cores are idle. The 192-core and 64-core EPYC parts above are the two ends of that trade in one family.
  • Bandwidth per core. Delivered throughput is min⁡(N,B/d)\min(N, B/d) core-equivalents for memory bandwidth BB and per-core demand dd: a roofline in core count. The 192-core EPYC above has 12 DDR5 channels, one for every 16 cores, and the Berkeley report counted this memory wall as one of the three walls. Intel’s recent platform answers with more channels and multiplexed DIMMs aimed at more bandwidth per thread. The same pressure, far stronger, shapes AI accelerators; see The memory wall.
  • Interconnect scaling. Rings stop scaling at a few dozen stops; meshes scale further but make LLC hit latency a function of distance; chiplets add a die-to-die tier.
  • SMT. Gains depend on how much one thread leaves idle, and threads contend for caches and queues; some designs drop SMT for isolation and predictable performance.
  • Vector width. Wider registers raise peak throughput per instruction; fixed widths fragment the software ecosystem, and length-agnostic ISAs move the cost into compilers.
  • Software maturity. A new vector extension is only useful once compilers and libraries target it, as the SG2042’s pre-standard RVV 0.7.1 showed.
0048489696144144192192useful work (cores’ worth)cores N →memory limit: 576 GB/s ÷ 4 = 14464 of 64 · 9.0 GB/s each
Per-core demand

64 cores × 4 GB/s = 256 GB/s of 576 GB/s: all cores are fed. Each core could have 9.0 GB/s.

Useful work versus core count when every core streams data from memory. Illustrative socket: 12 DDR5 channels, 576 GB/s shared by all cores.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

1. A power-constrained throughput model

Take NN identical cores sharing one voltage and clock. Per-core power is switching plus leakage, and the chip adds a fixed uncore:

P(N,f)=N(CV(f)2f+IleakV(f))+Punc≤Pmax⁡.P(N, f) = N\left(C V(f)^2 f + I_{\mathrm{leak}} V(f)\right) + P_{\mathrm{unc}} \le P_{\max}.

In the range where V∝fV \propto f, the switching term is ∝Nf3\propto N f^3, so the largest feasible clock is f∗∝(Pmax⁡/N)1/3f^* \propto (P_{\max}/N)^{1/3} and peak throughput Nf∗∝N2/3Pmax⁡1/3N f^* \propto N^{2/3} P_{\max}^{1/3}. Two things end the gain. Below the voltage floor, VV stops falling, power becomes linear in ff, and NfN f is capped at roughly Pmax⁡/(CVmin⁡2)P_{\max}/(C V_{\min}^2); meanwhile each extra core adds leakage and needs a share of the uncore whether or not it is busy. So there is an optimum NN for a given budget even for perfectly parallel work, and it rises with the budget. In the sim (Expert view, mixed job) the best core count goes 6 → 12 → 24 as the budget goes 60 → 120 → 240 W, at almost the same clock.

Now add Amdahl. With serial fraction 1−p1 - p run on one core at fsf_s and the rest on NN cores at f∗(N)f^*(N):

T(N)=1−pfs+pNf∗(N),fs={f∗(N)no boostmin⁡ ⁣(fmax⁡, f1∗)boostT(N) = \frac{1-p}{f_s} + \frac{p}{N f^*(N)}, \qquad f_s = \begin{cases} f^*(N) & \text{no boost} \\ \min\!\left(f_{\max},\ f^*_{1}\right) & \text{boost} \end{cases}

where f1∗f^*_1 is the clock one busy core can reach with the others power-gated. Without boost, raising NN lowers fsf_s and the serial term grows, which is why “a few fast cores” wins for serial code. With boost, fsf_s is decoupled from NN, and the optimum moves to many cores. That is the hardware form of Hill and Marty’s asymmetric chip, achieved in time (one core fast while the others sleep) rather than in area.

2. Hill and Marty’s three chips

With nn BCEs and perf(r)=r\mathrm{perf}(r) = \sqrt{r}, the symmetric, asymmetric and dynamic speedups are

Ssym=[1−fperf(r)+frperf(r) n]−1,Sasym=[1−fperf(r)+fperf(r)+n−r]−1,Sdyn=[1−fperf(r)+fn]−1.\begin{aligned} S_{\mathrm{sym}} &= \left[\frac{1-f}{\mathrm{perf}(r)} + \frac{f r}{\mathrm{perf}(r)\,n}\right]^{-1}, \\ S_{\mathrm{asym}} &= \left[\frac{1-f}{\mathrm{perf}(r)} + \frac{f}{\mathrm{perf}(r) + n - r}\right]^{-1}, \\ S_{\mathrm{dyn}} &= \left[\frac{1-f}{\mathrm{perf}(r)} + \frac{f}{n}\right]^{-1}. \end{aligned}

The dynamic chip, which can gang rr BCEs into one core for serial code and split them for parallel code, bounds the other two. Boost clocks are a partial, power-based version of it: they can’t make a core wider, but they can give it the power that the idle cores aren’t using.

3. The same loop, length-agnostic

DAXPY, yi←axi+yiy_i \leftarrow a x_i + y_i, in SVE, from the SVE paper. There is no remainder loop and no constant for the vector length anywhere in the code.

daxpy_sve.s (adapted from Stephens et al., IEEE Micro 2017)text
// x0 = &x[0], x1 = &y[0], x2 = &a, x3 = &n
daxpy_:
    ldrsw  x3, [x3]                    // x3 = *n
    mov    x4, #0                      // x4 = i = 0
    whilelt p0.d, x4, x3               // p0 = lanes with i+k < n
    ld1rd  z0.d, p0/z, [x2]            // broadcast *a
.loop:
    ld1d   z1.d, p0/z, [x0, x4, lsl #3] // x[i..]
    ld1d   z2.d, p0/z, [x1, x4, lsl #3] // y[i..]
    fmla   z2.d, p0/m, z1.d, z0.d       // y += a*x, active lanes only
    st1d   z2.d, p0, [x1, x4, lsl #3]   // store active lanes
    incd   x4                           // i += VL/64
.latch:
    whilelt p0.d, x4, x3               // recompute the predicate
    b.first .loop                      // loop while any lane is active
    ret
  1. 1L5The predicate covers the lanes still in range, so the last pass needs no special code.
  2. 2L12Advance by however many 64-bit lanes this hardware has; the binary never names a width.
  3. 3L15Branch on the predicate: the loop ends when no lane is active.

RISC-V V expresses the same idea with vsetvli, which returns how many elements the hardware will process this time; the loop subtracts that from the remaining count. From the specification’s vector-add example:

vvaddint32.s (adapted from the RISC-V V specification, v1.0)text
# z[i] = x[i] + y[i], 32-bit integers; a0 = n, a1 = x, a2 = y, a3 = z
vvaddint32:
    vsetvli t0, a0, e32, ta, ma  # t0 = elements this pass (<= n)
    vle32.v v0, (a1)             # load x
      sub a0, a0, t0             # n -= t0
      slli t0, t0, 2             # bytes = t0 * 4
      add a1, a1, t0             # bump x pointer
    vle32.v v1, (a2)             # load y
      add a2, a2, t0             # bump y pointer
    vadd.vv v2, v0, v1           # z = x + y
    vse32.v v2, (a3)             # store z
      add a3, a3, t0             # bump z pointer
      bnez a0, vvaddint32        # loop while elements remain
      ret
  1. 1L3The hardware chooses vl from the request and its own VLEN: strip-mining in one instruction.
  2. 2L5The remaining count drops by whatever the hardware took, so the last pass is just shorter.

A CPU’s vector unit does one math step on a row of 4 to 16 numbers. The program has to say “do these 8 together.” A GPU takes the same idea much further. It runs thousands of little workers, and each worker’s program handles just one number. The chip groups the workers into teams that move in step.

Step through the figure to see the same job both ways. Then go on to SIMT and GPUs.

Vector units and GPUs share the hardware idea, one instruction driving many lanes, but differ in who sees the lanes. In the program is written in vectors: the width is in the instruction set, and the compiler or programmer builds the masks for conditionals and loop tails. On a GPU, in , the program is written for a single thread, and the hardware groups 32 threads into a and masks off the threads on the other side of a branch. In the CUDA guide’s words, “SIMD vector organizations expose the SIMD width to the software, whereas SIMT instructions specify the execution and branching behavior of a single thread.”

A GPU also takes the other ideas of this chapter much further: thousands of hardware threads per chip instead of two per core, used to hide memory latency rather than to fill issue slots. That is the subject of SIMT and GPUs, the next chapter.

Length-agnostic SIMD and SIMT converge from opposite ends. SVE and RVV take the width out of the binary but keep the vector in the program; SIMT keeps a scalar-thread program and puts width, masks and reconvergence in hardware, so correctness never depends on warp size while performance does. The GPU then multiplies this chapter’s other levers: many simple cores per die, dozens of resident warps per core instead of SMT’s two to four threads, and matrix instructions that amortize per-instruction overhead even further than wide vectors. Continue with SIMT and GPUs; for how CPUs and accelerators are joined in a server, see Board and server.

CPU: SIMD (vector code)v = load(x[i : i+8])m = (v < 0)y = v * 2y = select(m, -v, y)store(out[i : i+8], y)width and mask are in the codeGPU: SIMT (one thread’s code)x = data[tid]if (x < 0) y = -xelse y = 2 * xout[tid] = yhardware builds the maskCPU3-14-15-92-6GPU3-14-15-92-6
1 / 6

The same job both ways: flip the negative numbers, double the others. Step through.

SIMD on a CPU versus SIMT on a GPU, for the same 8 numbers. The masked lanes look the same; in SIMD the program builds the mask, in SIMT the hardware does.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 4 correct
  1. Q1Dynamic power is about C·V²·f, and above a certain clock the voltage must rise with the clock. Why does that favor several slower cores over one fast one?

  2. Q2A program is 95% parallel. Using Amdahl’s law, about how fast does it run on 64 cores compared with 1?

  3. Q3What does simultaneous multithreading add to a core?

  4. Q4What is the main advantage of a vector-length-agnostic extension such as Arm SVE or RISC-V V?

Sources

Show Hide 28 sources
  1. The Landscape of Parallel Computing Research: A View from Berkeley (UCB/EECS-2006-183)Krste Asanović, Ras Bodik, Bryan Catanzaro, Joseph Gebis, Parry Husbands, Kurt Keutzer, David Patterson, William Plishker, John Shalf, Samuel Williams, Katherine Yelick · University of California, Berkeley · 2006The industry changed course in 2005 when Intel followed IBM’s Power4 and Sun’s Niagara to multiple cores; the ‘power wall’ (more transistors than power to turn them on), ‘memory wall’ and ‘ILP wall’ (diminishing returns on finding more ILP) add up to a ‘brick wall’; uniprocessor performance grew 52% a year from 1986 to 2002 and less than 20% a year after; leakage can be 40% of total power.
  2. The Free Lunch Is Over: A Fundamental Turn Toward Concurrency in SoftwareHerb Sutter · Dr. Dobb’s Journal 30(3) (author’s own copy, gotw.ca) · 2005Processor makers ‘have run out of room’ with clock speed and straight-line instruction throughput and are turning to hyperthreading and multicore; multicore on PowerPC and SPARC, coming from Intel and AMD in 2005; software must become concurrent to keep getting faster.
  3. Clock Rate versus IPC: The End of the Road for Conventional MicroarchitecturesVikas Agarwal, M. S. Hrishikesh, Stephen W. Keckler, Doug Burger · ISCA 2000 (author copy, University of Texas at Austin) · 2000Performance had grown 50–60% a year from faster clocks and higher IPC; with diminishing clock-rate gains and poor wire scaling, no scaling strategy for a conventional out-of-order core permits annual improvements better than 12.5% from 250 nm to 35 nm.
  4. Complexity-Effective Superscalar ProcessorsSubbarao Palacharla, Norman P. Jouppi, J. E. Smith · ISCA 1997 (University of Wisconsin–Madison digital repository) · 1997Rename, wakeup/select and bypass logic modeled against issue width and window size; rename and wakeup delays have quadratic components in issue width; a fully bypassed design needs 2·IW²·S bypass paths and bypass delay grows quadratically with issue width; window logic and bypasses are the critical structures for wide-issue machines.
  5. Simultaneous Multithreading: Maximizing On-Chip ParallelismDean M. Tullsen, Susan J. Eggers, Henry M. Levy · ISCA 1995 (author copy, University of California San Diego) · 1995An 8-issue superscalar sustains under 1.5 instructions per cycle (19% of issue slots); fine-grained multithreading uses only about 40% of a wide superscalar; SMT lets several threads issue in the same cycle and has the potential to reach 4× a superscalar’s throughput and 2× fine-grained multithreading’s; vertical versus horizontal waste.
  6. Reevaluating Amdahl’s LawJohn L. Gustafson · Communications of the ACM 31(5) (author’s own copy, johngustafson.net) · 1988Amdahl’s 1967 argument: serial fraction s caps speedup at 1/s, speedup = 1/(s + p/N); in practice problem size scales with processors, so scaled speedup = s′ + p′N; speedups of 1016–1021 on a 1024-processor hypercube for applications with 0.4–0.8% serial time.
  7. Amdahl’s Law in the Multicore EraMark D. Hill, Michael R. Marty · IEEE Computer (author copy, University of Wisconsin–Madison) · 2008A chip of n base-core equivalents (BCEs); a core built from r BCEs has perf(r) = √r; symmetric, asymmetric and dynamic multicore speedup formulas; asymmetric chips can beat symmetric ones (best 125.0 versus 51.2 at f = 0.975, n = 256).
  8. Dark Silicon and the End of Multicore ScalingHadi Esmaeilzadeh, Emily Blem, Renée St. Amant, Karthikeyan Sankaralingam, Doug Burger · ISCA 2011 (author copy, University of Wisconsin–Madison) · 2011Projected that 21% of a fixed-size chip must be powered off at 22 nm and more than 50% at 8 nm, and only 7.9× average multicore speedup through 2024 under ITRS scaling.
  9. Intel 64 and IA-32 Architectures Software Developer’s Manual, Volume 1: Basic Architecture (253665-084US)Intel · Intel · 2024SSE introduced with the Pentium III and SSE2 with the Pentium 4 (128-bit XMM registers); AVX adds 256-bit vectors; AVX-512 state is 32 512-bit ZMM registers and eight opmask registers k0–k7; Hyper-Threading makes one core appear as two or more logical processors, each with its own architectural state, sharing the execution engine; Turbo Boost converts thermal headroom into higher performance.
  10. Intel Advanced Vector Extensions 10.2 Architecture Specification (361050-007US)Intel · Intel · 2026AVX10 is the first major new vector ISA since AVX-512 was introduced in 2013; a converged vector ISA supported on all future processors, both P-cores and E-cores, with 32 vector registers and eight mask registers at vector lengths of 128, 256 and 512 bits.
  11. The ARM Scalable Vector ExtensionNigel Stephens, Stuart Biles, Matthias Boettcher, Jacob Eapen, Mbou Eyole, Giacomo Gabrielli, Matt Horsnell, Grigorios Magklis, Alejandro Martinez, Nathanael Premillieu, Alastair Reid, Alejandro Rico, Paul Walker · arXiv:1803.06185 (IEEE Micro 37(2), 2017) · 2017Advanced SIMD (NEON) has thirty-two 128-bit registers; SVE lets each implementation choose a vector length from 128 to 2048 bits in 128-bit steps, with a vector-length-agnostic programming model that runs unchanged on any length; 16 predicate registers, predicate-driven loop control (whilelt), first-fault loads; the DAXPY example.
  12. RISC-V “V” Vector Extension, Version 1.0RISC-V International · GitHub (riscv/riscv-v-spec release)Each hart (hardware thread) defines VLEN, the bits per vector register, a power of 2 no greater than 2^16; vsetvli takes the application vector length and returns how many elements the hardware will process (stripmining); the V extension for application processors requires VLEN ≥ 128; threads with live vector state cannot migrate between harts whose VLEN differs; the vector-vector add example.
  13. Intel Xeon Processor Scalable Family Technical OverviewDavid L. Mulnix (Intel) · Intel Developer Zone (technical article) · 2017, updated 2022As core counts rose, ring access latency increased and bandwidth per core diminished; the mesh lets traffic hop along a column then a row by a shortest path; up to 28 cores versus 22; 1 MB L2 per core and a shared, non-inclusive 1.375 MB LLC per core (versus 256 KB and 2.5 MB inclusive before); sub-NUMA clustering splits the socket into two domains; AVX-512 doubles the elements per FMA from 256 to 512 bits.
  14. Architecting for Flexibility and Value with Next Gen Intel Xeon Processors (Hot Chips 2023)Chris Gianos (Intel) · Hot Chips 35 · 2023Separate compute and I/O chiplets joined with EMIB; a logically monolithic mesh; the last-level cache is shared by all cores and can be partitioned into per-die sub-NUMA clusters; each core tile holds a core with its L2, an LLC slice with coherence agent, and a mesh router; P-core tiles have 2 MB L2 and a 4 MB LLC slice.
  15. The Next Generation of High Performance, Energy-Efficient Computing: Intel Xeon Processors Built on Efficient-Core (Hot Chips 2023)Don Soltis, Stephen Robinson (Intel) · Hot Chips 35 · 2023Modules of 2 or 4 cores share a 4 MB L2 and one frequency and voltage domain; each core is single threaded, for performance isolation; a mesh of core tiles with one shared LLC; up to 288 cores in two sockets; 205 W and higher per socket.
  16. AMD “Zen 4” EPYC Family Processor Architecture (Hot Chips 2023)Ravi Bhargava (AMD) · Hot Chips 35 · 2023AVX-512 on a 256-bit datapath, most 512-bit operations over two cycles; up to 50% fewer instructions than AVX2, lower power and more effective frequency in power-limited systems (+17% for HPL); SMT throughput uplift 34% on Zen 4 versus 25% on Zen 3; Zen 4c has the same IPC and features, lower maximum frequency and about 35% less core+L2 area, 16 cores per compute die instead of 8; compute dies around an I/O die.
  17. Arm Neoverse V2 Platform (Hot Chips 2023)Magnus Bruce (Arm) · Hot Chips 35 · 2023Quad 128-bit vector datapath; a typical 5 nm implementation with 2 MB L2 is 2.50 mm² and 1.4 W at 2.8 GHz and 0.75 V; the CMN-700 mesh interconnect supports up to 256 cores per die, up to 512 MB of system-level cache and up to 4 TB/s cross-sectional bandwidth.
  18. Arm Neoverse V2 Core Technical Reference Manual (102375_0002_03_en, r0p2)Arm · Arm documentation · 2022Implements SVE ‘with a 128-bit vector length’ and SVE2; ‘The Neoverse V2 core implements a scalable vector length of 128 bits’ (chapter 14).
  19. What is NUMA? (Linux kernel documentation)The Linux kernel developers · docs.kernel.org (open-source project docs)Memory access time and effective bandwidth depend on how far the cell containing the CPU is from the cell containing the memory; Linux maps cells to nodes and by default allocates from the local node, keeping accesses off the system interconnect.
  20. Capacity Aware Scheduling (Linux kernel documentation)The Linux kernel developers · docs.kernel.org (open-source project docs)Heterogeneous platforms have CPUs with different performance; CPU capacity is performance normalized to the most performant CPU, capacity = work per Hz × maximum frequency; Arm big.LITTLE big CPUs have more pipeline stages, bigger caches and smarter predictors and reach higher operating points.
  21. Energy Aware Scheduling (Linux kernel documentation)The Linux kernel developers · docs.kernel.org (open-source project docs)EAS works only on heterogeneous CPU topologies such as Arm big.LITTLE and picks an energy-efficient CPU for each task with minimal impact on throughput, using an energy model of each CPU’s power at each performance state.
  22. Single-ISA Heterogeneous Multi-Core Architectures: The Potential for Processor Power ReductionRakesh Kumar, Keith I. Farkas, Norman P. Jouppi, Parthasarathy Ranganathan, Dean M. Tullsen · MICRO-36 (author copy, University of California San Diego) · 2003Cores at different power/performance points running the same ISA, with system software choosing among them; 39% average energy reduction for 3% performance loss across 14 SPEC benchmarks, beyond what chip-wide voltage and frequency scaling achieves.
  23. Understanding Chiplets Today to Anticipate Future Integration Opportunities and LimitsGabriel H. Loh, Samuel Naffziger, Kevin Lepak (AMD) · Design, Automation and Test in Europe (DATE) 2021, proceedings archive · 2021First-generation EPYC: four 8-core chiplets, 32 cores, estimated 41% cheaper than a monolithic equivalent; second generation: up to eight 8-core compute dies (CCDs) on 7 nm around a 12 nm I/O die, giving 16- to 64-core products from two chip designs.
  24. AMD Launches 5th Gen AMD EPYC CPUs (press release)AMD · AMD investor relations · 2024October 10, 2024. EPYC 9965: 192 Zen 5c cores, 2.25 GHz base, 3.7 GHz boost, 500 W, 384 MB L3; EPYC 9755: 128 cores, 2.7/4.1 GHz, 500 W; EPYC 9575F: 64 Zen 5 cores, 3.3/5.0 GHz, 400 W, 256 MB L3; AVX-512 with the full 512-bit data path across the series; 12 channels of DDR5 memory per CPU.
  25. NVIDIA Grace CPU Superchip Architecture In Depth (NVIDIA Technical Blog)Jonathon Evans, Ian Finder, Ivan Goldwasser, John Linford, Vishal Mehta, Daniel Ruiz, Mathias Wagner · NVIDIA Developer Technical Blog · 202372 Arm Neoverse V2 cores per Grace CPU, each with 4×128-bit SVE2; 117 MB of L3 per CPU on the Scalable Coherency Fabric, a mesh with distributed cache and over 3.2 TB/s of bisection bandwidth; 500 W for the two-CPU superchip including memory.
  26. Is RISC-V Ready for HPC Prime-Time: Evaluating the 64-Core Sophon SG2042 RISC-V CPUNick Brown, Maurice Jamieson, Joseph Lee, Paul Wang · arXiv:2309.00381 · 202364 XuanTie C920 cores at 2 GHz in clusters of four sharing 1 MB of L2, 64 MB of L3 shared by all cores; RVV v0.7.1 at 128 bits, needing a vendor compiler fork; four NUMA regions, and thread placement that cycles across NUMA regions and clusters matters for performance.
  27. CUDA C++ Programming Guide, v12.4 (SIMT architecture)NVIDIA · NVIDIA CUDA Toolkit documentation archive · 2024SIMD vector organizations expose the SIMD width to software, whereas SIMT instructions specify the execution and branching behavior of a single thread; warps of 32 threads.