One AI chip, however fast, can’t train a modern AI model alone. The model is too big. So it’s split across many chips. The catch is that the chips must keep talking. After almost every chunk of math, each chip needs to hear what the others found.
A is a small team of chips joined by their own very fast, very short links. It can be a handful of chips or a whole rack of them. Inside the team, chips swap data so quickly that together they act a lot like one giant chip. A slower network joins many teams together. That’s the next chapter.
This chapter looks at how a team is wired, and at its most common job: everyone adding up their numbers so that everyone gets the total.
Large AI models are split across many accelerators, the specialized chips such as GPUs and TPUs that do the math. Some ways of splitting create heavy, frequent traffic. In , for example, every layer of a transformer model is divided among several chips, and they must combine their partial results four times per layer on each training step.1 That traffic sits directly in the path of the computation: until it finishes, nobody can move on.
A is a group of accelerators (8, 16, 64 or 72 in current systems) joined by a dedicated fabric with far more bandwidth per chip than the ordinary datacenter network. Chips in the domain can read and write each other’s memory directly and run group operations, called , quickly. The (next chapter) then connects many such domains. Training systems put the chattiest kind of splitting inside the domain and the more tolerant kinds across the scale-out network.2
This chapter covers:
- why accelerators need links of their own instead of the server’s general-purpose bus;
- direct (point-to-point) versus switched wiring, and what each does to bandwidth and latency;
- memory semantics: treating another chip’s memory almost like your own;
- collectives such as , and the algorithms that run them;
- rack-scale domains, and why copper sets their size.
A scale-up domain is the set of accelerators that share a memory-semantic fabric with roughly an order of magnitude more bandwidth per chip than the scale-out NIC. It exists because some forms of parallelism put communication on the critical path of every layer. Megatron-style needs two all-reduces in the forward pass and two in the backward pass of each transformer layer,1 and in practice it stays within the high-bandwidth domain, because across slower inter-server links its overhead becomes impractical.2 The domain’s size therefore bounds the tensor- (and expert-) parallel degree; the mapping of parallelism onto the hierarchy is the subject of Why the network looks this way.
The design questions this chapter works through:
- Topology. How to divide a fixed budget: among direct neighbors (ring, , ) or across parallel planes, and what each does to per-pair bandwidth, hop count and domain size.
- Semantics. Load/store/atomic access to peer memory, its ordering model, latency targets and link-level reliability.
- Collectives. The , the bandwidth lower bound for all-reduce, and which algorithm (ring, tree, halving-doubling, direct, dimension-wise, in-network) fits which topology and message size.
- Physical reach. Why current domains stop at one to a few racks of copper, and what open specifications target next.
Place a tensor-parallel group inside one domain or across two, and watch the traffic. Tap any part for details.
Every AI chip in a server already has a link to the server’s main processor. It’s the same kind of slot a graphics card plugs into, called . So why not let the chips talk through it?
Because it’s far too slow for this job. On one recent AI chip, the team links carry about ten times more data per second than its PCIe slot.3 On another chip, they carry about fourteen times more.4 Using the slot would be like serving a whole dinner rush through one tiny window.
The general-purpose bus is too narrow
In a server, accelerators connect to the host CPU through , the general-purpose expansion bus covered in Board and server. PCIe is built for feeding a device from the host, not for heavy device-to-device traffic. Two published numbers show the gap:
- One accelerator offers 1,200 GB/s of chip-to-chip networking (counting both directions) but 128 GB/s on its PCIe Gen5 x16 host link, about a ninth.3
- A current GPU has 18 dedicated links totaling 1.8 TB/s (both directions), roughly fourteen times that same Gen5 x16 figure.43
Before dedicated links, GPU-to-GPU traffic shared the same PCIe tree as CPU-to-GPU traffic, and that tree was often the bottleneck.5
The traffic sits on the critical path
In tensor parallelism, each layer’s matrices are split across, say, eight chips. Each chip multiplies its share, and then the chips must sum their partial results so every chip has the complete output before the next layer starts. That sum is an , and a transformer layer needs four of them per training step (two going forward, two going backward).1 With dozens of layers and many steps per second, the chips spend real time waiting unless the links are fast.
That is why a training job is usually arranged so that tensor parallelism stays inside a server’s fast fabric, and only less chatty kinds of splitting cross the slower network between servers.2
Scale-up versus scale-out
| Scale-up fabric | Scale-out network | |
|---|---|---|
| Reach | One server to a few racks | A whole datacenter |
| Chips | Tens, up to about a thousand in open specifications | Thousands to hundreds of thousands |
| Bandwidth per chip | Highest | Several times lower |
| How chips talk | Direct memory reads and writes | Messages through a network card |
Bandwidth ratios
The case for a separate fabric is mostly arithmetic. Published per-chip figures (all bidirectional) put scale-up an order of magnitude above the host link:
- Intel Gaudi 3: 1,200 GB/s of integrated Ethernet versus 128 GB/s of PCIe Gen5 x16, against 3.67 TB/s of HBM.3
- NVIDIA NVLink: 600 GB/s (12 links) for Ampere, 900 GB/s (18 links) for Hopper and 1,800 GB/s (18 links of 2 lanes at 200 Gb/s PAM4) for Blackwell.4
- Google TPU v4: six inter-chip links at 50 GB/s, against 1,200 GB/s of HBM.6
Even scale-up links remain several times slower than local , which is why the parallelism layout tries to keep each chip’s working set local and exchange only activations or gradients.
Why not PCIe peer-to-peer?
PCIe does support device-to-device transfers, but in a host-rooted tree they cross shared switches and, for devices under different sockets, the CPU interconnect. Measurements on 8-GPU servers found PCIe peer bandwidth limited by switch sharing, and the bus a common bottleneck when GPUs communicate.5 A dedicated fabric adds bandwidth that scales with the number of accelerators rather than with host lanes, and lets the fabric protocol be tuned for small memory transactions and collectives.
Communication on the critical path
Tensor-parallel all-reduce volume per layer is proportional to batch × sequence × hidden size and does not shrink as the tensor-parallel degree grows, while the compute per chip does. So the communication fraction rises with the degree, and on slow links it dominates. That is the quantitative reason Megatron-LM keeps tensor parallelism within the server and uses pipeline parallelism across servers.2 Collectives were already a major cost in classic HPC: a five-year study on a Cray T3E found over 40% of MPI time in allreduce and reduce.7
Send data from accelerator 1 to accelerator 4: through the host’s PCIe tree, or over dedicated links.
There are two basic ways to wire a team of chips.
Direct links
Run cables straight from chip to chip. With 8 chips, each one can have its own cable to all 7 others.8 But with 72 chips, each would need 71 cables, and a recent chip has only 18 link plugs.4
So big teams with direct links use simpler patterns. In a ring, each chip talks only to its two neighbors. In a grid, each chip talks to the chips beside, above and below it.
Through a switch
Or plug every chip into a , a hub chip that can connect any chip to any other. Now every pair of chips is the same short trip apart. The price is the switches themselves: more chips to build, power and cool.
The pattern of connections is the . A key constraint shapes it: each chip has a fixed number of high-speed lanes, set by how many circuits fit along the edge of the die and in its power budget. Topology design is mostly about how to divide those lanes.
Point-to-point topologies
- Full mesh. Every chip links directly to every other. With chips, each chip splits its lanes ways. AMD’s MI300X gives each GPU seven 16-lane links, one to each of the seven other GPUs in an eight-GPU server.8 Intel’s Gaudi 3 servers wire every pair of their eight cards together with Ethernet ports: 21 of each chip’s 24 ports go to the other seven (three each), and the remaining three go to the scale-out network.93 Every pair is one hop apart, but each pair only gets a slice of a chip’s bandwidth, and the pattern runs out of ports quickly as grows.
- Ring. Each chip links to two neighbors. Cheap and scales to any size, but data between distant chips passes through many others.
- Torus. Chips sit on a 2D or 3D grid, each linked to its 4 or 6 neighbors, with wraparound links joining opposite edges. AWS connects 16 Trainium2 chips as a 4 × 4 2D torus;10 Google’s TPU v4 builds a 3D torus of up to 4,096 chips.6
Irregular point-to-point layouts create uneven performance. In one earlier 8-GPU server, GPUs not directly linked had to relay data through a third GPU, which doubled or tripled latency compared with a direct link.5
Switched topologies
In a switched design, each chip spreads its links across several switch chips (parallel “planes”), and every switch connects to every chip. An early switched NVIDIA server, the DGX-2, joined 16 GPUs so that any GPU could reach any other at the same 300 GB/s, with nearly identical latency for every pair.5 Its current rack design joins 72 GPUs through 18 switch chips, each GPU sending one link to each switch.4
Switches cost extra chips, power and a little latency per trip, but they make every pair equal, let the whole bandwidth go to any one destination, and can even add up data as it passes through, a trick called .
Dividing a lane budget
Treat each chip as having ports of bandwidth per direction (per-chip ). The topology decides how is spent:
| Topology | Ports per neighbor | Per-pair bandwidth | Hops (max) | Size limit |
|---|---|---|---|---|
| Ring | to each of 2 | to a neighbor | none (latency grows) | |
| Full mesh | to each of | 1 | ||
| 2D torus () | to each of 4 | to a neighbor | none (hops grow as ) | |
| Switched, one level | 1 to each of switch planes | to any one peer | 1 switch |
The full mesh’s per-pair bandwidth falls as , and its size is capped by port count: MI300X’s seven Infinity Fabric links (x16 at up to 32 Gb/s per lane) make exactly an eight-GPU mesh, with a quoted 896 GB/s peak peer-to-peer aggregate.8 A one-level switched domain is capped by switch radix instead: NVIDIA’s 72-port NVLink Switch chip sets the 72-GPU NVL72 domain, with each GPU’s 18 links going one to each of 18 switch chips.4 Broadcom’s Scale-Up Ethernet specification sketches the same plane structure with Ethernet switches: 64 accelerators, each with twelve 800G interfaces, through twelve 64-port switches, giving any pair up to 9.6 Tb/s.11
Uniformity and routing
Point-to-point fabrics without forwarding expose topology to software. On the P100/V100 DGX-1 “hybrid cube mesh,” NVLink was not self-routed: traffic between GPUs without a direct link had to be relayed through an intermediate GPU chosen by the user, and measured latency rose from about 9 µs to roughly 2× (P100) or 3× (V100).5 The same study measured the switched DGX-2 as homogeneous: every pair reached 300 GB/s through NVSwitch, and the extra hop between baseboards cost almost nothing in latency.5 Irregularity also hurts collectives: when a scheduler hands a job an odd subset of an 8-GPU server, ring-based libraries leave some links unused, and packing spanning trees instead was up to 8× faster.12
Torus properties
A torus keeps degree constant, so it scales with short, cheap cables, but its grows only as the cut surface. Wraparound links matter: TPU v4 reports that they double both bisection bandwidth and the bandwidth of collectives such as all-reduce compared with a mesh, and it chose a 3D torus over its predecessor’s 2D torus partly for bisection.6 Collectives map well onto tori when they run dimension by dimension (see “Under the hood”); arbitrary all-to-all traffic does not, because it crosses many hops and shares links.
Each chip has 16 links. Pick a topology, then tap a chip to route data to it from chip 1.
Most computer networks work like the mail. You pack data into a message, address it and send it. Team links can work more like a shared whiteboard. A chip can read a number straight from a teammate’s memory, or write one into it, just like using its own memory.
That makes some jobs much simpler. A chip that needs one small number from a teammate just reaches over and grabs it. Some systems even describe a team of chips as sharing one big pool of memory.10
The hard part is the wait. Each reach across a link takes about a millionth of a second.13 That sounds tiny, but to a chip it’s a long time. So a chip keeps thousands of reads and writes going at once, like a kitchen cooking many orders at the same time.
Loads and stores instead of messages
A scale-out network is usually message-based: software asks a network card to send a block of data, and the card on the other side places it in memory. Scale-up fabrics usually offer instead. An accelerator issues ordinary memory operations to an address that happens to live on another accelerator:
- Read (load): fetch data from a peer’s memory.
- Write (store): place data directly in a peer’s memory.
- Atomic: read-modify-write a value in one indivisible step, for example add 1 to a counter, so chips can coordinate without races.
NVLink has supported direct reads, writes and atomics on peer GPU memory since its first generation.5 The open UALink specification is built around the same idea: direct read, write and atomic transactions, with the same ordering rules whether the memory is local, on the host or on a remote accelerator.13 Broadcom’s Scale-Up Ethernet framework likewise targets “one-sided” memory loads, stores and atomics, meaning the software on the remote chip doesn’t have to take part (the sender still gets an acknowledgment).11 AWS describes its 16-chip Trainium2 instances as pooling the memory of all 16 chips.10
Latency matters more
Because memory operations are small (a few hundred bytes each), the fabric has to be quick. UALink targets a request-to-response round trip under 1 µs;13 Scale-Up Ethernet targets under 2 µs.11 Even so, a chip must keep many requests outstanding to fill a fast link. By , data in flight equals bandwidth × latency: to read at 800 GB/s with a 1 µs round trip, 800 KB must be in flight at all times, about 3,000 requests of 256 bytes.
Reliability at the link
A dropped memory write is not something software can easily notice and resend, so these fabrics fix errors at each link. NVLink and NVSwitch check every transfer with a CRC (a checksum) and replay it when it fails.5 UALink specifies and credit-based flow control, in which a sender only transmits when the receiver has advertised free buffer space, so packets are never dropped for lack of room.13
Semantics and ordering
A memory-semantic fabric carries read, write and atomic transactions with an ordering model, rather than NIC-mediated messages. UALink 1.0 defines exactly that, with the same ordering model for host-attached, local and remote accelerator memory, a 64-byte transaction-layer flit packed into 640-byte data-link flits, and address compression that reaches about 95% protocol efficiency for 256-byte writes (20 of 21 flits carry data in its example).13 Scale-Up Ethernet keeps the transport generic (XPU-specific command/response transactions used for put, get and atomics) and lists PGAS-style shared memory as the memory architecture, leaving registration and coherence services outside its scope.11 The UALink white paper does not describe hardware across the domain either. Collective libraries instead order their transfers with memory fences and flags; NCCL’s protocols differ mainly in which of the two they use.14
Little’s law sets the queue depth
. At and , 800 KB must be in flight: 3,125 requests of 256 B. Halve the latency and the required depth halves. This is why scale-up specifications state latency targets in hundreds of nanoseconds to microseconds (UALink: under 1 µs request-to-response;13 SUE: under 2 µs round trip11) and why they shave FEC latency, for example UALink’s reduced codeword interleaving, which trades burst-error correction for latency.13
Reliability and flow control
Memory traffic has no end-to-end retransmission in software, so reliability is pushed down a layer. NVLink applies CRC with replay on each link, ECC on datapaths and routing state, and a fabric manager that guards routing tables against out-of-range access.5 UALink combines Ethernet-PHY FEC, a 32-bit CRC per 640-byte flit, and credit-based flow control, plus partitioning of a pod into isolated virtual pods.13 Lossless credit flow avoids drops but means congestion backs up hop by hop, so incast (many senders, one receiver) must be handled without head-of-line blocking, which SUE lists among its requirements.11
1. B’s software copies the value into a send buffer and asks its network card to send it.
100 reads in flight × 256 B ÷ 1 µs = 26 GB/s, 3% of the link. Filling it takes 3,125.
The job these links do most is a team effort where every chip takes part at once. The most important one is called an . Every chip starts with its own list of numbers. At the end, every chip holds the sum of everyone’s lists.
The ring trick
The obvious way is for every chip to send its whole list to every other chip. That floods the links. A classic trick is to pass the numbers around a ring instead, adding as they go.
Two things set the time
Every pass costs two things. There’s a fixed delay just to get started, however little you send. Then there’s the time to move the data, which grows with the amount you send.
With a little data, the delays matter most. A ring of many chips has many passes, so it’s slow. With lots of data, link speed matters most. Smart software picks a different method for each case.7
The common collectives
Libraries such as NCCL (for GPUs) and MPI (for supercomputers) provide a standard set of .15 For chips each holding a buffer:
- Broadcast: one chip’s buffer is copied to all.
- Reduce: buffers are summed element by element; one chip gets the result.
- : summed, and every chip gets the result. Used to add up gradients in training and partial sums in tensor parallelism.
- : summed, but each chip gets a different slice of the result.
- : each chip contributes a block; every chip ends with all blocks.
- : every chip sends a different block to every other chip.
A useful identity: a reduce-scatter followed by an all-gather is an all-reduce.15 The best all-reduce algorithms are built exactly that way.
Ring all-reduce, step by step
Take chips, each with a buffer cut into 4 slices (A, B, C, D):
- Reduce-scatter, 3 steps. In each step, every chip sends one slice to its right-hand neighbor, which adds its own copy of that slice. After steps, each chip holds one slice that contains all 4 chips’ contributions.
- All-gather, 3 more steps. Each chip passes its finished slice to the right; after 3 more steps every chip has all 4 finished slices.
Each chip sent slices of 1/4 of the buffer: 1.5 times the buffer. In general a chip sends of the buffer, which approaches 2 for large and does not grow with the number of chips. That is the least any algorithm can manage when the chips do all the adding themselves.16
A simple cost model
The says sending bytes costs : a fixed latency per step plus the bytes divided by bandwidth .7 For the ring, adding up its steps gives:
A worked example, using the simulator’s defaults: 8 chips, per step, per chip.
- 100 MB buffer: latency ; data ; total . Mostly moving data: bandwidth-bound.
- 10 KB buffer: latency still 28 µs; data about 0.02 µs. Almost all waiting: latency-bound.
Because the latency term grows with , a ring of 72 chips pays 142 steps of latency on every all-reduce. For small messages, libraries switch to algorithms with fewer steps, such as trees, whose step count grows only with .717
Definitions and the decomposition
NCCL’s semantics: AllReduce leaves the element-wise reduction on every rank; ReduceScatter leaves of it per rank; AllGather concatenates contributions on every rank; AlltoAll sends chunk of rank ’s input to rank .15 ReduceScatter + AllGather = AllReduce, and that decomposition is what makes all-reduce bandwidth-optimal.
The model and the lower bound
Thakur, Rabenseifner and Gropp model a message as ( per byte), with per byte for the local reduction, assuming full-duplex links and single-ported nodes.7 Patarasuk and Yuan prove that, without compression, a one-item all-reduce on processes needs at least partial results communicated, so some process must move at least of the data; a contention-free logical ring achieves it.16 The ring’s cost is then
(the paper’s form for its ring allreduce).7 Its bandwidth term is optimal; its latency term is the worst of the common algorithms. Equating the two and terms (ignoring ) gives a crossover at : below it the ring is latency-bound. With , and , , which is larger than many tensor-parallel messages.
Choosing by message size
MPICH switches algorithms on size: recursive doubling ( steps, but bytes) for short all-reduces and Rabenseifner’s reduce-scatter + all-gather ( steps, bytes) for long ones.7 NCCL makes the same trade with both algorithms (ring and double binary tree) and protocols: its LL protocol has about 1 µs per hop but reaches only 25–50% of peak bandwidth, LL128 about 2 µs and roughly 95%, and Simple about 6 µs and near peak.14 Ring latency grows linearly with GPU count; double binary trees keep full bandwidth with logarithmic latency, which NVIDIA measured as up to 180× lower latency at 24,576 GPUs.17
Reporting: algbw and busbw
nccl-tests reports algorithm bandwidth and a “bus bandwidth” that multiplies it by for AllReduce and for ReduceScatter and AllGather, so that a bandwidth-optimal implementation’s busbw approaches the per-rank link bandwidth regardless of rank count.18 Comparing busbw with the hardware’s per-chip bandwidth is the quickest sanity check of a collective’s efficiency.
All-reduce: summed, and every chip gets the result. Training uses it to add up gradients and partial sums. Showing the inputs; press After.
For years, a tight team was 8 chips on one board. Today’s biggest switched teams fill a whole rack.19 Grid-style designs go further and join whole racks of chips.6
What sets the limit is the wire. At these speeds, a copper cable carries a clean signal for only a few meters. Going farther means turning the signal into light and sending it down a glass fiber. That costs far more money and power. One company found these light links cost over ten times as much as copper ones.6 So designers squeeze as many chips as they can within copper’s reach, which is about a rack.
From one board to one rack
On NVIDIA’s previous HGX H200 baseboard, an NVLink domain was eight GPUs. The GB200 NVL72 design stretches it across a rack: 18 compute trays holding 72 GPUs and nine switch trays, joined by four NVLink “cartridges” of more than 5,000 copper cables mounted at the back of the rack.19 Google’s TPU v4 uses passive electrical cables to wire 64 chips into a 4 × 4 × 4 block within each rack.6 AWS joins four 16-chip Trainium2 servers into a 64-chip “UltraServer.”10
Why copper, and why a rack
Copper cables need no lasers or light detectors at their ends, which keeps them cheap and low-power, but at today’s lane speeds they are short: the UALink specification targets cables under 4 meters and domains of one to four racks.13 Optical links reach much farther but need converters at each end. When TPU v3’s 2D torus grew across many racks, some of its wraparound links became so long that they had to be optical, and optical links cost over ten times as much as electrical ones.6 The trade-offs between copper and the various kinds of optics are the subject of Optics.
The cost of density
Packing 72 high-power chips into one rack to keep their links short drives up power and heat. The NVL72 rack needs a bus bar for 1,400 amps, over 100 pounds of added steel to withstand 6,000 pounds of connector mating force, and direct liquid cooling for about 120 kW.19 The consequences for the facility are covered in Power and cooling.
Reach sets the domain
At 100–200 Gb/s per lane, electrical reach is short, and the specifications say so: UALink, built on IEEE 802.3 200G-per-lane electrical PHYs (KR/CR), targets cables under 4 m, 1–4 racks and up to 1,024 endpoints;13 Scale-Up Ethernet allows up to 10 m from accelerator to switch.11 Past that, optics enter: TPU v3’s 2D-torus wraparound links between distant racks had to be optical, at over 10× the cost of electrical links, which pushed TPU v4 to electrically cabled blocks per rack (passive cables) with optical circuit switches only between blocks.6
Anatomy of a copper rack-scale domain
NVL72: 72 GPUs, each with 18 NVLink5 links (2 lanes at 200 Gb/s PAM4, 100 GB/s bidirectional per link), one link to each of 18 NVLink Switch chips of 72 ports; two switch chips per switch tray, nine trays.4 That is GPU-side links, carried by four cartridges of more than 5,000 copper cables.19 The single switch level is set by radix: 72 ports per switch chip, 72 GPUs. The same slides show the domain growing from an 8-GPU hybrid cube mesh (2016) to an 8-GPU switched baseboard (2022) to 72 GPUs (2024).4
Torus designs reach rack scale differently. TPU v4 builds 4 × 4 × 4 blocks with passive electrical cables; each block exposes 16 links on each of its 6 faces (96 optical links) to 48 optical circuit switches, which assemble up to 64 blocks (4,096 chips) and can route around failed blocks.6 Trn2 uses a 4 × 4 2D torus inside a server at 1,024 GB/s per chip and rings between four servers at 256 GB/s per chip, a 4:1 bandwidth step at the server boundary.10
1 rack: longest cable ≈ 2.0 m, under UALink’s 4 m copper target. Cheap, low-power, no lasers.
A rack-scale domain. Tap a part to read about it, or trace one compute tray’s links.
Pick a wiring pattern along the top, set the number of chips and press Play to watch an all-reduce. Amber arrows carry numbers that get added when they arrive. Blue arrows hand out the finished totals. On small teams, the squares show what each chip holds.
Then slide “Amount of data” from small to large. Watch the bar flip from mostly waiting to mostly moving data.
Each chip has 16 links. Choose a topology and the simulator divides those links among neighbors, peers or switch planes, picks the fastest available algorithm, and splits the all-reduce time into its latency part () and bandwidth part (). The chart plots time against message size for all four topologies; the shaded region is where the current setup is latency-bound. Try a 64-chip ring against a 64-chip 2D torus at 10 KB, and a full mesh above 17 chips.
An model over a 16-port chip ( per direction). Pick the algorithm per topology: ring (two counter-rotating rings), binomial tree, halving-doubling (with whole-buffer fold-in for non-powers of two), direct one-shot on a full mesh, dimension-wise 2D rings on the torus, and, on the switched fabric, in-network reduction. The chart shows every algorithm on the current topology plus the endpoint bandwidth floor ; the readouts give busbw and the crossover size. Check that for the ring, that a single ring on a full mesh wastes most links, and that in-switch reduction beats the floor.
Here is what the numbers below mean in everyday terms.
- Chips in a team: from 8 in one server up to 72 in one rack. One design joins 64-chip blocks into teams of thousands.
- Speed per chip: hundreds of billions to nearly two trillion bytes a second. That’s many times more than each chip’s link to the main processor.
- Wiring: some teams use direct cables, some use switches and some use a grid. No single answer has won yet.
- Largest one-level switched GPU domain (NVL72)
- 72 GPUs
- Copper cables in that rack’s NVLink spine
- 5,000+
- UALink 1.0 target: accelerators per pod
- 1,024
- UALink 1.0 target: request-to-response
- < 1 µs
Reading the table
- Three wiring strategies coexist. Switched domains (NVL72, UALink, SUE) give every pair the same path; full meshes (MI300X, Gaudi 3) skip the switch but stop at 8 chips; tori (TPU, Trainium) use short neighbor links and reach large sizes with more hops.
- Ethernet is moving into scale-up. Gaudi 3 already uses standard Ethernet ports for its in-server mesh.9 UALink reuses Ethernet’s physical layer, and with it existing cables, connectors and retimers,13 while Scale-Up Ethernet runs its transport over Ethernet switches.11
- Memory semantics are the norm. NVLink, UALink and Scale-Up Ethernet all describe direct memory reads, writes and atomics, not just messages.51311
Derived checks
- NVL72 aggregate: , matching the quoted 130 TB/s all-to-all; the quoted 260 TB/s all-reduce is twice that, consistent with in-switch reduction halving per-GPU all-reduce traffic.4
- MI300X: x16 at 32 Gb/s = 64 GB/s per direction per link; , the quoted peer-to-peer figure. One peer gets one link: 64 GB/s each way.8
- Gaudi 3: , the quoted node scale-up bandwidth; one peer gets each way.39
- TPU v4: of ICI against 1,200 GB/s of HBM, a 4:1 ratio.6
Bigger teams are harder
A bigger team lets a model be split more ways. But every extra chip adds cables, power, heat and parts that can break. When one chip in a tight team fails, the whole team’s job may have to stop.
Switch or no switch
Skipping the switch saves money and power, and works well for 8 chips. Past that, direct wiring runs out of plugs, or data has to hop through many chips. Switches fix that but cost more.
Waiting versus moving
No single way of adding up numbers is best. Lots of small messages want the fewest hops. Huge messages want every link busy. Good software picks a method for each case.7
Direct versus switched
| Direct (mesh, ring, torus) | Switched | |
|---|---|---|
| Extra chips | None | Switch chips, trays and their power |
| Bandwidth between any two chips | A slice (mesh) or shared over hops (ring, torus) | A chip’s full bandwidth |
| Distance | 1 hop (mesh) to many hops (ring) | Always through one switch |
| Size limit | Ports per chip (mesh) or latency (ring) | Ports per switch chip |
| Extras | — | Can add up data inside the switch |
What goes wrong
- Uneven paths. In irregular direct topologies, chips that aren’t neighbors pay extra latency, and performance depends on which chips a job happens to get.512
- Idle links. An algorithm that doesn’t match the wiring, like a single ring on a full mesh, leaves most links unused.12
- Latency at scale. Ring algorithms add a step per chip, so large rings become slow for all but huge messages.17
- Failures. A bigger domain has more parts that can fail mid-job. TPU v4’s optical switches let the system route around failed blocks rather than lose the whole machine.6
- Power and heat. Short copper means dense racks: about 120 kW and liquid cooling for one 72-GPU rack.19
Topology trade-offs, quantified
- Mesh: no switch, shrinking pairs. Per-pair bandwidth . Collectives that use every link at once (direct reduce-scatter/all-gather) still reach near the floor, but point-to-point traffic between two chips (pipeline hand-offs, KV-cache moves) sees only one link: 64 GB/s per direction on MI300X.8
- Switched: uniform, but a hop and a radix. Any chip can push to any one peer, and in-network reduction becomes possible. Costs: switch silicon, trays, power and one switch traversal per transfer, and a one-level domain capped at the switch radix (72 ports → 72 GPUs).4
- Torus: cheap cables, latency. Neighbor bandwidth (2D) or (3D); dimension-wise collectives recover full bandwidth, but all-to-all and irregular traffic share links over multiple hops. Wraparound links double bisection compared with a mesh but are the longest cables.6
Failure modes in practice
- Topology-blind allocation. Schedulers hand jobs fragments of a server (3, 5, 6 or 7 of 8 GPUs are common), and ring libraries then strand links; spanning-tree packing recovers up to 8×.12
- Relayed traffic. Without hardware routing, non-adjacent pairs relay through a peer at 2–3× latency.5
- Wrong regime. A bandwidth-tuned protocol on small messages pays ~6 µs per hop; a latency-tuned one on large messages wastes 50–75% of bandwidth (NCCL LL).14
- Incast and backpressure. Lossless, credit-based fabrics never drop, so congestion propagates; scale-up transports must spread traffic across planes and avoid head-of-line blocking.11
- Blast radius. Larger domains raise the chance that some component fails during a job; reconfigurable optics let TPU v4 skip failed blocks and schedule around unavailable hosts.6
8 chips: a full mesh gives each pair 2 links, one hop, with no switch. Rings and grids are still short.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Ring all-reduce
Split the buffer into chunks. In reduce-scatter step (), rank sends chunk to rank , which adds it to its own copy. After steps rank holds the fully reduced chunk . All-gather repeats the pattern without the addition. Each step moves bytes per rank on one link in each direction:7
The bandwidth term meets the Patarasuk–Yuan lower bound for algorithms in which only endpoints reduce.16 Running two rings in opposite directions on half the buffer each uses both directions of every link, which is how the simulator charges a physical ring.
Trees and recursive algorithms
- Binomial tree reduce + broadcast: steps each way, but the whole buffer every step: (plus on the reduce).7 Latency-optimal, bandwidth-poor.
- Recursive doubling: , MPICH’s short-message allreduce.7
- Rabenseifner (halving-doubling): reduce-scatter by recursive vector halving and distance doubling, then all-gather by vector doubling and distance halving: . Optimal in both terms for powers of two, when the network carries the butterfly pattern without contention.716 For other , MPICH first folds the extra ranks into neighbors with a half-buffer exchange and sends them the result at the end.7 (The simulator uses a simpler whole-buffer fold: 2 extra steps and extra bytes.)
- Double binary trees: two complementary binary trees, arranged so that no rank is interior in both; pipelined, they approach full bandwidth with logarithmic latency. NCCL uses them alongside rings.1417 The simulator’s “Tree” is the unpipelined binomial version, to show the classic latency-for-bandwidth trade.
Direct (one-shot) on a full mesh
With links per chip, reduce-scatter can be a single step: every rank sends chunk to rank over the dedicated link, and all transfers proceed at once; all-gather mirrors it. With per-peer bandwidth ( ports per peer):
Since , this is within rounding of the floor, with only two steps. A single ring on the same mesh uses 2 of the links, so its bandwidth term is times worse, the mismatch the simulator’s expert view exposes, and the reason libraries pack many trees or rings onto point-to-point graphs.12 (This formula is our derivation under the model; real one-shot kernels also pay for concurrent streams per chip.)
One bidirectional ring drives 2 of each chip’s 7 links: 14 steps of α, and a bandwidth term (p − 1)/2 = 3.5× the floor.
Dimension-wise algorithms on a torus
Mikami et al.’s 2D-torus all-reduce on an grid does reduce-scatter along rows, all-reduce along columns on the -sized result, and all-gather along rows; they count GPU-to-GPU operations against for a single ring over all GPUs.20 Ying et al. found TPU v3’s 1-D ring latency-limited at pod scale and ran two concurrent rings on halves of the payload, one along and one along , so both dimensions’ links stay busy, using wraparound links bidirectionally; they report twice the throughput of the 1-D algorithm.21 For an torus with neighbor bandwidth , a half that goes rows first costs
and the simulator takes the slower of the two halves. On 8 × 8 the bandwidth term equals a 64-chip ring’s (about ) but the latency term is instead of . A snake-shaped single ring on the same torus touches only 2 of 4 neighbor links and pays both the and twice the bandwidth term.
In-network reduction
With , each endpoint sends its bytes into the switch planes once and receives the reduced bytes once. With full-duplex links and pipelining, , below the endpoint floor by nearly 2×. NCCL’s NVLS algorithms use NVLink SHARP in NVSwitch for exactly this, and its CollNet algorithms use SHARP-capable network switches.14 The NVL72 switch chip lists 3.6 TFLOPS of in-network compute, and the system’s quoted all-reduce bandwidth (260 TB/s) is twice its all-to-all bandwidth (130 TB/s).4 The catch: arithmetic precision and format support are fixed by the switch, and the reduction needs buffer space and synchronization inside it.
Model assumptions
The model ignores contention, assumes every rank sends and receives at once, treats latency as fixed per step, and charges nothing for chunking, protocol overhead or the reduction arithmetic ().7 Real libraries tune chunk sizes, channels and protocols, and the crossover sizes they choose are measured, not derived: MPICH, for example, set a 2 KB short/long boundary for reduce experimentally.7 Use the model to reason about scaling and regimes, and busbw from nccl-tests to measure the real thing.18
Q1Eight chips are wired as a full mesh, and each chip has 16 links in total. How many links can go to each other chip?
Q2A ring all-reduce runs on 8 chips. Roughly how much data does each chip send, compared with the size of the buffer being summed?
Q3What does it mean that a scale-up fabric has “memory semantics”?
Q4Why does a 2D torus all-reduce on 64 chips take far fewer steps than a single ring through all 64?
Sources
Show Hide 21 sources
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismTensor-parallel transformer layers need two all-reduces in the forward pass and two in the backward pass, four per layer per step.
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMTensor parallelism’s all-reduces are kept inside a multi-GPU server on the high-bandwidth links; across slower inter-server links they become impractical (“Takeaway #1”).
- Intel Gaudi 3 AI Accelerator: Architected for Gen AI Training and Inference (Hot Chips 2024 slides)24× 200 GbE RoCE ports on 48 SerDes; 1,200 GB/s bidirectional networking vs 128 GB/s bidirectional PCIe Gen5 x16; 3.67 TB/s HBM; point-to-point links between every pair of the 8 cards in a node; the same NICs serve scale-up and scale-out.
- NVIDIA Blackwell Platform (Hot Chips 2024 slides)NVLink generations (12 × 50 GB/s, 18 × 50 GB/s, 18 × 100 GB/s per GPU, bidirectional); the 72-port NVLink Switch chip with in-network reduction; the NVL72 rack with 72 GPUs on 18 switch chips; 130 TB/s all-to-all and 260 TB/s all-reduce bandwidth.
- Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirectMeasured latency and bandwidth on DGX-1, DGX-2 and Summit: NUMA effects in the point-to-point hybrid cube-mesh (no self-routing; ~9 µs direct, 2–3× when relayed), uniform access through NVSwitch, direct remote reads, writes and atomics, CRC with replay.
- TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings4,096 chips in a 3D torus; 6 ICI links at 50 GB/s per chip; 4×4×4 blocks wired with passive electrical cables inside a rack and optical links between racks; optical links over 10× the cost of electrical; wraparound links double bisection and all-reduce bandwidth versus a mesh.
- Optimization of Collective Communication Operations in MPICHThe α + nβ cost model; ring, recursive-doubling, Bruck, binomial-tree and Rabenseifner algorithms with their costs; size-based algorithm switching; over 40% of MPI time on a Cray T3E spent in allreduce and reduce.
- AMD CDNA 3 Architecture (white paper)MI300X: seven x16 Infinity Fabric links at up to 32 Gb/s per lane form a fully connected 8-GPU node, plus x16 PCIe Gen5 to the host; 896 GB/s peak peer-to-peer aggregate.
- Intel Gaudi 3 AI Accelerator Cluster Reference Design (white paper)Of each accelerator’s 24 × 200 GbE RoCE ports, 21 are used for scale-up inside the eight-card node and 3 for scale-out.
- Amazon EC2 Trn2 Architecture16 Trainium2 chips in a 4×4 2D torus over NeuronLink-v3 at 1,024 GB/s per chip, memory pooling across the 16 chips; a 64-chip UltraServer joins four instances in rings at 256 GB/s per chip.
- Scale-Up Ethernet Framework Specification (Scale-Ethernet-RM104)Ethernet-based scale-up transport for one-sided memory load/store/atomic transactions; up to 1,024 XPUs in a single switch hop; under 2 µs round trip; up to 10 m to the switch; a 64-XPU example with twelve switch planes; also allows direct mesh deployments.
- Blink: Fast and Generic Collectives for Distributed MLOn irregular point-to-point GPU topologies, rings leave links unused; packing spanning trees uses them, up to 8× faster synchronization than NCCL.
- Introducing UALink 200G 1.0 Specification (white paper)An open scale-up standard: memory semantics (read, write, atomics) with one ordering model; 200 Gb/s lanes on an Ethernet PHY, 800 Gb/s per four-lane station each way; up to 1,024 accelerators per pod; targets under 4 m cables, under 1 µs request-to-response, 1–4 racks; link-level retry and credit-based flow control.
- Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and AlgorithmsNCCL’s ring and double-binary-tree algorithms; the Simple, LL and LL128 protocols (about 6, 1 and 2 µs per hop; near-peak, 25–50% and about 95% of peak bandwidth); NVLS algorithms that reduce inside NVLink switches.
- Collective Operations (NCCL User Guide)Definitions of AllReduce, Broadcast, Reduce, AllGather, ReduceScatter and AlltoAll; ReduceScatter followed by AllGather is equivalent to AllReduce.
- Bandwidth Optimal All-reduce Algorithms for Clusters of WorkstationsLower bound: some process must communicate at least 2(N−1)/N of the data; a contention-free ring reduce-scatter plus all-gather meets it.
- Massively Scale Your Deep Learning Training with NCCL 2.4Ring latency grows linearly with GPU count; double binary trees give full bandwidth with logarithmic latency, up to 180× lower latency at 24,576 GPUs.
- Performance reported by NCCL tests (PERFORMANCE.md)Algorithm bandwidth S/t and bus bandwidth, corrected by 2(n−1)/n for AllReduce and (n−1)/n for ReduceScatter and AllGather.
- NVIDIA Contributes NVIDIA GB200 NVL72 Designs to Open Compute Project18 compute trays, nine switch trays and four NVLink cartridges with over 5,000 copper cables in one rack; 120 kW of liquid cooling; before NVL72, an NVLink domain was limited to eight GPUs on an HGX H200 baseboard.
- Massively Distributed SGD: ImageNet/ResNet-50 Training in a Flash2D-torus all-reduce: reduce-scatter along rows, all-reduce along columns, all-gather along rows; 2(X−1) steps instead of the ring’s 2(N−1).
- Image Classification at Supercomputer ScaleOn a 1,024-chip TPU v3 pod, a 1-D ring all-reduce was latency-limited; a 2-D algorithm runs two concurrent rings on halves of the data along X and Y, using wraparound links bidirectionally, for 2× the throughput.