Systems · Chapter 3 of 7 · Between the boxes

Scale-up fabrics

To act like one giant chip, a small group of AI chips is wired together with very fast, short links inside one computer or rack.

Scale-up links connect accelerators directly or through switch chips, at far higher bandwidth than an ordinary network. Accelerators in the group can read each other’s memory and run group operations quickly.

Point-to-point and switched scale-up topologies, link bandwidth and latency, memory semantics across the domain, collectives such as all-reduce, and rack-scale designs that extend the domain over copper backplanes.

One AI chip, however fast, can’t train a modern AI model alone. The model is too big. So it’s split across many chips. The catch is that the chips must keep talking. After almost every chunk of math, each chip needs to hear what the others found.

A is a small team of chips joined by their own very fast, very short links. It can be a handful of chips or a whole rack of them. Inside the team, chips swap data so quickly that together they act a lot like one giant chip. A slower network joins many teams together. That’s the next chapter.

This chapter looks at how a team is wired, and at its most common job: everyone adding up their numbers so that everyone gets the total.

Large AI models are split across many accelerators, the specialized chips such as GPUs and TPUs that do the math. Some ways of splitting create heavy, frequent traffic. In , for example, every layer of a transformer model is divided among several chips, and they must combine their partial results four times per layer on each training step. That traffic sits directly in the path of the computation: until it finishes, nobody can move on.

A is a group of accelerators (8, 16, 64 or 72 in current systems) joined by a dedicated fabric with far more bandwidth per chip than the ordinary datacenter network. Chips in the domain can read and write each other’s memory directly and run group operations, called , quickly. The (next chapter) then connects many such domains. Training systems put the chattiest kind of splitting inside the domain and the more tolerant kinds across the scale-out network.

This chapter covers:

  • why accelerators need links of their own instead of the server’s general-purpose bus;
  • direct (point-to-point) versus switched wiring, and what each does to bandwidth and latency;
  • memory semantics: treating another chip’s memory almost like your own;
  • collectives such as , and the algorithms that run them;
  • rack-scale domains, and why copper sets their size.

A scale-up domain is the set of accelerators that share a memory-semantic fabric with roughly an order of magnitude more bandwidth per chip than the scale-out NIC. It exists because some forms of parallelism put communication on the critical path of every layer. Megatron-style needs two all-reduces in the forward pass and two in the backward pass of each transformer layer, and in practice it stays within the high-bandwidth domain, because across slower inter-server links its overhead becomes impractical. The domain’s size therefore bounds the tensor- (and expert-) parallel degree; the mapping of parallelism onto the hierarchy is the subject of Why the network looks this way.

The design questions this chapter works through:

  • Topology. How to divide a fixed budget: among direct neighbors (ring, , ) or across parallel planes, and what each does to per-pair bandwidth, hop count and domain size.
  • Semantics. Load/store/atomic access to peer memory, its ordering model, latency targets and link-level reliability.
  • Collectives. The , the bandwidth lower bound for all-reduce, and which algorithm (ring, tree, halving-doubling, direct, dimension-wise, in-network) fits which topology and message size.
  • Physical reach. Why current domains stop at one to a few racks of copper, and what open specifications target next.
Domain 1Domain 2scale-out switchscale-up linksscale-up links
Place tensor parallelism

Place a tensor-parallel group inside one domain or across two, and watch the traffic. Tap any part for details.

Two scale-up domains joined by the scale-out network. Place the tensor-parallel group and watch where its traffic goes. Schematic, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

Every AI chip in a server already has a link to the server’s main processor. It’s the same kind of slot a graphics card plugs into, called . So why not let the chips talk through it?

Because it’s far too slow for this job. On one recent AI chip, the team links carry about ten times more data per second than its PCIe slot. On another chip, they carry about fourteen times more. Using the slot would be like serving a whole dinner rush through one tiny window.

The general-purpose bus is too narrow

In a server, accelerators connect to the host CPU through , the general-purpose expansion bus covered in Board and server. PCIe is built for feeding a device from the host, not for heavy device-to-device traffic. Two published numbers show the gap:

  • One accelerator offers 1,200 GB/s of chip-to-chip networking (counting both directions) but 128 GB/s on its PCIe Gen5 x16 host link, about a ninth.
  • A current GPU has 18 dedicated links totaling 1.8 TB/s (both directions), roughly fourteen times that same Gen5 x16 figure.

Before dedicated links, GPU-to-GPU traffic shared the same PCIe tree as CPU-to-GPU traffic, and that tree was often the bottleneck.

The traffic sits on the critical path

In tensor parallelism, each layer’s matrices are split across, say, eight chips. Each chip multiplies its share, and then the chips must sum their partial results so every chip has the complete output before the next layer starts. That sum is an , and a transformer layer needs four of them per training step (two going forward, two going backward). With dozens of layers and many steps per second, the chips spend real time waiting unless the links are fast.

That is why a training job is usually arranged so that tensor parallelism stays inside a server’s fast fabric, and only less chatty kinds of splitting cross the slower network between servers.

Scale-up versus scale-out

Scale-up fabricScale-out network
ReachOne server to a few racksA whole datacenter
ChipsTens, up to about a thousand in open specificationsThousands to hundreds of thousands
Bandwidth per chipHighestSeveral times lower
How chips talkDirect memory reads and writesMessages through a network card

Bandwidth ratios

The case for a separate fabric is mostly arithmetic. Published per-chip figures (all bidirectional) put scale-up an order of magnitude above the host link:

  • Intel Gaudi 3: 1,200 GB/s of integrated Ethernet versus 128 GB/s of PCIe Gen5 x16, against 3.67 TB/s of HBM.
  • NVIDIA NVLink: 600 GB/s (12 links) for Ampere, 900 GB/s (18 links) for Hopper and 1,800 GB/s (18 links of 2 lanes at 200 Gb/s PAM4) for Blackwell.
  • Google TPU v4: six inter-chip links at 50 GB/s, against 1,200 GB/s of HBM.

Even scale-up links remain several times slower than local , which is why the parallelism layout tries to keep each chip’s working set local and exchange only activations or gradients.

Why not PCIe peer-to-peer?

PCIe does support device-to-device transfers, but in a host-rooted tree they cross shared switches and, for devices under different sockets, the CPU interconnect. Measurements on 8-GPU servers found PCIe peer bandwidth limited by switch sharing, and the bus a common bottleneck when GPUs communicate. A dedicated fabric adds bandwidth that scales with the number of accelerators rather than with host lanes, and lets the fabric protocol be tuned for small memory transactions and collectives.

Communication on the critical path

Tensor-parallel all-reduce volume per layer is proportional to batch × sequence × hidden size and does not shrink as the tensor-parallel degree grows, while the compute per chip does. So the communication fraction rises with the degree, and on slow links it dominates. That is the quantitative reason Megatron-LM keeps tensor parallelism within the server and uses pipeline parallelism across servers. Collectives were already a major cost in classic HPC: a five-year study on a Cray T3E found over 40% of MPI time in allreduce and reduce.

host CPUPCIe switchPCIe laneschip-to-chip links1234Bandwidth, both directionsPCIe Gen5 x16128 GB/saccelerator A links1,200 GB/sGPU links (18)1,800 GB/s
Path

Send data from accelerator 1 to accelerator 4: through the host’s PCIe tree, or over dedicated links.

A server with four accelerators. Bars are per-chip bandwidth counting both directions, from the vendors’ published figures.Share freely with credit: ‘Figure from chipfieldguide.com’

There are two basic ways to wire a team of chips.

Direct links

Run cables straight from chip to chip. With 8 chips, each one can have its own cable to all 7 others. But with 72 chips, each would need 71 cables, and a recent chip has only 18 link plugs.

So big teams with direct links use simpler patterns. In a ring, each chip talks only to its two neighbors. In a grid, each chip talks to the chips beside, above and below it.

Through a switch

Or plug every chip into a , a hub chip that can connect any chip to any other. Now every pair of chips is the same short trip apart. The price is the switches themselves: more chips to build, power and cool.

The pattern of connections is the . A key constraint shapes it: each chip has a fixed number of high-speed lanes, set by how many circuits fit along the edge of the die and in its power budget. Topology design is mostly about how to divide those lanes.

Point-to-point topologies

  • Full mesh. Every chip links directly to every other. With pp chips, each chip splits its lanes p−1p - 1 ways. AMD’s MI300X gives each GPU seven 16-lane links, one to each of the seven other GPUs in an eight-GPU server. Intel’s Gaudi 3 servers wire every pair of their eight cards together with Ethernet ports: 21 of each chip’s 24 ports go to the other seven (three each), and the remaining three go to the scale-out network. Every pair is one hop apart, but each pair only gets a slice of a chip’s bandwidth, and the pattern runs out of ports quickly as pp grows.
  • Ring. Each chip links to two neighbors. Cheap and scales to any size, but data between distant chips passes through many others.
  • Torus. Chips sit on a 2D or 3D grid, each linked to its 4 or 6 neighbors, with wraparound links joining opposite edges. AWS connects 16 Trainium2 chips as a 4 × 4 2D torus; Google’s TPU v4 builds a 3D torus of up to 4,096 chips.

Irregular point-to-point layouts create uneven performance. In one earlier 8-GPU server, GPUs not directly linked had to relay data through a third GPU, which doubled or tripled latency compared with a direct link.

Switched topologies

In a switched design, each chip spreads its links across several switch chips (parallel “planes”), and every switch connects to every chip. An early switched NVIDIA server, the DGX-2, joined 16 GPUs so that any GPU could reach any other at the same 300 GB/s, with nearly identical latency for every pair. Its current rack design joins 72 GPUs through 18 switch chips, each GPU sending one link to each switch.

Switches cost extra chips, power and a little latency per trip, but they make every pair equal, let the whole bandwidth go to any one destination, and can even add up data as it passes through, a trick called .

Dividing a lane budget

Treat each chip as having PP ports of bandwidth bb per direction (per-chip B=PbB = Pb). The topology decides how PP is spent:

TopologyPorts per neighborPer-pair bandwidthHops (max)Size limit
RingP/2P/2 to each of 2B/2B/2 to a neighbor⌊p/2⌋\lfloor p/2 \rfloornone (latency grows)
Full mesh⌊P/(p−1)⌋\lfloor P/(p-1) \rfloor to each of p−1p - 1≈B/(p−1)\approx B/(p-1)1p≤P+1p \le P + 1
2D torus (k×kk \times k)P/4P/4 to each of 4B/4B/4 to a neighbor≈k\approx knone (hops grow as p\sqrt{p})
Switched, one level1 to each of PP switch planesBB to any one peer1 switchp≤switch radixp \le \text{switch radix}

The full mesh’s per-pair bandwidth falls as 1/(p−1)1/(p-1), and its size is capped by port count: MI300X’s seven Infinity Fabric links (x16 at up to 32 Gb/s per lane) make exactly an eight-GPU mesh, with a quoted 896 GB/s peak peer-to-peer aggregate. A one-level switched domain is capped by switch radix instead: NVIDIA’s 72-port NVLink Switch chip sets the 72-GPU NVL72 domain, with each GPU’s 18 links going one to each of 18 switch chips. Broadcom’s Scale-Up Ethernet specification sketches the same plane structure with Ethernet switches: 64 accelerators, each with twelve 800G interfaces, through twelve 64-port switches, giving any pair up to 9.6 Tb/s.

Uniformity and routing

Point-to-point fabrics without forwarding expose topology to software. On the P100/V100 DGX-1 “hybrid cube mesh,” NVLink was not self-routed: traffic between GPUs without a direct link had to be relayed through an intermediate GPU chosen by the user, and measured latency rose from about 9 µs to roughly 2× (P100) or 3× (V100). The same study measured the switched DGX-2 as homogeneous: every pair reached 300 GB/s through NVSwitch, and the extra hop between baseboards cost almost nothing in latency. Irregularity also hurts collectives: when a scheduler hands a job an odd subset of an 8-GPU server, ring-based libraries leave some links unused, and packing spanning trees instead was up to 8× faster.

Torus properties

A torus keeps degree constant, so it scales with short, cheap cables, but its grows only as the cut surface. Wraparound links matter: TPU v4 reports that they double both bisection bandwidth and the bandwidth of collectives such as all-reduce compared with a mesh, and it chose a 3D torus over its predecessor’s 2D torus partly for bisection. Collectives map well onto tori when they run dimension by dimension (see “Under the hood”); arbitrary all-to-all traffic does not, because it crosses many hops and shares links.

12345678910111213141516Full meshchip 1’s 16 links1 to each of 15bandwidth per link used1/16hops to the destination–worst-case hops1extra chips: 0
Topology

Each chip has 16 links. Pick a topology, then tap a chip to route data to it from chip 1.

Sixteen chips with an illustrative 16 links each, as in the simulator below. The topology decides how those links are divided.Share freely with credit: ‘Figure from chipfieldguide.com’

Most computer networks work like the mail. You pack data into a message, address it and send it. Team links can work more like a shared whiteboard. A chip can read a number straight from a teammate’s memory, or write one into it, just like using its own memory.

That makes some jobs much simpler. A chip that needs one small number from a teammate just reaches over and grabs it. Some systems even describe a team of chips as sharing one big pool of memory.

The hard part is the wait. Each reach across a link takes about a millionth of a second. That sounds tiny, but to a chip it’s a long time. So a chip keeps thousands of reads and writes going at once, like a kitchen cooking many orders at the same time.

Loads and stores instead of messages

A scale-out network is usually message-based: software asks a network card to send a block of data, and the card on the other side places it in memory. Scale-up fabrics usually offer instead. An accelerator issues ordinary memory operations to an address that happens to live on another accelerator:

  • Read (load): fetch data from a peer’s memory.
  • Write (store): place data directly in a peer’s memory.
  • Atomic: read-modify-write a value in one indivisible step, for example add 1 to a counter, so chips can coordinate without races.

NVLink has supported direct reads, writes and atomics on peer GPU memory since its first generation. The open UALink specification is built around the same idea: direct read, write and atomic transactions, with the same ordering rules whether the memory is local, on the host or on a remote accelerator. Broadcom’s Scale-Up Ethernet framework likewise targets “one-sided” memory loads, stores and atomics, meaning the software on the remote chip doesn’t have to take part (the sender still gets an acknowledgment). AWS describes its 16-chip Trainium2 instances as pooling the memory of all 16 chips.

Latency matters more

Because memory operations are small (a few hundred bytes each), the fabric has to be quick. UALink targets a request-to-response round trip under 1 µs; Scale-Up Ethernet targets under 2 µs. Even so, a chip must keep many requests outstanding to fill a fast link. By , data in flight equals bandwidth × latency: to read at 800 GB/s with a 1 µs round trip, 800 KB must be in flight at all times, about 3,000 requests of 256 bytes.

Reliability at the link

A dropped memory write is not something software can easily notice and resend, so these fabrics fix errors at each link. NVLink and NVSwitch check every transfer with a CRC (a checksum) and replay it when it fails. UALink specifies and credit-based flow control, in which a sender only transmits when the receiver has advertised free buffer space, so packets are never dropped for lack of room.

Semantics and ordering

A memory-semantic fabric carries read, write and atomic transactions with an ordering model, rather than NIC-mediated messages. UALink 1.0 defines exactly that, with the same ordering model for host-attached, local and remote accelerator memory, a 64-byte transaction-layer flit packed into 640-byte data-link flits, and address compression that reaches about 95% protocol efficiency for 256-byte writes (20 of 21 flits carry data in its example). Scale-Up Ethernet keeps the transport generic (XPU-specific command/response transactions used for put, get and atomics) and lists PGAS-style shared memory as the memory architecture, leaving registration and coherence services outside its scope. The UALink white paper does not describe hardware across the domain either. Collective libraries instead order their transfers with memory fences and flags; NCCL’s protocols differ mainly in which of the two they use.

Little’s law sets the queue depth

Sustained bandwidth=outstanding bytes÷round-trip time\text{Sustained bandwidth} = \text{outstanding bytes} \div \text{round-trip time}. At B=800 GB/sB = 800\,\mathrm{GB/s} and W=1 μsW = 1\,\mu\mathrm{s}, 800 KB must be in flight: 3,125 requests of 256 B. Halve the latency and the required depth halves. This is why scale-up specifications state latency targets in hundreds of nanoseconds to microseconds (UALink: under 1 µs request-to-response; SUE: under 2 µs round trip) and why they shave FEC latency, for example UALink’s reduced codeword interleaving, which trades burst-error correction for latency.

Reliability and flow control

Memory traffic has no end-to-end retransmission in software, so reliability is pushed down a layer. NVLink applies CRC with replay on each link, ECC on datapaths and routing state, and a fabric manager that guards routing tables against out-of-range access. UALink combines Ethernet-PHY FEC, a 32-bit CRC per 640-byte flit, and credit-based flow control, plus partitioning of a pod into isolated virtual pods. Lossless credit flow avoids drops but means congestion backs up hop by hop, so incast (many senders, one receiver) must be handled without head-of-line blocking, which SUE lists among its requirements.

Accelerator AAccelerator Bprogrammemory (HBM)programmemory (HBM)NICNICnetworkscale-up fabric
Mechanism
1 / 5

1. B’s software copies the value into a send buffer and asks its network card to send it.

Message passing versus memory semantics, for one value that lives on chip B. Step through each way.Share freely with credit: ‘Figure from chipfieldguide.com’
Chip AreadsChip Bmemoryread requests →← 256-byte replies3% of the pipe in usebandwidth26 GB/slink: 800 GB/s
Round-trip time

100 reads in flight × 256 B ÷ 1 µs = 26 GB/s, 3% of the link. Filling it takes 3,125.

Little’s law with the chapter’s example: an 800 GB/s link and 256-byte reads. Round trips of 1 µs and 2 µs are the UALink and Scale-Up Ethernet targets. Animation slowed about a million times.Share freely with credit: ‘Figure from chipfieldguide.com’

The job these links do most is a team effort where every chip takes part at once. The most important one is called an . Every chip starts with its own list of numbers. At the end, every chip holds the sum of everyone’s lists.

The ring trick

The obvious way is for every chip to send its whole list to every other chip. That floods the links. A classic trick is to pass the numbers around a ring instead, adding as they go.

Two things set the time

Every pass costs two things. There’s a fixed delay just to get started, however little you send. Then there’s the time to move the data, which grows with the amount you send.

With a little data, the delays matter most. A ring of many chips has many passes, so it’s slow. With lots of data, link speed matters most. Smart software picks a different method for each case.

The common collectives

Libraries such as NCCL (for GPUs) and MPI (for supercomputers) provide a standard set of . For pp chips each holding a buffer:

  • Broadcast: one chip’s buffer is copied to all.
  • Reduce: buffers are summed element by element; one chip gets the result.
  • : summed, and every chip gets the result. Used to add up gradients in training and partial sums in tensor parallelism.
  • : summed, but each chip gets a different 1/p1/p slice of the result.
  • : each chip contributes a block; every chip ends with all blocks.
  • : every chip sends a different block to every other chip.

A useful identity: a reduce-scatter followed by an all-gather is an all-reduce. The best all-reduce algorithms are built exactly that way.

Ring all-reduce, step by step

Take p=4p = 4 chips, each with a buffer cut into 4 slices (A, B, C, D):

  1. Reduce-scatter, 3 steps. In each step, every chip sends one slice to its right-hand neighbor, which adds its own copy of that slice. After p−1=3p - 1 = 3 steps, each chip holds one slice that contains all 4 chips’ contributions.
  2. All-gather, 3 more steps. Each chip passes its finished slice to the right; after 3 more steps every chip has all 4 finished slices.

Each chip sent 2(p−1)=62(p - 1) = 6 slices of 1/4 of the buffer: 1.5 times the buffer. In general a chip sends 2(p−1)/p2(p - 1)/p of the buffer, which approaches 2 for large pp and does not grow with the number of chips. That is the least any algorithm can manage when the chips do all the adding themselves.

A simple cost model

The says sending nn bytes costs α+n/B\alpha + n/B: a fixed latency α\alpha per step plus the bytes divided by bandwidth BB. For the ring, adding up its 2(p−1)2(p - 1) steps gives:

T=2(p−1) α+2(p−1)p⋅nBT = 2(p-1)\,\alpha + \frac{2(p-1)}{p} \cdot \frac{n}{B}

A worked example, using the simulator’s defaults: 8 chips, α=2 μs\alpha = 2\,\mu\mathrm{s} per step, B=800 GB/sB = 800\,\mathrm{GB/s} per chip.

  • 100 MB buffer: latency 14×2 μs=28 μs14 \times 2\,\mu\mathrm{s} = 28\,\mu\mathrm{s}; data 1.75×100 MB÷800 GB/s≈219 μs1.75 \times 100\,\mathrm{MB} \div 800\,\mathrm{GB/s} \approx 219\,\mu\mathrm{s}; total ≈247 μs\approx 247\,\mu\mathrm{s}. Mostly moving data: bandwidth-bound.
  • 10 KB buffer: latency still 28 µs; data about 0.02 µs. Almost all waiting: latency-bound.

Because the latency term grows with pp, a ring of 72 chips pays 142 steps of latency on every all-reduce. For small messages, libraries switch to algorithms with fewer steps, such as trees, whose step count grows only with log⁡2p\log_2 p.

Definitions and the decomposition

NCCL’s semantics: AllReduce leaves the element-wise reduction on every rank; ReduceScatter leaves 1/k1/k of it per rank; AllGather concatenates kk contributions on every rank; AlltoAll sends chunk jj of rank ii’s input to rank jj. ReduceScatter + AllGather = AllReduce, and that decomposition is what makes all-reduce bandwidth-optimal.

The α-β\alpha\text{-}\beta model and the lower bound

Thakur, Rabenseifner and Gropp model a message as α+nβ\alpha + n\beta (β=1/B\beta = 1/B per byte), with γ\gamma per byte for the local reduction, assuming full-duplex links and single-ported nodes. Patarasuk and Yuan prove that, without compression, a one-item all-reduce on NN processes needs at least 2(N−1)2(N - 1) partial results communicated, so some process must move at least 2(N−1)/N2(N - 1)/N of the data; a contention-free logical ring achieves it. The ring’s cost is then

Tring=2(p−1)α+2(p−1)p nβ+p−1p nγT_{\mathrm{ring}} = 2(p-1)\alpha + \frac{2(p-1)}{p}\, n\beta + \frac{p-1}{p}\, n\gamma

(the paper’s form for its ring allreduce). Its bandwidth term is optimal; its latency term is the worst of the common algorithms. Equating the two α\alpha and β\beta terms (ignoring γ\gamma) gives a crossover at n∗=pαBn^* = p\alpha B: below it the ring is latency-bound. With p=72p = 72, α=2 μs\alpha = 2\,\mu\mathrm{s} and B=800 GB/sB = 800\,\mathrm{GB/s}, n∗≈115 MBn^* \approx 115\,\mathrm{MB}, which is larger than many tensor-parallel messages.

Choosing by message size

MPICH switches algorithms on size: recursive doubling (lg⁡p\lg p steps, but nlg⁡pn \lg p bytes) for short all-reduces and Rabenseifner’s reduce-scatter + all-gather (2lg⁡p2 \lg p steps, 2(p−1)/p⋅n2(p-1)/p \cdot n bytes) for long ones. NCCL makes the same trade with both algorithms (ring and double binary tree) and protocols: its LL protocol has about 1 µs per hop but reaches only 25–50% of peak bandwidth, LL128 about 2 µs and roughly 95%, and Simple about 6 µs and near peak. Ring latency grows linearly with GPU count; double binary trees keep full bandwidth with logarithmic latency, which NVIDIA measured as up to 180× lower latency at 24,576 GPUs.

Reporting: algbw and busbw

nccl-tests reports algorithm bandwidth S/tS/t and a “bus bandwidth” that multiplies it by 2(n−1)/n2(n-1)/n for AllReduce and (n−1)/n(n-1)/n for ReduceScatter and AllGather, so that a bandwidth-optimal implementation’s busbw approaches the per-rank link bandwidth regardless of rank count. Comparing busbw with the hardware’s per-chip bandwidth is the quickest sanity check of a collective’s efficiency.

Chip 1Chip 2Chip 3Chip 4A1B1C1D1A2B2C2D2A3B3C3D3A4B4C4D4
Collective
Show

All-reduce: summed, and every chip gets the result. Training uses it to add up gradients and partial sums. Showing the inputs; press After.

The standard collectives on p = 4 chips. Each buffer has 4 slices; color marks the chip a piece came from, and striped cells are sums of all four.Share freely with credit: ‘Figure from chipfieldguide.com’

For years, a tight team was 8 chips on one board. Today’s biggest switched teams fill a whole rack. Grid-style designs go further and join whole racks of chips.

What sets the limit is the wire. At these speeds, a copper cable carries a clean signal for only a few meters. Going farther means turning the signal into light and sending it down a glass fiber. That costs far more money and power. One company found these light links cost over ten times as much as copper ones. So designers squeeze as many chips as they can within copper’s reach, which is about a rack.

From one board to one rack

On NVIDIA’s previous HGX H200 baseboard, an NVLink domain was eight GPUs. The GB200 NVL72 design stretches it across a rack: 18 compute trays holding 72 GPUs and nine switch trays, joined by four NVLink “cartridges” of more than 5,000 copper cables mounted at the back of the rack. Google’s TPU v4 uses passive electrical cables to wire 64 chips into a 4 × 4 × 4 block within each rack. AWS joins four 16-chip Trainium2 servers into a 64-chip “UltraServer.”

Why copper, and why a rack

Copper cables need no lasers or light detectors at their ends, which keeps them cheap and low-power, but at today’s lane speeds they are short: the UALink specification targets cables under 4 meters and domains of one to four racks. Optical links reach much farther but need converters at each end. When TPU v3’s 2D torus grew across many racks, some of its wraparound links became so long that they had to be optical, and optical links cost over ten times as much as electrical ones. The trade-offs between copper and the various kinds of optics are the subject of Optics.

The cost of density

Packing 72 high-power chips into one rack to keep their links short drives up power and heat. The NVL72 rack needs a bus bar for 1,400 amps, over 100 pounds of added steel to withstand 6,000 pounds of connector mating force, and direct liquid cooling for about 120 kW. The consequences for the facility are covered in Power and cooling.

Reach sets the domain

At 100–200 Gb/s per lane, electrical reach is short, and the specifications say so: UALink, built on IEEE 802.3 200G-per-lane electrical PHYs (KR/CR), targets cables under 4 m, 1–4 racks and up to 1,024 endpoints; Scale-Up Ethernet allows up to 10 m from accelerator to switch. Past that, optics enter: TPU v3’s 2D-torus wraparound links between distant racks had to be optical, at over 10× the cost of electrical links, which pushed TPU v4 to electrically cabled 434^3 blocks per rack (passive cables) with optical circuit switches only between blocks.

Anatomy of a copper rack-scale domain

NVL72: 72 GPUs, each with 18 NVLink5 links (2 lanes at 200 Gb/s PAM4, 100 GB/s bidirectional per link), one link to each of 18 NVLink Switch chips of 72 ports; two switch chips per switch tray, nine trays. That is 72×18=1,29672 \times 18 = 1{,}296 GPU-side links, carried by four cartridges of more than 5,000 copper cables. The single switch level is set by radix: 72 ports per switch chip, 72 GPUs. The same slides show the domain growing from an 8-GPU hybrid cube mesh (2016) to an 8-GPU switched baseboard (2022) to 72 GPUs (2024).

Torus designs reach rack scale differently. TPU v4 builds 4 × 4 × 4 blocks with passive electrical cables; each block exposes 16 links on each of its 6 faces (96 optical links) to 48 optical circuit switches, which assemble up to 64 blocks (4,096 chips) and can route around failed blocks. Trn2 uses a 4 × 4 2D torus inside a server at 1,024 GB/s per chip and rings between four servers at 256 GB/s per chip, a 4:1 bandwidth step at the server boundary.

longest cable ≈ 2.0 mcopper≈ 0.6 m
Racks in the domain
Traffic

1 rack: longest cable ≈ 2.0 m, under UALink’s 4 m copper target. Cheap, low-power, no lasers.

The longest cable in a row of racks. Lengths are illustrative (about 0.6 m per rack plus 2 m of vertical run); the 4 m copper limit is UALink’s target.Share freely with credit: ‘Figure from chipfieldguide.com’
computeswitchcomputeCompute trays72 GPUsin 18 trays

A rack-scale domain. Tap a part to read about it, or trace one compute tray’s links.

The GB200 NVL72 rack, with the numbers NVIDIA published. Schematic: tray order simplified, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

Pick a wiring pattern along the top, set the number of chips and press Play to watch an all-reduce. Amber arrows carry numbers that get added when they arrive. Blue arrows hand out the finished totals. On small teams, the squares show what each chip holds.

Then slide “Amount of data” from small to large. Watch the bar flip from mostly waiting to mostly moving data.

Each chip has 16 links. Choose a topology and the simulator divides those links among neighbors, peers or switch planes, picks the fastest available algorithm, and splits the all-reduce time into its latency part (steps×α\text{steps} \times \alpha) and bandwidth part (bytes÷bandwidth\text{bytes} \div \text{bandwidth}). The chart plots time against message size for all four topologies; the shaded region is where the current setup is latency-bound. Try a 64-chip ring against a 64-chip 2D torus at 10 KB, and a full mesh above 17 chips.

An α-β\alpha\text{-}\beta model over a 16-port chip (B=16bB = 16b per direction). Pick the algorithm per topology: ring (two counter-rotating rings), binomial tree, halving-doubling (with whole-buffer fold-in for non-powers of two), direct one-shot on a full mesh, dimension-wise 2D rings on the torus, and, on the switched fabric, in-network reduction. The chart shows every algorithm on the current topology plus the endpoint bandwidth floor 2(p−1)/p⋅n/B2(p-1)/p \cdot n/B; the readouts give busbw and the crossover size. Check that n∗=pαBn^* = p\alpha B for the ring, that a single ring on a full mesh wastes most links, and that in-switch reduction beats the floor.

Loading simulation…

Here is what the numbers below mean in everyday terms.

  • Chips in a team: from 8 in one server up to 72 in one rack. One design joins 64-chip blocks into teams of thousands.
  • Speed per chip: hundreds of billions to nearly two trillion bytes a second. That’s many times more than each chip’s link to the main processor.
  • Wiring: some teams use direct cables, some use switches and some use a grid. No single answer has won yet.
Largest one-level switched GPU domain (NVL72)
72 GPUs
Copper cables in that rack’s NVLink spine
5,000+
UALink 1.0 target: accelerators per pod
1,024
UALink 1.0 target: request-to-response
< 1 µs

Reading the table

  • Three wiring strategies coexist. Switched domains (NVL72, UALink, SUE) give every pair the same path; full meshes (MI300X, Gaudi 3) skip the switch but stop at 8 chips; tori (TPU, Trainium) use short neighbor links and reach large sizes with more hops.
  • Ethernet is moving into scale-up. Gaudi 3 already uses standard Ethernet ports for its in-server mesh. UALink reuses Ethernet’s physical layer, and with it existing cables, connectors and retimers, while Scale-Up Ethernet runs its transport over Ethernet switches.
  • Memory semantics are the norm. NVLink, UALink and Scale-Up Ethernet all describe direct memory reads, writes and atomics, not just messages.

Derived checks

  • NVL72 aggregate: 72×1.8 TB/s=129.6 TB/s72 \times 1.8\,\mathrm{TB/s} = 129.6\,\mathrm{TB/s}, matching the quoted 130 TB/s all-to-all; the quoted 260 TB/s all-reduce is twice that, consistent with in-switch reduction halving per-GPU all-reduce traffic.
  • MI300X: x16 at 32 Gb/s = 64 GB/s per direction per link; 7 links×64 GB/s×2 directions=896 GB/s7\ \text{links} \times 64\,\mathrm{GB/s} \times 2\ \text{directions} = 896\,\mathrm{GB/s}, the quoted peer-to-peer figure. One peer gets one link: 64 GB/s each way.
  • Gaudi 3: 8 cards×21 ports×200 Gb/s×2 directions=67.2 Tb/s8\ \text{cards} \times 21\ \text{ports} \times 200\,\mathrm{Gb/s} \times 2\ \text{directions} = 67.2\,\mathrm{Tb/s}, the quoted node scale-up bandwidth; one peer gets 3×200 Gb/s=75 GB/s3 \times 200\,\mathrm{Gb/s} = 75\,\mathrm{GB/s} each way.
  • TPU v4: 6×50 GB/s=300 GB/s6 \times 50\,\mathrm{GB/s} = 300\,\mathrm{GB/s} of ICI against 1,200 GB/s of HBM, a 4:1 ratio.

Bigger teams are harder

A bigger team lets a model be split more ways. But every extra chip adds cables, power, heat and parts that can break. When one chip in a tight team fails, the whole team’s job may have to stop.

Switch or no switch

Skipping the switch saves money and power, and works well for 8 chips. Past that, direct wiring runs out of plugs, or data has to hop through many chips. Switches fix that but cost more.

Waiting versus moving

No single way of adding up numbers is best. Lots of small messages want the fewest hops. Huge messages want every link busy. Good software picks a method for each case.

Direct versus switched

Direct (mesh, ring, torus)Switched
Extra chipsNoneSwitch chips, trays and their power
Bandwidth between any two chipsA slice (mesh) or shared over hops (ring, torus)A chip’s full bandwidth
Distance1 hop (mesh) to many hops (ring)Always through one switch
Size limitPorts per chip (mesh) or latency (ring)Ports per switch chip
Extras—Can add up data inside the switch

What goes wrong

  • Uneven paths. In irregular direct topologies, chips that aren’t neighbors pay extra latency, and performance depends on which chips a job happens to get.
  • Idle links. An algorithm that doesn’t match the wiring, like a single ring on a full mesh, leaves most links unused.
  • Latency at scale. Ring algorithms add a step per chip, so large rings become slow for all but huge messages.
  • Failures. A bigger domain has more parts that can fail mid-job. TPU v4’s optical switches let the system route around failed blocks rather than lose the whole machine.
  • Power and heat. Short copper means dense racks: about 120 kW and liquid cooling for one 72-GPU rack.

Topology trade-offs, quantified

  • Mesh: no switch, shrinking pairs. Per-pair bandwidth ≈B/(p−1)\approx B/(p-1). Collectives that use every link at once (direct reduce-scatter/all-gather) still reach near the 2(p−1)/p⋅n/B2(p-1)/p \cdot n/B floor, but point-to-point traffic between two chips (pipeline hand-offs, KV-cache moves) sees only one link: 64 GB/s per direction on MI300X.
  • Switched: uniform, but a hop and a radix. Any chip can push BB to any one peer, and in-network reduction becomes possible. Costs: switch silicon, trays, power and one switch traversal per transfer, and a one-level domain capped at the switch radix (72 ports → 72 GPUs).
  • Torus: cheap cables, p\sqrt{p} latency. Neighbor bandwidth B/4B/4 (2D) or B/6B/6 (3D); dimension-wise collectives recover full bandwidth, but all-to-all and irregular traffic share links over multiple hops. Wraparound links double bisection compared with a mesh but are the longest cables.

Failure modes in practice

  • Topology-blind allocation. Schedulers hand jobs fragments of a server (3, 5, 6 or 7 of 8 GPUs are common), and ring libraries then strand links; spanning-tree packing recovers up to 8×.
  • Relayed traffic. Without hardware routing, non-adjacent pairs relay through a peer at 2–3× latency.
  • Wrong regime. A bandwidth-tuned protocol on small messages pays ~6 µs per hop; a latency-tuned one on large messages wastes 50–75% of bandwidth (NCCL LL).
  • Incast and backpressure. Lossless, credit-based fabrics never drop, so congestion propagates; scale-up transports must spread traffic across planes and avoid head-of-line blocking.
  • Blast radius. Larger domains raise the chance that some component fails during a job; reconfigurable optics let TPU v4 skip failed blocks and schedule around unavailable hosts.
link share of Bworst-case hopsextra chipsFull mesh2/1610Ring8/16402D torus 2×44/1630Switched16/16116

8 chips: a full mesh gives each pair 2 links, one hop, with no switch. Rings and grids are still short.

How the four wirings scale with team size, for chips with an illustrative 16 links. Link share is one link’s fraction of a chip’s bandwidth: to any peer (mesh), to a neighbor (ring, torus) or to any one chip through all switch planes.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

Ring all-reduce

Split the buffer into pp chunks. In reduce-scatter step tt (t=0,…,p−2t = 0, \dots, p-2), rank ii sends chunk (i−t) mod p(i - t) \bmod p to rank i+1i + 1, which adds it to its own copy. After p−1p - 1 steps rank ii holds the fully reduced chunk (i+1) mod p(i + 1) \bmod p. All-gather repeats the pattern without the addition. Each step moves n/pn/p bytes per rank on one link in each direction:

Tring=2(p−1)α+2(p−1)p nβ+p−1p nγT_{\mathrm{ring}} = 2(p-1)\alpha + \frac{2(p-1)}{p}\, n\beta + \frac{p-1}{p}\, n\gamma

The bandwidth term meets the Patarasuk–Yuan lower bound for algorithms in which only endpoints reduce. Running two rings in opposite directions on half the buffer each uses both directions of every link, which is how the simulator charges a physical ring.

Trees and recursive algorithms

  • Binomial tree reduce + broadcast: ⌈lg⁡p⌉\lceil \lg p \rceil steps each way, but the whole buffer every step: T=2⌈lg⁡p⌉(α+nβ)T = 2\lceil \lg p \rceil (\alpha + n\beta) (plus γ\gamma on the reduce). Latency-optimal, bandwidth-poor.
  • Recursive doubling: lg⁡p⋅α+nlg⁡p⋅β+nlg⁡p⋅γ\lg p \cdot \alpha + n \lg p \cdot \beta + n \lg p \cdot \gamma, MPICH’s short-message allreduce.
  • Rabenseifner (halving-doubling): reduce-scatter by recursive vector halving and distance doubling, then all-gather by vector doubling and distance halving: T=2lg⁡p⋅α+2(p−1)/p⋅nβ+(p−1)/p⋅nγT = 2 \lg p \cdot \alpha + 2(p-1)/p \cdot n\beta + (p-1)/p \cdot n\gamma. Optimal in both terms for powers of two, when the network carries the butterfly pattern without contention. For other pp, MPICH first folds the r=p−p′r = p - p' extra ranks into neighbors with a half-buffer exchange and sends them the result at the end. (The simulator uses a simpler whole-buffer fold: 2 extra steps and 2n2n extra bytes.)
  • Double binary trees: two complementary binary trees, arranged so that no rank is interior in both; pipelined, they approach full bandwidth with logarithmic latency. NCCL uses them alongside rings. The simulator’s “Tree” is the unpipelined binomial version, to show the classic latency-for-bandwidth trade.

Direct (one-shot) on a full mesh

With p−1p - 1 links per chip, reduce-scatter can be a single step: every rank sends chunk jj to rank jj over the dedicated link, and all p−1p - 1 transfers proceed at once; all-gather mirrors it. With per-peer bandwidth Bpair=kbB_{\mathrm{pair}} = kb (kk ports per peer):

Tdirect=2α+2(n/p)BpairT_{\mathrm{direct}} = 2\alpha + \frac{2(n/p)}{B_{\mathrm{pair}}}

Since k(p−1)≈Pk(p - 1) \approx P, this is within rounding of the 2(p−1)/p⋅n/B2(p-1)/p \cdot n/B floor, with only two steps. A single ring on the same mesh uses 2 of the p−1p - 1 links, so its bandwidth term is (p−1)/2(p-1)/2 times worse, the mismatch the simulator’s expert view exposes, and the reason libraries pack many trees or rings onto point-to-point graphs. (This formula is our derivation under the α-β\alpha\text{-}\beta model; real one-shot kernels also pay for p−1p - 1 concurrent streams per chip.)

01234567links busy per chip2 of 7latency steps (× α)2(p − 1) = 14bandwidth term ÷ floor3.5×dashed: idle links
Schedule

One bidirectional ring drives 2 of each chip’s 7 links: 14 steps of α, and a bandwidth term (p − 1)/2 = 3.5× the floor.

Eight chips, fully meshed, with B_pair ≈ B/7. Ratios are the α-β bandwidth term against the endpoint floor 2(p − 1)/p · n/B; every busy link carries traffic both ways.Share freely with credit: ‘Figure from chipfieldguide.com’

Dimension-wise algorithms on a torus

Mikami et al.’s 2D-torus all-reduce on an X×YX \times Y grid does reduce-scatter along rows, all-reduce along columns on the 1/X1/X-sized result, and all-gather along rows; they count 2(X−1)2(X - 1) GPU-to-GPU operations against 2(N−1)2(N - 1) for a single ring over all NN GPUs. Ying et al. found TPU v3’s 1-D ring latency-limited at pod scale and ran two concurrent rings on halves of the payload, one along XX and one along YY, so both dimensions’ links stay busy, using wraparound links bidirectionally; they report twice the throughput of the 1-D algorithm. For an R×CR \times C torus with neighbor bandwidth B/4B/4, a half that goes rows first costs

TA=2(C−1)α+2(R−1)α+[2(C−1)C⋅n2+2(R−1)R⋅n2C]/B2\begin{aligned} T_{\mathrm{A}} = {} & 2(C-1)\alpha + 2(R-1)\alpha \\ & + \left[\frac{2(C-1)}{C} \cdot \frac{n}{2} + \frac{2(R-1)}{R} \cdot \frac{n}{2C}\right] \Big/ \frac{B}{2} \end{aligned}

and the simulator takes the slower of the two halves. On 8 × 8 the bandwidth term equals a 64-chip ring’s (about 1.97 n/B1.97\,n/B) but the latency term is 28α28\alpha instead of 126α126\alpha. A snake-shaped single ring on the same torus touches only 2 of 4 neighbor links and pays both the 126α126\alpha and twice the bandwidth term.

In-network reduction

With , each endpoint sends its nn bytes into the switch planes once and receives the reduced nn bytes once. With full-duplex links and pipelining, T≈2α+n/BT \approx 2\alpha + n/B, below the endpoint floor 2(p−1)/p⋅n/B2(p-1)/p \cdot n/B by nearly 2×. NCCL’s NVLS algorithms use NVLink SHARP in NVSwitch for exactly this, and its CollNet algorithms use SHARP-capable network switches. The NVL72 switch chip lists 3.6 TFLOPS of in-network compute, and the system’s quoted all-reduce bandwidth (260 TB/s) is twice its all-to-all bandwidth (130 TB/s). The catch: arithmetic precision and format support are fixed by the switch, and the reduction needs buffer space and synchronization inside it.

Model assumptions

The α-β\alpha\text{-}\beta model ignores contention, assumes every rank sends and receives at once, treats latency as fixed per step, and charges nothing for chunking, protocol overhead or the reduction arithmetic (γ\gamma). Real libraries tune chunk sizes, channels and protocols, and the crossover sizes they choose are measured, not derived: MPICH, for example, set a 2 KB short/long boundary for reduce experimentally. Use the model to reason about scaling and regimes, and busbw from nccl-tests to measure the real thing.

Novice · 0 of 4 correct
  1. Q1Eight chips are wired as a full mesh, and each chip has 16 links in total. How many links can go to each other chip?

  2. Q2A ring all-reduce runs on 8 chips. Roughly how much data does each chip send, compared with the size of the buffer being summed?

  3. Q3What does it mean that a scale-up fabric has “memory semantics”?

  4. Q4Why does a 2D torus all-reduce on 64 chips take far fewer steps than a single ring through all 64?

Sources

Show Hide 21 sources
  1. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismMohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catanzaro · arXiv:1909.08053 · 2019Tensor-parallel transformer layers need two all-reduces in the forward pass and two in the backward pass, four per layer per step.
  2. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMDeepak Narayanan et al. · arXiv:2104.04473 (SC21) · 2021Tensor parallelism’s all-reduces are kept inside a multi-GPU server on the high-bandwidth links; across slower inter-server links they become impractical (“Takeaway #1”).
  3. Intel Gaudi 3 AI Accelerator: Architected for Gen AI Training and Inference (Hot Chips 2024 slides)Roman Kaplan, Intel · Hot Chips 36 · 202424× 200 GbE RoCE ports on 48 SerDes; 1,200 GB/s bidirectional networking vs 128 GB/s bidirectional PCIe Gen5 x16; 3.67 TB/s HBM; point-to-point links between every pair of the 8 cards in a node; the same NICs serve scale-up and scale-out.
  4. NVIDIA Blackwell Platform (Hot Chips 2024 slides)Ajay Tirumala and Raymond Wong, NVIDIA · Hot Chips 36 · 2024NVLink generations (12 × 50 GB/s, 18 × 50 GB/s, 18 × 100 GB/s per GPU, bidirectional); the 72-port NVLink Switch chip with in-network reduction; the NVL72 rack with 72 GPUs on 18 switch chips; 130 TB/s all-to-all and 260 TB/s all-reduce bandwidth.
  5. Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirectAng Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan Tallent, Kevin Barker · arXiv:1903.04611 (IEEE TPDS) · 2019Measured latency and bandwidth on DGX-1, DGX-2 and Summit: NUMA effects in the point-to-point hybrid cube-mesh (no self-routing; ~9 µs direct, 2–3× when relayed), uniform access through NVSwitch, direct remote reads, writes and atomics, CRC with replay.
  6. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsNorman P. Jouppi et al. · arXiv:2304.01433 (ISCA 2023) · 20234,096 chips in a 3D torus; 6 ICI links at 50 GB/s per chip; 4×4×4 blocks wired with passive electrical cables inside a rack and optical links between racks; optical links over 10× the cost of electrical; wraparound links double bisection and all-reduce bandwidth versus a mesh.
  7. Optimization of Collective Communication Operations in MPICHRajeev Thakur, Rolf Rabenseifner, William Gropp · International Journal of High Performance Computing Applications (author copy, Argonne National Laboratory) · 2005The α + nβ cost model; ring, recursive-doubling, Bruck, binomial-tree and Rabenseifner algorithms with their costs; size-based algorithm switching; over 40% of MPI time on a Cray T3E spent in allreduce and reduce.
  8. AMD CDNA 3 Architecture (white paper)AMD · AMD Instinct technical documentationMI300X: seven x16 Infinity Fabric links at up to 32 Gb/s per lane form a fully connected 8-GPU node, plus x16 PCIe Gen5 to the host; 896 GB/s peak peer-to-peer aggregate.
  9. Intel Gaudi 3 AI Accelerator Cluster Reference Design (white paper)Intel · Intel · 2024Of each accelerator’s 24 × 200 GbE RoCE ports, 21 are used for scale-up inside the eight-card node and 3 for scale-out.
  10. Amazon EC2 Trn2 ArchitectureAmazon Web Services · AWS Neuron documentation16 Trainium2 chips in a 4×4 2D torus over NeuronLink-v3 at 1,024 GB/s per chip, memory pooling across the 16 chips; a 64-chip UltraServer joins four instances in rings at 256 GB/s per chip.
  11. Scale-Up Ethernet Framework Specification (Scale-Ethernet-RM104)Broadcom · Broadcom technical documentation · 2025Ethernet-based scale-up transport for one-sided memory load/store/atomic transactions; up to 1,024 XPUs in a single switch hop; under 2 µs round trip; up to 10 m to the switch; a 64-XPU example with twelve switch planes; also allows direct mesh deployments.
  12. Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and AlgorithmsZhiyi Hu, Siyuan Shen, Tommaso Bonato, Sylvain Jeaugey, et al. · arXiv:2507.04786 · 2025NCCL’s ring and double-binary-tree algorithms; the Simple, LL and LL128 protocols (about 6, 1 and 2 µs per hop; near-peak, 25–50% and about 95% of peak bandwidth); NVLS algorithms that reduce inside NVLink switches.
  13. Collective Operations (NCCL User Guide)NVIDIA · NCCL documentationDefinitions of AllReduce, Broadcast, Reduce, AllGather, ReduceScatter and AlltoAll; ReduceScatter followed by AllGather is equivalent to AllReduce.
  14. Bandwidth Optimal All-reduce Algorithms for Clusters of WorkstationsPitch Patarasuk and Xin Yuan · Journal of Parallel and Distributed Computing (author copy, Florida State University) · 2009Lower bound: some process must communicate at least 2(N−1)/N of the data; a contention-free ring reduce-scatter plus all-gather meets it.
  15. Massively Scale Your Deep Learning Training with NCCL 2.4Sylvain Jeaugey, NVIDIA · NVIDIA Technical Blog · 2019Ring latency grows linearly with GPU count; double binary trees give full bandwidth with logarithmic latency, up to 180× lower latency at 24,576 GPUs.
  16. Performance reported by NCCL tests (PERFORMANCE.md)NVIDIA · nccl-tests (GitHub)Algorithm bandwidth S/t and bus bandwidth, corrected by 2(n−1)/n for AllReduce and (n−1)/n for ReduceScatter and AllGather.
  17. NVIDIA Contributes NVIDIA GB200 NVL72 Designs to Open Compute ProjectAmr Elmeleegy, NVIDIA · NVIDIA Technical Blog · 202418 compute trays, nine switch trays and four NVLink cartridges with over 5,000 copper cables in one rack; 120 kW of liquid cooling; before NVL72, an NVLink domain was limited to eight GPUs on an HGX H200 baseboard.
  18. Massively Distributed SGD: ImageNet/ResNet-50 Training in a FlashHiroaki Mikami, Hisahiro Suganuma, Pongsakorn U-chupala, Yoshiki Tanaka, Yuichi Kageyama · arXiv:1811.05233 · 20182D-torus all-reduce: reduce-scatter along rows, all-reduce along columns, all-gather along rows; 2(X−1) steps instead of the ring’s 2(N−1).
  19. Image Classification at Supercomputer ScaleChris Ying, Sameer Kumar, Dehao Chen, Tao Wang, Youlong Cheng · arXiv:1811.06992 · 2018On a 1,024-chip TPU v3 pod, a 1-D ring all-reduce was latency-limited; a 2-D algorithm runs two concurrent rings on halves of the data along X and Y, using wraparound links bidirectionally, for 2× the throughput.