Systems · Chapter 2 of 7 · Inside the box

Board and server

A server is a powerful computer kept in a datacenter. Inside, the chips sit on boards with memory, power supplies, cooling, and a small helper computer that keeps watch.

Accelerators connect to host CPUs over PCIe, and CXL adds memory sharing over the same wires. Voltage regulators turn rack power into the low voltages chips need. A baseboard management controller monitors and controls the machine.

PCIe generations and lane budgets, CXL device types, power delivery from busbar to point of load, board and module form factors including open OCP specifications, and system management.

An AI server is a box of parts that work as a team. One or two ordinary processors run the show. Several , chips built for AI math, do the heavy lifting. Network cards link the box to other servers. Fast drives hold the data, and power supplies feed it all.

Building one is like sticking to a budget. Each processor has only so many fast wires coming out of it. The power supplies can give only so much electricity. And the box can get rid of only so much heat. Adding one more accelerator spends a bit of each budget.

The previous chapter, Packaging and chiplets, ended at the edge of one package. This chapter puts packages onto boards and boards into a server. A typical AI training server combines:

  • one or two , which boot the machine, run the operating system and feed the rest;
  • several (often eight), each with its own high-bandwidth memory;
  • network cards (), often one per accelerator, and flash drives;
  • power supplies and voltage converters, fans or liquid cooling, and a management controller.

One commercial 8-accelerator server, for example, pairs eight GPUs with two CPUs, ten network cards, ten NVMe drives and six 3.3 kW power supplies, and can draw up to 10.2 kW. Three questions shape every such design, and they are the subject of this chapter: how the parts are wired together ( and CXL), how power gets from the rack to each chip, and how the machine is managed when something goes wrong.

A server is a set of budgets that have to close at the same time: CPU and switch ports, link bandwidth on each path the data actually takes, power from the busbar through several conversion stages, airflow or coolant, and board area. An accelerator-dense 8U system illustrates the scale: 8 GPUs, 2 CPUs, 8 single-port 400 Gb/s ConnectX-7 adapters for the compute fabric plus 2 dual-port ones for storage and management, 2 M.2 boot drives and 8 U.2 NVMe drives, 2 TB of DRAM, 6 × 3.3 kW supplies in 4+2 redundancy and a 10.2 kW maximum input.

The chapter works through each budget in turn:

  • the host–accelerator split and the PCIe hierarchy (root ports, switches, endpoints), lane budgets and oversubscription, and how per-lane rates evolved from 8b/10b to PAM4 flits;
  • CXL’s three protocols and device types, and what memory expansion costs in latency;
  • power delivery from a 48 V-class busbar through intermediate and point-of-load conversion to sub-1 V rails at hundreds of amps;
  • module and rack form factors from the Open Compute Project, and out-of-band management with a BMC.

Accelerator-to-accelerator fabrics are only touched on here; they have their own chapter, Scale-up fabrics.

PSU ×6 (4+2)BMCNIC ×10CPU ×2accelerator ×8NVMe ×8fansfront ↓
Budget

Tap a part to read its job, or pick a budget to see which parts supply it and which use it.

A generic 8-accelerator server from above, with the part counts of the chapter’s example system. Layout schematic, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

A server’s main processors are like the one in a laptop, just bigger. They’re good at doing lots of different jobs, one after another. Accelerators are the opposite. They do one kind of job, giant piles of multiplication, extremely fast.

So the work is split. The is the boss processor. It loads the data, gets it ready and hands it to the accelerators. Then it tells them what to compute. If the host can’t keep up, the costly accelerators sit around waiting.

Who does what

The owns the machine: its memory, the operating system and the list of attached devices. An is a device attached to it. During training, the CPU’s jobs are to read samples from storage or the network, decode and transform them (decompressing images, cropping, tokenizing text), copy batches to the accelerators, and launch and synchronize their work. The accelerators’ jobs are the matrix math, using their own memory, and the exchanges with their peers.

The split creates a pipeline, and the slowest stage sets the pace. In one study of image and audio training, models needed between 3 and 24 CPU cores per GPU just for data preparation, and spent up to 65% of each pass over the data waiting on it. That is why server designers balance CPU cores, memory and drive bandwidth against the number of accelerators rather than simply adding accelerators.

Two kinds of link

Accelerators have two separate connections. The host link, usually PCIe, ties each one to a CPU for data and control. A second, much faster set of links ties the accelerators to each other. One open accelerator-module design, for example, gives each module one or two 16-lane host links and up to seven links to its peers. The host link is the subject of this chapter; the peer links are covered in Scale-up fabrics.

Division of labor

The owns the PCIe hierarchy, the I/O memory map and the main DRAM; each is an endpoint in that hierarchy with its own HBM, its own DMA engines and usually a dedicated scale-up fabric. The host link carries input batches, checkpoints, kernel launches and synchronization, plus network traffic when NICs exchange data with accelerator memory. Gradient exchange inside the server normally rides the scale-up links instead. The OAM base specification makes the split explicit: one or two x16 host links per module, and up to seven module-to-module links of up to x16–x20 lanes each.

The input pipeline is a host-side budget

Mohan et al. profiled nine models on 8-GPU servers and separated fetch stalls (waiting on I/O) from prep stalls (waiting on CPU decoding and augmentation). Their headline numbers: 3–24 CPU cores per GPU needed for pre-processing, up to 65% of epoch time spent in it, and an OS page cache that thrashes under the random access pattern of training. The implication for server design is that “balanced SKUs” matter: cores per accelerator, DRAM capacity for caching decoded data, and drive and NIC bandwidth per accelerator are all first-class parameters, not leftovers.

Host-link topologies

Three host-link patterns recur. In the flat pattern, each device takes a root port directly and lanes run out quickly. In the switched pattern, PCIe switches fan out lanes and keep accelerator, NIC and drive traffic local; vendor documentation for one 8-GPU OAM platform shows the modules reaching dual-socket hosts through optional PCIe switches and retimers. In the coherent pattern, a CPU and accelerator share a coherent link wider than PCIe, or the accelerator speaks CXL. The rest of this chapter quantifies the first two and introduces the third.

storagesampleshost CPU8 coresacceleratormatrix mathidle 67%PCIeslowest stage: CPU prepthis prep needs ≈ 24 cores per accelerator
Data preparation per sample

8 cores per accelerator, but this prep needs about 24: the CPU is the slowest stage, and the accelerator sits idle 67% of the time.

The training input pipeline: the slowest stage sets the pace. The 3 and 24 cores per accelerator are the ends of the range one study measured; the model is illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Inside the server, parts are joined by , a standard kind of super-fast wiring. The wires come in sets called . Each lane works like a two-way road: one side for sending, one for receiving. An accelerator usually gets 16 lanes, and a storage drive gets 4.

Every few years a new version of PCIe comes out. Each one doubles the speed of every lane.

The catch is that each processor has only so many lanes. When there are more parts than lanes, designers add chips. A switch lets several parts share one set of lanes, like a power strip shares one wall socket. That works well when parts take turns. If they all send at once, they slow each other down.

Lanes and links

(PCI Express) connects each device to the CPU over its own point-to-point link: no shared bus. A link is made of , and each lane is two pairs of wires, one pair sending and one receiving, so a lane carries data in both directions at full speed at the same time. An accelerator module’s host-link pinout, for instance, lists 16 transmit pairs and 16 receive pairs for its x16 link. The number of lanes is the : x1, x4, x8 or x16. Doubling the width doubles the bandwidth.

Generations

Each generation doubles the raw rate per lane, measured in gigatransfers per second (GT/s, roughly billions of bits per second on each wire pair). Not every transferred bit is data. Generations 1 and 2 sent 10 bits for every 8 bits of data, a 20% overhead; generations 3 to 5 send 130 bits for every 128, about 1.5%.

GenerationYearRate per lanex16, one direction
3.020108 GT/s≈ 16 GB/s
4.0201716 GT/s≈ 32 GB/s
5.0201932 GT/s≈ 63 GB/s
6.0202264 GT/s≈ 128 GB/s raw
7.02025128 GT/s≈ 256 GB/s raw

Rates and years for 3.0–5.0 from PCI-SIG’s generation table; 6.0 from a PCI-SIG blog post and 7.0 from its specification page. A newer device in an older slot works: both ends agree on the fastest speed they share.

The lane budget

Every link starts at a root port on the CPU’s , the logic that connects PCIe to the processor and its memory. A CPU package has a fixed number of lanes, so the first budget in a server is simple arithmetic. Eight accelerators at x16 need 128 lanes, eight x16 NICs another 128, and eight NVMe drives at x4 another 32: 288 lanes. Two CPUs with 80 usable lanes each offer 160. The design does not fit.

The usual fix is a : a chip with one upstream link to the CPU and several downstream links to devices. Four switches, each with a x16 uplink, use only 64 CPU lanes, and each can host two accelerators, two NICs and two drives (72 lanes) below it. The price is : 72 lanes below share 16 above, a 4.5:1 ratio. If all six devices pull data from host memory at once, each gets well under a quarter of its own link speed.

Often that is fine, because the busiest traffic never goes to the host. A NIC can write straight into an accelerator’s memory, and a drive can send data straight to an accelerator, when both sit behind the same switch; the traffic turns around inside the switch. That is why AI servers pair each accelerator with a NIC under a shared switch.

Reach and retimers

Faster signals fade faster along copper. At PCIe 5.0 speed, a good circuit board carries the signal only about 5 inches (13 cm) before it is too weak. Longer paths, to a riser card or the far side of a big board, need a , a chip that receives the signal and sends a fresh copy, splitting the path in two.

Hierarchy

PCIe is a tree. Root ports on the originate links; a is an upstream port and several downstream ports, each a logical PCI-PCI bridge, joined by internal routing; endpoints sit at the leaves. Each is a dual-simplex pair of differential pairs, so a x16 link is 64 signal wires (an OAM host link’s pinout lists 16 TX and 16 RX pairs). Links train up from the lowest common rate and width; a x16 PCIe 5.0 port interoperates with a x1 device of any earlier generation.

Per-lane bandwidth by generation

Usable per-direction bandwidth per lane is the transfer rate times the coding efficiency, divided by 8. Through PCIe 5.0 the coding is NRZ with 8b/10b (gens 1–2) or 128b/130b (gens 3–5); PCI-SIG’s table gives x16 per-direction bandwidth after encoding of 32, 64, 126, 252 and 504 Gb/s for generations 1 to 5. PCIe 6.0 keeps the 16 GHz Nyquist frequency of 5.0 but sends 2 bits per symbol with PAM4. The penalty is a raw error rate far above the 10−1210^{-12} of earlier generations, so 6.0 adds a light-weight forward error correction (FEC), designed for a first-bit-error rate of 10−610^{-6}, in front of CRC and link-level retry. FEC needs fixed-size codewords, hence mode.

GenGT/sCodingEfficiencyGB/s per lane per directionx16 per direction
3.08128b/130b98.5%0.9815.8 GB/s
4.016128b/130b98.5%1.9731.5 GB/s
5.032128b/130b98.5%3.9463.0 GB/s
6.064PAM4, 256 B flit236/256=92.2%236/256 = 92.2\%7.38118 GB/s
7.0128PAM4, 256 B flit92.2%14.8236 GB/s

Computed from the rates and flit layout in the PCI-SIG sources. Neither column counts transaction headers: in 128b/130b mode each packet also carries framing, sequence number and link CRC, which flit mode removes. PCI-SIG’s analysis puts flit mode ahead of 128b/130b in packet efficiency for payloads up to 512 bytes, for up to about 3× the effective throughput of 5.0 on small transfers, falling to about 0.98 efficiency at 4 KB payloads. Headline figures such as “256 GB/s for x16” at 64 GT/s are raw and bidirectional.

Lane budgets and switches

Write the budget per root complex as ∑wi≤LCPU\sum w_i \le L_{\mathrm{CPU}} over directly attached devices, and per switch as ∑wi≤Ldown\sum w_i \le L_{\mathrm{down}}, with the switch’s uplink width wupw_{\mathrm{up}} charged to its CPU. The oversubscription ratio is ρ=∑wi/wup\rho = \sum w_i / w_{\mathrm{up}}. For host-bound streaming with all devices active, each device gets min⁡(1,1/ρ)\min(1, 1/\rho) of its link. Peer-to-peer traffic between devices under one switch does not touch the uplink; when access control allows, “the transaction can route entirely within the PCIe hierarchy and never reach the root port,” while forwarding between root ports is not defined by the specification and the Linux kernel blocks it by default. NVIDIA’s GPUDirect RDMA guidance ranks the paths the same way: PCIe switches only is best, a path through one CPU is worse, and a path that crosses the inter-socket link may be severely limited. Two practical consequences:

  • Pair each accelerator with its NIC (and ideally its drives) under one switch, so RDMA and storage DMA stay local.
  • Spread switches across both sockets, and keep each process on the socket that owns its devices, so host staging doesn’t cross the socket link.
lspci -tv (illustrative, trimmed): one switch under one root porttext
-[0000:00]-+-01.0-[01-06]----00.0-[02-06]--+-00.0-[03]----00.0  Accelerator (x16)
           |                               +-01.0-[04]----00.0  NIC, 400G (x16)
           |                               +-02.0-[05]----00.0  NVMe SSD (x4)
           |                               \-03.0-[06]----00.0  NVMe SSD (x4)
           \-03.0-[07]----00.0  CXL memory expander (x16)
  1. 1L1Root port 01.0 leads to a switch upstream port (bus 01) whose internal bus 02 fans out to four downstream ports.
  2. 2L2Accelerator and NIC share the switch: RDMA between them can turn around inside it.
  3. 3L5A device on its own root port: traffic to or from it always passes through the root complex.

Signal reach

Loss budgets grew from 22 dB at 8 GT/s to 28 dB at 16 GT/s and 36 dB at 32 GT/s, yet at 32 GT/s low-loss board material still allows only about 5 inches of system-board trace. A splits the channel into two segments, each with a full budget. PCIe 6.0 holds reach roughly at 5.0 levels: PAM4 keeps the Nyquist frequency, though choosing a 10−610^{-6} FBER rather than networking’s 10−410^{-4} costs about 2–4 inches; in return FEC latency stays under 2 ns.

CPUroot portdeviceendpointtransmit →← receivex16 · 5.0 · ≈ 63 GB/s each wayPCIe 5.0 · 32 GT/s · 2019
Link width
Generation

x16 PCIe 5.0: 32 GT/s per lane, ≈ 3.9 GB/s per lane each way after encoding overhead, ≈ 63 GB/s per direction.

Each lane is two wire pairs, one per direction, both running at once. Width multiplies the bandwidth; each generation doubles the rate per lane. Dot speed not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’
socket linkDRAMDRAMCPU 0switchaccel.NICNVMeCPU 1switchaccel.NICNVMeNIC → switch → accelerator
Path NIC → accelerator

Peer-to-peer: the data turns around inside the switch. It uses none of the CPU’s lanes and none of its memory, which is why each accelerator gets a NIC under the same switch.

PCIe is a tree: root complexes in the CPUs, a switch under each, devices at the leaves. Three ways for a NIC’s data to reach an accelerator, best first.Share freely with credit: ‘Figure from chipfieldguide.com’

A server’s main memory plugs into slots next to the processor, and there’s only room for so many. AI jobs often want more. is a newer standard that uses the same PCIe wires to add memory on a plug-in card. The processor can use that memory as if it were its own.

The trade-off is distance. Memory on a CXL card is a bit farther away. So each trip to it takes a little longer. For data that isn’t needed all the time, that’s a fine deal.

(Compute Express Link) is an open standard that runs on the same physical lanes as PCIe. A CXL device starts out as an ordinary PCIe device and then negotiates CXL. On top of the PCIe-style protocol (CXL.io) it adds two more:

  • CXL.cache lets a device keep a cached copy of the CPU’s memory, with the hardware keeping the copies consistent.
  • CXL.mem lets the CPU read and write a device’s memory with ordinary load and store instructions, as if it were main memory.

Devices come in three types, depending on which protocols they use:

TypeExampleProtocols
Type 1Smart NIC without its own memory poolCXL.io + CXL.cache
Type 2GPU or FPGA with local memoryCXL.io + CXL.cache + CXL.mem
Type 3Memory expanderCXL.io + CXL.mem

Today the most common use in servers is the : a card full of DRAM that adds capacity without adding memory channels to the CPU. A device using 8 PCIe 5.0 lanes needs about a third as many pins as a DDR5 memory channel, and at a typical mix of reads and writes it carries about as much as one DDR5-4800 channel.

The cost is latency. The CXL designers estimate a CPU load from a directly attached expander at around 170 ns, about the same as reaching memory attached to the other CPU socket. Measured devices vary: in one study, load latency ranged from 35% longer to about three times longer than that remote-socket memory, depending on the controller chip. Software therefore treats CXL memory as a slower tier, a separate node with no processors of its own.

The standard has moved quickly. CXL 1.0 and 1.1 (2019) connected one host to its devices; 2.0 (2020) added a switch level so memory can be pooled among several hosts; 3.0 (2022) doubled the rate to 64 GT/s and allowed larger fabrics.

Protocols and device types

multiplexes three protocols on the PCIe PHY. Link training starts at 2.5 GT/s as PCIe, then negotiates CXL through alternate protocol negotiation. CXL.io reuses PCIe semantics for discovery, configuration, address translation and DMA; CXL.cache lets a device coherently cache host memory; CXL.mem lets the host (and other devices) load and store to device-attached memory as cacheable “host-managed device memory.” Type 1 devices implement .io + .cache, Type 2 all three, Type 3 .io + .mem. CXL 1.x/2.0 use a 68-byte flit (2 B protocol ID, 64 B payload, 2 B CRC); CXL 3.0 adopts PCIe 6.0’s 256-byte flit at 64 GT/s with no added latency, plus a latency-optimized variant.

VersionReleasedRateScope
1.0 / 1.1Mar / Sep 201932 GT/ssingle host
2.0Nov 202032 GT/sone switch level; pooling across ~2–16 hosts; hot-plug
3.0Aug 202264 GT/smulti-level switching, fabrics, memory sharing

From Das Sharma, Blankenship and Berger.

Latency budget

The CXL port costs about 21–25 ns round trip on each side, and with about 15 ns of flight time through retimers the designers estimate a 57 ns end-to-end adder for a memory access across a link; the specification’s pin-to-pin target is 80 ns for a CXL.mem access. Their table estimates about 170 ns load latency from a CPU to a directly attached Type 3 device and about 250 ns through one CXL switch, and describes the direct case as similar to a remote-socket DDR access. Pond, a memory-pooling design for a public cloud, estimates 70–90 ns over NUMA-local DRAM for pools of 8–16 sockets and more than 180 ns for rack-scale pools. Measured hardware spreads widely: on a 4th-generation Xeon host, three CXL devices showed 35%, about 2× and about 3× the load latency of remote-socket DDR5, and the authors also found that emulating CXL with remote-socket memory can misstate latency.

Bandwidth and pins

A x16 CXL link at 32 GT/s carries 64 GB/s raw per direction (128 GB/s at 64 GT/s). Pond notes that a bidirectional x8 port at a 2:1 read:write mix matches one DDR5-4800 channel, and Sun et al. note a PCIe 5.0 x8 memory device uses about 3× fewer pins than a DDR5 channel. The pin argument is the strategic one: CPU packages are pin-limited, and serial lanes deliver more bandwidth per pin than a parallel DDR bus, at the cost of latency.

What it is used for

  • Capacity tiering. Hot pages in local DDR, colder pages on a expander, managed by the OS’s policies.
  • Pooling. Memory that would be stranded on one host is lent to another without a reboot; Pond reports a 7% DRAM cost reduction with performance within 1–5% of NUMA-local allocations.
  • Coherent accelerators. Type 2 devices share data with the host without explicit copies. On today’s AI servers the bulk of accelerator data still moves by DMA over PCIe semantics, so this is the least deployed of the three.
PCIe lanessocket linkslots fullDRAMCPU 0CPU 1Type 3 cardswitchpooled memoryshared by hostsload latency →local DRAMno linkother socket≈ CXL cardCXL card≈ 170 ns (est.)via CXL switch≈ 250 ns (est.)
Load from

A Type 3 memory expander on PCIe lanes: more capacity with no new memory channels. Estimated ≈ 170 ns per load; measured devices were 35% to about 3× slower than other-socket memory.

One CPU loading data from four places. CXL times are the CXL designers’ estimates; the local and other-socket bars are drawn from those estimates, so they are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Chips run on less than one volt. Voltage is the push that makes electricity flow, and a phone charger gives about 5 volts. Yet a big AI chip uses as much power as a few space heaters. With such a weak push, it needs a huge flow of electricity to get that much power. That flow is called current.

Sending a big current down a long wire wastes energy as heat. So power travels most of the way at a higher voltage, and is turned down in steps. A thick metal in the rack carries about 50 volts. Right next to each chip, a last converter called a brings it below one volt.

The chain

  1. Power supplies (PSUs) convert the building’s alternating current (AC) to direct current (DC). In the Open Rack V3 design, a power shelf of 3 kW rectifiers produces about 51 V and feeds a up the back of the rack. Each rectifier must exceed 97.5% peak efficiency.
  2. The busbar carries 48 V-class DC to every server; equipment must accept 46–52 V or 52–56 V depending on the option.
  3. An intermediate converter in the server steps that down, often by a fixed 4:1 to about 12 V.
  4. Point-of-load regulators () next to each chip make the final core voltage, at or below 1 V, at hundreds of amps.

Why higher voltage on the way

Power is voltage times current, and the energy lost in a wire grows with the square of the current. A server drawing 6 kW pulls 500 A at 12 V but only about 120 A at 50 V. Quadrupling the voltage cuts the current to a quarter and the wire losses to a sixteenth. It also lets the chip’s module accept fewer, thinner power pins: one open accelerator-module standard supplies up to 700 W at about 48–54 V, against 350 W for its 12 V version.

The last inch

A 700 W accelerator at 0.8 V needs about 875 A. No single converter handles that, so the regulator is a : many identical stages in parallel, each switching in turn and each handling a share of the current. One application guide recommends keeping each phase to about 30–40 A, so a rail like this needs dozens of phases, surrounding the chip on the board.

Each stage wastes some power as heat. A 3 kW rectifier above 97% efficient loses under 100 W; a point-of-load stage may be 85–95% efficient depending on load. One research regulator converting 48 V to 1 V at 640 A peaks at 95.2% and falls to 84.4% at full load. Losses multiply down the chain, so a server’s wall power is noticeably higher than the sum of its chips’ power.

Redundancy

Power supplies fail, so servers carry spares. With redundancy, one more supply than needed is installed; with N+N, twice as many, often split across two independent feeds. One 8-GPU server uses six 3.3 kW supplies arranged 4+2. The capacity you can count on is that of the N supplies, not all of them.

Stages and their efficiencies

The chain multiplies: Pwall=Pload/(ηPoL ηIBC ηPSU)P_{\mathrm{wall}} = P_{\mathrm{load}} / (\eta_{\mathrm{PoL}}\, \eta_{\mathrm{IBC}}\, \eta_{\mathrm{PSU}}). Open Rack V3 rectifiers are specified at 3 kW, a fixed 51 V output (switchable to 48 V), peak efficiency above 97.5%, and at least N+1 redundancy per shelf. The rack base specification requires IT gear to operate from 46–52 V (51 V nominal) or 52–56 V (54 V nominal), and rates the IT-gear input connector’s power path at 100 A continuous, about 5 kW per connector at 50 V. The PSU specification’s rationale for a narrow voltage range is explicit: it “widely enable[s] 4:1 fixed ratio converters and downstream conventional 12V PoL converters.” A fixed-ratio converter only divides the voltage; holding the output steady is left to the point-of-load stage after it, which is why the bus voltage itself has to stay in a narrow range.

Distribution loss

For a conductor of resistance RR carrying power PP at voltage VV, loss=(P/V)2R\text{loss} = (P/V)^2 R. At 6 kW, 12 V means 500 A; 50 V means 120 A, and 1/17 of the loss in the same copper. The OAM pinout shows the same scaling at module level: its 48 V rail is 16 pins rated 16 A at 44 V, about 700 W; its 12 V rail is 27 pins rated 27 A, about 300 W at the 11 V minimum (the spec’s 12 V limit is 350 W).

Point-of-load regulation

Modern processors draw “hundreds of amperes of current at very low voltage (i.e., ≤ 1 V).” A parallels nn phases interleaved at 360∘/n360^\circ/n with shared input and output capacitors. Interleaving cancels ripple current, which shrinks the capacitor banks; spreading loss across phases eases thermals; and during a load step the phases act in parallel, cutting effective output inductance by nn so the controller can slew current faster. Controllers add and drop phases with load to keep efficiency high. The application note’s guideline of 30–40 A per phase gives n≈I/35 An \approx I / 35\,\mathrm{A}: about 25 phases for an 875 A rail (700 W at 0.8 V), before transient margin.

Efficiency is load-dependent. TI’s five-phase 12 V→1.8 V example holds above 90% from 5 to 200 A. Direct 48 V→1 V conversion is harder; the virtual-intermediate-bus regulator of Chen et al. (a 2:1 switched-capacitor stage to 24 V, then series-capacitor buck modules) reached 95.2% peak and 84.4% at its 640 A full load (power stage only). The drop at full load is the I2RI^2 R term in switches, inductors and board copper. It matters doubly in AI servers, whose accelerators can run near full load for long stretches of training.

Failure modes

  • Transient droop. An accelerator stepping from idle to full load demands hundreds of amps in microseconds. Too few phases or too little capacitance and the rail dips below the chip’s minimum voltage; the chip errors or throttles. On-chip, the same problem continues inside the die (see Power planning).
  • Current limits upstream. Eight accelerators synchronized by the training loop start and stop together, so the whole server’s load swings in lockstep, which the PSUs and busbar must ride through. Open Rack V3 rectifiers (3 kW rated) specify a pulse-power envelope and over-power protection that trips above 3.45 kW for 10 s or 3.6 kW for 100 ms.
  • Redundancy that isn’t. Usable capacity is (installed−redundant)×rating(\text{installed} - \text{redundant}) \times \text{rating}. A system loaded beyond N supplies has lost its redundancy without anyone noticing until a supply fails. One 8-GPU system with 4+2 supplies continues at reduced performance if three PSUs lose power, and will only boot with at least three working.
ACPSU4:1VRMchipbusbarAC≈ 50 V120 A12 V500 A0.8 V875 Aone accel.voltage falls, current rises →loss in the busbar1× (≈ 120 A)
Busbar voltage

48 V-class bus: the 6 kW server pulls about 120 A from the busbar, then the current grows at each step down. Tap a stage.

From building to chip, the voltage steps down and the current grows (wire thickness ∝ √current). The 6 kW server and 700 W accelerator are the chapter’s examples.Share freely with credit: ‘Figure from chipfieldguide.com’

Servers come in standard sizes, so they stack neatly in tall metal racks, like books on a shelf. A big AI server is about as tall as a microwave oven.

Accelerators come in standard shapes too. Some look like thick graphics cards. The biggest are flat modules screwed onto a large shared board. A group of companies called the shares open designs for these shapes. That way, parts from different makers fit together.

Racks and units

Most servers mount in racks measured in rack units (U) of 44.45 mm (1.75 inches). Open Rack V3, from the (OCP), supports both that unit and its own 48 mm “OpenU,” and replaces per-server power supplies with the shared 48 V-class busbar described above.

Accelerator modules

Accelerators began as plug-in cards, in the same kind of slot as a graphics card. As power and accelerator-to-accelerator wiring grew, the OCP defined the (OAM): a 102 × 165 mm board that mounts flat onto a baseboard through two high-speed connectors. The specification explains why: plug-in cards suffered from signal loss through connectors, cabling between cards, and limits on how cards could be wired together. A universal baseboard holds eight modules with the wiring between them built in; one 8-GPU platform puts eight such modules on one.

Network cards and drives

  • OCP NIC 3.0 defines network-card shapes: the small version carries up to 16 PCIe lanes, the large one up to 32.
  • The EDSFF “E3” family defines data-center drive shapes for 1U and 2U servers, with connectors for x4 or x8 PCIe. The smallest air-cooled version is recommended to stay within 25 W. These drives speak , the standard command set for flash over PCIe.

Cooling sets the shape too

The OAM specification recommends air cooling only up to about 450 W per module; above that, it suggests other solutions such as liquid cooling. That one line explains much about why AI servers have become taller, louder or liquid-cooled; the full story is in Power and cooling.

Why accelerators left the slot

The PCIe card electromechanical form factor was the quick path to market, but the specification lists its problems for multi-accelerator systems: “excessive signal insertion loss from ASIC to PCIe connectors and baseboard, inter-card cabling complexity reducing robustness and serviceability, and limits the supported inter-ASIC topologies.” OAM v1.5 answers with two 688-pin mezzanine connectors (rated to 56 Gb/s NRZ or 112 Gb/s PAM4), a 102 × 165 mm module, a 44–59.5 V input for up to 700 W (or 11–13.2 V for up to 350 W), one or two x16 host links, and up to seven module-to-module links that may be split into sub-links. The baseboard then fixes the scale-up topology (fully connected, hybrid cube mesh and others), which is why module and baseboard are specified together.

Open specifications that matter here

SpecWhat it fixesKey numbers
Open Rack V3Rack frame, busbar, IT-gear input48 mm OpenU or 44.45 mm RU; 46–52 V or 52–56 V; 100 A connector
ORv3 48V PSURack-level rectifiers3 kW; 51 V; > 97.5% peak; ≥ N+1
OAM 1.5Accelerator module700 W at 48/54 V; x16 host link(s); ≤ 7 peer links
OCP NIC 3.0Network adapterSFF ≤ x16, LFF ≤ x32
DC-SCM 2.0Management and security moduleBMC + hardware root of trust off the host board
EDSFF E3 (SNIA)Drive form factor1C = x4, 2C = x8; E3.S 25 W air-cooled

The common thread is modularity: each spec draws a boundary (rack, module, card, management module) so that the parts on either side can evolve on different schedules and come from different suppliers.

1U = 44.45 mm8U AI serverPSU shelfbusbarbaseboard: peer links built in8 × OAM, 102 × 165 mm each
Accelerator form factor

Tap a part to read about it, and switch between the two accelerator shapes.

Left: a rack in rack units, with a shared power shelf and busbar as in Open Rack V3. Right: a plug-in card, or eight OAM modules on a baseboard. Rack contents illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Datacenters hold so many servers that nobody can visit each one. So every server has a tiny extra computer inside, called the , with its own network connection. It reads the temperature sensors, runs the fans and keeps a log of problems.

The BMC keeps working when the main processors are off or crashed. So staff can fix a server from far away, without ever touching it.

A (BMC) is a small processor on standby power with its own network port. The open-source OpenBMC firmware project lists what a BMC typically does: power the host on and off, control fans, read sensors, keep event logs, offer a remote console (keyboard, video and mouse over the network) and update firmware.

Software talks to the BMC through standard interfaces. The older one is IPMI. The newer one, from the DMTF standards body (2015), works like a website’s data interface: a program sends a web request and gets back a structured text document (JSON) describing, say, a power supply or a temperature sensor. It was designed as a secure replacement for IPMI over the network. One 8-GPU server’s BMC, for instance, offers Redfish, IPMI, SNMP, a remote console and a web interface on its own 1 Gb Ethernet port.

The BMC is also a security boundary: it can update every firmware image in the machine, so it must be trusted. The OCP’s DC-SCM design moves the BMC and a hardware “root of trust” chip, which checks that firmware hasn’t been tampered with, onto a small separate module. One open design pairs a BMC chip running OpenBMC with a separate root-of-trust chip and uses under 20 W.

What the BMC owns

The is an out-of-band management SoC on standby power with its own NIC (or a sideband into a host NIC). OpenBMC’s feature list is a good inventory: chassis power control and power-state management, fan control, sensors, LEDs, inventory, logging, host watchdog, serial-over-LAN, remote KVM, full IPMI 2.0 with DCMI, a Redfish service and firmware update. On an accelerator server it also reads accelerator and VRM telemetry, enforces power caps, and is the first responder for thermal events.

Redfish

was introduced in 2015 as a “RESTful interface over HTTPS in JSON format based on OData v4” and “a secure, multi-node capable replacement for IPMI-over-LAN,” covering health, sensors, power supplies and thresholds, reboot and console access. Resources are linked by URI. DMTF’s public mockup of a 1U server exposes a CPU power sensor like this:

GET /redfish/v1/Chassis/1U/Sensors/CPU1Power (DMTF mockup, excerpt)json
{
  "@odata.type": "#Sensor.v1_8_1.Sensor",
  "Id": "CPU1Power",
  "ReadingType": "Power",
  "Reading": 90,
  "ReadingUnits": "W",
  "SensingInterval": "PT0.01S",
  "PhysicalContext": "CPU",
  "Thresholds": {
    "UpperCriticalUser": { "Reading": 115, "Activation": "Increasing", "DwellTime": "PT0.03S" },
    "UpperCautionUser": { "Reading": 82, "DwellTime": "PT1S" }
  },
  "@odata.id": "/redfish/v1/Chassis/1U/Sensors/CPU1Power"
}
  1. 1L2Every resource declares its schema type and version, so clients can validate it.
  2. 2L5The reading and its units: 90 W at the time of sampling.
  3. 3L7ISO 8601 duration: this sensor is sampled every 10 ms.
  4. 4L10A critical threshold that must be exceeded for 30 ms before it triggers; a power-capping policy would act on it.
  5. 5L13The resource’s own URI; collections link to members the same way.

DC-SCM: management as a module

The OCP Datacenter-ready Secure Control Module (DC-SCM) puts the BMC, a hardware root of trust (HWRoT) and supporting logic on a module separate from the host processor board, connected by a standard interface. The Project Argus implementation uses an ASPEED AST2600 BMC running OpenBMC, an AST1060 HWRoT for secure firmware authentication, recovery and update, and a CPLD for I/O bridging, at under 20 W; the same module is meant to work across many host boards and generations. The benefit is that security and management firmware are qualified once and reused, while host boards turn over with each CPU generation.

operatormgmt networkBMCalways onPSUstandbyhost: CPUs, acceleratorssensors✕ crashedfans
Host state
BMC job

The host is crashed, but the BMC runs on standby power with its own network port. Pick a job for it.

The BMC on standby power, with its own network port, does its jobs whatever the host’s state. Schematic.Share freely with credit: ‘Figure from chipfieldguide.com’

Build a server. Add accelerators, network cards and drives. Each line in the drawing is a connection to a processor. When the processors run out of wires, parts turn red and dashed. Add a second processor or switch chips to fit them in.

Then watch the power bar. If the colored bar passes the thin line marked “supply limit”, the power supplies can’t keep up. Try the “8 accel, no switch” and “8 accel, switched” buttons to compare.

Pick CPUs, lanes per CPU, PCIe generation and device counts. The lane bars show each CPU’s budget; devices that don’t fit are drawn dashed. Turn on switches to fan out lanes and watch the uplink ratio appear on the switch links. The bandwidth panel shows per-lane and per-link speeds in each direction, and the power panel adds a fixed conversion efficiency to the device power to get wall power, compared with the supplies’ capacity with one spare. Try raising the accelerator TDP until the server no longer fits its supplies.

The full model: lane budgets per root complex and per switch (illustrative 144-lane switches, x16 up and 128 lanes down), per-direction bandwidth from GT/s × coding efficiency (128b/130b, or 236/256 for flits), and a power chain Pwall=Pload/(ηVRM⋅0.975⋅0.96)P_{\mathrm{wall}} = P_{\mathrm{load}} / (\eta_{\mathrm{VRM}} \cdot 0.975 \cdot 0.96) against N+1 or N+N supply capacity. Expert controls add PCIe 7.0, VRM efficiency, a CXL Type 3 expander and a data-loading readout: drives → (switch uplinks, when data bounces through host memory) → accelerator links. On “8 accel, switched” (one switch per CPU, 8:1), raise the demand per accelerator past about 16 GB/s until the uplinks starve, then enable direct DMA.

Loading simulation…
Lanes for one accelerator
16
Speed boost per new PCIe version
2×
Power bar in the rack
≈ 50 V
Big AI server, most power
≈ 10 kW

What these numbers mean:

  • An accelerator usually gets 16 lanes, the widest standard size. A drive gets 4.
  • Each new version of PCIe doubles the speed of every lane. The newest, PCIe 7.0, was finished in 2025.
  • Racks carry about 50 volts on a metal bar. That’s ten times a phone charger, which keeps the current and the wasted heat low.
  • One server with 8 accelerators can use up to 10 kilowatts (10,000 watts). That’s like five or six electric kettles boiling at once.
PCIe 5.0 x16, one direction
≈ 63 GB/s
Open Rack V3 rectifier
3 kW, > 97.5% peak
OAM 1.5 module power
≤ 700 W
CXL expander load latency (est.)
≈ 170 ns

Bandwidth. A PCIe 5.0 x16 link carries 504 Gb/s, about 63 GB/s, in each direction after encoding overhead. That is far below an accelerator’s own memory bandwidth, which is why training keeps data on the accelerator and uses the host link mainly for loading and control.

Power. Open Rack V3 rectifiers are 3 kW each and must exceed 97.5% peak efficiency. OAM 1.5 modules can draw up to 700 W. A point-of-load regulator can be well below that efficiency at full load: one research design measured 84.4%.

Memory over CXL. The CXL designers estimate about 170 ns for a CPU to load from a directly attached memory expander, similar to reaching the other socket’s memory.

PCIe 6.0 flit: TLP bytes / total
236 / 256
PCIe 5.0 channel loss budget
36 dB
48 V→1 V, 640 A regulator: peak / full load
95.2% / 84.4%
CXL link adder (est.) / switch path
57 ns / ≈ 250 ns

Links

  • x16 per-direction bandwidth after encoding: 126 / 252 / 504 Gb/s for PCIe 3.0 / 4.0 / 5.0 (2010, 2017, 2019). PCIe 6.0: 64 GT/s PAM4, 256 B flits with 236 B TLP, 6 B DLP, 8 B CRC, 6 B FEC; FEC latency under 2 ns; expected retry time about 100 ns. PCIe 7.0: 128 GT/s, released June 2025.
  • Channel loss budgets: 22 / 28 / 36 dB at 8 / 16 / 32 GT/s; about 5 in of low-loss trace at 32 GT/s.

Data path

  • GPUDirect Storage: without it, data takes an extra copy through a in CPU memory; on some systems a direct path through a PCIe switch or a NIC acting as one “offers at least twice the peak bandwidth as compared to taking a data path through the CPU.”
  • Input pipelines: 3–24 CPU cores per GPU for pre-processing; up to 65% of epoch time in it.

CXL

  • Estimated: 57 ns link adder; ~170 ns direct-attached Type 3 load; ~250 ns through a switch. Pond’s estimate for 8–16-socket pools: +70–90 ns over NUMA-local.
  • Measured: 35%, ~2× and ~3× remote-socket DDR5 load latency across three devices.
  • Sharing wires. Switch chips let more parts connect. But parts that share a connection slow each other down when they all talk at once.
  • Faster links, shorter reach. Each new PCIe version is twice as fast, but the signal fades over a shorter distance. So boards need better materials or helper chips along the way.
  • More memory, a bit slower. CXL memory cards add space, but each trip to them takes longer.
  • Spare power supplies. A spare mostly sits idle, which costs money. But losing a server when a supply breaks costs more.
  • Hotter chips. Past a few hundred watts per accelerator, fans struggle to keep up. So servers move to liquid cooling.
ChoiceYou getYou give up
PCIe switchesMore devices than CPU lanes; direct accelerator-to-NIC and drive-to-accelerator pathsShared uplink bandwidth, extra latency, switch power and cost
Newer PCIe generationTwice the bandwidth per lane, so fewer lanes per deviceShorter reach; retimers, better board material
Two CPU socketsTwice the lanes, memory channels and coresTraffic crossing between sockets is slower; more power
CXL memory expansionMore capacity and bandwidth without more DDR channelsHigher latency; software must place data well
48 V distributionA sixteenth the wire loss of 12 V; thinner barsAn extra conversion stage on each board
N+N power suppliesSurvives losing a whole power feedHalf the installed capacity is held in reserve

What goes wrong

  • A link trains narrow or slow. A marginal connector or long trace makes a x16 Gen5 link come up as x8 or Gen4. Everything works, at half speed, and nobody notices until a benchmark comes up short.
  • Traffic crosses sockets. A NIC on one CPU and its accelerator on the other forces data across the inter-socket link, which NVIDIA’s guidance warns can severely limit peer-to-peer performance.
  • Data loading stalls. Too few CPU cores or too little drive bandwidth per accelerator leaves accelerators idle.
  • Lost redundancy. Adding accelerators or raising their power can quietly push the load past what N supplies carry; the spare is now needed just to run.

Topology trade-offs

  • Flat vs switched. Flat attachment gives every device an unshared path to host memory but exhausts root ports and pushes all peer traffic through the root complex, where forwarding between root ports is not defined by the specification. Switches buy fan-out and local peer-to-peer at the cost of ρ:1\rho{:}1 host-bound bandwidth, added latency per hop, switch power, and a single point of failure per group.
  • Generation vs reach. Each doubling tightens the loss budget relative to frequency; at 32 GT/s about 5 inches of low-loss trace is the practical limit before a retimer. Retimers add latency, power and board area; PCIe 6.0 trades 2–4 inches of reach for an FEC that stays under 2 ns.
  • Coherence vs simplicity. CXL.cache and Type 2 devices avoid explicit copies but couple the device to the host’s coherence protocol; most accelerator traffic still uses DMA with explicit synchronization.

Power trade-offs

  • Two-stage vs direct 48 V→PoL. A fixed-ratio bus converter plus 12 V PoL reuses a mature ecosystem; direct conversion removes a stage but must handle a large step-down ratio and the high input voltage stress, which is why research designs split it into merged stages.
  • Phase count. More phases mean lower per-phase loss and better transients but more board area next to the package, which competes with memory, decoupling and signal escape.
  • Redundancy vs utilization. N+N survives a feed loss but strands half the installed capacity; N+1 per shelf (Open Rack V3’s minimum) is cheaper but covers only a single rectifier failure.

Failure modes worth instrumenting

  • Link width and speed below capability (check negotiated vs capable link status on every boot), and rising correctable-error counts that precede a link dropping to a lower rate.
  • Rail droop and VRM thermal throttling during synchronized load steps; PSU load per unit above the N-supply share.
  • Data stalls: accelerator idle time correlated with input-pipeline CPU saturation or storage queue depth.
PSU ×4x16, sharedswitchPCIe 4.0 linksCPUaccel.NICaccel.NICPCIe switch+ 4 devices fit behind one x16 port− shared uplink: 4:1, ≤ 25% each
Design choices

PCIe switch: each choice buys something and pays for it somewhere else.

Design choices drawn into a small server, with what each gives and costs. Numbers from the chapter; the drawing is schematic.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

1. Link bandwidth model

Per-lane, per-direction usable bandwidth is b=Rη/8b = R\eta/8 (GB/s), with RR in GT/s and η\eta the coding or flit efficiency: 0.8 for 8b/10b, 128/130 for 128b/130b, and 236/256 if you count TLP bytes in a PCIe 6.0 flit. A link of width ww gives B=wbB = wb. Transaction-layer overhead reduces this further for small payloads, and differently in each mode: in 128b/130b mode every TLP carries its own framing, sequence number and link CRC, while flit mode amortizes them, which is why PCI-SIG reports flit mode ahead for payloads up to 512 B. The simulator uses the plain Rη/8R\eta/8 and ignores headers, flow-control credits and retries.

2. Lane allocation and oversubscription

The simulator assigns devices round-robin to ports (CPU root complexes, or switches interleaved across sockets), in order accelerators → NICs → drives, and admits a device only if its width fits the port’s remaining lanes:

allocation (as implemented in the simulator)text
for each switch k: cpu[socket(k)].used += 16            # x16 uplink
for kind in [accelerator, NIC, NVMe]:
  for i in 0 .. count(kind)-1:
    p = ports[i mod |ports|]
    root = p.isSwitch ? p : cpu[p.socket]
    if root.used + width(kind) <= capacity(root): root.used += width(kind)
    else: mark device i unconnected
rho(switch) = switch.used / 16                          # downstream / upstream lanes
hostShare   = min(1, 1 / rho)
  1. 1L1Switch uplinks are charged to the CPU first; the illustrative switch has 128 downstream lanes.
  2. 2L5Interleaving switches across sockets spreads devices evenly between CPUs.
  3. 3L9With every device streaming from host memory at once, each gets this fraction of its link.

Real platforms add constraints the model leaves out, for example root ports that split only in fixed patterns (x16, 2 × x8, 4 × x4) and lanes reserved for boot or management devices. Check the platform’s own documentation before trusting a lane count.

3. Data-loading bottleneck

Model the path from local NVMe to accelerators as stages in series, each with a capacity, and take the minimum:

  • drives: Cd=nd⋅min⁡(4b,media limit)C_{\mathrm{d}} = n_{\mathrm{d}} \cdot \min(4b, \text{media limit}) (the simulator uses 14 GB/s as an illustrative media limit);
  • switch uplinks, only when data bounces through host memory and devices sit behind switches: Cu=nsw⋅16bC_{\mathrm{u}} = n_{\mathrm{sw}} \cdot 16b per direction, because each byte goes up one uplink and down another;
  • accelerator links: Ca=na wa bC_{\mathrm{a}} = n_{\mathrm{a}} \, w_{\mathrm{a}} \, b.

Feed=min⁡(Cd,Cu,Ca)\text{Feed} = \min(C_{\mathrm{d}}, C_{\mathrm{u}}, C_{\mathrm{a}}); each accelerator receives min⁡(D,feed/na)\min(D, \text{feed}/n_{\mathrm{a}}) for a demand DD and is idle a fraction 1−that/D1 - \text{that}/D. With a , host DRAM sees 2×feed2 \times \text{feed} of traffic (one write, one read). Direct DMA removes both the uplink stage and the host-memory traffic, which is the mechanism behind GPUDirect Storage’s reported “at least twice the peak bandwidth” on switch-based systems. The model assumes drives and accelerators are spread evenly over switches, so direct traffic stays local; it ignores CPU decode cost, which Mohan et al. show is often the real limit.

4. Power chain

Pload=(∑iTDPi)(1+ffan)+PBMCPwall=PloadηVRM ηIBC ηPSU\begin{aligned} P_{\mathrm{load}} &= \Big(\sum_i \mathrm{TDP}_i\Big)(1 + f_{\mathrm{fan}}) + P_{\mathrm{BMC}} \\ P_{\mathrm{wall}} &= \frac{P_{\mathrm{load}}}{\eta_{\mathrm{VRM}}\, \eta_{\mathrm{IBC}}\, \eta_{\mathrm{PSU}}} \end{aligned}

The simulator uses illustrative values: ffan=0.06f_{\mathrm{fan}} = 0.06, PBMC=20 WP_{\mathrm{BMC}} = 20\,\mathrm{W} (the order of the DC-SCM design’s under-20 W budget), ηIBC=0.975\eta_{\mathrm{IBC}} = 0.975, ηPSU=0.96\eta_{\mathrm{PSU}} = 0.96 and a default ηVRM\eta_{\mathrm{VRM}} of 0.90. Treating ηVRM\eta_{\mathrm{VRM}} as constant is the largest simplification: real point-of-load efficiency peaks at partial load and falls at full load (95.2% vs 84.4% in Chen et al.). Supply capacity is (n−r)×3.3 kW(n - r) \times 3.3\,\mathrm{kW} with r=1r = 1 for N+1 and r=⌈n/2⌉r = \lceil n/2 \rceil for N+N.

Currents follow from I=P/VI = P/V at each stage: the readout reports the 54 V input current, the 12 V intermediate-bus current and one accelerator’s core current at 0.8 V. The phase-count estimate n≈I/Iphasen \approx I / I_{\mathrm{phase}} with Iphase≈30–40 AI_{\mathrm{phase}} \approx 30\text{–}40\,\mathrm{A} is a sizing rule of thumb; transient response, thermal limits and current-sense accuracy set the real number.

5. CXL latency, as a sum

Das Sharma et al. build the CXL.mem load latency bottom-up: CPU-side load-to-use under 100 ns (including the DRAM access the device performs), plus a CXL port round trip of 21–25 ns at each end, plus ~15 ns of flight time with retimers, which gives the 57 ns adder and an estimate near 170 ns; a switch adds 2×(21–25)+10+10≈62–70 ns2 \times (21\text{–}25) + 10 + 10 \approx 62\text{–}70\,\mathrm{ns}. The model’s assumption is that controllers hit the port-latency targets; the measured spread across real devices (35% to ~3× remote DDR5) is the size of the error when they don’t.

050100150200250300ns≈ 100 nsCPU side, incl. DRAM access: < 100 ns
1 / 5

Running total ≈ 100 ns. The CPU-side work plus the DRAM access itself.

CXL.mem load latency, built up from Das Sharma et al.’s component estimates (ranges at their midpoints, whiskers show the range). Dashed lines: their table’s end-to-end estimates. Not measurements.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 4 correct
  1. Q1A PCIe 5.0 lane carries about 3.9 GB/s in each direction. Roughly how much does a x16 PCIe 5.0 link carry in each direction?

  2. Q2Four accelerators with x16 links sit behind one PCIe switch whose uplink to the CPU is x16. What happens when all four copy data from host memory at the same time?

  3. Q3What does a CXL Type 3 device add to a server?

  4. Q4Why do newer racks distribute power at about 48–54 V instead of 12 V?

Sources

Show Hide 28 sources
  1. Introduction to NVIDIA DGX H100/H200 Systems (DGX H100/H200 System User Guide)NVIDIA documentationHardware overview: 8 GPUs, 2 Xeon 8480C CPUs, 8 single-port and 2 dual-port ConnectX-7 cards, NVMe drives, 6 × 3.3 kW PSUs in 4+2 redundancy, 10.2 kW max input, an 8U chassis, and a BMC speaking Redfish, IPMI, SNMP and KVM.
  2. AMD Instinct MI300X Platform data sheetAMD · AMD technical documentationEight MI300X OAMs on a universal baseboard (UBB 2.0), each with a PCIe Gen 5 x16 host link and seven Infinity Fabric links to its peers; 750 W per GPU.
  3. AMD Instinct MI300 Series microarchitectureAMD ROCm documentationNode-level architecture: MI300X OAMs attach to the host over PCIe Gen 5 x16 links, through optional PCIe switches and retimers, with dual-socket EPYC hosts.
  4. The Evolution of the PCI Express Specification: On its Sixth Generation, Third Decade and Still Going StrongDebendra Das Sharma · PCI-SIG blog · 2022PCIe 6.0 at 64.0 GT/s with PAM4; BER of 10⁻¹² in the first five generations; light-weight FEC under 2 ns; 256-byte flit (236 B TLP, 6 B DLP, 8 B CRC, 6 B FEC); FBER 10⁻⁶ costs 2–4 inches of reach; ~100 ns retry; CXL and NVMe run on the PCIe PHY.
  5. PCI Express 6.0 Specification at 64.0 GT/s with PAM-4 signaling: a Low Latency, High Bandwidth, High Reliability and Cost-Effective InterconnectDebendra Das Sharma · PCI-SIG (Hot Interconnects slides)Generation table with data rate, encoding (8b/10b, 128b/130b, PAM-4 flit), x16 bandwidth per direction after encoding and year; mixed-generation interoperability; latency targets and flit layout.
  6. The PCIe 7.0 Specification, Version 1.0 is Now Available to MembersPCI-SIGPCIe 7.0 at 128.0 GT/s doubles PCIe 6.x (64.0 GT/s); released to members on June 11, 2025.
  7. Retimers to the Rescue: PCI Express Specifications Reach Their Full Potential (PCI-SIG Educational Webinar)Kurt Lender and Casey Morrison · PCI-SIG · 2019Channel loss budgets of 22, 28 and 36 dB at 8, 16 and 32 GT/s; about 5 inches of low-loss board trace at PCIe 5.0 speed; retimers split the channel into two segments.
  8. The PCI Express Port Bus Driver Guide HOWTOThe Linux Kernel documentationA Root Port originates a PCIe link from the Root Complex; switch upstream and downstream ports are logical PCI-PCI bridges.
  9. PCI Peer-to-Peer DMA SupportThe Linux Kernel documentationBehind PCIe switches, peer-to-peer transactions can route entirely within the hierarchy and never reach the root port; forwarding between hierarchy domains is undefined and blocked by default.
  10. GPUDirect RDMANVIDIA CUDA documentationDirect GPU-to-peer data path over standard PCIe; devices must share an upstream root complex; a path through PCIe switches only performs best, through one CPU is worse, across CPU sockets may be severely limited.
  11. GPUDirect Storage Overview GuideNVIDIA documentationWithout GDS, data takes an extra copy through a bounce buffer in CPU memory; a direct path through a PCIe switch can offer at least twice the peak bandwidth of a path through the CPU on some systems.
  12. Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay Chidambaram · arXiv (VLDB 2021) · 2021Training time is often dominated by data stalls; DNNs need 3–24 CPU cores per GPU for pre-processing and spend up to 65% of epoch time on it.
  13. NVM Express SpecificationsNVM ExpressThe NVMe specifications define how host software talks to non-volatile memory over PCIe, RDMA, TCP and other transports.
  14. SFF-TA-1008: Enterprise and Datacenter Standard Form Factor (E3), Rev 3.0aSNIA SFF Technology Affiliate · 2026E3 drive form factors for 1U and 2U systems; 1C connector = x4 PCIe, 2C = x8; recommended 25 W max for an air-cooled E3.S device, 79.2 W connector limit at 12 V.
  15. An Introduction to the Compute Express Link (CXL) InterconnectDebendra Das Sharma, Robert Blankenship, Daniel S. Berger · arXiv · 2023CXL runs on the PCIe physical layer; CXL.io/.cache/.mem; Type 1/2/3 devices; CXL 1.0–3.0 dates; 57 ns end-to-end link adder; ~170 ns estimated CPU-to-Type-3 load latency, similar to remote-socket DDR; 250 ns through a switch.
  16. Compute Express Link 3.0 (white paper)Debendra Das Sharma and Ishwar Agarwal · CXL ConsortiumCXL 3.0 is based on PCIe 6.0, doubling the rate to 64 GT/s with no added latency, for up to 256 GB/s aggregate raw bandwidth on x16.
  17. Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, et al. · arXiv (ASPLOS 2023) · 2022Estimates CXL adds 70–90 ns over same-NUMA-node DRAM for 8–16-socket pools and more than 180 ns at rack scale; a bidirectional x8 CXL port at 2:1 read:write matches a DDR5-4800 channel.
  18. Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, et al. · arXiv (MICRO 2023) · 2023Measured three real CXL memory devices: load latency 35% to about 3× longer than remote-socket DDR5, depending on the controller; a PCIe 5.0 x8 device uses about 3× fewer pins than DDR5.
  19. Open Rack V3 Base Specification, Revision 1.0Glenn Charest, Steve Mills, Loren Vorreiter · Open Compute Project · 202248 V busbar; 48 mm OpenU or 44.45 mm EIA-310 rack units; IT gear must accept 46–52 V (51 V nominal) or 52–56 V (54 V nominal); IT-gear input connector power path rated 100 A continuous.
  20. Open Rack V3 48V PSU Specification, Rev 1.0Hamid Keyhani, Ted Tang, Dmitriy Shapiro, John Fernandes, Ben Kim, Tiffany Jin, Rommel Mercado · Open Compute Project · 2022Downloads as a .docx; the site may require a regular browser. 3 kW single-phase rectifiers with at least N+1 redundancy in the power shelf; fixed 51 V output, switchable to 48 V; peak efficiency above 97.5%; over-power protection above 3.45 kW for 10 s or 3.6 kW for 100 ms, plus a pulse-power envelope; the narrow range enables 4:1 fixed-ratio converters feeding conventional 12 V point-of-load converters.
  21. Multiphase Buck Design From Start to Finish (Part 1), SLVA882BCarmen Parisi · Texas Instruments application report · 2021Multiphase bucks are parallel phases interleaved at 360°/n; keep each phase to about 30–40 A; benefits in ripple, thermal spread and transients; a 5-phase 12 V→1.8 V design stays above 90% from 5 to 200 A.
  22. Virtual Intermediate Bus CPU Voltage RegulatorYenan Chen, Ping Wang, Hsin Cheng, Gregory Szczeszynski, Stephen Allen, David M. Giuliano, Minjie Chen · IEEE Transactions on Power Electronics (author copy, Princeton University) · 2021Processors draw hundreds of amperes at ≤ 1 V; a 48 V→1 V, 640 A two-stage regulator reaches 95.2% peak and 84.4% full-load power-stage efficiency.
  23. OCP Accelerator Module (OAM) Design Specification v1.5OCP OAI workstreams · Open Compute Project · 2022102 × 165 mm module; up to 700 W at 44–59.5 V or 350 W at 12 V; one or two x16 host links; up to 7 module-to-module links; air cooling recommended up to 450 W; host-link pins listed as 16 transmit and 16 receive differential pairs.
  24. OCP NIC 3.0 Design Specification, Version 1.00OCP Server Workgroup, OCP NIC subgroup · Open Compute Project · 2019Small form factor (SFF) cards carry up to 16 PCIe lanes, large form factor (LFF) up to 32.
  25. Project Argus DC-SCM 2.0 Module Design Specification, Rev 1.0Lenovo and Cloudflare · Open Compute Project · 2023A DC-SCM 2.0 module that separates management and security from the host board: ASPEED AST2600 BMC running OpenBMC, AST1060 hardware root of trust, under 20 W.
  26. OpenBMC features (openbmc/docs)OpenBMC project (GitHub)Feature list of the open-source BMC firmware: Redfish, full IPMI 2.0 with DCMI, remote KVM, SSH-based serial-over-LAN, fan control, chassis power control, sensors, LEDs, inventory, logging, host watchdog, firmware update.
  27. Introduction to RedfishJeff Autor · DMTF · 2015RESTful interface over HTTPS in JSON based on OData v4; a secure, multi-node capable replacement for IPMI-over-LAN; covers sensors, power supplies, power thresholds, reboot and console access.
  28. Redfish-Mockup-Server: public-rackmount1 Sensors/CPU1PowerDMTF (GitHub)DMTF’s sample Redfish Sensor resource for a CPU power reading, with thresholds.