An AI server is a box of parts that work as a team. One or two ordinary processors run the show. Several , chips built for AI math, do the heavy lifting. Network cards link the box to other servers. Fast drives hold the data, and power supplies feed it all.
Building one is like sticking to a budget. Each processor has only so many fast wires coming out of it. The power supplies can give only so much electricity. And the box can get rid of only so much heat. Adding one more accelerator spends a bit of each budget.
The previous chapter, Packaging and chiplets, ended at the edge of one package. This chapter puts packages onto boards and boards into a server. A typical AI training server combines:
- one or two , which boot the machine, run the operating system and feed the rest;
- several (often eight), each with its own high-bandwidth memory;
- network cards (), often one per accelerator, and flash drives;
- power supplies and voltage converters, fans or liquid cooling, and a management controller.
One commercial 8-accelerator server, for example, pairs eight GPUs with two CPUs, ten network cards, ten NVMe drives and six 3.3 kW power supplies, and can draw up to 10.2 kW.1 Three questions shape every such design, and they are the subject of this chapter: how the parts are wired together ( and CXL), how power gets from the rack to each chip, and how the machine is managed when something goes wrong.
A server is a set of budgets that have to close at the same time: CPU and switch ports, link bandwidth on each path the data actually takes, power from the busbar through several conversion stages, airflow or coolant, and board area. An accelerator-dense 8U system illustrates the scale: 8 GPUs, 2 CPUs, 8 single-port 400 Gb/s ConnectX-7 adapters for the compute fabric plus 2 dual-port ones for storage and management, 2 M.2 boot drives and 8 U.2 NVMe drives, 2 TB of DRAM, 6 × 3.3 kW supplies in 4+2 redundancy and a 10.2 kW maximum input.1
The chapter works through each budget in turn:
- the host–accelerator split and the PCIe hierarchy (root ports, switches, endpoints), lane budgets and oversubscription, and how per-lane rates evolved from 8b/10b to PAM4 flits;
- CXL’s three protocols and device types, and what memory expansion costs in latency;
- power delivery from a 48 V-class busbar through intermediate and point-of-load conversion to sub-1 V rails at hundreds of amps;
- module and rack form factors from the Open Compute Project, and out-of-band management with a BMC.
Accelerator-to-accelerator fabrics are only touched on here; they have their own chapter, Scale-up fabrics.
Tap a part to read its job, or pick a budget to see which parts supply it and which use it.
A server’s main processors are like the one in a laptop, just bigger. They’re good at doing lots of different jobs, one after another. Accelerators are the opposite. They do one kind of job, giant piles of multiplication, extremely fast.
So the work is split. The is the boss processor. It loads the data, gets it ready and hands it to the accelerators. Then it tells them what to compute. If the host can’t keep up, the costly accelerators sit around waiting.
Who does what
The owns the machine: its memory, the operating system and the list of attached devices. An is a device attached to it. During training, the CPU’s jobs are to read samples from storage or the network, decode and transform them (decompressing images, cropping, tokenizing text), copy batches to the accelerators, and launch and synchronize their work. The accelerators’ jobs are the matrix math, using their own memory, and the exchanges with their peers.
The split creates a pipeline, and the slowest stage sets the pace. In one study of image and audio training, models needed between 3 and 24 CPU cores per GPU just for data preparation, and spent up to 65% of each pass over the data waiting on it.12 That is why server designers balance CPU cores, memory and drive bandwidth against the number of accelerators rather than simply adding accelerators.
Two kinds of link
Accelerators have two separate connections. The host link, usually PCIe, ties each one to a CPU for data and control. A second, much faster set of links ties the accelerators to each other. One open accelerator-module design, for example, gives each module one or two 16-lane host links and up to seven links to its peers.23 The host link is the subject of this chapter; the peer links are covered in Scale-up fabrics.
Division of labor
The owns the PCIe hierarchy, the I/O memory map and the main DRAM; each is an endpoint in that hierarchy with its own HBM, its own DMA engines and usually a dedicated scale-up fabric. The host link carries input batches, checkpoints, kernel launches and synchronization, plus network traffic when NICs exchange data with accelerator memory. Gradient exchange inside the server normally rides the scale-up links instead. The OAM base specification makes the split explicit: one or two x16 host links per module, and up to seven module-to-module links of up to x16–x20 lanes each.23
The input pipeline is a host-side budget
Mohan et al. profiled nine models on 8-GPU servers and separated fetch stalls (waiting on I/O) from prep stalls (waiting on CPU decoding and augmentation). Their headline numbers: 3–24 CPU cores per GPU needed for pre-processing, up to 65% of epoch time spent in it, and an OS page cache that thrashes under the random access pattern of training.12 The implication for server design is that “balanced SKUs” matter: cores per accelerator, DRAM capacity for caching decoded data, and drive and NIC bandwidth per accelerator are all first-class parameters, not leftovers.
Host-link topologies
Three host-link patterns recur. In the flat pattern, each device takes a root port directly and lanes run out quickly. In the switched pattern, PCIe switches fan out lanes and keep accelerator, NIC and drive traffic local; vendor documentation for one 8-GPU OAM platform shows the modules reaching dual-socket hosts through optional PCIe switches and retimers.3 In the coherent pattern, a CPU and accelerator share a coherent link wider than PCIe, or the accelerator speaks CXL. The rest of this chapter quantifies the first two and introduces the third.
8 cores per accelerator, but this prep needs about 24: the CPU is the slowest stage, and the accelerator sits idle 67% of the time.
Inside the server, parts are joined by , a standard kind of super-fast wiring. The wires come in sets called . Each lane works like a two-way road: one side for sending, one for receiving. An accelerator usually gets 16 lanes, and a storage drive gets 4.
Every few years a new version of PCIe comes out. Each one doubles the speed of every lane.
The catch is that each processor has only so many lanes. When there are more parts than lanes, designers add chips. A switch lets several parts share one set of lanes, like a power strip shares one wall socket. That works well when parts take turns. If they all send at once, they slow each other down.
Lanes and links
(PCI Express) connects each device to the CPU over its own point-to-point link: no shared bus. A link is made of , and each lane is two pairs of wires, one pair sending and one receiving, so a lane carries data in both directions at full speed at the same time. An accelerator module’s host-link pinout, for instance, lists 16 transmit pairs and 16 receive pairs for its x16 link.23 The number of lanes is the : x1, x4, x8 or x16. Doubling the width doubles the bandwidth.
Generations
Each generation doubles the raw rate per lane, measured in gigatransfers per second (GT/s, roughly billions of bits per second on each wire pair). Not every transferred bit is data. Generations 1 and 2 sent 10 bits for every 8 bits of data, a 20% overhead; generations 3 to 5 send 130 bits for every 128, about 1.5%.5
| Generation | Year | Rate per lane | x16, one direction |
|---|---|---|---|
| 3.0 | 2010 | 8 GT/s | ≈ 16 GB/s |
| 4.0 | 2017 | 16 GT/s | ≈ 32 GB/s |
| 5.0 | 2019 | 32 GT/s | ≈ 63 GB/s |
| 6.0 | 2022 | 64 GT/s | ≈ 128 GB/s raw |
| 7.0 | 2025 | 128 GT/s | ≈ 256 GB/s raw |
Rates and years for 3.0–5.0 from PCI-SIG’s generation table; 6.0 from a PCI-SIG blog post and 7.0 from its specification page.546 A newer device in an older slot works: both ends agree on the fastest speed they share.5
The lane budget
Every link starts at a root port on the CPU’s , the logic that connects PCIe to the processor and its memory.8 A CPU package has a fixed number of lanes, so the first budget in a server is simple arithmetic. Eight accelerators at x16 need 128 lanes, eight x16 NICs another 128, and eight NVMe drives at x4 another 32: 288 lanes. Two CPUs with 80 usable lanes each offer 160. The design does not fit.
The usual fix is a : a chip with one upstream link to the CPU and several downstream links to devices. Four switches, each with a x16 uplink, use only 64 CPU lanes, and each can host two accelerators, two NICs and two drives (72 lanes) below it. The price is : 72 lanes below share 16 above, a 4.5:1 ratio. If all six devices pull data from host memory at once, each gets well under a quarter of its own link speed.
Often that is fine, because the busiest traffic never goes to the host. A NIC can write straight into an accelerator’s memory, and a drive can send data straight to an accelerator, when both sit behind the same switch; the traffic turns around inside the switch.910 That is why AI servers pair each accelerator with a NIC under a shared switch.
Reach and retimers
Faster signals fade faster along copper. At PCIe 5.0 speed, a good circuit board carries the signal only about 5 inches (13 cm) before it is too weak.7 Longer paths, to a riser card or the far side of a big board, need a , a chip that receives the signal and sends a fresh copy, splitting the path in two.7
Hierarchy
PCIe is a tree. Root ports on the originate links; a is an upstream port and several downstream ports, each a logical PCI-PCI bridge, joined by internal routing; endpoints sit at the leaves.8 Each is a dual-simplex pair of differential pairs, so a x16 link is 64 signal wires (an OAM host link’s pinout lists 16 TX and 16 RX pairs).23 Links train up from the lowest common rate and width; a x16 PCIe 5.0 port interoperates with a x1 device of any earlier generation.5
Per-lane bandwidth by generation
Usable per-direction bandwidth per lane is the transfer rate times the coding efficiency, divided by 8. Through PCIe 5.0 the coding is NRZ with 8b/10b (gens 1–2) or 128b/130b (gens 3–5); PCI-SIG’s table gives x16 per-direction bandwidth after encoding of 32, 64, 126, 252 and 504 Gb/s for generations 1 to 5.5 PCIe 6.0 keeps the 16 GHz Nyquist frequency of 5.0 but sends 2 bits per symbol with PAM4. The penalty is a raw error rate far above the of earlier generations, so 6.0 adds a light-weight forward error correction (FEC), designed for a first-bit-error rate of , in front of CRC and link-level retry.4 FEC needs fixed-size codewords, hence mode.
| Gen | GT/s | Coding | Efficiency | GB/s per lane per direction | x16 per direction |
|---|---|---|---|---|---|
| 3.0 | 8 | 128b/130b | 98.5% | 0.98 | 15.8 GB/s |
| 4.0 | 16 | 128b/130b | 98.5% | 1.97 | 31.5 GB/s |
| 5.0 | 32 | 128b/130b | 98.5% | 3.94 | 63.0 GB/s |
| 6.0 | 64 | PAM4, 256 B flit | 7.38 | 118 GB/s | |
| 7.0 | 128 | PAM4, 256 B flit | 92.2% | 14.8 | 236 GB/s |
Computed from the rates and flit layout in the PCI-SIG sources.546 Neither column counts transaction headers: in 128b/130b mode each packet also carries framing, sequence number and link CRC, which flit mode removes. PCI-SIG’s analysis puts flit mode ahead of 128b/130b in packet efficiency for payloads up to 512 bytes, for up to about 3× the effective throughput of 5.0 on small transfers, falling to about 0.98 efficiency at 4 KB payloads.4 Headline figures such as “256 GB/s for x16” at 64 GT/s are raw and bidirectional.16
Lane budgets and switches
Write the budget per root complex as over directly attached devices, and per switch as , with the switch’s uplink width charged to its CPU. The oversubscription ratio is . For host-bound streaming with all devices active, each device gets of its link. Peer-to-peer traffic between devices under one switch does not touch the uplink; when access control allows, “the transaction can route entirely within the PCIe hierarchy and never reach the root port,” while forwarding between root ports is not defined by the specification and the Linux kernel blocks it by default.9 NVIDIA’s GPUDirect RDMA guidance ranks the paths the same way: PCIe switches only is best, a path through one CPU is worse, and a path that crosses the inter-socket link may be severely limited.10 Two practical consequences:
- Pair each accelerator with its NIC (and ideally its drives) under one switch, so RDMA and storage DMA stay local.
- Spread switches across both sockets, and keep each process on the socket that owns its devices, so host staging doesn’t cross the socket link.
-[0000:00]-+-01.0-[01-06]----00.0-[02-06]--+-00.0-[03]----00.0 Accelerator (x16)
| +-01.0-[04]----00.0 NIC, 400G (x16)
| +-02.0-[05]----00.0 NVMe SSD (x4)
| \-03.0-[06]----00.0 NVMe SSD (x4)
\-03.0-[07]----00.0 CXL memory expander (x16)- 1L1Root port 01.0 leads to a switch upstream port (bus 01) whose internal bus 02 fans out to four downstream ports.
- 2L2Accelerator and NIC share the switch: RDMA between them can turn around inside it.
- 3L5A device on its own root port: traffic to or from it always passes through the root complex.
Signal reach
Loss budgets grew from 22 dB at 8 GT/s to 28 dB at 16 GT/s and 36 dB at 32 GT/s, yet at 32 GT/s low-loss board material still allows only about 5 inches of system-board trace.7 A splits the channel into two segments, each with a full budget.7 PCIe 6.0 holds reach roughly at 5.0 levels: PAM4 keeps the Nyquist frequency, though choosing a FBER rather than networking’s costs about 2–4 inches; in return FEC latency stays under 2 ns.4
x16 PCIe 5.0: 32 GT/s per lane, ≈ 3.9 GB/s per lane each way after encoding overhead, ≈ 63 GB/s per direction.
Peer-to-peer: the data turns around inside the switch. It uses none of the CPU’s lanes and none of its memory, which is why each accelerator gets a NIC under the same switch.
A server’s main memory plugs into slots next to the processor, and there’s only room for so many. AI jobs often want more. is a newer standard that uses the same PCIe wires to add memory on a plug-in card. The processor can use that memory as if it were its own.
The trade-off is distance. Memory on a CXL card is a bit farther away. So each trip to it takes a little longer. For data that isn’t needed all the time, that’s a fine deal.
(Compute Express Link) is an open standard that runs on the same physical lanes as PCIe. A CXL device starts out as an ordinary PCIe device and then negotiates CXL.15 On top of the PCIe-style protocol (CXL.io) it adds two more:15
- CXL.cache lets a device keep a cached copy of the CPU’s memory, with the hardware keeping the copies consistent.
- CXL.mem lets the CPU read and write a device’s memory with ordinary load and store instructions, as if it were main memory.
Devices come in three types, depending on which protocols they use:15
| Type | Example | Protocols |
|---|---|---|
| Type 1 | Smart NIC without its own memory pool | CXL.io + CXL.cache |
| Type 2 | GPU or FPGA with local memory | CXL.io + CXL.cache + CXL.mem |
| Type 3 | Memory expander | CXL.io + CXL.mem |
Today the most common use in servers is the : a card full of DRAM that adds capacity without adding memory channels to the CPU. A device using 8 PCIe 5.0 lanes needs about a third as many pins as a DDR5 memory channel,18 and at a typical mix of reads and writes it carries about as much as one DDR5-4800 channel.17
The cost is latency. The CXL designers estimate a CPU load from a directly attached expander at around 170 ns, about the same as reaching memory attached to the other CPU socket.15 Measured devices vary: in one study, load latency ranged from 35% longer to about three times longer than that remote-socket memory, depending on the controller chip.18 Software therefore treats CXL memory as a slower tier, a separate node with no processors of its own.
The standard has moved quickly. CXL 1.0 and 1.1 (2019) connected one host to its devices; 2.0 (2020) added a switch level so memory can be pooled among several hosts; 3.0 (2022) doubled the rate to 64 GT/s and allowed larger fabrics.15
Protocols and device types
multiplexes three protocols on the PCIe PHY. Link training starts at 2.5 GT/s as PCIe, then negotiates CXL through alternate protocol negotiation. CXL.io reuses PCIe semantics for discovery, configuration, address translation and DMA; CXL.cache lets a device coherently cache host memory; CXL.mem lets the host (and other devices) load and store to device-attached memory as cacheable “host-managed device memory.” Type 1 devices implement .io + .cache, Type 2 all three, Type 3 .io + .mem.15 CXL 1.x/2.0 use a 68-byte flit (2 B protocol ID, 64 B payload, 2 B CRC); CXL 3.0 adopts PCIe 6.0’s 256-byte flit at 64 GT/s with no added latency, plus a latency-optimized variant.1516
| Version | Released | Rate | Scope |
|---|---|---|---|
| 1.0 / 1.1 | Mar / Sep 2019 | 32 GT/s | single host |
| 2.0 | Nov 2020 | 32 GT/s | one switch level; pooling across ~2–16 hosts; hot-plug |
| 3.0 | Aug 2022 | 64 GT/s | multi-level switching, fabrics, memory sharing |
From Das Sharma, Blankenship and Berger.15
Latency budget
The CXL port costs about 21–25 ns round trip on each side, and with about 15 ns of flight time through retimers the designers estimate a 57 ns end-to-end adder for a memory access across a link; the specification’s pin-to-pin target is 80 ns for a CXL.mem access. Their table estimates about 170 ns load latency from a CPU to a directly attached Type 3 device and about 250 ns through one CXL switch, and describes the direct case as similar to a remote-socket DDR access.15 Pond, a memory-pooling design for a public cloud, estimates 70–90 ns over NUMA-local DRAM for pools of 8–16 sockets and more than 180 ns for rack-scale pools.17 Measured hardware spreads widely: on a 4th-generation Xeon host, three CXL devices showed 35%, about 2× and about 3× the load latency of remote-socket DDR5, and the authors also found that emulating CXL with remote-socket memory can misstate latency.18
Bandwidth and pins
A x16 CXL link at 32 GT/s carries 64 GB/s raw per direction (128 GB/s at 64 GT/s).15 Pond notes that a bidirectional x8 port at a 2:1 read:write mix matches one DDR5-4800 channel,17 and Sun et al. note a PCIe 5.0 x8 memory device uses about 3× fewer pins than a DDR5 channel.18 The pin argument is the strategic one: CPU packages are pin-limited, and serial lanes deliver more bandwidth per pin than a parallel DDR bus, at the cost of latency.
What it is used for
- Capacity tiering. Hot pages in local DDR, colder pages on a expander, managed by the OS’s policies.
- Pooling. Memory that would be stranded on one host is lent to another without a reboot; Pond reports a 7% DRAM cost reduction with performance within 1–5% of NUMA-local allocations.17
- Coherent accelerators. Type 2 devices share data with the host without explicit copies. On today’s AI servers the bulk of accelerator data still moves by DMA over PCIe semantics, so this is the least deployed of the three.
A Type 3 memory expander on PCIe lanes: more capacity with no new memory channels. Estimated ≈ 170 ns per load; measured devices were 35% to about 3× slower than other-socket memory.
Chips run on less than one volt. Voltage is the push that makes electricity flow, and a phone charger gives about 5 volts. Yet a big AI chip uses as much power as a few space heaters. With such a weak push, it needs a huge flow of electricity to get that much power. That flow is called current.
Sending a big current down a long wire wastes energy as heat. So power travels most of the way at a higher voltage, and is turned down in steps. A thick metal in the rack carries about 50 volts. Right next to each chip, a last converter called a brings it below one volt.
The chain
- Power supplies (PSUs) convert the building’s alternating current (AC) to direct current (DC). In the Open Rack V3 design, a power shelf of 3 kW rectifiers produces about 51 V and feeds a up the back of the rack. Each rectifier must exceed 97.5% peak efficiency.20
- The busbar carries 48 V-class DC to every server; equipment must accept 46–52 V or 52–56 V depending on the option.19
- An intermediate converter in the server steps that down, often by a fixed 4:1 to about 12 V.20
- Point-of-load regulators () next to each chip make the final core voltage, at or below 1 V, at hundreds of amps.22
Why higher voltage on the way
Power is voltage times current, and the energy lost in a wire grows with the square of the current. A server drawing 6 kW pulls 500 A at 12 V but only about 120 A at 50 V. Quadrupling the voltage cuts the current to a quarter and the wire losses to a sixteenth. It also lets the chip’s module accept fewer, thinner power pins: one open accelerator-module standard supplies up to 700 W at about 48–54 V, against 350 W for its 12 V version.23
The last inch
A 700 W accelerator at 0.8 V needs about 875 A. No single converter handles that, so the regulator is a : many identical stages in parallel, each switching in turn and each handling a share of the current. One application guide recommends keeping each phase to about 30–40 A,21 so a rail like this needs dozens of phases, surrounding the chip on the board.
Each stage wastes some power as heat. A 3 kW rectifier above 97% efficient loses under 100 W; a point-of-load stage may be 85–95% efficient depending on load. One research regulator converting 48 V to 1 V at 640 A peaks at 95.2% and falls to 84.4% at full load.22 Losses multiply down the chain, so a server’s wall power is noticeably higher than the sum of its chips’ power.
Redundancy
Power supplies fail, so servers carry spares. With redundancy, one more supply than needed is installed; with N+N, twice as many, often split across two independent feeds. One 8-GPU server uses six 3.3 kW supplies arranged 4+2.1 The capacity you can count on is that of the N supplies, not all of them.
Stages and their efficiencies
The chain multiplies: . Open Rack V3 rectifiers are specified at 3 kW, a fixed 51 V output (switchable to 48 V), peak efficiency above 97.5%, and at least N+1 redundancy per shelf.20 The rack base specification requires IT gear to operate from 46–52 V (51 V nominal) or 52–56 V (54 V nominal), and rates the IT-gear input connector’s power path at 100 A continuous, about 5 kW per connector at 50 V.19 The PSU specification’s rationale for a narrow voltage range is explicit: it “widely enable[s] 4:1 fixed ratio converters and downstream conventional 12V PoL converters.”20 A fixed-ratio converter only divides the voltage; holding the output steady is left to the point-of-load stage after it, which is why the bus voltage itself has to stay in a narrow range.
Distribution loss
For a conductor of resistance carrying power at voltage , . At 6 kW, 12 V means 500 A; 50 V means 120 A, and 1/17 of the loss in the same copper. The OAM pinout shows the same scaling at module level: its 48 V rail is 16 pins rated 16 A at 44 V, about 700 W; its 12 V rail is 27 pins rated 27 A, about 300 W at the 11 V minimum (the spec’s 12 V limit is 350 W).23
Point-of-load regulation
Modern processors draw “hundreds of amperes of current at very low voltage (i.e., ≤ 1 V).”22 A parallels phases interleaved at with shared input and output capacitors. Interleaving cancels ripple current, which shrinks the capacitor banks; spreading loss across phases eases thermals; and during a load step the phases act in parallel, cutting effective output inductance by so the controller can slew current faster. Controllers add and drop phases with load to keep efficiency high.21 The application note’s guideline of 30–40 A per phase21 gives : about 25 phases for an 875 A rail (700 W at 0.8 V), before transient margin.
Efficiency is load-dependent. TI’s five-phase 12 V→1.8 V example holds above 90% from 5 to 200 A.21 Direct 48 V→1 V conversion is harder; the virtual-intermediate-bus regulator of Chen et al. (a 2:1 switched-capacitor stage to 24 V, then series-capacitor buck modules) reached 95.2% peak and 84.4% at its 640 A full load (power stage only).22 The drop at full load is the term in switches, inductors and board copper. It matters doubly in AI servers, whose accelerators can run near full load for long stretches of training.
Failure modes
- Transient droop. An accelerator stepping from idle to full load demands hundreds of amps in microseconds. Too few phases or too little capacitance and the rail dips below the chip’s minimum voltage; the chip errors or throttles. On-chip, the same problem continues inside the die (see Power planning).
- Current limits upstream. Eight accelerators synchronized by the training loop start and stop together, so the whole server’s load swings in lockstep, which the PSUs and busbar must ride through. Open Rack V3 rectifiers (3 kW rated) specify a pulse-power envelope and over-power protection that trips above 3.45 kW for 10 s or 3.6 kW for 100 ms.20
- Redundancy that isn’t. Usable capacity is . A system loaded beyond N supplies has lost its redundancy without anyone noticing until a supply fails. One 8-GPU system with 4+2 supplies continues at reduced performance if three PSUs lose power, and will only boot with at least three working.1
48 V-class bus: the 6 kW server pulls about 120 A from the busbar, then the current grows at each step down. Tap a stage.
Servers come in standard sizes, so they stack neatly in tall metal racks, like books on a shelf. A big AI server is about as tall as a microwave oven.1
Accelerators come in standard shapes too. Some look like thick graphics cards. The biggest are flat modules screwed onto a large shared board. A group of companies called the shares open designs for these shapes. That way, parts from different makers fit together.
Racks and units
Most servers mount in racks measured in rack units (U) of 44.45 mm (1.75 inches). Open Rack V3, from the (OCP), supports both that unit and its own 48 mm “OpenU,” and replaces per-server power supplies with the shared 48 V-class busbar described above.19
Accelerator modules
Accelerators began as plug-in cards, in the same kind of slot as a graphics card. As power and accelerator-to-accelerator wiring grew, the OCP defined the (OAM): a 102 × 165 mm board that mounts flat onto a baseboard through two high-speed connectors. The specification explains why: plug-in cards suffered from signal loss through connectors, cabling between cards, and limits on how cards could be wired together.23 A universal baseboard holds eight modules with the wiring between them built in; one 8-GPU platform puts eight such modules on one.2
Network cards and drives
- OCP NIC 3.0 defines network-card shapes: the small version carries up to 16 PCIe lanes, the large one up to 32.24
- The EDSFF “E3” family defines data-center drive shapes for 1U and 2U servers, with connectors for x4 or x8 PCIe. The smallest air-cooled version is recommended to stay within 25 W.14 These drives speak , the standard command set for flash over PCIe.13
Cooling sets the shape too
The OAM specification recommends air cooling only up to about 450 W per module; above that, it suggests other solutions such as liquid cooling.23 That one line explains much about why AI servers have become taller, louder or liquid-cooled; the full story is in Power and cooling.
Why accelerators left the slot
The PCIe card electromechanical form factor was the quick path to market, but the specification lists its problems for multi-accelerator systems: “excessive signal insertion loss from ASIC to PCIe connectors and baseboard, inter-card cabling complexity reducing robustness and serviceability, and limits the supported inter-ASIC topologies.”23 OAM v1.5 answers with two 688-pin mezzanine connectors (rated to 56 Gb/s NRZ or 112 Gb/s PAM4), a 102 × 165 mm module, a 44–59.5 V input for up to 700 W (or 11–13.2 V for up to 350 W), one or two x16 host links, and up to seven module-to-module links that may be split into sub-links.23 The baseboard then fixes the scale-up topology (fully connected, hybrid cube mesh and others), which is why module and baseboard are specified together.
Open specifications that matter here
| Spec | What it fixes | Key numbers |
|---|---|---|
| Open Rack V3 | Rack frame, busbar, IT-gear input | 48 mm OpenU or 44.45 mm RU; 46–52 V or 52–56 V; 100 A connector19 |
| ORv3 48V PSU | Rack-level rectifiers | 3 kW; 51 V; > 97.5% peak; ≥ N+120 |
| OAM 1.5 | Accelerator module | 700 W at 48/54 V; x16 host link(s); ≤ 7 peer links23 |
| OCP NIC 3.0 | Network adapter | SFF ≤ x16, LFF ≤ x3224 |
| DC-SCM 2.0 | Management and security module | BMC + hardware root of trust off the host board25 |
| EDSFF E3 (SNIA) | Drive form factor | 1C = x4, 2C = x8; E3.S 25 W air-cooled14 |
The common thread is modularity: each spec draws a boundary (rack, module, card, management module) so that the parts on either side can evolve on different schedules and come from different suppliers.
Tap a part to read about it, and switch between the two accelerator shapes.
Datacenters hold so many servers that nobody can visit each one. So every server has a tiny extra computer inside, called the , with its own network connection. It reads the temperature sensors, runs the fans and keeps a log of problems.
The BMC keeps working when the main processors are off or crashed. So staff can fix a server from far away, without ever touching it.
A (BMC) is a small processor on standby power with its own network port. The open-source OpenBMC firmware project lists what a BMC typically does: power the host on and off, control fans, read sensors, keep event logs, offer a remote console (keyboard, video and mouse over the network) and update firmware.26
Software talks to the BMC through standard interfaces. The older one is IPMI. The newer one, from the DMTF standards body (2015), works like a website’s data interface: a program sends a web request and gets back a structured text document (JSON) describing, say, a power supply or a temperature sensor. It was designed as a secure replacement for IPMI over the network.27 One 8-GPU server’s BMC, for instance, offers Redfish, IPMI, SNMP, a remote console and a web interface on its own 1 Gb Ethernet port.1
The BMC is also a security boundary: it can update every firmware image in the machine, so it must be trusted. The OCP’s DC-SCM design moves the BMC and a hardware “root of trust” chip, which checks that firmware hasn’t been tampered with, onto a small separate module. One open design pairs a BMC chip running OpenBMC with a separate root-of-trust chip and uses under 20 W.25
What the BMC owns
The is an out-of-band management SoC on standby power with its own NIC (or a sideband into a host NIC). OpenBMC’s feature list is a good inventory: chassis power control and power-state management, fan control, sensors, LEDs, inventory, logging, host watchdog, serial-over-LAN, remote KVM, full IPMI 2.0 with DCMI, a Redfish service and firmware update.26 On an accelerator server it also reads accelerator and VRM telemetry, enforces power caps, and is the first responder for thermal events.
Redfish
was introduced in 2015 as a “RESTful interface over HTTPS in JSON format based on OData v4” and “a secure, multi-node capable replacement for IPMI-over-LAN,” covering health, sensors, power supplies and thresholds, reboot and console access.27 Resources are linked by URI. DMTF’s public mockup of a 1U server exposes a CPU power sensor like this:28
{
"@odata.type": "#Sensor.v1_8_1.Sensor",
"Id": "CPU1Power",
"ReadingType": "Power",
"Reading": 90,
"ReadingUnits": "W",
"SensingInterval": "PT0.01S",
"PhysicalContext": "CPU",
"Thresholds": {
"UpperCriticalUser": { "Reading": 115, "Activation": "Increasing", "DwellTime": "PT0.03S" },
"UpperCautionUser": { "Reading": 82, "DwellTime": "PT1S" }
},
"@odata.id": "/redfish/v1/Chassis/1U/Sensors/CPU1Power"
}- 1L2Every resource declares its schema type and version, so clients can validate it.
- 2L5The reading and its units: 90 W at the time of sampling.
- 3L7ISO 8601 duration: this sensor is sampled every 10 ms.
- 4L10A critical threshold that must be exceeded for 30 ms before it triggers; a power-capping policy would act on it.
- 5L13The resource’s own URI; collections link to members the same way.
DC-SCM: management as a module
The OCP Datacenter-ready Secure Control Module (DC-SCM) puts the BMC, a hardware root of trust (HWRoT) and supporting logic on a module separate from the host processor board, connected by a standard interface. The Project Argus implementation uses an ASPEED AST2600 BMC running OpenBMC, an AST1060 HWRoT for secure firmware authentication, recovery and update, and a CPLD for I/O bridging, at under 20 W; the same module is meant to work across many host boards and generations.25 The benefit is that security and management firmware are qualified once and reused, while host boards turn over with each CPU generation.
The host is crashed, but the BMC runs on standby power with its own network port. Pick a job for it.
Build a server. Add accelerators, network cards and drives. Each line in the drawing is a connection to a processor. When the processors run out of wires, parts turn red and dashed. Add a second processor or switch chips to fit them in.
Then watch the power bar. If the colored bar passes the thin line marked “supply limit”, the power supplies can’t keep up. Try the “8 accel, no switch” and “8 accel, switched” buttons to compare.
Pick CPUs, lanes per CPU, PCIe generation and device counts. The lane bars show each CPU’s budget; devices that don’t fit are drawn dashed. Turn on switches to fan out lanes and watch the uplink ratio appear on the switch links. The bandwidth panel shows per-lane and per-link speeds in each direction, and the power panel adds a fixed conversion efficiency to the device power to get wall power, compared with the supplies’ capacity with one spare. Try raising the accelerator TDP until the server no longer fits its supplies.
The full model: lane budgets per root complex and per switch (illustrative 144-lane switches, x16 up and 128 lanes down), per-direction bandwidth from GT/s × coding efficiency (128b/130b, or 236/256 for flits), and a power chain against N+1 or N+N supply capacity. Expert controls add PCIe 7.0, VRM efficiency, a CXL Type 3 expander and a data-loading readout: drives → (switch uplinks, when data bounces through host memory) → accelerator links. On “8 accel, switched” (one switch per CPU, 8:1), raise the demand per accelerator past about 16 GB/s until the uplinks starve, then enable direct DMA.
- Lanes for one accelerator
- 16
- Speed boost per new PCIe version
- 2×
- Power bar in the rack
- ≈ 50 V
- Big AI server, most power
- ≈ 10 kW
What these numbers mean:
- An accelerator usually gets 16 lanes, the widest standard size. A drive gets 4.23
- Each new version of PCIe doubles the speed of every lane. The newest, PCIe 7.0, was finished in 2025.6
- Racks carry about 50 volts on a metal bar. That’s ten times a phone charger, which keeps the current and the wasted heat low.19
- One server with 8 accelerators can use up to 10 kilowatts (10,000 watts). That’s like five or six electric kettles boiling at once.1
- PCIe 5.0 x16, one direction
- ≈ 63 GB/s
- Open Rack V3 rectifier
- 3 kW, > 97.5% peak
- OAM 1.5 module power
- ≤ 700 W
- CXL expander load latency (est.)
- ≈ 170 ns
Bandwidth. A PCIe 5.0 x16 link carries 504 Gb/s, about 63 GB/s, in each direction after encoding overhead.5 That is far below an accelerator’s own memory bandwidth, which is why training keeps data on the accelerator and uses the host link mainly for loading and control.
Power. Open Rack V3 rectifiers are 3 kW each and must exceed 97.5% peak efficiency.20 OAM 1.5 modules can draw up to 700 W.23 A point-of-load regulator can be well below that efficiency at full load: one research design measured 84.4%.22
Memory over CXL. The CXL designers estimate about 170 ns for a CPU to load from a directly attached memory expander, similar to reaching the other socket’s memory.15
- PCIe 6.0 flit: TLP bytes / total
- 236 / 256
- PCIe 5.0 channel loss budget
- 36 dB
- 48 V→1 V, 640 A regulator: peak / full load
- 95.2% / 84.4%
- CXL link adder (est.) / switch path
- 57 ns / ≈ 250 ns
Links
- x16 per-direction bandwidth after encoding: 126 / 252 / 504 Gb/s for PCIe 3.0 / 4.0 / 5.0 (2010, 2017, 2019).5 PCIe 6.0: 64 GT/s PAM4, 256 B flits with 236 B TLP, 6 B DLP, 8 B CRC, 6 B FEC; FEC latency under 2 ns; expected retry time about 100 ns.4 PCIe 7.0: 128 GT/s, released June 2025.6
- Channel loss budgets: 22 / 28 / 36 dB at 8 / 16 / 32 GT/s; about 5 in of low-loss trace at 32 GT/s.7
Data path
- GPUDirect Storage: without it, data takes an extra copy through a in CPU memory; on some systems a direct path through a PCIe switch or a NIC acting as one “offers at least twice the peak bandwidth as compared to taking a data path through the CPU.”11
- Input pipelines: 3–24 CPU cores per GPU for pre-processing; up to 65% of epoch time in it.12
CXL
- Sharing wires. Switch chips let more parts connect. But parts that share a connection slow each other down when they all talk at once.
- Faster links, shorter reach. Each new PCIe version is twice as fast, but the signal fades over a shorter distance. So boards need better materials or helper chips along the way.7
- More memory, a bit slower. CXL memory cards add space, but each trip to them takes longer.15
- Spare power supplies. A spare mostly sits idle, which costs money. But losing a server when a supply breaks costs more.
- Hotter chips. Past a few hundred watts per accelerator, fans struggle to keep up. So servers move to liquid cooling.23
| Choice | You get | You give up |
|---|---|---|
| PCIe switches | More devices than CPU lanes; direct accelerator-to-NIC and drive-to-accelerator paths | Shared uplink bandwidth, extra latency, switch power and cost |
| Newer PCIe generation | Twice the bandwidth per lane, so fewer lanes per device | Shorter reach; retimers, better board material |
| Two CPU sockets | Twice the lanes, memory channels and cores | Traffic crossing between sockets is slower; more power |
| CXL memory expansion | More capacity and bandwidth without more DDR channels | Higher latency; software must place data well |
| 48 V distribution | A sixteenth the wire loss of 12 V; thinner bars | An extra conversion stage on each board |
| N+N power supplies | Survives losing a whole power feed | Half the installed capacity is held in reserve |
What goes wrong
- A link trains narrow or slow. A marginal connector or long trace makes a x16 Gen5 link come up as x8 or Gen4. Everything works, at half speed, and nobody notices until a benchmark comes up short.
- Traffic crosses sockets. A NIC on one CPU and its accelerator on the other forces data across the inter-socket link, which NVIDIA’s guidance warns can severely limit peer-to-peer performance.10
- Data loading stalls. Too few CPU cores or too little drive bandwidth per accelerator leaves accelerators idle.12
- Lost redundancy. Adding accelerators or raising their power can quietly push the load past what N supplies carry; the spare is now needed just to run.
Topology trade-offs
- Flat vs switched. Flat attachment gives every device an unshared path to host memory but exhausts root ports and pushes all peer traffic through the root complex, where forwarding between root ports is not defined by the specification.9 Switches buy fan-out and local peer-to-peer at the cost of host-bound bandwidth, added latency per hop, switch power, and a single point of failure per group.
- Generation vs reach. Each doubling tightens the loss budget relative to frequency; at 32 GT/s about 5 inches of low-loss trace is the practical limit before a retimer.7 Retimers add latency, power and board area; PCIe 6.0 trades 2–4 inches of reach for an FEC that stays under 2 ns.4
- Coherence vs simplicity. CXL.cache and Type 2 devices avoid explicit copies but couple the device to the host’s coherence protocol; most accelerator traffic still uses DMA with explicit synchronization.
Power trade-offs
- Two-stage vs direct 48 V→PoL. A fixed-ratio bus converter plus 12 V PoL reuses a mature ecosystem;20 direct conversion removes a stage but must handle a large step-down ratio and the high input voltage stress, which is why research designs split it into merged stages.22
- Phase count. More phases mean lower per-phase loss and better transients but more board area next to the package, which competes with memory, decoupling and signal escape.21
- Redundancy vs utilization. N+N survives a feed loss but strands half the installed capacity; N+1 per shelf (Open Rack V3’s minimum) is cheaper but covers only a single rectifier failure.20
Failure modes worth instrumenting
- Link width and speed below capability (check negotiated vs capable link status on every boot), and rising correctable-error counts that precede a link dropping to a lower rate.
- Rail droop and VRM thermal throttling during synchronized load steps; PSU load per unit above the N-supply share.
- Data stalls: accelerator idle time correlated with input-pipeline CPU saturation or storage queue depth.12
PCIe switch: each choice buys something and pays for it somewhere else.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
1. Link bandwidth model
Per-lane, per-direction usable bandwidth is (GB/s), with in GT/s and the coding or flit efficiency: 0.8 for 8b/10b, 128/130 for 128b/130b, and 236/256 if you count TLP bytes in a PCIe 6.0 flit.54 A link of width gives . Transaction-layer overhead reduces this further for small payloads, and differently in each mode: in 128b/130b mode every TLP carries its own framing, sequence number and link CRC, while flit mode amortizes them, which is why PCI-SIG reports flit mode ahead for payloads up to 512 B.4 The simulator uses the plain and ignores headers, flow-control credits and retries.
2. Lane allocation and oversubscription
The simulator assigns devices round-robin to ports (CPU root complexes, or switches interleaved across sockets), in order accelerators → NICs → drives, and admits a device only if its width fits the port’s remaining lanes:
for each switch k: cpu[socket(k)].used += 16 # x16 uplink
for kind in [accelerator, NIC, NVMe]:
for i in 0 .. count(kind)-1:
p = ports[i mod |ports|]
root = p.isSwitch ? p : cpu[p.socket]
if root.used + width(kind) <= capacity(root): root.used += width(kind)
else: mark device i unconnected
rho(switch) = switch.used / 16 # downstream / upstream lanes
hostShare = min(1, 1 / rho)- 1L1Switch uplinks are charged to the CPU first; the illustrative switch has 128 downstream lanes.
- 2L5Interleaving switches across sockets spreads devices evenly between CPUs.
- 3L9With every device streaming from host memory at once, each gets this fraction of its link.
Real platforms add constraints the model leaves out, for example root ports that split only in fixed patterns (x16, 2 × x8, 4 × x4) and lanes reserved for boot or management devices. Check the platform’s own documentation before trusting a lane count.
3. Data-loading bottleneck
Model the path from local NVMe to accelerators as stages in series, each with a capacity, and take the minimum:
- drives: (the simulator uses 14 GB/s as an illustrative media limit);
- switch uplinks, only when data bounces through host memory and devices sit behind switches: per direction, because each byte goes up one uplink and down another;
- accelerator links: .
; each accelerator receives for a demand and is idle a fraction . With a , host DRAM sees of traffic (one write, one read). Direct DMA removes both the uplink stage and the host-memory traffic, which is the mechanism behind GPUDirect Storage’s reported “at least twice the peak bandwidth” on switch-based systems.11 The model assumes drives and accelerators are spread evenly over switches, so direct traffic stays local; it ignores CPU decode cost, which Mohan et al. show is often the real limit.12
4. Power chain
The simulator uses illustrative values: , (the order of the DC-SCM design’s under-20 W budget25), , and a default of 0.90. Treating as constant is the largest simplification: real point-of-load efficiency peaks at partial load and falls at full load (95.2% vs 84.4% in Chen et al.).22 Supply capacity is with for N+1 and for N+N.
Currents follow from at each stage: the readout reports the 54 V input current, the 12 V intermediate-bus current and one accelerator’s core current at 0.8 V. The phase-count estimate with 21 is a sizing rule of thumb; transient response, thermal limits and current-sense accuracy set the real number.
5. CXL latency, as a sum
Das Sharma et al. build the CXL.mem load latency bottom-up: CPU-side load-to-use under 100 ns (including the DRAM access the device performs), plus a CXL port round trip of 21–25 ns at each end, plus ~15 ns of flight time with retimers, which gives the 57 ns adder and an estimate near 170 ns; a switch adds .15 The model’s assumption is that controllers hit the port-latency targets; the measured spread across real devices (35% to ~3× remote DDR5)18 is the size of the error when they don’t.
Running total ≈ 100 ns. The CPU-side work plus the DRAM access itself.
Q1A PCIe 5.0 lane carries about 3.9 GB/s in each direction. Roughly how much does a x16 PCIe 5.0 link carry in each direction?
Q2Four accelerators with x16 links sit behind one PCIe switch whose uplink to the CPU is x16. What happens when all four copy data from host memory at the same time?
Q3What does a CXL Type 3 device add to a server?
Q4Why do newer racks distribute power at about 48–54 V instead of 12 V?
Sources
Show Hide 28 sources
- Introduction to NVIDIA DGX H100/H200 Systems (DGX H100/H200 System User Guide)Hardware overview: 8 GPUs, 2 Xeon 8480C CPUs, 8 single-port and 2 dual-port ConnectX-7 cards, NVMe drives, 6 × 3.3 kW PSUs in 4+2 redundancy, 10.2 kW max input, an 8U chassis, and a BMC speaking Redfish, IPMI, SNMP and KVM.
- AMD Instinct MI300X Platform data sheetEight MI300X OAMs on a universal baseboard (UBB 2.0), each with a PCIe Gen 5 x16 host link and seven Infinity Fabric links to its peers; 750 W per GPU.
- AMD Instinct MI300 Series microarchitectureNode-level architecture: MI300X OAMs attach to the host over PCIe Gen 5 x16 links, through optional PCIe switches and retimers, with dual-socket EPYC hosts.
- The Evolution of the PCI Express Specification: On its Sixth Generation, Third Decade and Still Going StrongPCIe 6.0 at 64.0 GT/s with PAM4; BER of 10⁻¹² in the first five generations; light-weight FEC under 2 ns; 256-byte flit (236 B TLP, 6 B DLP, 8 B CRC, 6 B FEC); FBER 10⁻⁶ costs 2–4 inches of reach; ~100 ns retry; CXL and NVMe run on the PCIe PHY.
- PCI Express 6.0 Specification at 64.0 GT/s with PAM-4 signaling: a Low Latency, High Bandwidth, High Reliability and Cost-Effective InterconnectGeneration table with data rate, encoding (8b/10b, 128b/130b, PAM-4 flit), x16 bandwidth per direction after encoding and year; mixed-generation interoperability; latency targets and flit layout.
- The PCIe 7.0 Specification, Version 1.0 is Now Available to MembersPCIe 7.0 at 128.0 GT/s doubles PCIe 6.x (64.0 GT/s); released to members on June 11, 2025.
- Retimers to the Rescue: PCI Express Specifications Reach Their Full Potential (PCI-SIG Educational Webinar)Channel loss budgets of 22, 28 and 36 dB at 8, 16 and 32 GT/s; about 5 inches of low-loss board trace at PCIe 5.0 speed; retimers split the channel into two segments.
- The PCI Express Port Bus Driver Guide HOWTOA Root Port originates a PCIe link from the Root Complex; switch upstream and downstream ports are logical PCI-PCI bridges.
- PCI Peer-to-Peer DMA SupportBehind PCIe switches, peer-to-peer transactions can route entirely within the hierarchy and never reach the root port; forwarding between hierarchy domains is undefined and blocked by default.
- GPUDirect RDMADirect GPU-to-peer data path over standard PCIe; devices must share an upstream root complex; a path through PCIe switches only performs best, through one CPU is worse, across CPU sockets may be severely limited.
- GPUDirect Storage Overview GuideWithout GDS, data takes an extra copy through a bounce buffer in CPU memory; a direct path through a PCIe switch can offer at least twice the peak bandwidth of a path through the CPU on some systems.
- Analyzing and Mitigating Data Stalls in DNN TrainingTraining time is often dominated by data stalls; DNNs need 3–24 CPU cores per GPU for pre-processing and spend up to 65% of epoch time on it.
- NVM Express SpecificationsThe NVMe specifications define how host software talks to non-volatile memory over PCIe, RDMA, TCP and other transports.
- SFF-TA-1008: Enterprise and Datacenter Standard Form Factor (E3), Rev 3.0aE3 drive form factors for 1U and 2U systems; 1C connector = x4 PCIe, 2C = x8; recommended 25 W max for an air-cooled E3.S device, 79.2 W connector limit at 12 V.
- An Introduction to the Compute Express Link (CXL) InterconnectCXL runs on the PCIe physical layer; CXL.io/.cache/.mem; Type 1/2/3 devices; CXL 1.0–3.0 dates; 57 ns end-to-end link adder; ~170 ns estimated CPU-to-Type-3 load latency, similar to remote-socket DDR; 250 ns through a switch.
- Compute Express Link 3.0 (white paper)CXL 3.0 is based on PCIe 6.0, doubling the rate to 64 GT/s with no added latency, for up to 256 GB/s aggregate raw bandwidth on x16.
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsEstimates CXL adds 70–90 ns over same-NUMA-node DRAM for 8–16-socket pools and more than 180 ns at rack scale; a bidirectional x8 CXL port at 2:1 read:write matches a DDR5-4800 channel.
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesMeasured three real CXL memory devices: load latency 35% to about 3× longer than remote-socket DDR5, depending on the controller; a PCIe 5.0 x8 device uses about 3× fewer pins than DDR5.
- Open Rack V3 Base Specification, Revision 1.048 V busbar; 48 mm OpenU or 44.45 mm EIA-310 rack units; IT gear must accept 46–52 V (51 V nominal) or 52–56 V (54 V nominal); IT-gear input connector power path rated 100 A continuous.
- Open Rack V3 48V PSU Specification, Rev 1.0Downloads as a .docx; the site may require a regular browser. 3 kW single-phase rectifiers with at least N+1 redundancy in the power shelf; fixed 51 V output, switchable to 48 V; peak efficiency above 97.5%; over-power protection above 3.45 kW for 10 s or 3.6 kW for 100 ms, plus a pulse-power envelope; the narrow range enables 4:1 fixed-ratio converters feeding conventional 12 V point-of-load converters.
- Multiphase Buck Design From Start to Finish (Part 1), SLVA882BMultiphase bucks are parallel phases interleaved at 360°/n; keep each phase to about 30–40 A; benefits in ripple, thermal spread and transients; a 5-phase 12 V→1.8 V design stays above 90% from 5 to 200 A.
- Virtual Intermediate Bus CPU Voltage RegulatorProcessors draw hundreds of amperes at ≤ 1 V; a 48 V→1 V, 640 A two-stage regulator reaches 95.2% peak and 84.4% full-load power-stage efficiency.
- OCP Accelerator Module (OAM) Design Specification v1.5102 × 165 mm module; up to 700 W at 44–59.5 V or 350 W at 12 V; one or two x16 host links; up to 7 module-to-module links; air cooling recommended up to 450 W; host-link pins listed as 16 transmit and 16 receive differential pairs.
- OCP NIC 3.0 Design Specification, Version 1.00Small form factor (SFF) cards carry up to 16 PCIe lanes, large form factor (LFF) up to 32.
- Project Argus DC-SCM 2.0 Module Design Specification, Rev 1.0A DC-SCM 2.0 module that separates management and security from the host board: ASPEED AST2600 BMC running OpenBMC, AST1060 hardware root of trust, under 20 W.
- OpenBMC features (openbmc/docs)Feature list of the open-source BMC firmware: Redfish, full IPMI 2.0 with DCMI, remote KVM, SSH-based serial-over-LAN, fan control, chassis power control, sensors, LEDs, inventory, logging, host watchdog, firmware update.
- Introduction to RedfishRESTful interface over HTTPS in JSON based on OData v4; a secure, multi-node capable replacement for IPMI-over-LAN; covers sensors, power supplies, power thresholds, reboot and console access.
- Redfish-Mockup-Server: public-rackmount1 Sensors/CPU1PowerDMTF’s sample Redfish Sensor resource for a CPU power reading, with thresholds.