Guide 4 of 4 · 7 chapters

Systems

One chip is never enough. This guide follows AI hardware outward: from the chip to its package, the computer it sits in, tall racks of computers, and the network that joins thousands of chips.

How chips are packaged, mounted in servers, wired into tightly coupled groups and connected across a datacenter network, and why the way models are split, along with power and cooling, shapes every layer.

Chiplets and HBM integration, PCIe and CXL, scale-up fabrics, scale-out topologies and optics, parallelism-driven traffic and rack power density, using public specifications and papers.

The Expert level assumes you know a datacenter is racks of servers joined by a network, and not much more.

  1. Inside the box · package and server

    01Packaging and chiplets
    A modern AI chip is often several pieces of silicon in one package, placed side by side or stacked like pancakes.
    Chiplets split a design into smaller dies that are joined on a silicon interposer (2.5D) or stacked (3D). Stacks of HBM memory sit beside the compute die. Die-to-die standards such as UCIe define how chiplets talk to each other.
    2.5D interposers and bridges, hybrid bonding and 3D stacking, bump pitch and die-to-die bandwidth density, UCIe, HBM integration, and the thermal and yield trade-offs of chiplet designs.
    Full chapter
  2. 02Board and server
    A server is a powerful computer kept in a datacenter. Inside, the chips sit on boards with memory, power supplies, cooling, and a small helper computer that keeps watch.
    Accelerators connect to host CPUs over PCIe, and CXL adds memory sharing over the same wires. Voltage regulators turn rack power into the low voltages chips need. A baseboard management controller monitors and controls the machine.
    PCIe generations and lane budgets, CXL device types, power delivery from busbar to point of load, board and module form factors including open OCP specifications, and system management.
    Full chapter
  3. Between the boxes · fabrics and networks

    03Scale-up fabrics
    To act like one giant chip, a small group of AI chips is wired together with very fast, short links inside one computer or rack.
    Scale-up links connect accelerators directly or through switch chips, at far higher bandwidth than an ordinary network. Accelerators in the group can read each other’s memory and run group operations quickly.
    Point-to-point and switched scale-up topologies, link bandwidth and latency, memory semantics across the domain, collectives such as all-reduce, and rack-scale designs that extend the domain over copper backplanes.
    Full chapter
  4. 04Scale-out networking
    Teaching a big AI takes thousands of chips in many racks, joined by cables, like a very fast internet inside one building.
    Scale-out networks connect servers through layers of switches. RDMA lets one machine write straight into another’s memory. InfiniBand and Ethernet compete here, and the topology decides how much bandwidth can cross the whole network at once.
    Fat-tree, rail-optimized and other topologies, bisection bandwidth and oversubscription, RDMA and RoCE, congestion control and load balancing, InfiniBand versus Ethernet, and the Ultra Ethernet Consortium’s specifications.
    Full chapter
  5. 05Optics
    Over long distances, copper wires lose too much signal. So networks send the data as flashes of light through thin glass threads.
    Pluggable optical modules turn electrical signals into light at the edge of a switch or server. Linear-drive optics remove some electronics to save power, and co-packaged optics move the optical parts right next to the switch chip.
    Reach and energy per bit for copper and optics, SerDes and DSP power, pluggable, linear-drive (LPO) and co-packaged optics (CPO), and the reliability and serviceability trade-offs among them.
    Full chapter
  6. Why it’s built this way

    06Why the network looks this way
    A big AI model is too large for one chip, so it’s split across many. How you split it decides how much the chips must talk to each other.
    Data parallelism copies the model and splits the data. Tensor, pipeline and expert parallelism split the model itself. Each creates different traffic, from huge group exchanges to small hand-offs between neighbors.
    Data, tensor, pipeline, sequence and expert parallelism, their communication volumes and collective patterns, overlapping communication with compute, and why tensor parallelism stays inside the scale-up domain while data parallelism crosses the scale-out network.
    Full chapter
  7. 07Power and cooling
    One rack of AI computers can use as much electricity as dozens of homes, and it gets very hot. Getting power in and heat out limits how many racks you can use.
    Rack power has climbed past what air cooling can handle, so many AI racks use liquid cooling. Power flows from the grid through transformers, backup systems and distribution to the rack, losing a little at every step.
    Rack power density trends, the power conversion chain and its efficiency, direct-to-chip liquid and immersion cooling, PUE and its limits, and the facility constraints that now shape cluster design.
    Full chapter