Guide 4 of 4 · 7 chapters
Systems
One chip is never enough. This guide follows AI hardware outward: from the chip to its package, the computer it sits in, tall racks of computers, and the network that joins thousands of chips.
How chips are packaged, mounted in servers, wired into tightly coupled groups and connected across a datacenter network, and why the way models are split, along with power and cooling, shapes every layer.
Chiplets and HBM integration, PCIe and CXL, scale-up fabrics, scale-out topologies and optics, parallelism-driven traffic and rack power density, using public specifications and papers.
The Expert level assumes you know a datacenter is racks of servers joined by a network, and not much more.
Inside the box · package and server
01Packaging and chipletsA modern AI chip is often several pieces of silicon in one package, placed side by side or stacked like pancakes.Chiplets split a design into smaller dies that are joined on a silicon interposer (2.5D) or stacked (3D). Stacks of HBM memory sit beside the compute die. Die-to-die standards such as UCIe define how chiplets talk to each other.2.5D interposers and bridges, hybrid bonding and 3D stacking, bump pitch and die-to-die bandwidth density, UCIe, HBM integration, and the thermal and yield trade-offs of chiplet designs.- 02Board and serverA server is a powerful computer kept in a datacenter. Inside, the chips sit on boards with memory, power supplies, cooling, and a small helper computer that keeps watch.Accelerators connect to host CPUs over PCIe, and CXL adds memory sharing over the same wires. Voltage regulators turn rack power into the low voltages chips need. A baseboard management controller monitors and controls the machine.PCIe generations and lane budgets, CXL device types, power delivery from busbar to point of load, board and module form factors including open OCP specifications, and system management.
Between the boxes · fabrics and networks
03Scale-up fabricsTo act like one giant chip, a small group of AI chips is wired together with very fast, short links inside one computer or rack.Scale-up links connect accelerators directly or through switch chips, at far higher bandwidth than an ordinary network. Accelerators in the group can read each other’s memory and run group operations quickly.Point-to-point and switched scale-up topologies, link bandwidth and latency, memory semantics across the domain, collectives such as all-reduce, and rack-scale designs that extend the domain over copper backplanes.- 04Scale-out networkingTeaching a big AI takes thousands of chips in many racks, joined by cables, like a very fast internet inside one building.Scale-out networks connect servers through layers of switches. RDMA lets one machine write straight into another’s memory. InfiniBand and Ethernet compete here, and the topology decides how much bandwidth can cross the whole network at once.Fat-tree, rail-optimized and other topologies, bisection bandwidth and oversubscription, RDMA and RoCE, congestion control and load balancing, InfiniBand versus Ethernet, and the Ultra Ethernet Consortium’s specifications.
- 05OpticsOver long distances, copper wires lose too much signal. So networks send the data as flashes of light through thin glass threads.Pluggable optical modules turn electrical signals into light at the edge of a switch or server. Linear-drive optics remove some electronics to save power, and co-packaged optics move the optical parts right next to the switch chip.Reach and energy per bit for copper and optics, SerDes and DSP power, pluggable, linear-drive (LPO) and co-packaged optics (CPO), and the reliability and serviceability trade-offs among them.
Why it’s built this way
06Why the network looks this wayA big AI model is too large for one chip, so it’s split across many. How you split it decides how much the chips must talk to each other.Data parallelism copies the model and splits the data. Tensor, pipeline and expert parallelism split the model itself. Each creates different traffic, from huge group exchanges to small hand-offs between neighbors.Data, tensor, pipeline, sequence and expert parallelism, their communication volumes and collective patterns, overlapping communication with compute, and why tensor parallelism stays inside the scale-up domain while data parallelism crosses the scale-out network.- 07Power and coolingOne rack of AI computers can use as much electricity as dozens of homes, and it gets very hot. Getting power in and heat out limits how many racks you can use.Rack power has climbed past what air cooling can handle, so many AI racks use liquid cooling. Power flows from the grid through transformers, backup systems and distribution to the rack, losing a little at every step.Rack power density trends, the power conversion chain and its efficiency, direct-to-chip liquid and immersion cooling, PUE and its limits, and the facility constraints that now shape cluster design.