Every chapter so far has been about moving data. This last one is about two things every chip needs: electricity going in, and heat coming out.
A rack is a tall metal cabinet full of computers. An average American home uses a bit over 1 kilowatt (1,000 watts), day and night.21 An ordinary rack uses several times that.4 A rack of today’s AI chips can use 120 kilowatts. That is as much as about a hundred homes, packed into a box the size of a closet.16 And nearly all of that electricity turns into heat.
Power and cooling now decide how big an AI cluster can be. New sites often wait years to get enough electricity.3
A datacenter is, physically, a machine for turning electricity into computation and heat. Every component in the earlier chapters, from the package to the optics, draws power, and essentially all of that power becomes heat inside the room. This chapter follows the energy both ways: in from the grid, through a chain of converters, to the chips; and out from the chips, through air or liquid, to the outdoors.
The scale changed quickly. The most common rack in today’s datacenters draws under 10 kilowatts (kW).4 Racks built for AI training exceed 100 kW, and designs aimed at several hundred kW to 1 megawatt (MW) are in development.514 U.S. datacenters used about 4.7% of the country’s electricity in 2024, and the Department of Energy’s reference projection more than triples that energy by 2030.1
The chapter covers five ideas:
- Rack power density: why AI racks draw ten times more than ordinary ones.
- The power conversion chain: the steps from the grid to about 1 volt at the chip, and what each costs.
- Air versus liquid: why air runs out and how cold plates, rear doors and immersion work.
- Measuring efficiency: PUE, the industry’s headline number, and what it can’t see.
- The facility as a limit: grid connections, water and power swings.
You know a cluster is racks of servers joined by a network. This chapter treats it as a thermodynamic and electrical system: a fixed amount of silicon power, multiplied up by conversion losses on the way in and by cooling overhead on the way out. The numbers that matter are rack (kW per rack), the efficiency of each stage of the , the heat each cooling method can remove per rack, and facility-level metrics such as .
The pressure comes from one direction. Accelerator power has gone from around 100 W to more than 1,000 W per chip14, and scale-up domains want as many of those chips as possible within copper reach, which pushes them into one rack. The result is liquid-cooled racks at 120–140 kW today, roadmaps above 500 kW per rack before 2030, and a shift from 48 V to ±400 V DC distribution to feed them.161514
We work through the chain stage by stage with real efficiency data, compare air, rear-door, direct-to-chip and immersion cooling with the heat-transfer equations behind them, take PUE apart using the industry’s measurement guidelines and the ITUE/TUE extension, and end with facility constraints: grid interconnection, water, and the power swings of synchronous training.
Rack off: nothing flows in, nothing has to be cooled. Switch it on.
Most racks still use under 10 kilowatts, about as much as a few homes.4
AI racks are different. Each AI chip now uses over 1,000 watts, as much as a microwave oven. Chips once used about a tenth of that.14 And designers cram as many chips as they can into one rack, because chips that sit close together can talk much faster (see the scale-up chapter). Put 72 of them in one cabinet and you get well over 100 kilowatts.16
These racks are heavy, too. Packed with metal, chips and water pipes, some weigh more than a small car.5
is the power one rack draws, in kilowatts. Because electronics turn nearly all their input into heat, it is also the heat that rack releases. It is the number that sizes everything else: the cables and breakers feeding the rack, the airflow or water flow cooling it, and how many racks one electrical room can serve.
Most of the industry is still at modest densities:
- In Uptime Institute’s 2025 survey of operators, the most common rack density averaged almost 9 kW, and more than 80% of operators had no racks above 30 kW at all. About one site in eight had some racks of 30–59 kW; racks above 100 kW were rare.4
- Most traditional datacenters were built to support peak densities of about 10–20 kW per rack.5
AI training racks sit far above that, at over 100 kW.5 Two things drive it. First, each accelerator’s (the sustained power its cooling must handle) has climbed past 1,000 W.14 Second, accelerators that work on one model need very fast links between them, and the fastest links are short copper ones, so designers cram as many accelerators as possible into one rack.5
Cooling and density feed each other. An air-cooled server needs tall heat sinks: one white paper’s example fits four 10U, 8-GPU air-cooled servers (32 GPUs) in a 42U rack, versus twenty-one 2U liquid-cooled servers (168 GPUs) once cold plates replace the heat sinks.5 (A “U” is a rack unit, 44.45 mm of height.)
Separate three numbers that are often blurred: the modal density of a fleet (what most racks draw), the provisioned density of a design (what the busbar, breakers and cooling are sized for), and the measured draw of a rack under a given workload.
- Fleet reality. Uptime’s 2025 survey puts the average modal density at almost 9 kW (7.5 kW excluding sites whose typical rack is 30 kW or more), up from 8.3 kW the year before. Over 80% of operators report no rack above 30 kW; about one in eight has racks in the 30–59 kW band; cabinets above 100 kW exist but are rare. HPC and AI training are concentrated in relatively few sites.4
- Design targets. Liquid-cooled AI racks are provisioned at 120–140 kW today1615, HPC compute racks have exceeded 125 kW8, and vendor and hyperscaler roadmaps run from 100 kW toward 1 MW.514
- Measured draw. AI servers rarely sit at nameplate. LBNL’s 2024 analysis of training runs found average node power of about 78% of the manufacturer rating for computationally saturated workloads.2 Inference fleets swing more with demand.
Two forces set the density. The first is the per-accelerator power trend: socket power keeps rising while allowable case temperature falls, so the required case-to-coolant thermal resistance keeps shrinking.7 The second is network topology: the incentive to cut accelerator-to-accelerator latency.5 A scale-up domain built on passive copper reaches only a few meters, so its accelerators must share a rack (or a pair of racks), and the domain size times the accelerator power becomes the rack power. A 72-accelerator domain at around 1 kW per package, plus CPUs, switches and conversion loss, lands above 100 kW per rack whether or not the facility is ready.
The physical side effects compound: racks above 100 kW weigh over 1,800 kg (4,000 lb), need deeper and wider cabinets to fit manifolds and power hardware (one guide recommends at least 750 × 1,200 mm and 48U), and require floor and lift ratings older buildings may not have.57
4 servers × 8 = 32 accelerators at 1,000 W: 32 kW of chip power alone (CPUs, memory, fans and losses come on top).
Voltage is the push that makes electricity flow. Power lines bring electricity in with a huge push, many thousands of volts. A chip needs a gentle push of about 1 volt. So the electricity steps down in stages, like a river splitting into smaller and smaller streams. A lowers it for the building. Tiny converters next to each chip bring it down to about 1 volt. The picture below walks through every step.
Each step is good but not perfect. Each one loses a little as heat, and the losses add up. By the time electricity reaches the chips, about a fifth may be lost. And that lost energy is more heat to remove.13
The takes utility power, usually medium-voltage AC (thousands to tens of thousands of volts), down to the roughly 1 V a processor core runs on. A typical chain:1013
| Stage | What it does | Typical efficiency |
|---|---|---|
| Medium-voltage AC to 480 V (or 400 V) AC for the building | About 99%; a 500 kVA low-voltage unit must reach 99.14% at 35% load under U.S. rules | |
| Rides through grid disturbances with batteries until generators start | 95% or better for modern double-conversion units; about 99% in “eco” mode | |
| Splits power into circuits for each rack | Very high (mostly cable loss); lower if it has its own transformer | |
| AC to DC: 12 V in older servers, a 48 V-class bus (46–52 V) in modern racks | About 94–96% for the best (80 PLUS Titanium) near half load | |
| DC bus down to about 1 V at hundreds of amperes, right next to the chip | Roughly 80–90% |
Sources for the table: transformer rules from the U.S. Code of Federal Regulations12, UPS figures from DOE’s best-practice guide8, power-supply levels from the 80 PLUS program table11, and the regulator range from published estimates for server boards (about 78–84% implied by one paper’s server model, 84% for a supercomputer’s two board-level stages combined).10
Efficiencies multiply. If five stages are 99%, 96%, 99.5%, 96% and 90% efficient, the share that reaches the chip is , or 82%. About 18% of the electricity bought is lost before any computing happens, and all of it becomes heat that also has to be cooled.13
Why 48 volts? Power is voltage times current, and the heat lost in a wire grows with the square of the current. Older servers distributed 12 V; modern racks use a 48 V-class (a thick copper bar up the back of the rack), which needs one-quarter the current for the same power and so loses one-sixteenth as much in the same copper.1413 At 100+ kW per rack even 48 V means thousands of amperes, which is why ±400 V DC is being developed next.14
Load matters. Most converters are least efficient when lightly loaded. Power supplies are usually best around half load11, and a UPS paired with an identical spare for redundancy runs at 50% load or less by design.8 The server-side stages are covered in more depth in Board and server.
A concrete chain, from the ITUE/TUE paper’s study of the Jaguar supercomputer: the site distributes 13.8 kV; building transformers step down to 480 V AC; switchboards feed each cabinet; a cabinet PSU converts 480 V AC to 52 V DC; an intermediate bus converter on each blade makes 12 V DC; point-of-load regulators produce 1.3 V for processors and 1.8 V for memory. The authors estimated the two board-level stages at 84% combined.10 Modern AI racks keep the shape but move the PSUs into shelves feeding a 48 V-class busbar, with OCP Open Rack v3 specifying a 46–52 V DC interface for IT gear.13
Stage by stage, with what sets each loss:
- Transformers. Core (no-load) loss is paid at any load; copper loss scales with . U.S. minimums are specified at 35% load: 99.14% for a 500 kVA three-phase low-voltage dry-type unit built 2016–2029, rising to 99.31% after April 2029.12 Many transformers run most efficiently at about 20–50% load, so sizing them to the real load matters.8
- UPS. Double conversion (AC→DC→AC continuously) improved from 85–90% in the 1990s to 95% or more; eco/bypass modes reach about 99%.8 Across facility types the spread is wide: LBNL modeled 77–85% for small sites and 90–99% for hyperscale and AI.2 Redundancy is the hidden cost: N+1 pairs sit at 50% load or less, and one guide’s example gains about 1.2 points of efficiency from a 10-point higher load factor.8 Training clusters checkpoint, which makes full UPS coverage less critical; some rack designs carry battery backup units instead.515
- PSU / power shelf. Power-factor correction then isolated DC/DC. 80 PLUS Titanium (230 V internal redundant) requires 90/94/96/91% at 10/20/50/100% load; on a 1,000 W supply the gap between 92% and 96% is 87 W versus 42 W of heat.11 Research front ends report 99% peak for the power-factor-correction stage, and a 400-to-50 V isolated stage 98.5% peak (97.5–98% at full load).13
- Intermediate bus and point of load. 48 V to about 1 V at hundreds of amperes is an extreme conversion ratio, so designs split it into a fixed-ratio stage and a multiphase regulator. Loss is dominated by conduction: package resistance, board copper, vias and inductors, which is why regulators keep moving closer to the package.13 On-die power delivery continues the same problem at chip scale; see Power planning.
Two consequences of the series structure. The end-to-end efficiency is the product , so input power is . And every watt lost downstream is carried, with its own losses, by every stage upstream: a regulator loss is also a small PSU, PDU, UPS and transformer loss, and then a cooling load.13
Distribution voltage is the other lever. At a fixed conductor resistance, loss is . A 120 kW rack at about 50 V draws 2,400 A, more than one busbar rated for 1,400 A carries16; a 1 MW rack at 50 V would need 20,000 A. Google’s move from 12 V to 48 V in the rack and now to ±400 V DC is this arithmetic; its first step, an AC-to-DC “sidecar” power rack, is reported to improve end-to-end efficiency by about 3% and frees the IT rack for accelerators.14 Higher-voltage DC buses shift the hard problems to insulation, DC fault interruption and protection, whose ecosystems are still developing.13
Utility: Medium-voltage AC arrives from the utility. Follow 1,000 W through the chain.
I = P / V = 120,000 W / 48 V ≈ 2,500 A. Heat in the same copper ∝ I²: 1× of the 48 V loss. 2 busbars of 1,400 A.
Getting the heat out is the other half. There are a few main ways:
- Air. Fans blow cool air through the computers, and big air conditioners cool the room. It’s simple, like a fan on a hot day.
- Water to the chip. A metal block with water flowing through it sits right on top of each hot chip. This is called .
- Dunking. Whole computers sit in a tank of special liquid that doesn’t carry electricity. The liquid soaks up all the heat.
Why water? It carries about 4,000 times more heat than the same amount of air.14 Think of cooling a hot pan. Blowing on it is slow, but running it under the tap cools it in seconds.
As racks get hotter, air can’t keep up. You need faster and faster fans, which get loud and use lots of power themselves. That’s why the hottest AI racks today are cooled with water.75
All cooling is the same equation: heat removed = mass flow × specific heat × temperature rise, written . To move more heat you need more flow, a fluid that holds more heat, or a bigger temperature rise. Water holds about four times more heat per kilogram than air and, being about 800 times denser, a few thousand times more per liter.2214
Air
Servers pull room air in at the front (the cold aisle) and blow it out the back (the hot aisle). Computer room air handlers cool it, usually with water from a or, when the weather allows, with an (“free cooling” from outside air or cooling towers). ASHRAE recommends 18–27 °C (65–80 °F) at server inlets.8
Air runs out on volume. A 40–50 kW rack can need up to 5,000 cubic feet per minute (cfm) of air, while the best perforated floor tile delivers about 1,900 cfm.7 Server fans, which once used up to 20% of server power and fell below 2% with good controls, are climbing back to 10–20% in dense servers: at least 5 kW of fans in a 50 kW rack.7
Rear-door heat exchangers
A is a water coil on the back of the rack. The servers stay air-cooled, but the hot exhaust is cooled before it reaches the room. Passive doors use the servers’ own fans; active doors add fans.8 One vendor publishes sizing curves for 20–40 kW racks depending on water temperature and flow, with the water kept above the room’s dew point so the coil doesn’t sweat.17
Direct-to-chip (cold plates)
In , cold plates (metal blocks with fine internal channels) replace heat sinks on the processors, and sometimes memory. A water-glycol mix flows through them. It is a hybrid: most of the heat goes into liquid, the rest still into air.8 Because the chip coolant needs strict chemistry and filtration (untreated water corrodes and clogs the fine channels), a (CDU) keeps it in its own loop and passes the heat to the building’s water through a heat exchanger.56
Immersion
In the electronics sit in an electrically insulating fluid. In single-phase systems an oil-like fluid is pumped through a heat exchanger; in two-phase systems the fluid boils on hot components and condenses on a coil above.8 It captures nearly all the heat with no server fans, but servicing means lifting hardware out of fluid (sometimes with a crane), and the fluid must be compatible with every material in the server.7
Liquid also changes the rest of the plant. Cold plates can run on warm water, so the facility can reject heat with dry coolers or cooling towers much of the year instead of running chillers.28
Every cooling method is a chain of thermal resistances from junction to outdoor air. ASHRAE frames the requirement as a case-to-coolant thermal resistance, : as socket power rises and allowable case temperature falls, the required drops, forcing either colder coolant or a better path.7 The capacity of each method is set by its weakest link.
| Method | Where heat goes | What limits it |
|---|---|---|
| Air (with containment) | Room air → air handler coil → chilled water or economizer | Volumetric airflow per rack (≈ 5,000 cfm at 40–50 kW vs 1,900 cfm per best tile), fan power, acoustics, heat-sink height |
| Rear-door heat exchanger | Exhaust air → door coil → facility or CDU water | Coil area, water temperature and flow, server airflow; dew point |
| Direct-to-chip (single-phase) | Cold plate → technology cooling system (TCS) loop → CDU → facility water | Cold-plate , flow and pressure drop per manifold, CDU capacity, residual air heat |
| Immersion | Fluid → tank heat exchanger or condenser → facility water | Fluid properties (lower and higher viscosity than water), heat sinks in fluid, serviceability |
Sources: airflow and fan figures from ASHRAE7, method descriptions from DOE’s guide8, loop structure from Schneider’s liquid-cooling architecture paper.6
Liquid loops. A direct-to-chip system has three loops: the technology cooling system (TCS) through the servers, the facility water system (FWS), and a condenser water system to the outdoors where there is a chiller. The CDU between TCS and FWS does temperature control, flow control, pressure control (some run the TCS below atmospheric pressure so a leak draws air in rather than spraying water), fluid treatment, and heat exchange with isolation. Rack-mounted CDUs cover roughly 20–40 kW (rejecting to room air) or 40–80 kW (to facility water); floor-mounted liquid-to-liquid units run from 300 kW to over 1 MW.6
Water temperature classes. ASHRAE names liquid-cooling classes by their maximum facility supply temperature: W17, W27, W32, W40, W45 and W+.7 LBNL modeled liquid-cooled AI facilities at W45.2 Warm supply lets dry coolers replace chillers for most of the year, but ASHRAE warns that rising socket power and memory pre-heat in shared loops may push future equipment back toward W32 or W27, colder and less efficient.7
Residual heat. Cold plates capture the processors (and sometimes memory), not the PSUs, NICs, drives and switch ASICs. Whatever is not captured must be removed by air, so a “liquid” rack can still exceed the room’s air capacity; a rear door is a common complement.87
Immersion specifics. One vendor’s single-phase immersion module has a CDU rated for up to 200 kW with warm water and 368 kW with chilled water.18 As component power rises, single-phase natural convection hits its limit and designs need forced flow or two-phase.7 Many two-phase fluids are fluorinated chemicals. 3M announced it would exit all PFAS manufacturing, including fluorinated fluids, by the end of 202519, and new dielectric fluids are being sought that address PFAS and global-warming concerns.5
Air cooling: cold aisle in, hot aisle out, air handlers cool the room. A 40–50 kW rack can need more than twice the air one floor tile delivers. Tap a part.
Datacenters are often judged by one score, the . It compares all the electricity the building uses with the part that reaches the computers.9
- A PUE of 2 means cooling and other things use as much again as the computers do.
- A PUE of 1.1 means only 10% extra.
- A perfect score is 1. You can’t go below it.
The average has been stuck around 1.5 for years. The newest, best-run sites get close to 1.1.417
There’s a catch. PUE counts everything inside the computer as useful, even its fans. So a site can improve its score just by moving fans from the room into the computers. It hasn’t saved any electricity at all.10
(PUE) is total facility energy divided by IT equipment energy, over a year.98 It was created by The Green Grid industry consortium and is now an international standard.8 The overhead it captures is everything that supports the IT: cooling (chillers, pumps, fans in air handlers, cooling towers), power delivery losses (transformers, UPS, PDUs) and lighting.9
Where things stand:
- Uptime’s operator survey: a weighted average of 1.54 in 2025, almost unchanged for six years after falling from 2.5 in 2007 to 1.65 in 2014. Facilities built since about 2020 average 1.48.4
- DOE/LBNL’s national model: an average of 1.45 in 2024 (down from 1.55 in 2018), and about 1.15 for facilities serving AI equipment.1
What PUE can’t tell you:
- Nothing about useful work. PUE is not a productivity metric: it measures only the supporting infrastructure, so servers doing useful work and servers sitting idle at the same power score the same.89
- Losses inside the servers are invisible. IT energy is measured where power enters the IT equipment, so server fans, power supplies and count as IT.910
- It can move the wrong way. Install more efficient servers without changing the building and PUE goes up, because the IT part shrank.10
- Water and carbon are outside it. Evaporative cooling lowers PUE but uses water, measured separately as , and PUE gives no credit for waste heat reused outside the datacenter.29
To fix the second problem, the Energy Efficient HPC Working Group proposed , the same ratio applied inside the server (power into the server divided by power used by processors, memory and storage), and , which measures everything from the meter to the silicon.10
The measurement guidelines that PUE’s sponsors (The Green Grid, ASHRAE, Uptime Institute, DOE, EPA and others) agreed on are precise about boundaries. Total energy is everything that crosses the data-center boundary: the utility meters of a dedicated building, or the datacenter’s sub-metered share of a mixed-use one. IT energy is measured at one of three points: UPS output, PDU output, or IT equipment input. Annual energy is strongly recommended, and the guidelines caution that comparing different data centers first needs analysis of factors such as reliability level and climate.9 By construction PUE runs from 1.0 to infinity.8
Four structural limits follow from that definition:
- Boundary shifting. Everything inside the IT boundary is “useful.” The ITUE/TUE paper’s thought experiment: three identical sites, one with building fans and server fans, one with only server fans, one with only building fans. Their PUEs rank b best, then a, then c, although nothing says which uses least energy.10 The guidelines list in-rack and in-chassis cooling as supporting infrastructure, yet the UPS- and PDU-output readings they allow may still include non-IT devices installed in IT racks, such as fans.9 Direct-to-chip racks with in-rack CDUs and immersion tanks are exactly this case.
- Denominator effects. More efficient IT raises PUE unless the infrastructure turns down with it.10 Raising inlet temperature lowers chiller energy but speeds server fans; with fan power scaling as speed cubed, a 30% speed-up more than doubles fan power, which PUE books as IT.8
- Model versus reality. LBNL notes its simulated PUEs assume systems are commissioned and run as designed, which “is rarely the case” (air handlers fighting each other over humidity, for example).2
- Scope. PUE excludes water and IT efficiency4, gives no credit for heat reused outside the datacenter or for on-site renewable generation9, and says nothing about work per joule.
, with fans, PSUs, voltage regulators and the baseboard controller as overhead; . In the paper’s Jaguar case study, the switchboards fed the cabinets 5,259.56 kW, PSU output 4,209.90 kW and board-level conversion 84% efficient, so compute received about 3,536 kW: , and with the building’s PUE of 1.25, . For every kilowatt reaching compute, another 0.49 kW went to cooling and conversion inside the cabinets, nearly twice the 0.25 kW of overhead per IT kilowatt that the building’s PUE reported.10 The trade-off: TUE needs component-level power telemetry, and a cleaner definition of what counts as “compute” (the authors count storage; status LEDs and management controllers are overhead).10
PUE = 155 / 115 = 1.35. Total facility power 155 kW.
A big AI datacenter can use as much electricity as 100,000 homes. The largest ones being built will use 20 times more.3 That much power doesn’t just plug into the wall.
- Power lines. The area may need new power lines. In some places, the wait for a new connection is three to five years.5
- Big equipment. The wait for large transformers has doubled in a few years.3
- Water. Some cooling systems save electricity by letting water evaporate, like sweat. That uses a lot of water, which is scarce in many places.2
So companies now pick where to build based on where they can get power.
Datacenters are now a noticeable share of electricity demand. Worldwide they used about 415 terawatt-hours (TWh) in 2024, about 1.5% of all electricity, and the International Energy Agency expects that to roughly double to 945 TWh by 2030.3 In the U.S. the share was 4.7% in 2024; DOE’s reference case reaches 11.8% by 2030, within a range of 9.5–15.3%.1
Constraints outside the rack now shape cluster design:
- Grid interconnection. Training clusters want hundreds of megawatts in one place, because the accelerators need low-latency links to each other. In energy-constrained regions, grid-connection queues have reached 3–5 years.5 The IEA estimates about 20% of planned datacenter projects could be delayed.3
- Electrical gear. Some North American datacenters still distribute 120/208 V and will need higher voltage; the “blocks” of switchgear and UPS must grow for many 100 kW racks.5 Lead times for transformers and cables have doubled.3
- Water versus energy. Evaporative cooling (cooling towers, adiabatic assist) is usually more energy-efficient than waterless options, so low PUE and low water use can pull in opposite directions.2
- Power swings. In synchronous training, thousands of accelerators compute together, then pause together to exchange data (see Why the network looks this way), so the whole site’s power rises and falls together, by tens or hundreds of megawatts.20
At campus scale the binding constraints are interconnection capacity, lead time, and load shape, more than PUE.
- Demand. U.S. datacenter use was 192 TWh in 2024 (4.7%); the 2030 reference case is 649 TWh (11.8%), bounded by 521–843 TWh under compounded uncertainty, and datacenters would account for about a third of U.S. load growth over 2024–2030.1 Infrastructure (mostly cooling) fell from 36% of datacenter electricity in 2018 to 31% in 2024 as load shifted to low-PUE, liquid-cooled AI facilities.1
- Interconnection. Latency-sensitive training is why developers procure hundreds of megawatts at one site rather than spreading load; queues of 3–5 years follow.5 The IEA puts about 20% of planned projects at risk of delay, with transformer and cable lead times doubled in three years.3
- Load dynamics. A training iteration alternates a compute phase near TDP with a communication phase near idle. Across a site the swing reaches tens to hundreds of MW, with energy concentrated around 0.2–3 Hz, overlapping frequencies at which grid components such as turbine-generator shafts and long transmission lines can resonate. Proposed fixes span software (inject filler work when power drops), GPU firmware (ramp-rate limits and a minimum power floor, for example 65% of TDP) and rack energy storage.20 A power floor trades energy for grid stability, a cost no PUE figure shows.
- Water. LBNL estimated an average site WUE just over 0.36 L/kWh through 2023, and stressed the coupling: water-cooled chillers and evaporative economizers lower PUE but consume water, while air-cooled chillers use none but more energy.2
A campus and the constraints outside the rack. Tap a part.
Compute, then communicate: the site swings by about 160 MW every iteration, in the 0.2–3 Hz band where grid equipment can resonate.
Build a rack and try to keep it powered and cool. Pick a way to cool it, how many chips it holds, and how much power each uses. In the colored bar, the first, biggest part (labeled “the chips” below it) is the electricity that reaches the chips. The rest is lost or spent on cooling. Add chips until the cooling can’t keep up, then switch to water to rescue it.
The simulation builds a rack budget from the chips outward. You set the accelerators per rack, their power, a per-server overhead for CPUs, memory and networking, and the number of racks. It walks the power chain (regulators, server fans, power supplies, PDU, UPS, transformer), applies the PUE, and shows where every kilowatt goes. It then checks the rack’s heat against an illustrative limit for the chosen cooling mode, and works out the airflow or water flow needed. Things to try:
- Start with “Air, 8-GPU servers” and add accelerators. Where does air cooling give out?
- Lower the UPS or PSU efficiency by a few points. How much does facility power change?
- Switch the same rack between air and direct-to-chip. Why does the IT power change too?
The Expert view exposes every chain efficiency, the cold-plate capture fraction for direct-to-chip, and the arithmetic. It also brackets PUE’s “IT” against TUE’s “compute” on the bar and reports ITUE and TUE, so you can see how much loss PUE files under useful work. Try:
- Hold the chips fixed and move between air (8% fan share) and immersion (no fans). Compare the change in PUE with the change in TUE.
- With direct-to-chip, lower the capture fraction until the residual air heat, not the liquid loop, fails the rack.
- Set the PUE below what the electrical losses alone allow and see what the model does.
- Halve and watch the flow double: .
Cooling limits per mode (35, 60, 250 and 200 kW, plus 35 kW of residual air for direct-to-chip) and fan shares (8%, 8%, 3%, 0%) are round illustrative values chosen to sit inside the published ranges discussed above, not ratings of any product. The flow uses for water and for air; annual energy assumes 8,766 hours at the chosen average load; the Beginner view converts it to homes at 10,800 kWh a year.21
- Share of U.S. electricity used by datacenters, 2024
- about 1 in 20
- Most common rack today
- about 9 kilowatts
- A water-cooled AI rack
- 120–140 kilowatts
- Average PUE score, 2025
- about 1.5
- Almost one in every twenty units of electricity used in the U.S. goes to datacenters. That share may more than double by 2030.1
- A typical rack uses about as much as a handful of homes.4 An AI rack uses more than ten times that.1615
- An average datacenter uses half as much again on cooling and losses as its computers use.4
- One AI rack running all year uses about as much electricity as a hundred American homes.21
- World datacenter electricity, 2024
- ≈ 415 TWh
- U.S. datacenters, 2024 → 2030 (ref.)
- 192 → 649 TWh
- Weighted average PUE (Uptime 2025)
- 1.54
- Water vs air, heat per volume
- ≈ 4,000×
Sources: IEA3, LBNL1, Uptime4, Google.14
Rack power, from typical to frontier
| Rack | Power | Cooling |
|---|---|---|
| Typical (modal) enterprise or colocation rack | ≈ 7.5–9 kW | Air |
| Peak design density of a traditional datacenter | 10–20 kW | Air |
| Dense air-cooled rack (≈ 5,000 cfm) | 40–50 kW | Air, often with rear doors |
| AI training rack | Over 100 kW | Direct-to-chip liquid |
- PUE, U.S. facilities serving AI (2024, modeled)
- ≈ 1.145
- Jaguar ITUE / TUE
- 1.49 / 1.86
- 80 PLUS Titanium at 50% load (230 V redundant)
- 96%
- Server fan share, dense air-cooled
- 10–20%
Sources: LBNL1, the ITUE/TUE paper10, 80 PLUS table11, ASHRAE.7
The ITUE/TUE worked example
Two sites with the same building (1.98 MW of infrastructure) and the same 10,000 servers, differing only in server power hardware. Per-server watts from the paper’s platform model:10
| Low-efficiency server | High-efficiency server | |
|---|---|---|
| PSU loss | 58 W | 18 W |
| Voltage regulator loss | 56 W | 38 W |
| Fans | 18 W | 12 W |
| Processor, memory, other (compute) | 198 W | 198 W |
| Total per server | 330 W | 266 W |
| PUE | 1.60 | 1.74 (worse) |
| ITUE | 1.67 | 1.34 |
| TUE | 2.67 | 2.34 (better) |
The efficient site draws about 4.64 MW instead of 5.3 MW for the same output, yet its PUE is worse. Only TUE ranks them correctly.10
Stage efficiencies used as defaults in the simulation
| Stage | Default | Published anchor |
|---|---|---|
| Transformer | 99.0% | ≥ 99.14% at 35% load for a 500 kVA LV dry-type unit (U.S. minimum) |
| UPS | 96.0% | ≥ 95% double conversion; ≈ 99% eco mode; 90–99% modeled for hyperscale and AI |
| PDU | 99.5% | Mostly cable and any transformer loss |
| PSU | 96.0% | 80 PLUS Titanium 96% at 50% load |
| VRMs | 90.0% | 84% for Jaguar’s two board stages; about 78–84% implied by the ITUE paper’s server model. The default assumes a newer 48 V design. |
Anchors: eCFR12, DOE8, LBNL2, 80 PLUS11, and the ITUE/TUE paper.10 The PDU and VRM defaults are assumptions.
- Water cooling is strong but fussy. It removes far more heat, but it adds pumps, pipes and the risk of leaks.7
- Packing chips close helps speed but makes heat. Close chips talk faster. But a hotter, heavier rack needs a stronger floor, thicker power cables and water cooling.5
- Saving electricity can cost water. Cooling by evaporating water saves electricity but uses water that may be scarce.2
- Backup isn’t free. Spare equipment, kept ready in case something breaks, makes the site more reliable. But machines running half-empty waste more energy.8
- A good score can hide waste. A low PUE doesn’t mean the computers themselves are efficient.10
What you give up for what you get
- Density vs facility readiness. Packing accelerators into one rack shortens links and raises throughput, but needs liquid cooling, higher-voltage distribution, bigger electrical blocks and floors rated for over 1,800 kg per rack.5
- Warm water vs chip temperature. Warmer coolant lets the site skip chillers for more of the year, but leaves less temperature margin for the chips; ASHRAE warns future parts may need colder water.7
- Energy vs water. Evaporative cooling lowers PUE and raises WUE; dry cooling does the reverse.2
- Redundancy vs efficiency. N+1 and 2N UPS designs keep each unit lightly loaded, where efficiency is lowest. Training jobs that checkpoint can tolerate less backup than inference services.85
- Immersion vs serviceability. Immersion removes all fans and captures nearly all heat, but complicates repairs, warranties and fluid supply.719
Ways designs go wrong
- Forgetting the leftover air heat. Cold plates capture most heat, not all; the residual can overwhelm a room planned as “liquid-cooled.”8
- Running plant water through the chips. Untreated water corrodes, fouls and clogs cold plates, making accelerators throttle or shut down; a CDU keeps the loops separate.5
- Chasing PUE. Warmer rooms lower chiller energy but speed up server fans, and the fans count as IT.810
- Designs that don’t run as designed. Air handlers fighting each other over humidity are common enough that LBNL expects real PUEs to be worse than modeled ones.2
Engineering trade-offs
- Coolant vs flow vs chip margin. For fixed , flow scales as and pump power, at fixed geometry, roughly as flow cubed. A large saves pumping and suits heat reuse, but raises the outlet temperature that the last cold plate in a series loop sees. Series loops minimize flow; parallel loops avoid pre-heat; ASHRAE expects designs to mix both.7
- Where to put conversion. Pulling AC/DC conversion into a sidecar rack frees IT space and gains about 3% end to end14; moving regulators closer to the package cuts conduction loss but raises local heat flux next to the hottest part.13
- Measurement boundary vs incentive. Moving fans or pumps across the IT boundary changes PUE without changing energy; use TUE, or at least report temperatures and boundaries with PUE.10
- Grid stability vs energy. Power floors and filler workloads smooth training swings but burn energy doing nothing useful; rack batteries cost capital and space.20
Failure modes
- Condensation. Water below the room’s dew point sweats on coils and pipes; rear doors must stay above it unless the supply tracks the measured dew point.17
- Undersized margin. Rear-door sizing guidance adds 5–10% extra heat removal for airflow uncertainty; racks that are unevenly populated or have unusual airflow fall off the published curves.17
- Fluid chemistry and materials. Galvanic corrosion between mixed metals, biological growth and fouling in cold-plate channels; immersion fluids that attack plastics or void warranties.57
- Hidden UPS capacity loss. When server fan power rises from 2% to 10% of server power, the same fans eat about 8% of UPS capacity, because they ride on the same backed-up feed.7
- Double counting. When estimating savings from a better converter, add the avoided loss and its cooling share (); multiplying by PUE as well counts the cooling overhead twice.13
2N: each PSU at 20% load, 94.0% efficient: 64 W of heat per 1,000 W delivered (one PSU alone: 49 W).
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
1. The rack energy balance
The simulation, and most first-cut facility models, build power outward from the silicon. With accelerators at , servers with overhead (CPUs, memory, NICs), fan share and stage efficiencies :
The matters: must at least cover the electrical losses between the meter and the rack, so with a 99% transformer, 96% UPS and 99.5% PDU, no facility can honestly report PUE below about 1.06 before it spends a watt on cooling. The general form, , and its sensitivity , show that in a purely series chain a point of efficiency is worth about the same at any stage; in real, branched distribution the stages carrying the most power give the most leverage.13
Worked default: 72 accelerators at 1,000 W and 18 servers at 1,200 W give 93.6 kW of silicon. A 90% VRM raises that to 104.0 kW, 3% fans to 107.1 kW, a 96% PSU to 111.6 kW of IT power. The PDU, UPS and transformer bring it to 118.0 kW at the meter, and PUE 1.15 to 128.3 kW, leaving 10.3 kW for cooling and lighting. ; . PUE suggests 87% of facility power is “useful”; 73% reaches silicon.
2. Heat transport:
For water, and density ≈ 1 kg/L; for air, the volumetric heat capacity is about at room conditions.22 Their ratio, about 3,500 per unit volume, is the “roughly 4,000×” in Google’s comparison.14 For a 120 kW rack:
- Water, : ≈ 172 L/min.
- Air, : ≈ 14,000 cfm, more than seven best-case floor tiles at 1,900 cfm each.7
The same equation explains the 40–50 kW air rack at about 5,000 cfm: 45 kW at a 15 K rise needs 2.5 m³/s, about 5,250 cfm.7
3. Fan and pump affinity laws
For a given fan or pump and system, , , . DOE’s example: a 30% fan speed increase when inlet air rises from 77 °F to 91 °F more than doubles fan power ().8 Applied to liquid loops, halving doubles flow and, through the same pipes and cold plates, multiplies pump power by roughly eight. That is why is a first-class design variable, not a detail.
4. Thermal resistance budget
The cooling requirement for a socket is .7 An illustrative case: a 1,000 W package with an 80 °C case limit cooled by 40 °C water needs across the thermal interface, cold-plate base, channels and the fluid’s convective film. Doubling power to 2,000 W at the same temperatures halves the allowed resistance to 0.02 K/W; keeping 0.04 K/W instead would need coolant at . In practice the designer gets the budget back from better interfaces, larger cold-plate area, more flow, or colder (less efficient) water: ASHRAE’s warning about W45 sliding back toward W32 or W27.7 The thermal side of the package itself is covered in Packaging and chiplets, and where the watts come from in Speed and power.
T_case = 40 + 1000 × 0.04 = 80 °C ≤ 80 °C. Within budget: θ_ca must be ≤ (80 − 40) / 1000 = 0.040 K/W, or coolant ≤ 40 °C at this θ.
5. PUE, ITUE and TUE
All three are computed on annual energy.109 Because ITUE and PUE share , moving a fan from the building into the server lowers PUE and raises ITUE by offsetting amounts; TUE moves only if total energy does. Neither metric contains work. For comparing chips and systems on work per watt and cost, see Comparing chips.
6. Incremental savings from a better stage
For a fixed delivered load, improving a stage reduces conversion loss by . If that heat would have been removed by mechanical cooling with coefficient of performance COP, the facility saving is about , with the fraction of the avoided heat that the active cooling would have handled.13 The review’s example: a 1 MW load through (95.96%) improved to (96.25%) cuts input by about 3.2 kW; at COP 4 the cooling saving adds about 0.8 kW.13
7. Annual energy
Here is the average load as a fraction of the provisioned power. A 100 MW facility at uses about 701 GWh a year, the consumption of roughly 65,000 average U.S. homes at 10,791 kWh each.21 Real depends on workload: LBNL found about 78% of rated power for computationally saturated training nodes.2
That’s the end of the Systems guide. You’ve followed electricity from a switch smaller than a virus to a rack that uses as much power as a hundred homes. What’s left is the software that keeps all those chips busy. That’s the Software guide, coming soon.
This chapter closes the Systems guide and the four-guide path from transistor to datacenter. Power and cooling explain several choices you met earlier: why scale-up domains fit in a rack, why optics fight for every picojoule per bit, and why the busiest racks are liquid cooled. The next layer up is software: compilers that map models onto chips, kernels that keep them busy, runtimes, and the collective libraries that drive the network. Those are planned for the Software guide, coming soon.
Every layer in these guides ends up as a term in the same budget: picojoules per operation in the transistor, watts per accelerator in the package, kilowatts per rack, megawatts per site. Software decides how much useful work each of those joules buys, and also how the load swings: the training-loop power oscillations above are a software-shaped problem with a grid-scale consequence. Compilers, kernels, runtimes and collective libraries are the subject of the Software guide, coming soon.
One accelerator package: over 1,000 W, from billions of transistors (the Transistors guide) in one package.
Q1Electricity passes through a transformer (99%), a UPS (96%), a PDU (99.5%), a power supply (96%) and voltage regulators (90%). About what fraction reaches the chips?
Q2Where do the standard PUE guidelines say the “IT” part of PUE should be measured?
Q3Direct-to-chip cooling captures about 80% of a 120 kW rack’s heat in liquid. What happens to the rest?
Q4Why do hyperscale and AI-focused facilities tend to report lower PUE than small server rooms?
Sources
Show Hide 22 sources
- United States Data Center Energy Usage Report: 2025 Update192 TWh (4.7% of U.S. electricity) in 2024; 649 TWh (11.8%) reference case for 2030, range 9.5–15.3%; national average PUE 1.55 (2018) to 1.45 (2024); about 1.145 for facilities serving AI equipment; infrastructure 31% of datacenter electricity in 2024.
- 2024 United States Data Center Energy Usage ReportUPS efficiency ranges by datacenter type (77–85% small sites, 90–99% hyperscale and AI); liquid-cooled AI modeled at ASHRAE class W45; simulated PUEs assume systems run as designed, which is rarely the case; trade-offs between low PUE and low water use.
- Energy and AI: Executive summaryDatacenters used about 415 TWh (around 1.5% of world electricity) in 2024, heading to about 945 TWh by 2030; a typical AI-focused datacenter uses as much as 100,000 households; about 20% of planned projects could face delays; wait times for transformers and cables have doubled.
- Uptime Institute Global Data Center Survey 2025 (Keynote Report 180)Weighted average annual PUE 1.54, flat for six years (2.5 in 2007, 1.65 in 2014); 1.48 for facilities built since about 2020; modal rack density almost 9 kW; over 80% of operators have no racks above 30 kW; racks above 100 kW exist but are rare.
- How 6 AI Attributes Change Data Center Design (White Paper 110, version 3)Traditional datacenters support peak rack densities of about 10–20 kW; AI training racks exceed 100 kW and 1,800 kg; trend toward 1 MW racks and higher distribution voltages; grid connection queues of 3–5 years; air-cooled vs liquid-cooled GPU counts per rack; CDUs isolate chip coolant from facility water.
- Navigating Liquid Cooling Architectures for Data Centers with AI Workloads (White Paper 133, version 2.1)Technology cooling system, facility water system and condenser water loops; the five functions of a coolant distribution unit; CDU capacities (rack-mounted liquid-to-air 20–40 kW, liquid-to-liquid 40–80 kW, floor-mounted liquid-to-liquid 300 kW to over 1 MW).
- Emergence and Expansion of Liquid Cooling in Mainstream Data Centers (white paper)Water classes W17–W+; thermal resistance as the cooling requirement; server fan power fell from about 20% to under 2% and is rising again to 10–20% in dense servers; a 40–50 kW rack can need up to 5,000 cfm against 1,900 cfm from a best-in-class floor tile; H1 class; immersion service challenges.
- Best Practices Guide for Energy-Efficient Data Center Design (revised July 2024)ASHRAE recommended inlet range 65–80 °F; fan affinity law; double-conversion UPS efficiency from 85–90% in the 1990s to 95% or more, eco mode up to 99%; N+1 UPS load factors; rear doors, cold plates and immersion; HPC racks of 60 kW in 2013 and over 125 kW recently; PUE defined on annual energy.
- Recommendations for Measuring and Reporting Overall Data Center Efficiency, Version 2: Measuring PUE for Data CentersPUE = total datacenter energy ÷ IT energy, annual energy strongly recommended; IT energy measured at UPS output, PDU output or IT equipment input; in-rack and in-chassis cooling listed as supporting infrastructure; caution when comparing datacenters; does not address IT efficiency; heat reused outside the datacenter and on-site renewables do not change PUE.
- TUE, a new energy-efficiency metric applied at ORNL’s JaguarPUE ignores fan, PSU and voltage-regulator losses inside IT equipment; defines ITUE and TUE = ITUE × PUE; worked example of 330 W vs 266 W servers; Jaguar’s 13.8 kV to 1.3 V chain, 84% board-level conversion, ITUE 1.49 and TUE 1.86.
- What do the different PSU (power supply units) ratings mean?Official 80 PLUS efficiency tables by load, Standard through Titanium and Ruby; 230 V internal redundant Titanium requires 90/94/96/91% at 10/20/50/100% load.
- 10 CFR 431.196: Energy conservation standards for distribution transformersLow-voltage dry-type units are rated at 35% load, liquid-immersed and medium-voltage dry-type units at 50%. At 35% load, a 500 kVA three-phase low-voltage dry-type transformer must reach 99.14% (2016–2029) and 99.31% from April 2029.
- GaN Power Devices and Converter Architectures for AI Data Centers: Efficiency, Reliability, and Deployment PathwaysThe grid-to-load chain (PFC, isolated DC/DC, 48 V bus, point of load); cascaded efficiency is multiplicative; the 48 V bus cuts I²R loss; OCP ORV3 46–52 V interface; point-of-load conversion dominated by conduction loss; ±400 V and 800 V DC emerging; avoid multiplying loss savings by PUE.
- Enabling 1 MW IT racks and liquid cooling at OCP EMEA Summit48 VDC in the rack replaced 12 VDC about ten years earlier; ±400 VDC for up to 1 MW per rack; ML racks above 500 kW before 2030; a sidecar power rack improving end-to-end efficiency by about 3%; liquid cooling across 2,000+ TPU pods at about 99.999% availability; water carries about 4,000× more heat per volume than air.
- Meta’s open AI hardware visionCatalina introduces an ORv3 high-power rack supporting up to 140 kW; fully liquid cooled, with a power shelf and battery backup unit.
- NVIDIA Contributes NVIDIA GB200 NVL72 Designs to Open Compute Project120 kW of cooling capacity per rack, handled with direct liquid cooling; 72 GPUs in 18 compute trays and 9 switch trays; a busbar rated for 1,400 A, twice the earlier standard.
- Heat exchanger performance (rear door heat exchanger)Performance curves for rack heat loads of 20 to 40 kW versus water temperature and flow; size for 105–110% heat removal; water must stay above the room dew point.
- GRC Pushes Density Limits with Support for 200 kW Immersion RacksFallback source (no openly available primary found): GRC’s ICEraQ Series 10 single-phase immersion module, whose CDU supports up to 200 kW with warm water and 368 kW with chilled water.
- 3M to Exit PFAS Manufacturing by the End of 20253M will stop making fluoropolymers, fluorinated fluids and PFAS-based additives by the end of 2025.
- Power Stabilization for AI Training DatacentersSynchronous training alternates compute-heavy and communication-heavy phases, so cluster power swings by tens to hundreds of megawatts at frequencies that can excite grid components; mitigations in software, GPU firmware (ramp limits, power floors) and rack energy storage.
- How much electricity does an American home use?In 2022 the average U.S. residential utility customer used 10,791 kWh of electricity.
- Table of specific heat capacitiesAir at room conditions: c_p ≈ 1.012 J/(g·K), volumetric ≈ 0.00121 J/(cm³·K); liquid water at 25 °C: c_p ≈ 4.181 J/(g·K).