A chip has billions of tiny switches called transistors, and every one of them runs on electricity. Power planning designs the wires that bring electricity to all of them. Those wires are called the .
When the push gets weaker, transistors switch more slowly. If it sags too much in a busy spot, signals arrive late and the chip makes mistakes. Power planning keeps that sag small.
It happens right after the floorplan, before the other parts are placed. The power wires claim their space first.
A chip is built from billions of transistors, tiny electrically controlled switches, grouped into small prebuilt building blocks called (logic gates, and memory bits called flip-flops). Every cell needs two power connections: , the supply voltage, and VSS, ground. A modern chip runs on less than one volt but can draw tens or even hundreds of amps, and all of that current has to reach cells spread across the whole chip.
It travels on metal. Above the transistors sit ten or more layers of copper wiring, the , numbered from M1 at the bottom (thin, closely packed wires) to around M10 at the top (thick, wide ones). Power planning designs the part of the that lives on these layers: a mesh of power wires that carries VDD and VSS from the chip’s connections to its package down to every cell. That network has several jobs: carry both the average current and the peaks, keep the supply steady with little electrical noise, give signals a path for their return current, avoid wires slowly wearing out (an effect called ) or overheating, and use as little chip area and wiring as possible.1
Why it matters comes down to one rule, Ohm’s law: the voltage lost along a wire equals the current through it times its resistance, . The loss is called . Take a processor that uses 100 W at 1 V. It draws 100 A, so a resistance of just 1 mΩ (a thousandth of an ohm) on the way in loses 100 mV, 10% of the supply. Transistors switch more slowly at lower voltage, so a drop that large makes gates slower and costs the chip speed.20
The clock is the chip’s metronome: a signal that ticks billions of times a second, with each step of work having to finish before the next tick. If a sag in voltage makes some logic too slow, its result arrives after the tick and the chip computes a wrong answer, unless the whole chip runs at a slower clock.
Power planning comes right after floorplanning, which fixes the chip’s outline, where large blocks such as memories go, and where the connections to the package sit. It comes before placement, which positions the millions of individual cells. The power wires go in first because they claim space that placement and signal routing must work around. Commercial place-and-route tools (Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler) build the grid from templates, and power-analysis tools such as Ansys RedHawk-SC and Cadence Voltus check it. The open-source OpenROAD flow uses a generator called pdngen to build the grid and an analyzer called PSM to compute the voltage drop.23 The final checks on the finished layout are covered in Signoff.
You know that a design becomes a netlist of logic gates and is then laid out. Every one of those gates, the , needs a supply connection () and a ground connection (VSS), and power planning builds the wiring that provides them: the on-chip part of the . The wiring is a mesh of metal on the chip’s ten or so stacked wiring layers. Its quality is measured by how much voltage is lost on the way to each cell, the (current times resistance ), and by whether any wire carries enough current to wear out over the product’s life, the limit.
The useful way to think about the stage is as a budget in three currencies. The first is routing tracks: each metal layer has evenly spaced lanes that wires must run along, and every lane given to power is taken from signal and clock wires. The second is chip area, for capacitor cells that steady the supply and switch cells that turn blocks off. The third is voltage margin, how far the supply may sag before logic becomes too slow to meet the clock. A denser grid has lower resistance and makes the drop and wear-out limits easier to meet, but it uses more tracks.6 At the most advanced manufacturing nodes power wiring takes at least 20% of routing resources, while the power rails inside every cell limit how small the cells can be made.5
The grid is not built once. A first version is built from predefined patterns (templates) at floorplan time, using estimates of each block’s power and the positions of the package connections. It is refined after placement, when the real current and routing-congestion maps are known.6 Final verification on the routed layout is part of Signoff. Two things complicate the picture. Blocks that can be switched off, or that run at their own voltage, each need their own grid, switches and boundary cells. And at the newest manufacturing nodes the whole grid can move to the back of the silicon wafer. The sections below take each piece in turn: the budget for voltage sag, where the worst point sits and why, how wire wear-out is checked, switched domains, the package, and finally the solvers that compute all of it.
- Routing resources taken by power interconnect at advanced nodes
- ≥20%
- IR drop reduction from backside power with buried rails (imec and Arm simulation)
- 7×
Source: Horiguchi and Beyne at imec.5
60 A × 1 mΩ = 60 mV lost on the way, so the cells see 0.94 V (6% drop). Within the 10% goal.
Power planning starts with the floorplan and a guess of how much electricity each part will use. It ends with the power wires drawn in. It also makes a color map that shows where the push sags.
Chip design tools pass work to each other as files, most of them in a few standard text formats. The ones that matter here:
- (Design Exchange Format) records the physical state of the chip: its outline, the rows that cells sit in, and where every block and wire is.
- (Library Exchange Format) describes the building materials: each metal layer’s allowed widths and spacings, and the outline and connection points (pins) of every cell and large block.
- Liberty (
.lib) lists each cell’s speed and power use. - SAIF and VCD record how often each signal switched while the design was simulated running real work. Switching is what draws current, so these files say where the current will be.
- (Unified Power Format, the standard IEEE 1801) says which parts of the chip can be switched off or run at a different voltage.
| Direction | What | Typical format |
|---|---|---|
| In | Floorplan: the logic area’s outline, cell rows, positions of large blocks such as memories, and positions of the connections to the package | DEF or the tool’s own database |
| In | Metal layers: widths, spacings, rules for the vertical connectors (vias) | Technology LEF |
| In | Outlines of cells and large blocks, with the shapes of their power pins | Cell and block LEF |
| In | How much power each cell uses when it switches and when it sits idle | Liberty .lib |
| In | How often each signal switches, once simulations exist | SAIF or VCD |
| In | Which blocks can switch off or run at other voltages, and the special cells that requires | UPF (IEEE 1801) |
| In | The most current each layer and via may carry without wearing out | Rule files from the chip factory (foundry) |
| Out | The power wiring: rings, straps, rails and via stacks | DEF SPECIALNETS |
| Out | Power-switch and decap cells, where used | DEF component list |
| Out | Maps of voltage drop, the worst-drop number, and current in each wire | Text reports and per-cell voltage files |
The power wiring is stored apart from ordinary signal wiring. DEF has a separate SPECIALNETS section for it: ordinary signal routers don’t connect to its pins, and each shape is labeled by its role, such as RING, STRIPE or FOLLOWPIN (the rails that connect the cells to the power structure).4
The inputs are the floorplan, the technology’s layer rules and the cell outlines (in the and file formats), each cell’s power data (Liberty), switching activity from simulation (SAIF or VCD files) once it exists, the power intent (), and the foundry’s current limits. The outputs are the wiring, written to DEF as special nets, plus any switch and capacitor cells, and analysis reports. Two of those inputs decide whether an early voltage-drop number means anything, and both are easy to get wrong.
Where the current is. The analysis needs to know how much current each point of the grid draws. Before placement no cell has a position yet, so the current is estimated per block and spread evenly over the block’s area. After placement, OpenROAD’s analyzer computes each cell’s power itself and lets you override individual cells by hand.3 A uniform spread hides the hot spots that placement will create, so early numbers are a sanity check, not a verdict.
Where the supply comes from. Power enters the chip through solder bumps on its face ( packaging) or pads around its edge. The analysis has to model those entry points as voltage sources, and OpenROAD’s analyzer offers three choices: treat every point on the top metal layer as a perfect supply, model an array of bumps, or model straps on an imaginary layer above the top. Its bump model defaults to a 140 µm spacing, can leave out bumps to account for the ones that carry signals instead of power, and can add an external resistance to stand in for the package.3 The “every point is a perfect supply” option is the optimistic one: it ignores the voltage lost traveling sideways from a bump, and the points between bumps are often where the worst drop is. Record which model a reported number used.
Tap or hover a file to see what it holds.
Think of the power wires as roads. Electricity comes into the chip through thousands of tiny solder dots on its underside.
From there it travels on highways, wide , and then on streets, thin that touch every part. The picture below builds them one layer at a time.
The push sags most far from the solder dots, and in busy spots that use a lot of electricity. It also dips when lots of parts switch at the same instant, like everyone in a building turning on the shower at the same moment.
For those sudden dips, engineers sprinkle near busy parts. They store a little electricity close by, like a water tower in a neighborhood.
From package to transistor
- Package and bumps or pads. A chip sits in a package, the protective case that connects it to the circuit board. In most large chips the connections are tiny solder balls called bumps, spread over the chip’s whole face, and the chip is mounted face-down on the package (). Older and smaller chips connect through pads around their edge instead, joined to the package by fine bond wires.1 Many bumps or pads carry nothing but power and ground.
- Rings. A is a loop of wide supply and ground wires around the logic area (the core) or around a large block such as a memory. Straps inside the loop extend out to connect to it.2
- Straps. are wide supply and ground wires on the upper metal layers, repeated every so many micrometers (µm, thousandths of a millimeter). Each layer’s wires run in one direction, so straps on neighboring layers cross at right angles and together form a mesh. A strap pattern is described by its layer, width, pitch (the repeat distance) and offset (where the first one starts).2
- Rails. Cells are placed side by side in rows, like books on shelves. Thin on the lowest metal layer run along each row, and every cell touches them; their spacing is set by the rows.2
- Vias. A is a tiny vertical plug joining a wire to the layer above or below. Wherever a supply strap crosses a supply wire on the next layer, a joins them, and stacks of vias step the current down to the rails.2
IR drop
Every stretch of wire and every via is a small resistor, so current flowing to a cell loses some voltage on the way (). Suppose a block of logic draws 1.5 A and the wiring between it and the bumps adds up to 20 mΩ. The drop is : 30 mV, or 4% of a 0.75 V supply. Engineers separate two cases. Static IR drop uses each cell’s average current, and depends only on resistance. Dynamic IR drop is the brief, deeper dip caused by local bursts of demand when many cells switch at once, and it has to be kept within bounds for the chip to meet its timing.7 Dynamic dips also include noise from the package, explained below. A common goal is to keep the total sag under about 10% of the supply.1 (Where the current comes from in the first place, the charging of tiny capacitances each time a transistor switches plus a constant trickle of leakage, is covered from the transistor side in Speed and power.)
Why the drop matters: each cell’s delay is measured in advance at a stated supply, say 0.75 V, and stored in its Liberty file. The timing analysis that checks whether every step of work finishes before the next clock tick uses those numbers. A cell that really sees 0.71 V is slower than its file says. The drop eats into noise margins and adds gate delay, so it costs real performance.20 So the timing team and the power team have to agree on how much drop the timing numbers allow for.
Electromigration
Current in a metal wire is a flow of electrons, and at the current levels on a chip the electrons push metal atoms along with them. Over years this opens holes (voids) that can break the wire. Power wires carry current almost always in the same direction, which is the harder case, and the limits are set as a maximum current density (current per unit of the wire’s cross-section) for each layer.8 A wider wire carries the same current at lower density, and a hotter wire fails sooner.9 Vias are a weak point: the via arrays in a power grid are a known site of this wear, so the number of via cuts matters as much as the width of the wire.8
Decoupling capacitance
Two more electrical ideas are needed here. A capacitor stores a small amount of electric charge and can release it almost instantly. Inductance is the property of any current path that resists a change in current: the voltage lost across it is , the inductance times how fast the current changes. The package’s wiring has inductance, so a sudden jump in demand makes the supply dip.
When a block suddenly draws more current, the regulator on the board can’t respond within nanoseconds (billionths of a second), and the package’s inductance resists the change. The first dip has to be covered by capacitance on the chip itself.1 Designers add : cells with no logic function that hold charge between VDD and VSS. At high frequency that charge only reaches logic within a short distance, so decap belongs next to the busy logic.12 OpenROAD’s insert_decap command adds decap cells where the analyzed drop is worst, up to a chosen total capacitance.3 Decap built from transistors costs chip area and adds leakage, a small current that flows all the time.11
Choosing the grid
The main choices are which layers carry straps, how wide the straps are, how far apart they repeat, and how big the via arrays are. Wires on each layer run along evenly spaced lanes called , and every lane a power strap covers is one a signal can’t use. With illustrative numbers: a supply strap and a ground strap, each 1.6 µm wide, every 20 µm, cover 3.2 µm of every 20 µm, or 16% of that layer before spacing rules. Halving the pitch puts twice as many straps side by side, which roughly halves that layer’s resistance and doubles the lanes lost. A grid that uses more of the wiring has lower resistance between the transistors and the supply, but may leave too little room for important signal and clock wires.6 So grids are often denser where current is high and sparser where wiring is crowded, and large blocks such as memories get their own grid pattern with a keep-out margin (a halo) around them.62
Splitting the droop budget
Start from the number the timing analysis assumes. Each cell’s delay was characterized at some supply voltage, and the chip only meets its clock if the voltage each cell really sees is not much lower. The total allowed sag, often held under about 10% of , has two main sources: IR drop through the resistance of the grid, and noise, the voltage lost when current changes quickly through the inductance of the package and bumps.1 They depend on different things:
- Static drop is set by grid resistance and average current. You fix it with wider or denser straps and more via cuts.
- Dynamic drop is set by how fast current changes, how much charge-storing capacitance (decap) sits close enough to help, and the package. You fix it with decap placement, package design, and limiting how fast activity ramps up.
Treat each as a line item, and check that their sum matches the voltage the timing libraries assume. Dynamic drop also depends on what the chip is assumed to be doing. Vector-based analysis replays switching activity recorded from simulating real workloads (a VCD waveform file). It is accurate for the time windows it covers, but covering a large chip needs a huge number of patterns, and good patterns arrive late in the project. assumes the activity instead. It is faster and available earlier, but harder to make accurate.7
Where the worst node is
Picture the grid as a network of resistors with supply connections at fixed points. The voltage a cell sees is the supply minus the drop along its path from the supply points, so the worst point is wherever the paths are longest and carry the most current. In a wire-bond chip, power enters only around the edge, so the center of the die has the longest path. In a flip-chip design the bumps form a regular array, and the points farthest from any bump sit at the center of the square formed by four neighboring bumps.22 Not every bump site carries power, since some carry signals, so the power bumps are thinned out and sit farther apart than the bump spacing alone suggests.3 Zoom in and the same logic repeats. Between two straps, current has to travel sideways along the thin, resistive rails on the lowest layer to reach the nearest via stack, so the low points sit midway between straps. The simulation below shows exactly this.
Via stacks and track cost
To get from a strap on M7 down to a rail on M1, current passes through a via on every layer in between. Each intermediate layer needs a small metal landing pad for the vias above and below it, and a large via array grows those pads into patches that block the signal lanes next to them. OpenROAD’s grid generator lets you cap the rows and columns of each via array, snap intermediate layers onto the routing lanes, and keep pass-through layers at minimum width, so the stack uses a single row of cuts across that layer and doesn’t spread into neighboring lanes.2 A second lever is to vary the pitch by region. Template-based grid synthesis picks a pitch for each region from a small library of patterns designed to join cleanly at their edges. On average it frees thousands of routing tracks in crowded regions compared with one uniform grid, while still meeting the IR and electromigration limits.6
Checking wire wear-out
Electromigration checks follow a recipe. Compute the current in every wire segment and via, from the same resistor-network solution that gives the voltage drop. Divide by the cross-section to get current density. Compare it with a limit derived from Black’s equation, an empirical formula for how a wire’s expected life falls as current density and temperature rise (written out in Under the hood). Then clear any short segment that passes the Blech test: in a short enough wire, the mechanical stress that builds up pushes atoms back as hard as the current pushes them forward, and the wire never fails.8 Power grids are redundant, with many parallel paths. A failed segment raises the voltage drop elsewhere without immediately breaking the chip, but it pushes more current, and so a higher chance of failure, into the paths that remain.8
There is a convenient link between the two checks. The stress that builds up along a segment is proportional to its current density times its length, and that product is proportional to the segment’s voltage drop. So optimizing a grid for IR drop also tends to improve its resistance to electromigration.8 OpenROAD’s analyzer can write out the current in every grid segment (its -enable_em option), which is enough to spot undersized via arrays early.3
Decap placement
When current changes very fast, a capacitor far away is little help: its charge has to travel through the grid, and the grid’s inductance limits how quickly it can arrive. At high frequencies the supply of charge is highly localized, and the useful amount of decap is set mainly by the grid’s inductance.12 Decap far from a hotspot does little for the first dip, even if the chip-wide total looks generous. Flows therefore reserve room for decap near clusters of busy logic, or insert it where the analysis shows the worst drop.3 How much to add is limited by the area and leakage it costs.11 Decap also adds to the chip’s total capacitance, which forms a resonant circuit with the package’s inductance (see Package, board and the backside below).10
vias: Via arrays join straps wherever two wires of the same net cross, and stacks of vias step the current down to the rails.
Worst dip 97 mV (13% of 0.75 V) against a 10% goal; the steady IR drop afterwards is only 12 mV. Only the block’s own capacitance helps.
A phone chip has parts that sit idle most of the time. Even idle transistors leak a little electricity, like a dripping tap. Billions of drips add up and drain the battery.
So chips can switch idle parts off completely. This is called : big switches sit between the power wires and the part, like a light switch for a whole room.
Waking a part back up pulls a sudden rush of electricity. If all the switches closed at once, the push would dip for the rest of the chip. So they close a few at a time.
Other parts run on a weaker push to save power. Each part that switches off or runs on its own push needs its own set of power wires.
Power gating
Even a transistor that is switched off lets a tiny current leak through, and billions of them add up to real battery drain. cuts the supply to a block that isn’t in use. A is a transistor used as an on/off switch between the real supply and the block’s own local supply wire, its “virtual” supply. A low-leakage PMOS transistor on the VDD side is called a header; an NMOS transistor on the VSS side is called a footer. (PMOS and NMOS are the two kinds of transistor, introduced in The switch.) In the common coarse-grain style, many switch cells are spread across the block, all feeding one shared virtual supply.14 In OpenROAD’s grid generator the upper mesh stays on the always-present supply, the cell rails carry the switched supply, and a switch cell is inserted wherever the rails meet the mesh.2
Waking a block recharges all the tiny capacitances inside it. If a large part of the circuit is switched on at once, the resulting can disturb the supply that working neighbors share. Designs limit it by turning the switches on in stages, for example by passing the “on” signal from switch to switch in a chain (a daisy chain).14 OpenROAD can wire the control as a star, with one signal driving every switch, or as a daisy chain, which needs each switch cell to have an acknowledge output that passes the signal on.2 The trade-off is between a lower surge and a shorter wake-up: turning on in stages keeps the surge down but takes longer.15
Two more cell types make gating safe. hold a powered-down block’s outputs at a fixed 0 or 1, so drifting, unknown values don’t reach logic that is still running. are flip-flops (the cells that store one bit from one clock tick to the next) with a low-leakage backup that keeps their value during shutdown, so the block can resume quickly.14 Both need power while the block is off, so an supply must reach into the switched block.
Multiple voltages and DVFS
Blocks that don’t need top speed can run at a lower voltage to save power. A logic 1 from a 0.7 V block may not register as a 1 in a 1.0 V block, so signals crossing between them pass through .16 Each group of logic that shares a supply is a . It gets its own region of the chip, a , and its own grid; in OpenROAD’s grid generator that is set_voltage_domain -region.2 (dynamic voltage and frequency scaling) changes a block’s voltage and clock speed while the chip runs. Switching power, the power spent charging and discharging capacitance, scales as (capacitance, voltage squared, the share of it switching each cycle, and clock frequency), so a lower voltage saves a lot. But the voltage needed for reliable operation depends on the clock frequency, so the clock has to slow down too.18 For the grid, that means meeting the drop budget at every operating point and the electromigration limits at the highest-current one.
UPF
All of this intent, which blocks switch off, which run at which voltage, and where isolation, level shifters and retention go, is written in (Unified Power Format, the standard IEEE 1801) as commands in the Tcl scripting language.17 Simulation, synthesis, place-and-route and verification tools all read it alongside the design, so the grid, the switch cells and the boundary cells come from one description.16
Switch sizing and sequencing
A power switch is a resistor in series with everything in its block while it is on, so the voltage lost across the switches comes out of the block’s drop budget. When it is off, it still leaks a little. Switches must be sized for the block’s switching current: more total width lowers the on-state drop, but adds leakage and area.14 More drive current also shortens wake-up but raises the surge.15 A wake-up sequence divides the switches into groups and turns the groups on one by one, in an order chosen to keep the surge under a limit while keeping wake-up short.15 In UPF a switch can have several control inputs, one per stage, and acknowledge outputs that report back to the power controller when each stage is on.16
A small always-on power controller runs the sequence, and the order follows from what each cell protects against. Isolation must be active whenever the block’s outputs might be unknown, so it goes on before power-down and comes off last, after the switches have acknowledged and the state is back. Retention registers save their state before the supply goes and restore it once the supply is stable again. Each of those control signals must exist in the design, and UPF names them in set_isolation_control and set_retention_control.16
The always-on supply inside switched blocks
Some cells inside a switched block must stay powered: isolation cells placed there, the backup latches of retention registers, and buffers that carry always-on signals across the block. They need the unswitched supply inside the block, which OpenROAD’s grid generator models as a secondary power net of the voltage domain.2 That net needs its own straps or pin connections, sized for everything that stays awake. It is easy to undersize because those cells are few, yet the retention cells all switch together during save and restore.
Power states drive analysis
A power state table lists the legal combinations of supply states, for example run, sleep and hibernate.16 Each row is a separate scenario for drop and electromigration analysis. The worst static drop is often in the all-on, highest-voltage state, but the worst dynamic event can be a wake-up: one block’s rush current hitting a grid it shares with a neighbor running at full speed.1415
Off: the header switches are open, the virtual VDD has drained to 0 V and leakage is cut. Isolation cells hold the block’s outputs at 0.
Power planning doesn’t stop at the chip’s edge. Electricity comes from the circuit board, then through the chip’s package, the case the chip sits in. A weak spot anywhere along that path shows up as sag on the chip. So the chip and package teams plan it together.
The newest chips go further. They put the power wires on the back of the chip, using . That leaves the top almost entirely for signal wires.
The whole supply chain
The chip’s supply is made by a voltage regulator on the circuit board. From there it crosses the board’s copper layers, the package, the bumps, and only then the on-chip grid. Capacitors sit at every stage: large ones on the board, smaller ones in the package, and the on-chip decap. Each piece of wiring adds resistance and inductance.1
Target impedance
Engineers describe that whole chain with one number at each frequency: its impedance, which is resistance generalized to currents that change over time. The is , the allowed ripple divided by the largest sudden change in current.13 For a 0.75 V supply with 5% allowed ripple and a 10 A step, . No single part can hold that at every frequency. Large capacitors near the regulator work for slow changes; small ones near and on the chip work for fast ones.1 The fastest changes are handled on the chip and in the package, slower ones on the board.13
Backside power
In a conventional chip the transistors sit at the surface of the silicon, and all the power and signal wiring is built in layers above them. Backside power moves the power wiring underneath. In Intel’s version, called PowerVia, deep, narrow holes filled with metal (nano-TSVs, “through-silicon vias”) are made alongside the transistors. After the front-side wiring is built, a blank carrier wafer is bonded on top and the stack is flipped. The original wafer is then ground away until the nano-TSVs are exposed, and thick power wiring is built on that side.19 In a simulation of an Arm processor core, published with imec, backside delivery with power rails buried below the transistors cut IR drop by 7× compared with conventional front-side delivery.5 Intel’s PowerVia test cores ran more than 6% faster and lost 30% less power in delivery.19
The impedance profile
Chip, package and board form one electrical network, and its impedance (resistance generalized to changing currents) varies with frequency. Board components set it at low frequencies, the package in the middle, and on-chip elements at high frequencies.11 Two consequences follow. First, the , (the allowed ripple divided by the largest current step), has to be met across the whole range, not at one frequency.13 Second, the on-chip capacitance and the package inductance together form a resonant circuit, like a mass on a spring, with a main resonance at about 100–200 MHz in one published analysis of a 90 nm processor. At that frequency die and package act as one, and decap anywhere on the die helps; above it, decap helps only locally.10
Harris’s Pentium 4 example shows the result: an impedance spike near 100 MHz caused by package inductance. When current jumps in a step, the supply dips three times, and each dip is answered by a different stage: the first by on-chip capacitance, the second by the package’s capacitors, the third by the board’s.1 The chip team controls the first dip through decap and grid design, and that is where power planning meets package design. Because the resonance frequency depends on the die capacitance (see Under the hood), every added nanofarad of decap moves the resonance and changes its peak, so an accurate analysis of the on-chip grid needs a package model.10 For the static (DC) analysis, OpenROAD’s external-resistance setting is the simplest way to put the package in the loop.3
Backside delivery in practice
moves the grid under the transistors. imec’s approach combines buried power rails, which replace the cell rails on the lowest metal layer with metal lines about 30 nm wide at about 100 nm pitch sunk below the transistors, with nano-TSVs at 200 nm pitch that reach them from the back without taking any standard-cell area.5 Taking the power wiring off the front relieves signal congestion: PowerVia test cores packed some regions up to 95% full.19 The planning problem changes with it. The backside stack has its own layers, widths and via rules; front-side straps and their track cost largely disappear; and cell libraries change, because the rails move below the transistors.5 Heat removal needed new design rules, and debugging needed new methods, because the transistors are sandwiched between two wiring stacks.19
Highest peak 50 mΩ near 84 MHz, over the target. More on-chip capacitance lowers the peak and moves it down in frequency.
Front-side: power takes 14 of 60 tracks in this slice, and rails on the lowest layer inside every cell.
The simulation shows a small piece of chip as a grid. Electricity enters at the top and bottom edges. Colors show how much the push sags at each point.
Change how close together the thick power wires are, and how wide they are. Then turn on the “Hotspot”, a busy patch that uses much more electricity. The chip is fine until the hotspot appears. Can you fix it?
The simulation models a small patch of chip as a 24 × 24 grid of points, each standing for a group of cells that draws the same small current, with a 0.75 V supply. Thin horizontal rails on the lowest metal layer join each point to its neighbors. Thicker vertical straps on an upper layer connect to the supply at the top and bottom edges and feed the rails through vias. You set the strap pitch (a strap every 2–12 points, default 8) and the strap width (1–4 units, default 2). The readout gives the worst IR drop in millivolts and as a share of the supply, against a budget of 5%, which is 37.5 mV. The defaults pass. Now turn on the Hotspot, a 6 × 6 patch drawing four times the current, and they fail. Find the cheapest change that brings the drop back within budget, and watch how much of the upper layer’s routing tracks the straps use up.
Behind the map, the simulation writes one current-balance equation per grid point (the system explained in Under the hood) and solves it by successive over-relaxation, a simple method that sweeps the grid over and over, nudging each point’s voltage toward the value its neighbors imply, until nothing changes. It reports how many sweeps that took, along with the worst drop and the share of tracks the straps use. With the hotspot on, compare two fixes: a narrower pitch everywhere, or wider straps. Both cost tracks; note which buys more millivolts per percent of tracks spent. Then look at where the worst drop sits relative to the straps and the edge connections, and think about how a real flow would respond: denser straps only over the hotspot, decap next to it, or a different bump arrangement. Watch the sweep count too. This kind of simple relaxation slows down as grids grow, which is why real tools use the solvers in Under the hood.20
After the first check, the engineer sees the chip painted like a weather map. It is green where the push is healthy and red where it sags.
First they check the reddest spot. It is allowed to sag only a little, often about 5 percent. Then they ask why it is red. Red over a busy block needs more power wires right there. Red far from the solder dots means the package plan needs work.
Engineers don’t draw the power grid by hand. They write a short script of rules, and a tool generates the millions of wire shapes. Below is an illustrative script for the open-source OpenROAD tools. It builds a grid over the logic area and a separate grid over the on-chip memory blocks (SRAM “macros”, large prebuilt blocks), then checks it and computes the voltage drop. Layer names (M1 at the bottom to M10 at the top), widths and pitches are generic; lengths are in micrometers. The notes explain each line.23
# pdn.tcl (illustrative, OpenROAD pdngen and psm syntax; units in um)
add_global_connection -net VDD -pin_pattern {^VDD$} -power
add_global_connection -net VSS -pin_pattern {^VSS$} -ground
set_voltage_domain -power VDD -ground VSS
define_pdn_grid -name core_grid
add_pdn_ring -grid core_grid -layers {M9 M10} -widths 5.0 -spacings 2.0 -core_offsets 4.0
add_pdn_stripe -grid core_grid -layer M1 -followpins
add_pdn_stripe -grid core_grid -layer M4 -width 0.48 -pitch 20.0 -offset 2.0
add_pdn_stripe -grid core_grid -layer M7 -width 1.60 -pitch 40.0 -offset 10.0
add_pdn_stripe -grid core_grid -layer M9 -width 3.20 -pitch 80.0 -offset 20.0 -extend_to_core_ring
add_pdn_stripe -grid core_grid -layer M10 -width 3.20 -pitch 80.0 -offset 20.0 -extend_to_core_ring
add_pdn_connect -grid core_grid -layers {M1 M4}
add_pdn_connect -grid core_grid -layers {M4 M7}
add_pdn_connect -grid core_grid -layers {M7 M9}
add_pdn_connect -grid core_grid -layers {M9 M10}
define_pdn_grid -macro -name sram_grid -cells {SRAM_2KX64} -grid_over_pg_pins -halo {2.0}
add_pdn_stripe -grid sram_grid -layer M5 -width 0.96 -pitch 10.0 -offset 1.0
add_pdn_connect -grid sram_grid -layers {M4 M5}
add_pdn_connect -grid sram_grid -layers {M5 M7}
pdngen
check_power_grid -net VDD -floorplanning
# After global placement, when per-instance power is available:
set_pdnsim_source_settings -bump_dx 140 -bump_dy 140
analyze_power_grid -net VDD -source_type BUMPS -enable_em -em_outfile vdd_em.rpt \
-voltage_file vdd_inst_voltage.rpt
analyze_power_grid -net VSS -source_type BUMPS- 1L2Every cell has pins named VDD and VSS. These two lines tie all of them to the supply and ground nets; without them the grid would have nothing to connect to.
- 2L4One voltage domain covering the whole logic area. Blocks that switch off or run at another voltage would each get their own domain and region.
- 3L7A ring of supply and ground wires around the logic area, on the two top layers (one horizontal, one vertical).
- 4L8The thin rails on M1 that run along every row of cells. Their spacing comes from the rows, so no width or pitch is given.
- 5L9The first strap layer above the rails. Straps close together mean current travels only a short way along the thin rails to reach a via stack.
- 6L10Mid-level straps: 2 × 1.60 µm of metal in every 40 µm, which uses 8% of M7 before spacing rules.
- 7L11Top straps on the thickest layers, wide and far apart. The bumps feed these directly.
- 8L13Each connect rule adds via arrays wherever supply wires (or ground wires) on the two layers cross.
- 9L18Memory blocks get their own pattern over their power pins. The halo is a keep-out margin that holds cells and core straps away from the block’s edge.
- 10L19M5 straps over the memory. In this example the memory’s power pins are on M4 and it uses M1 to M4 inside, so the grid must not cut through those layers.
- 11L23Builds the grid and writes it to the design as special (power) wiring.
- 12L24Before placement: checks for straps that connect to nothing and for unconnected pins, ignoring cells that don’t have a final position yet.
- 13L27Models the package bumps as a grid with 140 µm spacing, the analyzer’s default.
- 14L28Static IR drop on VDD, also writing the current in every segment for the electromigration check. Analyze VSS too: a rise on ground eats into the same budget.
The result is a worst-drop number and a color map of the chip. Read the map before the number: where the worst point sits tells you which layer or which connection to fix.
Here are three files from one illustrative design: a chip block containing a DSP (a digital signal processor, a unit full of multiply-and-add hardware) that can be switched off when idle. First, the summary report from static IR-drop and electromigration (EM) analysis after placement, run twice: before and after a local fix. Read it top to bottom: the setup, the result, where the drop accumulates layer by layer, and what changed in the second run.
==== Static IR / EM summary (illustrative) ====
Net: VDD Nominal 0.750 V Budget 5.0% = 37.5 mV Sources: bumps, 140 um pitch
Total current 1.84 A Grid: M1 rails, M4/M7/M9/M10 straps
RUN 1 (uniform grid)
Worst drop 41.2 mV (5.5%) u_dsp/u_mac_array (612.4, 488.0)
Instances > budget 3,418
Drop by layer at worst node:
bump + M10/M9 6.1 mV
M7 7.9 mV
M4 9.6 mV
M1 rails + vias 17.6 mV
EM: 2 via arrays V1-V3 over Javg limit (1.21x, 1.08x) at u_mac_array
RUN 2 (M4 pitch 20 -> 10 over u_dsp only, V1-V3 arrays 2x3 -> 3x3, decap added)
Worst drop 29.8 mV (4.0%) u_dsp/u_mac_array (618.0, 492.6)
Instances > budget 0
EM: 0 segments over limit
M4 track usage 4.8% core-wide, 9.6% inside u_dsp region
Congestion delta +2.1% overflow in u_dsp region (global route estimate)- 1L2The setup: nominal supply, the budget (5% of 0.75 V), and the source model. Always record the source model; treating the whole top layer as a perfect supply would hide the drop between bumps.
- 2L6The worst point is 41.2 mV down, 5.5% against a 5.0% budget, inside the DSP’s multiplier array (coordinates in µm). One busy block, not the whole chip.
- 3L7How many cells see more drop than the budget. All of them are near the hotspot.
- 4L8The drop at the worst point, split by where it happens on the way down. The top of the stack is healthy.
- 5L12The largest share, over 40%, is lost in the M1 rails and via stacks: current travels too far along thin M1 to reach a stack. So the fix belongs low in the stack.
- 6L13Two via arrays (stacks through via levels V1 to V3, near the bottom of the metal stack) carry more than the allowed average current density (Javg): 21% and 8% over. Same place as the drop hotspot, so widening for IR and fixing EM pull the same way.
- 7L15The fix: M4 straps twice as close together over the DSP only, via arrays grown from 2 × 3 to 3 × 3 cuts, plus decap. Changing the pitch everywhere would cost tracks in regions that never needed them.
- 8L16Now 4.0%, inside budget, with no cells over and no EM violations.
- 9L19The cost of the fix: M4 tracks used by power doubled, but only inside the DSP region.
- 10L20A quick routing estimate shows 2.1% more overflow (more wires wanting a region than it has tracks for) in that region. Check this before accepting a grid change: a clean drop map with unroutable signals is no win.
Second, the power intent for the DSP, written in UPF (IEEE 1801). It declares two power domains, the always-on top level (PD_TOP) and the switchable DSP (PD_DSP), and the supply nets that feed them. The DSP switches off in sleep, keeps some state in retention registers, and holds its outputs at 0 through isolation cells while it is off. Its power switch has two control inputs, so it can turn on in two stages, and a power state table lists which supply combinations are legal.16
# dsp_power.upf (illustrative IEEE 1801 power intent)
create_power_domain PD_TOP
create_power_domain PD_DSP -elements {u_dsp}
create_supply_port VDD
create_supply_port VSS
create_supply_net VDD -domain PD_TOP
create_supply_net VSS -domain PD_TOP
create_supply_net VSS -domain PD_DSP -reuse
create_supply_net VDD_DSP_SW -domain PD_DSP
connect_supply_net VDD -ports VDD
connect_supply_net VSS -ports VSS
set_domain_supply_net PD_TOP -primary_power_net VDD -primary_ground_net VSS
set_domain_supply_net PD_DSP -primary_power_net VDD_DSP_SW -primary_ground_net VSS
create_power_switch SW_DSP -domain PD_DSP \
-input_supply_port {vin VDD} \
-output_supply_port {vout VDD_DSP_SW} \
-control_port {stage1_en u_pmu/dsp_stage1_en} \
-control_port {stage2_en u_pmu/dsp_stage2_en} \
-ack_port {stage1_ack u_pmu/dsp_stage1_ack {stage1_en}} \
-ack_port {stage2_ack u_pmu/dsp_stage2_ack {stage2_en}} \
-on_state {full_on vin {stage1_en & stage2_en}} \
-off_state {full_off {!stage1_en & !stage2_en}}
set_isolation ISO_DSP -domain PD_DSP -applies_to outputs -clamp_value 0 \
-isolation_power_net VDD -isolation_ground_net VSS
set_isolation_control ISO_DSP -domain PD_DSP \
-isolation_signal u_pmu/dsp_iso -isolation_sense high -location parent
set_retention RET_DSP -domain PD_DSP \
-retention_power_net VDD -retention_ground_net VSS
set_retention_control RET_DSP -domain PD_DSP \
-save_signal {u_pmu/dsp_ret posedge} -restore_signal {u_pmu/dsp_ret negedge}
add_port_state VDD -state {ON 0.75}
add_port_state SW_DSP/vout -state {ON 0.75} -state {OFF off}
add_port_state VSS -state {GND 0.0}
create_pst pst_main -supplies {VDD VDD_DSP_SW VSS}
add_pst_state RUN -pst pst_main -state {ON ON GND}
add_pst_state SLEEP -pst pst_main -state {ON OFF GND}- 1L3The DSP instance becomes its own domain. Everything else stays in PD_TOP.
- 2L10The switched supply. Only the switch drives it; the domain’s cell rails connect to this net.
- 3L14The DSP’s primary supply is the switched net, so its standard cells lose power in sleep.
- 4L16Header switch on VDD. The physical switch cells are placed by the power-grid tool to implement this.
- 5L19Two control ports let the power controller turn the switch cells on in two stages instead of all at once. Staging limits rush current.
- 6L21Acknowledge ports report when each stage has finished, so the power controller can sequence safely.
- 7L26Clamp DSP outputs to 0 while it is off. The clamp runs on the always-on VDD.
- 8L28Isolation cells sit in the parent domain, outside the switched region. The control signal must exist in the RTL.
- 9L31Retention registers keep their shadow latches on unswitched VDD.
- 10L33Save on the rising edge of dsp_ret before power-down, restore on the falling edge after power-up.
- 11L37Both domains run at 0.75 V, so no level shifters are needed. A different voltage here would add set_level_shifter.
- 12L39The power state table. Each row is a scenario for IR, EM, timing and power-aware verification.
Third, the physical side of the same intent, in OpenROAD’s grid generator (pdngen): a region of the chip whose cell rails carry the switched supply, with header switch cells wherever those rails meet the always-on mesh above, chained so they turn on one after another.2
# pdn_dsp.tcl (illustrative, OpenROAD pdngen syntax; units in um)
define_power_switch_cell -name HDRSW_X4 -control SLEEP -acknowledge SLEEP_ACK \
-power_switchable VDDV -power VDD -ground VSS
set_voltage_domain -name DSP -region dsp_region -power VDD -ground VSS \
-switched_power VDD_DSP_SW
define_pdn_grid -name dsp_grid -voltage_domains {DSP} \
-power_switch_cell HDRSW_X4 -power_control u_pmu/dsp_stage1_en \
-power_control_network DAISY
add_pdn_stripe -grid dsp_grid -layer M1 -followpins
add_pdn_stripe -grid dsp_grid -layer M4 -width 0.48 -pitch 10.0 -offset 1.0
add_pdn_connect -grid dsp_grid -layers {M1 M4}
add_pdn_connect -grid dsp_grid -layers {M4 M7}
pdngen- 1L2Describes the header cell’s pins. The acknowledge output is what makes a daisy chain possible.
- 2L3Unswitched VDD in, switched virtual VDD out.
- 3L5The region is the DSP voltage area from the floorplan. Its rails carry the switched supply.
- 4L9The chain starts from the UPF’s first-stage enable, dsp_stage1_en. The daisy chain is how the staged turn-on happens physically.
- 5L10DAISY passes the control from switch to switch, so they close in sequence. STAR would close them all at once.
- 6L12M4 at a 10 µm pitch inside the domain: the IR hotspot from the summary, fixed by construction.
- 7L13Where the switched M1 rails meet the unswitched M4 mesh, pdngen inserts a switch cell.
- 8L14M4 up to the core’s M7 straps. Everything above the rails stays on unswitched VDD.
Worst drop 43.0 mV (5.7%), over the 5% budget; straps use 8% of the layer. The hot spot is centered on the busy block, and bumps surround it: a local current problem, not a supply problem.
- Planning for an average day. Power wires that cope with normal use can sag badly when every part gets busy at once.
- Too many power wires. Thick power wires everywhere fix the sag, but leave no room for signal wires.
- Waking up too fast. Switching a sleeping part on all at once makes the push dip for the rest of the chip.
- Forgetting the package. The chip’s wires can be perfect and the supply still shaky if the package wasn’t planned with it.
- Straps connected to nothing, or missing vias. A strap that never reaches a supply connection, or a forgotten rule for joining two layers, leaves a region fed only through the thin rails, with a large drop. Automatic connectivity checks, such as OpenROAD’s
check_power_grid, catch these.3 - A memory block left unpowered. A grid pattern for a large block that doesn’t land on the block’s power pins leaves it without supply. Check every large block’s power connections, not just the standard cells.
- An optimistic model of the package. Treating the whole top layer as a perfect supply hides the voltage lost between bumps. Model the bumps or pads as they will actually be built.3
- Checking only the supply side. Ground wiring has resistance too, and a cell works on the difference between its local VDD and its local VSS. Analyze both.
- Static passes, dynamic fails. A grid sized for average current can still dip when a whole block switches at once. Run the time-varying (dynamic) analysis once switching activity is available.7
- A grid that crowds out the signals. Dense straps everywhere leave too few routing tracks, and the problem shows up only when the signal wires are drawn. Check routing congestion whenever the grid changes.6
- Decap in the wrong place. Plenty of total decap on paper, but far from the hotspot, does little for the first and fastest dip, because at high frequency charge only travels a short distance.12
- Vectorless results taken as truth, or ignored. Vectorless analysis assumes switching activity rather than measuring it. Taken literally, its pessimism drives overdesign; dismissed, it hides real risk. Calibrate it against waveform-based runs on known stress patterns.7
- Wake-up never analyzed. One block’s rush current at wake-up is a dynamic drop event for its neighbors. Simulate it with the real switch chain and schedule.15
- An undersized always-on supply. The unswitched supply inside a switchable block feeds few cells and gets few straps, then dips when all the retention registers save or restore at once.
- Via arrays as the weak link. Wide straps pass the electromigration check while the via arrays that feed them fail. Check via current as well as wire current.8
- Late grid changes. The grid is meant to be set from templates at floorplan time and refined incrementally after placement.6 Changing strap pitch wholesale after placement moves the blocked tracks under cells that are already placed, so earlier routing estimates no longer hold. Keep later changes local.
- Decap budgets without a package model. The die’s capacitance and the package’s inductance resonate together, so sizing decap without a package model misjudges that resonance.10
A strap with no vias to the rails (a missing connect rule): that region is fed only sideways through thin M1, so the drop there is large.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
The grid as a circuit
Every analysis tool starts by turning the layout into a circuit. Each wire segment and via becomes a resistor (plus capacitance for time-varying analysis); the network is linear, doesn’t change over time, and contains no sources of its own. Each bump or pad becomes an ideal voltage source, and each cell or block becomes a current that varies over time. Writing Kirchhoff’s current law (current in equals current out) at every node gives a matrix equation:
Here holds the node voltages, the conductances () between nodes, the capacitances, and the sources and cell currents.20 Kozhaya and colleagues leave out on-chip inductance as too small to matter, and Pant and colleagues found the on-die grid behaves as an RC network except for very fast, very local effects.10 Qian and colleagues argue that in the top few metal layers inductance can no longer be ignored, and include it in their solver.22 The challenge is scale: for a modern microprocessor the network can easily include millions of nodes and tens of millions of elements.20
Static analysis:
For static (DC) analysis, drop the capacitance. At a node with neighbors , Kirchhoff’s law says
the currents flowing in from the neighbors, each conductance times a voltage difference, add up to the current the node’s cells draw. Writing that for every node gives one big linear system, . The matrix is sparse (each node touches only a few neighbors) and symmetric positive definite, a property that the fast solvers below rely on.22
A two-node example makes it concrete. A rail is fed from through two segments of 2 Ω each (conductance 0.5 S), and each of the two nodes draws 1 mA. The equations are
Solving gives and . The far node sees 6 mV of drop: 4 mV in the first segment, which carries both nodes’ 2 mA, and 2 mV in the second. A real grid is the same thing with millions of rows, which is why the solver matters.
Direct solution: Cholesky factorization
The textbook way to solve is to factor into a lower-triangular matrix times its transpose (, Cholesky factorization) and then solve two easy triangular systems. Because is symmetric positive definite, Cholesky needs roughly half the computation of general LU factorization and no row swapping (pivoting) for stability. Reordering the nodes first, with heuristics such as minimum degree or nested dissection, limits fill-in, the new nonzero entries that factoring creates.21 For time-varying analysis with a fixed time step , the matrix to factor is at every step, and it never changes, so one factorization serves the whole simulation and each step costs only the two triangular solves. Kozhaya and colleagues found 100 steps per clock cycle sufficient for digital power grids.20 The weakness is size: solve time grows faster than linearly with the matrix, which pushes the largest grids toward iterative methods.20
Preconditioned conjugate gradient
Conjugate gradient is an iterative method for symmetric positive definite systems. It improves a guess step by step using only matrix–vector products, and the number of steps it needs depends on how badly conditioned the matrix is. A preconditioner is a cheap approximation of the matrix’s inverse that makes the problem better conditioned. A common choice is an incomplete Cholesky factorization, which limits fill-in by level or drops entries below a size threshold. Power-grid matrices are often badly conditioned, so this preconditioning is critical.21 The threshold is a dial: keeping more fill-in gives a more effective preconditioner but costs more to build and apply.21
Multigrid
Simple iterations such as Jacobi, Gauss–Seidel and the over-relaxation used in the simulation above update each node from its neighbors. Think of the error in the current guess as a pattern across the grid. These methods quickly flatten jagged, node-to-node error (high-frequency error) but remove smooth, slowly varying error very slowly, so they need more sweeps as the grid gets finer. Multigrid alternates the two things that work: a few relaxation sweeps to smooth the error, then a correction computed on a coarser grid, where smooth error looks jagged and is cheap to remove. A system of equations is then solved in work, proportional to its size.20 Kozhaya, Nassif and Najm adapted the idea to irregular power grids: reduce the grid to a coarser one, solve that, and map the solution back. Their method was fast for both static and time-varying analysis, with a worst-case error of about 16% in the voltage drop on their test designs, a trade that suits early planning better than final checks.20
Random walks
Rearranged, the node equation says
a node’s voltage is a weighted average of its neighbors’, minus a term for the current it draws. Qian, Nassif and Sapatnekar read this as a game. A walker starts at node and at each step moves to neighbor with probability , paying a “motel” cost of at every node it passes. Voltage sources are “homes”: reaching one ends the walk with a payout of . The expected money left at the end equals . Averaging walks estimates it, with an error whose variance falls as , so accuracy is a dial you trade against runtime.22
The method is local: one node’s voltage needs only walks that start there, which suits quick what-if checks on part of a grid, and nodes already solved can become new homes that shorten later walks. It struggles where homes are rare, as in wire-bond grids with supply pads only on the edges, where a walk from the center wanders for a very long time. It also struggles where an upper layer’s wires are much more resistive than the vias beneath it, so walkers keep wandering in the well-connected lower layers instead of climbing to the sources. Hierarchical variants address both. The paper reports a 71K-node flip-chip grid solved in 4.16 s, and time-varying analysis with inductance of a 642K-node grid at 2.1 s per time step.22
Each walk starts at the ringed node, steps to a random neighbor, pays I/Σg = 0.5 mV at every inner node (more at the edges, where Σg is smaller), and ends at a bump. The average spend is the drop. Exact answer: 16.63 mV.
Dynamic analysis with and without vectors
Vector-based dynamic analysis takes cell currents from simulating the design on real input patterns (recorded in VCD files), so it is only as complete as those patterns.7 Kouroussis and Najm framed the alternative by analogy with static timing analysis, which checks timing without any input patterns. They verify the grid by giving each current source an upper bound (a local constraint) and each group of sources, such as a block or the whole chip, a combined upper bound from its known power (a global constraint). Linear programming then finds the worst drop over every combination of currents that satisfies those bounds. With local constraints alone, every source can peak at once, and the result is very pessimistic; the global constraints bring it back toward reality.23 Practical vectorless analysis works from assumed switching activity instead. It is fast and early, but harder to make accurate than vector-based analysis.7
Black’s equation and the Blech limit
Electromigration limits come from Black’s equation:
The time to failure falls as a power of the current density , and exponentially as the temperature rises ( is an activation energy, Boltzmann’s constant, and a fitted constant). Black derived ; later work puts between 1 and 2. and the statistical spread of failure times come from accelerated tests on single wires at high current and temperature. Reliability engineers turn a failure-rate target and a product lifetime into a minimum for each wire, and so into a maximum allowed current density.8 A worked example: with , 25% more current density leaves of the original life. Blech found that short wires resist electromigration. If times the wire length stays below a technology constant, , the stress that builds up pushes atoms back as hard as the electron flow pushes them forward, and the wire is “immortal”.89
Target impedance and resonance
In the frequency domain the requirement is at every frequency .13 One of the hardest points to meet is the resonance of the die’s capacitance with the package’s inductance.10 An inductance and a capacitance resonate at
With illustrative values of 30 pH of effective package inductance and 100 nF of die capacitance, , close to the 100 MHz spike in Harris’s Pentium 4 example.1 Adding decap raises and so lowers and changes the peak, and the resistance in the loop (its damping) sets how high the peak rises. That is why chip and package teams size decap and package inductance together.
Q1In a chip whose solder bumps cover its whole face, which order does current follow from the package to a logic cell?
Q2A block draws 2 A through an effective grid resistance of 20 mΩ. is 0.75 V and the budget is 5%. Does it pass?
Q3What is the difference between static and dynamic IR drop?
Q4A block that runs at 0.6 V and can be switched off sends signals into a block that runs at 0.9 V and never switches off. Which cells does that crossing need?
Q5Why are a block’s power switches often wired in a chain, so each one passes the “on” signal to the next, instead of all receiving it at once?
Sources
Show Hide 23 sources
- Lecture 19: Packaging, Power, & Clock (slides for CMOS VLSI Design, 4th ed.)PDN functions; droop target under ±10% of VDD; IR drop and L·di/dt worked examples; regulator → board → package → bumps → on-chip capacitance; first droop handled by on-chip bypass caps, spike near 100 MHz from package L.
- Power Distribution Network Generator (pdn)pdngen commands: set_voltage_domain (-switched_power, -secondary_power), define_pdn_grid (-macro, -halo, -power_switch_cell, STAR/DAISY control), add_pdn_ring, add_pdn_stripe (-followpins, -width, -pitch), add_pdn_connect (via arrays, -min_width_layers), define_power_switch_cell.
- IR Drop Analysis (psm, based on PDNSim)Static IR analyzer: worst IR drop, worst current density, floating-stripe check; analyze_power_grid, check_power_grid, -source_type FULL/BUMPS/STRAPS, 140 µm default bump pitch, bump depopulation, external resistance; insert_decap near the highest IR drop.
- LEF/DEF 5.7 Language Reference (DEF syntax chapters)SPECIALNETS defines nets with special pins and special wiring; regular routers do not route to special pins; SHAPE RING, STRIPE, FOLLOWPIN (connects standard cells to power structures).
- How to power chips from the backsidePower interconnect takes at least 20% of routing resources; rails limit cell height scaling; 10% margin for loss between regulator and transistors; buried rails ~30 nm wide at ~100 nm pitch; nTSVs at 200 nm pitch; 7× lower IR drop in an Arm CPU simulation with imec (IEDM 2019).
- OpeNPDN: A Neural-network-based Framework for Power Delivery Network SynthesisPDN must meet IR, EM and congestion constraints; denser PDN lowers resistance but takes tracks from signals; region-wise template pitches free thousands of routing tracks versus a uniform grid.
- PowerNet: Transferable Dynamic IR Drop Estimation via Maximum Convolutional Neural NetworkDynamic IR drop definition; vectorless vs. VCD-based analysis: vector-based needs many patterns and arrives late, vectorless is faster and earlier but harder to make accurate.
- Invited: Toward Accurate, Large-scale Electromigration Analysis and Optimization in Integrated SystemsDC EM in power grids; Black’s equation (n revised to 1–2) and the Blech jL criterion in rule-based checks; via arrays as EM weak points; grid redundancy; jL ∝ IR so IR-driven sizing also helps EM.
- ElectromigrationBlack’s equation and its terms; temperature dependence; Blech length; wider wires lower current density; via current crowding.
- Power Grid Physics and Implications for CADDecaps act globally at the main resonance, set by their interaction with package inductance at about 100–200 MHz, and more locally above it; accurate on-die grid analysis needs a package model.
- Leveraging Symbiotic On-Die Decoupling CapacitanceLow-frequency impedance is set by the board, middle frequencies by the package and high frequencies by on-die elements; intentional MOS-gate decap needs die area and increases leakage current.
- On-die Decoupling Capacitance: Frequency Domain Analysis of Activity RadiusTarget impedance achieved by decap at board, package and die; at high frequencies decoupling charge is highly localized, defining an effective decoupling radius around each switching element.
- Power integrityTarget impedance Z = ΔV/ΔI; capacitors of several sizes for low impedance across frequency; high frequencies handled on die and package, lower ones on the PCB.
- Power gatingPMOS header and NMOS footer switches; fine- vs. coarse-grain; rush current and staged turn-on; daisy-chained switch control; retention registers; isolation cells.
- An Efficient NBTI-Aware Wake-Up Strategy for Power-Gated DesignsWaking a power-gated design can cause excessive surge current; a wake-up sequence divides the sleep transistors into groups that turn on one by one, chosen to minimize wake-up time while keeping surge current under a constraint.
- Low Power Design, Verification, and Implementation with IEEE 1801 UPF (tutorial)Power gating, multi-voltage and DVFS; level shifter, isolation and retention cells; UPF commands create_power_domain, create_supply_net, create_power_switch (control and ack ports), set_isolation, set_retention, set_level_shifter, add_port_state, create_pst.
- Unified Power FormatUPF is IEEE 1801, a set of Tcl extensions describing supplies, power states, switches, isolation, level shifters and retention.
- Voltage and frequency scalingSwitching power C·V²·A·f; the voltage needed for stable operation depends on the clock frequency and can be reduced if frequency is reduced; DVFS adjusts both together.
- Intel Is All-In on Backside Power DeliveryPowerVia test cores: more than 6% frequency gain, 30% less power loss, some regions up to 95% filled; nano-TSVs, carrier wafer bonding and thinning; new thermal rules and debug methods.
- A Multigrid-like Technique for Power Grid AnalysisGrid as RC(L) network with voltage sources and current drains, Gx + Cẋ = u; constant step reuses one factorization; classical iterations kill high-frequency error only; multigrid is O(N); coarsen-solve-map back with ~16% worst error in drop.
- A Technical Survey of Sparse Linear Solvers in Electronic Design AutomationPower integrity analysis solves large sparse systems; DC power-grid analysis often yields symmetric positive-definite matrices solved by Cholesky or conjugate gradient; power-grid matrices are often severely ill-conditioned, so preconditioning is critical.
- Power Grid Analysis Using Random WalksRandom-walk game equivalent to G·V = E; motel costs and home awards; error variance falls as 1/M; local computation; inductive effects in the top few layers; wire-bond grids and layers whose wires are far more resistive than the vias below them are slow; hardest node (most walks) is the one farthest from C4 pads; hierarchical and RKC extensions.
- A Static Pattern-Independent Technique for Power Grid Voltage Integrity VerificationVectorless verification as optimization under user-supplied current constraints (local and global), solved with linear programming; local-only constraints are very pessimistic.