By now the big blocks on the chip are in place, and the power wires are drawn. What’s left is up to millions of tiny building blocks called . Placement decides exactly where each one goes.
Every connection becomes a real wire, and long wires are slow. But if too many cells crowd one spot, the wires won’t fit. Try the three layouts below and watch both meters.
By this stage the design exists as a : a long list of small prebuilt circuits and the connections between them. Each circuit is a , such as a NAND gate or a (a one-bit memory that updates once per tick of the chip’s clock). A mid-sized block has hundreds of thousands of them. All standard cells have the same height, so they can be lined up in , like books on shelves. Standard cells, in the Transistors guide, shows what is inside one.
The step before, floorplanning, fixed the chip’s outline, the positions of large blocks such as memories (called ), the rows and the power wiring. Placement now gives every standard cell an exact spot in a row and the direction it faces. After it come clock tree synthesis, which wires the clock signal to every flip-flop, and routing, which draws all the other wires.
The placer balances four goals at once:1
- Short wires. Every connection becomes a metal wire, and longer wires are slower and use more power.
- Timing. A signal leaving one flip-flop has to pass through its gates and wires and reach the next flip-flop before the next clock tick. Paths that barely make it need their cells especially close together.
- Room for wiring. If too many connected cells crowd into one spot, their wires won’t all fit over it. This is called .
- Power. Wires that switch often waste the most energy, so they get priority for being short.
Placement is also where the netlist starts to change for physical reasons. Once the tool knows how far apart cells are, it adds (cells that re-drive a weak signal), swaps gates for stronger versions and reorders the test wiring, as the sections below explain.
Every chip-design flow has a placement engine: Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler, and the open-source OpenROAD project, whose global placer (gpl) grew out of the RePlAce research tool and whose detailed placer (dpl) grew out of OpenDP.23
Placement turns a gate-level netlist into coordinates. It decides how far apart connected cells sit, and so how long every wire will be. A wire’s resistance and capacitance grow with its length, so placement settles much of the remaining delay and power, and how hard routing will be. In the RePlAce authors’ words, placement quality directly affects timing closure, die utilization, routability and design turnaround time.4 Clock tree synthesis and routing can only work with the positions they are given.
Inside the tool, placement is four interleaved jobs:
- treats every cell’s as a continuous number and optimizes them all together: short estimated wirelength, with a penalty for any area that gets too full.
- snaps each cell onto the row-and-site grid without overlaps, moving each as little as possible.
- improves the legal result with small local moves.
- Physical synthesis edits the netlist using the new positions: it buffers long or heavily loaded nets, resizes gates, swaps cells to versions with a different (), faster or less leaky, and restructures logic. It works on estimated parasitics: wire resistance and capacitance worked out from the placed positions, because no real wires exist yet.25
Commercial tools (Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler) run these jobs inside their own placement-optimization commands. OpenROAD exposes each as a separate command, which makes it a good model of what happens inside. Finding the best placement is computationally intractable: global placement is NP-complete, and most placement sub-problems are at least NP-hard, so no known method finds the optimum in reasonable time as designs grow. Every engine is therefore a stack of heuristics and continuous relaxations.61 The academic state of the art for global placement is nonlinear electrostatic placement (ePlace, RePlAce, DREAMPlace), and OpenROAD’s gpl uses it.62
- Default target density, OpenROAD global_placement
- 0.70
- Default stopping overflow, OpenROAD global_placement
- 0.10
- DREAMPlace global placement speedup on GPU vs. multi-threaded RePlAce
- >30×
Those numbers come from the OpenROAD gpl documentation and the DREAMPlace paper. Target density is the most any small area may be filled; overflow is the share of cell area still in over-full areas when global placement stops.27
Random positions: total wirelength 451 µm, the longest of the three. Long wires are slow and power-hungry, and they also crowd the routing.
Placement starts with a parts list that says how the cells connect, and the floorplan. It also needs to know how big and how fast each kind of cell is.
It ends with an exact spot for every cell. Each one sits neatly in a row, and nothing overlaps.
| Direction | What | Format |
|---|---|---|
| In | The netlist: every cell and connection, after synthesis and after test logic (scan chains) was added | Verilog |
| In | The floorplan: chip and core outline, rows, placed macros, input/output pins, power wiring, tap and endcap cells, blockages, regions | DEF, or the tool’s own database (OpenROAD’s .odb) |
| In | Physical library: metal layers, the site size, each cell’s outline, pin shapes and no-go areas | LEF (a technology LEF and a cell LEF) |
| In | Timing library: each cell’s delay, power and area | Liberty (.lib) |
| In | Timing constraints: clock speeds, input and output timing, exceptions | SDC |
| In | Scan chain description: which flip-flops may be reordered | SCANCHAINS section of a DEF |
| Out | The placed design: every cell’s location and orientation, marked PLACED or FIXED | DEF or .odb |
| Out | Updated netlist with new buffers, resized cells and reordered scan chains | Verilog |
| Out | Reports: utilization, timing on estimated wires, congestion, legality | Text logs and GUI heat maps |
A word on the formats. Verilog is a hardware description language; here it is used only as a list of cell instances and the named connections, or nets, between their pins. and are plain-text formats for physical shapes: LEF describes the library (what each cell looks like from outside), and DEF describes one design (where everything is). holds each cell’s measured delay and power tables, and says how fast the clock runs.
LEF defines the , the smallest step a cell can be placed at, and every cell is a whole number of sites wide. DEF ROW statements line those sites up across the core.89 In the output DEF, each cell is marked PLACED, meaning later tools may still move it, or FIXED, meaning automatic tools must leave it alone.9
Two inputs are easy to overlook. The first is the resistance and capacitance per micron of wire on the metal layers (OpenROAD’s set_wire_rc). Before routing, every wire delay is estimated from the placed positions and these values, so a wrong value skews every timing decision the placer makes.25 The second is the scan DEF. Its SCANCHAINS section splits each chain into FLOATING segments, whose flip-flops may be reordered, and ORDERED segments, which must keep their order.9
Physical-only cells (taps, endcaps, fillers, decoupling capacitors) carry + SOURCE DIST in DEF. They connect only to power and ground and never appear in the logical netlist, so any check that compares netlist and layout has to account for them.9
Before placement: rows, a macro, and FIXED tap and endcap cells from floorplanning. Tap an input or output.
Placement software works in three passes, from rough to exact.
- Rough layout. Each connection acts like a rubber band, pulling cells together. Crowding pushes cells apart. The cells slide around until the pulls and pushes balance. Some still overlap a little.
- Snap into rows. Each cell moves to the nearest free slot in a row, so nothing overlaps. This step is called .
- Polish. The software tries small swaps between neighbors. It keeps the ones that make the wires shorter.
To keep score, the software adds up a quick guess of every wire’s length, and tries to make that total small.
Some signals have the least time to spare. The software gives their cells an extra strong pull, so they end up closer together.
Three passes
finds a rough position for every cell that is free to move. To measure how full each area is, the tool lays a coarse grid of square bins over the core and adds up the cell area in each bin. It then shrinks the total estimated wirelength (, explained below) while keeping the cell area in every bin under a limit called the target density.6 Modern global placers treat positions as continuous numbers rather than slots, and nudge all cells at once in whatever direction improves the score. In OpenROAD’s gpl, wires pull connected cells together while an pushes cells out of crowded bins, as if each cell were an electric charge. The placer stops when the cell area still sitting in over-full bins, the , falls below a target: 10% by default.2
then moves each cell to a free site in a row, removing every overlap and moving each cell as little as possible.3 polishes the result with local moves, such as swapping two cells, trying a different order for a few neighbors in a row, or flipping a cell left to right, and keeps only the changes that shorten wires.103
Scoring a placement: HPWL
A net is one electrical connection: the output pin of one cell joined to the input pins of one or more others. Half-perimeter wirelength scores a net by drawing the smallest rectangle around all its pins and adding the rectangle’s width and height, which is half its perimeter. Take a net with pins at , and µm. The box runs from to 5 and from to 7, so . Move the third cell to and the box shrinks to to 4, so HPWL drops to .
Why this number? Chip wires run only horizontally and vertically, so a two-pin net needs at least the width plus the height of its box, and HPWL is exactly that. A net with more pins still has to reach across the full width and the full height of its box, so its real wire is at least as long as HPWL, and usually longer. HPWL is an optimistic estimate, but it is quick to compute and it rises and falls with timing and routability, which is why placers use it as their main score.6
Why rough first, then snap
Choosing a slot for each of a million cells directly is an enormous search: every cell has thousands of possible slots. Treating positions as continuous numbers lets the placer move all cells at once, a little at a time, in the direction that improves the score, which works even for millions of cells. Snapping to the grid afterwards costs a little wirelength, and that cost stays small only if global placement left enough free space near every cell.
What else it watches
- Timing. The clock ticks at a fixed rate, and every timing path (from one flip-flop, through gates and wires, to the next flip-flop) must deliver its signal before the next tick. The time to spare is the path’s ; negative slack means the path is too slow. pauses now and then to compute slack, then gives the nets on the worst paths a higher weight, so shortening them counts for more. In OpenROAD, the worst 10% of nets get weights of up to 5 by default.2
- Room for wires. The tool estimates congestion inside the loop. OpenROAD uses , which spreads each net’s expected wire evenly over its bounding box and adds up all the nets tile by tile. Cells in congested tiles are temporarily treated as bigger, or “inflated,” so they push their neighbors away and open up room for wires.211
- Density. The target density caps how full any local area gets. Overall , the share of the core covered by cells, is kept well under 100%, because resizing gates, adding buffers and the small moves of detailed placement all need free space in every part of the chip.12
- Power. Each time a wire switches between 0 and 1 it charges or drains its capacitance, which grows with its length. Placers can give wires that switch often a higher priority for being short.1
Fixing the netlist while placing
Once cells have positions, the tool can estimate each wire’s length, resistance and capacitance, run a real timing analysis and repair what it finds. OpenROAD’s resizer shows the usual moves:5
- Buffering. A copies its input with fresh strength. Breaking up a long wire, or a wire that feeds many inputs, with buffers keeps signal edges sharp and delays low.
- Sizing. Each gate comes in several strengths (X1, X2, X4…). A bigger version drives its wire faster but takes more area and power.
- Threshold-voltage swap. Libraries offer each gate built with faster but leakier transistors (a low ) or slower, less leaky ones. The tool uses the fast ones only where a path needs them.
- Pin swapping and cloning. Moving the latest-arriving signal to a gate’s faster input, or duplicating a gate so each copy drives part of the load.
OpenROAD’s documentation recommends running its main timing repair after clock tree synthesis, once real clock delays are known.5 Logic restructuring can also rebuild a small group of gates to be faster or smaller.13 Every change is legalized again, because new or bigger cells need free sites.3
Reordering the scan chains
So the chip can be tested after manufacturing, test logic links all its flip-flops into long shift registers called . A tester shifts a pattern in through the chain, runs the chip for a clock tick, and shifts the result out.14 The chains were wired before anyone knew where the flip-flops would land, so neighbors in a chain may sit far apart. The chip’s normal operation doesn’t depend on which flip-flop follows which in the chain, so re-stitches each chain so consecutive flip-flops are close. Finding the shortest order is a version of the classic traveling-salesman problem.14
The objective, step by step
Written as an optimization: give every movable cell coordinates and minimize total HPWL, subject to every bin’s density staying at or below the target .6 Two things make that hard to solve directly. HPWL is built from max and min, so it has no smooth gradient; placers replace it with a smooth approximation (see Under the hood). And there may be a million density constraints, one per bin, so they are folded into a single penalty that is zero only when every bin is under target.6 The problem becomes:
The weight is the dial between the two goals. It starts small, so wirelength dominates and cells clump near the center of their connections, with overflow near 100%. The placer raises step by step, which spreads the cells out at the cost of longer wires, until overflow reaches the stopping target.4 That is why HPWL climbs during a healthy run: the cells are spreading.
HPWL is exact for two-pin nets and a lower bound for nets with more pins, since any tree joining the pins has to span the box in both directions. It ignores detours, which metal layer a wire uses, and the vias between layers. So routability-driven mode adds a congestion estimate on top. In OpenROAD gpl, once overflow drops to a check point (0.3 by default), each iteration runs , computes a routing-congestion (RC) metric from the most congested 0.5% and 1% of tiles, and inflates the cells in congested tiles. It aims for an RC of 1.01 by default and stops if RC has not decreased for three consecutive iterations.2 RePlAce introduced the same idea, with congestion from a real global router (NCTU-GR) driving the inflation.4
Timing-driven net weighting
In gpl’s mode the placer pauses at set overflow points, 64% and 20% by default, and does the following:2
- Run repair_design on the current rough positions, so badly overloaded nets get buffers before timing is measured.
- Run timing analysis on estimated wires to find the slack of every net.
- Sort nets by slack. Give the worst net the maximum weight (5 by default), scale down to 1.0 at the 10% point, and leave the rest at 1.
- Resume placement. The weight multiplies that net’s wirelength term, so the optimizer pulls its cells together harder.
By default the resizer’s changes from step 1 are kept rather than thrown away (keep_resize_below_overflow defaults to 1.0).2 Net weighting is coarse. One number per net can’t tell a net that is critical through one path from one that is critical through thousands, and the placer only sees timing at those few checkpoints.
Utilization and whitespace
Two different numbers describe fullness. Core is total cell area divided by available core area: an average. Target density is the cap for each bin, and gpl defaults to 0.7.2 A block can be 60% utilized overall and still have bins at the cap next to a macro. The free space, or whitespace, also has to be local: a buffer or an upsized gate belongs next to the net it fixes, not in a distant empty area.12 How much to leave is a project choice. A 2003 paper reported that whitespace in real designs varied from about 20% to about 70% with methodology and design time, and named guarding against routing congestion as a major reason for it.15 In practice a team picks a starting point, looks at the first congestion and timing results, and adjusts.
Physical synthesis moves
Optimization during placement uses the same moves as late timing fixes, but it is cheaper, because cells can still move to make room. OpenROAD’s resizer splits the work in two:5
- repair_design fixes electrical limits. It buffers nets that break the maximum slew (how slowly a signal may rise or fall), capacitance or fanout (how many inputs one output drives), buffers long wires to cut their RC delay, and resizes gates to even out slews.
- repair_timing fixes setup and hold violations and is meant to run after CTS. Its moves include removing buffers, swapping to a faster threshold voltage, upsizing gates, swapping pins, rebuffering, cloning gates and splitting loads. With -recover_power it then downsizes gates on paths with positive slack to save power.
Restructuring cuts out a cloud of logic and re-synthesizes it with the ABC logic optimizer for delay or area.13
Scan reordering
works only on the FLOATING lists of the DEF SCANCHAINS section; ORDERED segments stay intact.9 Classic reorderers treat it as a traveling-salesman problem and minimize the Manhattan distance between consecutive flip-flops. Gupta, Kahng and Mantik pointed out that this misjudges the cost: a scan connection only has to reach the net already leaving the previous flip-flop, not its output pin. Their routing-aware version uses that incremental connection as the cost and checks timing slack at every sink the change affects.14 After reordering, the test patterns must be regenerated for the new chain order.
Advanced-node effects
In older processes, any two cells that didn’t overlap could sit side by side: each cell was drawn to be correct next to any neighbor. Below about 10 nm that stops being true. The front-end-of-line (FEOL) layers, which define the transistors themselves, bring rules such as a minimum width for implant regions (the doped areas that set a transistor’s threshold voltage), drain-to-drain abutment and oxide-diffusion jogs, and two cells that are each legal can break them together. Padding every library cell so that all neighbors are safe costs too much area, so Han, Kahng and Lee argued for a final, rule-aware legalization pass after normal placement.16
Multi-row-height cells (large flip-flops, multiplexers, high-drive gates) span two or more rows. They must land where the power and ground rails line up with their own, which ties rows together during legalization and brings extra rules for edge spacing and for pins that short or become unreachable.17
is the other big one. Pins close to each other, or close to a cell’s edge, can block each other’s access points, so detailed routing starts with many rule violations (DRCs) that placement never saw. Two placements with the same cell and pin density can differ here, because what matters is how the pins line up with the routing tracks. Remedies include padding cells, pin-access-aware detailed placement, and small cell moves during routing.18
Start of global placement: With only wirelength counted, cells pile up near the middle of their connections. Most of their area is in over-full bins.
The software doesn’t get a blank floor. Some areas are marked “no parts here,” like the space right around a big memory block. That keeps room for the wires that reach it.
A few small helper cells are put down first, in a regular pattern, like the posts in a parking lot. Engineers also scatter some spare cells around. If a bug turns up late, they can fix it by changing only the wires, which is much cheaper than redoing the whole chip.
The placer never starts from an empty floor. Floorplanning has already marked areas that are off limits or reserved, and some special cells go down before any logic.
Blockages and halos
A is a rectangle where the placer may put no cells, or only some. DEF has three kinds:9
- Hard: no cells at all.
- Soft: kept empty during the main placement, but later steps, such as timing repair and clock tree synthesis, may put cells there. It saves space, for example in a narrow channel between two macros, for buffers added later.
- Partial: cells may fill only a set percentage of the area, which thins out a crowded spot.
A is a no-cells border attached to a macro. It moves with the macro and keeps standard cells from crowding right up against it.9
Regions
A ties a group of cells to an area. A guide region is a preference that wirelength or timing can override. A fence region is strict: its cells must stay inside, and no other cells may enter.9
Cells placed before the logic
- . Transistors sit in regions of silicon called wells, which must be tied firmly to power or ground, and a tap cell makes that connection.8 Without enough taps, a stray path through the silicon can switch on and short power to ground, a failure called latch-up, so foundries set a design rule on the spacing between a transistor and the nearest tap.19 OpenROAD’s tapcell command places taps in a checkerboard at a set distance.20 Standard cells explains wells and taps in more detail.
- (boundary cells) sit at both ends of every row and along macro edges to finish off the rows.820
Taps and endcaps are FIXED before standard-cell placement starts, so the placer works around them. are added during placement: unconnected gates scattered across the chip. If a bug turns up late, an engineering change order () can wire spares into the design by changing only the metal wiring layers.21
DEF encodes the constraint vocabulary directly. BLOCKAGES - PLACEMENT is hard. + SOFT is honored only in initial placement. + PARTIAL maxDensity limits standard-cell area in the rectangle, and later clock buffers and timing buffers ignore it. Component halos use + HALO [SOFT] left bottom right top. REGIONS take TYPE FENCE or GUIDE.9 A partial blockage is the usual way to thin out a locally congested area, such as a macro corner or a narrow channel, while keeping room there for buffers added later.9 Cell padding (set_placement_padding, or global_placement -pad_left/-pad_right) is the per-cell version: it reserves empty sites beside chosen cells, usually pin-dense or high-fanout ones, to leave room for routing.32
pitch comes from the foundry’s latch-up rule on the spacing between any active area and the nearest tap.19 tapcell -distance sets the pitch, and -halo_width_x/-y set the margin around macros when rows are cut.20 Taps and are LEF CLASS CORE WELLTAP and CLASS ENDCAP, and they appear in DEF with SOURCE DIST.89 Because they are FIXED at a regular pitch, they chop rows into segments, and legalization has to treat each segment separately. are spread across the design during placement so that a metal-only ECO always has gates within reach. A patch built from spares that are too far away has long wires, which can cause timing violations and routing congestion of their own.21
Before placement: macros, blockages, the fence, and the FIXED taps and endcaps. Tap a part to read about it.
Below is a tiny chip with 18 cells. Lines join the cells that connect. Drag cells around and watch the “Total wire length” number. Try lining up each colored group between its pins.
Then press Run annealer to let the computer try thousands of random swaps. At first it even keeps some bad swaps, to get out of dead ends. Later it gets pickier.
The grid is 12 × 8 sites, with 18 cells and 6 fixed input/output pins; the readout is total HPWL in µm. Drag cells to beat the starting HPWL by hand. Then press Run annealer to watch : it makes random moves and swaps, always keeps improvements, and keeps a worse move with probability , where is a “temperature” that drops over time. The chart plots HPWL and . Early on, while is high, HPWL bounces around; as drops it settles. Toggle the congestion map (a estimate) and notice that the shortest-wire arrangement can still have a hot spot where several connections’ boxes overlap. Press Scatter for a new random start.
Drag any of the 18 cells on the 12 × 8 site grid and total HPWL updates live. Toggle a RUDY congestion map, or press Run annealer to watch simulated annealing work, with HPWL and temperature plotted per cooling epoch. The annealer is a toy simulated-annealing placer: Metropolis acceptance (keep a worse move with probability ) and geometric cooling,22 plus a move window that shrinks with temperature so late moves stay local. Watch the accept-rate readout fall as drops; if cooling is too fast, the run freezes in a local minimum. Run it from several scatters and compare final HPWL. The spread is run-to-run variance, and placement stability matters in a timing-closure loop.12 Overlay the RUDY map after annealing: making wires very short can push local routing demand past local supply,1 which is what routability-driven cell inflation is meant to break up.2
After a placement run, an engineer first checks that every cell has a proper spot. Then they check how full the floor is. The cells should fill only about 70 percent, which leaves room for wires.
Next they look at a map, a picture of where too many wires want to pass. It looks like a weather map. A few warm spots are normal. A bright red blob means the wiring will get stuck there.
Below is an illustrative placement script in the style of OpenROAD. A script is simply the list of commands the tool runs, top to bottom. This one loads the floorplanned design, runs global placement, fixes electrical problems using estimated wires, legalizes and polishes, and checks timing. You don’t need to follow every line; the numbered notes explain the important ones. Real flows, such as OpenROAD-flow-scripts, split these steps across several scripts and tune the settings for each design.
Below are an illustrative OpenROAD-style script, the DEF it produces, and an annotated log. The log is modeled on OpenROAD output but is not verbatim. Read the log for three things: the overflow curve (HPWL should rise as cells spread, then flatten), the cost of legalization (displacement and the change in HPWL), and where congestion overflow sits relative to macros and fences.
# place.tcl: illustrative OpenROAD-style placement step, not a complete flow
read_db results/2_floorplan.odb ;# rows, macros, IO pins, PDN, tap and endcap cells
read_liberty lib/stdcells_typ.lib
read_sdc constraints/top.sdc
set_wire_rc -signal -layer M3
set_wire_rc -clock -layer M5
# 1. Global placement: analytical, timing- and routability-driven
global_placement -timing_driven -routability_driven \
-density 0.65 -overflow 0.10 -pad_left 1 -pad_right 1
# 2. Fix electrical rules using placement-based wire estimates
estimate_parasitics -placement
repair_design -max_wire_length 400
report_design_area
# 3. Legalize, then refine
set_placement_padding -global -left 1 -right 1
detailed_placement
improve_placement
optimize_mirroring
check_placement -verbose
# 4. Check timing before clock tree synthesis
estimate_parasitics -placement
report_worst_slack -max
report_tns
write_db results/3_place.odb
write_def results/3_place.def- 1L2Load the floorplanned design. It already holds the rows, macros, power wiring and the FIXED tap and endcap cells, so the placer only moves standard cells.
- 2L5Resistance and capacitance per micron of wire, taken from metal layer M3. Until real wires exist, every delay estimate depends on these values.
- 3L9Timing-driven: give the nets on the slowest paths more weight. Routability-driven: inflate cells in tiles that RUDY flags as congested.
- 4L10No bin may be more than 65% full; stop when only 10% of cell area is still in over-full bins. One empty site on each side of every cell leaves room to reach its pins.
- 5L13Estimate every wire’s resistance and capacitance from the placed positions.
- 6L14Add buffers on long wires and fix limits on signal slope, load and fanout. New buffers need free space nearby, which is why utilization matters.
- 7L19Legalization: snap every cell to a site in a row, moving each as little as possible.
- 8L20Detailed placement: local moves that cut wirelength without breaking legality.
- 9L21Flip cells left to right where that shortens wires.
- 10L22Fails on overlaps, cells off the site grid, cells in blockages or outside their fences.
- 11L26Worst setup slack: the time to spare on the slowest path, using estimated wires. The clock is still ideal (it reaches every flip-flop at the same instant), so a small negative number is often fixed later, in CTS and after routing.
- 12L27Total negative slack: the shortfalls of all failing paths added up. It shows whether one path is failing or thousands.
Here is part of the output DEF. Each line under COMPONENTS is one cell: its instance name, its library cell, its status and its location, followed by its orientation. Locations are in database units, 1,000 per micron here, so ( 48200 10000 ) means , . N means the cell sits upright; FS means it is mirrored top to bottom, which is how every other row is placed.
VERSION 5.8 ;
DESIGN top ;
UNITS DISTANCE MICRONS 1000 ;
DIEAREA ( 0 0 ) ( 200000 200000 ) ;
ROW ROW_0 CoreSite 10000 10000 N DO 900 BY 1 STEP 200 0 ;
ROW ROW_1 CoreSite 10000 12000 FS DO 900 BY 1 STEP 200 0 ;
# ... 88 more rows ...
REGIONS 1 ;
- alu_fence ( 110000 40000 ) ( 180000 90000 ) + TYPE FENCE ;
END REGIONS
COMPONENTS 7 ;
- u_sram SRAM_1KX32 + FIXED ( 20000 120000 ) N + HALO 5000 5000 5000 5000 ;
- ENDCAP_0 ENDCAP_X1 + SOURCE DIST + FIXED ( 10000 10000 ) N ;
- TAP_0 TAPCELL_X1 + SOURCE DIST + FIXED ( 30000 10000 ) N ;
- U1023 NAND2_X2 + PLACED ( 48200 10000 ) N ;
- U1024 INV_X1 + PLACED ( 48800 12000 ) FS ;
- u_alu/U88 DFF_X1 + REGION alu_fence + PLACED ( 130000 60000 ) FS ;
- spare_17 NOR2_X1 + SOURCE USER + PLACED ( 90000 34000 ) N ;
END COMPONENTS
BLOCKAGES 3 ;
- PLACEMENT RECT ( 10000 100000 ) ( 18000 190000 ) ;
- PLACEMENT + SOFT RECT ( 60000 112000 ) ( 66000 160000 ) ;
- PLACEMENT + PARTIAL 40.0 RECT ( 140000 140000 ) ( 170000 170000 ) ;
END BLOCKAGES- 1L5A row of 900 sites, each 0.2 µm wide, starting at (10, 10) µm. The row is one cell tall.
- 2L6The next row is 2 µm higher and flipped (FS), so neighboring rows share a power or ground rail.
- 3L9A fence: cells assigned to alu_fence must stay inside, and no other cells may enter.
- 4L12A FIXED macro (a memory) with a 5 µm halo on every side. The placer treats the halo as a blockage that moves with the macro.
- 5L13SOURCE DIST marks physical-only cells (taps, endcaps, fillers) that are not in the logical netlist.
- 6L14A well tap, FIXED before placement. The next tap in this row is a set distance to the right.
- 7L15An ordinary standard cell. PLACED means later tools may still move it, for example during CTS or an ECO.
- 8L17This flip-flop belongs to the fence region. It sits in a flipped row, hence FS.
- 9L18A spare gate, unconnected, waiting for a possible metal-only ECO.
- 10L21Hard blockage: no standard cells at all.
- 11L22Soft blockage in a channel between macros: empty during initial placement, available to buffers and clock cells later.
- 12L23Partial blockage: standard cells may fill at most 40% of this area, which thins out a congestion hot spot.
The log below follows one run. Overflow, the share of cell area still in over-full bins, starts near 1 because the cells begin piled together, and the run stops at 0.1. HPWL rises while the cells spread, which is expected. In this run, legalization adds 1.8% to HPWL.
If you read only three things, read the final overflow (did global placement finish spreading?), the legalization displacement and HPWL change (did snapping undo much of the work?), and the congestion table (will the wires fit?). A hot spot reported next to a fence or a macro usually points back to a floorplan or constraint choice.
[gpl] Target density 0.650, placeable area 15,220.4 um^2, core util 58.6%
[gpl] Instances 7,180 (movable 6,872, fixed 308), nets 7,140
[gpl] Initial placement done, HPWL 18,060 um
[gpl] Iter Overflow HPWL(um)
[gpl] 1 0.974 17,980
[gpl] 100 0.712 59,930
[gpl] Timing-driven (overflow 0.64): reweighted worst 10% of nets, max weight 5
[gpl] 250 0.301 95,410
[gpl] Routability: RUDY RC metric 1.18 > target 1.01, inflating cells in 41 tiles (+3.1% area)
[gpl] Timing-driven (overflow 0.20): reweighted worst 10% of nets, max weight 5
[gpl] 420 0.198 102,540
[gpl] 515 0.099 104,490
[gpl] Converged: overflow 0.099 at iteration 515
[rsz] Inserted 62 buffers on 43 nets, resized 163 instances
[dpl] Legalized 6,934 instances
[dpl] Displacement: average 0.96 um, max 14.2 um (U8817)
[dpl] HPWL 104,490 -> 106,370 um (+1.8%)
[dpl] check_placement: 0 overlaps, 0 off-site, 0 fence violations
[grt] Estimated congestion (GCell = 15 M3 pitches)
Layer Usage Max H/V overflow Total overflow
M2 60.9% 0 / 0 0
M3 75.7% 2 / 0 4
M4 71.2% 0 / 3 3
M5 48.3% 0 / 0 0
[grt] Hotspot: 3 GCells over capacity near (142.0, 61.5) um, inside alu_fence- 1L1Average utilization is 58.6%, and the placer pushes every bin toward the 65% target. A local area can be fuller than the average. The final overflow (0.099) measures how much cell area is still above the target.
- 2L5Overflow near 1: the initial placement puts cells on top of each other near the middle of their connections.
- 3L7At an overflow checkpoint (0.64, then 0.20), timing analysis on estimated wires raises the weight of the most critical nets.
- 4L8HPWL rises as the density penalty spreads cells. That is normal. A placer whose HPWL stays flat here is not spreading.
- 5L9RUDY found tiles over the target congestion metric. Cells there are inflated so neighbors are pushed out and wire room opens up.
- 6L13Stop at the overflow target (0.1). Lower targets spread more evenly but take longer and raise HPWL.
- 7L14Physical synthesis changed the netlist: 62 new buffers must now be legalized too (6,872 + 62 = 6,934).
- 8L16An average move under a micron is healthy. A large maximum usually means a cell was pushed out of a full fence or around a blocked area.
- 9L17The cost of legalization. A much bigger jump suggests global placement ended too dense.
- 10L19A trial global route. The chip is divided into routing tiles (GCells), each 15 wiring tracks of M3 wide.
- 11L22The worst M3 edge between two GCells is 2 tracks over capacity; 4 tracks of overflow in total. The router may detour around it, at a cost in timing or rule violations.
- 12L25The hot spot is inside the fence. The fence is too small for its logic: enlarge it, add a partial blockage, or lower the density there.
Converged: Overflow fell from 0.97 to 0.099, under the 0.1 target. HPWL rose as cells spread, then flattened: a healthy run.
- Too crowded. If cells are packed too tightly, the wires can’t all fit. Give the cells more room.
- Too spread out. If cells are far apart, signals travel too far and the chip is slower. It also costs more, because it’s bigger.
- Jams near big blocks. Lots of wires head for the edges of memory blocks. Without a clear border around them, the wiring gets stuck.
- Problems found late. A placement can look fine and still fail when the wires are drawn. A quick practice run of the wiring catches this early.
- Congestion hot spots. Packing many cells with many pins together, or squeezing cells into a narrow gap between two macros, can demand more wires than the area holds. A trial global route after placement shows the hot spots in a congestion report. Fixes include partial blockages, padding (empty sites kept beside chosen cells), halos, or a lower target density.2393
- Timing that looked fine before placement. Synthesis, the step that turned the design’s code into gates, could only guess wire lengths. Real distances between placed cells can make a path that passed synthesis fail. The cures are timing-driven placement and checking slack with placement-based wire estimates before moving on to clock tree synthesis.25
- No room to fix things. When cells fill nearly all the space, there is nowhere to put buffers or bigger gates, and legalization has to push cells far from where global placement wanted them.12
- Large legalization moves. If global placement ends too dense, legalization has to move cells a long way to find free sites, which undoes the wirelength and timing work. The detailed placement report lists how far cells moved.3
- Over-tight fences. A fence must hold all its cells and nothing else.9 One that is too small for its logic packs that logic to the limit and creates a congestion island.
- A bad floorplan. Placement can’t rescue a floorplan that puts two heavily connected macros on opposite sides of the chip. The fix is to go back to floorplanning.
- Optimistic HPWL, pessimistic routing. HPWL ignores detours and layer limits, so a placement can win on HPWL and lose on routed rule violations. Compare a global route’s overflow per layer, not just RUDY.2311
- Pin-access trouble. Cell- and pin-density models can miss pin-access problems. Two placements with the same cell and pin density can differ in how pins line up with routing tracks, and the bad one starts detailed routing with many DRCs.18 Pad hard-to-reach cells, or run a pin-access-aware detailed placement pass.
- Net-weight overshoot. Aggressive timing weights pull critical clusters tight. That can create density peaks and push non-critical nets onto longer routes, so total negative slack (the sum over all failing paths) gets worse even as worst negative slack improves. Tune the weight cap and the share of nets reweighted (OpenROAD exposes both) for each design.2
- Legalization damage. A large maximum displacement usually means a full fence, rows fragmented by taps and endcaps, or multi-height cells with too few matching rows. Check the maximum and the distribution of displacement as well as the average.317
- Advanced-node adjacency violations. Drain-drain abutment, implant width and oxide-diffusion jog rules are too complex for the normal placement flow to fully consider, so they need a dedicated legalization pass at the end.16
- Instability. Some placers, notably min-cut and annealing ones, produce very different placements from run to run. Fixes aimed at one placement may not apply to the next, which breaks run-to-run comparison and slows timing-closure loops.12
- Scan reorder surprises. Reordering changes which flip-flop drives which scan input, so it changes scan-path timing and the test patterns. Keep ORDERED segments and chain partitions intact.914 Check hold timing on the new, often very short scan connections (a signal that arrives too soon after the clock edge can corrupt the capture), and rerun ATPG on the reordered chains.
Balanced: no tile over capacity, 30% free space left for buffers and resizing.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
The problem
Model the netlist as a hypergraph : each cell is a vertex, and each net is a hyperedge, an edge that can join any number of vertices. Choose for every movable cell to minimize total HPWL, where for each net over its pins. The constraints: every cell sits in a row, on a site, overlapping nothing.6 Global placement is NP-complete.6 Even the simpler job of cutting a netlist into two balanced halves with the fewest nets crossing (min-cut bipartitioning) has no known polynomial-time exact algorithm, so practical tools rely on heuristics.24 The standard decomposition is global placement on a relaxed problem (overlaps allowed, every bin at or under ), then legalization, then detailed placement.6
Historically there are four families of placer: stochastic (simulated annealing), min-cut partitioning, quadratic, and nonlinear analytical.6
Simulated annealing, and why it lost at scale
borrows from metalworking: cool slowly and atoms settle into a low-energy order. Each step moves a randomly chosen cell or swaps two cells. A move that lowers the cost is always kept; one that raises it by is kept with probability , so uphill moves become rarer as the temperature falls. is lowered geometrically, with .22 TimberWolf, from the 1980s, is the representative annealing placer.6 Annealing gives good quality, but it needs a huge number of moves and converges slowly, and that poor scalability pushed it out of global placement once designs reached millions of cells.61
Min-cut partitioning
Top-down placers cut the netlist and the die in half, again and again: each half of the cells goes to one half of the region, choosing the split that cuts the fewest nets. Fiduccia–Mattheyses (FM) improves a split by moving one cell at a time to the other side while keeping the halves balanced in area; careful data structures make each pass linear in netlist size.24 Multilevel partitioning (hMETIS) first merges cells into ever coarser clusters, splits the smallest version, then expands it back level by level, refining the split at each level. Its cuts were on average 6–23% better than earlier tools, and it often ran 4–10 times faster.25 Min-cut placers such as Capo were among the best in the 2000s. But a bad early cut can’t be undone later, and keeping enough whitespace in every region as the cuts get smaller is hard.615
Quadratic and force-directed placement
Treat each connection as a spring whose energy is the squared distance between its ends. Total energy is then a quadratic function of the cell coordinates, and its minimum, where all spring forces balance, is found by solving a large sparse system of linear equations, typically with the conjugate gradient (CG) method. Minimizing wirelength alone piles many cells into dense regions, typically around the center, so force-directed placers add pseudo-pins and pseudo-nets that pull cells out of crowded areas.26 SimPL keeps two placements. The lower-bound one is the quadratic solution, using the Bound2Bound net model (each pin of a multi-pin net is linked to the net’s two extreme pins). The upper-bound one comes from a fast look-ahead legalization of it. Upper-bound positions become anchors that pull the next lower-bound solve, and the two converge.26 ComPLx generalizes this as a primal-dual Lagrange optimization.27 Quadratic placers are fast, but their quality and robustness usually trail nonlinear placers.6
Smooth wirelength: log-sum-exp and weighted-average
A nonlinear placer needs a wirelength with a gradient, but a net’s -span is , and max and min have kinks. Both popular fixes replace max with a “soft” max controlled by a smoothing parameter . Log-sum-exp (LSE) approximates the -span as
with error at most for an -pin net. The weighted-average (WA) model takes a weighted mean of the pin coordinates that favors the largest ones (weights ) minus one that favors the smallest (weights ); its error bound is roughly half that of LSE. Smaller is more accurate but less smooth, and it can’t be made arbitrarily small because of floating-point limits.6 ePlace, RePlAce and DREAMPlace all use WA.647
γ = 8 µm, 4 pins: true span 76.0 µm; LSE 76.6 (+0.6, bound γ ln n = 11.1); WA 74.1 (-1.9).
Electrostatic placement: ePlace, RePlAce, DREAMPlace
ePlace’s model works in four steps on each iteration:6
- Treat each cell as a positive charge equal to its area, and spread the charges onto a grid of bins to get a charge density.
- Solve Poisson’s equation for the electric potential of that density. On a regular grid this can be done with fast Fourier transforms in , with boundary conditions that keep cells inside the region.
- The density penalty is the system’s potential energy, and its gradient on each cell is the electric force , where is the local field: cells in crowded bins are pushed toward emptier ones.
- Take a step on the whole objective, , for all cells at once. The netlist is placed flat, without clustering cells first.
For the step, ePlace uses Nesterov’s accelerated gradient method instead of nonlinear conjugate gradient with line search. Nesterov’s method adds momentum: each step starts from a point pushed a little further along the previous step. Its step length is the inverse of a Lipschitz constant (a bound on how fast the gradient can change) predicted from recent iterations, which avoids the costly line search and made placement more than 2× faster than CG. A preconditioner approximates the Hessian so that big and small cells move at sensible rates.6 RePlAce adds a per-bin density multiplier that grows exponentially with that bin’s overflow, plus dynamic step-size control. It reports 2.00% lower HPWL than the best published results on the ISPD 2005 and 2006 benchmarks, and it extends to routability through congestion-driven cell inflation.4 OpenROAD’s gpl is built on RePlAce.2
DREAMPlace recasts the same problem as neural-network training. Cell positions play the role of the network’s weights, each net plays a training sample, the WA wirelength plays the prediction error, and the density cost plays the regularization term. A forward pass computes the objective and a backward pass computes the gradients. Built on PyTorch with custom GPU kernels for wirelength and density, it reports over 30× speedup in global placement against multi-threaded RePlAce with no quality loss, about one minute for a million-cell design, and nearly linear scaling to 10 million cells.7
Legalization: Tetris and Abacus
Tetris-style legalization sorts cells by and legalizes them one at a time, greedily dropping each into the nearest free position. Abacus also sorts by and legalizes one cell at a time, trying it in nearby rows. For each trial row, its PlaceRow step shifts the cells already legalized in that row to minimize their total movement, and it keeps the row with the lowest cost. That gives lower total displacement than Tetris.28 OpenROAD’s dpl started from OpenDP’s diamond search and now defaults to a negotiation-based legalizer. It handles 1×–4× multi-height cells, fence regions and fragmented rows.3 Mixed-cell-height legalizers must also match each multi-row cell to rows with the right power and ground rails; one approach uses window-based insertion plus network-flow clean-up for displacement, edge spacing and pin access.17 At 10 nm and below, a final pass based on mixed integer-linear programming can fix adjacency rules (drain-drain abutment, implant width, oxide-diffusion jogs) in windows that are solved independently. One such pass fixed 99% of violations on an abstracted 7 nm library.16
Detailed placement moves
FastPlace’s detailed placer shows the standard move set. Global swap finds each cell’s optimal region (where its HPWL would be smallest with everything else held still) and swaps it with a cell or an empty space in that region if that helps most. Vertical swap moves a cell up or down a row toward its optimal region. Local reordering tries every order of three consecutive cells in a segment. Single-segment clustering shifts cells within a row segment while keeping their order. The passes repeat until improvement stalls.10 OpenROAD adds improve_placement, and optimize_mirroring, which flips cells about their vertical axis where that lowers HPWL.3 Pin-access-aware refinement makes small local cell moves during routing to clear access conflicts that placement models miss.18
Q1A connection joins three pins at (1, 1), (4, 3) and (2, 6) µm. What is its HPWL: the width plus the height of the smallest box around the pins?
Q2Global placement leaves cells slightly overlapping and between slots. What does legalization do next?
Q3An area is marked with a soft placement blockage. What does that mean?
Q4For testing, a chip’s flip-flops are linked into long chains. Why can the tool change the order of flip-flops in a chain after placement?
Sources
Show Hide 28 sources
- Placement (electronic design automation)Placement goals (timing, wirelength, congestion, power), most sub-problems at least NP-hard, short wires can exceed local routing supply, high-activity nets kept short, early algorithm families.
- Global Placement (gpl)RePlAce-based Nesterov placer; -density default 0.7, -overflow default 0.1; timing-driven reweighting at overflow 64% and 20%, max weight 5, top 10% of nets; RUDY-based cell inflation from overflow 0.3, target RC 1.01, stop after three non-improving iterations.
- Detailed Placement (dpl)Legalization to sites/rows minimizing displacement; OpenDP origin; mixed-cell-height and fence support; padding, fillers, improve_placement.
- RePlAce: Advancing Solution Quality and Routability Validation in Global PlacementPlacement quality drives timing closure, utilization and routability; W + λD objective with λ raised over time; per-bin local density penalty, dynamic step size; NCTU-GR-driven cell inflation; 2.00% HPWL gain on ISPD 2005/2006.
- Gate Resizer (rsz)repair_design buffering and sizing, repair_timing with sizing, pin swap, cloning, buffering and VT swap; estimate_parasitics -placement.
- ePlace: Electrostatics-Based Placement Using Fast Fourier Transform and Nesterov’s MethodProblem formulation, HPWL correlating with timing and routability, NP-completeness, LSE and WA models, eDensity via Poisson/FFT, Nesterov’s method, survey of four placer families.
- DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI PlacementPlacement cast as neural-network training in PyTorch on GPUs (wirelength as prediction error, density as regularization); over 30× faster than multi-threaded RePlAce; CPU partition-based parallelism saturates near 5×.
- LEF/DEF 5.7 Language Reference (LEF syntax chapters)SITE definitions and MACRO classes including WELLTAP and ENDCAP.
- LEF/DEF 5.7 Language Reference (DEF syntax chapters)ROWS, COMPONENTS (PLACED/FIXED, HALO, REGION, SOURCE DIST), BLOCKAGES (SOFT, PARTIAL), REGIONS (FENCE/GUIDE), SCANCHAINS.
- FastPlace 2.0: An Efficient Analytical Placer for Mixed-Mode Designs (slides)Detailed placement moves: global swap, vertical swap, local re-ordering, single-segment clustering.
- On Robustness and Generalization of ML-Based Congestion Predictors to Valid and Imperceptible PerturbationsDefines RUDY: each net’s wire volume spread uniformly over its bounding box, used as a congestion indicator.
- On Whitespace and Stability in Mixed-Size Placement and Physical SynthesisGate sizing, buffering and detailed placement need local whitespace in every region of the die.
- Restructure (rmp)Logic restructuring for area or delay by handing logic cones to ABC.
- A Proposal for Routing-Based Timing-Driven Scan Chain OrderingScan chain ordering from placement data, cast as a traveling-salesman problem; affects routability, wirelength and timing.
- Hierarchical Whitespace Allocation in Top-Down PlacementReports (2003, 180–130 nm era) that whitespace varied from about 20% to about 70% with methodology and design time; guardbanding for routing congestion a major contributor.
- Scalable Detailed Placement Legalization for Complex Sub-14nm ConstraintsDrain-drain abutment, minimum implant width and OD jog rules break correct-by-construction placement at 10 nm and below.
- Pin-Accessible Legalization for Mixed-Cell-Height CircuitsMulti-row-height cells, power/ground rail alignment, edge spacing and pin access in legalization.
- In-Route Pin Access-Driven Placement Refinement for Improved Detailed Routing ConvergenceNeighboring and cell-boundary pins degrade pin access and cause routing DRCs at advanced nodes.
- Latch-upLatch-up as a low-impedance path between supply rails; substrate taps lower the resistance that triggers it; fabs set design rules on spacing from active area to the nearest tap.
- Tapcell (tap)Tap cells in a checkerboard at a set distance; endcap and boundary cells at row ends and around macros; halo around macros when rows are cut.
- Resource-Aware Functional ECO Patch GenerationSpare cells are spread over a design during placement so later metal-only ECOs can rewire them.
- VLSI Physical Design: From Graph Partitioning to Timing Closure, Chapter 4 slides (Global and Detailed Placement)Simulated-annealing placement: random cell moves and swaps, Metropolis acceptance exp(−Δcost/T), geometric cooling T = α·T with 0 < α < 1; very slow.
- Global Routing (grt)GCells, capacity, overflow, congestion reports viewable in the GUI.
- Hypergraph Partitioning and ClusteringBalanced hypergraph partitioning is NP-hard, so heuristics are used; the Fiduccia–Mattheyses (FM) algorithm moves every vertex once per pass under balance constraints, with gain buckets giving linear time per pass.
- Multilevel Hypergraph Partitioning: Applications in VLSI DomainMultilevel coarsen, bisect, project and refine; the algorithm behind hMETIS.
- SimPL: An Effective Placement AlgorithmForce-directed quadratic placement; lower-bound and upper-bound placements with look-ahead legalization and anchors.
- ComPLx: A Competitive Primal-dual Lagrange Optimization for Global PlacementGeneralizes SimPL as a primal-dual Lagrange optimization.
- Abacus: Fast Legalization of Standard Cell Circuits with Minimal Movement (slides)Tetris-style greedy legalization vs. Abacus, which re-places already-legal cells in a row to minimize total movement.