A chip is built from millions of tiny parts, and before any of them are put in place, someone has to plan the space. That plan is the . It sets how big the chip is and where the big pieces go.
The big pieces are called . Most of them are blocks of memory. Like a kitchen, each one is large and hooked up to a lot of other things. Put them in the wrong spot and wires get long, the chip gets slower, and some areas get too crowded to wire up.
This is the first step where real distances matter. Every later step builds on the plan, so a mistake here spreads everywhere.
A chip starts as a description of what its logic should do. Earlier steps turn that description into a : a list of millions of small, standard building blocks and the wires between them. The building blocks are , each a tiny circuit with one simple job, such as an AND gate or a (a cell that stores one bit and updates it on each tick of the chip’s clock). Before a tool can place those cells on the silicon, someone has to plan the space.
That plan is the . It fixes the chip’s size and shape, the rows the small cells will sit in, where the connections to the outside world are, and where the go. Macros are large blocks whose layout is already finished, most often memories. The floorplan may also lay out the start of the power wiring.1
In the flow, floorplanning comes after logic synthesis (which produces the netlist) and design for test (which adds logic for testing chips at the factory), and before power planning and placement. It is the first step that works with real distances, measured in micrometers (µm, thousandths of a millimeter).
It matters out of proportion to the time it takes. Once the macros are in place they are rarely moved in later steps, so a poor floorplan can leave speed on the table or make the design hard to finish.16 Every later step, from placing cells to building the clock network to drawing the wires, works inside the space the floorplan leaves.
Every chip layout tool has floorplanning commands: commercial ones such as Cadence Innovus and Synopsys IC Compiler II, and the open-source OpenROAD project, which has separate commands for the outline and rows, pin placement, macro placement and the small “tap” cells that rows need.2467 This chapter uses OpenROAD for its examples because its documentation is public.
You know the design arrives here as a gate-level netlist that now has to be laid out. Floorplanning is the first layout step. It turns the netlist, plus the physical descriptions of the cells and macros, into decisions that are locked before anything else is placed:
- the die and core outline, and the rows that standard cells will sit in;
- where the I/O pins or package bumps go;
- each macro’s position and orientation, with a keep-out margin (halo) around it;
- blockages that keep cells out of, or thin them in, chosen areas;
- regions for power domains that run at their own voltage or switch off;
- in designs built as separate blocks, each block’s outline, pins and timing budget.
Each decision bounds what placement, clock-tree synthesis (CTS, which builds the network that delivers the clock) and routing can achieve later. Macro placement in particular is rarely revisited once fixed, and it is still mostly done by hand: a back-end designer works out the design’s structure and dataflow, often by talking to the RTL designers, and places the memories to match.16 Some physical-synthesis flows, such as Synopsys Fusion Compiler and Cadence Genus iSpatial, even expect a floorplan DEF with macro and I/O placement as an input to synthesis itself.17
The underlying packing problems are NP-hard, so tools rely on heuristics.1 The pressure on automation keeps growing: auto-generated RTL for machine-learning accelerators can put several hundred macros into one block, which breaks the traditional practice of lining macros up around the edge.17 Commercial tools ship automatic macro placers; Cadence’s, for example, places macros and standard cells together.24 OpenROAD’s rtl_macro_placer implements Hier-RTLMP, a published multilevel macro placer, so this chapter uses it as the worked model of what happens inside such a tool.6
The die is the whole rectangle of silicon. The core inside it holds the logic, with a margin around it.
Floorplanning starts with a parts list, the size and shape of each big block, and how fast the chip must run.
It ends with a map of the chip, plus a few quick checks.
| Direction | What | Format |
|---|---|---|
| In | The netlist: every cell and every connection, after synthesis and test insertion | Verilog |
| In | Manufacturing layers, placement sites and routing tracks; the outline, pins and blocked areas of every cell and macro | Tech LEF, cell and macro LEF |
| In | How fast each cell and macro is, and how much power it uses | Liberty (.lib) |
| In | Timing targets: clock speeds, timing of signals at the chip’s edge, exceptions | SDC |
| In | Power intent: which logic runs at which voltage or can switch off, and the cells needed between them | UPF (IEEE 1801) |
| In | Package and pin requirements: pad order, bump map, which pins go on which side | Tcl constraints, spreadsheets, or a parent floorplan |
| Out | Die and core area, rows, tracks, pins, locked macros with halos, tap and endcap cells, blockages, regions | DEF or a tool database (OpenDB .odb) |
| Out | For designs built as separate blocks: block outlines, block pins, per-block timing targets | DEF per block, per-block SDC |
| Out | Utilization, trial congestion and timing reports | Text reports and maps in the tool’s viewer |
The formats are plain text files with short names. Verilog is a language for describing hardware; after synthesis the Verilog file is just a long list of cells and the wires joining them. files describe the building blocks from the outside only: a macro’s LEF gives its size, where its pins are, and which layers of wiring it blocks, which is all a floorplanner needs to know about it. files give each cell’s delay and power. files give the timing targets, such as how fast the clock ticks.
, the IEEE 1801 standard, describes power intent: which groups of logic have their own supply, which can be switched off, and which special cells are needed where they meet.9 The main output is a file, which records the chip outline, the rows, each placed part with its position and the way it faces, the pins, the blocked areas and the regions.10
In the DEF, the field that downstream tools obey is each component’s placement status. FIXED means automatic tools may not move it (an engineer still can, by hand), PLACED means they may, and COVER marks parts of a cover macro that nothing may move.10 That is how a floorplan locks its macros while leaving standard cells free for the placer.
When the chip is built as separate blocks, planning starts top-down: each block receives from the top level a fixed outline, pin locations and its own constraints,16 and time budgeting carves the top-level timing into per-block targets: an arrival time at each block input and a required time at each block output.14 In open-source flows, OpenROAD-flow-scripts exposes the same choices as flow variables: either CORE_UTILIZATION, CORE_ASPECT_RATIO and CORE_MARGIN, or explicit DIE_AREA and CORE_AREA, plus MACRO_PLACEMENT_TCL for hand-placed macros, IO_CONSTRAINTS for pins and TAPCELL_TCL for tap and endcap cells. It can also start from a ready-made FLOORPLAN_DEF.3
Tap or hover an input to see what part of the floorplan it decides.
Floorplanning answers a few big questions, roughly in this order.
- How big? Engineers add up the space all the parts need, then leave extra. How full the floor ends up is called . Too full, and there’s no room for wires. Too empty, and the chip costs more than it should.
- Where do the big blocks go? Usually along the edges, near the things they talk to. That leaves one big open area in the middle for the small parts.
- How wide are the aisles? The gap between two big blocks is a . Wires to those blocks must squeeze through it, so it has to be wide enough.
Then the engineer tries a quick, rough version of the next step to look for crowded spots, and if it looks bad, they move blocks and try again. A few hours here can save weeks later.
Sizing the die and core
The is the whole rectangle of silicon. The core is the inner rectangle that holds the logic, with a margin around it. The core is filled with : horizontal strips of equal height, like the lines on ruled paper. Standard cells all share that height, so they can snap into a row side by side. Macros sit on top of the rows, and the rows are cut around them.
To size the core, add up the area of the standard cells (synthesis reports it), add the macros, and leave headroom. The headroom is set by a target : the share of the available area the cells will actually fill. The rest is space for wires and for the extra cells later steps add, such as buffers (small cells that strengthen a signal on a long wire). Definitions differ. The rule of thumb usually quoted, a target of around 60–70%, counts the macros as filled area too: (macros + cells) ÷ core.1 This chapter’s examples use cells ÷ the area the macros leave free instead, which gives a lower number for the same core. OpenROAD’s initialize_floorplan command sizes the core from three numbers: the utilization, the () and the margin between core and die. It computes , or takes explicit die and core rectangles instead, and then builds the rows.2
A worked example, done by hand:
- Synthesis reports 400,000 µm² of standard cells (0.4 mm²). At a 70% target, counted against the free area only, they need of free area.
- The design also has four (on-chip memory blocks) of 300 × 180 µm, a ROM (read-only memory) of 200 × 120 µm, and a (the circuit that generates the clock) of 100 × 100 µm. Each gets a 10 µm clear border on every side, called a . With halos they take 301,200 µm².
- So the core needs about 873,000 µm². As a square, that is about 934 µm on a side. Round up to 940 µm and add a 20 µm margin on each side: the die is 980 × 980 µm, just under 1 mm square.
- Check: the core is . Take away 301,200 for macros and halos and 582,400 µm² is free; the cells fill of it.
- Counted the other way, with macros as filled area, the same core is full: above the 60–70% rule of thumb. Always check which definition a target uses before comparing numbers.
- Core (illustrative)
- 940 × 940 µm
- Macros + halos (illustrative)
- 301,200 µm²
- Standard-cell utilization of free area (illustrative)
- 68.7%
- Utilization counting macros (illustrative)
- 79%
Chips are usually kept close to square.11 A long, thin core makes wires that run along its long side longer.
I/O, pads and pins
A chip talks to the outside world through its I/O (inputs and outputs). In a traditional package, thin bond wires join pads on the chip to the package, so the pads sit in a ring of I/O cells around the core: the . If the pads need more edge than the logic needs area, the die is : it has to be bigger than the logic requires, and room inside goes spare. If the logic is large and the pads are few, the die is core-limited, and filler cells close the gaps between pads.11
packaging works differently. Solder bumps go on the chip’s top surface and the chip is mounted face-down, so the bumps connect straight to the package. The short connections reduce inductance (a wire’s opposition to fast changes in current), which allows faster signals.12 OpenROAD’s pad module can build either an I/O ring with wirebond pads or a grid of bumps.5
A block inside a larger chip has pins instead of pads: points on its edge where a wire enters or leaves. chooses where each one goes. OpenROAD’s place_pins puts pins on the (the evenly spaced lanes that wires must follow on each metal layer), choosing spots that keep wires short. A separate command can restrict a group of pins to one edge or part of an edge. If the cells haven’t been placed yet, the tool treats them all as sitting at the center of the die.4 That is why many flows place pins once, run a rough trial placement, and place them again.
Placing macros
A ’s inside is finished, so the floorplan can only choose where it goes and which way it faces. Usually it can only be mirrored or turned 180°, not turned sideways, because the transistor gates inside must keep running in the same direction as everywhere else on the chip.17 Experts follow a few rules of thumb, and automatic macro placers now build them in:1617
- Push macros toward the edges. A macro is an obstacle for wiring, especially when the chip has few metal layers to route over it. Lining macros up along the edges keeps the center open for standard cells, and neat tiles of macros make the power wiring easier to build.17
- Follow the dataflow. Put each macro near the logic and the pins it exchanges data with. For example, an SRAM that buffers data from the external memory interface belongs on the same side as that interface’s pins.
- Face the pins toward the logic. A macro’s pins usually line one edge. Mirror it so that edge faces the logic it connects to, not the die edge or another macro.
- Avoid notches. Small, odd-shaped pockets of space left between macros are hard to use and tend to collect crowded wiring.
- Don’t block the pins. Keep macros from covering the paths from the chip’s pins into the core.
Analog blocks and the PLL add a noise concern. Millions of digital cells switching at once send electrical noise through the shared silicon underneath (the substrate). Distance from noisy logic, substrate contacts, and guard rings (rings of contacts around a sensitive block that soak up the noise) are standard defenses.18 So the PLL gets some distance from busy logic; in the example later in this chapter it sits in a corner next to the pin that brings in its reference clock.
Channels, halos and blockages
Every between macros has to carry the wires to their pins, plus any power straps (thick wires carrying the supply) and buffers, so it has to be wide enough for what passes through it.15 To keep cells from crowding in next to the macros, planners add placement blockages around them; OpenROAD’s block_macro_channels command creates soft blockages from each macro’s halo.6
A is a keep-out border that travels with a macro. In DEF it is part of the macro’s entry. A SOFT halo is obeyed only during the first placement, so a small channel stays free then but can take buffers added later.10 A does the same for any rectangle you choose. OpenROAD’s macro placer takes a minimum channel width. Its tapcell command cuts the rows around each macro’s halo and adds endcap cells where rows end, plus at a set spacing along every row.67 Tap cells tie the silicon under the transistors to the supply, so it stays at a fixed voltage.
Power domains
Many chips save power by running part of their logic at a lower voltage, or by switching it off when idle. Each such group of logic is a , described in the file.9 The floorplan gives each domain a , a fenced region its cells must stay inside. OpenROAD sets it with set_domain_area.8
Signals that cross the fence need help. A is needed where a low-voltage domain drives a high-voltage one, and level shifters add enough area and wire that they should be planned during floorplanning.19 The UPF also asks for , which hold a signal leaving a switched-off domain at a fixed value, so the logic that is still on doesn’t see garbage.8 Both kinds of cell need room near the boundary.
Hierarchy and partitions
Very large chips are split into blocks that are laid out separately, often by different teams: . Top-down planning gives each block an outline, pins and its own timing targets.16 A signal that leaves one block and enters another still has to arrive within one tick of the clock (at 1 GHz, one nanosecond). The splits that time between the blocks, and each block’s constraint file records its share.
Blocks can be laid out with channels between them, which hold top-level wiring and buffers,15 or , packed edge to edge. Since wires can run over the blocks on upper metal layers, layouts without channels have become the norm.13 Abutting saves the channel area, but with no space between blocks, every connection between them, including the clock, has to pass through the blocks themselves and be planned into them up front.
Is it feasible?
A floorplan is a guess until something is placed in it. Planners run a quick , sometimes called a prototype, to check that the plan can be built,16 then look at two quick estimates.15 The first is : places where more wires need to pass than there are routing tracks. The second is timing: whether signals can get from one flip-flop to the next within a clock tick, using rough estimates of wire delay and of the buffers that will be added. These numbers are trends, not final results, but a crowded channel or a signal that has to cross the whole chip points straight at a macro, a pin or a channel to change.
Outline and rows
OpenROAD’s initialize_floorplan takes either a target (-utilization, -aspect_ratio, and -core_space as one value or as bottom, top, left, right) or explicit -die_area and -core_area, plus the -site that defines a row’s unit width and height. Its documented arithmetic is short:
and the die is the core plus the margins. -additional_sites builds rows for site types the netlist doesn’t use yet, such as double-height cells expected later, and -flip_sites flips the row orientation for named sites.2 Rows alternate between normal (N) and flipped (FS) orientation so that neighboring rows share a power or ground rail along their common edge. DEF DIEAREA can be a right-angled polygon, which matters for an L-shaped block carved out of a parent floorplan.10
Keep the core edges on the site and track grids. Pins are placed on routing tracks,4 and an off-grid core edge leaves rows and pins that don’t line up with the tracks, which wastes routing along that edge.
Picking utilization: a typical target is 60–70%, in a definition that counts macros as filled area, chosen to leave room for routing.1 The headroom isn’t waste. Later steps add cells: clock buffers, buffers inserted to fix timing, and larger versions of slow cells. Teams usually start from what worked on earlier projects and adjust after the first trial placement.
A pad-limited die changes the arithmetic. sets the die size, and spare core area appears whether or not you need it.11 Area-array remove that coupling, but then bump planning becomes its own task: defining the bump grid, assigning each net to a bump, and routing bumps to I/O cells on the redistribution layer (RDL).5 A pad ring may also be more than one ring deep, with corner and filler cells to close each ring.5
Pins and bumps
OpenROAD’s pin placer (ppl) puts pins on the track grid to minimize wirelength, optionally refining them with simulated annealing. set_io_pin_constraint restricts pins to an edge or an interval of an edge, mirrors pairs of pins, and keeps groups together and in order; place_pins adds corner avoidance and minimum spacing. For chips stacked face to face, define_pin_shape_pattern defines a grid of pin positions across the die for micro bumps and .4
The catch is in the cost function. With cells unplaced, ppl computes wirelength as if every cell sat at the die center,4 so a first pin assignment is close to arbitrary. Iterate: place pins, trial-place, place pins again, and lock them only when the block owner and the top-level owner agree. In a design built in blocks, block pins come from the top-level view so that neighboring blocks line up.
Macro placement in practice
The hand recipe is: read the dataflow, group memories by function, tile each group, and put the tiles along the periphery. Periphery also matters when there are few routing layers to cross macros, and tiling eases building the power grid.17 The recipe breaks down when the sum of macro perimeters is large compared with the floorplan perimeter: macros then stack many deep, the core ends up far from many of them, and wires stop following the dataflow. Hier-RTLMP lets macro groups move into the core by tiling them along the internal boundaries of physical clusters.17 RTL-MP adds stability: macro locations should stay put when RTL changes are small, so it accepts preferred locations and blockages from the user.16
Macro planning also has to handle groups of macros with identical footprints, which tile with less wasted space when kept together; macros whose orientation is limited to mirroring and 180° rotation; and notches, small leftover regions that standard cells can’t use well and that hurt routability.1716 Analog IP and PLLs want distance, substrate contacts and guard rings between them and noisy logic.18
Channels and keep-outs
Size a from its routing demand: the pins on the facing macro edges, the nets that pass through, the power straps that must cross, and the buffers and standard cells you intend to allow there.15 A useful rule is to make a channel either wide enough to hold real logic or close it completely by abutting the macros. A channel in between is too narrow for useful logic but still attracts cells, which then can’t get their wires out.
DEF gives three tools for those slivers.10 A hard placement blockage keeps all cells out. A SOFT blockage keeps the initial placement out but lets clock-tree synthesis and timing optimization drop buffers in later. A PARTIAL blockage caps standard-cell density inside a rectangle, say at 50%, which is a common treatment along macro edges. OpenROAD’s macro placer weights boundary and notch penalties alongside wirelength, and its -pin_aware_channels option trims channels based on where macro pins actually are.6
Voltage areas
A is a hard region: the domain’s cells must stay inside it.8 Level shifters belong on each crossing where a low-voltage domain drives a high-voltage one; without them the receiver suffers short-circuit current and leakage. Yu et al. measure the cost as interconnect length overhead: a shifter placed outside its net’s bounding box lengthens the net. That is why they assign voltages and shifters during floorplanning.19
The UPF also names isolation strategies with their clamp values (the fixed level an isolated signal is held at) and power switches for domains that shut off.89 The switches are physical cells, so a domain that shuts off needs room for them, and each domain needs room inside it, along its boundary, for the shifters and isolation cells.
Hierarchical planning and budgets
Top-down hierarchical flows on a fixed die turn each block into a fixed-outline problem: minimize wirelength inside a given outline, with blocks whose aspect ratio may vary within limits.22 The top level hands each block a fixed outline, pin locations and constraints,16 with set as arrival times at block inputs and required times at block outputs.14 Budgets rest on the top-level floorplan’s early estimates, which are crude and sensitive to block shapes, top-level routing and pin assignment, so treat them as provisional and revisit them as blocks mature.13 Large designs must also be planned while the RTL is still changing: designers prototype from earlier revisions and expect macro locations to stay put when the RTL changes only a little.16
Feasibility before commitment
A prototype can work on a clustered netlist rather than individual cells, which shrinks the problem and splits closure into a global step on the clusters and a detailed step on the gates.16 Planners also estimate timing and routing congestion at this stage, to find channel and block hot spots.15
Two readings drive most decisions. The first is overflow: in each small tile of the global-routing grid, how many more wires are needed than there are tracks, usually concentrated in channels and at macro corners. The second is where the worst timing paths run. Worst negative slack (WNS) is how late the single worst signal arrives against its deadline. If the worst paths cross the die between a macro and logic that should have been its neighbor, the floorplan is at fault: a path from an SRAM on one edge to its consumer on the far side is a floorplan bug, not a placement bug.
A moderate fill: room left for wires and for buffers added later. Core 934 µm square, die 974 µm. Counting macros as filled: 80%.
A narrow, open channel attracts cells that then can’t get their wires out. A congestion hot spot.
Below is an empty chip with four big blocks. Three are memories. The fourth, the PLL, makes the chip’s clock, the steady beat that keeps every part in step. Drag the blocks around. Lines show what connects to what, and thicker lines carry more wires.
Press Scattered, then Edges, and watch the “wire length” number drop. Then try Tight channels to see which gaps are too narrow for wires.
Below is a die with four macros: two SRAMs, a ROM and a PLL. The memory interface pins are on the left edge and connect to both SRAMs; the reference-clock pin is on the top edge and connects to the PLL; general-purpose pins on the right connect to the logic. The standard-cell logic isn’t drawn as separate cells. It is treated as one soft region centered in whatever space the macros leave, and every macro connects to it too.
The straight lines (flylines) show what connects to what; thicker lines stand for wider buses. The wirelength readout adds up each line’s length, measured horizontally plus vertically because chip wires only run in those two directions, weighted by the line’s width. Utilization is the cell area divided by the core area left after macros and halos, so turning halos on raises it. The narrowest channel readout flags gaps that are too narrow to route, and overlapping macros show as errors. Start from Scattered, compare it with Edges, then rotate an SRAM with the R button and see how its lines change.
Same die, macros and pins, with one extra readout: a congestion total. For every channel under 100 µm it divides the length of the facing macro edges (a stand-in for how many pins must squeeze through) by the channel’s width, and sums the results. Load Tight channels, find the worst gap, and try three fixes: widen it, rotate a macro so long edges stop facing each other across a narrow gap, or push the macros together to remove the channel. Turn halos on and watch overlaps appear where macros sat legally a moment ago.
The model is a toy. It rotates macros by 90°, which real macros usually can’t do, and wirelength is Manhattan distance to the center of the free area with no routing, so a layout can win on wirelength and still lose on congestion. Hier-RTLMP’s cost function makes the same trade-off explicit with separate wirelength, boundary and notch terms.17
When engineers check a floorplan, they look at the picture first. Are the big blocks on the edges, close to what they connect to? Is there one big open area for the small parts, not a maze of leftover corners? Are the gaps between blocks wide enough for the wires?
Floorplans are built by scripts, so they can be rerun whenever the design changes. Below is an illustrative script in Tcl, the scripting language most chip tools use: each line is one command, and text after # is a comment. It is written for OpenROAD and builds the design from the worked example above. In order, it reads the library files and the netlist, sets the outline and a power-domain region, places the pins and the macros, adds tap and endcap cells, saves the floorplan, and then runs a throwaway trial placement to check it. The notes beside the code explain the lines that matter. Real flows split this into several scripts.
Below are an illustrative script, the DEF it produces, and an annotated log modeled on OpenROAD output (not verbatim). Read the log for three things: utilization of the free area and of each voltage area, the narrowest channel the macro placer accepted, and where the trial placement’s congestion and worst timing path sit relative to the macros.
# floorplan.tcl: illustrative OpenROAD-style floorplan step, not a complete flow
read_lef tech/tech.lef
read_lef lib/stdcells.lef
read_lef lib/macros.lef ;# abstracts for SRAM_2KX64, ROM_4KX32, PLL_CORE
read_liberty lib/stdcells_typ.lib
read_liberty lib/macros_typ.lib
read_verilog results/1_synth.v
link_design soc_top
read_sdc constraints/soc_top.sdc
read_upf -file constraints/soc_top.upf
# 1. Die, core and rows (microns)
initialize_floorplan -die_area {0 0 980 980} -core_area {20 20 960 960} -site CoreSite
# or size it from a target: -utilization 70 -aspect_ratio 1.0 -core_space 20
make_tracks
set_domain_area PD_ACC -area {560 400 900 760}
# 2. I/O pins on the die edges
set_io_pin_constraint -pin_names {ddr_*} -region left:*
set_io_pin_constraint -pin_names {clk_ref} -region top:850-950
set_io_pin_constraint -pin_names {gpio_*} -region right:*
place_pins -hor_layers M3 -ver_layers M4 -corner_avoidance 20 -min_distance 2
# 3. Macros: pin the PLL by hand, let the macro placer do the rest
place_macro -macro_name u_pll -location {850 850} -orientation R0 -exact
rtl_macro_placer -min_channel_size 40 -max_num_level 2
# 4. Well taps and endcaps; rows are cut around macro halos
tapcell -distance 20 -tapcell_master TAPCELL_X1 -endcap_master ENDCAP_X1 \
-halo_width_x 10 -halo_width_y 10
report_design_area
write_def results/2_floorplan.def
write_db results/2_floorplan.odb
# 5. Feasibility check on a throwaway copy
global_placement -density 0.70
global_route -congestion_report_file reports/fp_congestion.rpt
estimate_parasitics -placement
report_worst_slack -max- 1L4Macro LEF files: each macro’s size, pin shapes and blocked areas. That is all the floorplanner sees of an SRAM.
- 2L10Power intent (UPF). Each power domain declared here needs a physical region below.
- 3L13Explicit outline: a 980 µm die with a 940 µm core and 20 µm margin, rows built from CoreSite.
- 4L14The alternative: let the tool size the core from utilization, aspect ratio (height ÷ width) and core spacing. Check whether the tool’s utilization counts macro area before picking the number.
- 5L16The voltage area for the power-gated accelerator domain. Its cells must stay inside.
- 6L19Keep the DDR pins on the left edge, next to the SRAMs that use them. Pin lists are shortened here.
- 7L20The reference clock enters on the top edge right above where the PLL will sit.
- 8L22Left- and right-edge pins go on metal layer M3 (horizontal wires), top and bottom on M4 (vertical), snapped to routing tracks. Cells are unplaced, so the tool treats them as sitting at the die center.
- 9L25Hand-place the PLL in the top-right corner. R0 is OpenDB’s name for DEF orientation N; MY is FN, MX is FS, R180 is S.
- 10L26Hier-RTLMP: clusters by hierarchy and dataflow, then anneals cluster and macro positions. Channels narrower than 40 µm are not allowed.
- 11L29Tap cells every 20 µm in a checkerboard pattern; endcap cells at the ends of rows, including where rows are cut around macro halos.
- 12L33The floorplan hand-off (DEF and .odb). Everything below is a throwaway trial.
- 13L36Trial placement at 70% target density: is there enough room, and where does it get crowded?
- 14L37A quick global route (a coarse plan of every wire over a grid of tiles) gives a congestion map per layer. Hot spots near macro corners and channel mouths point at the floorplan.
- 15L39Worst setup slack: how much time the slowest path has to spare, with an ideal clock and estimated wires. A large negative number on a path that crosses the die is a floorplan problem.
Here is part of the DEF file the script writes. Each line describes one thing on the chip, and coordinates are in database units, 1,000 per µm, so 980000 means 980 µm. The standard cells have no location yet; only the macros, the tap and endcap cells, and the pins do. “N” and “FN”, “FS” and “S” say which way a part faces: normal, mirrored, flipped upside down, or turned 180°.
VERSION 5.8 ;
DESIGN soc_top ;
UNITS DISTANCE MICRONS 1000 ;
DIEAREA ( 0 0 ) ( 980000 980000 ) ;
ROW ROW_0_1 CoreSite 340000 20000 N DO 2000 BY 1 STEP 200 0 ;
ROW ROW_1_1 CoreSite 340000 22000 FS DO 2000 BY 1 STEP 200 0 ;
# ... more row segments, cut around macro halos ...
ROW ROW_469 CoreSite 20000 958000 FS DO 4100 BY 1 STEP 200 0 ;
REGIONS 1 ;
- PD_ACC ( 560000 400000 ) ( 900000 760000 ) + TYPE FENCE ;
END REGIONS
COMPONENTS 8 ; # count shortened to this excerpt; the real file also lists every cell
- u_sram0 SRAM_2KX64 + FIXED ( 30000 30000 ) FN + HALO 10000 10000 10000 10000 ;
- u_sram1 SRAM_2KX64 + FIXED ( 30000 250000 ) FN + HALO 10000 10000 10000 10000 ;
# ... u_sram2 and u_sram3 stacked above at the same x ...
- u_rom ROM_4KX32 + PLACED ( 750000 30000 ) N + HALO 10000 10000 10000 10000 ;
- u_pll PLL_CORE + FIXED ( 850000 850000 ) N + HALO 10000 10000 10000 10000 ;
- ENDCAP_469_0 ENDCAP_X1 + SOURCE DIST + FIXED ( 20000 958000 ) FS ;
- TAP_469_0 TAPCELL_X1 + SOURCE DIST + FIXED ( 40000 958000 ) FS ;
END COMPONENTS
PINS 212 ; # 3 of the 212 shown
- clk_ref + NET clk_ref + DIRECTION INPUT + USE CLOCK
+ LAYER M4 ( -100 -400 ) ( 100 0 ) + PLACED ( 900000 980000 ) N ;
- ddr_dq[0] + NET ddr_dq[0] + DIRECTION INOUT + USE SIGNAL
+ LAYER M3 ( 0 -100 ) ( 400 100 ) + PLACED ( 0 120000 ) N ;
- gpio[0] + NET gpio[0] + DIRECTION INOUT + USE SIGNAL
+ LAYER M3 ( -400 -100 ) ( 0 100 ) + PLACED ( 980000 500000 ) N ;
END PINS
BLOCKAGES 4 ;
- PLACEMENT + SOFT RECT ( 20000 220000 ) ( 340000 240000 ) ;
# ... the same for the two other gaps between stacked SRAMs ...
- PLACEMENT + PARTIAL 50.0 RECT ( 340000 20000 ) ( 400000 880000 ) ;
END BLOCKAGES- 1L4The die outline: 980 × 980 µm. DIEAREA can also be a rectilinear polygon.
- 2L5Rows are cut around macros. This bottom row starts right of the SRAM column’s halo (x = 340 µm) and stops at the ROM’s halo: 2,000 sites of 0.2 µm.
- 3L6The next row is 2 µm higher and flipped (FS) so neighboring rows share a power or ground rail.
- 4L8The top row runs from the left core edge to the PLL’s halo.
- 5L10The voltage area as a FENCE region: PD_ACC cells must stay inside, and no other cells may enter.
- 6L13FIXED: automatic tools may not move it. FN mirrors the SRAM so its pin edge, on the left in the LEF abstract, faces the core logic.
- 7L14Stacked 40 µm above the first SRAM (210 to 250 µm). With two 10 µm halos, 20 µm of free space remains between them.
- 8L16PLACED: the ROM has a spot but tools may still nudge it. The team has not locked it yet.
- 9L17The PLL in the top-right corner, its halo touching the core corner, next to the reference-clock pin.
- 10L18Physical-only cells carry SOURCE DIST. Endcaps close each row segment; taps follow at the set distance.
- 11L22Pins carry a layer shape relative to their location. This one sits on the top edge above the PLL.
- 12L30A soft blockage in the 20 µm sliver between SRAM halos: no cells during initial placement, but buffers may use it later.
- 13L32A partial blockage along the macro column caps standard-cell density at 50% where routes to SRAM pins converge.
And a log of the run: the messages the tool prints as it works, each tagged with the command that printed it. The useful lines are the utilization of the free area and of the power domain, the narrowest channel, and, in the trial placement at the end, where wiring got crowded and which signal path was slowest. Negative slack means a signal arrives late: −0.142 ns means 0.142 billionths of a second too late for the clock.
[ifp] Die (0, 0) (980, 980) um, core (20, 20) (960, 960) um, margin 20 um
[ifp] Added 470 rows of 4,700 CoreSite sites (0.20 x 2.00 um)
[upf] Power domain PD_ACC area (560, 400) (900, 760) um
[ppl] Placed 212 pins: left 96, top 4, right 112, bottom 0
[mpl] Physical hierarchy: 2 levels, 11 clusters (3 macro, 2 mixed, 6 std-cell)
[mpl] Placed 5 macros; u_pll fixed by user
[mpl] Narrowest macro-to-macro channel 40.0 um (u_sram0 / u_sram1)
[tap] Cut rows around 6 macros with 10 um halo: 470 rows -> 512 row segments
[tap] Inserted 1,024 endcaps and 7,296 tapcells
[util] Core area 883,600 um^2; macros + halos 301,200 um^2
[util] Standard-cell area 400,012 um^2 -> 68.7% of free core area
[util] Power domain PD_ACC: 122,400 um^2, cells 81,550 um^2 (66.6%)
---- feasibility check (placement discarded) ----
[gpl] Converged: overflow 0.099 at iteration 488
[grt] Overflow H 0, V 37; worst GCell +4 near (345.0, 230.0) um
[sta] Ideal clocks, placement-based parasitics
[sta] WNS -0.142 ns, TNS -3.81 ns, 61 failing endpoints
[sta] Worst: u_sram3/DO[17] -> u_ls_acc_17 -> u_acc/mac_7/acc_reg[31]/D- 1L2940 µm of core height ÷ 2 µm rows = 470 rows; 940 µm ÷ 0.2 µm sites = 4,700 sites per full row.
- 2L4Pins on the edges the constraints asked for. Bottom is empty because nothing was constrained there.
- 3L5The macro placer turned the logical hierarchy into physical clusters before annealing their positions.
- 4L7Exactly at the minimum channel size. Check it in the congestion report below.
- 5L8Each macro halo cuts the rows it crosses, so row segments outnumber rows.
- 6L11The headline number: cells over free area, not over the whole core. Counting macros as filled area, the same core is 79% full, so say which definition a target uses.
- 7L12Check each voltage area separately. A fence that is too full can’t be relieved by spreading cells outside it.
- 8L14Global placement stops when cells still overlap by under about 10% of their area; the detailed legalization that removes the rest is skipped in a trial.
- 9L15Overflow counts wires that don’t fit the tracks available: none horizontally, 37 vertically. The worst tile needs 4 more tracks than it has. It sits just right of the SRAM column (its halo edge is at x = 340 µm), beside the channel between u_sram0 and u_sram1: vertical routes running along the stacked SRAMs’ pin edges.
- 10L17WNS (worst negative slack) is the latest single path, 0.142 ns late; TNS (total negative slack) sums the lateness of all 61 late endpoints. Clocks are ideal and wires are estimated, so treat this as a trend, not a verdict.
- 11L18The worst path runs from the top SRAM on the left edge through a level shifter into the accelerator domain. Moving PD_ACC closer to the SRAMs, or the SRAMs it uses toward PD_ACC, is a floorplan fix.
The floorplan the script above writes, to scale. Tap a check to see the part of the map it’s about.
- Big blocks in the middle. They chop the floor into pieces, and wires have to go around them. Push them to the edges.
- Aisles too narrow. The wires can’t all fit through. Make the gap wider, or close it so the blocks touch.
- Plugs on the wrong side. If a block’s connections face the chip’s edge, every wire must wrap around it. Turn the block around.
- Planning with an old design. The design keeps changing. Check the plan again each time it does.
- Macros scattered in the core. They break the space for standard cells into pieces, leave awkward notches and force wires to detour. Line them up along the edges unless the flow of data argues otherwise.1617
- Starved channels. Narrow channels and notches between macros attract cells that then can’t get their wires out, which hurts routability.16 Size channels from an estimate of the wiring, or keep cells out of them with blockages.610
- Pins assigned blind. Before the cells are placed, the pin placer treats them all as sitting at the center of the die, so every pin aims there. Place the pins again after a trial placement.4
- Forgotten crossing cells. Level shifters and isolation cells need space at the edges of voltage areas. Without it they land far from their wires and make them longer.19
- Noisy neighbors. A PLL or analog block next to busy digital logic picks up noise through the shared silicon. Give it distance and guard rings.18
- Mixing up utilization. One common definition counts macro area as filled;1 another, like the worked example here, counts only standard cells over the free area. Comparing a number of one kind with a number of the other makes a full design look half-empty, or the reverse. Check how your tool computes it.
- Chopped-up rows. Every macro and halo cuts the rows it crosses into shorter pieces, each needing its own endcap cells, and the tool drops pieces shorter than a minimum width.7 Many macros can leave the cells only slivers to sit in.
- Peripheral placement past its limit. With hundreds of macros, edge stacking goes deep, the core ends up far from many macros, and wiring ignores the dataflow. Cluster and tile macros inside the core instead.17
- Unstable macro placement. An automatic placer that moves every macro on a small RTL change makes runs impossible to compare and breaks late fixes that assumed the old layout. Lock the macros you trust, give the placer preferred locations and blockages, and prefer placers designed for stability.16
- Optimizing the wrong proxy. Placers optimize cheap stand-ins for the routed result: half-perimeter wirelength (HPWL; for each net, half the perimeter of the smallest box around its pins), cell density and estimated congestion. These can disagree with what full place and route produces. Cheng et al. found weak rank correlation between Circuit Training’s proxy cost and post-route metrics among good solutions.24 Always confirm a floorplan change with a trial placement and global route.
- Budgets that don’t add up. Block budgets derived from the top level rest on early area and timing estimates, which are crude and sensitive to block shapes, top-level routing and pin assignment.13 Set before those settle, budgets can over-constrain one side of an interface and under-constrain the other, so revisit them as the blocks mature.
- Abutment without a plan. Abutted blocks have no top-level channels, so every crossing signal and the clock must be planned through the blocks up front.
- Pad-limited surprise. A core-limited die sits close to the line when its pads nearly fill the perimeter, and late I/O or power pads can push it over and grow the die. Track pad count against perimeter from the start.11
- Floorplanning a moving target. The RTL keeps changing during planning, and designers re-prototype from earlier revisions as it does.16 Keep the floorplan scriptable and re-runnable rather than hand-edited.
The macro’s pin edge faces the die edge. Every net detours around the macro, through the channels at its ends.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
The classical problem
Classical floorplanning places a set of modules (blocks) in the plane to minimize a weighted sum of the area of the bounding rectangle and an estimate of total wirelength. Modules are rigid, like pre-designed macros, or flexible, with a fixed area but a range of shapes.1315 The rectangle-packing variants are NP-hard.1 Positions are real numbers, so the space of possible placements is infinite. The key design choice is therefore a finite encoding of a layout that still contains an optimal packing and is cheap to change a little at a time.20 Floorplanners pair such an encoding with a search heuristic, most often .16 Annealing makes a random small change, always keeps it if the cost falls, and sometimes keeps it even if the cost rises, with that chance shrinking as a “temperature” is lowered, so the search can climb out of poor local optima early and settles down later.
Slicing floorplans and Polish expressions
A comes from cutting rectangles again and again with straight lines that run all the way across, so it maps to a binary tree of cuts. Written bottom-up in postfix order, the tree becomes a Polish expression: block names are operands and the two cut directions are operators, so blocks take symbols.15 For example, 12+43*+ combines blocks 1 and 2 with one kind of cut, blocks 4 and 3 with the other, and then joins the two results. Wong and Liu searched over a normalized form of these expressions, which speeds up the search.21
Evaluating a candidate is cheap. Each flexible block’s legal shapes form a shape function, and the shape functions combine bottom-up through the slicing tree, so the minimum area of a given slicing floorplan is found in polynomial time; for non-slicing floorplans the same problem is NP-hard.15 The limit is expressiveness: this encoding cannot handle non-slicing floorplans, where no single cut crosses the layout without splitting a block.21
Sequence pairs
Murata et al. encode any packing, slicing or not, as a : two orderings of the module names. If comes before in both, is left of . If comes before in but after it in , is above . Every pair of modules thus gets a horizontal or a vertical relation, so no two overlap. Take : and keep the same order in both lists, so is left of ; so are and ; but comes before in the first list and after it in the second, so sits above . The result is and stacked in a column with to their right.
Decoding turns the relations into two constraint graphs, one horizontal and one vertical, with each node weighted by its module’s width or height. The longest path to each node gives its or coordinate, and the longest path overall gives the chip’s width and height. The set of all sequence pairs is finite and contains at least one optimal packing, so annealing over it can in principle find the best answer. Murata placed the 49-module MCNC ami49 benchmark by annealing over sequence pairs.20 Sequence pairs are still the workhorse: OpenROAD’s macro placer anneals over them today.17
(abcd, bacd) decodes to a 9 × 5 packing, 31% of it dead space. Tap two names in one sequence to swap them.
B*-trees
Chang et al.’s B*-tree encodes a compacted placement, one in which no module can slide further down or left. The root is the bottom-left module. A node’s left child is the lowest unvisited module touching it on the right; its right child is the module directly above it with the same coordinate. Decoding is a linear-time walk of the tree with no constraint graphs, and after a swap only the modules later in the depth-first order need new coordinates. The annealer’s moves rotate a module, move one, swap two, or reshape a soft module. On MCNC benchmarks it ran about 4.5× faster and used about 60% less memory than the earlier O-tree encoding.21
Fixed-outline and mixed-size
Classical formulations leave the outline free and minimize area. Top-down hierarchical flows hand each block an outline instead, and the goal becomes minimizing wirelength while fitting inside it. Adya and Markov showed these fixed-outline instances are significantly harder than classical ones: classical methods failed on fixed-outline instances built from standard benchmarks. Adding new objectives and slack-based moves, which push modules with spare room toward free space, let their Parquet floorplanner find acceptable fixed-outline floorplans across a range of aspect ratios.22
The same paper describes a sequential flow: run a standard-cell placer on everything, cluster the cells into soft blocks by position, solve a fixed-outline floorplan so the macros no longer overlap, lock the macros, and place the cells again.22 Concurrent flows instead move macros and cells in one analytical placer. RePlAce models every instance as an electric charge and its density penalty as potential energy, spreading instances along the gradient; Nvidia’s AutoDMP adds concurrent macro and cell placement to the GPU placer DREAMPlace, with automatic parameter tuning by Bayesian optimization.24
RTL-MP and Hier-RTLMP
RTL-MP turns the netlist into a clustered netlist using the logical hierarchy, connection signatures between modules, the regularity of macro arrays and the number of flip-flop stages between clusters. It then places clusters and macros with sequence-pair annealing and handles pin access, notch avoidance and the push toward the periphery. Its objective weights are autotuned per design. On industrial designs it beat a commercial macro placer and came close to handcrafted floorplans after standard-cell place and route.16
Hier-RTLMP extends this to a multilevel physical hierarchy for blocks with hundreds of macros. Each level is placed with multi-start sequence-pair annealing. The moves swap two clusters in the first sequence, the second, or both (each with probability 0.3), or reshape a cluster (0.1). The cost is a weighted sum of area, HPWL, and penalties for leaving the outline, missing the peripheral bias, blocking pin access, ignoring guidance regions and creating notches, each normalized by its starting value. Inside a leaf macro cluster, a similar annealer places the individual macros, with a move that flips them.17 OpenROAD exposes those weights directly as options to rtl_macro_placer.6
Reinforcement learning and the debate around it
In 2020 Google researchers proposed training a reinforcement-learning (RL) agent, a program that learns by trial and reward, to place the nodes of a netlist onto a chip canvas. A network trained to predict placement quality serves as the encoder of the agent’s policy, so it improves as it trains on more blocks and transfers to unseen ones. They reported placements superhuman or comparable on modern accelerator netlists in under six hours, against baselines that needed human experts in the loop for several weeks.23
The work appeared in Nature in 2021, and Google released an open-source implementation, Circuit Training (CT). CT clusters the standard cells, lets the agent place macros one by one at the centers of grid cells, places the cell clusters by force-directed placement, and rewards the agent with a proxy cost: a weighted sum of wirelength, density and congestion terms.24
Cheng, Kahng et al. reimplemented the undocumented parts and evaluated CT on open testcases against simulated annealing, RePlAce, AutoDMP, a commercial macro placer and human experts, using a commercial flow for post-route results. They reported that poor initial placement locations in CT’s input could degrade routed wirelength by up to 10%; that SA reached better proxy cost on most testcases; that human experts beat CT on large, macro-heavy designs; that analytical placers produced better routed wirelength; and that proxy cost correlated poorly with post-route metrics among low-cost solutions.24 Markov’s meta-analysis went further and concluded that the RL method lagged behind human designers, simulated annealing and commercial software.25
The original authors dispute these results. They argue the reproductions skipped pre-training, used 20× fewer experience collectors and half as many GPUs, did not train to convergence, and used testcases unrepresentative of modern chips. They report that the method, now called AlphaChip, has produced layouts used in three generations of Google’s TPU.26 The disagreement is unresolved in the literature. What both sides’ work shows clearly is narrower: proxy objectives need to be checked against post-route results, comparisons need strong baselines such as well-tuned annealing and analytical placers, and open benchmarks and flows are what let anyone check.24
Q1A chip’s core has 800,000 µm² of area. The macros and their keep-out borders take 200,000 µm². The small standard cells add up to 330,000 µm². What share of the remaining free area do the cells fill?
Q2What does it mean that a chip is pad-limited?
Q3Why do engineers run a quick trial placement before they lock the floorplan?
Q4Part of a chip runs at a lower voltage to save power. Why does the floorplan leave room along the edge of that region?
Q5A large chip is split into blocks. What is the main trade-off between leaving channels between the blocks and packing them edge to edge (abutted)?
Sources
Show Hide 26 sources
- Floorplan (microelectronics)What floorplanning covers (chip area and aspect ratio, I/O pads, macros, standard-cell rows; power/ground network not always included); core utilization (macros + cells) ÷ core of 60–70% typically targeted, leaving room for routing; variants of rectangle packing are NP-hard.
- Initialize Floorplan (ifp)initialize_floorplan by utilization, aspect ratio (height/width) and core space, or by explicit die and core areas; core_area = design_area / (utilization / 100); rows and tracks.
- Flow VariablesCORE_UTILIZATION, CORE_ASPECT_RATIO, CORE_MARGIN, DIE_AREA, CORE_AREA, MACRO_PLACEMENT_TCL, IO_CONSTRAINTS, TAPCELL_TCL, FLOORPLAN_DEF.
- Pin Placer (ppl)Pins placed on the die boundary on the track grid to minimize wirelength; unplaced cells assumed at the die center; edge/interval constraints; simulated annealing; pin grids for micro bumps and hybrid bonding.
- Chip-level Connections (pad)ICeWall-based: an I/O ring around the chip boundary connected with wirebond pads or a bump array; IO sites, corner and filler cells, bump arrays (rows × columns at a pitch), net-to-bump assignment, RDL routing.
- Hierarchical Macro Placement (mpl)rtl_macro_placer implements Hier-RTLMP; min channel size, annealing weights (wirelength, boundary, notch, guidance); place_macro with R0/MY/MX/R180 orientations.
- Tapcell (tap)tapcell inserts endcaps and tapcells; tapcells in a checkerboard at a set distance; halo widths used when cutting rows around macros; corner cells at macro corners; minimum row width during cut rows.
- Unified Power Format (upf)read_upf, create_power_domain, set_isolation, set_level_shifter, set_domain_area for a domain’s physical area.
- Unified Power FormatUPF is IEEE 1801; describes supplies, power switches, isolation, level shifters and retention.
- LEF/DEF 5.7 Language Reference (DEF syntax chapters)DIEAREA (rectangle or polygon), ROW, COMPONENTS (FIXED, PLACED, COVER, HALO [SOFT]), PINS, BLOCKAGES (SOFT, PARTIAL), REGIONS (FENCE, GUIDE).
- Pad Ring and Floor Planning (ELEC6231 lecture notes)Core surrounded by a pad ring; pad-limited vs. core-limited dies; block arrangement affects chip area, cost and yield; chips near square at the top level.
- Flip chipSolder bumps on the chip pads, chip flipped face-down; short connections cut inductance compared with wire bonding.
- Classical Floorplanning Harmful?Classical floorplanning: hard, soft and semi-soft blocks, minimum area with an optional wirelength estimate; channeled block layouts gave way to channelless layouts with over-the-block routing; pre-synthesis area and performance estimates are crude and sensitive to block shapes, top-level routing and pin assignment.
- Bounded Potential Slack: Enabling Time Budgeting for Dual-Vt Allocation of Hierarchical Design (slides)Hierarchical design adds floorplanning, pin assignment, wire resource allocation and time budgeting; time budgeting sets timing assertions at each block boundary: arrival time at block inputs, required arrival time at block outputs.
- VLSI Physical Design: From Graph Partitioning to Timing Closure, Chapter 3 slides (Chip Planning)Chip planning also covers I/O pad placement, defining channels between blocks for routing and buffering, power and ground networks, and estimation of timing and routing congestion; objective α·area + (1 − α)·wirelength; hard library blocks vs. blocks with flexible shapes; slicing trees and Polish expressions (2n − 1 symbols); shape functions combined bottom-up give a slicing floorplan’s minimum area in polynomial time (NP-hard for non-slicing).
- RTL-MP: Toward Practical, Human-Quality Chip Planning and Macro PlacementFloorplanning done manually by back-end designers capturing dataflow; macro placement rarely changes after it is fixed; clustered netlist from logical hierarchy and dataflow; pin access, notch avoidance, periphery; beats a commercial macro placer, close to handcrafted floorplans.
- Hier-RTLMP: A Hierarchical Automatic Macro Placer for Large-scale Complex IP BlocksHundreds of macros per block; floorplan DEF feeds physical synthesis; limits of peripheral placement; multilevel clustering; sequence-pair annealing moves and cost terms.
- Substrate Coupling in Digital Circuits in Mixed-Signal Smart-Power SystemsSubstrate noise mitigation by physical separation of aggressors and victims, substrate contacts and guard rings.
- Voltage and Level-Shifter Assignment Driven FloorplanningLevel shifters are needed where a low-voltage module drives a high-voltage one; their area and wirelength overhead should be handled during floorplanning.
- Rectangle Packing and Its Applications to VLSI Module Placement ProblemPackings are uncountably infinite, so the key is a finite solution space that contains an optimal packing; sequence-pair left-of/above relations; horizontal and vertical constraint graphs with longest paths; (n!)² sequence-pairs, at least one optimal (Theorem 4); simulated annealing with sequence-pair moves places MCNC ami49.
- B*-Trees: A New Representation for Non-Slicing FloorplansOrdered binary tree over a compacted placement; left child = adjacent module to the right, right child = module above; no constraint graphs; annealing moves; about 4.5× faster and 60% less memory than O-tree. Notes that Wong and Liu’s normalized Polish expressions speed up search but cannot handle non-slicing floorplans.
- Fixed-Outline Floorplanning: Enabling Hierarchical DesignFixed-outline formulation is harder than classical floorplanning; slack-based moves; mixed-size flow: place cells, cluster, floorplan macros, fix them, re-place cells.
- Chip Placement with Deep Reinforcement LearningRL agent places netlist nodes onto a canvas; learns from prior designs; claims superhuman or comparable placements in under six hours.
- Assessment of Reinforcement Learning for Macro PlacementOpen reimplementation and evaluation of Circuit Training against SA, RePlAce, AutoDMP, a commercial macro placer and human experts; proxy cost poorly correlated with post-route metrics.
- The False Dawn: Reevaluating Google’s Reinforcement Learning for Chip Macro PlacementMeta-analysis concluding that the RL method underperformed humans, simulated annealing and commercial software.
- That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip DesignResponse arguing the reproductions skipped pre-training, used less compute, did not train to convergence and used unrepresentative test cases; reports use in three TPU generations.