Design Flow · Stage 5 of 13 · Front end

Logic synthesis

Software turns the code into a parts list: thousands of small, ready-made circuit pieces and the wires between them.

A synthesis tool turns the design code into a netlist: a list of ready-made circuit blocks from a library, called standard cells, and the wires between them, chosen to meet a speed target while keeping area and power low.

Elaboration, technology-independent optimization, technology mapping, timing-driven sizing and restructuring, clock gating, retiming, physical-aware synthesis, and equivalence checking against the RTL.

Builds The gate netlist

By now the chip exists as a written description. It says what the chip should do, in a special language a bit like programming code. But a factory can’t build from a description. It needs to know exactly which tiny parts to use and how to wire them together.

Logic synthesis fills that gap. A synthesis tool reads the description and writes a : a long list of parts and the wires between them. The parts come from a catalog of ready-made building blocks called . One block might answer “yes” only when both of its inputs say “yes.” Another might remember a single 1 or 0.

The tool also has goals. The circuit must be fast enough, small, and careful with power, so batteries last. Engineers call this trio : power, performance (speed) and area (size).

A chip design starts life as code. Engineers describe what the chip does in a such as Verilog or VHDL. The style of that code is called (RTL): it says which values the chip stores, how they change on each tick of its clock, and what calculations happen in between. RTL reads like a program, but nothing in it says which physical circuits to build.

Logic synthesis is the step that decides. A synthesis tool reads the RTL and produces a : a list of circuit components and the wires that connect them. Every component is a , a small pre-designed circuit from a library made for one particular manufacturing process. Most cells are , which compute simple functions of 0s and 1s (an AND gate outputs 1 only when both its inputs are 1), or , which each store one bit.

A few words this chapter uses throughout:

  • . A signal that switches between 0 and 1 at a fixed rate, such as a billion times a second (1 GHz). Every flip-flop updates on each rising edge, or “tick.” The time between ticks is the clock period: 1 nanosecond (ns) at 1 GHz.
  • Register. A group of flip-flops holding a multi-bit value, such as a 32-bit number.
  • . The route a signal takes from one flip-flop, through a chain of gates, to the next flip-flop. It has to finish within one clock period.

Most synthesis tools work in three broad phases:

  1. Elaboration: read the RTL and build a generic circuit of adders, comparators, selectors and registers that isn’t tied to any library yet.
  2. Technology-independent optimization: simplify that circuit by removing unused logic, sharing hardware and rewriting logic into fewer steps.
  3. Technology mapping and timing optimization: build the circuit from real library cells, then change cell sizes and reshape logic until every timing path fits in the clock period.

Widely used tools include Synopsys Design Compiler and Fusion Compiler, Cadence Genus and the open-source Yosys. Yosys runs the early phases itself and calls ABC, a logic optimizer from UC Berkeley, for optimization and mapping.

Synthesis comes after the RTL has been written and tested in simulation, and before test circuitry is added and the physical layout begins. It is the first point where a design gets numbers for : power, performance (speed) and area. They are estimates, because nothing has a physical position yet, but they show early whether the design can meet its goals.

Synthesis is where the RTL first meets a real cell library, real timing goals and a real cost function. The tool’s job is to produce a gate-level (a Verilog file listing every library cell instance and the wires between them) that does exactly what the RTL does, meets the timing constraints written in with some margin, and leaves the layout tools a manageable problem. Synopsys Design Compiler and Fusion Compiler, Cadence Genus and open-source Yosys with ABC all follow the same broad order:

  1. Elaborate the RTL into a generic circuit.
  2. Optimize it without reference to any library.
  3. Map it onto library cells.
  4. Optimize incrementally for timing and power: resize cells, restructure logic, add buffers.

The quality of results here (QoR, in tool reports) sets a ceiling for the rest of the flow. Layout tools can resize cells, add buffers and restructure logic locally, but they rarely recover from a poorly structured arithmetic block or a badly encoded state machine. Most of the risk at this stage sits in the inputs and estimates rather than in the algorithms:

  • Constraints. SDC that asks for too little leaves real paths unoptimized and unchecked. SDC that asks for too much wastes area and power.
  • Wire estimates. Nothing is placed yet, so wire delay is a guess. When the guess disagrees with what placement later finds, timing moves. Physical-aware synthesis exists to narrow that gap.
  • Verification. Transformations such as retiming (moving flip-flops through the logic) and re-encoding state machines make it harder to prove, with , that the netlist still matches the RTL.
RTL (Verilog)always @(posedge clk) y <= (a & b) | c;goal (SDC)clock period 2.0 nssynthesisnetlistabcyclkn1n2AOI21_X1U1INV_X1U2DFF_X1y_regareadelaypower
Clock goal

Two lines of RTL became three library cells and the nets between them. Tap a part, or tighten the goal.

RTL plus a timing goal in, a netlist of library cells out. A tighter goal buys speed with bigger cells. Cell sizes and bars are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Synthesis needs three things: the checked design description, the catalog of building blocks, and the goals, such as how fast the chip must run. It hands back the list of parts and wires. It also reports how fast, how big and how power-hungry the result is.

DirectionWhatFormat
InDesign source: the RTL code, plus its file list and settingsVerilog / SystemVerilog / VHDL
InCell library: each cell’s function, area, delay and powerLiberty .lib, one per process/voltage/temperature corner
InTiming constraints: clock speed, budgets at the inputs and outputs, exceptionsSDC, a list of Tcl commands
In (physical-aware only)Cell outlines and a rough floorplanLEF, DEF
OutThe mapped designGate-level Verilog netlist
OutConstraints for later toolsSDC rewritten to use the netlist’s names
OutQuality of resultsTiming, area, power and summary reports (text)

Those formats in plain words:

  • Verilog is the language the RTL is written in. The output netlist uses it too, in its simplest form: one line per cell, naming the cell and what each of its pins connects to.
  • Liberty is a text file describing every cell in the library. A library is measured at a set of conditions called a : how fast the transistors came out in manufacturing (process), the supply voltage and the temperature. A chip has to work at its slowest corner, so synthesis usually checks speed there.
  • SDC holds the timing goals: how fast the clock ticks and how much time signals have at the block’s edges.
  • LEF and DEF describe physical shapes and positions. They only matter when synthesis is asked to take layout into account.

The netlist goes next to , which adds circuitry so each manufactured chip can be tested, and then to floorplanning, the first layout step. A separate tool such as OpenSTA can read the same netlist, Liberty files and SDC and check the timing independently.

Real runs take more inputs than the table shows:

  • A dont_use list of cells the tool must not pick.
  • Modules or instances to keep intact rather than optimize across.
  • Power intent, usually a UPF file, if the design has more than one voltage domain.
  • Library sets that cover several threshold-voltage flavors (the same cells built to be faster and leakier, or slower and more frugal) and at least one slow corner for checking setup timing.
  • For physical-aware runs, the technology LEF, the cell LEFs and a floorplan DEF (die size, positions of large blocks, blocked areas), so the tool can run a coarse placement while it optimizes.

The teams downstream need more than the netlist. Design for test needs to know which flip-flops sit behind clock-gating cells, and whether those cells’ test-enable pins are tied off or left for the scan chains. Floorplanning wants big blocks kept as separate modules. The equivalence checker needs the RTL, the netlist and any notes the synthesis tool wrote about flip-flops it renamed, merged or moved.

inputsoutputsRTL (.v/.sv)Liberty (.lib)SDCLEF / DEFNetlist (.v)SDC (renamed)Reportssynthesis→ next: design for test
Physical-aware synthesis

Three inputs, three outputs. Tap a file to see what it holds.

Synthesis inputs (left) and outputs (right). LEF and DEF matter only in physical-aware runs.Share freely with credit: ‘Figure from chipfieldguide.com’

Synthesis happens in five big steps. It reads the description, simplifies it, picks real parts, checks the speed, and proves nothing changed. Step through them in the first picture below.

Every choice is a trade-off. Faster parts are usually bigger and use more power. To save power, the tool can also pause parts of the chip that have nothing to do right now, like turning off the lights in empty rooms.

Elaboration

turns the RTL into a generic circuit. RTL is organized into modules, self-contained pieces that use one another much as functions do in a program. The tool fills in settings the code left open (a bus width written as a parameter, say), copies out repeated structures, connects the modules into one hierarchy and converts each piece of behavior into hardware: adders, comparators, registers and , selectors that pass one of several inputs to their output. Yosys, for example, first keeps these as coarse word-level cells, such as one 32-bit adder, and later breaks them down into single-bit gates.

This is also where storage is inferred. Code that updates a value on the clock tick becomes flip-flops. Code meant to be pure logic that forgets to assign a value in some case becomes a , a storage element nobody asked for: in that case the hardware has to remember the old value, so the tool builds something that can.

Technology-independent optimization

Before any library cell appears, the tool cleans up the generic circuit. It replaces logic whose output is always the same with a constant, deletes logic that drives nothing, and shrinks operators to the bit widths actually used: an adder whose top bits are never read loses them. Then come bigger changes:

  • Resource sharing. Two additions in different branches of an if-statement never happen in the same clock cycle. builds one adder with a multiplexer in front to select its inputs. Yosys’s share pass uses a (a program that decides whether any combination of inputs can make a logic formula true) to prove the two operations can never be active together.
  • State machine encoding. Control logic is often a : a register that records which step of a procedure the circuit is on, plus logic that picks the next step. The tool can find such a machine and change how its steps are numbered. Yosys can recode it as , with one flip-flop per state. A machine with 8 states needs 3 flip-flops when the states are counted in binary but 8 in one-hot; one-hot spends the extra flip-flops to make the next-step logic simpler and faster.
  • Boolean optimization. The logic is rewritten into fewer and shallower gates, much like simplifying an algebra expression. Modern tools do this on an And-Inverter Graph, a network made only of two-input AND gates and inverters (NOT gates).

Technology mapping with Liberty

rebuilds the optimized logic out of cells from the target library. A library offers many functions, from inverters, NANDs, NORs and XORs to combined gates such as AOI21, which computes NOT((A AND B) OR C) in a single cell, plus flip-flops. Each comes in several sizes, called . With generic names, NAND2_X1 is the smallest two-input NAND and NAND2_X2 is twice as strong: quicker when it has a lot to drive, but bigger and hungrier for power.

The tool learns about the cells from the file. For each cell it lists the logic function, the area, the leakage power (drawn even when nothing switches), the capacitance of each input pin, timing tables and the energy used each time the cell switches. Capacitance is the electrical load: the more pins and wire a cell output drives, the more charge it has to move and the longer that takes. Delay is stored as an table. Its rows are input transition times (how long the incoming signal takes to swing between 0 and 1) and its columns are output loads; the tool reads between the entries. The Transistors guide’s cell-library chapter shows how these tables are measured. A trimmed, illustrative entry:

generic_typ.lib (excerpt, illustrative)liberty
cell (NAND2_X1) {
  area : 0.80;
  cell_leakage_power : 12.4;
  pin (A1) { direction : input; capacitance : 0.0016; }
  pin (A2) { direction : input; capacitance : 0.0017; }
  pin (ZN) {
    direction : output;
    function : "!(A1 & A2)";
    timing () {
      related_pin : "A1";
      cell_rise (delay_template_3x3) {
        index_1 ("0.01, 0.05, 0.20");
        index_2 ("0.001, 0.005, 0.020");
        values ("0.012, 0.025, 0.071", \
                "0.018, 0.032, 0.079", \
                "0.035, 0.050, 0.098");
      }
    }
  }
}
  1. 1L2Cell area, usually in square micrometers (µm²). Synthesis adds these up for its area report.
  2. 2L3Leakage: power drawn while the cell sits idle, in the library’s power unit.
  3. 3L4Input pin capacitance, here in picofarads (pF). It adds to the load on whatever cell drives this pin.
  4. 4L8The logic function: output ZN is NOT(A1 AND A2). The mapper matches logic against this.
  5. 5L12index_1: input transition times in ns, one per row.
  6. 6L13index_2: output loads in pF, one per column.
  7. 7L14Delays in ns. They grow to the right (more load) and downward (slower input). A 0.05 ns input driving 0.005 pF gives 0.032 ns, or 32 picoseconds (ps).

Flip-flops are mapped separately from the rest of the logic. In Yosys, dfflibmap maps registers onto the library’s flip-flop cells and abc -liberty maps the logic.

Timing-driven optimization with SDC

Without goals, the tool has no idea how fast the design must be. The file supplies them as a list of commands.

  • create_clock names a clock and gives its period, for example 1.0 ns.
  • set_input_delay and set_output_delay say how much of each period is used up outside the block. The tool treats each input as if it came from a flip-flop on the same clock somewhere outside, and each output as if it fed one.
  • set_false_path tells the tool to ignore a path for timing. A typical case is a signal passing between two unrelated clocks through a synchronizer, a small circuit built to make that crossing safe.
  • set_multicycle_path gives a path more than one period, when the design only reads its result every few cycles.

The tool then checks every timing path. It adds up the cell and wire delays along the path to get the arrival time, the moment the signal reaches the next flip-flop. It also works out the required time: the next clock tick, minus a small safety margin, minus the receiving flip-flop’s (how long its input must be steady before the tick to be captured reliably). is the required time minus the arrival time.

A worked example. With a 1.0 ns period, a 0.08 ns margin and a 0.05 ns setup time, the signal is required by 1.0−0.08−0.05=0.871.0 - 0.08 - 0.05 = 0.87 ns. If it arrives at 0.80 ns, the slack is +0.07 ns: 70 ps to spare. If it arrives at 0.90 ns, the slack is −0.03 ns and the path fails. The path with the worst slack is the . On failing paths the tool swaps in bigger cells, reshapes the logic and adds buffers (cells that simply re-drive a signal); on paths with time to spare it swaps in smaller cells to win back area.

Area, power and timing trade-offs

Each of these choices moves . A tighter clock produces bigger cells, more buffers and more leakage. Two techniques go after power directly:

  • . A flip-flop with an enable (an input that says “load a new value this cycle” or “keep the old one”) is normally built with a multiplexer that feeds its old value back in. It still receives every clock tick and uses power on each. When a group of flip-flops shares one enable, synthesis can instead stop their clock with a single integrated clock-gating (ICG) cell. Yosys’s clockgate pass does exactly this and can pick the ICG from the Liberty file. While the group is idle its flip-flops see no ticks, so they don’t switch and use less power.
  • Multibit flip-flops. Packing several one-bit flip-flops into one shares the clock circuitry inside the cell and lowers the load on the clock wiring.

moves flip-flops through the logic to even out the work between them, without changing what the circuit outputs on each tick. Picture two stages of a pipeline, sections of logic separated by registers like stations on an assembly line. If the first takes 1.2 ns and the second 0.6 ns, a 1.0 ns clock fails. Moving some logic from the first stage across the register into the second can bring both under 1.0 ns, and results still take the same number of ticks to come out.

Power estimates at this stage have three parts. Leakage comes straight from each cell’s Liberty leakage value. Internal power comes from the Liberty energy tables, counted each time a cell switches. Switching power is the energy spent charging and discharging the wires and pins each cell drives, so it depends on their capacitance and on how often each signal changes. Without activity data from simulating real workloads, the tool assumes a default rate, so treat its power number as a rough guess. The clock is the busiest signal in the design, changing every cycle, which is why stopping it saves so much.

Hierarchy: flatten or keep

Flattening dissolves the module boundaries into one big circuit so optimization can work across them, which usually improves speed and area. Keeping the hierarchy makes the netlist easier to debug, to lay out block by block and to patch late. Flows often mix the two. The open-source OpenROAD-flow-scripts, for example, can synthesize hierarchically, keeping modules above an area threshold and flattening smaller ones.

Proving it’s still the same design

Synthesis rewrites almost everything, so teams check the result with (LEC) between the RTL and the netlist. LEC is formal: instead of trying sample inputs, as a test does, it proves mathematically that both give the same outputs for every possible input. The open-source tool EQY does this with Yosys.

Elaboration and the generic netlist

fixes parameter values, expands generate loops, connects the module hierarchy and turns each process (an always block) into word-level operators, multiplexer trees and storage elements. Yosys keeps these as coarse cells in its internal format, RTLIL (one cell for a whole 32-bit adder, say), and lowers them to single-bit gates later. Its default synth script shows the order most tools follow:

  1. proc converts processes into multiplexers and flip-flops.
  2. fsm finds and re-encodes state machines, wreduce trims operators to the bit widths actually used, and alumacc gathers arithmetic into adder and multiply-accumulate units.
  3. share merges operators that are never active together, and the memory passes recognize memories.
  4. techmap lowers everything to generic single-bit gates, and abc optimizes and maps them.

Commercial tools add datapath synthesis, which picks an adder or multiplier architecture (a small, slow ripple-carry adder or a large, fast parallel-prefix one, for example) from the timing context. Review elaboration warnings before anything else. Latches nobody intended (from a combinational block that leaves a signal unassigned on some branch), truncated widths, nets driven from two places and missing modules turned into black boxes all show up here first.

Technology-independent optimization

The logic is restructured on an (AIG), in which every function is a network of two-input ANDs whose connections may be inverted. That uniform form makes small local rewrites cheap. Two kinds alternate. Rewriting and refactoring replace small pieces of the graph with smaller equivalents, cutting the node count (a rough proxy for area) without adding levels. Balancing reshapes chains into trees, cutting the number of levels (a rough proxy for delay) without adding nodes. “Under the hood” shows how. Structural transformations also run at this level:

  • merges operators that are never active in the same cycle, such as two adders in different branches of a case, into one with a multiplexer on its inputs. It saves area but adds that multiplexer’s delay, so a timing-driven tool may decline to share, or un-share, on critical paths. Yosys proves the operators are mutually exclusive with a .
  • State machines can be re-encoded. Yosys documents (one flip-flop per state) as its recoding target. Binary or Gray encodings use fewer flip-flops but deeper next-state logic. Either way, recoding changes the set of flip-flops, which matters for equivalence checking later.

Mapping against Liberty

covers the optimized graph with library cells: every node must end up inside some cell, and the cover should be fast on critical paths and small elsewhere. The cell data comes from the file:

  • area and leakage power;
  • input pin capacitance, which sets the load each pin puts on whatever drives it;
  • tables of delay and output transition, indexed by input transition (also called slew) and output load;
  • internal-energy tables for power;
  • timing checks such as setup and hold for flip-flops, and limits such as max_capacitance on outputs.

The Transistors guide’s cell-library chapter walks through how those tables are measured and read. Design-rule limits (max_transition, max_capacitance, max_fanout) work as rules rather than costs: a net that breaks one gets fixed even if its timing was fine. That is why a net with a very high fanout, one driver feeding thousands of pins, ends up as a tree of buffers. OpenROAD’s repair_design command, for example, inserts buffers specifically to fix max slew, capacitance and fanout violations. Huge nets such as reset and scan enable are usually left alone in synthesis, marked ideal (assumed to have no delay), and buffered once placement shows where their loads sit.

Timing-driven synthesis

The tool times every path against the . For a path from one flip-flop to another it computes two numbers:

  • Arrival time: the launching flip-flop’s clock-to-output delay plus every cell and wire delay along the path.
  • Required time: the next clock edge, minus the capturing flip-flop’s setup time, minus any clock uncertainty.

is required minus arrival. Paths from the block’s inputs and to its outputs use the budgets set by set_input_delay and set_output_delay in place of the missing flip-flop. No clock tree exists yet, so clocks are : every flip-flop is assumed to see the edge at the same instant. In silicon the edge reaches different flip-flops at slightly different times (skew) and wobbles from cycle to cycle (jitter), so set_clock_uncertainty subtracts a margin to cover them. Exceptions need care:

  • A removes a path from both optimization and checking.
  • A with -setup N moves the setup check NN cycles out. By default the hold check moves with it, to one cycle before the new setup edge, so you add -hold N−1 to put the hold check back on the original edge.

Optimization tracks two numbers per group of paths: WNS (worst negative slack, the single worst path) and TNS (total negative slack, the sum over every failing endpoint). On paths with negative slack the tool:

  • upsizes cells to stronger drive strengths;
  • swaps cells to a faster, leakier threshold-voltage flavor;
  • restructures logic so late-arriving signals pass through fewer levels;
  • duplicates drivers that have many loads, and inserts buffers.

On paths with positive slack it does the reverse, downsizing cells to reclaim area and leakage. Yosys’s ABC integration shows the open-source version of this. A delay target (-D, in picoseconds) adds retiming toward that target, and a constraint file (-constr) adds buffering and up- and down-sizing steps.

Power-oriented transformations

  • . A flip-flop with an enable is normally built with a feedback multiplexer: when the enable is low it reloads its own output, but it still clocks every cycle. When many flip-flops share one enable, synthesis removes their multiplexers and puts one integrated clock-gating (ICG) cell on their clock instead. The ICG holds the enable in a latch while the clock is active, so the gated clock can’t glitch. Yosys’s clockgate pass groups flip-flops by clock and enable, can pick ICGs from Liberty, skips groups below a minimum size, and can tie a named ICG pin low for later connection to scan enable. The cost is timing. The ICG sits upstream of the flip-flops it feeds, so its clock edge arrives earlier than theirs, and the enable has to reach the ICG before that earlier edge. Some enables become new critical paths.
  • Multibit banking. pack two, four or more bits into one cell that shares its internal clock buffering and presents a smaller clock load. Merging works best on flip-flops that end up close together, so many flows bank at or after placement instead of, or as well as, in synthesis. OpenROAD-flow-scripts, for example, clusters flip-flops during placement.

Retiming

moves flip-flops across logic, forward or backward, to minimize the clock period or the number of flip-flops, while keeping the circuit’s cycle-by-cycle behavior at its inputs and outputs. A simple case: one pipeline stage takes 1.2 ns and the next 0.6 ns. Moving the register between them earlier, past the last 0.3 ns of the first stage’s logic, gives two 0.9 ns stages that both fit a 1.0 ns clock, with no added latency.

It is powerful on deep, unbalanced pipelines. The price is that it renames, splits and merges flip-flops, which complicates equivalence checking, debugging and any later change that refers to register names. Open flows treat it with caution. OpenROAD-flow-scripts marks its module retiming option experimental, warns that nothing checks the retimed design is equivalent, and notes that its objective ignores the SDC, even the clock period.

Hierarchy

Flattening dissolves module boundaries so optimization can reach across them: constants propagate into blocks, logic moves across ports, and outputs nobody reads disappear. Keeping hierarchy preserves block boundaries for hierarchical floorplanning, limits how far a late change spreads, and makes equivalence checking and timing debug easier. A common compromise keeps large blocks and flattens small ones. OpenROAD-flow-scripts exposes exactly that threshold, measured in multiples of a basic NAND2 gate’s area.

Physical-aware synthesis

A wire’s delay depends on its length, which isn’t known until placement. Classic synthesis guessed it with a , a table in the Liberty file that maps fanout (the number of pins a net drives) to an estimated length, capacitance and resistance. As processes scaled and wire delay grew, that broke down. A wire-load model guesses a net’s length from its pin count alone, but real nets with the same pin count vary widely, and many turn out longer than the model predicts, so synthesis underestimates their delay.

Physical-aware synthesis replaces the table with distances from a real, if rough, placement: the tool places cells quickly and estimates each net from where its pins sit. Coupling synthesis to placement has its own trap. Each logic change, such as upsizing a cell to fix timing, can disturb the placement and so the wire estimates, and iterating synthesis and placement this way may never converge. Integrated flows therefore update the placement incrementally, keeping it close to the final one throughout.

Equivalence checking

(LEC) works in two steps. First it matches state points (flip-flops, ports and black boxes) between the RTL, the reference, and the netlist, the implementation, mostly by name. That cuts the design into cones: the pure logic feeding each state point. Then it proves each matched pair of cones equivalent. Anything that breaks one-to-one matching needs extra handling: retiming, state re-encoding, and flip-flops merged because they held duplicate or constant values. EQY, for instance, lets you declare a new state encoding in a recode section, or exclude state registers from matching and use a sequential strategy that reasons across several cycles. Commercial synthesis tools can write guidance files describing such changes for the LEC tool.

RTLif (sel) y = a + b;else y = a + c;1 Elaborategeneric blocksabc++10sely
1 / 5

Elaboration: the if-else becomes two adders and a multiplexer selected by sel. Generic blocks, no library yet.

A two-way if-else through the main phases of synthesis. Cell names are generic and the slack values are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
tap a gate to upsize itFFFFNAND2X10.15AOI21X10.17OAI21X10.20NOR2X10.14XOR2X10.130next edge 1.0 nsarrival 0.90 ns required 0.87 nsslack −0.03 ns area minimum

Arrival 0.90 ns, required 0.87 ns: slack −0.03 ns, violated. Tap a gate to upsize it.

One timing path with the chapter’s worked numbers: 1.0 ns clock, 0.08 ns margin, 0.05 ns setup, so the data is required by 0.87 ns. Gate delays and areas are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
Clock pins hit each tick4 of 44 muxes, no ICGq0q1q2q3muxclken = 0
Clock gating
Enable en

No gating: every flip-flop receives every clock tick. With en = 0, each one reloads its own value through its mux, and still switches inside.

Clock gating replaces four enable multiplexers with one integrated clock-gating (ICG) cell. Start the clock, then switch gating on and change en.Share freely with credit: ‘Figure from chipfieldguide.com’
logic between registersFFFF0.30.30.30.30.30.3FFstage 11.2 ns ✗stage 20.6 ns ✓clock period 1.0 ns

Stage 1: 1.2 ns, stage 2: 0.6 ns. Stage 1 misses the 1.0 ns clock. Move the register.

Retiming with the chapter’s numbers: 1.2 ns and 0.6 ns stages on a 1.0 ns clock, drawn as six 0.3 ns pieces of logic.Share freely with credit: ‘Figure from chipfieldguide.com’

Type a short logic rule using four inputs: a, b, c and d. Each input is either 1 (yes) or 0 (no). Use & for AND, | for OR, ^ for “one or the other, but not both” and ! for NOT. The simulator shows the parts a tool would pick to build your rule, and how much smaller it can make them.

Try (a & b) | (a & c). Does the tool notice that a appears in both halves? It can build the same rule with fewer parts.

Enter a Boolean expression, a logic rule, with up to four inputs (a, b, c, d). & means AND, | OR, ^ XOR (one or the other but not both) and ! NOT. The simulator prints the rule’s (the output for every combination of inputs), a minimized (AND terms joined by OR) and a netlist built from a tiny generic library (INV, NAND2, NOR2, AND2, OR2, XOR2, XNOR2, AOI21, OAI21). It reports total area and a delay estimate that adds a fixed delay per cell along the longest path, so it tracks and ignores loads and wires. Three to try:

  • (a & b) | (a & c): the sum of products has two terms, but the factored form a⋅(b+c)a \cdot (b + c) needs fewer gates. Does the mapped netlist find it?
  • !((a & b) | c): this is exactly one AOI21 cell (AND-OR-INVERT). Compare its area and delay with an AND2 + OR2 + INV version.
  • a ^ b ^ c ^ d: this is parity (1 when an odd number of inputs are 1). Its sum of products has eight terms of four inputs each and nothing to merge, yet it maps to three XOR2 cells.

The simulator is a toy version of the front of the flow: truth table, two-level minimization, then a cover with a generic library (INV, NAND2, NOR2, AND2, OR2, XOR2, XNOR2, AOI21, OAI21) costed by area and a fixed per-cell delay summed along the longest path, which tracks .

  • Use a ^ b ^ c ^ d to see why two-level minimization lost to multi-level synthesis. No two input combinations that give a 1 differ in a single bit, so nothing merges: the SOP is eight 4-literal terms, while a balanced XOR tree is two levels deep.
  • Use !((a & b) | c) to watch a complex gate absorb three levels of generic logic.
  • Use (a & b) | (a & c) | (b & c) (majority) to compare covers chosen for area and for depth.

Real mappers replace the fixed per-cell delay with NLDM lookups that depend on load and input transition.

Loading simulation…

After a synthesis run, an engineer checks three numbers first.

  1. Is it fast enough? The report lists the slowest paths and how much spare time each has. A negative number means that path is too slow, so something must change.
  2. How big is it? If it is much bigger than planned, the chip may not fit or may cost too much.
  3. How much power? This is only an estimate so far. But a big surprise is easier to fix now than later.

Then they read the warnings. Sometimes the tool deletes a part because nothing uses what it makes. A warning like that often points to a mistake in the design.

Below are four illustrative files from a small open-source flow, with generic cell names: a Yosys script (the commands that run synthesis), an SDC constraints file, a piece of the resulting netlist and one timing path from the timing report. The notes beside each explain the lines that matter.

The artifacts below follow an open-source flow (Yosys and ABC, with OpenSTA-style timing reports). Commercial tools lay their reports out differently but carry the same information. All numbers are illustrative.

synth.ys (illustrative Yosys script)tcl
read_verilog -sv rtl/alu.sv rtl/top.sv
hierarchy -check -top top
synth -top top -flatten
dfflibmap -liberty lib/generic_typ.lib
abc -liberty lib/generic_typ.lib -D 1000
opt_clean
stat -liberty lib/generic_typ.lib
write_verilog -noattr out/top_netlist.v
  1. 1L1Read the RTL source files.
  2. 2L2Elaborate and check the module hierarchy under the chosen top module.
  3. 3L3Run Yosys’s generic synthesis script (proc, fsm, share, techmap, ABC and more), flattening the hierarchy first.
  4. 4L4Map the generic flip-flops onto the library’s flip-flop cells.
  5. 5L5ABC optimizes the logic and maps it onto the Liberty cells. -D sets a 1000 ps (1 ns) delay target and enables retiming toward it.
  6. 6L7Print cell counts and total area, using the areas from the Liberty file.
  7. 7L8Write the gate-level netlist for the next stages.
constraints.sdc (illustrative)sdc
create_clock -name clk -period 1.0 [get_ports clk]
set_clock_uncertainty 0.08 [get_clocks clk]
set_input_delay  0.40 -clock clk [get_ports {a_in[*] b_in[*] en}]
set_output_delay 0.35 -clock clk [get_ports {sum_q[*]}]
set_driving_cell -lib_cell BUF_X2 [get_ports {a_in[*] b_in[*] en}]
set_load 0.005 [all_outputs]
set_false_path -from [get_ports rst_n]
set_multicycle_path 2 -setup -from [get_cells u_div/*] -to [get_cells u_acc/*]
set_multicycle_path 1 -hold  -from [get_cells u_div/*] -to [get_cells u_acc/*]
  1. 1L1A 1.0 ns clock (1 GHz) on the input port clk. Time units follow the Liberty file, here ns.
  2. 2L2An 80 ps safety margin for clock skew and jitter, because the clock is treated as perfect during synthesis.
  3. 3L30.40 ns of each cycle is used up outside the block before these inputs arrive.
  4. 4L4The logic outside needs these outputs 0.35 ns before the next tick.
  5. 5L5Inputs are driven as if by a BUF_X2 buffer cell, so the tool sees realistic input transition times.
  6. 6L6Each output drives a 0.005 pF load (capacitance units also follow the library).
  7. 7L7Exclude the raw reset input, which can change at any moment, from timing. Reset is usually synchronized on-chip, and the synchronized reset is still timed.
  8. 8L8The divider’s result is only read every second cycle, so its setup check gets two cycles.
  9. 9L9Moves the hold check back to the original edge. Without it, hold would be checked a cycle late.
top_netlist.v (excerpt, illustrative)verilog
module top (clk, rst_n, en, a_in, b_in, sum_q);
  input clk, rst_n, en;
  input [1:0] a_in, b_in;
  output [1:0] sum_q;
  wire n1, n2, gclk;
  wire [1:0] sum_d;

  XOR2_X1  U1 (.A(a_in[0]), .B(b_in[0]), .Z(sum_d[0]));
  NAND2_X1 U2 (.A1(a_in[0]), .A2(b_in[0]), .ZN(n1));
  XOR2_X1  U3 (.A(a_in[1]), .B(b_in[1]), .Z(n2));
  XNOR2_X1 U4 (.A(n2), .B(n1), .ZN(sum_d[1]));

  ICG_X1   clk_gate_sum_q (.CK(clk), .E(en), .SE(1'b0), .GCK(gclk));
  DFFR_X1  sum_q_reg_0_ (.D(sum_d[0]), .CK(gclk), .RN(rst_n), .Q(sum_q[0]));
  DFFR_X1  sum_q_reg_1_ (.D(sum_d[1]), .CK(gclk), .RN(rst_n), .Q(sum_q[1]));
endmodule
  1. 1L5Internal wires (nets) the tool created. Their names mean nothing; only the connections matter.
  2. 2L8One cell: type XOR2_X1, instance name U1, then each pin and the net it connects to. Bit 0 of a + b is a XOR b.
  3. 3L9NAND2 gives the carry out of bit 0, inverted. The mapper chose it because a NAND is smaller than an AND.
  4. 4L11XNOR with the inverted carry equals XOR with the true carry, so no separate inverter is needed.
  5. 5L13Clock-gating cell, inserted because both flip-flops load only when en is 1: it passes clk to gclk only then. SE (scan enable) is tied to 0 until test logic connects it.
  6. 6L14A flip-flop with a reset pin (RN, active when 0). Its enable multiplexer is gone; it runs on the gated clock instead.
report_checks -path_delay max (illustrative)text
Startpoint: u_alu/op_a_reg_3_ (rising edge-triggered flip-flop clocked by clk)
Endpoint: u_alu/acc_reg_31_ (rising edge-triggered flip-flop clocked by clk)
Path Group: clk
Path Type: max

  Delay    Time   Description
---------------------------------------------------------------
  0.000   0.000   clock clk (rise edge)
  0.000   0.000   clock network delay (ideal)
  0.000   0.000 ^ u_alu/op_a_reg_3_/CK (DFF_X1)
  0.112   0.112 v u_alu/op_a_reg_3_/Q (DFF_X1)
  0.046   0.158 ^ u_alu/U812/ZN (NAND2_X1)
  0.071   0.229 v u_alu/U813/ZN (AOI21_X1)
  0.064   0.293 ^ u_alu/U840/ZN (OAI21_X1)
  0.069   0.362 v u_alu/U841/ZN (AOI21_X1)
  0.058   0.420 ^ u_alu/U902/ZN (OAI21_X2)
  0.073   0.493 v u_alu/U903/ZN (AOI21_X1)
  0.066   0.559 ^ u_alu/U955/ZN (OAI21_X1)
  0.070   0.629 v u_alu/U956/ZN (AOI21_X1)
  0.067   0.696 ^ u_alu/U1010/ZN (OAI21_X1)
  0.072   0.768 v u_alu/U1011/ZN (AOI21_X1)
  0.088   0.856 ^ u_alu/U1102/ZN (XNOR2_X1)
  0.034   0.890 v u_alu/U1150/ZN (NOR2_X1)
  0.000   0.890 v u_alu/acc_reg_31_/D (DFF_X1)
          0.890   data arrival time

  1.000   1.000   clock clk (rise edge)
  0.000   1.000   clock network delay (ideal)
 -0.080   0.920   clock uncertainty
          0.920 ^ u_alu/acc_reg_31_/CK (DFF_X1)
 -0.052   0.868   library setup time
          0.868   data required time
---------------------------------------------------------------
          0.868   data required time
         -0.890   data arrival time
---------------------------------------------------------------
         -0.022   slack (VIOLATED)
  1. 1L1The path starts at a flip-flop, launched by the clock tick.
  2. 2L2It ends at the data input of another flip-flop, which captures it on the next tick.
  3. 3L6Delay is each step’s own delay; Time is the running total, in ns.
  4. 4L9Ideal clock: no clock wiring exists yet, so the tick is assumed to reach every flip-flop at once.
  5. 5L11Clock-to-Q: time from the tick until the launching flip-flop’s output changes, from its Liberty table. ^ marks a rising signal, v a falling one.
  6. 6L13Alternating AOI21 and OAI21 cells are a typical mapped carry chain.
  7. 7L16The tool already upsized this stage to X2 trying to speed the path up.
  8. 8L25Total time for the data to reach the capturing flip-flop.
  9. 9L29The SDC’s uncertainty margin is subtracted from the time available.
  10. 10L31Setup time: the capturing flip-flop needs its data this long before the tick.
  11. 11L37Negative slack: the data arrives 22 ps late. Fix by restructuring the adder, retiming, or relaxing the clock.

The netlist is plain structural Verilog: no behavior or arithmetic is left, only cells and connections. Each line is one cell: the library cell’s name, a name for this copy of it (an instance) and its pin connections. The two-bit adder became four gates, and the enable on the output register became one clock-gating cell feeding both flip-flops. Names such as U2 are made up by the tool, while flip-flop names such as sum_q_reg_0_ keep a trace of the RTL signal they came from, which helps debugging and equivalence checking.

Read a timing path from top to bottom. The upper half adds up delays to get the arrival time. The lower half starts from the next clock tick and subtracts the margin and the setup time to get the required time. Slack is the difference. Here the path is the carry chain of a 32-bit adder: each bit’s carry depends on the bit below, so the signal ripples through a long chain of gates, and it misses by 22 ps.

Typical responses:

  • Let the tool choose a faster adder design, one that computes carries in parallel.
  • Turn on retiming, so the tool can move flip-flops to even out the stages.
  • Move part of the addition into the previous pipeline stage in the RTL.
  • Accept a slower clock.

A summary of the whole run condenses those reports. Formats vary by tool; the fields don’t.

qor_summary.txt (illustrative)text
Design                    : top
Corner                    : slow (setup)
Clock clk period          : 1.000 ns
WNS (setup)               : -0.022 ns
TNS (setup)               : -0.311 ns
Violating endpoints       : 27
Max transition violations : 3
Combinational cells       : 41,206
Flip-flop bits            : 9,730  (4,914 single-bit, 1,204 x 4-bit multibit)
Clock-gating cells        : 212   (gated flop bits: 88.4%)
Total cell area           : 48,930.6 µm²
Leakage power             : 1.84 mW
Dynamic power (estimate)  : 63.2 mW  (default toggle rate)
  1. 1L2Setup timing is judged at the slow corner. Hold timing (data must not change too soon after the edge) mostly gets fixed after the clock tree is built, so synthesis rarely reports it seriously.
  2. 2L4Worst negative slack: the single worst path, the one in the report above.
  3. 3L5Total negative slack: the sum over all failing endpoints. A small WNS with a large TNS means many paths fail, which usually points to the architecture or the constraints.
  4. 4L7Design-rule violations get fixed before slack. A non-zero count usually means a high-fanout net without buffers.
  5. 5L9Multibit banking packed about half the bits into 4-bit cells, lowering the clock-pin load.
  6. 6L10Share of flip-flop bits behind a clock-gating cell. Higher usually means lower clock power, if the enables really are off often.
  7. 7L13With a default toggle rate, dynamic power is a rough guess. Feed it activity recorded in simulation (SAIF or VCD files) for a real number.

Read the slack distribution before chasing the worst path. Here the worst path misses by 22 ps, but 27 endpoints fail for a total of 311 ps, about 12 ps each on average. Thousands of endpoints failing by a few picoseconds usually means the clock target, the uncertainty or a missing exception is wrong. A handful of deep failures points at specific RTL structures. Before the hand-off, time the same netlist and SDC again in a standalone timing tool, and run equivalence checking.

Worst slack (WNS)+0.01 nsCell area48,931 µm²Power (estimate)65 mWendpoints by slack0 failing−0.30+0.4 ns
Synthesis run

WNS +0.01 ns: every endpoint meets timing. Next: read the warnings, then re-check timing in a separate tool.

Three illustrative runs of the same design. The bars count timing endpoints by their slack (square-root scale); red ones fail.Share freely with credit: ‘Figure from chipfieldguide.com’
  • Wrong goals. If the speed goal is missing or mistyped, the tool may build something far too slow, or far too big. Teams check the goals as carefully as the design.
  • Code that means two things. Some design code acts one way in testing and another way once it becomes real parts. Rerunning the tests on the parts list catches this.
  • Missing pieces. If part of the design isn’t connected to anything useful, the tool deletes it. That usually means there is a bug.
  • Guessing the wires. Nothing has been laid out yet, so the tool can only guess how long the wires will be. A design that barely passes now may fail later.
  • Unconstrained paths. If the SDC is missing a create_clock for some clock, or an input or output has no delay set, the paths through it are never timed, and the tool doesn’t optimize what it doesn’t time. Every flow should report unconstrained endpoints (path ends with no timing check at all) and treat any as an error.
  • Inferred latches. A block of RTL meant to be pure logic must assign every output in every case. Miss one case and the hardware has to remember the old value, so the tool builds a latch. Latches complicate timing and are rarely intended. Read the elaboration warnings, or write such blocks with SystemVerilog’s always_comb, which tools check for exactly this.
  • Simulation and synthesis disagree. RTL is tested by simulating it before synthesis, and some code behaves differently in a simulator than in hardware. Examples: initial blocks, which set starting values in simulation but build nothing on a chip; delays written as #5, which synthesis ignores; incomplete sensitivity lists (the signals that trigger a block in older Verilog); and assigning X, “unknown.” Simulating key tests on the netlist itself and running equivalence checking catch most cases.
  • Logic optimized away. If an output is left unconnected or tied to a constant, synthesis deletes all the logic that feeds it. Treat a sudden drop in area or flip-flop count as a red flag.
  • Over-constraining. Setting the clock period tighter than needed “for safety” makes the tool upsize and buffer everywhere, which costs area and power and can make layout harder. Add margin with clock uncertainty instead.
  • Wire estimates. Without placement, wire delays are guesses, and wire-load models in particular tend to guess too short. Compare synthesis timing with timing after placement early, while there’s still time to adjust.

Most teams automate these checks. A synthesis run fails if it finds inferred latches, unconstrained endpoints, unmapped cells or a failed equivalence check, and each run’s reports are compared with the previous one, so a jump in area, cell count or negative slack gets investigated the same day. Catching a problem here costs minutes. Catching it after layout costs days.

  • Exceptions that hide real paths. A wildcard set_false_path, or a set_clock_groups declaration that calls two clocks unrelated when a real synchronous path runs between them, removes that path from optimization and from signoff, the final checks before manufacturing. Lint the SDC with a constraint checker, and review every exception with the RTL owner.
  • Multicycle hold mistakes. A setup multicycle without the matching hold adjustment moves the hold check a cycle late. That produces large hold violations that aren’t real, which later stages may then “fix” by inserting delay cells, cells whose only job is to slow a signal down.
  • Correlation gaps. Synthesis with wire-load models, or with no wire delay at all, followed by a real placement can shift slack a lot, because a wire-load model guesses each net’s length from its pin count and underestimates many nets. Physical-aware synthesis narrows the gap, but only if the placement it works from stays close to the final one.
  • Equivalence checks that give up or fail. Large multipliers and deep arithmetic can make the proof for a single cone so hard that the checker aborts. Retiming, state re-encoding and merged flip-flops break the matching of state points. Keep the synthesis tool’s guidance, use the checker’s settings for arithmetic, and don’t stack aggressive sequential transformations without a plan to verify them.
  • Clock-gating side effects. An enable must reach its ICG before the ICG’s clock edge, which comes earlier than the flip-flops’ edge, and the ICG’s test-enable pin must be connected or tied correctly for scan testing. Many tiny gating groups waste area and complicate clock tree synthesis, the later step that builds the clock wiring; set a minimum group size.
  • Boundary optimization and late changes. Flattened blocks, or blocks optimized across their boundaries, lose the meaning of their ports. That makes engineering change orders (small, late edits made directly to the netlist) and partial re-synthesis harder. Decide which blocks to keep before the first full run.
  • Library hygiene. Flows keep a dont_use list so the mapper avoids cells it shouldn’t pick, such as cells whose pins are hard for the router to reach. Loading the wrong corner or threshold-voltage set skews every result.
constraints.sdccreate_clock -period 1.0 [get_ports clk]# no set_input_delay on dindinU1U2U3FFclknot timed
Trap
Show

No input delay on din, so the path from din is unconstrained: never timed, never optimized. Report unconstrained endpoints and treat any as an error.

Four common synthesis traps, each as written and fixed (or, for wire estimates, before and after placement). Slack values are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

NPN classes of 4-input functions
222
Typical 5-input cuts per AIG node
20–30
Node types in an AIG
AND2 + inverted edges

Sources: Mishchenko et al. 2006 for the NPN class count, and Mishchenko et al. 2005 for the typical cut count per node. Both terms are explained below.

From two-level to multi-level logic

Early logic minimization targeted form, a sum of products (an OR of AND terms), which maps directly onto programmable logic arrays. Quine–McCluskey solves it exactly in two steps: list every prime implicant (an AND term that can’t be shortened without covering a 0), then choose the smallest set that covers every 1. It starts from the truth table, so its time and memory roughly double with every input added. Espresso (Brayton and colleagues at IBM, later refined at UC Berkeley) avoids expanding the function into its minterms, the individual rows of the truth table. It works heuristically on cubes (product terms) of the ON-set, the don’t-care set and the OFF-set, and its results come very close to minimal and are always free of redundant terms.

Two-level form is the wrong target for standard cells. Some common functions, parity among them, need exponentially many product terms. MIS and then SIS moved to multi-level Boolean networks: a graph in which each node holds a small SOP, simplified with Espresso against local don’t-cares (input combinations that can’t occur at that node), while algebraic operations (extraction, kerneling, resubstitution) find logic that several nodes can share. The SIS scripts that drove this worked well but were slow and needed hand-tuning.

And-Inverter Graphs

An is a directed acyclic graph (connections run one way and never loop back) in which every internal node is a two-input AND and any edge may carry an inversion. Primary inputs have no incoming edges. Each register is cut in two: its output acts as an extra input and its input as an extra output, so only the logic between registers remains. Structural hashing during construction keeps a table of every node’s pair of inputs, so no two AND nodes share the same pair and trivially duplicated logic merges for free. AIGs came out of formal verification and replaced SOP- and BDD-based representations in ABC’s synthesis flow. ABC now handles designs with millions of nodes that SIS could not finish.

Rewriting, refactoring and balancing

Two ideas make rewriting work.

  • A cut of a node is a small set of nodes that every path from the inputs to that node must pass through. The logic between the cut and the node is then a small function of the cut’s nodes. A cut is 4-feasible if it has at most four nodes, so its function fits in a 16-entry truth table.
  • Two functions are in the same NPN class if one becomes the other by negating inputs (N), permuting them (P) or negating the output (N). The 65,536 functions of four inputs fall into just 222 classes.

Rewriting is a greedy local search. It visits nodes in topological order, inputs first. For each node it enumerates the 4-feasible cuts, computes each cut’s function as a 16-bit truth table, looks up its NPN class and tries the precomputed AIG subgraphs stored for that class. For each candidate it counts how many nodes would disappear with the old subgraph and how many the new one needs, reusing nodes that already exist elsewhere in the graph (this is what “DAG-aware” means). It keeps the replacement that saves the most nodes without adding delay.

Refactoring takes one larger cut per node, collapses the logic inside it and factors it again algebraically. Balancing reduces depth by algebraic tree-height reduction: the chain ((a⋅b)⋅c)⋅d((a \cdot b) \cdot c) \cdot d, three levels deep, becomes (a⋅b)⋅(c⋅d)(a \cdot b) \cdot (c \cdot d), two levels deep, with the same number of nodes. ABC’s resyn2 script interleaves these. It balances first, then alternates rewriting and refactoring (less area, no more delay) with balancing (less delay, no more area). Each pass is cheap, so many passes add up to a global effect. The authors report it is several orders of magnitude faster than SIS and MVSIS, with comparable or better quality after mapping.

Cut enumeration and cut-based technology mapping

Classical mappers broke logic into a subject graph of NAND2 gates and inverters, described each library cell as a small pattern graph of the same gates, and covered the subject graph with those patterns using tree covering (as in SIS) or DAG covering. Because matching compared graph shapes, the result inherited the subject graph’s structure, good or bad. This is called structural bias.

Cut-based mapping with Boolean matching, from the same Berkeley group, compares functions instead of shapes. It works in five steps:

  1. Enumerate the kk-feasible cuts (at most kk nodes) of every AIG node.
  2. Compute each cut’s truth table.
  3. Look the truth table up in a hash table of library cell functions. This is Boolean matching; it also settles which cell pin connects to which cut node, and in which polarity.
  4. In topological order, compute the best arrival time at each node using the fastest matching cell.
  5. In reverse order, from the outputs back, select the cover, then recover area on paths that aren’t critical.

Boolean matching is faster than structural matching and finds more matches, especially for industrial libraries with large, complex gates. Two extensions attack structural bias further:

  • Supergates are small precomputed networks of library gates treated as single gates, so larger functions can match.
  • Choice nodes store several equivalent structures for the same logic in one AIG, so the mapper can pick among them.

The number of cuts grows quickly with kk: the same authors found most test cases have 20 to 30 five-input cuts per node. For mapping into FPGA lookup tables, priority cuts keep only a few of the best cuts per node, which makes memory and runtime linear in circuit size while preserving quality; the authors note that similar cut-based methods exist for standard cells. The same paper extends mapping to search combined mapping-and-retiming solutions. The FPGA flow built around LUT mapping is the subject of From hardware code to bitstream.

yn1n2n3abcdChainnodes 3depth 3
1 / 5

The chain ((a·b)·c)·d as written: three AND nodes, three levels deep.

An AIG for y = a·b·c·d. Step through balancing, one cut, and two covers Boolean matching finds for that cut. Cell names are generic.Share freely with credit: ‘Figure from chipfieldguide.com’

Gate sizing and logical effort

After mapping, which cells to use is fixed but their sizes are not.

Logical effort gives the intuition. The delay of one stage, in units of a basic inverter’s delay, is:

d=gh+pd = gh + p
  • gg, the logical effort, says how much worse than an inverter the gate is at delivering current for the same input capacitance: 1 for an inverter, 4/3 for NAND2, 5/3 for NOR2.
  • hh, the electrical effort, is the load capacitance divided by the gate’s own input capacitance. A bigger gate gets a smaller hh for the same load.
  • pp, the parasitic delay, is the gate’s delay from its own internal capacitance, even with no load.

For example, a NAND2 driving four copies of itself has h=4h = 4 and p=2p = 2, so d=(4/3)⋅4+2≈7.3d = (4/3) \cdot 4 + 2 \approx 7.3 units. A path’s delay is smallest when every stage carries the same effort ghgh. This explains why covers built from NANDs beat covers built from NORs, and why a net with a large fanout wants a tree of buffers that grow gradually in size.

Gain-based (constant-delay) synthesis builds on this. Gain is the hh above: the load a stage drives divided by its input capacitance. The tool fixes each gate’s delay during mapping; as placement changes the wire loads, it resizes gates to hold those delays, so what changes is area rather than timing. Its limits are real:

  • It assumes continuous sizing, while real libraries offer a few discrete drive strengths.
  • It ignores input slope.
  • Real wires force another pass of buffering or remapping.

Mapping a continuous-sizing solution onto a discrete library can leave a sub-optimal netlist. That is why mapped netlists still go through discrete sizing against the library’s NLDM tables.

Retiming

Leiserson and Saxe model a circuit as a graph. Vertices are blocks of logic, each with a delay; edges are connections, each labeled with the number of registers on it. Retiming gives every vertex vv an integer lag r(v)r(v), the number of registers moved from its outgoing edges to its incoming edges. An edge from uu to vv that held ww registers then holds w+r(v)−r(u)w + r(v) - r(u), which must not go negative. Along any path the inner lags cancel, and the inputs and outputs keep a lag of 0, so every input-to-output path keeps its register count and the circuit’s I/O behavior. They give an efficient algorithm for the minimum clock period, show that minimizing the register count is solvable in polynomial time, and characterize optimal retiming as a mixed-integer linear program. In a full synthesis flow retiming interacts with mapping, and searching both together covers a larger space than mapping first and retiming afterward.

Equivalence checking: BDDs vs SAT

Combinational equivalence checking builds a miter: tie the two circuits’ matching inputs together, XOR each pair of matching outputs, and OR all the XORs into one output. That output can only be 1 for an input that makes the circuits disagree, so they are equivalent exactly when it is constant 0.

One way to prove that uses binary decision diagrams. A reduced ordered BDD represents a function as a decision graph that tests variables in a fixed order and shares identical subgraphs. For a given order each function has exactly one such graph (it is canonical), so once both sides are built, equivalence is a pointer comparison. The catch is size. It depends heavily on the variable order, and for the outputs of an integer multiplier it grows exponentially with word size for every order.

Modern checkers lean on SAT solvers instead, which search for an input that sets the miter output to 1 and either find one (a counterexample) or prove none exists. They:

  1. Convert the miter to an AIG and structurally hash it.
  2. Simulate random and guided input vectors to find internal nodes that look equivalent.
  3. Prove or refute those candidates in topological order with SAT, or with BDDs under small resource limits.
  4. Merge the nodes proven equal, interleaving light AIG rewriting to shrink the problem before attacking the outputs.

Structural similarity between RTL and netlist is what makes this work: each proven internal equivalence simplifies the next SAT call. ABC’s equivalence checker grew out of exactly this need to verify its own synthesis results. It also explains the pain points. Retiming and state re-encoding remove the internal correspondences these methods feed on, and deep arithmetic has few internal equivalences to exploit.

abcreference (RTL)(a·b) + cimplementationINV(AOI21)110miteragree
a
b
c
Netlist

Both give 1, so the XOR is 0. One input proves nothing; equivalence means the miter is 0 for every input.

A miter: both versions share inputs and an XOR compares their outputs. With several outputs, their XORs are ORed into one. The circuits are equivalent exactly when the miter output is constant 0.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 5 correct
  1. Q1Which input tells the synthesis tool each cell’s area, delay and power?

  2. Q2The clock period is 1.0 ns, and the SDC says set_input_delay 0.4 -clock clk on an input. How much time is left for logic inside the block, from that input to the first flip-flop?

  3. Q3A synthesis timing report shows slack −0.022 ns (VIOLATED) on a path. What does that mean?

  4. Q4What does retiming do?

  5. Q5A group of 32 flip-flops loads a new value only when a signal called en is 1. What does clock gating do to them?

Sources

Show Hide 28 sources
  1. synth – generic synthesis script (Yosys command reference)YosysHQ · Yosys documentationThe begin/coarse/fine/check steps of Yosys’s default synthesis script, including proc, fsm, alumacc, share, techmap, abc and the -flatten option.
  2. Internal cell libraryYosysHQ · Yosys documentationRTLIL represents a design with coarse-grain word-level cells and fine-grain gate-level cells.
  3. share – perform SAT-based resource sharing (Yosys command reference)YosysHQ · Yosys documentationThe share pass merges shareable resources and uses a SAT solver to decide whether two resources can be shared.
  4. FSM handlingYosysHQ · Yosys documentationFSM detection, extraction, optimization and one-hot recoding in Yosys.
  5. Mapping to cell librariesYosysHQ · Yosys documentationExample flow: dfflibmap and abc -liberty map a design onto a Liberty cell library.
  6. Technology mapping commands: abc, clockgate, dfflibmapYosysHQ · Yosys documentationDefault ABC scripts, the -D delay target (which adds retiming), -constr buffering and sizing, and clock-gating insertion with ICG cells.
  7. Chapter 8: Cell Characterization (figures), Digital VLSI Chip Design with Cadence and Synopsys CAD ToolsErik Brunvand · University of Utah (book companion site)Liberty file structure: area, leakage, lookup-table templates indexed by input transition and output load, cell_rise tables, internal power and wire-load models.
  8. SDC CommandsVerilog-to-Routing project · VTR documentationMeaning of create_clock (including virtual clocks), set_input_delay/set_output_delay, set_false_path, set_multicycle_path and set_clock_uncertainty.
  9. OpenSTA: Parallax Static Timing AnalyzerParallax Software · GitHubA gate-level static timing verifier that reads Verilog netlists, Liberty libraries and SDC, and supports false-path and multicycle exceptions.
  10. Flow variablesThe OpenROAD Project · OpenROAD-flow-scripts documentationSynthesis options in an open flow: hierarchical vs flat synthesis, keep-size threshold, ABC area/speed strategy, experimental retiming, and multibit flop clustering at placement.
  11. Gate Resizer (rsz)The OpenROAD Project · OpenROAD documentationrepair_design inserts buffers to repair max slew, max capacitance and max fanout violations and builds buffer trees for high-fanout nets.
  12. Addressing the Timing Closure Problem by Integrating Logic Optimization and Placement (Memorandum UCB/ERL M00/66)Wilsin Gosti, Sunil P. Khatri, Alberto L. Sangiovanni-Vincentelli · University of California, Berkeley, Electronics Research Laboratory · 2000Timing closure fails when synthesis and layout timing estimates disagree; wire-load models estimate length from pin count and underestimate many nets; synthesis–placement iterations may not converge; incremental placement kept close to the final one.
  13. Improving Placement under the Constant Delay ModelKolja Sulimma, Ingmar Neumann, Lukas van Ginneken, Wolfgang Kunz · Design, Automation and Test in Europe (DATE) 2002 (open proceedings archive) · 2002Constant delay model: cell delay fixed, cell area a function of load; the critical path never changes; an approximation, since library cell sizes are quantized with minimum and maximum sizes; remapping after placement steps.
  14. Power Reduction via Near-Optimal Library-Based Cell-Size SelectionMohammad Rahman, Hiran Tennakoon, Carl Sechen · Design, Automation and Test in Europe (DATE) 2011 (open proceedings archive) · 2011Rounding continuous cell sizes to the nearest library sizes destroys much of the benefit of continuous optimization.
  15. Clock gatingWikipediaClock gating removes the clock from idle logic to cut dynamic power; ICG cells use an internal latch for a glitch-free gated clock.
  16. FF-Bond: Multi-bit Flip-flop Bonding at PlacementChang-Cheng Tsai, Yiyu Shi, Guojie Luo, Iris Hui-Ru Jiang · ACM International Symposium on Physical Design (ISPD), hosted by Peking University CECA · 2013Multibit flip-flops present a smaller clock load through shared clock logic; merging is often done at or after placement.
  17. EQY: Getting StartedYosysHQ · EQY documentationFormal equivalence checking to ensure a synthesis tool has not changed a design’s function.
  18. Reference for .eqy file formatYosysHQ · EQY documentationGold vs gate designs, match rules for net names, and recode sections for FSM state encodings changed by synthesis.
  19. Lecture 6: Logical Effort (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd CollegeDelay model d = gh + p; logical effort 1 for an inverter, 4/3 for NAND2, 5/3 for NOR2.
  20. Retiming Synchronous Circuitry (MIT-LCS-TM-309)Charles E. Leiserson, James B. Saxe · MIT Laboratory for Computer Science (DSpace@MIT) · 1986Retiming as a graph problem: algorithms for minimum clock period and polynomial-time minimum register count.
  21. SIS: A System for Sequential Circuit Synthesis (Memorandum UCB/ERL M92/41)Ellen M. Sentovich, Kanwar Jit Singh, Luciano Lavagno, Cho Moon, Rajeev Murgai, Alexander Saldanha, Hamid Savoj, Paul R. Stephan, Robert K. Brayton, Alberto Sangiovanni-Vincentelli · University of California, Berkeley, Electronics Research Laboratory · 1992Multi-level synthesis with Espresso for node simplification, state assignment, tree-covering technology mapping and retiming.
  22. Espresso heuristic logic minimizerWikipediaHistory of Espresso (IBM, then UC Berkeley) and why Quine–McCluskey’s exponential growth made a heuristic necessary.
  23. Graph-Based Algorithms for Boolean Function ManipulationRandal E. Bryant · IEEE Transactions on Computers (author’s copy, Carnegie Mellon University) · 1986Ordered BDDs are canonical; size depends on variable ordering; multiplier outputs grow exponentially for every ordering.
  24. ABC: An Academic Industrial-Strength Verification ToolRobert Brayton, Alan Mishchenko · Computer Aided Verification (CAV), UC Berkeley author copy · 2010ABC’s origins in SIS and MVSIS, its AIG-based synthesis that replaced SIS scripts, and its equivalence checker.
  25. DAG-Aware AIG Rewriting: A Fresh Look at Combinational Logic SynthesisAlan Mishchenko, Satrajit Chatterjee, Robert Brayton · Design Automation Conference (DAC), UC Berkeley author copy · 2006AIG definition, 4-feasible cuts, 222 NPN classes, rewrite/refactor/balance and the resyn2 script.
  26. Technology Mapping with Boolean Matching, Supergates and ChoicesAlan Mishchenko, Satrajit Chatterjee, Robert Brayton, Xinning Wang, Timothy Kam · UC Berkeley EECS technical report · 2005Cut-based standard-cell mapping with Boolean matching, structural bias, supergates, choice nodes and area recovery.
  27. Combinational and Sequential Mapping with Priority CutsAlan Mishchenko, Sungmin Cho, Satrajit Chatterjee, Robert Brayton · International Conference on Computer-Aided Design (ICCAD), UC Berkeley author copy · 2007Keeping a few priority cuts per node instead of all K-cuts gives linear memory and runtime; sequential mapping combines mapping with retiming.
  28. Improvements to Combinational Equivalence CheckingAlan Mishchenko, Satrajit Chatterjee, Robert Brayton, Niklas Eén · International Conference on Computer-Aided Design (ICCAD), UC Berkeley author copy · 2006Miter construction, simulation plus BDD/SAT sweeping of internal equivalences, and interleaving SAT with AIG rewriting.