By now the chip is a complete drawing. Every tiny switch is placed and every wire is drawn. is the final exam before that drawing goes to the factory.
Why so careful? An app can get a fix after it ships. A chip can’t. Once the factory starts making it, a mistake means starting over. That costs months of time and millions of dollars.
So engineers run a set of separate checks on the finished drawing. Will every signal arrive on time, even on a slow chip on a hot day? Does every shape follow the factory’s rules? Is the drawing really the circuit the designers meant?1
The name is old. In the late 1960s, chip designs were drawn by hand, and the engineer who drew one checked it and signed it.1 Today each check still has an owner, who signs off when it passes.
A chip design ends as a layout: an exact drawing, layer by layer, of every transistor (the microscopic electrical switches that do the computing) and every metal wire that connects them. The factory, called a foundry, turns that drawing into a set of stencils called masks and uses them to print the chip onto silicon wafers. Handing over the final file is called . Once the masks are made, a mistake costs a new mask set and months of delay, so the layout gets one last, thorough inspection first.
That inspection is : a set of independent checks the layout must pass before tapeout. The standard list:1
- Timing. A chip is paced by a clock, a signal that ticks billions of times a second. On each tick, tiny one-bit storage cells called take in new values. checks that every signal traveling from one flip-flop to the next arrives in time for the tick, now using the delays of the real, finished wires.
- Interference. Wires running side by side can disturb each other. Signal-integrity analysis checks that this crosstalk doesn’t make signals late or cause false blips.
- Manufacturing rules. Design rule checking confirms every shape is one the factory can actually make: no wire too thin, no gap too narrow.
- Same circuit. Layout versus schematic rebuilds the circuit from the drawing and confirms it is exactly the one the designers intended.
- Power. IR-drop analysis checks that the power wiring delivers enough voltage everywhere; electromigration checks that no wire carries so much current that it wears out within the chip’s lifetime.
- Same logic. Equivalence checking proves mathematically that the final list of gates still computes exactly what the design did before layout.
Signoff comes after routing (the step that draws the wires) and before tapeout, but in practice it starts weeks earlier and runs as a loop: run every check, collect the violations (the places where a check fails), fix them with small, targeted changes called , and run again, until everything is clean or each remaining item has a formal , a written and approved reason why it is harmless.
The checks are run by specialized software. Most chip companies mix tools from several vendors rather than using one suite for everything, and foundries publish recommended flows that list the tools, versions and scripts they have evaluated.1 Who signs? Typically the timing lead, the physical verification lead, the power owner and the owners of each part of the chip each approve their part of a checklist, and a project lead signs the release to the foundry.
Signoff is the final verification of the routed layout before tapeout, in four families: timing (including crosstalk), physical verification against the foundry’s rules, power integrity, and logical equivalence. What sets it apart from the checks run during implementation is accuracy. Every earlier stage uses simplified models so it can run quickly and iterate. Signoff replaces each simplification with the most accurate model the schedule allows:
- Wires: estimated lengths become resistance and capacitance extracted from the routed geometry, once per manufacturing extreme of the metal.
- Power: an ideal supply at every cell becomes a solved model of the power grid that shows where the voltage sags.
- Conditions: a few corners become the full set of scenarios, each pairing a way the chip is used (a mode) with a combination of process, voltage and temperature (a corner).
- Variation: a flat safety margin becomes per-cell statistical variation.
- Rules: “the router followed the rules” becomes a full run of the foundry’s own rule files on the final merged layout file (GDSII or OASIS).
The placer and router contain their own timers, and they never agree exactly with the signoff timer. When they disagree, the signoff number wins.
The scope keeps growing. Device and interconnect corners multiply: three device corners, five interconnect corners, two temperatures and two voltages already make 60 combinations, and signoff at 40 and 28 nm meant hundreds of mode-corner combinations.11 Checking only corners is both pessimistic at the corners it checks and risky at the ones it skips, which is the argument for statistical methods.9 Path-based timing, the more exact mode used to close the last violations by re-timing each failing path with its own conditions, costs four times the runtime or more.8 So most of the engineering is judgment: choosing which scenarios matter, which pessimism is safe to remove, and which violations are real.
Each check reads the same finished layout and looks at something different. Tap one.
- Analysis corners in one 16 nm foundry example
- 58
- Runtime cost of path-based vs. graph-based timing
- ≥4×
- Unknowns in power-grid sparse solves
- 10⁶–10⁹
In: the finished drawing, the circuit the designers meant to build, how fast it must run, and the factory’s rule book.
Out: a stack of reports. Each one either says “clean” or lists every problem and where it is. When all of them are clean, the approved drawing goes to the factory.
| Direction | What | Typical format |
|---|---|---|
| In | The final layout, with wires and metal fill | GDSII or OASIS (plus DEF from the layout tool) |
| In | The final list of gates and connections, and a transistor-level version for LVS | Verilog; SPICE/CDL |
| In | Timing and power data for each building block, one library per PVT corner | Liberty .lib (often with LVF variation tables) |
| In | Resistance and capacitance of every wire, one file per RC corner | SPEF |
| In | Timing goals and clocks, one set per mode | SDC |
| In | The foundry’s rule decks and extraction technology files | Tool-specific files |
| In | How often each signal switches, for power analysis | VCD or SAIF |
| Out | Timing, noise, DRC, LVS, IR and EM reports and marker files | Text reports, marker databases |
| Out | ECO changes | Tcl scripts, updated Verilog/DEF |
| Out | Signed checklist and the released layout | Checklist; GDSII/OASIS |
What the formats are:
- GDSII and OASIS are the layout files themselves: long lists of polygons, each tagged with the layer it belongs to. DEF is the layout tool’s own text description of where every gate and wire sits.
- Verilog here holds the : the list of every gate in the design and which wires (nets) connect them. LVS needs a version one level lower, listing transistors instead of gates, in SPICE or CDL format.
- Liberty files describe the , the pre-designed gates the chip is built from: how much delay each adds and how much power it uses, measured at one corner per file.
- (Synopsys Design Constraints, a format read by nearly all timing tools) states the timing goals: how fast each clock ticks and when signals enter and leave the chip.
- are the factory’s rules, written so a checking tool can run them.
- VCD and SAIF are recordings from simulation of how often each signal switched, which tells the power tools where current is drawn.
Two inputs are new at this stage. A file lists, for every net, how resistance is spread along it, its capacitance, and how strongly it is coupled to each neighboring wire. Wires vary in manufacturing too, so extraction is calibrated separately for each wire extreme (each ), and the open-source extractor OpenRCX, for example, takes the number of corners to extract as an option.4 The timer then reads one set of wire values per corner. And LVS compares the layout against a SPICE netlist, because a layout contains transistors, not gates.18 For a digital design that reference is generated from the final gate-level netlist plus each standard cell’s transistor-level description.
Two details cause most input problems.
Parasitic annotation. “Annotation” means attaching each net’s extracted values to that net in the timer. Several choices hide in that step. OpenSTA’s read_spef, for example, can keep coupling capacitors as connections between two nets (needed for crosstalk analysis) or tie them to ground with a multiplying factor; accepts separate min and max SPEFs per scenario; stitches SPEFs from separately extracted blocks together at the block boundaries; and ignores some SPEF fields (it uses the library’s pin capacitances instead of the SPEF’s *I values). If a net is missing from the SPEF, the timer quietly falls back to an estimate, and the path looks fine. So every run includes report_parasitic_annotation, which lists unannotated nets.2
Version control of foundry data. The rule deck, the extraction technology file and the Liberty set are all versioned by the foundry, and an update can turn a clean layout dirty. The checklist records exactly which versions were run.
Inputs on the left, outputs on the right. Tap any box for what it holds.
Signoff is a set of separate checks. These are the main ones:
- On time. Software adds up the delay along every path a signal can take. Now it counts the real wires too. It checks that each signal arrives before the next clock tick. It does this for a slow chip and a fast one, on a hot day and a cold one. Each of these cases is called a .
- Factory rules. (DRC) measures every shape against the factory’s rule book. Are wires wide enough? Are the gaps between them big enough?15
- Same circuit. (LVS) works out the circuit from the drawing and compares it with the one the designers meant. It looks for wires that touch by mistake, or never connect.16
- Power. Checks that every part of the chip gets enough power, with no weak spots. Busy wires also must not wear out after years of use.1
Timing with real wires
A starts at one flip-flop, runs through a chain of logic gates and wires, and ends at the next flip-flop. Every gate and every wire adds a little delay. (STA) adds up the delays along every path and compares the total with the time available. It does this with arithmetic on known delays rather than by simulating the chip with test data, which is what “static” means.
Each path must pass two checks:
- : the signal must arrive a short time before the next tick, so the flip-flop can capture it. Too slow is a setup failure.
- : the new signal must not arrive so soon after a tick that it disturbs the value the flip-flop is still capturing. Too fast is a hold failure.
A small example: the clock ticks every 1.000 ns (a billion times a second). A path’s gates and wires add up to 0.850 ns, and the receiving flip-flop needs its input ready 0.050 ns before the tick. The signal is ready at 0.850 ns but needed by 0.950 ns, so there are 0.100 ns to spare. That spare time is the ; negative slack is a failure.
At signoff the wire delays come from . Every wire has resistance, which slows the current, and capacitance, a small store of charge that must be filled or emptied every time the signal changes. Both add delay. The extractor reads the finished layout and computes each wire’s resistance and capacitance, including the capacitance between neighboring wires, using a calibrated technology file generated once for each process and each manufacturing extreme. OpenROAD’s open-source extractor, OpenRCX, models each wire from its distance to the nearest neighbor and how densely the layers above and below are wired, and writes the result as .4 Commercial extractors include Synopsys StarRC and Cadence Quantus. The extracted numbers can differ a lot from the estimates used earlier: a finished wire may detour around a crowded area, hop between layers, or run for a long way next to a busy neighbor. That is why slack often moves after routing, and why only extracted timing counts for signoff.
Corners and modes
No two chips come out of the factory exactly alike, and each runs at varying voltage and temperature. A is one extreme combination of process (how fast the factory happened to make the transistors), voltage and temperature. Wires vary too, which gives . A chip also runs in different modes: normal operation, a factory test mode, a low-power mode. Each mode has its own timing goals (its own SDC file). One mode at one PVT corner with one RC corner is a scenario, and checking all of them is called analysis. Corners are paired on purpose: slow transistors with the slowest wires, and fast transistors with the fastest wires.6 The first pairing is where setup, the too-slow check, is usually worst; the second is where hold, the too-fast check, is.
What else timing checks
- . A reset signal forces flip-flops to a known starting value. Many resets act immediately, without waiting for the clock. Releasing such a reset too close to a tick can leave a flip-flop stuck between 0 and 1 for a while, a condition called , so the release is checked like a setup and a hold.13
- Electrical limits. How sharply each signal switches (its transition time, or slew), how much load each gate drives (capacitance), and how many inputs one output feeds (fanout) must stay within the library’s limits. Timing also checks gates that switch the clock off to save power, and that clock pulses are never too short.2
Safety margins for variation
Even on one chip, two identical gates are never exactly alike. Timing analysis allows for this (OCV) by assuming that when a signal needs to be early it might be a few percent late, and vice versa: for example, delays ×1.05 on the path that must be fast enough and ×0.95 on the path it is racing. One catch: the clock signals to the sending and receiving flip-flops usually share part of their route, and a shared wire can’t be fast and slow at the same moment. The fix, called , gives back the margin counted twice on the shared part. One flat percentage for the whole chip turned out to be far too cautious, which forced designs to be slower than they needed to be. Random differences partly cancel out along a long chain of gates, so flows moved to , which sets the margin by path length and distance, and then to , which gives each cell its own measured spread, stored in the cell library in a format called LVF.7
Crosstalk
On a modern chip, neighboring wires are so close that most of a wire’s capacitance is to its neighbors rather than to the ground below.10 So when a neighbor (the “aggressor”) switches while a wire (the “victim”) is switching at the same time, it pulls on it and can slow it down or speed it up.3 -aware timing adds this extra delay, called delta delay, to the setup and hold checks. An aggressor can also put a brief false blip, a glitch, on a victim that should be quiet; noise analysis checks that glitches stay too small to flip the gates they reach.1
Physical verification
(DRC) runs the foundry’s over the final layout. Rules cover minimum width, minimum spacing, how far metal must extend around each via (a vertical connection between layers), minimum area, metal density and much more.15 (LVS) works out from the layout shapes which transistors exist and which wires connect them, then compares that circuit with the intended one. Typical errors are shorts (two wires touching that shouldn’t), opens (a wire broken in two), missing devices, wrong device types and devices of the wrong size.16 Alongside them run:
- (ERC): for example, inputs connected to nothing, or outputs wired together.
- checks: long wires that could collect enough electric charge during manufacturing to damage the transistor they connect to.
- Density checks, confirming that every region of each layer, after , has the amount of metal the polishing step needs.14
A DRC run writes its violations, filed under the name of each rule, to a report database. A width or spacing violation is recorded as an edge pair, the two edges that are too close, which the layout viewer can display on top of the layout.17 Unlike timing, DRC and LVS have no slack: a rule passes or it doesn’t, and the only ways past a violation are to fix it or to get a waiver approved.
Power integrity
The power grid is a mesh of metal stripes that carries the supply from the chip’s connections to every cell. By Ohm’s law, current flowing through resistance loses voltage: . For example, 2 A flowing through 0.005 Ω of grid loses , about 1% of a 0.8 V supply. Cells that receive less voltage switch more slowly, so drop eats into timing. IR-drop analysis solves the grid as a network of resistors with each cell drawing current. Static analysis uses average currents. Dynamic analysis looks at the short, deeper dips caused by local bursts of switching, and those dips have to be kept small for the chip to meet its timing.25 Dynamic analysis is either vector-based, driven by recorded switching activity from simulation (VCD files), or vectorless, where the tool assumes switching statistically. Vectorless is faster and available before realistic test workloads exist; vector-based is more accurate for the activity it covers.25 Power-grid tools also report the current density in each segment of the grid, for the checks.22 Over years, a heavy current pushes metal atoms along a wire until it thins and breaks, so each wire’s current must stay below a limit for its width.
Equivalence
Layout rewrites the netlist: it adds buffers (gates that simply strengthen a signal), swaps gates for stronger or weaker versions, builds the clock distribution, reorders the chains used for factory testing, and adds delay cells to fix hold failures. proves, by mathematical reasoning rather than by trying inputs, that the final netlist computes the same outputs as the netlist that came out of synthesis, for every possible input.1
ECOs and the closure loop
Violations come back as small fixes, not a rerun of the whole flow. A timing swaps a gate for a stronger version, adds a buffer, or inserts a delay cell to fix a hold failure. A functional ECO patches a late logic bug. A chip is built from dozens of mask layers; the lowest ones form the transistors and are the most expensive. Once those are frozen, a metal-only ECO uses , unused gates scattered across the chip during placement, and rewires them so only the upper wiring layers change.28 Every ECO goes back through extraction, timing, DRC, LVS and equivalence, because a fix for one check can break another.
The checks, in the order they run
- Extraction. The reads the routed layout and computes, for every net, a resistor network, its capacitance to ground and its coupling capacitance to each neighbor, using a calibrated technology file generated once per process and per interconnect (RC) corner, and writes the results as .4
- Static timing. The timer reads the gate-level netlist, one Liberty library per process-voltage-temperature corner, the SPEFs and one SDC constraint file per mode. For every pin it computes when signals arrive and when they are required, then checks setup and hold on every , recovery and removal on asynchronous reset pins, the library’s limits on slew, capacitance and fanout, clock-gating checks and minimum pulse widths.2
- Signal integrity. The coupling capacitances let switching neighbors change a net’s delay or put a glitch on it. SI-aware timing adds both.
- Physical verification. DRC, LVS, ERC, antenna and density checks run on the final merged layout with the foundry’s rule deck.14
- Power integrity. The power grid is solved for static and dynamic IR drop, and wire currents are checked against electromigration limits.
- Equivalence. Logical equivalence checking (LEC) proves the final netlist matches the post-synthesis netlist.
- Fix and repeat. Violations become ECOs, and every ECO round reruns all of the above on what changed.
The sections below take each in turn and explain where the difficulty lies.
Corners, modes and dominance
A device corner is named for how fast the factory made the two transistor types, NMOS and PMOS: slow-slow, typical, fast-fast, and the skewed slow-fast and fast-slow. An interconnect corner is a manufacturing extreme of the metal. Cworst is wires printed wide and thick, so close to their neighbors that capacitance is maximal; RCworst is wires printed narrow and thin, so resistance rises and the product of resistance and capacitance is maximal. A process may define 3 to 6 device corners and 5 to 11 interconnect corners before voltage and temperature. Every independent supply domain multiplies the set: one engineer quoted in an industry article puts a typical design with two power supplies at 30 or more corners.6
Nobody signs off the raw cross product. Step by step, teams:
- Pair device and RC corners in the direction that hurts: slow devices with maximum RC, fast devices with minimum RC.6 The first pairing targets setup, the second hold.
- Add the cross pairings that catch different net types. A short net’s delay is mostly its driver charging capacitance, so it is worst at Cworst; a long net’s resistance matters, so it can be worst at RCworst.
- Run a dominance analysis: a scenario dominates if it sets the worst slack for at least one endpoint (the flip-flop input or chip output where a path ends). Scenarios that never do are dropped.
Each mode brings its own SDC. In scan-shift mode, for example, test data is shifted through long chains of flip-flops on a slow test clock; the paths are short flop-to-flop hops, so hold matters more than in functional mode. Low-power modes change which supplies, and so which libraries, apply. OpenSTA expresses this directly: read_sdc -mode loads a mode’s constraints, and define_scene binds a mode to Liberty files and named min and max SPEFs.2
Variation models and pessimism
Every margin for variation costs slack, so the question is how much margin is honest. The methods, in the order flows adopted them:
- Flat . Multiply late paths by one global derate such as 1.05 and early paths by 0.95, whatever the cell, slew or load. Simple, but the excess pessimism led to over-design and real performance limits.7
- . Look the derate up by path depth (the number of stages) and by distance across the die, which accounts for random variation cancelling out along deeper paths and reduces the pessimism of one global value.7
- with LVF. Model each cell’s local variation as a function of its own delay and load. The variation (a sigma, or standard deviation) is characterized by Monte Carlo circuit simulation at each slew and load point and stored in the library in Liberty Variation Format (LVF).7 The timer then treats each stage’s delay as a mean plus a spread and combines the spreads statistically along the path.
A worked example of why statistics helps: ten stages, each with a sigma of 1 ps. A 3-sigma margin applied stage by stage adds . If the stages vary independently, their sigmas combine by root-sum-square instead: for the path, so a 3-sigma margin of about 9.5 ps. Random variation on long paths largely cancels.
Two more sources of pessimism stay under the engineer’s control. must be on for every check: when the launch and capture clocks share a stretch of clock tree, derating that stretch late for one and early for the other counts a physically impossible difference, and CPPR credits back the shared delay × (late derate − early derate). The other is the choice between . Graph-based analysis keeps one worst arrival and slew per pin; at a gate with several inputs, the output inherits the worst input’s slew even for a path that enters through a sharper one. Path-based analysis re-times a single path with its own slews. On one published 28 nm design signed off at a 1.2 ns clock period, graph-based analysis was up to 110 ps more pessimistic than path-based on an endpoint, and path-based cost 4× or more in runtime. The usual compromise is graph-based everywhere and path-based on at least the paths that fail graph-based, to avoid fixing paths that aren’t really failing.8
Signal integrity in practice
Crosstalk analysis has a circular dependency. A net’s timing window is the interval in which it can switch. An aggressor only changes a victim’s delay if their windows overlap, and the size of the change depends on how they align; but the windows depend on delays, including the crosstalk delays being computed. Tools break the circle by iterating: compute windows, compute delta delays, update windows, and repeat until nothing changes.10 The worst alignment is the one that delays the signal at the victim’s receiver the most, not the one that maximizes the victim stage’s own delay. Current-source driver models handle that better than the classic slew-and-load lookup tables, because they don’t assume ramp-shaped inputs or a single lumped load capacitor.10
In practice crosstalk shows up after routing as new setup failures on long parallel runs, and new hold failures where an aggressor speeds up a short path. The basic fix is a routing ECO that moves neighboring wires away from the critical net.3 Glitch noise is checked separately: a glitch on a quiet net must stay too small to cross the switching threshold of the gates it reaches.1
Physical verification at scale
Rule decks no longer consist of simple width and spacing checks. Advanced processes add many contextual rules, where the allowed value depends on the surroundings (for example, a larger spacing next to a wide wire, or extra enclosure where a wire ends). These rules drive up run time and push tools to cut the layout into tiles and exploit its hierarchy to run in parallel.21 Signoff DRC runs on the final merged layout file with fill and all third-party blocks included, because a rule can fail at the boundary between two blocks that are each clean on their own.
For LVS of a digital block, the reference is a SPICE (or CDL, a SPICE dialect) netlist built from the final gate-level Verilog plus each cell’s transistor netlist. Large designs are compared hierarchically rather than flattened into one huge transistor list; the open-source Netgen, for example, supports hierarchical LVS.19 Check power and ground connectivity first: an unconnected well tap or a missing supply label on a large block fails LVS as surely as a signal short. Antenna violations are fixed with protection diodes or a jump to a higher layer near the gate.14 After any metal-only change, an XOR of the old and new layouts (a geometric comparison that shows only the shapes that differ) confirms that only the intended shapes moved.14
Power integrity and timing interact
Dynamic IR drop is the dip in supply voltage caused by local demand and switching patterns, and it has to be bounded for the circuit to meet timing.25 The two ways to analyze it trade off. Vector-based analysis replays switching activity from simulation, but a large chip needs many patterns to cover its behavior and good patterns arrive late in the project. Vectorless analysis assigns switching statistically, so it is faster and available earlier, but harder to make accurate. Run time is significant: in one study, a commercial tool spent about 2.5 hours on the dynamic analysis of a design with about two million cells, and more than four hours in total.25
The drop then has to be fed back into timing. The usual approach is a budget: power integrity work keeps the drop below a specified limit, and timing is signed off with cell delays at a worst-case voltage corner that assumes that limit. But the real grid never meets the limit exactly, so some cells see more drop than budgeted and slow down. The more exact approach annotates each cell with its own simulated supply voltage, re-times, and fixes the paths that fail, for example by resizing cells in an ECO.26 On the grid side, the fixes are a stronger grid (wider stripes, more vias) and decoupling capacitors: decap cells, which store charge right next to busy logic, inserted where the drop is highest.22
Electromigration
checks today are mostly rule-based. Step by step: estimate each wire’s current density; flag wires above a limit derived from Black’s equation, a semi-empirical formula for a wire’s mean time to failure as a function of current density and temperature; then use the Blech criterion to clear short wires. Migrating atoms pile up at one end of a wire and the resulting mechanical stress pushes back against them; if the product of current density and wire length stays below a technology-specific value, the two balance and the wire is effectively immortal. A widely used variant simply gives wires below a threshold length a higher current limit.27 Power-grid wires carry mostly one-way (DC) current. Signal and clock wires carry current that reverses each transition, which partly undoes the damage, so they are checked with an effective current density instead.27
ECO discipline
Late ECOs are limited by what can still change. Once placement is frozen, a functional patch must be mapped onto nearby spare cells. Too few spares, or long wires to reach them, create new timing and congestion problems, so patch generation itself should take spare locations into account.28 Timing ECOs come from the signoff timer, which sees every scenario, and go back to the routing tool as sizing and buffering changes. Every round ends with re-extraction, the full scenario set, physical verification and LEC. Teams track violation counts per check per round: a flat or rising count means the fixes are fighting each other across scenarios.
The checklist
A typical tapeout checklist requires:
- zero setup, hold, recovery/removal and max-transition violations in every signed scenario, with path-based analysis on any remaining marginal endpoints;
- crosstalk delta delay included and glitch analysis clean;
- DRC, LVS, ERC, antenna and density clean with the current foundry deck, plus approved ;
- static and dynamic IR drop within budget;
- EM clean at the target lifetime and temperature;
- LEC passing;
- the exact tool, deck, library and SPEF versions recorded.
Typical transistors (TT): setup slack +0.150 ns, hold slack +0.015 ns.
Opposite-direction aggressor: delta delay +18 ps on a 28 ps transition. A setup risk.
Two nets in the schematic, two in the layout, each joining the same pins. LVS reports the layout correct.
Worst drop 53 mV (6.7% of 0.8 V) at the busy block: over the 5% budget, so cells there switch more slowly than timing assumed.
The first simulation is a real-looking timing report for one path between two memory cells. Hover over or tap any line to see what it means. Times are in ns, short for nanoseconds. A nanosecond is a billionth of a second.
The bottom line, the slack, is the spare time. At first there is about a tenth of a nanosecond to spare. Turn on “OCV”, a safety margin for tiny factory differences. Watch the spare time shrink. Then speed up the clock to 0.75 and turn off “CPPR”, a correction that gives some of that margin back. The spare time drops below zero, and the path fails.
The first simulation is a timing report for one path, laid out the way the open-source timer OpenSTA prints it, with times in nanoseconds (ns, billionths of a second). Hover or tap a line to see what it means. Read it top to bottom. The clock tick reaches the sending flip-flop, acc_reg[7], after 0.412 ns of travel through the clock wiring. The data leaves at 0.530 and, after five gates, arrives at the receiving flip-flop’s input, addr_reg[15]/D, at 0.976. The receiving side starts at the next tick (0.800 ns later), adds the 0.398 ns its own clock takes to arrive, subtracts 0.050 for clock uncertainty (a small allowance for jitter in the clock) and 0.074 for the flip-flop’s setup time, giving a required time of 1.074. .
- Turn on the variation margin (OCV ±5%). With CPPR on, slack falls to +0.051 ns. Turn CPPR off and it falls to +0.029 ns.
- Set the clock period to 0.75 ns with OCV on. With CPPR on, the path barely passes (+0.001 ns); with CPPR off it fails (−0.021 ns, VIOLATED).
The report is in OpenSTA’s format. Check the arithmetic yourself. OCV multiplies the launch clock latency (0.412) and the data path (0.564 from the launch flop’s clock pin to the capture flop’s D pin) by 1.05, and the capture clock latency (0.398) by 0.95. Arrival: . Required: . Slack: +0.029. The two clock paths share 0.220 ns of tree before they split, so CPPR credits , restoring +0.051. At 0.75 ns the CPPR credit alone separates +0.001 ns from −0.021 ns: once variation is on, removing double-counted pessimism decides pass or fail. Before signing such a path off, ask two questions: what path-based analysis would report, and whether the library’s setup time carries its own variation margin (from LVF) too.
The second simulation is a tiny piece of chip drawing with two wiring layers, M1 and M2, and the plugs that join them. Press “Run DRC” and the checker lists five problems, like wires that are too thin or too close.
Sizes are in nm, short for nanometers. A human hair is about 80,000 nanometers wide.
Pick a problem to see two possible fixes. One fixes it cleanly. The other doesn’t, or it causes a new problem. Keep going until the checker says “DRC clean”.
The second simulation is a tiny piece of layout with two metal layers, M1 and M2, joined by vias (V1), with sizes in nanometers (nm). The starting layout breaks five rules, one each:
M1.W.1: every M1 wire at least 30 nm wide.M1.S.1: M1 shapes at least 30 nm apart.M2.S.1: M2 shapes at least 40 nm apart.V1.EN.1: M1 extends at least 10 nm beyond each via on every side (“enclosure”).M1.A.1: every M1 shape at least 2400 nm² in area.
Press “Run DRC” and pick a fix for each. Watch for the fix that solves one rule by breaking another, such as widening a wire until it is too close to its neighbor.
The Expert view shows each rule as a line of a KLayout-style deck. Sort the checks into three kinds: single-layer edge checks (width and spacing), two-layer checks (enclosure of V1 by M1) and polygon measurements (area). Then think about how each is computed: spacing and enclosure as queries for pairs of edges closer than a distance, area per merged polygon. That split is the basis of the sweepline algorithm in Under the hood.
In the last weeks before the chip goes to the factory, the signoff team meets every day. They look at a scoreboard like the one below, with a count of problems for each check.
Small fixes go in overnight, and every check runs again. When every row reads zero, or each leftover problem has an approved reason, the owners sign.
Signoff tools are driven by scripts written in Tcl, a simple command language. Here is an illustrative timing script in the syntax of the open-source timer OpenSTA. It loads the design and its data, defines four scenarios, and asks for reports. The file names, corner names and libraries are generic. In the library names, ss means slow transistors and ff fast ones, 0p72v means 0.72 V, and 125c and m40c mean 125 °C and −40 °C.
# sta_signoff.tcl (illustrative, OpenSTA syntax; times in ns)
read_liberty lib/stdcell_ss_0p72v_125c.lib
read_liberty lib/stdcell_ff_0p88v_m40c.lib
read_verilog results/top_final.v
link_design top
# Parasitics from extraction: one SPEF per RC corner
read_spef -name cworst results/top_cworst.spef.gz
read_spef -name rcworst results/top_rcworst.spef.gz
read_spef -name cbest results/top_cbest.spef.gz
report_parasitic_annotation -report_unannotated
# Modes: one SDC per way the chip is used
read_sdc -mode func constraints/func.sdc
read_sdc -mode scan constraints/scan_shift.sdc
# Scenes: mode x PVT corner x RC corner
define_scene func_ss_cw -mode func -liberty lib/stdcell_ss_0p72v_125c.lib -spef cworst
define_scene func_ss_rcw -mode func -liberty lib/stdcell_ss_0p72v_125c.lib -spef rcworst
define_scene func_ff_cb -mode func -liberty lib/stdcell_ff_0p88v_m40c.lib -spef cbest
define_scene scan_ff_cb -mode scan -liberty lib/stdcell_ff_0p88v_m40c.lib -spef cbest
# Flat OCV derates are SDC commands, so set them in each mode
foreach m {func scan} {
set_mode $m
set_timing_derate -late 1.05
set_timing_derate -early 0.95
}
report_checks -path_delay max -scenes {func_ss_cw func_ss_rcw} \
-format full_clock_expanded -digits 3
report_checks -path_delay min -scenes {func_ff_cb scan_ff_cb} -digits 3
report_check_types -recovery -removal -violators
report_check_types -max_slew -max_capacitance -max_fanout -violators
report_wns -max
report_tns -max
report_wns -min- 1L2Timing data for the cells at one corner: slow transistors, low voltage (0.72 V), hot (125 °C). This is where setup is worst.
- 2L3The opposite corner: fast transistors, high voltage, cold. This is where hold is worst.
- 3L4The netlist: the final list of gates and connections.
- 4L8Wire resistance and capacitance from extraction, one file per wire extreme. The capacitance between neighboring wires in these files is what makes crosstalk analysis possible.
- 5L11Lists any wire with no extracted values. Such a wire would silently use an estimate and could hide a failure.
- 6L14Each SDC file holds the timing goals for one mode. Normal operation and test mode have different clocks and rules.
- 7L18A scene (scenario) combines a mode with one corner’s libraries and wire data. Setup at slow transistors with maximum wire capacitance.
- 8L19Same transistors with the maximum-resistance wire corner: long wires can be slower here than at maximum capacitance.
- 9L20Hold at fast transistors with minimum wire capacitance.
- 10L21Test mode shifts data through chains of flip-flops with very short hops between them, a classic hold risk.
- 11L26The variation margin: paths that must be fast enough are made 5% slower, and the paths they race 5% faster.
- 12L30The worst setup paths, with every clock buffer listed so the clock timing is visible.
- 13L33Reset release checks. Easy to forget, and a failure can leave flip-flops undecided as the chip comes out of reset.
- 14L34Electrical limits: how sharply signals switch, how much load each gate drives, how many inputs each output feeds.
- 15L35Summary numbers for the daily dashboard: WNS is the worst negative slack (the single worst failure), TNS the total of all negative slack.
A real signoff run has more scenarios than this, for example hold at the slow corners too and every mode at every relevant corner. The worst setup path from a run like this looks like the report in the first simulation above.
Physical verification is driven by a rule deck. Here is an illustrative deck for the open-source KLayout tool, written in its Ruby-based DRC language, with generic layer numbers and sizes in micrometers (µm; 0.040 µm = 40 nm).17
# signoff.drc (illustrative KLayout DRC; generic layers, units in um)
source("top_final.oas")
report("Signoff DRC", "top_drc.lyrdb")
m2 = input(17, 0)
v2 = input(18, 0)
m3 = input(19, 0)
m2.width(0.040).output("M2.W.1", "M2 width >= 0.040")
m2.space(0.040).output("M2.S.1", "M2 spacing >= 0.040")
v2.enclosed(m3, 0.010).output("V2.EN.1", "M3 encloses V2 by >= 0.010")
m2.with_area(nil, 0.0030).output("M2.A.1", "M2 area >= 0.0030 um2")
m3.with_density(0.0 .. 0.20, tile_size(50.um)).output("M3.DEN.1", "M3 density >= 20% per 50 um tile")- 1L2The layout to check. Real flows usually pass the file name in from the command line so the same deck runs on any design.
- 2L3Violations go to a report file that the layout viewer can step through, marker by marker.
- 3L5Each layer is read by its number in the layout file. The foundry’s layer map says which number is which layer.
- 4L9Width check. Each violation is reported as the pair of edges that are too close across the shape.
- 5L10Spacing check between separate shapes on the same layer, also reported as edge pairs.
- 6L11A two-layer check: every via (V2) must sit inside an M3 shape with at least 10 nm of metal around it.
- 7L12Picks out M2 shapes smaller than the minimum area. Tiny islands are hard to print and can peel off.
- 8L13Density check on 50 µm squares. Squares with less than 20% metal need metal fill so the layer polishes flat.
Three artifacts from late signoff, all illustrative. First, the physical verification summary after a round of ECOs. The layout is in OASIS format (.oas), and the LVS reference is a CDL netlist.
==== Physical verification summary (illustrative) ====
Layout: top_final.oas Netlist: top_final.cdl Deck: generic_10M rev 1.4
DRC rules run: 4,812 rules violated: 3 results: 7
M2.S.3 M2 space to wide M2 (width > 0.300) >= 0.080 4
V3.EN.2 M3 end-of-line enclosure of V3 >= 0.030 2
M7.DEN.1 M7 density per 50 um window >= 20% 1
ANT rules run: 10 violations: 0
ERC floating gates: 0 wells without tap: 0 supply shorts: 0
LVS result: INCORRECT
layout schematic
instances 1,284,551 1,284,551
nets 1,302,117 1,302,118
ports 412 412
mismatched nets: 2
u_lsu/n_4471 + u_lsu/n_4472 -> one layout net (short)
short path: M3 (812.40, 1204.66) - V3 - M4 (812.40, 1204.70)- 1L2The CDL netlist is built from the final Verilog plus each cell’s transistor netlist. The deck revision goes on the checklist.
- 2L5A contextual rule: the required spacing depends on how wide the neighbor is. Often caused by an ECO wire placed next to a wide power stripe.
- 3L6End-of-line enclosure: a via needs more metal beyond the end of a wire than along its sides.
- 4L7One window short of metal after an ECO removed wiring. Rerun fill locally, then re-extract, because fill adds capacitance.
- 5L8Antenna clean: the router already inserted protection diodes or layer jumps.
- 6L11Any LVS mismatch blocks tapeout. Shorts are never waived.
- 7L14One fewer net in the layout: two intended nets became one, the signature of a short.
- 8L17The two nets the extractor merged. Both connect to the same layout polygon.
- 9L18Where they touch (coordinates in µm): an ECO via landed on the wrong M3 wire. Fix the route, then rerun DRC, LVS and STA on the affected nets.
Second, the power integrity summary for the worst functional scenario. Javg and Jrms are the average and root-mean-square current densities that EM limits are written in.
==== Power integrity summary (illustrative) ====
Scene: func_ss_0p72v_125c VDD nominal 0.720 V PDN: M1-M10 + bumps
STATIC (average current per cycle)
Total VDD current 2.41 A
Worst IR drop VDD+VSS 14.2 mV (2.0%) u_core/u_alu region
DYNAMIC, vectorless (data toggle 0.20, clock 1.0, 1.25 ns window)
Worst effective drop 68.4 mV (9.5%) u_core/u_mul/U8812
Instances above 8% budget 1,906
DYNAMIC, vector-based (VCD 18.40-18.65 us, mul_burst test)
Worst effective drop 52.7 mV (7.3%) u_core/u_mul/U9034
EM (125 C, 10-year target)
Power grid: 3 segments over Javg limit M2 rail at (1410.2, 655.0)
Signal: 1 net over Jrms limit clk_core trunk, M8
Via arrays: 0 over limit- 1L2IR drop is analyzed per scenario. Slow, hot and low voltage leaves the least timing margin. PDN is the power delivery network; bumps are the chip’s connections to the package.
- 2L6Static drop uses average current, so it tests the grid’s resistance. 2% (drop on VDD plus rise on ground) says the grid is not under-built.
- 3L8Vectorless: the tool assumes 20% of data nets and every clock net switch in each window. Broad coverage, usually pessimistic.
- 4L10The multiplier block switches all at once. Candidates for decap cells or extra grid stripes.
- 5L11Vector-based: a real stress pattern from simulation. Lower than vectorless here, but it only covers this 250 ns window.
- 6L15Power-grid EM: mostly one-way current. Widen the rail or add parallel stripes and vias.
- 7L16Signal EM on a clock trunk: high activity, current in both directions, checked with RMS current. Fix with a wider wire (a non-default routing rule) or by splitting the load.
Third, ECOs. The signoff timer can try fixes in place, with every scenario loaded, before handing them to the routing tool. These are OpenSTA netlist-edit commands; real flows replay the same edits in the place-and-route tool, then re-extract.2
# eco_fix.tcl (illustrative; OpenSTA netlist-edit commands)
# Timing ECO: upsize a driver on a failing setup path
replace_cell u_core/u_lsu/U78 NOR2_X2
# Hold ECO: insert a delay cell in front of a capture pin
make_net u_core/eco_n1
make_instance u_core/eco_hold1 DLY_X1
disconnect_pin u_core/n_2291 {u_core/u_lsu/addr_reg[3]/D}
connect_pin u_core/n_2291 u_core/eco_hold1/A
connect_pin u_core/eco_n1 u_core/eco_hold1/Z
connect_pin u_core/eco_n1 {u_core/u_lsu/addr_reg[3]/D}
# Metal-only functional ECO: use a pre-placed spare NAND2
disconnect_pin u_core/spare_tielo u_core/spare_nand2_17/A1
disconnect_pin u_core/spare_tielo u_core/spare_nand2_17/A2
disconnect_pin u_core/n_771 u_core/u_dec/U12/A1
connect_pin u_core/n_771 u_core/spare_nand2_17/A1
connect_pin u_core/n_772 u_core/spare_nand2_17/A2
make_net u_core/eco_n2
connect_pin u_core/eco_n2 u_core/spare_nand2_17/ZN
connect_pin u_core/eco_n2 u_core/u_dec/U12/A1
report_checks -path_delay min_max -through {u_core/u_lsu/addr_reg[3]/D} -digits 3
write_verilog results/top_eco1.v- 1L3Swap the gate for a stronger version of the same function and footprint. replace_cell requires matching ports.
- 2L7A new cell changes the transistor layers, so this hold fix is not metal-only. If those layers were already frozen, it would need a spare delay cell instead.
- 3L8Bus bits go in braces so Tcl doesn’t treat [3] as a command.
- 4L14A spare cell’s inputs are tied to a constant so it sits idle. Free them first.
- 5L16Break the old connection and route the signal through the spare gate instead.
- 6L21The spare’s output now drives the original pin: a logic patch made with only metal and via changes.
- 7L23The new nets have no extracted wiring yet, so this check is optimistic until the router builds the wires and extraction reruns.
- 8L24The ECO netlist goes to place and route, then to equivalence checking against the pre-ECO netlist plus the intended change.
Round 1: the first full signoff run. Each night, ECOs go in and every check reruns. Step through the rounds.
- Fixing one thing breaks another. Speeding up a slow path can make a fast path too fast, or push a wire too close to its neighbor. That’s why every fix goes through all the checks again.
- Testing the wrong conditions. A chip checked only at room temperature can fail in a freezer or a hot car.
- Too many problems, too late. If thousands of problems show up in the last week, small fixes can’t keep up. So teams start running these checks early.
- Waving problems through. Approving a problem without truly looking at it can leave a real bug in the finished chip.
- Missing wire data. Wires with no extracted resistance and capacitance are timed with estimates and can look fine when they aren’t. Caught by an annotation report on every run.2
- Wrong timing goals. The constraint files can tell the tool to ignore a path (a “false path”) because it never matters in practice. A rule written too broadly hides real failures, and so does a clock that was never defined. Caught by reviewing the constraints and by listing any path endpoint that has no timing goal at all.
- Hold checked at the wrong corner. Hold is usually worst at the fast corner, but differences in clock arrival can cause hold failures at slow corners too. Check hold in every scenario.
- Crosstalk surprises. Timing that passed before crosstalk analysis fails once it is included. Run it as soon as the wiring is stable.
- Power connections in LVS. Regions of silicon not tied to the supply, or power pins left unlabeled on large pre-built blocks, make LVS fail on its first run. Check power connections early.
- Fill added after timing. Metal fill adds capacitance. Extract with the fill in place (or a model of it) before the final timing run.
- Corner pruning that drops the dominant scenario. A scenario pruned early because it never set the worst slack can become the worst one after an ECO changes which paths are critical. Re-run the dominance analysis after large ECO rounds.
- Pessimism removal used as a fix. Path-based analysis and POCV are legitimate, but hundreds of endpoints passing by a few picoseconds only after path-based re-timing are still a yield risk. Track how much slack came from removing pessimism and how much from real fixes.
- Correlation gaps. The place-and-route tool’s built-in timer and the signoff timer disagree, so the optimizer works on the wrong paths. Compare the two on a sample of endpoints and adjust the implementation flow’s margins until they agree.
- IR drop left out of timing. A path with positive slack at nominal voltage fails inside a hotspot where the supply dips 9%. Either budget the drop in the timing run or re-time hotspot paths with their simulated supply.25
- Signal EM on clocks. Like other signal wires, clock wires carry current in both directions and get the signal-wire EM treatment.27 But a clock switches every cycle, far more often than a typical data net, so checking it with the switching rate assumed for data nets understates its current and misses violations.
- Spare cells in the wrong place. A metal-only ECO fails when the nearest suitable spare is far away and the long wire to it breaks timing.28
- Unreviewed waivers. DRC waivers carried over from a previous project or deck version, or reset timing exceptions copied without review, are where bugs escape to silicon.
The top path fails setup by 20 ps. Pick a setup fix and watch all three checks.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
STA as longest paths on a DAG
Cut the netlist at the flip-flops and the logic between them becomes a directed acyclic graph (DAG): nodes are pins, edges carry cell and wire delays, and no path loops back on itself. Timing is then two passes over that graph:3
- Forward. Visit nodes in topological order (every node after all of its inputs). Each node’s actual arrival time (AAT) is the maximum, over its fan-ins, of the fan-in’s arrival plus the edge delay.
- Backward. Visit nodes in reverse order. Each node’s required arrival time (RAT) is the minimum, over its fan-outs, of the fan-out’s required time minus the edge delay.
- Slack. At every node, . Negative slack anywhere marks a violation.
Both passes touch each edge once, so the work grows linearly with the size of the graph, which is why STA scales to billions of pins where circuit simulation cannot.3 Production timers keep separate rise and fall, early and late values at every node, propagate slews alongside arrivals, and update only the affected part of the graph after an ECO.
Pins are nodes, cell and wire delays are edges (ns). Step through the forward pass, then the backward pass.
Graph-based vs. path-based
Storing one value per node is the source of pessimism. At a gate with several inputs, graph-based analysis propagates the worst input slew to the output even if the reported path enters through a different pin with a sharper slew. A slower slew makes every following stage slower, so the error accumulates along the path and into the capture flop’s setup time. Path-based analysis takes the paths GBA reports and recomputes each with its own slews (and its own depth for AOCV), so its slack is never worse than GBA’s. Enumerating paths is the cost: in Kahng et al.’s measurements PBA ran 4× or more slower, which is why it runs only on violators.8
Statistical STA
Visweswariah’s method writes every delay, arrival time and slack in one canonical first-order form: a mean, plus a sensitivity to each global source of variation (each modeled as a unit Gaussian), plus an independent random term. Adding two such quantities is easy: the result is in the same form. Taking their max is not, because the max of two Gaussians isn’t Gaussian. So the method computes the tightness probability (the chance that one arrival is the later one) and the mean and variance of the max, then writes the result back in canonical form with that mean and variance matched. The result is approximate, but the traversal stays linear in circuit size and in the number of variation sources, and every slack comes out with sensitivities that show what makes a path fragile.9 POCV, the method signoff flows adopted, is a much simpler statistical model: each cell carries its own characterized variation, stored in the library as LVF tables.7 Treating the stages as independent, their sigmas combine by root-sum-square instead of adding linearly, as in the worked example under How it works.
Delay models
The classic library model, , describes each timing arc as tables of delay and output slew indexed by input slew and a single lumped output capacitance. Two of its assumptions break down at advanced nodes: the input is not a clean ramp, and a resistive wire does not behave like a single capacitor. Current-source models instead describe the driver’s output current as a function of the voltages on its input and output pins, so the tool can compute the actual output waveform, even with a noisy input, into the real wire network. Not assuming ramp inputs or lumped loads is also what makes them a better basis for crosstalk analysis.10
RC extraction: patterns vs. field solvers
A field solver computes capacitance from Maxwell’s equations on the real 3-D geometry. It is the most accurate method, but far too expensive for the billions of wire segments on a chip. Full-chip extractors instead cut the layout into small geometries and match each against a pattern library. The library is built once per process: enumerate sample geometries, solve each with a field solver, and fit formulas or tables to the results. The 2.5-D variant approximates 3-D effects by combining 2-D cross-sections taken in two directions.5 OpenRCX’s Extraction Rules file is this kind of calibration: capacitance and resistance as functions of the distance to the nearest wire and the wiring density above and below.4 The weaknesses are pattern coverage, since each new process node adds structures, and table size, which is why researchers are fitting neural-network capacitance models to field-solver data.5
LVS as graph matching
LVS has to decide whether two netlists with millions of devices are the same circuit, even though nothing in the layout says which device corresponds to which. The classic Gemini approach:20
- Model each netlist as a bipartite graph: device vertices on one side, net vertices on the other, an edge for each device terminal. Making nets explicit vertices cuts the edge count for a net with terminals from to .
- Give every vertex an initial label from invariants that don’t depend on naming, such as device type and net degree, and group vertices with equal labels into partitions.
- Refine iteratively: each vertex’s new label combines its old label with its neighbors’ labels. Integer labels stand in for exact ones and give the same result with very high probability.
- If both graphs refine identically into partitions of one vertex each, the pairing is the match.
Interchangeable terminals, such as a transistor’s source and drain, are grouped into equivalence classes so they don’t cause false mismatches. Symmetric circuits leave partitions that won’t split; the matcher guesses a pairing, keeps relabeling, and backtracks if the guess fails. Memory arrays are the slow case.20 Netgen’s author notes that, as far as he knows, all LVS tools use this same class-partitioning approach.19 KLayout’s LVS shows the pipeline before the comparison: derive device regions with Boolean layer operations, extract devices, declare which layers connect to which, simplify (merge transistors drawn as several fingers, drop floating nets), then compare against the SPICE reference.18
DRC as polygon geometry
A DRC engine runs computational-geometry operations on polygons and edges.21 Width and space checks return edge pairs, the two edges that violate the distance; Boolean operations derive helper layers (for example “M2 wider than 0.3 µm” for a width-dependent spacing rule); and the engine can work flat, tiled for memory and threads, or hierarchically so a cell checked once covers every instance.17
The core spacing check is a sweepline problem. For horizontal edges, move a vertical line across the layout from left to right, keeping the edges it currently crosses in a balanced search tree ordered by . When an edge enters, query the tree for edges within the spacing value and report each as a violation. With a heap of event points this runs in time for edges and violations, which is optimal. X-Check parallelizes this on GPUs and reports average speedups of 45× on space checks and 61× on enclosing checks against a multi-threaded CPU checker.21 Hierarchy is the other lever: a rule that can be decided inside a cell is checked once per cell, and only interactions across cell boundaries need context.
IR drop as a sparse linear system
For DC analysis the power network is a resistor network. Writing Kirchhoff’s current law at every node (nodal analysis) gives , where is the conductance matrix, the unknown node voltages and the cell currents and supply connections.23 The size is the problem: these systems commonly have to unknowns, heading toward for detailed full-chip models.24 And the currents are large: 100 W at 1 V means 100 A, and 100 A through just 1 mΩ already drops 100 mV, 10% of the supply.
Each node connects only to a few neighbors, so is sparse, and DC power-grid analysis typically gives a symmetric positive-definite matrix. Two solver families exploit that: Cholesky factorization, which factors once into triangular pieces, and the conjugate gradient method, which converges iteratively without factoring. Power-grid matrices are often badly conditioned (conductances vary over a huge range), so conjugate gradient needs a good preconditioner to converge quickly.24 Dynamic analysis adds capacitance (decap, cell and wire) and gives
Backward Euler with time step turns each step into23
The matrix on the left stays fixed for a fixed step, so one factorization serves every time point and only the right-hand side changes. The hard part becomes the currents themselves: vector-based analysis derives them from simulated switching activity, vectorless analysis from estimated activity.25
Q1Why does the final timing check read the parasitics file (SPEF) produced after routing, instead of the wire estimates used earlier?
Q2A hold check fails when a signal arrives too early, before the flip-flop has finished capturing the previous value. In which conditions is a hold failure most likely?
Q3LVS counts one fewer wire (net) in the layout than in the intended circuit, and two of the intended wires both match the same layout wire. What is the most likely cause?
Q4What makes an ECO “metal-only”?
Q5What does a recovery check verify?
Sources
Show Hide 28 sources
- Signoff (electronic design automation)Signoff check types (LVS, DRC, formal post-layout vs. post-synthesis netlist, IR drop, SI glitches vs. gate thresholds, STA, EM); origin of the term in hand-checked, signed-off vellum drawings in the late 1960s; mixing vendors’ tools; foundry-recommended tools, versions and scripts.
- OpenSTA Commandsread_spef, read_sdc -mode, define_scene (MCMM), set_timing_derate, report_checks, report_check_types (recovery, removal, max slew/cap/fanout), replace_cell and netlist edits.
- VLSI Physical Design: From Graph Partitioning to Timing Closure, Chapter 8 slides (Timing Closure)STA on a DAG: AAT as max over fan-in, RAT as min over fan-out, slack = RAT − AAT; coupled parallel wires can slow down or speed up transitions; push neighboring wires away from critical nets.
- OpenRCX parasitic extraction documentationExtraction from coupling distance and track-density context, calibrated Extraction Rules file per process corner, multi-corner extraction, write_spef.
- CNN-Cap: Effective Convolutional Neural Network Based Capacitance Models for Full-Chip Parasitic ExtractionPattern matching (2.5-D, pre-characterized with field solvers) is the full-chip method in commercial extractors; field solvers are most accurate but too costly full-chip.
- Process Corner ExplosionKept as a fallback for practitioner-quoted figures with no open primary source: 3–6 device corners and 5–11 interconnect corners; a typical two-supply design runs 30+ corners; SS with maxRC and FF with minRC pairing.
- Fast and Accurate Library Generation Leveraging Deep Learning for OCV ModellingA single global derate caused excessive pessimism and over-design; AOCV accounts for cancellation of random variation with path depth and distance; POCV models local variation per cell as a function of delay and load, characterized by Monte Carlo and represented in LVF at multiple slew/load points.
- Using Machine Learning to Predict Path-Based Slack from Graph-Based Timing AnalysisGBA propagates worst-case slew and is pessimistic; PBA uses path-specific slews at 4X or more runtime; GBA is always more pessimistic than PBA.
- System and method for statistical timing analysis of digital circuits (US Patent 7,428,716)Corner-based timing is risky and pessimistic at the same time; first-order delay model with a deterministic part, sensitivities to zero-mean unit-variance global sources and an independent random part; arrival tightness probabilities; max re-expressed in standard form with mean and variance matched; runtime linear in graph size and number of variation sources; sensitivities reported to flag fragile paths.
- Worst-Case Aggressor-Victim Alignment with Current-Source Driver ModelsCoupling capacitance dominates wire capacitance; delay noise depends on aggressor alignment within timing windows; noise and windows solved iteratively; current-source models vs. lookup tables.
- The Future of SignoffMoving to 40 or 28 nm means signoff methodologies must cope with hundreds of mode-corner combinations, complicated OCV deratings and growing best-worst margins.
- “Unobserved Corner” Prediction: Reducing Timing Analysis Effort for Faster Design Convergence in Advanced-Node DesignMuch design turnaround time goes into timing analysis across PVT corners; a 1M-instance foundry 16 nm example with 58 analysis corners (10 analyzed, 48 predicted).
- Asynchronous reset synchronization and distribution – challenges and solutionsKept as a fallback (no open primary source found). Reset release must meet setup/hold-like conditions (recovery and removal) or the flop can go metastable; reset distribution resembles CTS.
- Physical verificationDRC incl. density for CMP, LVS, XOR after a metal spin, antenna checks and fixes (diode, layer hop), ERC items.
- Design rule checkingRule deck/runset; width, spacing, enclosure, area, density and antenna rules; tools: Calibre, IC Validator, Pegasus, KLayout, Magic.
- Layout versus schematicDevice and connectivity extraction, netlist comparison; shorts, opens, component and parameter mismatches; KLayout and Netgen as free tools.
- DRC RunsetsRunset syntax: input, width, space, enclosed, with_area, output; checks return edge pairs; deep (hierarchical) and tiled modes.
- LVS IntroductionDevice extraction, connect statements for connectivity, netlist simplification and comparison against a SPICE reference.
- Netgen 1.5: netlist comparison (LVS) and format manipulationNetgen performs LVS; its author notes all LVS tools he knows use the same class-partitioning algorithm; hierarchical LVS since 1.4.35.
- SubGemini: Identifying SubCircuits using a Fast Subgraph Isomorphism AlgorithmGemini’s bipartite device/net graph, invariant-based partitioning, iterative relabeling from neighbors, guessing on symmetric ambiguity.
- X-Check: GPU-Accelerated Design Rule Checking via Parallel Sweepline AlgorithmsDRC as computational geometry on polygons and edges; explosion of contextual rules; tiling and hierarchy for parallelism; O(n log n + k) sweepline distance check.
- PDNSim: power grid analysis documentationOpen-source static IR analyzer; reports worst current density and per-segment current in the power grid for EM; insert_decap adds decap cells where IR drop is highest.
- SRAM-PG: Power Delivery Network Benchmarks from SRAM CircuitsDC PDN analysis as Gx = b; transient analysis via modified nodal analysis and backward Euler, (G + C/h)x(t + h) = (C/h)x(t) + b(t + h).
- A Technical Survey of Sparse Linear Solvers in Electronic Design AutomationPower integrity analysis solves large sparse systems; DC power-grid analysis often yields symmetric positive-definite matrices solved by Cholesky or conjugate gradient; power-grid matrices are often severely ill-conditioned, so preconditioning is critical.
- PowerNet: Transferable Dynamic IR Drop Estimation via Maximum Convolutional Neural NetworkDynamic IR drop definition; vectorless vs. VCD-based analysis and their trade-offs; commercial analysis of a ~2M-cell design took over 4 hours.
- IR-Aware ECO Timing Optimization Using Reinforcement LearningTiming is usually signed off at a worst-case voltage corner assuming the IR-drop limit; the real grid does not meet the limit exactly, causing timing failures; per-cell IR-annotated supply voltage and gate-sizing ECOs fix them.
- Invited: Toward Accurate, Large-scale Electromigration Analysis and Optimization in Integrated SystemsDC EM in power grids vs. AC EM with recovery in signal and clock wires; Black’s equation and the Blech criterion in rule-based EM checks.
- Resource-Aware Functional ECO Patch GenerationFunctional vs. timing ECO; metal-only ECO after placement is frozen; spare cells spread during placement and rewired; scarce or distant spares cause timing and congestion problems.