Earlier chapters built logic gates from transistors, one at a time. A real chip has billions of transistors, and nobody draws them one by one. Instead, designers pick from a catalog of ready-made building blocks called .
Each cell is a small group of transistors that does one simple job. It might flip a signal, combine two signals or remember one bit. Its shape is already drawn. Its speed and power use are already measured and stored in a file that the design software reads.
This chapter links the two halves of the site. Before it, everything is about single transistors. The Design Flow guide starts from this catalog of cells.
The earlier chapters built an inverter, NAND and NOR gates and memory cells out of individual transistors. A chip with millions of gates can’t be designed that way. Instead, a team designs each kind of gate once, draws its layout (the exact shapes of every layer that the factory will print), measures its behavior, and packages the result as a . A cell library is a collection of these, usually with several versions of each function.1 The open SKY130 high-density library has 428 of them.14
Each cell ships as several files, called views, each for a different tool:
- Layout (GDS): the full drawing of every layer, for the factory.
- Abstract (): just the outline, pin positions and blocked areas, for the tools that place and wire cells.
- Timing and power (,
.lib): the cell’s logic function, size, and tables of delay and power, for synthesis and timing analysis. - Transistor netlist (SPICE): the list of transistors and how they connect, for detailed simulation and for checking the layout.
- Verilog model: the cell’s behavior, for logic simulation.
This chapter uses a real, open library: sky130_fd_sc_hd, the high-density standard cells for SkyWater’s 130 nm process (SKY130). Its name reads as process (sky130), source (fd, the SkyWater foundry), type (sc, digital standard cells) and library (hd, high density).2 Every cell in it is published with exactly the views above.4
The library is where the design flow begins: synthesis builds the design out of these cells, and placement arranges them in rows.
A standard-cell library is the contract between device physics and the digital flow. Below it, everything is transistors, models and layout rules. Above it, synthesis, placement, routing and static timing analysis (STA) see only abstractions: a Boolean function, an outline with pins, and tables of delay, slew, power and timing checks.1 None of those tools simulates a transistor; whatever the earlier chapters showed about drive current, capacitance and threshold voltage reaches them only through the library.
Concretely, each cell is a bundle of views: GDS (the layout), (the abstract for place-and-route), a SPICE or CDL netlist (for characterization and for layout-versus-schematic, or LVS, checking that the drawn layout matches the circuit), Verilog (for simulation), and one model per process/voltage/temperature (PVT) corner. The open SKY130 high-density library, sky130_fd_sc_hd, publishes all of them per cell and per drive strength, including Liberty data at sixteen corners.413
This chapter walks through that contract in the order it’s built:
- The cell architecture (height in tracks, rails, sites, pin access).
- The layout of an inverter and a flip-flop in SKY130, with ASAP7 for a FinFET comparison.
- Characterization in SPICE.
- How a non-linear delay model (NLDM) table is read, including interpolation and extrapolation.
- How synthesis and placement consume the result.
The numbers come straight from the open library files, so you can check every one.
Layout (GDS): the exact shapes on every layer. Only the factory, and the final layout checks, need all of it.
Picture the chip’s logic area as a parking lot painted with long, straight lanes.
- Same height, different widths. Every cell is exactly one lane tall, so cells line up end to end in rows. A simple cell is narrow. A complicated one is wide.
- Power and ground strips. Each cell has a power strip along one edge and a ground strip along the other. Side by side, the strips join up, so every cell gets power with no extra wiring.
- Small, medium and large. The same job comes in several strengths. A small cell can drive one neighbor. A big one can push a signal to dozens of others quickly.
- Helper cells. Some cells do no logic at all. They fill gaps or cap the ends of rows.
Fixed height, variable width
Every cell in a library has the same height, so cells can be placed side by side in rows.1 Widths vary, but always in whole multiples of a small unit called a site. In sky130_fd_sc_hd the site is 0.46 µm wide and 2.72 µm tall.3 The chip’s core area is tiled with of these sites, and a cell must start on a site boundary.
Rails and wells
Along the top edge of each cell runs a strip of metal for the supply (VDD, called VPWR in SKY130); along the bottom runs ground (VGND). These line up when cells touch, which designers call abutting, so a whole row shares one continuous pair of rails. Inside the cell, the PMOS transistors (which connect the output to the supply) sit in the top half, inside a region of n-type silicon called the n-well; the NMOS transistors (which connect it to ground) sit in the bottom half.17 Neighboring rows are flipped upside down so that they share a rail.23
Height in tracks
Cell height is usually counted in : how many horizontal wires on the lowest metal layer fit across the cell. A wiring track is the width of one wire plus the gap to its neighbor.17 SKY130 offers three heights:3
| Libraries | Height | Intended for |
|---|---|---|
| hd, hdll | 9 tracks (2.72 µm) | High density; hdll adds lower leakage |
| ls, ms, hs, lp | 11 tracks (3.33 µm) | Low, medium and high speed, and low power |
| hvl | 14 tracks (4.07 µm) | 5 V high-voltage logic |
A shorter cell packs more logic into the same area, but leaves less room for wide transistors and for wires inside the cell.
Pins
A cell’s pins are small shapes on its lowest wiring layers where the router must connect a wire. They have to be reachable without breaking any spacing rule, a property called . In the SKY130 inverter, input A and output Y are drawn on a layer called li1 (local interconnect, a thin conducting layer just below the first metal layer), while the power pins are on both li1 and the first metal layer, met1.8
Families and drive strengths
Each logic function comes in several . The SKY130 inverter exists as inv_1, inv_2, inv_4, inv_6, inv_8, inv_12 and inv_16; the suffix grows with the size of the transistors.5 Bigger versions drive heavy loads faster but are wider and are harder to drive themselves. The same library offers NAND, NOR, AND-OR-invert and other combinations, adders, latches, flip-flops (one-bit memories that update on a clock edge), clock-gating cells and scan versions of the flip-flops for testing.314
Physical-only cells
Some cells have no logic function:
- tie the n-well and the substrate to the supplies. SKY130 hd has one called
tapvpwrvgnd.7 Without enough of them, a parasitic structure in the silicon can switch on and short the supply to ground, a failure called latch-up.26 - finish the ends of each row.
- Fill and cells fill leftover gaps; decaps also add a little local charge storage on the supply.
Cell architecture
A standard cell is a fixed-height, site-aligned layout with its supply rails on the top and bottom edges.1 The textbook template is:
- PMOS in an n-well along the VDD rail and NMOS along the GND rail.
- Horizontal diffusion strips with vertical polysilicon gates crossing them.
- Every gate designed to abut its neighbors without breaking a design rule.
That fixed template is what makes automated placement possible.17 In sky130_fd_sc_hd the unit site (unithd) is 0.46 × 2.72 µm, and the library is quoted as 9 met1 tracks tall.3
The inverter’s LEF shows how the rails work:8
- VPWR and VGND are 0.48 µm met1 rectangles centered on the cell’s top and bottom edges.
- Half of each rail belongs to the cell and half to the row above or below.
- Both pins carry
SHAPE ABUTMENT, telling the router that abutment connects them.
Flows create rows from the site definition (OpenROAD’s initialize_floorplan -site) and alternate their orientation between N and FS (flipped), so each rail is shared by two rows.23
Track height is the main density knob
Height in tracks sets three things: how wide the devices can be, how much wiring fits inside the cell, and how many access points each pin has. SKY130 shows the trade within one process:3
- hd (9 tracks, 2.72 µm) targets a routed density of 160 kGates/mm² or better and, compared with the 11-track ls library, lower dynamic power with comparable timing and leakage.
- hdll uses the same height and claims 5–10× lower leakage than hd.
- ls, ms and hs are 11 tracks (3.33 µm), graded by speed. hs cells are drop-in compatible with ms cells of the same function and drive strength.
Advanced nodes push height down to 7.5 or 6 tracks. The open ASAP7 FinFET kit ships both a 7.5-track and a 6-track library.15 At those heights each pin has very few legal access points, and becomes a placement constraint, not just a routing detail.
Pins and local interconnect
In SKY130 hd, signal pins sit on li1, a local-interconnect layer below met1. In the inverter, A is a single 0.33 × 0.24 µm rectangle and Y is a set of three rectangles.8 The router reaches them through a via from met1. Libraries keep pins on the lowest layers so that upper layers stay free for routing over the cells.
Families, footprints and swaps
Each function comes in a family of . In SKY130 that’s inv_1 through inv_16.5 Liberty groups them with a shared cell_footprint ("sky130_fd_sc_hd__inv" for every inverter), and the library header sets in_place_swap_mode : match_footprint.14 Optimizers use this to resize a cell without disturbing its neighbors. OpenROAD’s resizer has a -match_cell_footprint option for exactly that.25
Taps, endcaps and fill
The textbook template puts well and substrate contacts in every gate.17 SKY130 hd provides a separate substrate-and-well tap cell, tapvpwrvgnd, with no logic pins.7 The flow places it in a checkerboard at a fixed distance (OpenROAD’s tapcell -distance), along with at row ends, before any logic is placed.24 The spacing exists because taps keep the well and substrate resistance low enough that the parasitic p-n-p-n structure (a thyristor) formed by the wells and diffusions can’t switch on and latch up; SkyWater’s documentation lists latch-up protection among the tap grid’s jobs.263
Cells of one height sit on a grid of 0.46 µm sites. Add cells and watch the rails join.
Here are two real cells from SkyWater’s open library, which anyone can download and look at.
- The inverter flips a signal: 1 becomes 0, and 0 becomes 1. It has just two transistors. About fifty of them side by side would be as wide as a human hair.89
- The flip-flop remembers one bit. It only updates that bit when the chip’s clock ticks, a steady beat that keeps every part in step. It has 24 transistors and is more than five times wider than the inverter. Yet it is exactly the same height.1011
Newer factories make far smaller transistors. In a teaching kit for a modern factory process, the same inverter takes up about 86 times less room.16
The inverter: sky130_fd_sc_hd__inv_1
The smallest inverter has one input pin, A, and one output pin, Y.5 Its transistor netlist has exactly two devices: an NMOS 0.65 µm wide and a PMOS 1.0 µm wide, both with a 0.15 µm channel length.9 The PMOS is wider because holes move more slowly than electrons, so it needs extra width to pull the output up about as fast as the NMOS pulls it down (see CMOS logic).
The cell outline is 1.38 × 2.72 µm, which is three sites wide.8 Inside, the layout follows the classic pattern: a horizontal strip of p-type diffusion (the PMOS) in the top half, a strip of n-type diffusion (the NMOS) in the bottom half, and one vertical polysilicon line crossing both to form the shared gate.17
The flip-flop: sky130_fd_sc_hd__dfxtp_1
A stores one bit. When the clock rises, it copies its input D to its output Q; the rest of the time Q holds its value. It is built from two latches in a row, called master and slave, that take turns being transparent.17 SKY130 calls this cell a “delay flop”, with pins CLK, D and Q.6 It has 24 transistors and measures 7.36 × 2.72 µm, or 16 sites.1110
Side by side
| Cell | Transistors | Size (µm) | Area (µm²) | Input capacitance |
|---|---|---|---|---|
inv_1 | 2 | 1.38 × 2.72 | 3.75 | 2.3 fF |
inv_4 | — | 2.30 × 2.72 | 6.26 | 9.0 fF |
nand2_1 | — | 1.38 × 2.72 | 3.75 | 2.3 fF per input |
dfxtp_1 | 24 | 7.36 × 2.72 | 20.02 | 1.7 fF (D), 1.8 fF (CLK) |
Areas and input capacitances are from the library’s Liberty file; sizes are area divided by the 2.72 µm height.14 A femtofarad (fF) is farad. Input capacitance is how much charge a pin soaks up per volt, and it is what the driving cell has to charge.
The same cells at 7 nm
ASAP7 is an open “predictive” kit that models a 7 nm FinFET process for teaching and research (see Shrinking).15 Its 7.5-track inverter, INVx1, is 0.162 × 0.27 µm (0.044 µm²), and its flip-flop DFFHQNx1 is 1.08 × 0.27 µm (0.29 µm²).16 That’s about 86 times less area than the SKY130 inverter and about 69 times less than the SKY130 flip-flop. The cells are still one row tall, and the flip-flop is still several inverters wide.
inv_1 from netlist to abstract
The netlist is two devices:9
sky130_fd_pr__nfet_01v8, , .sky130_fd_pr__pfet_01v8_hvt, , . This is a high-threshold PMOS: the library mixes threshold flavors inside the cell, not only between libraries.
The cell is 3 sites (1.38 µm) wide. A is a single li1 rectangle and Y is three. The LEF also records ANTENNAGATEAREA on A and ANTENNADIFFAREA on Y for the router’s antenna checks, which stop long wires from collecting enough charge during manufacturing to damage a thin gate.8 Here is the abstract, trimmed:
MACRO sky130_fd_sc_hd__inv_1
CLASS CORE ;
SIZE 1.380000 BY 2.720000 ;
SYMMETRY X Y R90 ;
SITE unithd ;
PIN A
ANTENNAGATEAREA 0.247500 ;
DIRECTION INPUT ;
PORT
LAYER li1 ;
RECT 0.320000 1.075000 0.650000 1.315000 ;
END
END A
PIN VGND
DIRECTION INOUT ;
SHAPE ABUTMENT ;
USE GROUND ;
PORT
LAYER met1 ;
RECT 0.000000 -0.240000 1.380000 0.240000 ;
END
END VGND
END sky130_fd_sc_hd__inv_1- 1L2CORE: an ordinary standard cell that goes in a row (macros, pads and endcaps have other classes).
- 2L3Width 1.38 µm = 3 × 0.46 µm sites; height is the row height.
- 3L4The cell may be mirrored, which is how alternate rows are flipped to share rails.
- 4L7Gate area on this pin, used by the router’s antenna-effect checks.
- 5L11The only access shape for A: a 0.33 × 0.24 µm rectangle on local interconnect (li1).
- 6L16ABUTMENT: this pin connects simply by placing cells edge to edge.
- 7L20A 0.48 µm met1 rail centered on y = 0: half belongs to the row below.
dfxtp_1: a master-slave flop in 16 sites
dfxtp_1 is a positive-edge D flip-flop built from two latches.176 Its netlist has 24 transistors. CLK feeds an inverter that drives a second inverter, producing the complementary internal clocks, and the output stage drives Q with its own 0.65 µm NMOS and 1.0 µm PMOS, the same sizes as the inverter’s.11 The cell is 7.36 µm (16 sites) wide, 5.3 inverter-widths, with CLK marked USE CLOCK in the LEF so clock-tree tools treat it as a sink.10
Its Liberty model holds more than delay tables. There are setup and hold constraint tables indexed by clock transition and data transition, and a minimum-pulse-width check on CLK. At the fastest edges (10 ps on both pins), the typical-corner values are:14
- Setup: 50.8 ps for a rising D and 103.3 ps for a falling D.
- Hold: −28.5 ps and −46.8 ps.
A negative hold means D may change slightly before the clock edge and still be captured correctly. The numbers move with slew: at a 1.5 ns clock edge and a fast D, setup for a rising D drops to −45 ps.
The same cells in ASAP7
ASAP7’s 7.5-track library uses a 0.054 × 0.270 µm site. INVx1 is 3 sites (0.162 µm) and DFFHQNx1 is 20 sites (1.08 µm).16 The comparison with SKY130:
- Inverter area: 3.75 µm² versus 0.044 µm², about 86 times smaller.
- Flip-flop area: 20.0 µm² versus 0.29 µm², about 69 times smaller.
- Site width: 0.46 µm versus 54 nm; row height: 2.72 µm versus 0.27 µm.
The flop shrank less than the inverter: area scaling is uneven across cell types.
inv_1 and dfxtp_1 at true scale: one row tall, 3 and 16 sites wide. Tap a cell.
Before anyone can use a cell, someone has to measure its speed and power. This isn’t done on real chips. It’s done with a very detailed computer simulation of the cell’s transistors.
The simulation asks one question over and over: “If the input changes this fast, and the output has this much to push, how long does the output take to switch?” What the output has to push is called its load. The simulation tries many mixes, from a fast input to a slow one and from a light load to a heavy one. The answers go into a table.
Then it does it all again for a hot chip and a cold one, a strong power supply and a weak one, and transistors that came out of the factory a bit fast or a bit slow. Each set of conditions gets its own catalog file. This whole job is called .
turns a cell’s transistor netlist into the numbers in its Liberty file. It is done by simulating the netlist in , a circuit simulator that uses the process kit’s transistor models.18 Open-source characterizers such as LibreCell’s lctime take three inputs (a Liberty template, the transistor models and the cell’s netlist), sweep a list of output loads and input transition times, and write a Liberty file.19
What gets measured
- Function and pins: the logic function (for the inverter,
!A) and each input pin’s capacitance. - Delay: for each (input pin to output pin), the time from the input crossing half the supply voltage to the output crossing half the supply voltage.
- Output transition: how long the output edge takes, measured from 20% to 80% of the supply in SKY130.14
- Power: energy drawn inside the cell on each switch, and leakage while it sits still.
- Timing checks for flip-flops: and times, the windows around the clock edge in which D must stay steady.
The grid
Delay depends mostly on two things: the input (how sharp the incoming edge is) and the (how much capacitance the cell must charge).18 So each arc is simulated on a grid. In SKY130 hd the grid is 7 × 7:
- Input transitions from 10 ps to 1.5 ns.
- Loads from 0.5 fF up to a limit that depends on the cell: 181 fF for
inv_1, 563 fF forinv_4.14
For the one-arc inverter that is 49 points for each output direction, rise and fall, or 98 measurements of delay and transition per corner.
Corners
Transistors vary from wafer to wafer, supply voltage sags, and temperature changes, so each (process, voltage, temperature) gets its own Liberty file. SKY130 hd ships sixteen, named like tt_025C_1v80 (typical transistors, 25 °C, 1.80 V) or ss_100C_1v60 (slow transistors, 100 °C, 1.60 V).13
is SPICE run at scale. The inputs are the cell’s extracted netlist (with layout parasitics), the foundry’s device models, and a template that fixes the measurement setup: the index grid, the thresholds, and the input driver.1819 The outputs, per , are:
- The cell’s function, recognized from the transistor network or declared.
- Every , with
timing_sense(the inverter’s A→Y isnegative_unate) andtiming_type(combinational,rising_edge,setup_rising,hold_rising,min_pulse_width). cell_rise,cell_fall,rise_transitionandfall_transitiontables.- Internal-power tables, state-dependent leakage, and pin capacitances (separately for rise and fall).
Open characterizers (LibreCell’s lctime, CharLib, libretto) do the same job in the open.18
Measurement conventions
The SKY130 hd header fixes the measurement setup:14
- Delay is measured from the input’s 50% crossing to the output’s 50% crossing (
input_threshold_pctandoutput_threshold_pctare 50). - Transition is measured from 20% to 80% of the swing, with
slew_derate_from_library1.0. - The input stimulus is a linear ramp (
driver_model : "ramp"). Itsnormalized_driver_waveformconfirms the convention: the 10 ps index is a ramp that reaches full swing in 16.7 ps, because 20–80% is 0.6 of the swing.
Getting these conventions wrong when mixing libraries silently shifts every delay. Units are time_unit : "1ns" and capacitive_load_unit(1, "pf").
Grid design
Every SKY130 hd delay table uses del_1_7_7: input_net_transition on index_1 and total_output_net_capacitance on index_2.
- Transition axis: 10 ps to 1.5 ns in steps of about 2.3×.
- Load axis: per cell, from 0.5 fF up to the pin’s
max_capacitance.
For inv_1 the last load index is 0.181284 pF, and the output’s max_capacitance is the same number.12 The grid stops exactly at the design-rule limits, so a netlist that obeys max_transition and max_capacitance never needs to read past the top edge. Geometric spacing puts more points where the surface bends most, at small slews and loads. A rule of thumb from open characterization practice: choose the grid so that linear interpolation between points is good enough, and choose the bounds so that extrapolation is never needed.18
Constraint arcs
Setup and hold are found by bisection: move the data edge toward the clock until the output fails (or until clock-to-Q degrades by a set percentage; SKY130’s files record violation_delay_degrade_pct : "10"). Each point needs many simulations, which is why constraint tables are coarser, 3 × 3 in dfxtp_1.14
Corners and data volume
SKY130 hd publishes sixteen corners, from ff_n40C_1v95 to ss_100C_1v40, two of them with CCS noise data. The typical-corner Liberty file alone is about 12 MB for 428 cells.1314 Libraries offered in several threshold-voltage flavors multiply that again, which is why characterization is a compute-farm job and why a model update means re-characterizing everything.
Input transition 10 ps, load 0.5 fF: delay 20.3 ps (50% to 50%), output transition 14.5 ps (20% to 80%).
Open the catalog page for the small inverter, and you find a grid of numbers: 7 rows by 7 columns.
- Going down, the input changes more slowly.
- Going across, the output has more to push.
- Each box holds the delay for that mix: how long the output takes to switch.
The smallest numbers sit in the top-left corner. The biggest sit in the bottom-right.
Real circuits rarely land exactly on a box. So the design tool looks at the four nearest boxes and blends them, trusting the closest ones most. It’s like guessing the temperature at your house from the four nearest weather stations.
Past the edge of the table, the tool can only stretch the numbers outward, and that guess can be quite wrong. Good designs stay inside the table.
A Liberty delay table is an table (non-linear delay model): a grid of delays indexed by input transition (rows) and output load (columns).18 Here are the first four rows and columns of inv_1’s rising-output delay at the typical corner, converted to picoseconds and femtofarads:12
| transition ↓ / load → | 0.5 fF | 1.3 fF | 3.6 fF | 9.5 fF |
|---|---|---|---|---|
| 10 ps | 20.3 | 25.6 | 38.9 | 72.8 |
| 23 ps | 25.5 | 30.6 | 43.9 | 78.3 |
| 53 ps | 37.4 | 43.6 | 56.6 | 90.3 |
| 122 ps | 54.7 | 64.8 | 84.7 | 121.1 |
Read it like a road atlas’s distance chart. With a 53 ps input edge and a 9.5 fF load, the output rises 90.3 ps after the input changes. A 9.5 fF load is roughly what four other inv_1 inputs present (2.3 fF each), plus no wire at all.
In between: interpolation
Suppose the input transition is 100 ps and the load is 5 fF. Neither is in the table. The tool picks the two rows that bracket 100 ps (53 and 122 ps) and the two columns that bracket 5 fF (3.6 and 9.5 fF), and blends the four entries 56.6, 90.3, 84.7 and 121.1 ps. Each entry is weighted by how close the point is to it. This is .21 The answer is 84.2 ps.
Chaining cells
Each cell has a second table for its output transition, read the same way. That number becomes the input transition of the next cell down the path. A timing analyzer walks every path in the design this way, cell by cell, adding up delays.
Off the edge
Outside the grid, the tool keeps using the last two rows or columns and extends the straight line through them. This is extrapolation, and it can be far off because the real surface curves.1820 Libraries set the top edge at the cell’s limits on input transition and load, so a well-built design never goes there.
An arc is four 2-D tables:
cell_riseandcell_fallgive delay as a function of input transition and output load.rise_transitionandfall_transitiongive the output slew over the same grid.
For a negative-unate arc like the inverter’s, a rising input produces cell_fall, and STA picks the table that matches the output edge.12
The lookup
OpenSTA’s implementation is short enough to state in full.20
- Find such that by bisection, and the same way for load . Clamp to , so a point past the last index uses the last interval.
- Compute and .
- Return the weighted sum of the four corners:
Interpolation is linear in the raw index values, even though the grid is spaced geometrically.
Worked example
inv_1 cell_rise, , :12
- Rows 53.13 and 122.47 ps give .
- Columns 3.565 and 9.521 fF give .
- The four entries are , , and .
weights: (1-u)(1-v) = 0.246 u(1-v) = 0.513
(1-u)v = 0.078 uv = 0.163
delay = 0.246*56.6 + 0.513*84.7 + 0.078*90.3 + 0.163*121.1
= 13.9 + 43.5 + 7.0 + 19.7
= 84.2 ps
slew = same weights on rise_transition -> 59.6 ps- 1L1The weights always sum to 1. The heaviest sits on the entry nearest the point.
- 2L6The output slew is read with the same u and v; it becomes the next stage’s input slew.
What the surface looks like
- Along a row, delay is close to linear in load at large loads. The slope is an effective drive resistance: about 5.6 ps/fF (5.6 kΩ) for
inv_1rising, and 1.8 ps/fF forinv_4.14 - Down a column, slow inputs add delay because the cell spends longer partly on.
- Negative entries are real.
inv_4’scell_fallreaches −53.7 ps at a 1.5 ns input and 0.5 fF. A strong, lightly loaded inverter pulls its output through 50% before a slow input gets there. The same effect bendsinv_1’scell_fallback down in its last row at the lightest load (31.7 ps, below the 46.7 ps above it). - Some tables are nearly flat.
dfxtp_1’s CLK→Qrise_transitionbarely changes with clock slew (23.3–23.6 ps in every row at 0.5 fF), because Q is driven by an internal stage, not directly by CLK.
Extrapolation
Because the index is clamped and or is not, a point outside the grid gets a straight-line continuation of the edge interval. The open-source characterization literature warns that extrapolation can produce large errors depending on the grid bounds.18 In SKY130 hd the bounds are the design-rule limits, so the high side is also a max_transition or max_capacitance violation, which the flow repairs by buffering or upsizing rather than trusting the number.25 The low side (below 10 ps or 0.5 fF) extrapolates over a short distance and is usually benign.
Stage 3 (inv_1): input 29.2 ps, load 50.0 fF → delay 187 ps, output transition 226 ps. Path delay: 242 ps.
Once the catalog exists, the rest of chip design builds on it.
- Synthesis turns the designer’s plan into a shopping list of cells. It picks small cells where speed doesn’t matter and strong ones where it does.
- Placement drops every cell into the rows, like books onto shelves. Cells that talk to each other go close together.
- Wiring and checking connect the cells, then use the delay tables to make sure every signal arrives on time.
The whole Design Flow guide follows that journey step by step.
Every step of the design flow reads the library, each through a different view.1
- Synthesis turns register-transfer level code (a hardware description) into a of library cells. In the open-source tool Yosys,
dfflibmap -libertymaps the design’s registers to the library’s flip-flops andabc -libertymaps the remaining logic to its gates. This step is called .22 The tool reads area and delay from Liberty to choose between, say, aninv_1and aninv_4. - Floorplanning builds the rows from the library’s site.23 Then tap cells and endcaps go in.24
- Placement puts each cell on legal sites using the LEF outlines, and timing-driven placement reads Liberty to keep critical paths short.
- Optimization resizes cells and inserts buffers to fix nets whose transition or load exceeds the library’s limits.25
- Routing connects the LEF pins, and timing signoff runs static timing analysis with the Liberty tables and the extracted wire capacitances.
Each view feeds a different consumer, and a mismatch between views is a classic source of late surprises.
| Stage | Reads | Uses it for |
|---|---|---|
| Synthesis | Liberty | (dfflibmap, abc -liberty in Yosys22), sizing, area and power estimates, design-rule limits |
| Floorplan | Tech LEF, cell LEF | Rows from the site, tracks from layer pitches23, tap and endcap insertion24 |
| Placement | LEF, Liberty | Legal sites, orientation, pin positions for wirelength, timing-driven weights |
| Optimization | Liberty | Footprint-matched resizing, buffering for max slew, max capacitance and max fanout25 |
| Routing | LEF | Pin shapes, obstructions, antenna areas |
| STA and signoff | Liberty (all corners) | Arc delays and slews from NLDM (or CCS) tables, setup and hold checks |
| LVS and GDS merge | SPICE/CDL, GDS | Comparing layout to netlist, and replacing abstracts with full layouts |
Three practical consequences follow. The dont_use list (cells kept out of synthesis) is a library-quality decision, not a design one. A library update must keep LEF, Liberty and GDS in step, or timing and layout describe different cells. And the corner set you load at synthesis decides what “meets timing” means long before signoff.
Synthesis maps RTL to library cells (Yosys: dfflibmap and abc -liberty), choosing sizes from Liberty area and delay.
This is a real timing table from SkyWater’s open library. Move the two sliders. One makes the input change faster or slower. The other makes the load lighter or heavier. The four outlined boxes are the measured numbers the tool blends to get your answer.
Try the big inverter and see how much less the load slows it down. Then push a slider all the way right to fall off the table. Below, a simple drawing shows the cell sitting in its row.
The table is cell_rise or cell_fall for three SKY130 hd cells at the typical corner. Drag input transition and output load; the four bracketing entries light up and the readout shows the interpolated delay and output transition. Start at 100 ps and 5 fF (the example above), then:
- Switch to
inv_4and sweep the load: delay grows far more slowly. - Switch to
dfxtp_1: clock-to-Q delay starts near 270 ps even at the lightest load. - Push either slider past the last row or column to see the extrapolation warning.
The lookup reproduces OpenSTA’s: bisection on each axis, index clamped to the edge interval, bilinear weights shown with their arithmetic. Things to check:
- Compare the load slope of
inv_1andinv_4along the top row. - Find
inv_4fall’s negative corner at 1.5 ns and 0.5 fF. - Push load past
max_capacitanceand watch exceed 1 while the extrapolated delay keeps climbing linearly. - Read the output transition readout for
dfxtp_1at different clock slews.
- SKY130 hd row height
- 2.72 µm (9 tracks)
- Cells in the typical hd Liberty file
- 428
- PVT corners published for hd
- 16
- Points per NLDM table
- 7 × 7 = 49
What these numbers mean:
- Every cell in SkyWater’s most compact library is the same height, about 30 times thinner than a hair.3
- The catalog lists 428 cells.14 It is measured all over again for 16 different mixes of heat, power and factory luck.13
- With little to push, the small inverter switches in tens of trillionths of a second. Pushed to its limit, it takes more than a billionth of a second.12
- The same inverter in a modern teaching kit takes about 86 times less room.16
| Quantity | Value | Source |
|---|---|---|
| hd site | 0.46 × 2.72 µm, 9 met1 tracks | 3 |
| hd routed density | ≥ 160 kGates/mm² | 3 |
inv_1 area, input capacitance | 3.75 µm², 2.30 fF | 12 |
inv_1 max load | 181 fF | 12 |
inv_1 rise delay, 10 ps input, 0.5 fF | 20.3 ps | 12 |
inv_1 rise delay, 1.5 ns input, 181 fF | 1,697 ps | 12 |
dfxtp_1 CLK→Q rise, 53 ps clock, 3.4 fF | 307 ps | 14 |
dfxtp_1 area | 20.02 µm² (5.3 × inv_1) | 14 |
| ASAP7 INVx1, DFFHQNx1 area | 0.044 µm², 0.29 µm² | 16 |
The range in the delay rows is the point of the table: the same two transistors take anywhere from 20 ps to 1.7 ns depending on what drives them and what they drive. A single “gate delay” number would be wrong by nearly two orders of magnitude at one end or the other.
Corner spread
Interpolating inv_1 cell_rise at 50 ps and 5 fF in three corners gives:134
- 51 ps at
ff_n40C_1v95(fast transistors, cold, high voltage). - 63 ps at
tt_025C_1v80(typical). - 88 ps at
ss_100C_1v60(slow transistors, hot, low voltage).
That’s a 1.7× spread for one gate. Note that the slow corner uses different index values (its load axis runs to 410 fF), so comparisons need interpolation, not entry-by-entry reading.
Drive strength in numbers
Along the 10 ps row, inv_1’s cell_rise slope between its 25 and 68 fF columns is 5.6 ps/fF. inv_4’s between 54 and 175 fF is 1.8 ps/fF, 3.1 times lower.14 The price:
- 3.9× the input capacitance (9.0 versus 2.3 fF).
- 1.7× the area (6.26 versus 3.75 µm²).
The area grows less than the drive partly because the smallest sizes don’t fill their footprint: inv_2 fits in the same 3.75 µm² as inv_1.14
- Big catalog or small one. More kinds of cells give the tools more choices and better chips. But every cell must be drawn, measured and checked, for every set of conditions.
- Short cells or tall cells. Shorter cells pack more logic into the chip. Taller cells have room for stronger transistors and are easier to wire up.
- Small or strong. A strong cell pushes heavy loads fast. But it is bigger, uses more power, and is harder for the cell before it to push.
The most common mistake is using a cell outside the range it was measured for. The tools then have to guess, and the guess can be badly off.
- Density versus drive. SKY130 offers 9-track libraries for density and 11-track ones for speed.3 A shorter cell has narrower transistors and fewer wiring tracks inside, so it is weaker and harder to connect to.
- Sizing. Upsizing a cell speeds up its own output but adds input capacitance that slows the cell driving it.
inv_4loads its driver about four times as much asinv_1.14 The right size depends on the whole path (see Speed and power). - Speed versus leakage. Libraries built with lower-threshold transistors switch faster and leak more. SKY130’s hdll library claims 5–10× less leakage than hd at the same height.3
- Accuracy versus effort. More grid points, more corners and more detailed models cost characterization time and analysis runtime.
What goes wrong
- Extrapolation. Nets with too much load or too slow an edge fall off the table. The number is a guess, and the design breaks a library limit. Flows fix these by buffering or upsizing.25
- Wrong corner. Signing off with only the typical corner hides slow-corner failures.
- Pin access. Dense placement can leave a pin that the router can’t reach legally, which only shows up late, in routing.
- Track height. Fewer tracks buy density and lower capacitance at the cost of drive, internal routing and pin access points. SKY130 sells the trade as hd (9 tracks) versus hs (11).3 ASAP7 offers 7.5 and 6 tracks.15 Mixing heights in one block needs row planning, because the rails must line up.
- Library richness. Complex cells (AOI/OAI, muxes, multibit flops) let the mapper absorb logic levels, but each one adds characterization, layout and verification cost, and some are poor for pin access and end up on a
dont_uselist. - Grid resolution. NLDM error is largest where the surface curves most, at small slews and loads and wherever a crossover such as the negative-delay region sits between grid points. Denser grids cut interpolation error and multiply simulation count.18
- Model fidelity. NLDM gives a delay and a ramp slew into a lumped capacitance. Resistive wires shield far-end capacitance, and real waveforms aren’t ramps, so advanced nodes add that store output current.18 The SKY130 hd release includes CCS noise data at two corners.13 The cost is larger files and slower analysis.
- Design-rule limits as guard rails. Because SKY130 ends its grid at max_transition and max_capacitance, those limits double as “no extrapolation” fences. A library whose grid ends short of its limits lets a clean design extrapolate silently.12
Failure modes teams check for
- View mismatches. A LEF pin moved without a GDS update, or a Liberty function that disagrees with the netlist, shows up as layout-versus-schematic failures or wrong simulation.
- Corner gaps. Setup is usually worst at slow corners and hold at fast ones; signing off with too few corners hides failures that silicon will find.
- Threshold mixing. Merging libraries characterized with different delay or slew thresholds without a derate shifts every lookup.
- Characterization holes. A missing arc (an unlisted when-condition or a non-unate arc treated as unate) creates timing that STA never checks.
hd and hdll, 9 tracks (2.72 µm): 11 rows in 30 µm. The densest choice, at 160 kGates/mm² or better; hdll claims 5–10× lower leakage at the same height.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Bilinear interpolation, precisely
With the bracketing indices and , and the fractions and defined above, the lookup is2120
Expanding gives , with
Along a line of constant or constant the function is linear; along a diagonal it is quadratic.21 The twist term is the cross-sensitivity: how much more load costs when the input is slow. For the worked example it is , small compared with the first-order terms, which is why the surface looks nearly planar inside one cell.
Interpolation error
Linear interpolation between points apart has error bounded by roughly times the second derivative along that axis. Delay against load is close to linear at large loads, so the load-axis error is small there. It is largest at small slews and loads, where the cell’s intrinsic delay and the input slope effect dominate and the surface curves. Geometric spacing of the index points (about 2.3× steps on SKY130’s slew axis, 2.7× on inv_1’s load axis) concentrates points where curvature is highest.12 The weights, however, are linear in the raw index values. A point halfway between two load indices on a log scale (their geometric mean) gets a weight well below 0.5: for inv_1’s 3.57–9.52 fF interval, the geometric midpoint 5.83 fF gives .
Extrapolation, and why the clamp matters
OpenSTA’s index search returns for any value beyond the last index, and 0 for any value below the first.20 The fraction is not clamped, so or extends the bilinear surface of the edge cell. For inv_1 at 50 ps and 300 fF (1.65× max_capacitance) this gives about 1.73 ns. Whether that is optimistic or pessimistic depends on the curvature the grid never saw, which is why libraries end the grid at the design-rule limits, and why flows treat a limit violation as something to repair rather than a number to trust.25
Effective drive resistance from a table
Reading the table as recovers the simple RC model from the Speed and power chapter. Taking the slope between two large-load columns along the fastest row:
inv_1rise: , so .inv_1fall: . The NMOS pull-down is stronger than the PMOS pull-up even with the PMOS 1.5× wider.inv_4rise: .
14 The ratios tell you how the devices scale inside the cell, and the intercepts tell you how much of the delay is the input slope rather than the load.
inv_1 cell_rise, 10 ps row: C = 50.3 fF in [25.4, 67.9], v = 0.59 → 303 ps. Slope 5.62 ps/fF, so R_eff ≈ 5.6 kΩ, intercept d₀ ≈ 19.9 ps.
Why NLDM is not enough at advanced nodes
NLDM assumes the output drives a lumped capacitor and that a ramp with the tabulated slew is a good stand-in for the real waveform. Both fail when wire resistance is significant: the driver sees an effective capacitance smaller than the total, and the far-end waveform has a long tail. Current-source models store the output current as a function of time per slew/load point, so the tool can integrate it into the actual RC network.18 The cost of is tables of waveforms instead of scalars and a slower delay calculator.
Characterizing a constraint
Setup time is found by a search per (clock slew, data slew) point:
- Fix both slews.
- Launch D at a time offset from the clock edge, and bisect on that offset until the flop either fails to capture or its CLK→Q delay grows past a set degradation. SKY130 uses 10%, recorded as
violation_delay_degrade_pct.14 - Repeat for hold, searching from the other side.
Each table entry needs a whole search of transient simulations, which is why constraint tables are coarse (3 × 3 in dfxtp_1). The degradation threshold is a library-wide choice: a tighter one leaves more margin, a looser one more usable time.
Q1Which two numbers does a timing tool use to look up a cell’s delay in an NLDM table?
Q2In the SKY130 high-density library, an inverter is 1.38 µm wide and a flip-flop is 7.36 µm wide. What do they have in common?
Q3Where do the numbers in a Liberty timing table come from?
Q4Your point lies between table entries. What does the timing tool do?
Sources
Show Hide 26 sources
- Standard cellA standard cell is a group of transistors and interconnect giving a logic or storage function; fixed height lets cells sit in rows; Liberty carries timing and power; synthesis and place-and-route use the library.
- LibrariesLibrary naming: sky130 (process), fd (SkyWater foundry), sc (digital standard cells), hd (library name).
- SkyWater Foundry Provided Standard Cell LibrariesSeven libraries in three cell heights: hd and hdll at 9 met1 tracks (0.46 × 2.72 µm site), ls/ms/hs/lp at 11 tracks (0.48 × 3.33 µm), hvl at 14 tracks; hd density 160 kGates/mm² or better; hdll 5–10× lower leakage than hd; cells have no taps, and a separate staggered tap grid provides latch-up protection.
- skywater-pdk-libs-sky130_fd_sc_hd: cells/invEach inverter drive strength ships as GDS layout, LEF, SPICE/CDL netlist, Verilog model and per-corner Liberty data.
- sky130_fd_sc_hd__inv: InverterInverter with input A and output Y in drive strengths 1, 2, 4, 6, 8, 12 and 16; a schematic, and a GDS layout image of each drive strength.
- sky130_fd_sc_hd__dfxtp: Delay flop, single outputD flip-flop with inputs CLK and D and output Q, with GDS layout images for drive strengths 1, 2 and 4.
- sky130_fd_sc_hd__tapvpwrvgnd: Substrate and well tap cellA cell with no logic inputs or outputs that ties the substrate and well to the supplies.
- sky130_fd_sc_hd__inv_1.lefSIZE 1.38 BY 2.72, SITE unithd; signal pins A and Y on li1; VPWR and VGND on li1 and met1 with SHAPE ABUTMENT; met1 rails 0.48 µm wide centered on the cell edges.
- sky130_fd_sc_hd__inv_1.spiceTwo transistors: nfet_01v8 W = 0.65, L = 0.15 and pfet_01v8_hvt W = 1.0, L = 0.15 (µm).
- sky130_fd_sc_hd__dfxtp_1.lefSIZE 7.36 BY 2.72, SITE unithd; pins CLK (USE CLOCK), D, Q, VPWR, VGND.
- sky130_fd_sc_hd__dfxtp_1.spiceThe flip-flop’s transistor netlist: 24 transistors (X0–X23).
- sky130_fd_sc_hd__inv_1__tt_025C_1v80.lib.jsonTypical-corner Liberty data for inv_1: area 3.7536, input capacitance 0.002302 pF, max_capacitance 0.181284 pF, 7 × 7 cell_rise, cell_fall and transition tables.
- skywater-pdk-libs-sky130_fd_sc_hd: timingSixteen PVT corners from ff_n40C_1v95 to ss_100C_1v40, including two with CCS noise data.
- sky130_fd_sc_hd__tt_025C_1v80.lib (OpenROAD-flow-scripts platform copy)Full typical-corner Liberty file: 428 cells; ns and pF units; 50% delay thresholds and 20–80% slew thresholds; del_1_7_7 template; in_place_swap_mode match_footprint; inv_4 and dfxtp_1 timing, setup and hold tables.
- ASAP7 PDK and Cell LibrariesThe open ASAP7 7 nm FinFET predictive PDK with 7.5-track and 6-track standard-cell libraries.
- asap7sc7p5t_28_R_1x_220121a.lefASAP7 7.5-track cells: site 0.054 × 0.270 µm; INVx1 0.162 × 0.27 µm; DFFHQNx1 1.08 × 0.27 µm.
- Lecture 1: Circuits & Layout (CMOS VLSI Design, 4th ed. slides)Standard-cell methodology (abutting VDD and GND, nMOS at bottom and pMOS at top, well and substrate contacts), horizontal diffusion and vertical poly, wiring tracks, and the master-slave D flip-flop.
- Standard-cell characterizationTiming and power come from simulating the transistor netlist on a grid of input edge rate and output load; linear interpolation between grid points; extrapolation can cause large errors; CCS stores output current; open-source characterizers.
- lctime (LibreCell characterization kit)An open characterizer: takes a Liberty template, SPICE models and a cell’s extracted netlist, recognizes the cell’s function, sweeps output loads and input slews, and writes Liberty.
- TableModel.cc (OpenSTA)Two-dimensional table lookup by bilinear interpolation; out-of-range values use the edge interval, which extrapolates linearly.
- Bilinear interpolationRepeated linear interpolation in two variables; the four-corner weighted formula; linear along each axis but quadratic along other lines.
- Mapping to cell librariesdfflibmap -liberty maps flip-flops and latches to library cells; abc -liberty maps the remaining logic.
- Initialize Floorplaninitialize_floorplan builds rows from the -site argument; make_tracks adds routing tracks from the technology LEF.
- TapcellInserts endcaps at row ends and tap cells in a checkerboard with a set distance between them.
- Gate Resizerrepair_design buffers nets to fix max slew, max capacitance and max fanout and resizes gates; -match_cell_footprint obeys Liberty footprints when swapping.
- Latch-upParasitic thyristor in CMOS; substrate and well taps lower the resistances that let it trigger.