Design Flow · Stage 10 of 13 · Back end

Clock tree synthesis

The clock is the chip’s heartbeat, a signal that ticks millions or billions of times a second. This step delivers each tick everywhere at almost the same moment.

CTS builds a branching network of small amplifiers that carries the clock signal to every flip-flop, so each gets the tick at nearly the same moment and with a sharp edge. Timing is then re-checked with the real clock delays.

Skew groups, useful skew, non-default rules and shielding, multi-source and mesh topologies, clock gating, OCV and CPPR, and post-CTS hold fixing.

Builds The clock network

A chip is packed with tiny memory cells called flip-flops. Each one holds a single 1 or 0. They all work to the beat of one clock, the way a band follows its drummer. On every tick, each flip-flop grabs a new value.

The tick has to reach all of them at nearly the same moment. If some hear it early and others late, they grab values at the wrong time. Then the chip’s answers come out wrong.

One wire can’t do the job. Picture one coach yelling “go” to every runner in a city marathon. The runners at the back hear it faint and late. builds a branching network of small amplifiers instead. Each one repeats the tick loud and clear. The gap in when the tick arrives is called , and keeping it tiny is the main job.

The clock ticks billions of times a second, so this network is one of the most power-hungry parts of the chip.

A digital chip keeps its state in millions of , one-bit memory cells. A single signal, switching steadily between 0 and 1, tells them all when to update: on each rising edge (each “tick”), every flip-flop stores the value at its input. Between ticks, results flow from one flip-flop through a chain of logic gates to the next. Such a flip-flop-to-flip-flop route is called a timing path, and each must finish within one clock period.

By this stage the design is a , a list of every cell and the wires between them, and placement has given every cell a position on the chip. Until now, the timing tools have used an that reaches every flip-flop instantly. (CTS) replaces that assumption with a real network. It groups the (every pin the clock must reach), places , small amplifiers, near each group, and wires them back to the clock’s entry point. The goals are three numbers: small (the difference in arrival time between flip-flops that pass data to each other), short (travel time from source to flip-flop), and a sharp edge, or short . Power is the fourth concern.

Why a tree instead of one wire? Three reasons. Load: a block can have tens of thousands of sinks, far more than one driver can switch quickly. Wire delay: a wire’s resistance and capacitance both grow with its length, and its delay is roughly resistance times the capacitance it must charge, so an unbuffered wire’s delay grows roughly with the square of its length. Edge shape: a weak driver on a huge load produces a slow, smeared edge, which makes flip-flops slower and wastes power. Once the tree exists, the timing tools switch to a , calculated from the real delays. Real skew then shows up in every timing path, and a repair step fixes the “hold” problems it creates (explained below).

Until CTS, the clock exists only as a line in the timing constraints (the SDC file): a period and some estimated delays. CTS turns it into the largest and busiest net on the die. The tool works through each clock network in two phases. First it builds: it groups the sinks, buffers each group, and connects the groups to the root, aiming at targets for skew, latency and maximum transition set per (a set of sinks that must be balanced against each other). Then it optimizes with the real, propagated clock delays: it deliberately shifts some clock arrivals to rescue slow paths (useful skew, or concurrent clock and data optimization, CCD), inserts delay cells to fix hold failures, and resizes cells to fix slow edges and save power.

One consequence is easy to miss. Timing signoff assumes that any two clock paths may differ slightly because of manufacturing variation (), but it gives credit back for the stretch of tree that two flip-flops share, since a shared wire can’t be both fast and slow at once (). So the shape of the tree, specifically where each critical pair of flip-flops splits apart, decides how much margin signoff takes away.

Power is the other half of the objective. Every clock node switches on every cycle, while a data node switches only when its value changes. Clock networks have taken well over a quarter of total power in some designs, and 40–44% across several generations of the DEC Alpha processor. Tools: Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler, and the open-source OpenROAD TritonCTS, which first clusters sinks with a capacity-limited k-means method and then connects the clusters with a generalized H-tree.

clock pinArrival time at each flip-flop0time →
Clock network

One driver, one unbuffered wire threaded past 16 clock sinks. Send a clock edge to see when each one switches.

The same 16 clock sinks fed by one unbuffered wire and by a buffered H-tree. Arrival times are illustrative, not from a real chip.Share freely with credit: ‘Figure from chipfieldguide.com’
Clock network share of chip power, DEC Alpha 21164
40%
Clock network share of chip power, DEC Alpha 21264
44%
Extra power of a clock mesh vs. a conventional tree (one estimate)
+20–40%
Sinks per leaf clock buffer in this page’s example (illustrative)
~20

Power figures from Friedman and Toyama.

In: the chip with every part already placed, how fast the clock must tick, and a list of amplifier parts the tool may use.

Out: the same chip with the clock network added, plus a report card. It says how evenly the tick arrives and how much power the network will use.

Chip tools pass work along in a handful of standard text formats. The ones CTS reads and writes:

DirectionWhatTypical format
InThe placed design: every cell, its position, and the wires between cellsVerilog plus (a layout description) or an OpenROAD database file (.odb)
InTiming and power data for every cell in the library, including clock buffers, clock inverters and clock-gating cells (.lib)
InPhysical facts: cell outlines and pins, the metal wiring layers, and any special wide or widely spaced wiring rules
InThe timing constraints: each clock’s period, clocks derived from it, safety margins, estimated edge rates (Synopsys Design Constraints)
InCTS settings: which buffers to use, targets for skew, latency and transition, wiring rulesTool script in Tcl (e.g. options to clock_tree_synthesis)
OutThe design with the clock buffers added, placed in legal positionsVerilog, DEF or .odb
OutConstraints updated to use the real (propagated) clock delaysSDC
OutReports on skew, latency, transition, buffer count and timingText reports and logs

Beyond the basic files, three kinds of input shape the result. Scenarios. A chip is checked in several modes (normal operation, test) at several corners (slow, hot, low-voltage silicon; fast, cold silicon; and so on), a setup called multi-mode multi-corner. Which corner the tree is balanced in matters, because a tree balanced in one can be skewed in another. Balancing rules. Skew-group definitions say which sinks must align, and exclude or ignore pins say which needn’t. Macro latencies. A memory block or other macro may contain its own internal clock tree. Liberty can describe that internal delay with min/max_clock_tree_path timing groups, and OpenSTA includes it in latency and skew reports when asked to.

TritonCTS uses that information to balance the macro tree against the register tree by inserting delay buffers, and adds dummy loads (cells that exist only to add capacitance) to even out the leaves. Its inserted instances carry the name prefixes clkbuf_regs, delaybuf and clkload, which makes them easy to find in the netlist. Clock nets often leave CTS with routing guides or pre-routed wires and special wiring rules attached, and are routed before ordinary signal nets so they get the best tracks.

CTSPlaced designnetlist + DEFCell timingLiberty .libCell shapesLEFConstraintsSDCCTS settingsTcl scriptDesign + treenetlist + DEFNew constraintsSDC (propagated)Reportsskew, latencyinputsoutputs

Five inputs on the left. Tap a file to see what it holds, then run CTS to produce the outputs.

CTS inputs and outputs with their usual file formats. Tap a file to read what it holds.Share freely with credit: ‘Figure from chipfieldguide.com’

The tool builds the tree from the bottom up, in four moves:

  1. Group neighbors. Flip-flops that sit close together form small groups.
  2. Feed each group. A small amplifier, called a , sits next to each group and repeats the tick to it.
  3. Branch back to the start. Bigger buffers feed the small ones, level by level, until one trunk reaches the spot where the clock enters the chip.
  4. Even it out. Branches that would get the tick too early get a little extra delay, so everyone hears it together.

Each flip-flop has two deadlines. New data must arrive a little before the tick, or it misses its ride. And it must not change too soon after the tick, or it tramples the old value before that value is saved.

Data that arrives late can be fixed by slowing the clock down. Data that changes too soon cannot, because clock speed plays no part in that deadline. So engineers watch it closely.

The tree also saves power with . Switches on its branches stop the tick from reaching parts of the chip that have nothing to do right now. Parts that don’t tick don’t use that power.

The build, step by step

  1. Find the clock networks. Starting at each clock’s source, as named in the timing constraints, trace every wire and cell the clock passes through to every sink, and note where a derived clock (for example, one running at half speed) begins.
  2. Group the sinks. Gather nearby sinks into clusters small enough for one buffer to drive within its limits on load and edge rate. TritonCTS uses a k-means method (repeatedly assigning each sink to the nearest cluster center and moving the centers), with a cap on sinks per cluster and on cluster width.
  3. Build the upper tree. Connect the cluster buffers to the root with a branching structure (an H-tree in TritonCTS), adding buffers wherever a wire gets too long to keep the edge sharp.
  4. Balance. Branches that would deliver the tick early get extra delay buffers, dummy loads or slightly longer wire until arrival times match within the target.
  5. Legalize and optimize. Nudge the new cells into legal positions in the rows of cells, switch timing to the real clock delays, then repair the timing failures that appear.

From assumed to real clock delays

Before CTS, the constraints file (SDC) tells the timing tool what to assume: set_clock_latency gives an estimated travel time for the clock and set_clock_transition an estimated edge rate. After CTS, the command set_propagated_clock tells the tool to calculate those values from the actual buffers and wires instead. The tool that does this is a : it adds up delays along every timing path and checks each one against the clock, without simulating any input data.

What CTS measures

  • Skew. Global skew is the spread of arrival times over all sinks. Local skew is the difference between two flip-flops connected by a timing path. Only local skew affects whether the chip works: two flip-flops that never exchange data can be far apart in time without harm.
  • Insertion delay (latency). Source-to-sink travel time. Shorter is better: fewer buffers burn less power, and a shorter path varies less.
  • Transition. The clock’s rise and fall time at each pin, held under a maximum.
  • Duty cycle. The fraction of each period the clock spends high, ideally 50%. Buffers whose rising and falling delays differ shift that fraction a little at every level, and through a long chain a 50% clock can drift toward being stuck high.

Setup, hold and the sign of skew

Every timing path is checked twice. The check asks whether data launched on one tick arrives before the next tick, with a little time to spare. The check asks whether the new data arrives late enough not to spoil what the receiving flip-flop is capturing on the same tick. Take a launch flip-flop and a capture flip-flop. Call their clock arrival times LL and CC, the data delay dd (which includes the flip-flop’s own delay), the period TT, and the skew s=C−Ls = C - L. Then:

  • Setup =(T+C−tsetup)−(L+dmax)=T+s−dmax−tsetup= (T + C - t_{\mathrm{setup}}) - (L + d_{\mathrm{max}}) = T + s - d_{\mathrm{max}} - t_{\mathrm{setup}}
  • Hold slack =(L+dmin)−(C+thold)=dmin−s−thold= (L + d_{\mathrm{min}}) - (C + t_{\mathrm{hold}}) = d_{\mathrm{min}} - s - t_{\mathrm{hold}}

Here dmaxd_{\mathrm{max}} and dmind_{\mathrm{min}} are the slowest and fastest possible data delays, and slack is spare time: positive passes, negative fails. A small example: with T=1000 psT = 1000\,\mathrm{ps}, s=+40 pss = +40\,\mathrm{ps}, dmax=950 psd_{\mathrm{max}} = 950\,\mathrm{ps} and tsetup=80 pst_{\mathrm{setup}} = 80\,\mathrm{ps}, setup slack =1000+40−950−80=+10 ps= 1000 + 40 - 950 - 80 = +10\,\mathrm{ps}. Positive skew (capture later) adds to setup slack and subtracts the same amount from hold slack; negative skew does the reverse. The period appears only in the setup check, which is why a hold failure can’t be fixed by slowing the clock.

Useful skew

Zero skew is a convenience, not a requirement. Timing only needs each connected pair of flip-flops to pass its checks. So a tool can deliver the clock to the capture flip-flop of a failing path a little later, giving that path more time. The time is borrowed from the next path, which starts at the same flip-flop and now launches later. This is . Modern tools apply it after CTS as concurrent clock and data optimization: they adjust clock arrival times and resize logic cells together.

Clock cells

Cell libraries include dedicated clock buffers and inverters (here CLKBUF_X4, CLKINV_X8, where X4 or X8 is the drive strength) with strong drive and closely matched rise and fall delays, plus integrated clock-gating cells (ICGs). TritonCTS recognizes clock buffers by a substring of their name (by default CLKBUF) or by a library attribute, and also accepts an explicit list.

Shapes of clock network

TopologyIdeaTrade-off
Recursive symmetric branching, equal wire length to every endpointZero skew by symmetry; blockages and uneven sinks break it; used for top levels
Balanced (clustered) treeCluster sinks, buffer each cluster, balance upwardLeast power and wire; most exposed to variation on long unshared branches
Fishbone / spineWide trunk with side ribs feeding local treesRegular and easy to balance across a wide block; more wire than a tree
Grid of connected wires driven by many buffersLowest skew and variation; highest power and wiring cost
Multisource CTSH-tree feeding a sparse mesh, with local trees hanging from tap pointsMiddle ground: much of the mesh’s robustness at close to tree power

The H-tree’s exact zero skew makes it the usual choice for the top levels. Toyama estimates that a full mesh uses 20–40% more power than a conventional tree, with multisource designs much closer to the tree.

Wiring the clock

Wires next to each other are coupled electrically, so a signal switching on one can speed up or slow down an edge on its neighbor (). Clock wires therefore get a : double spacing to reduce coupling, and often double width to reduce resistance. TritonCTS applies a double-spacing rule to the clock nets above the leaves, by default on the first half of the levels counted from the root. Critical trunks may also be shielded by running power or ground wires alongside, as on the Alpha 21264’s global clock grid, which costs wiring tracks. The trunk usually runs on the upper metal layers, which are thicker and less resistive. Svensson and Afghahi showed that using wider-than-ordinary lines for global clock distribution keeps cross-chip clock delay low.

Clock gating, derived clocks and unrelated clocks

stops the clock to idle flip-flops to save power. The integrated clock-gating cell (ICG) that does it contains a latch, so its on/off (enable) signal can change only while the clock is inactive, and the gated clock never glitches. To CTS, an ICG is a node in the tree that must be balanced through, and its enable input gets its own setup and hold check (set_clock_gating_check). A clock divider (a flip-flop that toggles on every tick, producing a half-speed clock) is declared with create_generated_clock. The divider’s clock pin is a sink of the main tree, and its output is the root of a new tree. Clocks that have no fixed timing relationship are declared unrelated with set_clock_groups -asynchronous, so the timing tool doesn’t check paths between them. Data crossing between them is made safe by synchronizer circuits written into the design, and CTS does not try to balance one against the other.

After the tree: hold fixing

With real skew in place, some short paths into late-clocked flip-flops now fail hold. The fix is to slow those paths down with small delay buffers. OpenROAD’s repair_timing is meant to run after CTS with propagated clocks. It repairs setup first and then hold, and by default it won’t insert a hold buffer that would break setup.

What the targets really mean

Global skew is only a proxy. Whether the chip works depends on local skew between flip-flops joined by a timing path, which is why useful-skew methods work only on those pairwise constraints. Tools still keep skew tight within each , because a balanced tree is predictable and leaves room for later optimization to add deliberate offsets. What matters at signoff is slack after variation margins, and that depends on three properties of the tree: its latency, how much of each critical pair’s two clock paths is unshared, and its edge rates.

Variation, derates and CPPR

Static timing analysis allows for by scaling delays: for a setup check it multiplies the launch-side delays by a “late” factor (say 1.05) and the capture clock’s delays by an “early” factor (say 0.95), and does the reverse for hold. Clock paths are derated too (set_timing_derate -clock -early / -late). One flat factor everywhere turned out to be excessively pessimistic, because random variations partly cancel along a path of many cells. Advanced OCV (AOCV) therefore looks the factor up by path depth and distance; parametric OCV (POCV) goes further and gives each cell its own variation (a standard deviation, sigma) as a function of its delay and load, stored in the library in the Liberty Variation Format (LVF).

Now the catch. In a setup check the launch clock is treated as late and the capture clock as early. But if the two flip-flops’ clock paths share their first few buffers, those shared buffers would have to be late and early at the same instant, which is impossible. credits back the late-minus-early difference on the shared segment. A worked illustration: if two flip-flops split at the last buffer, nearly their whole clock path is shared, and they lose almost nothing to clock derates. If they split at the root, every buffer of both paths is derated against them. A 400 ps unshared path at ±5% costs about 20 ps on each side, around 40 ps in all. Shallow, low-latency trees, and clustering tightly connected flip-flops under one branch, both buy real slack.

Uncertainty before and after CTS

is a margin for jitter, the cycle-to-cycle wobble of the clock edge. Before CTS it usually also covers the skew the tree will add. The two are different things: jitter varies from one cycle to the next, while skew is a difference between locations. The ideal-mode set_clock_latency and set_clock_transition describe the tree only when the design is analyzed with ideal clocks. A typical flow deletes them after CTS and cuts setup uncertainty to jitter plus margin. Leaving the pre-CTS numbers in place counts skew twice, once as computed and once as margin, and sends the optimizer chasing violations that aren’t real.

Topology trade-offs in practice

An H-tree buys balance with capacitance: its total wirelength, and so its power, exceeds that of a buffered tree shaped around the actual sinks, and symmetric structures add latency. A puts branch resistances in parallel, and everything above the mesh is shared by all loads, so only the short stubs below it are exposed to uncorrelated variation. Multisource CTS drives a mesh that Toyama describes as one to two orders of magnitude sparser than a full clock mesh, fed by H-trees, with ordinary subtrees hanging off tap points. The Alpha 21264 shows the hybrid style at processor scale: a trunk and X- and H-trees fed 16 drivers of a gridded global clock, every grid wire was shielded by power or ground, locally gated clocks sat below, and 65 ps of skew was measured on silicon.

Clock gating inside the tree

An ICG’s clock pin sits upstream of the flip-flops it gates, so its clock arrives earlier than theirs. Its enable signal, though, is launched by ordinary flip-flops at full latency and must arrive before that earlier ICG clock, so the enable path gets less time than a normal path. Moving ICGs toward the root gates more of the tree and saves more power, but tightens enable timing; cloning an ICG so each copy sits near its flip-flops does the opposite. TritonCTS’s -balance_levels option tries to keep a similar number of levels across non-register cells such as clock gates and inverters.

Generated clocks, macros and skew groups

A divided clock inherits the main clock’s latency up to the divider flip-flop, plus the divider’s clock-to-Q delay, plus its own tree. So when a divide-by-2 domain exchanges data with its parent clock, the two trees must be balanced at the flip-flops where they meet. Skew groups say exactly which sinks must align: include the divider’s fanout and the main-clock flip-flops it talks to, and exclude scan-only or asynchronous sinks. Macros bring their own internal latency; TritonCTS inserts delay buffers to balance macro and register trees, with a derate option to insert only part of the delay needed.

Hold repair and the useful-skew budget

Post-CTS hold repair is where a bad tree shows. By default OpenROAD caps hold buffers at 20% of the design’s instance count and refuses buffers that would create setup violations unless told otherwise. Every picosecond of useful skew given to a setup path is a picosecond taken from hold on that pair and from setup on the next stage, so CCD methods look for clock-arrival changes that improve the worst negative slack (WNS) and total negative slack (TNS) together.

Power and electromigration

Clock nets are long, heavily loaded and switch every cycle. Switching also wears out wires. (EM) is metal atoms being pushed along a wire by current until it thins and fails; one form of it is driven by the heating from alternating (RMS) current, and its expected time to failure falls as the switching rate rises. Clock nets have the highest switching rate on the die, so clock drivers and trunk wires need EM checks at their true rate. The usual fixes are wider wires, extra vias (vertical connections between layers) on trunk connections, and splitting large loads across several drivers.

clock pinClock arrivalA?B?C?D?
1 / 5

1. Find the sinks: Trace from the clock pin to every sink. Until now timing used an ideal clock (dashed): zero delay to every pin.

The build, step by step, on 24 sinks. Arrival bars follow the drawn wire lengths and are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
Clustered treeLow skewHandles variationLow powerWire (vs. tree)1.0×
Topology

Clustered tree: Clusters, leaf buffers, then balancing upward. Least wire and power; long unshared branches are most exposed to variation.

Five clock network shapes over the same 20 sinks. Wire is measured from each drawing, relative to the clustered tree; the ratings restate the table.Share freely with credit: ‘Figure from chipfieldguide.com’

The simulator shows two flip-flops with a calculation between them. Sliders change when the tick reaches the sender, when it reaches the receiver, and how long the calculation takes.

Times are in ps, short for picoseconds. A picosecond is a trillionth of a second. In that time, light travels about the width of a grain of salt.

Two boxes answer two questions. “Fast enough?” asks if the data arrived in time. “Stable long enough?” asks if it stayed put until it was saved.

  1. Make the calculation slow, about 950, until “Fast enough?” says No. Now drag the receiver’s tick about 60 later. It says Yes again: the late tick gave the data more time.
  2. Make the calculation very short, about 150, and slide the receiver’s tick much later. “Stable long enough?” says No. Try slowing the clock. It still says No, because clock speed plays no part in this check.

The model is one sending (launch) flip-flop, the logic between, and one receiving (capture) flip-flop. All times are in picoseconds (ps, trillionths of a second). Setup slack = (period + capture arrival − setup time) − (launch arrival + data delay). Hold slack = (launch arrival + data delay) − (capture arrival + hold time). The data delay includes the launch flip-flop’s clock-to-Q delay.

  1. Useful skew fixes setup. Period 1000 ps, launch 300 ps, capture 300 ps, data 950 ps, setup 80 ps, hold 40 ps. Setup slack is −30 ps. Move capture to 360 ps: setup slack becomes +30 ps and hold slack drops from 910 ps to 850 ps.
  2. Hold ignores the period. Set data to 150 ps with launch and capture at 300 ps: hold slack is +110 ps. Move capture to 430 ps: hold slack is −20 ps. Now change the period from 1000 ps to 2000 ps and watch hold slack stay at −20 ps.
  3. Find the window. With data 950 ps you needed capture at 330 ps or later. Suppose the same capture flop also receives a 150 ps path from the same launch flop. Set data to 150 ps to stand in for it, and slide capture up from 330 ps until hold fails (above 410 ps). Any capture time between 330 and 410 ps satisfies both paths; outside that window, only a data-path change helps.

Use the sim to reason about margins the sliders don’t show directly. Fold uncertainty and variation derates into the sliders by hand, and watch how fast the window that useful skew opens closes again.

  1. Add uncertainty. Take period 1000 ps, launch 300, capture 360, data 950: setup slack +30 ps. Model 50 ps of setup uncertainty by raising setup from 80 to 130 ps: slack −20 ps. The skew that fixed the ideal-mode path is not enough with post-CTS margins.
  2. Add OCV. Derate the data path late by 5% (950 → 998 ps) and the part of the capture clock path not shared with the launch clock early by 5% (if 200 of its 360 ps is unshared, capture becomes 350 ps). Setup slack falls by about 58 ps, before even counting the launch clock’s own late derate. The CPPR credit applies only to the shared segment.
  3. Borrowing from the next stage. Suppose the downstream path from the capture flop has data 900 ps and launches at 300 ps into a flop clocked at 300 ps: setup slack +20 ps. After you moved this flop’s clock to 360 ps, that stage launches at 360 ps and its slack is −40 ps. Around any loop of stages the skews cancel, which is the core of Fishburn’s formulation (see Under the hood).

In the sim, a green “Yes” means the signal makes its deadline; a red “No” means it misses.

In the sim, a green slack number means the check passes with time to spare; red means it fails by that many picoseconds.

Loading simulation…

After the tool builds the tree, an engineer looks at a picture of the chip with the clock network drawn on top. They check how evenly the tick arrives and how much power the network will use.

A branch much longer than its neighbors stands out. The fix is often to move some flip-flops, or to tell the tool which parts belong together.

An engineer drives CTS with two text files. The first is the constraints file (SDC): a list of commands that tell the timing tools about each clock. This illustrative example (times in nanoseconds; 1 ns = 1000 ps) shows the clock setup used before CTS and what changes afterward. The numbered notes under the code explain the key lines.

constraints.sdcsdc
# constraints.sdc  (illustrative; times in ns)
create_clock -name core_clk -period 1.000 [get_ports clk]
create_clock -name io_clk   -period 4.000 [get_ports io_clk_in]
create_generated_clock -name div2_clk -source [get_ports clk] \
    -divide_by 2 [get_pins u_clkdiv/q_reg/Q]

# Pre-CTS only: estimates for a tree that does not exist yet
set_clock_latency 0.400 [get_clocks core_clk]
set_clock_transition 0.060 [get_clocks core_clk]
set_clock_uncertainty -setup 0.120 [get_clocks core_clk]
set_clock_uncertainty -hold  0.040 [get_clocks core_clk]

# Board-level delay before the clock reaches the chip pin
set_clock_latency -source 0.150 [get_clocks core_clk]

# Clock-gating cells must see a stable enable around the edge
set_clock_gating_check -setup 0.040 -hold 0.010

# io_clk is unrelated; CDC is handled by synchronizers in RTL
set_clock_groups -asynchronous -group {core_clk div2_clk} -group {io_clk}

# Post-CTS version replaces lines 8-11 with:
# set_propagated_clock [all_clocks]
# set_clock_uncertainty -setup 0.050 [get_clocks core_clk]
# set_clock_uncertainty -hold  0.020 [get_clocks core_clk]
  1. 1L2Defines a 1 GHz clock (period 1 ns) arriving on the chip pin named clk. By default it rises at 0 and falls at 0.5 ns.
  2. 2L4A half-speed clock made by a divider flip-flop. The divider’s clock input is a sink of core_clk’s tree; its output (Q) is the root of a new tree.
  3. 3L8Estimated travel time through the tree that doesn’t exist yet (0.4 ns). Ignored once clocks are propagated; delete it anyway so the file says what is real.
  4. 4L9Estimated clock edge rate, so flip-flop timing before CTS uses a realistic edge.
  5. 5L10Safety margin before CTS = jitter + expected skew + extra margin. It shrinks after CTS because real skew is then calculated; leaving the pre-CTS value in counts skew twice, a common bug.
  6. 6L14Delay before the clock reaches the chip pin. It stays after CTS; real delays replace only the on-chip part.
  7. 7L17Adds setup and hold checks on the enable inputs of clock-gating cells, against the clock at that cell.
  8. 8L20Paths between these groups are not timed, and CTS does not need to balance across them.
  9. 9L23After CTS: use delays calculated through the built tree.

The second is the script that runs the tools. This one uses the open-source OpenROAD tools and is illustrative. It loads the placed design, builds the tree, reports on it, then repairs timing. A “net” is one wire and everything connected to it; “parasitics” are the resistance and capacitance of the wires, which set their delay.

cts.tcltcl
# cts.tcl  (illustrative, OpenROAD syntax)
read_db results/3_place.odb
read_sdc constraints_cts.sdc

# RC per micron for clock and signal wires (clock trunk on upper metal)
set_wire_rc -clock  -layer M7
set_wire_rc -signal -layer M3

clock_tree_synthesis \
    -root_buf CLKBUF_X16 \
    -buf_list {CLKBUF_X4 CLKBUF_X8 CLKBUF_X16} \
    -sink_clustering_enable \
    -sink_clustering_size 20 \
    -sink_clustering_max_diameter 60 \
    -balance_levels \
    -apply_ndr half \
    -repair_clock_nets

set_propagated_clock [all_clocks]
estimate_parasitics -placement
report_cts
report_clock_skew -setup -digits 3
report_clock_skew -hold  -digits 3

# Fix setup first, then hold, then re-legalize the new cells
repair_timing -setup
repair_timing -hold -hold_margin 0.010
detailed_placement
report_checks -path_delay max -digits 3
report_checks -path_delay min -digits 3
  1. 1L6Tells the tool which metal layer’s resistance and capacitance to use when estimating clock wires: here thick upper metal (M7).
  2. 2L10Root buffer: the strongest cell, driving the trunk.
  3. 3L11Buffers the tool may use: clock-specific cells, not general-purpose BUF_X* cells.
  4. 4L12Group nearby sinks first; each group gets a leaf buffer, which becomes an endpoint of the H-tree.
  5. 5L13At most 20 sinks per leaf buffer, within a 60 µm diameter.
  6. 6L15Keep a similar number of levels through clock gates and inverters.
  7. 7L16Double-spacing wiring rule on the first half of the levels counted from the root; the leaf nets never get it.
  8. 8L17Buffer the long wire from the clock pin to the root before latency balancing.
  9. 9L19Switch timing from the ideal clock to delays calculated through the tree that was just built.
  10. 10L26Setup repair before hold, so hold buffers don’t undo setup fixes.
  11. 11L27Hold repair with 10 ps extra margin. Buffers that would break setup are refused by default.
  12. 12L28Snap the buffers inserted by CTS and by repair into legal positions.

Two artifacts tell you whether the tree is good. The first is the CTS summary plus a skew report: how many sinks and buffers, how deep the tree is, and the worst skew between two connected flip-flops. The second is the worst timing path, printed with the clock network expanded so you can see the clock delays on both sides. Both are illustrative, in OpenROAD/OpenSTA format, with times in ns.

cts_summary.loglog
[INFO CTS-0098] Clock net "core_clk"
[INFO CTS-0099]  Sinks 18432
[INFO CTS-0100]  Leaf buffers 912
[INFO CTS-0101]  Average sink wire length 41.27 um
[INFO CTS-0102]  Path depth 6 - 8
[INFO CTS-0207]  Dummy loads inserted 37
Total number of Clock Roots: 1.
Total number of Buffers Inserted: 1206.
Total number of Clock Subnets: 1206.
Total number of Sinks: 18432.

> report_clock_skew -setup -digits 3
Clock core_clk
  0.461 source latency u_lsu/addr_reg[3]/CK ^
 -0.389 target latency u_lsu/tag_reg[3]/CK ^
  0.050 clock uncertainty
 -0.018 CRPR
--------------
  0.104 setup skew
  1. 1L2Every clock pin reached by core_clk: flops, latches, ICG clock pins and macro clock pins.
  2. 2L3About 20 sinks per leaf buffer, matching -sink_clustering_size.
  3. 3L4Average leaf wire length. Long leaf wires mean slow slew and more coupling; watch for outliers near macros.
  4. 4L5Depth varies from 6 to 8 levels, usually because of extra levels through clock gates or into macro subtrees. Each extra level adds latency that isn’t shared with neighboring branches.
  5. 5L6Dummy loads (clkload*) even out leaf capacitance so sibling branches balance.
  6. 6L14Launch-side latency of the worst pair. This includes source latency and the propagated tree.
  7. 7L15Capture-side latency, subtracted. Here the capture clock arrives 72 ps before the launch clock, which hurts setup.
  8. 8L16Setup uncertainty is added to the reported skew so it reflects the effective penalty.
  9. 9L17CRPR credit for the shared part of the two clock paths reduces the effective skew.
  10. 10L19Note the sign: this format reports launch minus capture, the opposite of the capture-minus-launch convention used in the sim.

A timing path report reads top to bottom in two halves. The first half builds the arrival time: the clock edge at 0, the launch clock’s travel through the tree, then each cell on the data path with its delay and the running total. The second half builds the required time: the next clock edge one period later, the capture clock’s travel through the tree, the CPPR credit, and then the margins taken away (uncertainty and the flip-flop’s setup time). Slack is required minus arrival. The arrows ^ and v mark rising and falling edges.

worst_setup_path.rpttext
Startpoint: u_alu/acc_reg[12] (rising edge-triggered flip-flop clocked by core_clk)
Endpoint: u_alu/res_reg[31] (rising edge-triggered flip-flop clocked by core_clk)
Path Group: core_clk
Path Type: max

  Delay    Time   Description
---------------------------------------------------------
  0.000   0.000   clock core_clk (rise edge)
  0.458   0.458   clock network delay (propagated)
  0.000   0.458 ^ u_alu/acc_reg[12]/CK (DFF_X1)
  0.121   0.579 ^ u_alu/acc_reg[12]/Q (DFF_X1)
  0.143   0.722 v u_alu/U812/ZN (NAND2_X2)
  0.208   0.930 ^ u_alu/U1033/ZN (AOI22_X1)
  0.236   1.166 v u_alu/U1207/ZN (OAI21_X2)
  0.198   1.364 ^ u_alu/U1301/ZN (XNOR2_X1)
  0.028   1.392 ^ u_alu/res_reg[31]/D (DFF_X1)
          1.392   data arrival time

  1.000   1.000   clock core_clk (rise edge)
  0.443   1.443   clock network delay (propagated)
  0.011   1.454   clock reconvergence pessimism
  0.000   1.454 ^ u_alu/res_reg[31]/CK (DFF_X1)
 -0.050   1.404   clock uncertainty
 -0.042   1.362   library setup time
          1.362   data required time
---------------------------------------------------------
          1.362   data required time
         -1.392   data arrival time
---------------------------------------------------------
         -0.030   slack (VIOLATED)
  1. 1L9Launch clock latency, late-derated: 0.150 ns source latency + 0.308 ns through the tree.
  2. 2L11Data path starts at clock-to-Q. Everything from here to line 16 is data delay.
  3. 3L20Capture clock latency, early-derated: 0.150 + 0.293 ns. The capture clock is 15 ps earlier than launch, so skew costs this path 15 ps.
  4. 4L21CPPR: the two flops share the root and second-level buffers. Late minus early delay on that shared segment (11 ps) is credited back.
  5. 5L23Post-CTS setup uncertainty (jitter + margin), no longer a skew estimate.
  6. 6L24Setup time from the Liberty constraint table, looked up with the actual clock and data slews.
  7. 7L30−30 ps. Options: CCD delays res_reg[31]’s leaf by ≥30 ps if its downstream paths and hold allow; resize the data path; or regroup the two flops under a deeper common branch to gain CPPR.
memorymacroABCDArrival timeearlylateABCDSpread (skew)154 ps
Placement of group D

Colored by arrival time. D’s branch detours around the block and arrives 154 ps after the earliest group. Tap a group.

A layout view with the clock tree colored by arrival time, as engineers inspect it after CTS. Times are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
  • One branch runs late. A big block in the way forces a detour, so the tick reaches that corner late.
  • Data changes too soon. Once the real tree exists, some short paths deliver new data too early. The fix is to add tiny delay parts, which take up space and use power.
  • Noisy neighbors. A busy wire right next to the clock wire can nudge the tick early or late. Giving the clock wire extra room prevents it.
  • Wasted power. Too many strong amplifiers, or no clock gating, and the clock burns power even when the chip has nothing to do.
  • Pre-CTS estimates left in place. If the estimated clock delay and the large pre-CTS safety margin stay in the constraints after the real tree exists, skew is counted twice and the tools chase failures that aren’t real. Caught by reviewing the post-CTS constraints file and comparing slack before and after switching to real clock delays.
  • Hold explosion. Skew creates thousands of hold failures, and the repair adds thousands of delay buffers, costing area and power. Caught by tracking how many hold buffers were added; the root cause is often an unbalanced tree or a missing skew-group definition.
  • Missing derived clock. If nobody declares the half-speed clock coming out of a divider (create_generated_clock), the timing tool treats that output as ordinary data. The flip-flops it drives then go unchecked or are checked wrongly. Caught by reports that list flip-flops with no clock.
  • Slow edges at the leaves. Clusters that are too big or too spread out leave each leaf buffer driving too much wire, so the clock edge at the flip-flops is too slow (a max-transition violation). Seen in the CTS summary and fixed with smaller clusters.
  • Crosstalk on clock wires. Clock wires routed at normal spacing next to busy signal wires get their edges nudged by every neighboring switch, which shows up as extra skew only after the final wiring. Prevented with wider-spaced wiring rules and caught by timing analysis that models coupling between wires.
  • Balanced in the wrong corner. A tree balanced at one corner can be skewed at another, because buffer delay and wire delay change by different amounts with voltage and temperature: a tree whose branches mix buffer-heavy and wire-heavy paths drifts apart. Teams choose the balancing corner deliberately and check skew in every signoff scenario.
  • Long unshared clock paths. Critical pairs that split near the root pay the clock derate on almost the whole tree and get little CPPR credit back. Caught by sorting critical paths by their unshared clock latency; fixed by clustering the pair under one branch or with skew-group constraints.
  • Clock-gate enable timing. ICGs placed near the root save the most power, but their enable paths are checked against an early clock and fail. Look for clock-gating check violations, then clone the ICG or move copies toward their loads.
  • Duty-cycle drift. Long chains of buffers with mismatched rise and fall delays move the duty cycle at every stage. The minimum-pulse-width check catches a high or low phase that has become too short; run it on the routed clock rather than waiting for signoff. Use clock cells with matched edges, or pairs of inverters.
  • Scan mode surprises. For manufacturing test, scan chains link the flip-flops into long shift registers, with little or no logic between neighbors. Skew that is harmless in normal operation can then make data race through in test mode (a hold failure). Time the scan shift mode as its own scenario.
  • Electromigration on clock drivers. Wear-out from RMS current rises with switching activity, and clock nets switch every cycle. Run EM checks on clock nets with their true switching rate, not the default assumed for data nets.
logichold bufferslaunchcapturelate capture clocktick + hold timenew data arrives ✗time →
Failure
Show

The capture clock arrives late, and the data path is short, so new data arrives before the hold window closes: a hold failure.

Common CTS failures and their fixes. Pick one, then toggle the fix. Values are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

The problem

Given where the sinks are and how much load each presents, find a tree of wires from the source to all of them that meets a skew target with the least wire (and so the least capacitance and power). There are three formulations: a zero-skew tree (ZST), where every sink gets the clock at exactly the same time; a bounded-skew tree (BST), where skew only has to stay under a bound because signoff tolerates that much; and useful skew, which constrains only the local skew between sinks that exchange data. Most algorithms split the work into topology generation (which sinks merge with which) and geometric embedding (where each branch point goes on the chip).

H-trees and the method of means and medians

An is exactly zero-skew by symmetry but only for regular sink arrays, so it serves top-level distribution. The method of means and medians (Jackson, Srinivasan and Kuh, 1990) handles arbitrary sink positions top-down: split the sinks at the median into two equal halves, connect the center of mass of the whole set to the centers of mass of each half, and recurse with the split direction alternating. It balances sink counts and is fast, with no skew guarantee. Bottom-up recursive geometric matching (Kahng, Cong and Robins) pairs nearby subtrees and achieves near-zero pathlength skew in practice, still without a guarantee.

Tsay’s exact zero skew

Tsay (1991) was the first to guarantee exact zero skew, under the Elmore model. It merges two zero-skew subtrees bottom-up at a tapping point on the wire joining their roots. With subtree delays t1t_1, t2t_2, downstream capacitances C1C_1, C2C_2, a connecting wire of length LL, and per-unit resistance rr and capacitance cc, equalizing the Elmore delay from the tapping point to both sides gives the fraction zz of the wire on subtree 1’s side:

z=t2−t1+rL (C2+cL/2)rL (C1+C2+cL)z = \frac{t_2 - t_1 + rL\,(C_2 + cL/2)}{rL\,(C_1 + C_2 + cL)}

If 0≤z≤10 \le z \le 1 the tap lies on the wire. Otherwise one subtree is too slow, and the algorithm elongates (snakes) the wire on the faster side until delays match. Buffered trees apply the same merge per stage, since a buffer hides its subtree’s capacitance from upstream.

subtree 1t₁ 40 ps · C₁ 60 fFsubtree 2t₂ 40 ps · C₂ 60 fFtap · z = 0.5054 ps54 ps

z = 0.50: tap 100 µm from subtree 1. Elmore delay from the tap: 54.0 ps to subtree 1’s sinks, 54.0 ps to subtree 2’s. Skew 0.

Tsay’s exact zero-skew merge under the Elmore model. Subtree 1 is fixed at t₁ = 40 ps, C₁ = 60 fF; the wire is 200 µm with illustrative r = 2 Ω/µm, c = 0.2 fF/µm. Move subtree 2’s delay and load to push the tap along the wire, or off its end, where the faster side must be snaked.Share freely with credit: ‘Figure from chipfieldguide.com’

Deferred-merge embedding (DME)

Earlier methods fixed each branch point’s location as soon as it was computed. DME, found independently by three groups around 1991–92, defers that choice. Given a topology, a bottom-up pass computes for each internal node a merging segment: the locus of all points where the two child subtrees can join with minimum added wire and zero skew. Chip wires run only horizontally and vertically, so distance is measured as Δx+Δy\Delta x + \Delta y (Manhattan distance). Under that measure these loci are Manhattan arcs, line segments tilted at 45°, and merging two arcs yields another arc. A top-down pass then picks the root anywhere on its segment and places each child at the point of its segment nearest the parent. Kahng and Tsao summarize DME as a linear-time algorithm that optimally embeds any given topology with exact zero skew and minimum total wirelength. Its limitation is that it needs the topology as input, so it is paired with topology generators such as matching or greedy merging.

Bounded-skew trees (BST/DME)

Zero skew wastes wire when signoff tolerates some skew. BST/DME replaces merging segments with merging regions, the loci of feasible embedding points that keep skew within a bound BB, then embeds top-down as before. It works under both pathlength and Elmore delay, and sweeping BB gives a smooth trade-off between skew and wirelength; at B=0B = 0 it reduces to DME, and with no bound at all it approaches a Steiner tree, the shortest possible tree connecting the points.

Elmore delay and its limits

For an RC tree, the to a node sums, over each resistor on the path from the source, that resistance times all capacitance downstream of it. It is the first moment of the impulse response, which makes it additive along a path and cheap enough to evaluate inside merge loops. Gupta, Tutuianu and Pileggi showed that it acts as a delay bound for RC trees, even with general input signals. It is often inaccurate but faithful: reducing Elmore delay almost always reduces true delay, so it ranks choices well. It knows nothing about input slew, nonlinear driver behavior or inductance, and skew is a small difference of two large delays, so computing it needs far more accuracy than computing either delay alone. Modern CTS therefore builds with characterized lookup tables. TritonCTS characterizes buffered wire segments on the fly from Liberty data and selects segments from that table.

Buffer insertion: van Ginneken’s dynamic program

Van Ginneken (1990) finds optimal buffer positions on a fixed tree under Elmore wire delay and a linear buffer model. Working from sinks to source, each candidate solution at a node is a pair (Q,C)(Q, C): QQ is the slack (how late the signal may still arrive at that node) and CC is the capacitance the node’s driver will see downstream. At each legal buffer position the algorithm also considers a buffered option, merges candidate lists at branch points, and discards dominated candidates, those with more load and less required time than another. Run time is O(n2)O(n^2) in nn candidate positions. Lillis, Cheng and Lin extended it to bb buffer types in O(b2n2)O(b^2 n^2), and Li and Shi reduced that to O(bn2)O(b n^2) by showing the useful candidates lie on a convex hull in the (Q,C)(Q, C) plane. For clocks the objective changes from maximum slack to balance and power, but the candidate-and-prune structure carries over. TritonCTS uses dynamic programming over its generalized H-tree to choose the minimum-power topology that meets latency and skew targets, with capacitated k-means for sink clustering.

Clock mesh analysis

STA delay calculation assumes each net has one driver feeding a tree of wires. A mesh breaks that: many drivers are shorted together, paths split and rejoin, and the skew between mesh-buffer inputs feeds into the output. SPICE handles this accurately but slowly, and simplified models miss slew and input-skew effects, which is why research on fast mesh timing (including learned models trained on SPICE) continues.

Useful-skew scheduling as an LP

Fishburn (1990) posed skew scheduling as a linear program. Assign each register ii a latency xix_i. For every path i→ji \to j with delays dmind_{\mathrm{min}} and dmaxd_{\mathrm{max}}:

setup:xi+dmax+tsetup≤xj+Thold:xi+dmin≥xj+thold\begin{array}{ll} \text{setup:} & x_i + d_{\mathrm{max}} + t_{\mathrm{setup}} \le x_j + T \\ \text{hold:} & x_i + d_{\mathrm{min}} \ge x_j + t_{\mathrm{hold}} \end{array}

Minimize TT, or for a fixed TT maximize the smallest margin across all constraints. Fishburn named the two hazards zero-clocking (setup) and double-clocking (hold), and on a 4-bit ripple-carry adder with accumulator cut the minimum period from 9.5 ns with zero skew to 7.5 ns with scheduled skew. Every constraint bounds a difference xj−xix_j - x_i, so for a fixed TT the whole problem is a set of difference constraints. Draw each register as a node and each constraint xj−xi≤bx_j - x_i \le b as an edge from ii to jj with weight bb. A valid set of latencies exists exactly when no cycle of edges has a negative total weight, which the Bellman–Ford shortest-path algorithm checks. Summing the setup constraints around any cycle of kk registers cancels every xx, leaving T≥(∑(dmax+tsetup))/kT \ge \bigl(\sum (d_{\mathrm{max}} + t_{\mathrm{setup}})\bigr) / k. For example, two registers that feed each other through 900 ps and 1100 ps paths (ignoring setup time) can never run faster than a 1000 ps period, whatever the skew. Skew can move time around a loop but cannot shrink the loop’s average. Practical CCD methods solve incremental, physically constrained versions of this problem after CTS, because each latency change must be built from real buffers on a real tree.

Novice · 0 of 4 correct
  1. Q1Period 1000 ps, launch clock arrival 300 ps, capture clock arrival 340 ps, data-path delay (including clock-to-Q) 950 ps, setup 80 ps. What is the setup slack?

  2. Q2After CTS, a short path has −25 ps hold slack. Why won’t lowering the clock frequency fix it?

  3. Q3After CTS, the timing tool switches from an ideal clock to a propagated clock. What changes?

  4. Q4Why are clock wires often given double the normal spacing from their neighbors?

Sources

Show Hide 20 sources
  1. Clock Distribution Networks in Synchronous Digital Integrated CircuitsEby G. Friedman · Proceedings of the IEEE, vol. 89, no. 5 · 2001Skew vs. min/max path constraints, localized (useful) skew, H-tree vs. buffered tree, mesh, Fishburn’s LP skew scheduling; clock signals have the greatest fanout, longest distances and highest speeds; clock networks can take much more than 25% of power; Table 2: clock network share of power 40% (Alpha 21064), 40% (21164), 44% (21264).
  2. VLSI Physical Design: From Graph Partitioning to Timing Closure, Chapter 7 slides (Specialized Routing)Andrew B. Kahng, Jens Lienig, Igor L. Markov, Jin Hu · Book companion site, TU Dresden (ifte.de) · 2022Global vs. local skew, zero/bounded/useful skew formulations, H-tree, MMM, exact zero skew, DME, buffering under variation.
  3. Clock skewWikipediaPositive skew = receiving register clocked later; hold violations cannot be fixed by lengthening the period; skew vs. jitter.
  4. OpenSTA CommandsParallax Software / OpenSTA contributors · OpenSTA project (GitHub)Semantics of set_propagated_clock, set_clock_latency, set_clock_transition, set_clock_uncertainty, create_generated_clock, set_clock_groups, set_clock_gating_check, set_timing_derate, report_clock_skew, report_check_types -min_pulse_width.
  5. Clock Tree Synthesis (TritonCTS 2.0) documentationThe OpenROAD Project · OpenROAD documentationclock_tree_synthesis options: H-tree, CKMeans sink clustering, buffer lists, 2X-spacing NDR strategies, dummy loads, delay buffers for macros, report_cts.
  6. Gate Resizer documentation (repair_timing, repair_clock_nets)The OpenROAD Project · OpenROAD documentationrepair_timing runs after CTS with propagated clocks; setup before hold; hold buffer cap defaults to 20% of instances.
  7. Toward an Open-Source Digital Flow: First Learnings from the OpenROAD ProjectT. Ajayi, V. A. Chhabria, M. Fogaça, et al. · ACM/IEEE Design Automation Conference (DAC) · 2019TritonCTS: generalized H-tree, dynamic programming for minimum-power topology under latency/skew targets, capacitated k-means sink clustering.
  8. Duty cycle distortion correction circuitry (US 9,048,823 B2)J. H. Bui, L. H. Khoo, K. Nguyen, C. Sung, K. C. Sia (Altera Corp.) · U.S. Patent and Trademark Office, via Google Patents · 2015Unequal rise/fall delays in clock buffers push the duty cycle further from 50% at each stage; a 50% clock can approach 100% and end up stuck high.
  9. Clock gatingWikipediaGating stops flops from switching to save power; integrated clock-gating cells contain a latch for a glitch-free gated clock.
  10. What’s The Difference Between CTS, Multisource CTS, And Clock Mesh?Harvey Toyama · Electronic Design · 2012Kept as a fallback for the attributed estimates (no open primary source found). Mesh: near-zero skew at the fabric, 20–40% more power than conventional CTS; multisource CTS uses tap points on a sparse mesh fed by an H-tree.
  11. GATMesh: Clock Mesh Timing Analysis using Graph Neural NetworksMuhammad Hadir Khan, Matthew Guthaus · arXiv:2507.05681 · 2025Mesh analysis is hard because of reconvergent paths, multi-source driving and mesh-buffer input skew; SPICE is accurate but slow.
  12. The TAU 2014 Contest: Removing Pessimism during Timing Analysis (slides)Igor Keller, Jin Hu, Debjit Sinha · ISPD 2014 / TAU Workshop · 2014Early/late delay bounds from derating; a signal cannot be both early and late on the common clock path; CPPR credit.
  13. Fast and Accurate Library Generation Leveraging Deep Learning for OCV ModellingEunice Naswali, Namhoon Kim, Pravin Chandran · International Symposium on Quality Electronic Design (ISQED) · 2021A single global derate caused excessive pessimism; AOCV accounts for cancellation of random variation with path depth and distance; POCV models local variation per cell as a function of delay and load, represented in Liberty Variation Format (LVF).
  14. DiffCCD: Differentiable Concurrent Clock and Data OptimizationYuhao Ji, Yuntao Lu, Zuodong Zhang, Zizheng Guo, Yibo Lin, Bei Yu · IEEE/ACM International Conference on Computer-Aided Design (ICCAD) · 2025Traditional CTS minimizes skew; concurrent clock and data optimization treats skew as a timing resource post-CTS.
  15. Electromigration-Aware Architecture for Modern MicroprocessorsFreddy Gabbay, Avi Mendelson · arXiv:2005.01593 · 2020Peak, average and RMS (Joule-heating) EM; RMS-EM time to failure falls as switching rate rises.
  16. Elmore delayWikipediaElmore delay as the first moment of the impulse response; RC-ladder formula; inaccurate but faithful; bound result of Gupta, Tutuianu and Pileggi (1997).
  17. Planar-DME: A Single-Layer Zero-Skew Clock Tree RouterAndrew B. Kahng, Chung-Wen Albert Tsao · IEEE Transactions on Computer-Aided Design, vol. 15, no. 1 · 1996Reviews MMM, matching-based trees, Tsay’s exact zero skew with wire snaking, and the DME algorithm (merging segments, linear time).
  18. EE695K VLSI Interconnect, Lecture 9: High-Speed Clock RoutingCheng-Kok Koh · Purdue University, School of ECE (course notes) · 2005MMM, Tsay’s Elmore-based merge and tapping-point equation, DME, BST/DME, wire sizing for skew.
  19. Bounded-Skew Clock and Steiner RoutingJason Cong, Andrew B. Kahng, Cheng-Kok Koh, C.-W. Albert Tsao · ACM Transactions on Design Automation of Electronic Systems · 1998BST/DME: merging regions generalize DME’s merging segments; smooth trade-off between skew bound and wirelength.
  20. An O(bn²) Time Algorithm for Optimal Buffer Insertion with b Buffer TypesZhuo Li, Weiping Shi · DATE 2005; arXiv:0710.4691 · 2005Van Ginneken’s O(n²) dynamic program, (Q, C) candidates and dominance pruning, Elmore wire model; extension to b buffer types.