Design Flow · Stage 3 of 13 · Front end

RTL design

Engineers write what the chip should do as code. Unlike an app, this code describes circuits that all run at the same time.

Engineers write register-transfer level (RTL) code, in a language such as SystemVerilog, saying what each register (a group of one-bit memories) holds after every clock tick and what logic computes it. Automatic checks catch common mistakes early.

Code that simulates and synthesizes the same way, reset strategy, safe crossings between clock and reset domains (CDC and RDC), lint and CDC sign-off, and power intent (UPF) alongside the RTL. Generators such as Chisel and high-level synthesis produce RTL for some blocks.

Builds The behavior, as code

The spec says what a chip must do. The architecture picks its big blocks. Next, someone has to describe every block so exactly that a computer can build it. That description is called , short for register-transfer level.

Engineers write RTL in a special language. It looks like computer code, with names, if-statements and math. But each line describes a piece of circuit, and all the pieces work at the same time.

Every later step starts from this plan. A mistake left in the RTL gets built right into the chip.

A digital chip works with bits, values that are either 0 or 1. Two kinds of circuit handle them. compute: an AND gate outputs 1 only when both of its inputs are 1, and thousands of gates together make an adder or a comparator. remember: each one holds a single bit. A steady on-off signal called the , ticking perhaps a billion times a second, tells every flip-flop when to take in a new value. A group of flip-flops that holds a number together, say 32 bits, is a register.

(RTL) design describes a chip in exactly those terms: which registers exist, and what logic computes each register’s next value from the current register values and the inputs. On every clock tick, every register loads its new value at once.

Designers write RTL as text in a (HDL), today mostly SystemVerilog or VHDL. The code is used in two ways. A simulator runs it on an ordinary computer to check that it behaves correctly. A tool translates it into a , a list of real gates and flip-flops and the wires between them, which later stages place and wire up on the chip.

RTL design sits between architecture and verification. The architecture has already fixed the big blocks, how they talk to each other, which clocks they use and how fast they must run. The RTL designer turns each block into code that is correct, that synthesis can build, and that is fast enough. That same code is then tested by the verification team, turned into gates, checked for clock and reset hazards (covered below), and read by every engineer who works on the block afterwards.

RTL fixes the cycle-level behavior of a block: what every register holds after every clock edge. Synthesis converts it into a gate-level netlist, and placement and routing turn that into layout. Each of those steps may restructure gates, resize them or move registers, but must keep the RTL’s cycle behavior. Equivalence checking enforces that: a formal tool proves that two versions of a design, such as the RTL and the synthesized netlist, compute the same results for every possible input, without running any tests.

That makes the RTL two things at once. It is the golden reference for what the chip does. It is also the biggest lever on : power, performance and area, meaning how much energy the chip uses, how fast its clock runs and how much silicon it takes. The implementation tools optimize inside the structure the RTL chose. They rarely rescue a design with too much logic between two registers, a large multiplier squeezed into one cycle, or an unsafe clock crossing.

Beyond plain correctness, the engineering at this stage is about four things, each covered below:

  • Writing code whose meaning in simulation and in synthesis agree. The languages were designed for simulation, so the two can differ.
  • Structuring logic so the chip can reach its target clock speed.
  • Making every crossing between clocks, and between reset signals, safe by construction.
  • Keeping the code checkable: clean under lint and crossing checks, reviewable, and paired with a description of how the chip saves power.
in3?register A?register B?register C× 2+ 1logiclogic= ?= ?waits for the edgewaits for the edgeclockticks: 0

At power-up the registers hold unknown values (?). Press Clock tick to send a rising edge.

Registers and the logic between them, sharing one clock. RTL describes exactly this: what each register holds after every tick. Values are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

The RTL designer starts from the spec, the plan for the blocks, and some ready-made parts. Out comes the design code. Testers check that it works. Then a tool turns it into real circuit parts.

DirectionWhatFormat
InSpecification and micro-architecture: block diagrams, interfaces, register mapsDocuments, machine-readable register descriptions
InReused and third-party IPRTL or encrypted RTL, plus memory and macro models
OutDesign sourceSystemVerilog .sv, Verilog .v or VHDL .vhd, packages and a file list (.f)
OutBlock timing constraintsSDC: clocks, I/O delays, clock groups
OutPower intentUPF (IEEE 1801)
OutCheck results and waiversLint, CDC and RDC reports with reviewed waiver files

A few entries in the table need explaining.

  • IP (intellectual property) means a ready-made block, such as a USB interface or a memory, bought from another company or reused from an earlier chip. It sometimes arrives as encrypted RTL, which tools can read but people can’t.
  • A file list (.f) names every source file, in the order tools should read them.
  • files tell the timing tools how fast each clock runs and when signals arrive at the block’s edges, so synthesis knows how much time each path has.
  • UPF describes which parts of the chip can be switched off to save power, and what happens to their signals when they are.
  • Lint, CDC and RDC reports are the results of automatic code checks, covered under “Checks before hand-off” below. A waiver is a written, reviewed decision that one particular warning is harmless.

Verification reads the same source files to build its simulation. Synthesis reads the source, the file list and the SDC.

Hand-offs are gated by checks. A typical release of RTL to synthesis requires:

  • No lint errors, apart from reviewed waivers.
  • A clean structural CDC and RDC run: every crossing between clocks, or between resets, has a recognized safe circuit.
  • A trial synthesis with no inferred latches and no unexpected black boxes. A black box is a module the tool couldn’t find or read and leaves empty, which is how a missing file otherwise goes unnoticed.
  • SDC constraints that define every clock.

The SDC also tells timing analysis which paths to skip, typically those between unrelated clocks (declared as clock groups or false paths). Skipping them is only safe if the RTL has a synchronizer on each such path, so the RTL owner reviews those declarations. Register maps, the list of control registers software can read and write and their addresses, are often generated from one description; the generated RTL, software headers and documentation must all come from the same revision.

inSpecificationArchitectureReused IPRTL designoutSource .svSDCUPFReportsVerificationSynthesis
Show what is read by

Three kinds of input, four kinds of output. Tap any box to read about it.

Inputs and outputs of RTL design, and which later stage reads which file.Share freely with credit: ‘Figure from chipfieldguide.com’

An RTL designer’s code describes two kinds of things.

  • Memory. Flip-flops hold values, such as a score or a mode. They change only when the clock ticks.
  • Logic. Circuits that add, compare and choose work out new values from the stored ones. Logic reacts right away and remembers nothing.

A circuit that steps through a fixed list of modes is called a . A traffic light is one: green, then yellow, then red, then green again.

The language has traps. Say the code forgets to tell a signal what to do in one situation. Then the tool quietly adds a memory to hold the old value, which nobody asked for. Designers follow strict habits to avoid traps like this.

Languages and the synthesizable subset

Verilog and VHDL were created to simulate circuits and only later adopted to build them, so the official meaning of the code is whatever a simulator does with it. Synthesis tools accept only part of each language, the . Yosys, an open-source synthesis tool, does not support the non-synthesizable features listed in the standard IEEE 1364.1, and in SystemVerilog mode it accepts the keywords always_ff, always_comb, always_latch and logic used below. A delay such as #5 (“wait 5 time units”), reading files and printing messages all make sense in a simulation but have no hardware meaning. They belong in a testbench, the separate code that tests the design.

Modules, ports, parameters and generate

The unit of design is a module: a block with named inputs and outputs (its ports), like a component on a circuit board. Modules contain other modules, which builds a hierarchy: a chip contains a processor, which contains an adder. Parameters make a module reusable: a queue with a DEPTH parameter can be built 4 entries deep in one place and 64 in another. generate loops and conditionals copy or select hardware while the design is being assembled, for example one identical lane per bit of a bus. Style guides ask for every generate block to have a name, so the names of the signals inside it come out the same in every tool.

Clocked and combinational blocks

SystemVerilog writes the two kinds of circuit differently:

two_kinds.sv (illustrative)systemverilog
// Memory: runs once at each rising clock edge
always_ff @(posedge clk) begin
  count <= count + 1;
end

// Logic: re-evaluated whenever an input changes
always_comb begin
  y = sel ? a : b;
end
  1. 1L2“posedge clk” means “at each rising edge of the clock”. Every variable assigned here becomes flip-flops.
  2. 2L3<= is the non-blocking assignment, used for flip-flops. This line makes a counter that adds 1 on every tick.
  3. 3L7always_comb describes logic with no memory.
  4. 4L8“sel ? a : b” means “if sel is 1 then a, otherwise b”: a multiplexer, which picks one of two inputs.

always_ff @(posedge clk) describes : the block runs at each rising clock edge, and every variable it assigns becomes stored state. always_comb describes , logic with no memory, and re-runs whenever one of its inputs changes. The two use different assignment signs, and two rules prevent most problems:

  • Clocked blocks use the <=. The simulator reads every right-hand side first and updates the registers afterwards, all together, the way real flip-flops all take in their inputs at the same edge.
  • Combinational blocks use the =, which takes effect at once so the next line sees the new value, as in ordinary software.

Break the first rule and hardware disappears. Suppose a designer wants a three-stage shift register, where a value moves one stage further on each tick: q1 takes the input d, q2 takes the old q1, and q3 the old q2. Written with <=, that is what happens. Written with = as q1 = d; q2 = q1; q3 = q2;, each line sees the value the line above has just written, so d reaches q3 in a single tick and synthesis builds a single register.

The second trap is the accidental memory. A combinational block must give every output a value in every situation. always_comb if (en) q = d; says what q is when en is 1 but not when it is 0. Then q must keep its old value, keeping a value needs memory, and the tool builds an that nobody asked for. The fix is to assign a default value at the top of the block, and to give every case statement (a multi-way choice) a default branch.

State machines

Control logic, the part that decides what happens next, is usually a . A traffic light is the classic example: it is always in one of a few states (green, yellow, red) and moves to the next one on a timer. In hardware, a register holds the current state, and logic computes the next state from the current state and the inputs. In a Moore machine the outputs depend only on the state; in a Mealy machine they also depend on the inputs. The common style uses two blocks: always_ff for the state register and always_comb for the next state and the outputs.

Each state needs a bit pattern. With binary encoding, 10 states fit in 4 flip-flops, since 4 bits give 16 patterns. With encoding, each state gets its own flip-flop and exactly one is 1 at a time: 10 flip-flops, but asking “are we in state X?” means reading a single bit, so the next-state logic is simpler. Passing the outputs through a clocked block as well removes glitches (brief false pulses while the state bits change).

Coding for timing

Signals take time to pass through logic. Between two clock edges, the result of every path from one register to the next must arrive in time. The slowest such path, the , sets how fast the clock can run. An illustrative example: a 32-bit multiply followed by an add takes 1.5 ns, so the clock can tick at most every 1.5 ns, about 670 MHz. puts a register between the multiply and the add. Each half takes about 0.75 ns, so the clock can run nearly twice as fast (a bit less in practice, since the extra register adds a little delay of its own). Each result now takes two ticks instead of one, but a new result still finishes every tick. A processor pipeline, and the hazards that come with it, is the subject of Instructions and pipelines.

Other habits that help:

  • Put a register on each output of a block, so the next block gets a full clock period to work with.
  • Avoid long if / else-if chains when only one condition can be true at a time. A chain tests the conditions one after another; a parallel choice tests them all at once.
  • Keep wide comparisons and large multiplexers off paths that are already long.

Synthesis can shift existing registers a little to balance the work between them (), but it cannot add a clock cycle of delay. Only the designer knows whether the rest of the system can wait one more tick.

Semantics first

An HDL’s official meaning is its simulation meaning; synthesis infers hardware from a subset whose hardware meaning the tools agree on. RTL style rules exist to keep those two views aligned. SystemVerilog has three kinds of always block that state the designer’s intent, so tools can check it:

  • always_ff should contain only (<=) updates to flip-flops.
  • always_comb should contain only (=) assignments and must assign every output on every path.
  • always_latch marks the rare latch that is intended.

Two habits make simulation results depend on the order in which the simulator happens to run blocks: mixing = and <= in one block, and assigning the same variable from more than one block. Synthesis builds the same hardware whatever that order, so the simulation can disagree with the gates. A missing default in always_comb produces an . Lint tools flag it, and Verilator’s warning suggests always_latch in case the latch was meant.

unique case needs care. The keyword makes the simulator check that exactly one branch matches, which catches mistakes. It also lets synthesis rely on that promise. Yosys, for example, turns unique and priority into the full_case and parallel_case annotations, which tell synthesis that the listed branches cover every possible value and that no two can match at once, so it can drop the logic for anything else. lowRISC’s style guide forbids writing those annotations by hand because they so easily cause simulation/synthesis mismatches. So when a unique check fires in simulation, the gates may not do what the simulation did. Treat it as a real bug, not noise.

Structure and reuse

Parameterized modules, packages (shared files of type and constant definitions) and named generate blocks are SystemVerilog’s tools for reuse. They are limited: there is little real programming available while the design is being assembled. That is why many teams generate RTL from scripts, or write it in a language embedded in a full programming language, such as Chisel (see “Other ways to write RTL”). Naming conventions pay off in debugging. The lowRISC guide uses _d and _q for a register’s input and output (named after a flip-flop’s D and Q pins), _q2 for the same signal two cycles later, _i and _o for module inputs and outputs, and a clk_ prefix that names each extra clock and marks the signals that belong to it.

Control logic

Write as a state register plus a combinational next-state block, with the state values defined in one place (an enum, SystemVerilog’s type for a list of named values, or a set of parameters). The encoding can then change without touching the transition logic. Synthesis may change it anyway. Yosys’s fsm command finds state machines, extracts them and re-encodes them (at present always as one-hot), and can write a file listing the new encoding so that an equivalence checker can still match the netlist to the RTL. Register any FSM output that drives something glitch-sensitive, such as a clock, a reset or an interface to another clock domain, because outputs decoded from the state can glitch while the state bits change.

Timing-aware RTL

Synthesis optimizes the logic you write, inside the structure you chose. The structural decisions that matter most for clock speed are made here:

  • Where the pipeline registers go, and how the control signals that travel with the data follow them. A valid bit marks which stages hold real data; a stall signal holds a stage when the next one isn’t ready.
  • Whether a priority chain (test A, else B, else C) can become a parallel selection, when the conditions are known never to be true together.
  • Whether a wide comparison can be computed a cycle early and stored in a register.
  • Whether an enable that fans out to thousands of flip-flops is registered and copied, so that no single wire has to drive them all.

changes what happens on each cycle, so it is an RTL decision. in synthesis can only move registers that already exist, and it makes equivalence checking harder, because the register boundaries in the netlist no longer line up with the RTL.

Clock gating, the main power saving available at this level, is usually left to the tools. When a bank of flip-flops has an enable (“load a new value only when en is 1”), synthesis can replace the enable with a gate on the clock itself: an integrated clock-gating (ICG) cell that stops the clock to those flops whenever en is 0, so they stop switching and stop burning power. Yosys’s clockgate command does exactly this, with an option to tie off the cell’s test pin for scan testing. So write clean enables on register banks. Instantiate gating cells by hand only for coarse gating of whole blocks, and give each one a way for test mode to force the clock on.

State diagramtimerGREEN00YELLOW01RED10Hardwarestate register00next-statelogicloaded at each tickdecode
State encoding

State GREEN, stored as 00. Press Time’s up to send the timer pulse and a clock tick.

A traffic-light state machine: the state diagram (left) and its hardware (right): a state register, next-state logic, and lamps decoded from the state. Switch the encoding to see the flip-flops change.Share freely with credit: ‘Figure from chipfieldguide.com’

Big chips often have several clocks, each ticking at its own pace. Think of two drummers playing different beats. When a signal passes from one clock’s area to the other, it can arrive just as the receiver looks at it. The receiver can’t tell if it saw a 0 or a 1. This is a .

The fix is to give the signal a moment to settle. Two memory cells in a row hold it for an extra tick before the rest of the chip uses it.

Reset strategy

When a chip powers up, its flip-flops hold random values. A reset signal forces them into known starting values. There are two ways to wire it.

  • Synchronous reset is treated like any other input: the flip-flop loads its reset value at the next clock edge. It needs the clock to be running and adds a little logic in front of every flip-flop.
  • Asynchronous reset uses a dedicated reset pin on the flip-flop. It takes effect at once, even with no clock, and adds no logic in front of the flip-flop.

The danger with asynchronous reset is the moment it ends. If reset is released just before a clock edge, a flip-flop may not have enough time to recover (it misses its ) and can end up undecided between 0 and 1. Some flip-flops may also leave reset on this tick and others on the next, so the design starts in a state it was never meant to be in. The standard answer is a : a pair of flip-flops that passes the start of reset through immediately but lines its end up with a clock edge. Each clock gets its own. Many teams pick one convention for the whole chip. lowRISC, for example, uses active-low asynchronous resets everywhere (“active-low” means reset is on when the signal is 0).

Clock domain crossings

Large chips have several clocks: an interface may run at the rate its standard demands while the processor runs as fast as it can. All the flip-flops on one clock form a . Inside a domain, timing is predictable, and tools check that every signal arrives in time and stays steady while it is captured. That matters because a flip-flop needs its input steady for a short window around the clock edge, around 20 picoseconds (trillionths of a second) in a modern process.

Between unrelated clocks, nothing lines up, and sooner or later a signal will change inside the receiving flip-flop’s window. The flip-flop can then enter : an undecided state between 0 and 1 that usually resolves within a fraction of a nanosecond, but with no guaranteed limit. Timing tools can’t prevent this, because the clocks have no fixed relationship for them to check. The RTL has to include a safe circuit:

  • A single bit that changes rarely: a two-flop on the receiving clock. The first flip-flop may become undecided; the second reads it one tick later, after it has almost certainly settled.
  • An event, or a group of control bits: a request/acknowledge handshake. The sender holds its data steady and raises a request. Only the request and the reply pass through synchronizers, and the sender changes the data only after the reply has come back.
  • A stream of data: an (first in, first out), a small queue that one clock writes and the other reads. Only its two position counters cross between the clocks, in , a way of counting where only one bit changes per step.

Never send a multi-bit value through separate synchronizers, one per bit. Each bit may settle on a different tick, so for a moment the receiver sees a mix of old and new bits: a value that was never sent. If a counter steps from 7 (binary 0111) to 8 (1000), all four bits change, and the receiver might briefly read 15 (1111) or 0 (0000).

Reset domain crossings

Resets cause a similar problem. Parts of a chip often have their own reset, so one block can be reset while its neighbor keeps running. A is a path from a flip-flop on one reset to a flip-flop on another. When the first block’s reset hits, its outputs jump to their reset values at a moment unrelated to the neighbor’s clock, and the neighbor’s flip-flops can be caught mid-change, even though both blocks share one clock. Designers fix it by ordering the resets so the receiver enters reset first, or by blocking the receiver’s input while the source is in reset.

Resets

Neither reset style is free.

  • Synchronous reset puts reset logic in front of every flip-flop’s data input and needs a running clock, which fails if the clock happens to be gated off when reset arrives, and it takes effect only at a clock edge.
  • Asynchronous reset keeps the data path clean and works with no clock. Its release is asynchronous unless synchronized, and if the reset doesn’t come straight from a chip pin, the test logic must be able to control it. During scan testing (see Design for test) the flip-flops are chained into long shift registers, and a reset firing on its own would wipe the test pattern.

Either way, a reset net is a large tree, built much like a clock tree. Unlike the clock, it doesn’t need equal arrival times everywhere, but its release must reach every flop within one clock cycle and meet each flop’s times, the reset-pin equivalents of setup and hold. Each asynchronous clock domain needs its own , and if the order in which domains leave reset matters, that sequence has to be designed in explicitly. Not every flop needs a reset. Datapath registers whose contents only count when a valid bit says so can skip it, saving area and wiring, but then simulation must be able to show what their random start-up values do (see “What goes wrong”).

CDC schemes, step by step

A two-flop is only half a CDC design. The rest, following Ginosar’s tutorial:

  1. The source must hold the signal long enough to be sampled. A one-cycle pulse from a fast clock into a slow one can fall entirely between two receiving edges and vanish. The sender can’t know how long is long enough without knowing the other clock, so a request/acknowledge handshake closes the loop: the sender holds the request until the acknowledgment comes back.
  2. That handshake is slow. It can take two receiving-clock cycles to see the request, two sending-clock cycles to see the acknowledgment and a cycle on each side to react, then about as many again to lower both signals: around a dozen cycles per transfer.
  3. When throughput matters, the two-clock FIFO is the usual answer, and it synchronizes only its pointers. In Gray code a sampled pointer is either the old value or the new one, so the full and empty flags may be a cycle or two late, but are never wrong.
  4. The Gray value must come straight from a register, because the logic converting binary to Gray can pass through out-of-sequence codes while its input bits arrive at slightly different times.

Two more rules from the same tutorial. Never feed one asynchronous input into two separate synchronizers: one may resolve to 1 and the other to 0, and the logic downstream sees a contradiction. And place the two synchronizer flops next to each other, because any wire delay between them comes straight out of the settling time.

Timing analysis is normally told to skip paths between asynchronous clocks, through clock-group or false-path declarations in the SDC. That is right, since there is no fixed timing relationship to check, but it means no tool checks timing on those paths at all. The safety of each crossing rests entirely on the RTL scheme and on the structural CDC checks in the next section.

RDC

A needs no second clock. Suppose blocks A and B share a clock and B reads a register in A. If A’s asynchronous reset asserts while B is running, A’s register changes at an arbitrary moment relative to the clock, and B’s capturing flop can go metastable or load a corrupted value. The fixes:

  • Reset ordering: make sure B enters reset before A, so B is already held in reset when A’s outputs jump.
  • Isolation: an enable signal in B’s domain blocks B’s data input, or stops B’s clock, while A is in reset.
  • Staged resets: structures that assert and release the resets of neighboring blocks in a fixed sequence.

The numbers can be large. In the DVCon paper that describes this methodology, an RDC tool flagged 185 asynchronous reset-crossing issues within single clock domains on one design and 15,519 on another. The reset architecture (which resets exist, their order and the isolation enables) therefore belongs in the specification, so the RDC tool has something to check the design against.

windowclk Bdata (A)flopflop 2← logic reads this
Receiver

The change was outside the window around B’s edge, so the flop captured it cleanly. Slide it into the window.

A signal from clock domain A captured on clock B. The shaded bands are the window around each edge where the input must be steady (about 20 ps in a modern process; drawn far wider here). Times in clock periods T.Share freely with credit: ‘Figure from chipfieldguide.com’
before0111= 7after1000= 84 bits change (binary)syncreceiver1new1old1old1old= 15 ✗possible0123456789101112131415
Encoding

The receiver sees 1111, which is 15: a value that was never sent. The bits settled on different ticks.

A counter value crossing bit by bit, each bit through its own synchronizer. Tap the changing bits to choose which have settled. In Gray code only one bit changes per step, so the receiver can’t see a value that was never sent.Share freely with credit: ‘Figure from chipfieldguide.com’

Before anyone runs long tests, the design code goes through quick automatic checks.

  • works like a spell-checker for chip code. It flags patterns that look like mistakes, such as a memory added by accident.
  • Crossing checks look at every place a signal moves between clocks. They make sure a safe circuit is there.

Like an essay going to an editor, every change is also read by a teammate before it is accepted.

Lint

programs read the source code and report likely bugs without running it, the way a grammar checker reads an essay. Verilator, an open-source tool, does this in its --lint-only mode, and its -Wall option adds style warnings that are off by default. Each warning carries a short code. Typical ones:

  • LATCH: an accidental memory was inferred.
  • WIDTH: numbers of different sizes were combined, for example an 8-bit value stored into 4 bits, losing the top half.
  • CASEINCOMPLETE: a multi-way choice doesn’t cover every possible value.
  • BLKSEQ: a blocking = inside a clocked block.
  • UNUSEDSIGNAL: a signal that nothing reads.

Verible, another open-source project, adds a style linter and a formatter for SystemVerilog.

CDC and RDC checks

Ordinary simulation can’t catch crossing bugs. A simulated flip-flop never becomes undecided: when its input changes right at the clock edge, the simulation always captures one clean value, while real silicon might capture the other or settle a cycle late. Timing analysis misses them too. Instead, static CDC tools read the design, find every path between clock domains and check its structure: a synchronizer missing or in the wrong place, logic between domains that could glitch, and signals synchronized separately that later meet again. RDC tools do the same for paths between reset domains.

Power intent

To save power, many chips can switch parts of themselves off or run them at a lower voltage. The designer describes that in a separate file written in , the Unified Power Format (IEEE 1801), rather than in the RTL. It lists the power domains (regions powered together), their supplies and power switches, and the special cells needed at their edges: isolation cells that hold a switched-off block’s outputs at a fixed value, level shifters that translate signals between voltages, and retention registers that keep their contents while the power is off. The same RTL is verified and implemented together with this file. Power planning covers how that intent becomes physical.

Version control and review

RTL lives in a version-control system such as Git, which records every change and who made it. Every change is reviewed by a colleague, and an automated system (continuous integration) reruns lint, a quick set of simulations and often a trial synthesis on each one. Reviewers look for what tools miss: wrong reset values, unusual sequences of events at an interface, unclear names.

Lint as a gate

Run on every commit with a fixed rule set. When a warning is genuinely harmless, waive that one line, with a tool comment (Verilator’s lint_off) or a control file, and record the reason. Verilator’s documentation warns against switching a warning off for the whole design, because that removes the check everywhere, including from code not yet written. By default Verilator stops on any warning, which suits a continuous-integration gate. Whatever the tool, the classes that matter most are the same:

  • Latches.
  • Width truncation and extension: bits silently dropped or padded.
  • Multiple drivers: one signal assigned from two places.
  • Combinational loops: logic that feeds back to its own input with no register in between.
  • Incomplete case statements.
  • Naming and structure rules that keep the code reviewable.

Structural and functional CDC

A CDC sign-off typically runs in stages:

  1. Structural. A static tool reads the RTL, the clock definitions and constraints the designer writes: which paths are false, which signals are static (set once at start-up and never changed while running) or constant. It reports crossings with no synchronizer, synchronizers in the wrong place, combinational logic in front of a synchronizer (which can glitch, and the glitch can be captured), and signals that split into separate synchronizers and later reconverge.
  2. Functional. Assertions, properties checked during simulation or proved by a formal tool, confirm that the protocol around each synchronizer is followed: that a signal declared static really never changes, or that data is held long enough to be captured.
  3. Metastability injection. The simulation or formal model of each synchronizer is made to delay its output by an extra cycle at random, as silicon might, to show the logic downstream still works.

The constraints are themselves a risk. Writing them by hand is error-prone, and a wrong one, such as a signal declared static that actually changes, makes the structural tool pass a real crossing, and the bug reaches silicon. Applied to an Ethernet controller, an SPI block and shared library cells, the flow in that paper found missing synchronizers, logic on crossing paths and exactly that kind of wrong static declaration. Review constraints like code. The same discipline applies to RDC: declare the reset domains and their intended order, then check that every crossing is isolated or ordered.

Power intent alongside RTL

keeps power intent separate from functional intent. The RTL says what the logic does; the UPF says which power domains exist, how they are supplied and switched, and where isolation, level shifting and retention go. It can be written in steps. An IP provider ships constraints with a block, such as the value each output must be clamped to when the block is off; the chip team adds the configuration; implementation adds the choice of cells. For the RTL designer, power intent creates obligations in the code:

  • Every signal leaving a block that can be switched off has a defined clamp value, and the logic receiving it must behave sensibly when it sees that value.
  • Retention needs a controller that tells retention registers to save their contents before power-down and restore them after power-up, in the right order.
  • Resets and clocks must behave sensibly while a domain powers up.

Verification runs the RTL together with the UPF, so these sequences are exercised before implementation inserts any cells.

commitLintevery change✓CDC / RDCevery crossing✓Reviewintent, names✓✓mergedon to verification and synthesis
The change contains

Clean: lint passes, every crossing has a synchronizer, and the reviewer approves. The change is merged.

The checks a change passes before hand-off, in order: lint, clock- and reset-crossing checks, then review. Each catches a different kind of mistake.Share freely with credit: ‘Figure from chipfieldguide.com’

Not every team writes every line by hand. Some write a program that writes the design for them. Others describe the job in an ordinary programming language and let a tool work out the circuit. Either way, the result is the same kind of RTL, so the later steps don’t change.

  • (HLS) turns an algorithm written in C, C++ or SystemC (a C++ library for modeling hardware) into RTL. The input has no notion of clock ticks; the tool decides which operations happen in which tick and how much hardware to build.
  • Chisel builds hardware from inside Scala, a general-purpose programming language, and writes the result out as Verilog. It aims at generators: programs that produce many sizes and variants of a design from a few parameters.
  • SpinalHDL, also built on Scala, checks while assembling the design that no register takes input from a different clock domain unless the designer has added a synchronizer or marked the crossing as intended.
  • Amaranth is a Python library for RTL modeling. Its designs can be simulated, synthesized directly through Yosys, or written out as Verilog.

All of them end in Verilog or a netlist, so lint, synthesis and the rest of the flow apply unchanged. Debugging, though, happens on generated code, whose signal names may not match what the designer wrote.

works in three steps. Scheduling splits the algorithm into control steps, each fitting in one clock cycle. Allocation decides how many adders, multipliers, registers and memories to build. Binding assigns each operation and value to one of them. The same algorithm can take many cycles on little hardware or few cycles on a lot, and the designer’s directives pick the point. The designer still writes the block’s function and its interface protocol; the tool works out the micro-architecture in between. Logic whose value lies in exact cycle-by-cycle behavior, such as interfaces and tight control, is often still written by hand.

Hardware construction languages keep RTL’s model of registers and clock cycles but use a full programming language to build the design. Chisel brought object orientation, functional programming and parameterized types, and generates either a fast C++ simulator or Verilog for FPGA and ASIC flows. Because the design is built by a real program, that program can check rules as it goes. SpinalHDL raises a CLOCK CROSSING VIOLATION when a register depends on another clock domain without a tag or a synchronizer. Amaranth has first-class clock domains and ships clock domain crossing primitives and FIFOs in its standard library.

The costs are integration with the rest of the tool flow, and the readability of generated names wherever engineers read them: timing reports, late netlist patches (ECOs) and CDC waivers. Teams that use generators usually pin the generator version and check in the exact Verilog that went to synthesis. sv2v solves a narrower problem: it converts synthesizable SystemVerilog into the older Verilog-2005, and was first built so that Yosys could read SystemVerilog designs.

SystemVerilogwritten by handC / C++untimed algorithmHLS toolpicks cyclesChiselScala generatorElaborateemits VerilogAmaranthPython libraryElaborateemits VerilogVerilogRTL .v/.svLintsynthesisassign y = a + b;one adder, this route
Route

Hand-written SystemVerilog goes straight to lint and synthesis: what you wrote is what the tools read.

Four ways to produce RTL. Every route ends in Verilog, so lint, synthesis and the rest of the flow apply unchanged. The adder snippets are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Pick one of four small designs. Each shows a sketch of the circuit and a timeline of its last 16 steps. The short names, like en (enable, meaning “on”), are the circuit’s inputs and outputs.

  • Counter: press Clock tick and it counts from 0 to 15, then starts over. Flip rst (reset) to send it back to 0. Turn en off to freeze it.
  • Mux: a chooser. Flip a, b and sel. The output y changes at once, with no clock at all.
  • Latch bug: turn en off, and q keeps its old value. That is a memory added by accident.
  • Blocking vs non-blocking: two ways to write the same line of three boxes. One passes a value along one box per tick. The other fills all three boxes at once.

Four small designs, each with its SystemVerilog code (read-only), a sketch of the hardware and a waveform: a chart of every signal’s value over the last 16 steps.

  • Counter is an always_ff block that counts from 0 to 15. Press Clock tick, then toggle rst (reset, back to 0 on the next tick) and en (enable: count only while it is 1).
  • Mux is an always_comb two-input multiplexer. Toggle a, b and sel, and notice that y changes with no clock tick at all.
  • In Latch bug, set en to 0 and watch q hold its value: an .
  • In Blocking vs non-blocking, clock both shift registers. With <= a value needs three ticks to reach the end; with = all three stages load the input on the same tick.

The Expert view adds a lint panel for every preset and the synthesis result: the cells synthesis would build. Use it to connect the code, the simulation and the hardware:

  • In Latch bug (always_comb if (en) q = d;), the lint panel reports “Latch inferred for q”, and the cell list shows the latch.
  • In Blocking vs non-blocking, compare the number of registers inferred for the <= and = versions. With = the three stages collapse to one, the classic delay-line example.
  • In Counter, check how the synchronous reset and the enable turn into logic in front of the flip-flops’ data inputs.
Loading simulation…

Here is what an RTL designer’s day often looks like.

  1. Write or change code for one block, following the team’s style rules.
  2. Run the quick checks. Lint finishes in seconds. Any warning gets fixed right away.
  3. Run a few tests and read the timelines that show each signal flipping between 0 and 1.
  4. Send the change for review. A teammate reads and approves it.

Three illustrative files:

  • A small control state machine. It waits for start, stays busy until done, then raises an interrupt (irq, a signal that asks a processor for attention) for one tick.
  • A two-flop synchronizer, and the reset synchronizer its clock domain needs.
  • The output of Verilator’s lint on a buggy version of the state machine.

The names follow a common convention. Names ending in _q are register outputs (the current value), _d the value the register will take at the next tick, _i and _o mark inputs and outputs, and _n marks a signal that is active when it is 0, as in rst_ni: “reset, active low, input”.

Illustrative code in a lowRISC-like style (_i/_o ports, _d/_q registers, active-low asynchronous reset). The lint log follows Verilator’s format; exact text varies by version.

EditLintSimulateReviewfix→ mergectrl_fsm.svalways_comb begin state_d = state_q; busy_o = 1'b0; unique case (state_q) ...← irq_o default missing
1 / 5

Write the change: here, the output logic of the ctrl_fsm state machine. One output default has been forgotten.

The inner loop on the chapter’s example state machine: edit, lint, simulate, review. Lint catches the forgotten default in seconds. Illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
ctrl_fsm.sv (illustrative)systemverilog
module ctrl_fsm (
  input  logic clk_i,
  input  logic rst_ni,
  input  logic start_i,
  input  logic done_i,
  output logic busy_o,
  output logic irq_o
);
  typedef enum logic [1:0] {IDLE, RUN, FLUSH} state_e;
  state_e state_q, state_d;

  always_ff @(posedge clk_i or negedge rst_ni) begin
    if (!rst_ni) state_q <= IDLE;
    else         state_q <= state_d;
  end

  always_comb begin
    state_d = state_q;
    busy_o  = 1'b0;
    irq_o   = 1'b0;
    unique case (state_q)
      IDLE:    if (start_i) state_d = RUN;
      RUN: begin
        busy_o = 1'b1;
        if (done_i) state_d = FLUSH;
      end
      FLUSH: begin
        busy_o  = 1'b1;
        irq_o   = 1'b1;
        state_d = IDLE;
      end
      default: state_d = IDLE;
    endcase
  end
endmodule
  1. 1L3Active-low asynchronous reset, named with _n and _i suffixes.
  2. 2L9States defined once. Change the encoding here and nothing else moves.
  3. 3L12State register: always_ff, non-blocking assignments only.
  4. 4L17Next-state and output logic: always_comb, blocking assignments only.
  5. 5L18Defaults first. Every output is assigned on every path, so no latch can be inferred.
  6. 6L21unique adds a simulation check that exactly one item matches.
  7. 7L32The 2-bit encoding has a fourth, unused code. The default sends it back to IDLE.
sync_2ff.sv (illustrative)systemverilog
module sync_2ff #(
  parameter logic RESET_VAL = 1'b0
) (
  input  logic clk_dst_i,
  input  logic rst_dst_ni,
  input  logic d_async_i,
  output logic q_o
);
  logic meta_q;

  always_ff @(posedge clk_dst_i or negedge rst_dst_ni) begin
    if (!rst_dst_ni) begin
      meta_q <= RESET_VAL;
      q_o    <= RESET_VAL;
    end else begin
      meta_q <= d_async_i;
      q_o    <= meta_q;
    end
  end
endmodule

module rst_sync (
  input  logic clk_i,
  input  logic rst_async_ni,
  output logic rst_sync_no
);
  logic stage_q;

  always_ff @(posedge clk_i or negedge rst_async_ni) begin
    if (!rst_async_ni) {rst_sync_no, stage_q} <= 2'b00;
    else               {rst_sync_no, stage_q} <= {stage_q, 1'b1};
  end
endmodule
  1. 1L5This domain’s reset, itself produced by an rst_sync in the same clock domain.
  2. 2L6Must come straight from a flop in the source domain. Logic here could glitch, and the glitch could be captured.
  3. 3L16First stage: may go metastable when d_async_i changes near the clock edge.
  4. 4L17Second stage: samples a cycle later, after the first has almost certainly resolved. Keep the two flops adjacent.
  5. 5L30Reset asserts immediately, without waiting for a clock.
  6. 6L31Release shifts a 1 through two flops, so the domain leaves reset on a clock edge, two edges after the input deasserts.
verilator --lint-only -Wall (illustrative)log
$ verilator --lint-only -Wall rtl/ctrl_fsm.sv
%Warning-LATCH: rtl/ctrl_fsm.sv:17:3: Latch inferred for signal 'ctrl_fsm.irq_o' (not all control paths of combinational always assign a value)
                                    : ... Suggest use of always_latch for intentional latches
%Warning-CASEINCOMPLETE: rtl/ctrl_fsm.sv:21:5: Case values incompletely covered (example pattern 0x3)
%Warning-WIDTHTRUNC: rtl/ctrl_fsm.sv:44:15: Operator ASSIGN expects 4 bits on the Assign RHS, but Assign RHS's ADD generates 5 bits.
%Warning-BLKSEQ: rtl/ctrl_fsm.sv:44:13: Blocking assignment '=' in sequential logic process
                                      : ... Suggest using delayed assignment '<='
%Warning-UNUSEDSIGNAL: rtl/ctrl_fsm.sv:6:22: Signal is not used: 'dbg_i'
%Error: Exiting due to 5 warning(s)
  1. 1L1Lint only, no model built. -Wall turns on style warnings such as BLKSEQ and UNUSEDSIGNAL that are off by default.
  2. 2L2The buggy version dropped the irq_o = 1'b0 default, so irq_o holds its value in RUN and IDLE.
  3. 3L4The buggy version used a plain case with no default, leaving the fourth code (2'b11) uncovered. A unique case on an enum that lists every value would count as complete.
  4. 4L5The buggy version also has a 4-bit timeout counter that adds a 5-bit step, so the sum is truncated. Size the operands or cast the result.
  5. 5L6The same counter is updated with = inside always_ff. It simulates, but invites races.
  6. 6L8A debug input left unconnected. Remove it, or collect it into an intentional _unused signal.
  7. 7L9Warnings are fatal by default, so CI fails until each is fixed or waived with a reason.

Read the state machine top to bottom: one register block, then one combinational block that starts by giving every output a default value, then a case on the state with a default branch. The synchronizer file pairs a data synchronizer with the reset synchronizer for the same clock. Each lint warning carries a code (LATCH, WIDTHTRUNC) that you can look up in Verilator’s documentation. Fix the code where you can, and waive a warning only with a written reason.

Open-source tools cover the whole inner loop: Verilator for lint and fast simulation, Icarus Verilog for quick simulations (compile with iverilog, run with vvp), Verible for style, and Yosys for a trial synthesis.

The synchronizer here is written as two ordinary flip-flops. Many flows wrap it in one shared module used everywhere, so a single declaration tells the CDC tool which structure counts as a synchronizer, and one edit can switch to three stages for fast clocks or low voltage. The reset synchronizer is the classic asynchronous-assert, synchronous-release circuit. Its second flop cannot go metastable when reset releases, even if the release misses recovery time, because its input and output are both 0 at that moment: there is no change for it to capture.

  • The test says yes, the chip says no. Some code acts one way in a test but builds a different circuit. Strict habits and checks catch most of these.
  • Memories by accident. One forgotten situation in the code adds a memory nobody asked for.
  • Unsafe clock crossings. These bugs might strike once a month on a real chip and almost never in a test. Special checkers hunt for them.
  • Too slow. One step with too much work slows the whole chip. The fix is to split that step in two.
  • Simulation/synthesis mismatch. Some code simulates one way and builds different hardware. A blocking = in a clocked block can collapse a pipeline, and delays such as #5 and other non-synthesizable features are ignored or rejected by synthesis, so the tests ran something the chip won’t do.
  • Inferred latches. Caught by lint (LATCH) and by synthesis warnings.
  • Width bugs. Adding two 8-bit numbers can need 9 bits: 255 + 1 = 256. Store the sum in 8 bits and the carry is silently dropped, giving 0. Lint’s WIDTH warnings find most of these.
  • Missing or wrong synchronizers. Multi-bit values synchronized bit by bit, short pulses lost on the way into a slower clock, logic placed in front of a synchronizer.
  • Reset bugs. A register that should be reset isn’t, or reset ends on different ticks in different parts of a block. Simulators can hide the first: they mark a value nobody set as X (unknown), and some code treats X in a way that happens to look correct.
  • Timing found too late. Run a trial synthesis early. A path that misses its time budget by 30% won’t be rescued by tool settings; the RTL has to change, for example by adding a pipeline stage.
  • Scheduling races. Blocking assignments to related variables from separate clocked blocks give simulation results that depend on which block the simulator runs first, since the standard lets it run them in any order. The hardware has no such ambiguity, so RTL simulation and gate-level simulation can disagree.
  • X-optimism. RTL simulators usually use four values per bit: 0, 1, X (unknown) and Z (undriven). A register with no reset starts as X. if (x) then quietly takes the else branch, and a case on X matches no branch, so a missing reset can look harmless. Run some regressions in a two-state simulator that gives such registers real, random start-up values instead: Verilator with --x-assign unique and --x-initial unique, the +verilator+rand+reset+2 run-time option and a different seed each run.
  • unique and priority misuse. These keywords let synthesis drop logic for cases declared impossible (see “How it works”). If the simulation check would have fired, the gates differ from the RTL.
  • Reconvergence. Two signals synchronized separately can arrive a cycle apart, so logic that combines them sees a combination the source never produced. Synchronize one encoded value instead, or use a handshake.
  • CDC constraint rot. Static-signal declarations and clock-group exceptions written early and never revisited silence real crossings later. Review them at each release.
  • Reset domains nobody drew. Every block that can be reset on its own while its neighbors run creates reset crossings, and large chips have many asynchronous reset domains. Each crossing needs isolation or ordering.
  • Waiver creep. Global warning disables and bulk waivers hide new instances of a problem. Waive by line, with a reason, and review the waiver file.
  • Hierarchy drift. An unnamed generate block gets a name the tool makes up, and tools make up different names. CDC waivers, UPF scopes and SDC paths written against the RTL hierarchy then stop matching. Name every generate block.
reg32-bit ×+reg1.5 ns0123456nsclock: every 1.5 ns× +#1#2#3#4● result: 4 in 6 ns
Design

One stage: the multiply and add take 1.5 ns together, so the clock can tick at most every 1.5 ns, ≈ 670 MHz. 4 results in 6 ns.

Pipelining, with the chapter’s illustrative numbers: a multiply then an add between two registers. Top: the hardware and each path’s delay. Bottom: when each operation runs.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

Statement regions per SystemVerilog time slot
9
Synchronizer MTBF, 1 GHz, τ = 10 ps (Ginosar example)
≈ 4 × 10²⁹ years
Same circuit, stressed corner: τ = 100 ps, data at 1 kHz
≈ 1 minute
Verilator slowdown without activity gating (large designs)
≈ 100×

Sources: the LSU EE 4755 notes for the region count, Ginosar for the MTBF examples, and the Verilator internals document for the activity-gating figure.

Event-driven simulation and the stratified event queue

SystemVerilog’s meaning is defined by a reference simulation algorithm in the IEEE 1800 standard, which Accellera makes available at no charge through the IEEE GET program. Simulated time advances in steps called time slots, one per instant at which anything happens. Each slot is processed in ordered regions, and the simulator loops back through them until nothing is left to do at that instant. IEEE 1800-2012 defines nine regions for executing simulation code (the others are hooks for C-language callbacks):

  • Preponed: sample values for concurrent assertions.
  • Active, Inactive and NBA: the active region set, used by RTL.
  • Observed: evaluate concurrent assertions.
  • Reactive, Re-Inactive and Re-NBA: the reactive region set, for testbench code in program blocks.
  • Postponed: monitoring of the final values, for example by $monitor.

The Active region executes, in any order the simulator likes:

  • Blocking assignments.
  • Continuous assignments (assign).
  • Evaluation of built-in gate primitives.
  • The right-hand sides of non-blocking assignments, whose updates are queued into the NBA region.

A race exists when statements scheduled in the same time slot would give different results if executed in another order the standard allows. Take two clocked blocks: one says y1 = y2, the other y2 = y1, using blocking assignments, after a reset that set y1 to 0 and y2 to 1. If the first block runs first, both end up 1; if the second runs first, both end up 0. The hardware the designer meant, two flops swapping values, appears in neither. With assignments, every clocked block reads its inputs in the Active region before any register changes, and all the updates land together in the NBA region. Process order then cannot change the result, which is the same property real flip-flops get from sharing a clock edge. NBA updates can wake combinational logic, which runs in a fresh Active pass of the same slot. The Inactive region holds #0 assignments, which good RTL never needs.

PrepActiveInactNBAObsReactReInReNBAPost↑ the active region set, used by RTLalways @(posedge clk)// block 1y1 = y2;always @(posedge clk)// block 2y2 = y1;y10y21NBA queue (unused by =)pre-edge: y1 = 0, y2 = 1
Assignment
Simulator runs first
1 / 5

Before the edge: y1 = 0, y2 = 1 (from reset). Preponed samples them for assertions.

One time slot of the IEEE 1800 scheduler (regions abbreviated). Pick the assignment type and which block the simulator runs first, then step through the slot.Share freely with credit: ‘Figure from chipfieldguide.com’

Cycle-based simulation: how Verilator compiles RTL

Verilator is a compiler rather than a traditional simulator. It reads and lints the design, then turns it into a C++ or SystemC model, which is compiled together with a small user-written wrapper into an executable that runs the simulation. Inside, it builds a syntax tree, resolves parameters, inlines modules, folds constants and removes dead code. It then pseudo-flattens the hierarchy, so optimizations can cross module boundaries, orders the logic and writes C++. In the best case the result is a single eval function whose always blocks run top to bottom with no loops.

Scheduling is done once, at compile time. Verilator implements the Active and NBA regions of the IEEE 1800 model statically: it classifies logic as clocked or combinational and sorts it in data-flow order, so that each piece runs after everything that feeds it and one pass reaches a settled state. Combinational logic is evaluated only when one of the clocks or inputs driving it has triggered. The developers observed about a 100× slowdown on large designs with this activity gating turned off. Genuine combinational cycles are broken by marking the variables on the “back edges” of the dependency graph, then iterating the logic that reads them until the values settle.

That is the essence of : replace a run-time event queue with an order computed in advance. It is fast, but it changes some semantics:

  • Values are two-state, 0 or 1 only. Explicit X values and uninitialized variables get real values chosen by --x-assign and --x-initial; their unique modes vary those values from run to run, to expose reset bugs.
  • Delays and other timing controls are supported only with --timing, which needs a C++ compiler with coroutine support; without it they are ignored, as in earlier versions.

Race-free RTL gives the same cycle-by-cycle results in both kinds of simulator. RTL that depends on scheduling order, on X behavior or on delays may not, which is one more reason to follow the assignment rules.

Metastability and synchronizer MTBF

Ginosar derives the failure rate in two steps.

  1. How often a flop becomes metastable. A flip-flop goes metastable when its input changes inside a small window TWT_{\mathrm{W}} around the clock edge. With a clock at rate FCF_{\mathrm{C}} and data changing at rate FDF_{\mathrm{D}}, that happens TWFCFDT_{\mathrm{W}} F_{\mathrm{C}} F_{\mathrm{D}} times per second.
  2. How often it is still undecided when read. Inside, a metastable flop is a pair of cross-coupled inverters balanced near the middle; any imbalance grows exponentially, with time constant τ\tau. The chance it is still unresolved after a settling time SS is therefore e−S/τe^{-S/\tau}.

Multiplying the two gives the failure rate, and its inverse is the :

MTBF=eS/τTW FC FD\mathrm{MTBF} = \frac{e^{S/\tau}}{T_{\mathrm{W}}\, F_{\mathrm{C}}\, F_{\mathrm{D}}}

For a two-flop , SS is one receiving-clock period minus the first flop’s clock-to-output delay, the second flop’s setup time and any wire delay between them. Ginosar’s worked examples:

  • With τ=10 ps\tau = 10\,\mathrm{ps}, TW=20 psT_{\mathrm{W}} = 20\,\mathrm{ps}, a 1 GHz clock, data changing every tenth cycle and SS = one clock period, MTBF is about 4×10294 \times 10^{29} years. The universe is about 101010^{10} years old.
  • At low voltage and high temperature (τ=100 ps\tau = 100\,\mathrm{ps}, TW=200 psT_{\mathrm{W}} = 200\,\mathrm{ps}, FD=1 kHzF_{\mathrm{D}} = 1\,\mathrm{kHz}), the same circuit fails about once a minute: S/τS/\tau has dropped from 100 to 10, and that factor of e90e^{90} dwarfs the changes in TWT_{\mathrm{W}} and FDF_{\mathrm{D}}.
  • A third flop raises that to about a month, and a fourth to about 1,000 years. These are Ginosar’s rounded figures; worked exactly, as the figure below does, they come to about two minutes, 28 days and 1,700 years.

System MTBF falls roughly in proportion to the number of synchronizers, so a chip with 1,000 of them needs each to beat the system target by three orders of magnitude. The exponential dependence on SS is why stage count and keeping the flops adjacent matter far more than anything else in the circuit, and why τ\tau and TWT_{\mathrm{W}} must be taken at the slowest operating corner.

τ = 10 ps, T_W = 20 ps, F_D = 100 MHz, S = 1 ns4 × 1029 years1 s1 min1 day1 year1,000 yrage of universe1020 yr1030 yrmean time between failures (log scale) →
Corner
Flops in the synchronizer
Synchronizers on the chip

S/τ = 100, so a metastable flop is still unresolved with probability e⁻¹⁰⁰. Metastability events: 2 × 10⁶ per second. MTBF ≈ 4 × 10²⁹ years.

MTBF = e^(S/τ) / (T_W·F_C·F_D) on a log time axis, F_C = 1 GHz, S = one period per extra stage (flop delays ignored, as in Ginosar’s examples). Each stage multiplies MTBF by e^(T_C/τ).Share freely with credit: ‘Figure from chipfieldguide.com’

FSM encoding trade-offs

  • Binary. ⌈log⁡2N⌉\lceil \log_2 N \rceil flops for NN states (4 flops for 9 to 16 states), but every state test has to decode several bits, so that logic is denser and deeper.
  • . NN flops. Each state test is a single-bit check, so the next-state logic is shallow and the machine can usually run at a faster clock. That suits FPGAs, which have flip-flops to spare, and Yosys re-encodes the state machines it extracts as one-hot. The same logic-depth argument often holds for fast ASIC control. It also leaves 2N−N2^N - N illegal codes. Whether the hardware recovers from one, after a radiation upset for example, depends on what synthesis does with the default branch, so check it when it matters.
  • . Consecutive codes differ in one bit, so a machine that mostly steps through a fixed sequence toggles one state flop per transition. That reduces switching, and for a pure sequence it makes the state safe to sample from another clock domain, the same property that makes Gray counters safe to synchronize. It helps only when the transitions form something close to a chain.

Synthesis can extract the machine and re-encode it, as Yosys’s fsm command does, which is why defining the states in one place matters: the encoding becomes a tool setting rather than an RTL rewrite.

Novice · 0 of 5 correct
  1. Q1A designer wants a three-stage shift register, where a value moves one stage per clock tick. In a clocked block they write q1 = d; q2 = q1; q3 = q2; using the blocking = sign. What does synthesis build?

  2. Q2A block meant to have no memory says: always_comb if (en) q = d; and nothing else. What happens when en is 0?

  3. Q3Why can’t you pass an 8-bit counter value to a block on an unrelated clock through eight separate two-flop synchronizers, one per bit?

  4. Q4A design uses asynchronous reset. Why does it usually add a reset synchronizer?

  5. Q5Which line belongs in a testbench (code that tests the design) rather than in the design itself?

Sources

Show Hide 29 sources
  1. Register-transfer levelWikipediaRTL models a synchronous circuit as registers plus the combinational logic between them; HDL code at this level is translated to hardware by synthesis.
  2. Low Power Design, Verification, and Implementation with IEEE 1801 UPF (tutorial)Members of the IEEE P1801 working group (Erich Marschner, Jeffrey Lee, Qi Wang, John Biggs, Sushma Honnavara-Prasad) · Accellera Systems Initiative · 2013Power intent (domains, supply rails, switches, isolation, level shifters, retention) is captured in UPF, separate from the functional intent in RTL; RTL + UPF verification precedes implementation; successive refinement lets IP providers ship constraint UPF such as clamp values.
  3. Chisel: Constructing Hardware in a Scala Embedded LanguageJonathan Bachrach, Huy Vo, Brian Richards, Yunsup Lee, Andrew Waterman, Rimas Avižienis, John Wawrzynek, Krste Asanović · Design Automation Conference (DAC); author copy archived by the Wayback Machine (DOI 10.1145/2228360.2228584) · 2012Verilog and VHDL began as simulation languages and were later adopted for synthesis; Chisel builds parameterized generators in Scala and emits a C++ simulator or Verilog.
  4. Notes on Verilog support in YosysYosysHQ · Yosys documentationYosys does not support the non-synthesizable features defined in IEEE 1364.1; with -sv it accepts always_comb, always_ff, always_latch and logic; unique and priority are turned into full_case/parallel_case attributes.
  5. lowRISC Verilog Coding Style GuidelowRISC · GitHub (lowRISC/style-guides)Non-blocking for sequential logic, blocking for combinational; latches discouraged; default case items; unique case, and no full_case/parallel_case pragmas; _d/_q naming; active-low asynchronous resets; named generate blocks.
  6. L5: Simple Sequential Circuits and Verilog (6.111 lecture slides)MIT 6.111, Introductory Digital Systems Laboratory · MIT OpenCourseWare · 2006Blocking assignments take effect at once; non-blocking ones are deferred until all right-hand sides are evaluated; a three-flop delay line written with blocking assignments gives out = in (one flop); use non-blocking for sequential and blocking for combinational always blocks; always blocks run in parallel, so beware races.
  7. L6: Finite State Machines (6.111 lecture slides)MIT 6.111, Introductory Digital Systems Laboratory · MIT OpenCourseWare · 2006Moore outputs depend on state only, Mealy outputs on state and inputs; an FSM coded as a state register plus a combinational block; decoded outputs can glitch while state bits change; registered outputs are glitch-free.
  8. One-hotWikipedia contributors · WikipediaBinary-encoded state machines need a decoder to tell the state; one-hot needs one flip-flop per state, reads the state from one bit, suits FPGAs’ abundant flip-flops and usually allows a faster clock.
  9. FSM handlingYosysHQ · Yosys documentationThe fsm command identifies, extracts and re-encodes state machines (currently to one-hot) and can write a file of the changes for use in equivalence checking.
  10. Technology mapping (the clockgate command)YosysHQ · Yosys documentationclockgate turns each set of flip-flops sharing a clock and enable into an integrated clock-gating (ICG) cell plus enable-less flops, as an ASIC power saving; -tie_lo is intended for DFT scan-enable pins.
  11. Errors and WarningsWilson Snyder and contributors · Verilator documentationLATCH, BLKSEQ, CASEINCOMPLETE, WIDTH and UNUSEDSIGNAL warnings; lint_off metacomments; why global warning disables should be avoided.
  12. Asynchronous reset synchronization and distribution – challenges and solutionsRostislav (Reuven) Dobkin · Embedded.com · 2017Used as a fallback (no open primary source found). Synchronous reset needs an active clock; asynchronous reset works without one; release must meet recovery/removal or the flop can go metastable, and can reach flops on different cycles; trailing-edge synchronizer whose second flop sees no input change; release must reach all flops in one cycle; one synchronizer per clock domain; release order between domains; asynchronous resets must be directly accessible for DFT.
  13. Metastability and Synchronizers: A TutorialRan Ginosar · IEEE Design & Test of Computers (author copy, Technion) · 2011MTBF = e^(S/τ)/(T_W·F_C·F_D), two-flop synchronizers, req/ack handshakes, two-clock FIFOs with Gray pointers, and the bit-by-bit synchronization trap.
  14. A Specification-Driven Methodology for the Design and Verification of Reset Domain Crossing LogicPriya Viswanathan, Kurt Takara, Chris Kwok, Islam Ahmed · DVCon (Accellera DVCon proceedings)Defines RDC paths, where an asynchronous reset in one domain can make a register in another go metastable; mitigation by reset ordering, staged resets and isolation enables; RDC issue counts for two designs (Table I).
  15. Gray codeWikipediaGray codes change one bit per step; used for multi-bit counts crossing clock domains and for FIFO pointers; binary-to-Gray conversion must be reclocked.
  16. verilator ArgumentsWilson Snyder and contributors · Verilator documentation--lint-only, -Wall, -Wno-fatal, --timing (needs C++ coroutines) and the two-state --x-assign / --x-initial options (with +verilator+rand+reset+2) for finding reset bugs.
  17. VeribleCHIPS Alliance · Verible project siteA SystemVerilog parser with a style linter, formatter and syntax checker.
  18. Pragmatic Formal Verification Methodology for Clock Domain Crossing (CDC)Aman Kumar, Muhammad Ul Haque Khan, Bijitendra Mittra · DVCon Europe 2023 (arXiv:2406.06533) · 2024RTL simulation cannot show setup/hold violations and STA alone misses CDC bugs; structural CDC analysis looks for missing synchronizers, combinational logic on synchronizer paths and convergence; hand-written constraints can hide bugs; metastability injection.
  19. High-level synthesisWikipediaHLS turns untimed C/C++/SystemC into timed RTL through scheduling, allocation and binding, trading clock cycles against hardware resources.
  20. Clock crossing violationSpinalHDL project · SpinalHDL documentationSpinalHDL (Scala) checks at elaboration that every register depends only on registers in the same or a synchronous clock domain, and offers BufferCC for single-bit or Gray-coded crossings.
  21. IntroductionAmaranth project · Amaranth HDL documentationA Python library for RTL modeling of synchronous logic that can simulate, synthesize through Yosys, or emit Verilog; first-class clock domains and CDC primitives in the standard library.
  22. Getting started (EQY documentation)YosysHQ · YosysHQ EQY documentationEQY formally proves two designs equivalent, for example that a synthesis tool introduced no functional change or that a refactor preserved correctness.
  23. sv2v: SystemVerilog to VerilogZachary Snow · GitHubConverts SystemVerilog (IEEE 1800-2017) to Verilog-2005, focusing on synthesizable constructs; originally developed to target Yosys.
  24. Getting Started With Icarus VerilogIcarus Verilog project · Icarus Verilog documentationCompiling a Verilog file with iverilog and running the result with vvp.
  25. Simulator Timing Related Material (EE 4755 class notes)LSU EE 4755, Digital Design Using HDLs · Louisiana State University · 2019Time slots and regions of the stratified event scheduler (IEEE 1800-2012 clause 4.4), with nine simulation regions: Preponed, Active, Inactive, NBA, Observed, Reactive, Re-Inactive, Re-NBA, Postponed; #0 puts a process in the Inactive region; non-blocking updates come after inactive events; active events in any order.
  26. Proposal for Section 15: Scheduling Semantics (SystemVerilog LRM draft)Arturo Salz · Accellera SystemVerilog Enhancement Committee archive · 2005Draft LRM clause: the reference simulation algorithm compliant simulators implement; time slots divided into ordered regions; Active events in any order; #0 schedules into Inactive; a non-blocking assignment creates an NBA-region event; Preponed sampling for properties, Observed for property evaluation, Reactive for program blocks, Postponed for monitoring; the scheduler loops back to Active while regions are non-empty.
  27. Accellera Announces IEEE 1800™-2023 Standard Available Through IEEE GET ProgramAccellera Systems Initiative · Accellera · 2024IEEE 1800-2023 (SystemVerilog) is downloadable without charge through the IEEE GET Program; it covers behavioral, RTL and gate-level modeling plus verification.
  28. OverviewWilson Snyder and contributors · Verilator documentationVerilator is a compiler, not a traditional simulator: it lints SystemVerilog and converts it into a C++ or SystemC model that is compiled and run.
  29. Verilator InternalsWilson Snyder and contributors · GitHub (verilator/verilator, docs/internals.rst)Verilator implements the Active and NBA regions statically, orders logic in data-flow order, gates evaluation by activity, and iterates only to settle combinational loops.