An is a chip you can rewire after it is made. It is covered in thousands of tiny and switches, all left unconnected. The fabric chapter shows what is inside.
To use one, you write code that describes a circuit. Then a chain of programs turns that code into a file of settings. The file says what every tiny block does and which switches are on. Loading it into the chip takes a few seconds.
The steps look a lot like the ones for designing a chip. But there is one big difference. The FPGA already exists. The tools can’t add anything to it. They can only choose which parts to use and how to join them.
An is a grid of small programmable (LUTs), flip-flops, memory blocks and multiplier blocks, joined by programmable wiring. A configuration file called the sets every table and switch. What that fabric looks like is the subject of Lookup tables and the FPGA fabric; this chapter is about the software that fills it in.
The FPGA flow follows the same outline as the chip () flow in the Design Flow guide:
- Synthesis reads the hardware code and turns it into logic.
- Technology mapping covers that logic with LUTs, plus the chip’s fixed adders, memories and multipliers.
- Packing groups LUTs and flip-flops into the chip’s .
- Placement picks a site on the chip for every block.
- Routing chooses wires and switches to connect them.
- Timing analysis checks that every signal arrives before the next clock tick.
- Bitstream generation writes the configuration file.
What makes it different is that the silicon is already finished:
- Mapping targets one kind of logic part, the LUT, instead of a library of hundreds of gate types.
- Placement chooses among sites that already exist, each of a fixed type.
- Routing searches a fixed graph of wires and switches instead of drawing new wires.
- Timing problems are fixed by changing the design or the tool settings, never the transistors.
- A full build takes minutes to hours, and loading the result takes seconds.
Vendors ship their own complete tool suites. There is also a fully open-source flow built on and .12
You know what synthesis, placement, routing and static timing do for an ASIC; the Synthesis, Placement, Routing and Signoff chapters cover them. The FPGA versions solve the same problems with one constraint added: every resource, delay and wire already exists. That turns each step into an assignment problem onto a fixed resource graph, and it changes the algorithms in specific ways:
- Mapping is covering, not matching. A K-input LUT implements any of the functions of K inputs, so the mapper never asks which function a piece of logic computes, only whether it has at most K inputs.6 Depth-optimal mapping is then solvable in polynomial time; area-optimal mapping is not.
- Placement is discrete and typed. Logic, RAM and sit in columns of sites; the placer assigns instances to legal sites of the right type and cannot resize anything.10
- Routing is path search with fixed capacity. Every track and pin is a node of a routing-resource graph; the router picks paths so no node is used twice.7 Negotiated congestion was invented for exactly this case.9
- Timing closure has no gate sizing or buffering. The delay model is fixed per device and speed grade; the levers are logic structure, placement and constraints.20
- Bitstream generation needs a device database. For open tools that database came from reverse-engineering projects such as IceStorm, Trellis and X-Ray.11
Design in: Both flows start from the same register-transfer-level HDL. Much of it can target either one.
FPGA designers write their circuit in a , the same kind used for custom chips. It doesn’t list steps to run one after another, like a normal program. It describes parts that all work at the same time.
A program called a synthesis tool reads the code and looks for patterns. It spots adders, memories and multiplications. The FPGA has ready-made parts for those, so the tool uses them. Everything else is built from : tiny memories that can act as any small gate.
Some people skip the hardware language. They write normal code in a language like C++, and a tool turns it into hardware code for them. That is called high-level synthesis.27
The input is a , usually Verilog, SystemVerilog or VHDL, written at the : registers, plus the logic that computes their next values. Writing RTL is covered in the RTL design chapter; most of it applies to FPGAs unchanged.
for an FPGA then does three jobs:
- Elaborate and optimize. Build the design hierarchy, flatten it, and simplify the logic.
- Infer hard blocks. Recognize code patterns that match the chip’s dedicated resources: arrays become , multiply-accumulates become , adders use the , and registers become flip-flops.
- Map the rest to LUTs. Whatever is left is random logic, which covers with lookup tables (next section).
The open-source synthesis tool, which began as a student thesis project, shows the steps plainly.1 Its script for Lattice iCE40 FPGAs maps memories to the chip’s block RAM, optionally maps multipliers to DSP blocks, protects carry chains, hands the remaining logic to the ABC program for mapping into 4-input LUTs, legalizes the flip-flops, and writes a netlist for the place-and-route tool.2 Yosys has similar scripts for families from AMD, Altera, Lattice, Microchip, Gowin and others.3 Each vendor’s own suite contains its own synthesis step.
Starting from C or C++
(HLS) writes the RTL for you from C, C++ or SystemC. FPGAs suit it well, because a generated circuit can be tried on the chip, changed and reloaded with no manufacturing cost. HLS tools come from FPGA vendors, EDA companies and universities; one survey compared dozens of them, including the academic tools LegUp and Bambu.27
The key HLS idea is loop pipelining: starting a new loop iteration before the last one finishes. The number of clock cycles between iterations is the initiation interval, ideally 1. Memory ports and dependences between iterations set how low it can go. A loop with three loads and one store, from a memory with two ports, cannot start iterations faster than every two cycles.27
FPGA synthesis differs from ASIC synthesis mainly at the back end. Coarse-grained passes infer primitives before any gate-level mapping: memory inference matches an array’s ports, widths and read timing to the RAM primitive (or falls back to LUT RAM or flip-flops), DSP inference folds multiply, add and accumulator register into one hard block, and are wrapped so the LUT mapper leaves them alone.2 Only the residue becomes an AIG for ABC, whose FPGA mappers handle fixed and variable LUT sizes.4 A generic wide LUT can then be decomposed into the device’s physical LUTs plus dedicated multiplexers, as Yosys does for 7- and 8-input functions on Xilinx 7-series (two or four LUT6s plus MUXF7 and MUXF8).12
Two consequences for RTL style. First, inference is pattern matching, so a memory with an asynchronous read or an unusual reset can silently become thousands of flip-flops instead of one RAM block; check the utilization report. Second, vendor primitives instantiated directly in RTL make the code device-specific, which matters when an FPGA design is a prototype of an ASIC (see Verification, where is covered).
HLS adds scheduling, allocation and binding on top. Its central metric is the initiation interval , bounded below by resource limits (memory ports) and by loop-carried dependences. For a loop body with memory accesses on a memory with ports, . Partitioning arrays across banks raises .27
DSP block: A multiply-accumulate is inferred as a DSP block: a hard multiplier and adder built into the fabric.
Almost all of an FPGA’s logic is made of lookup tables. A lookup table is a tiny memory with a few inputs, often 4 or 6. For every pattern of inputs, it stores the answer.
Four inputs can make 16 patterns, so a 4-input table stores 16 answers. Fill them in one way and it acts like an AND gate. Fill them in another way and it is a totally different gate. One part can be anything!6
So the tool’s job is a puzzle. It chops the circuit into pieces that each need at most 4 inputs. Each piece becomes one table. It wants as few pieces as it can. It also wants short chains, because a signal slows down each time it passes through another table.
Sometimes the tool copies a small bit of logic into two tables. That uses more tables, but it can make the longest chain shorter, so the circuit runs faster.
For a custom chip, rebuilds the logic out of cells from a library, checking which cell computes each small piece (see Synthesis). An FPGA’s logic part is different. A K-input lookup table (LUT) is a memory of bits, addressed by its inputs. It can therefore compute any function of up to K inputs.6 So the mapper doesn’t care what a piece of logic does, only how many inputs it has. The iCE40 family, for example, uses 4-input LUTs, and Xilinx 7-series uses 6-input ones.212
The pieces are called cuts. A of a gate is a set of at most K signals that every path into the gate must pass through. Everything between those signals and the gate depends only on them, so one LUT can replace it. The mapper:
- lists the K-feasible cuts of every gate, building each gate’s list from its inputs’ lists;
- picks one cut per gate, depending on the goal;
- starts at the outputs, makes a LUT from each chosen cut, and repeats for every signal that cut needs.
There are two goals, and they pull against each other:
- Depth, the number of LUTs on the longest path (), sets speed: each level adds a LUT delay and a trip through the routing. Finding the minimum depth is a solved problem; the FlowMap algorithm (1994) does it in polynomial time.6
- Area, the number of LUTs, sets how big an FPGA you need. Minimizing it exactly is NP-complete once gates feed more than one place, so tools use heuristics.6
Minimum depth often comes from duplicating a shared gate into several LUTs. Real mappers such as ABC, which Yosys calls, therefore find a depth-optimal mapping first and then recover area without making the depth worse.5 Mapping is fast: in 2007, a million-gate network was mapped to 6-input LUTs in about a minute on a laptop.5
The subject graph is usually an . Cut enumeration runs in topological order: a primary input has only its trivial cut, and for an AND node with fanins ,
with dominated cuts (supersets of another cut) removed.5 Depth-oriented mapping then labels each node with the best level any of its cuts achieves, where the level of a cut is , and keeps that cut as the node’s representative.5 Because a K-LUT accepts any function, every enumerated cut is usable; a cell mapper would discard cuts whose function no library cell matches, which is why ASIC mapping needs Boolean matching and FPGA mapping does not.
Area recovery runs in passes under the required times of that depth. Area flow gives a global estimate that divides shared logic among its fanouts:
and a second pass with exact local area (the LUTs that selecting a cut would add, counted by dereferencing the cut’s fanout-free cone) refines it.5 Exhaustive enumeration blows up for large K; keeping only a few priority cuts per node (5 to 10) gives depth-optimal results on 95% of benchmarks even with a single cut, and with 8 cuts reduces memory 10× and runtime 5× for 6-LUTs at comparable quality.5 The FlowMap labeling itself, which computes the same depth labels with network flow instead of enumeration, is in Under the hood.
Library: 3 cells (AND2_X1, XOR2_X1, OR2_X1), found by matching patterns. LUT: one 4-LUT; only its 16 bits change.
Here is a small circuit drawn as gates, with signals flowing left to right. The colored blobs are lookup tables. Each blob covers the gates it can replace.
Slide the number of inputs per block from 2 to 6 and watch the blocks grow. Then switch between “Fastest” and “Fewest blocks.” Try the 4-bit adder with 4 inputs: one goal saves a step, the other saves blocks.
The network is made of two-input gates. Each colored region is one K-input LUT chosen by a cut-based mapper; solid wires are LUT inputs the router will have to connect, dashed wires disappear inside a LUT, and ×2 marks a gate copied into two LUTs. Things to try:
- With K = 2, every gate is its own LUT. How many levels does the 4-bit adder have?
- On the 4-bit adder at K = 4, compare depth first with area first. What does each give up?
- On the 8-to-1 mux, step K from 4 to 6. Why does the LUT count drop so sharply at 6?
- Use “Cover one LUT at a time” to watch the mapper work back from the outputs.
The mapper enumerates all non-dominated K-feasible cuts of a two-input-gate network. “Depth first” takes the depth-optimal labels and then recovers area with area flow under those required times; “Area first” minimizes area flow alone; “Depth only” keeps FlowMap’s largest-volume cut with no recovery.56 Try:
- The adder at K = 6 in “Depth only” versus “Depth first”: same depth, how many LUTs does recovery save, and where did the duplicated gates go?
- Turn on depth labels. Check that each LUT root’s label is one more than its deepest leaf.
- Compare “LUT bits used” across K for the comparator. Bigger LUTs cut depth but cost bits each, the architecture trade-off behind the choice of K.
The model counts LUTs and levels only. Real mappers also see carry chains, wide-function multiplexers and per-input delays, and work on AIGs that have been restructured first.
An FPGA is like a parking lot with marked spaces. Some spaces are for logic, some for memory and some for math blocks. Each part must park in a space of its own kind. The tool can’t paint new spaces.
First, the tool groups lookup tables that talk to each other into one logic block. That is called . Wires inside a block are short and fast.7
Next, it picks a space for each block. It wants blocks that share wires to sit close together. A common method starts with a messy layout, then tries swapping blocks again and again. Swaps that shorten the wires are kept. Early on, some bad swaps are kept too, so the layout doesn’t get stuck. This is called , after the way metal is cooled slowly.7
groups LUTs and flip-flops into the chip’s logic blocks, often several LUT-and-flip-flop pairs per block. A block has a fixed number of inputs, clock and control signals, so not every grouping is legal. Connections inside a block use fast local wiring, which takes pressure off the general routing, where most of an FPGA’s area and delay goes.8 The classic packers are greedy: start a block with one element, then keep adding the element that shares the most signals with it (T-VPack also favors timing-critical connections) until it is full. Research with these tools found that blocks of 7 to 10 LUTs gave the best trade-off, with about 30% less delay than one-LUT blocks.8 In nextpnr’s iCE40 flow, packing merges each LUT with the flip-flop it drives.12
Placement then assigns each block to a site. A modern FPGA is a grid of columns, each of one type: logic blocks, block RAM, DSP blocks, with I/O around the edge.10 The I/O pins are usually fixed by the circuit board. The goal is short, uncongested wiring and good timing; wirelength is estimated with each net’s bounding box, the familiar from ASIC placement.7
Two families of placer are used:
- Simulated annealing, as in the academic VPR tool: random swaps of blocks between sites, accepting worse moves with a probability that falls as a “temperature” is lowered.7
- Analytic placement, which solves a smooth wirelength problem for all blocks at once, spreads them out, then snaps them to legal sites. It scales better: one analytic placer ran 7.4× faster than VPR’s annealer with 6% better quality.10
nextpnr includes both.12 The ideas are the same ones taught in Placement for ASICs; what changes is that sites are discrete and typed, and nothing can be resized or buffered afterward.
Typical FPGA placement runs four steps: an initial placement or floorplan, global placement under resource constraints, packing and legalization to exact sites, then detailed placement to fix the worst cases. Macros complicate it: carry chains, cascaded DSPs and LUT-plus-mux groups must occupy adjacent sites in one column.10
VPR’s annealer is the reference design. Its cost is a bounding-box sum with a fanout correction (1 for nets of up to three terminals, rising to 2.79 at 50) and, for non-uniform channels, division by the average channel capacity :
Its adaptive schedule sets the starting temperature to 20 times the standard deviation of cost over random swaps, tries moves per temperature, and cools fastest when almost every move or almost no move is accepted. A range limiter shrinks the swap distance to keep the acceptance rate near 0.44. Cutting moves per temperature by 10× made placement 10× faster and only about 10% worse.7 The full schedule is in Under the hood.
Analytic placers (HeAP and its successors) minimize a quadratic or smoothed wirelength, then spread cells by region against each resource type’s supply, which annealing cannot match for runtime on large netlists. A 2023 open-source analytic placer reports critical-path delay within 2.2% and 0.59% of two Vivado releases at 14% and 8.5% more runtime.10 Packing order matters too: the Titan study traced much of VPR’s wirelength gap to Quartus II to VPR’s focus on dense packing.19
1. Mapped netlist: After technology mapping: 12 LUT + flip-flop pairs, one block-RAM instance and two I/O pins, joined by 20 connections.
A custom chip gets new wires drawn just for it. An FPGA doesn’t. Its wires are already built, cut into short pieces, with switches where the pieces meet.
So routing an FPGA is like planning trips on a subway map. Every station and track already exists. The tool picks a path for each signal and turns on the switches along it. Each switch it turns on becomes one bit in the settings file.7
Each wire piece can carry only one signal. When two signals want the same piece, one has to take a longer way. The tool lets them argue it out over many rounds, raising the price of busy pieces until everyone fits.9
Even a small FPGA has over a hundred thousand wire pieces and more than a million switches.12
An FPGA’s routing is a fixed network of wire segments in channels between the logic blocks, joined by programmable switches where channels cross and where block pins meet the channels. The router never sees the picture; it sees a graph. In VPR’s formulation, every routing track and every pin is a node of a routing-resource graph, and every allowed connection is an edge.7 Routing a net means finding a path of nodes from its source pin to each sink pin, then turning on the switches (edges) along it.
Every node can carry only one net. Routing nets one at a time makes the result depend on the order, so FPGA routers use , an algorithm called PathFinder that was invented for an FPGA in 1995. Every net is ripped up and rerouted each round; nodes that several nets want get more expensive, both now and in a history that never decreases, until each node has one net.97 How it works step by step is in Routing, which also shows the same idea in ASIC global routers.
The fixed graph changes what can go wrong:
- Capacity is fixed. If the router cannot resolve the overuse, the design is unroutable at that placement. The fix is a different placement or a less crowded design, not more tracks.
- Detours are slow. Every switch a signal passes adds delay, and on an FPGA routing already accounts for most of the delay.8
- The graph is huge. A small Lattice iCE40 UP5K has about 125,000 wires and 1.3 million switches; an 85K-element ECP5 has about 4 million wires and 28 million switches.12
PathFinder’s node cost combines a base cost (the node’s delay), a present-sharing term that grows with the number of other nets on the node, and a history term that grows every iteration the node is overused; every net is rerouted every iteration, and timing-critical connections weight delay over congestion.9 The Routing chapter’s Under the hood works through the cost function. FPGA-specific details:
- Graph search per net. Each net is a Dijkstra (maze) expansion over the graph, run times for a -terminal net. VPR keeps the wavefront when it reaches a sink and re-expands from the newly added wire at zero cost instead of restarting, which matters for high-fanout nets. It also limits each net to 3 channels outside its terminals’ bounding box and gives up after 45 iterations.7
- Lookahead. nextpnr’s router is timing-driven rip-up and reroute with A* search, which steers each expansion toward the sink with an estimate of the remaining cost.12
- Pin equivalence. All inputs of a LUT are functionally equivalent, so the architecture description can mark them interchangeable and the router may use whichever one is free.7
- Database size. The ECP5-85K graph (4M wires, 28M switches) takes 1 GB uncompressed; nextpnr deduplicates repeated tiles to 38 MB.12
Net 1 uses 2 wire segments and turns on 3 switches (pin → wire, wire → wire, wire → pin). Each switch is one or more bits in the bitstream.
Every circuit runs to the beat of a clock. Each signal must reach its next storage cell before the next tick. The tools add up the time a signal spends in every block and on every wire, along every path.
On an FPGA, those times are already known. The chip maker measured how slow each kind of block, wire and switch is. So the tools just look them up and add.
When a path is too slow, a custom-chip team can make parts bigger or add new wires. An FPGA user can’t. So they change the design instead. They split a long path with an extra storage step, help the tools place parts closer together, or run the clock a bit slower.
The FPGA tools use , the same method as for ASICs: add the delays along every path between flip-flops and compare the total with the clock period. The spare time is the ; the path with the least is the . The difference is where the delays come from. Every LUT, wire segment and switch is part of a finished chip, so its delay is a fixed number from the vendor’s device model, for each speed grade of the part.20 Open tools get theirs from the documentation projects, which also supply timing data.1213
is the loop of analyzing, changing and rebuilding until every path has non-negative slack. On an FPGA the levers are:
- Fewer LUT levels. The number of LUTs on a path has to suit the target clock and the speed grade.20
- and . Add registers to split long paths, or let synthesis move registers from short paths into long ones.20
- Hard blocks. Put arithmetic and wide multiplexers on the carry chains, DSP blocks and dedicated multiplexers instead of in LUTs.20
- Placement and tool settings. Floorplanning critical logic into a region, trying different implementation strategies in parallel runs, and physical optimization after placement.21
- The clock target itself, if the system can live with a slower clock.
What is missing is everything an ASIC team does in the Signoff loop: sizing gates up, inserting buffers anywhere, or widening wires. Because a rebuild takes minutes to hours, closure is iterative. Vendors speed up each turn: incremental implementation reuses an earlier result, and one vendor lets designers re-run timing analysis for “what-if” changes without re-running place and route.2125
A path’s delay is clock-to-Q, then alternating routing and LUT delays, then setup. With LUT levels the data arrival is roughly
and on an FPGA the net terms are usually larger than the LUT terms, because routing dominates delay.8 That makes the logic-level count a good first predictor, which is why vendor methodology asks designers to check the logic-level distribution against the target frequency and speed grade early.20
Closure practice that follows from a fixed fabric:
- Restructure, don’t resize. Retiming in synthesis can move registers from low-level paths into high-level ones, applied globally or only to the blocks that fail.20
- Separate logic delay from net delay. Too many levels is a design problem; long nets come from placement constraints or congestion and are fixed in placement or floorplanning.21
- Explore the tool’s search space. Placement and routing are heuristic, so different strategies and directives give different results; parallel runs and incremental flows from a good reference checkpoint trade compute for turnaround.21
Timing also feeds back into every earlier step: mapping can trade area for depth, T-VPack packs critical connections together, and placers and routers weight critical nets.89
Worst path 7.70 ns against a 5 ns clock: slack -2.70 ns. Routing is 71% of the logic-and-wire delay.
The last step writes the settings file, called the . It holds the 16 or 64 answers in every lookup table, and the setting of every switch. A big FPGA has millions of switches to set.12
For a long time, chip makers kept the meaning of those bits secret. Only their own tools could write the file. Then volunteers worked it out for some chips. They made many small test designs with the official tools and studied which bits changed.1514
Because of them, free tools can now take some FPGAs from code all the way to a working file. Students and hobbyists can build real hardware with no paid software at all.12
turns the placed and routed design into the configuration file. A device database maps every resource to the bits that control it: the LUT contents, flip-flop modes, routing switches, I/O settings and memory contents. Bits are loaded in fixed-size chunks called frames; on Xilinx 7-series a frame is 101 32-bit words.16
For years the vendors’ own closed tools were the only way to make a bitstream.12 Open documentation projects worked out what the bits mean by generating many small designs with the vendor tools and comparing the outputs, a method called fuzzing.1514 They are the foundation of the : Yosys for synthesis, for place and route and timing, and a bit packer from the project.11 The iCE40 flow is three commands:
yosys -p 'synth_ice40 -top blinky -json blinky.json' blinky.v
nextpnr-ice40 --hx1k --json blinky.json --pcf blinky.pcf --asc blinky.asc
icepack blinky.asc blinky.bin- 1L1Synthesis and LUT mapping; writes a JSON netlist of iCE40 primitives.
- 2L2Pack, place, route and check timing on an HX1K part; the .pcf file pins signals to package pins.
- 3L3IceStorm’s packer turns the text configuration into the binary bitstream.
There are two other families of tools. , inside the Verilog-to-Routing (VTR) project, is the academic flow: it targets FPGAs described in an architecture file, so researchers can test chips that don’t exist yet.18 The F4PGA project uses VPR or nextpnr for real devices.17 And every vendor ships a complete suite of its own, listed in the box below; these are what most production designs use.12
Bitstream generation is a serialization of the routing-resource graph’s chosen edges and every primitive’s configuration through the device database. On 7-series, configuration memory is addressed by frame: a 32-bit frame address, 101 words per frame (100 for the tiles in the column, one for the horizontal clock row), frames striped two words per tile, written in multiples of the frame size with auto-incrementing addresses and CRC checks on the register writes.16 F4PGA uses a textual intermediate, FASM, that names each enabled feature before it is packed into frames.17
The open databases come from fuzzers that sweep one feature at a time (logic, RAM, I/O, clocking, interconnect) through the vendor tools and correlate the changed bits.15 The same databases supply what nextpnr needs to place, route and time: locations, connectivity, and cell and interconnect timing.12 nextpnr models each architecture as an API implementation rather than a flat description file, which lets it handle the irregularities of real devices; VPR instead generates architectures from parameters, which is what architecture research needs (for example, finding the smallest FPGA or narrowest channel a benchmark fits).12
Lattice iCE40: Yosys (synth_ice40) → nextpnr → bitstream written with the database from Project IceStorm. Status: supported.
Testing a design in a computer program is slow. On the FPGA, the design runs at full speed. But now you can’t see inside it. The signals are buried in the chip.
So the tools can add a tiny recorder to the design. You pick which signals to watch and what event to wait for, like an error signal turning on. The recorder keeps saving the latest moments in a loop of memory. When the event happens, it saves a little more, then stops. Then it sends everything to the computer through the programming cable.22
The catch: adding or changing the recorder means building the whole design again, which can take a long time.
An is debug logic compiled into the design. Described generically, it has probe inputs on the chosen signals, a trigger condition, and a sample buffer in on-chip block RAM. Once armed, it writes one sample per clock into a circular buffer at the design’s full speed. When the trigger fires, it records the rest of the window and stops, and the samples are read out over the JTAG debug port and shown as waveforms.2223
Each vendor ships on-chip debug tools:
- AMD’s Integrated Logic Analyzer in Vivado, added in the RTL or inserted after synthesis.22
- Altera’s Signal Tap logic analyzer in Quartus Prime, with settings for sample depth and RAM type.23
- Lattice’s Reveal, with an inserter to choose signals and triggers and an analyzer to view the capture.24
- Microchip’s SmartDebug, its hardware debug tool in Libero SoC.26
It is not free. The buffer uses block RAM, the probes use routing, and changing what you watch means recompiling. So simulation still comes first (see Verification), and on-board debug catches what only shows up at speed or with real inputs. Custom chips use the same idea in silicon, a . Chip teams also run their RTL on FPGAs before tapeout, as or in larger systems.
Visibility is bounded by depth × width: each probed bit per sample costs block-RAM capacity, so a deep capture of a wide bus can exhaust the device’s RAM. Probes are attached either in RTL or on the synthesized netlist; the latter preserves the RTL but still needs place and route, and the added logic and routing can perturb timing near the probed paths.22 Trigger logic ranges from simple comparators to state-based sequences.23
The practical cost is compile time. Each probe change is a new implementation run, which is why vendors emphasize incremental implementation from a reference checkpoint.21 The capture also happens in a clock domain of your choosing; signals from other domains are sampled asynchronously, with the same caveats as any .
2. Recording: Armed, the core writes one sample per clock into a circular buffer, overwriting the oldest, at the design’s full speed.
- Bits in one 4-input / 6-input LUT
- 16 / 64
- iCE40 UP5K routing graph: wires / switches
- 125K / 1.3M
- Mapping a 1M-node network to 6-LUTs (2007 laptop)
- ≈ 1 min
- Longest compile in the Titan study (Quartus II)
- 36.5 h
What these numbers mean:
- 16 / 64: a lookup table with 4 inputs stores 16 answers, and one with 6 inputs stores 64. Two more inputs make it four times bigger.6
- 125,000 / 1.3 million: even a small, low-power FPGA has that many wire pieces and switches for the routing tool to choose from.12
- About 1 minute: turning a circuit of a million gates into lookup tables is quick. Placing and connecting them takes much longer.5
- 36.5 hours: in one study, the biggest test design took a day and a half to go from code to a finished layout. Small ones took about a minute.19
Three numbers frame the flow. Mapping is cheap: a million-node network maps to 6-input LUTs in about a minute.5 Routing is big: a small iCE40 UP5K’s graph has about 125,000 wire nodes and 1.3 million switch edges, and an 85K-element ECP5 has about 4 million and 28 million.12 Compile time spans orders of magnitude: in the 2013 Titan study, the largest circuit of an older benchmark set took 61 seconds in VPR and 96 seconds in Quartus II, while benchmarks of 90,000 to 1.8 million blocks took up to 36.5 hours in Quartus II, and some ran past a 48-hour limit in VPR.19
| Finding | Number | Source |
|---|---|---|
| Delay of 7- and 10-LUT logic blocks versus 1-LUT blocks | 30% and 34% less | T-VPack, 19998 |
| VPR placement made 10× faster by cutting moves per temperature 10× | about 10% worse | VPR, 19977 |
| Analytic placer (HeAP) versus VPR 5.0’s annealer | 7.4× faster, 6% better | as reported in AMF-Placer 2.010 |
| VPR versus Quartus II on large designs: runtime, memory, wire | 2.7×, 5.1×, 2.6× | Titan, 201319 |
| Xilinx 7-series configuration frame | 101 × 32-bit words | Project X-Ray16 |
Read the Titan numbers with their context: Stratix IV-class architecture models, 2013 tool versions, and VPR’s runtime dominated by packing (about 78%) while Quartus II spent the largest share, 49%, in placement. VPR’s placer was in fact faster than Quartus II’s, partly because it packed more densely into fewer blocks; its router was 3.4× slower and routed high-fanout clock nets that Quartus II put on dedicated networks.19 The gap between academic and commercial tools has narrowed since: a 2023 open analytic placer reports critical-path delay within 0.59% of Vivado 2021.2 at 8.5% more runtime.10
The routing-graph sizes explain why databases matter: the ECP5-85K graph needs 1 GB uncompressed and 38 MB after nextpnr deduplicates repeated tiles, and the deduplicated 85K database is only 13% larger than the 25K one.12
The best thing about the FPGA flow is speed of change. Fix a mistake, rebuild, and the new design is running the same day. A custom chip would need months and a new trip through the factory.
The price is paid in other ways. A big design can take hours to build, so each try costs real time.19 The chip’s wires are fixed, so a very full design may not fit at all. And signals spend most of their time on the switches and wires, so long paths are slow.8
Free tools can now do the whole job for some FPGAs. For most products, people still use the chip makers’ own tools.12
| You get | You give up |
|---|---|
| Rebuild and reload in minutes to hours | Compiles long enough to slow every debug turn on big designs |
| Simple mapping: any K-input function fits one LUT | Logic you could build from a few small gates still costs a whole LUT |
| Routing that already exists, with known delays | Fixed capacity: congestion can make a design unroutable |
| Hard blocks for memory, arithmetic and DSP | Code that doesn’t match their patterns falls back to slow, large LUT logic |
| An open flow for several families | Experimental or missing support for many large devices |
Common ways the flow goes wrong
- Too many LUT levels on a path for the target clock.20
- Congestion in a crowded region, giving long detours or a failed route.21
- Missed inference, where a memory or multiplier is built from LUTs and flip-flops instead of a hard block, wasting area and speed.2
- Slow iteration, when each change to a debug probe or a constraint means another long compile.19
- Depth versus area in mapping. Depth-optimal mapping duplicates logic; area recovery wins much of it back but not all, and exact area minimization is NP-complete.65
- Packing density versus wirelength. Dense packing minimizes block count but can lengthen wiring; Titan attributed much of VPR’s 2.6× wire gap to it.19
- Annealing versus analytic placement. Annealing handles irregular legality naturally and trades quality for runtime smoothly; analytic placement scales to millions of cells but needs careful spreading and legalization for typed sites and macros.710
- Closure by compute. Placement and routing are heuristics whose results depend on their settings, so teams buy closure with parallel strategy runs and incremental flows rather than design changes, which costs machines and time.21
- Open versus vendor flows. Open tools give full control and run anywhere, and their authors position them for research and experimentation rather than production, where certified vendor flows remain the norm.12
Whether any of this beats building an ASIC is a question of volume, speed and power, the subject of Where FPGAs win.
Huge build: Benchmarks of 90K to 1.8M blocks took up to 36.5 hours in Quartus II, and over 48 hours in VPR for some.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
1. Cut-based LUT mapping
The mapper in the simulation, and ABC’s classic flow, are four passes over a topologically ordered subject graph:5
for n in topological order: # 1. enumerate cuts
if n is a primary input: cuts[n] = {{n}}
else: cuts[n] = {{n}} ∪ merge(cuts[a], cuts[b], K) # drop dominated
for n in topological order: # 2. depth labels
best[n] = argmin over nontrivial c in cuts[n] of 1 + max(label[m] for m in c)
label[n] = that minimum
required = propagate(max output label) # 3. area recovery
for n in topological order:
best[n] = argmin area_flow(c) over cuts with depth(c) <= required[n]
cover = {}; frontier = outputs # 4. cover from outputs
while frontier: n = pop(); cover += LUT(n, best[n]); frontier += gates in best[n]- 1L3merge = all unions u ∪ v of fanin cuts with at most K members.
- 2L5Any cut is legal: a K-LUT implements any function of its K leaves.
- 3L7Required times keep the optimal depth while area recovery picks cheaper cuts.
- 4L11Gates absorbed into several cones are duplicated implicitly.
The number of cuts grows quickly with K, which is why practical mappers keep a handful of priority cuts per node instead of all of them.5
2. FlowMap: depth labels by network flow
FlowMap computes the same optimal labels without enumerating cuts. Let be the largest label among node ’s predecessors. The label of is either or . To test for , collapse together with every predecessor labeled into a single sink, give every node capacity 1, connect a source to the primary inputs, and run max-flow. If the flow is at most K, the min cut is a K-feasible cut of height , so gets label ; otherwise it gets . Augmenting paths stop as soon as the flow exceeds K, so each test costs and the whole labeling .6
Among minimum-height cuts FlowMap takes the one with the largest volume (found by searching from the source in the residual graph and taking the complement), which packs the most logic into each LUT. Covering then works back from the outputs and replicates shared logic implicitly, so a post-pass is needed to cut the LUT count.6 The simulation’s “Depth only” mode shows that raw result.
3. Area flow and exact area
Area recovery combines two views. Area flow, computed in one pass from inputs to outputs, spreads the cost of shared logic across its fanouts:
Exact local area then counts the LUTs a cut would really add: dereference the cut’s leaves, recurse into leaves whose reference count drops to zero, sum, and re-reference. The recursion only visits the cut’s maximum fanout-free cone, which is usually small.5 Duplication-free mapping, which restricts covering to fanout-free cones, can be solved optimally per cone, but the global area optimum may need duplication.6
4. Greedy packing
VPack and T-VPack build one cluster at a time: choose a seed (the unclustered element with the most used inputs, or in T-VPack the most critical), then repeatedly absorb the element with the highest attraction until the cluster is full or no legal element remains.8 In VPack, attraction is the number of nets shared with the cluster:
T-VPack ranks first by the highest criticality of any connection joining the element to the cluster, then uses shared nets to break ties, so critical connections end up inside clusters on the fast local wiring.8
5. VPR’s annealing schedule
T0 = 20 × stddev(cost over N_blocks random swaps)
moves(T) = 10 × N_blocks^1.33
T_next = α × T, where α = 0.5 if R_accept > 0.96
0.9 if 0.8 < R_accept ≤ 0.96
0.95 if 0.15 < R_accept ≤ 0.8
0.8 if R_accept ≤ 0.15
D_limit = D_limit × (1 − 0.44 + R_accept), clamped to [1, chip size]
stop when T < 0.005 × cost / N_nets- 1L1Hot enough that almost any move is accepted at first.
- 2L5Most time is spent where some but not all moves are accepted.
- 3L7The range limiter keeps the acceptance rate near 0.44 by shrinking swap distance.
- 4L8Below this, almost no cost-increasing move would be accepted.
Each move swaps two blocks (or a block and an empty site of the same type) within the range limit; the move is accepted if the cost falls, or with probability if it rises.7 Because a move’s cost change only involves the nets of the moved blocks, each evaluation is incremental, which is what makes millions of moves affordable.
6. Routing and analytic placement
Negotiated-congestion routing on the routing-resource graph is worked through in Routing; the FPGA-specific additions (incremental wavefronts, bounding-box limits, A* lookahead) are in the routing section above.712 Analytic placement for FPGAs follows the ASIC recipe in Placement (quadratic or smoothed wirelength, then spreading) with per-resource-type density, carry-chain and DSP-cascade macros, and a legalization step that maps cells to typed sites.10
You have seen how code becomes the settings file that wires up an FPGA. The next chapter asks when that is worth it. An FPGA is slower and hungrier for power than a chip built for one job, but it can change. Find out where that wins in Where FPGAs win.
This chapter compared each FPGA step with the chip flow. Next, Where FPGAs win weighs the result: FPGA versus ASIC on speed, power and cost, prototyping chips before tapeout, and jobs where low, predictable latency matters. If you want the hardware this flow targets, go back to Lookup tables and the FPGA fabric. The ASIC versions of each step are in Synthesis, Placement, Routing and Signoff, and the tool landscape in Tools.
Where FPGAs win turns this chapter’s costs into decisions: the area, speed and power gap to an ASIC, NRE and break-even volume, prototyping and emulation, and FPGAs beside processors. For how chip teams use this flow before tapeout, see Verification; for the fabric parameters (K, cluster size, routing architecture) that the algorithms here are tuned for, see Lookup tables and the FPGA fabric.
Uses (ch. 16): Next: FPGA versus ASIC on cost, speed and power, prototyping, low-latency jobs and FPGAs in datacenters.
Q1Why doesn’t an FPGA mapper need to match logic against a library of gate patterns?
Q2Packing puts LUTs and flip-flops into logic blocks. What does it gain?
Q3An FPGA design misses timing by a small margin. Which fix is NOT available to its designer?
Q4What did documentation projects such as IceStorm and Trellis make possible?
Sources
Show Hide 27 sources
- What is Yosys (Yosys documentation)Yosys began as a BSc thesis project by Claire Wolf; supports the synthesizable subset of Verilog-2005; uses ABC for gate-level optimization and mapping; maintained by YosysHQ.
- iCE40 technology library: synth_ice40 (Yosys command reference)The synth_ice40 script: flatten, coarse optimization, memory_libmap to SB_RAM40_4K block RAM, optional DSP mapping, ice40_wrapcarry for carry chains, abc9 mapping to 4-input LUTs, dfflegalize, and JSON output for nextpnr.
- Technology library commands (Yosys command reference)Per-family synthesis scripts, including synth_ice40, synth_ecp5, synth_nexus, synth_lattice, synth_gowin, synth_intel_alm, synth_xilinx and synth_microchip.
- ABC: A System for Sequential Synthesis and VerificationLogic optimization built on And-Inverter Graphs; FPGA technology mapping commands (fpga, if), including variable-LUT-size mapping.
- Combinational and Sequential Mapping with Priority CutsCut enumeration by merging fanin cuts and removing dominated cuts; depth-oriented mapping; area recovery with area flow, then exact local area; a 1M-node AIG mapped to 6-LUTs in about 1 minute with 150 MB; priority cuts give depth-optimum mappings in 95% of cases with one cut per node.
- ESE535 Electronic Design Automation, Day 3: Clustering (LUT Mapping, Delay)A K-LUT implements any K-input function (a library of 2^(2^K) gates); delay-optimal mapping is polynomial while area-optimal mapping with fanout is NP-complete; FlowMap labels each node with its predecessors’ height or one more, using max-flow min-cut, and picks the max-volume cut; covering replicates logic implicitly.
- VPR: A New Packing, Placement and Routing Tool for FPGA ResearchThe routing-resource graph (every track and pin a node, allowed connections the edges); the VPACK packer; simulated-annealing placement with a bounding-box cost and q(n) correction, initial temperature 20σ, 10·N^1.33 moves per temperature, α from the acceptance rate, a range limiter keeping acceptance near 0.44; a PathFinder-based router.
- Using Cluster-Based Logic Blocks and Timing-Driven Packing to Improve FPGA Speed and DensityRouting consumes most of an FPGA’s area and delay; T-VPack greedy packing (seed, then absorb by attraction and criticality); clusters of 7–10 LUTs give the best area-delay, 30–34% less delay than single-LUT blocks.
- Placement and Routing Tools for the Triptych FPGAMcMurchie and Ebeling’s negotiated-congestion router, PathFinder: node cost with present-sharing and history terms, every net rerouted each iteration until no resource is overused.
- AMF-Placer 2.0: Open Source Timing-driven Analytical Mixed-size Placer for Large-scale Heterogeneous FPGAColumnar heterogeneous FPGAs with CLB, DSP and BRAM sites; the placement steps (initial, global, packing and legalization, detailed); annealing placers are slow on large netlists; HeAP was 7.4× faster than VPR 5.0’s annealer with 6% better quality; AMF-Placer 2.0 is within 2.2% and 0.59% of Vivado 2020.2 and 2021.2 critical-path delay.
- nextpnr: a portable FPGA place and route tool (README)Vendor-neutral, timing-driven, open-source FPGA place and route; supported families and their bitstream projects (iCE40/IceStorm, ECP5/Trellis, Nexus/Oxide, Gowin/Apicula, GateMate, NG-Ultra; Cyclone V/Mistral, MachXO2 and Xilinx 7-series/X-Ray experimental); simulated-annealing and HeAP placers; the iCE40 example flow.
- Yosys+nextpnr: an Open Source Framework from Verilog to Bitstream for Commercial FPGAsA fully open flow for iCE40 (up to 8K LEs) and ECP5 (up to 85K); architectures as an API; iCE40 packing merges a LUT and its flip-flop; two timing-driven placers (annealing and HeAP-based analytic) and an A* rip-up-and-reroute router; IceStorm and Trellis supply timing and bitstream data; UP5K routing graph of 125K wires and 1.3M switches; ECP5-85K of 4M wires and 28M switches, a 1 GB database deduplicated to 38 MB; vendor flows remain for production.
- Project IceStorm: overviewDocuments the iCE40 bitstream format and provides tools to analyze and create bitstreams (IcePack, IceTime, IceProg, IceBox); the fully open Verilog-to-bitstream flow with Yosys and a place-and-route tool.
- Project Trellis documentationDocuments the Lattice ECP5 architecture so an open Verilog-to-bitstream toolchain can be built; its database is developed with fuzzers.
- Project X-Ray documentationDocuments the Xilinx 7-series architecture, found with fuzzers for logic, RAM, I/O, clocking and interconnect, to enable an open Verilog-to-bitstream toolchain.
- Configuration (Project X-Ray architecture documentation)Configuration memory is written in frames of 101 32-bit words (100 for the tiles, 1 for the clock row); frame addresses; frames striped across tiles, two words per tile; CRC checks; writes in multiples of the frame size with auto-incrementing addresses.
- F4PGA documentationAn open HDL-to-bitstream toolchain targeting Xilinx 7-series, Lattice iCE40 and ECP5, and QuickLogic EOS-S3; uses Yosys with VPR or nextpnr; defines the FASM format.
- Verilog-to-Routing (VTR) documentation: the VTR flowAn open framework for FPGA architecture and CAD research: elaboration and synthesis (Odin II, with Yosys options), logic optimization and technology mapping (ABC), then packing, placement, routing and timing analysis in VPR, from Verilog plus an architecture description.
- Titan: Enabling Large and Complex Benchmarks in Academic CAD23 benchmarks of 90K–1.8M blocks on Stratix IV: VPR at least 2.7× slower than Quartus II with 5.1× the memory and 2.6× the wire; packing ~78% of VPR’s runtime, placement 49% of Quartus II’s; up to 36.5 h in Quartus II and over 48 h in VPR; the largest MCNC20 circuit took 61 s in VPR and 96 s in Quartus II.
- Improving Logic Levels (UltraFast Design Methodology Guide for FPGAs and SoCs, UG949)The distribution of logic levels must fit the clock frequency goals for the device family and speed grade; synthesis retiming rebalances logic levels by moving registers from short paths into long ones; dedicated blocks and macro primitives as closure techniques.
- Design Closure (UltraFast Design Methodology Guide for FPGAs and SoCs, UG949)Timing-closure methods: implementation strategies and directives, parallel runs, incremental implementation from a reference checkpoint, physical optimization, floorplanning, and separate handling of logic delay, net delay and congestion.
- ILA (Vivado Design Suite User Guide: Programming and Debugging, UG908)The Integrated Logic Analyzer does in-system debugging of implemented designs: it triggers on hardware events and captures data at system speed; probes are added by HDL instantiation or inserted after synthesis; access over JTAG.
- Quartus Prime Pro Edition User Guide: Debug ToolsIn-system debugging tools, including the Signal Tap logic analyzer, which taps design signals into debug logic; trigger conditions; sample depth and RAM type; JTAG access.
- How to use Reveal with Soft JTAG for debugging the design? (Lattice answer database)Reveal Inserter: add signals to trace, set the sample clock, trigger unit and trigger expression; Reveal Analyzer configures the cores, waits for the trigger and shows the captured data.
- Lattice Radiant 3.0 design software (press release)June 2021: Radiant 3.0 for Lattice Nexus-platform devices; timing analysis can run on its own, so designers can try what-if scenarios without re-running place and route.
- Microchip’s first Libero SoC Design Suite release boosts FPGA designer productivity (press release)January 2019: Libero SoC v12.0, one design suite for PolarFire, IGLOO2, SmartFusion2 and RTG4 FPGAs; debug features through SmartDebug; reported runtime reductions for timing (60%) and place and route (25%).
- A Survey and Evaluation of FPGA High-Level Synthesis ToolsHLS turns C, C++ or SystemC into HDL; FPGAs suit it because implementations can be refined and replaced in the device; commercial and academic tools (LegUp, Bambu, DWARV) compared; loop pipelining and the initiation interval, limited by memory ports and loop-carried dependences.