Most chips are fixed. A factory builds the circuit, and it does the same job for the rest of its life. A CPU or a GPU looks flexible because it runs programs. But its wires never change. Only the list of steps it follows does.
An is different. It is a chip whose circuit you can rewire after it leaves the factory. Inside are thousands of small logic blocks and a huge web of wires and tiny switches. A file sets every block and every switch. Load a new file, and the same chip becomes a new circuit.1
The magic part is the lookup table. It is a tiny table of answers that can copy any small rule. The rest of this chapter shows how those tables, wires and switches fit together.
The chapters on CPUs and GPUs described fixed hardware that runs instructions. A field-programmable gate array () has no instruction stream at all. It is a prefabricated chip that can be electrically programmed to become almost any digital circuit. The configuration is the circuit, and it takes effect in under a second instead of the months a custom chip (an ) takes to make.1
An FPGA is a grid of a few kinds of building block, all set by stored bits:
- Lookup tables (): tiny truth tables that can hold any logic function of four to six inputs.
- Flip-flops beside every LUT, grouped with carry logic into .
- Programmable routing (the ): wire segments and switches that join everything up.
- Hard blocks: block RAM, DSP multipliers, and fixed circuits such as high-speed transceivers and processor cores.
- Configuration memory holding all of those settings, loaded from a file called the .
The idea of spatial computing, where the data flows through a circuit laid out for the job, should sound familiar from Dataflow and spatial meshes. An FPGA takes it to the extreme: the layout is down to the individual gate and wire. The cost of that freedom, and where it pays off, is the subject of Where FPGAs win.
An FPGA is a two-dimensional array of tiles: soft-logic clusters of K-input LUTs and flip-flops, columns of hard memory and multiply-accumulate blocks, an I/O ring and hard IP, all embedded in a programmable interconnect. Every programmable element is controlled by a configuration cell, almost always SRAM.1 That buys post-fabrication flexibility at a measured cost against a standard-cell ASIC in the same node: roughly 20–35× the area, 3–4× the critical-path delay and about 10× the dynamic power, and the survey attributes the gap largely to the routing fabric.1
The chapter follows the tile from the inside out:
- LUTs as -bit masks behind a mux tree, and why K settled at 4–6.
- Logic blocks: basic logic elements, clusters, carry chains, and each vendor’s names for them.
- Island-style routing: segments, connection boxes and switch boxes, and why they dominate area and delay.
- Hard blocks: block RAM, DSP slices and hard IP, and how much of the ASIC gap they close.
- Configuration: SRAM, flash and antifuse, bitstream structure, partial reconfiguration.
Products appear only as case studies, from AMD, Altera, Lattice and Microchip documentation and from the open bitstream-documentation projects IceStorm, Trellis and X-Ray. How a design becomes a bitstream (mapping, , placement, routing, timing) is the next chapter, From hardware code to bitstream.
Design A (a small counter): five logic blocks and the wires between them are switched in by the bitstream.
A logic gate follows a simple rule. An AND gate says “answer 1 only if both inputs are 1.” You can write any such rule as a : one row for each mix of inputs, and the answer for each row.
A just stores that table. Each row’s answer sits in its own tiny memory cell. The inputs act like a row number. They work a chooser that picks one cell and passes its answer out.1
Here is the surprise. The table doesn’t work anything out. It reads the answer it was given. Put different answers in the cells, and the same table follows a different rule. With four inputs there are 16 rows, and over 65,000 possible rules.
Why not one huge table with many inputs? Each extra input doubles the number of cells. A table with 20 inputs would need about a million cells! So real FPGAs use tables with four to six inputs, and join many of them up.1
Any combinational logic function of K inputs can be written as a with rows. A K-input stores the output column of that table in configuration cells and feeds them to a whose select lines are the K inputs. The inputs choose a row; the mux passes that row’s stored bit to the output.1
Three consequences follow:
- Every function costs the same. A 4-input AND and a 4-input XOR occupy one LUT each and have the same delay. There are functions of four inputs, and a 4-input LUT can hold any of them.
- The cells are configuration memory. In most FPGAs they are cells, the same six-transistor cell described in Memory cells, written once at power-up.
- The mux is built from switches. SRAM-based FPGAs typically build these multiplexers from pass transistors, which are compact but imperfect switches.1 Pass-transistor and transmission-gate muxes are covered in CMOS logic.
Why four to six inputs?
A bigger LUT does more work, so a circuit needs fewer of them and fewer slow trips through the routing between them. But the cost grows fast: the number of bits doubles with every input, and each extra input pin needs more routing tracks around the block.1 Architecture studies that mapped real benchmark circuits onto FPGAs of different LUT sizes found that total area first falls and then rises with K, while speed improves with diminishing returns, with significant gains up to about six inputs. Commercial FPGAs have moved from 4-input toward 6-input LUTs to capture those gains.1
Today both sizes are in use. Lattice’s iCE40 and ECP5 and Microchip’s PolarFire use 4-input LUTs.152023 AMD’s 7 series uses 6-input LUTs that can split into two 5-input LUTs sharing their inputs.4 Altera’s adaptive logic module has eight inputs feeding a fracturable LUT with a 64-bit mask, enough for one 6-input function or two smaller ones.10
A K-LUT is a -bit mask and a mux, built as a tree of two-input muxes in K levels. Two properties make it the universal soft-logic element: any K-input function fits (there are of them), and the delay through the tree depends only on K and the input pin used, not on the function.1
The LUT-size question is a classic architecture trade-off. Ahmed and Rose’s experiments, summarized in the Kuon–Tessier–Rose survey, swept K with a full synthesis, packing, placement and routing flow over 28 benchmarks. Area per block grows with bits and with the routing tracks K input pins demand, while the block count falls; the product has an interior minimum. Critical-path delay falls monotonically with diminishing returns, with significant gains up to and cluster sizes of three to four.1
Fracturable LUTs
Large LUTs waste bits on the many small functions in real netlists, so modern 6-input LUTs fracture. An AMD 7-series LUT is either one 6-input LUT or two 5-input LUTs with separate outputs and common inputs.4 Altera’s ALM has up to eight inputs shared by two combinational cells and a 64-bit LUT mask, so it can implement two independent 4-input functions, one 6-input function, and some 7-input functions.10 The bits that once fed one wide function can then feed two narrow ones, at the cost of extra input muxing.
LUT masks can also be written while the device runs. In SLICEM slices of the 7 series, about a third of the total, a LUT becomes 64 bits of distributed RAM or a 32-bit shift register.4 Every ECP5 slice can act as distributed RAM.19
C B A = 011 steers the three levels of 2:1 multiplexers to configuration bit 3, which holds 1. The stored pattern 11101000 is the function (Majority).
A 4-input LUT stores 2^4 = 16 bits and can hold any of 65,536 functions. It also needs 4 wires from the routing.
A lookup table answers at once. But many circuits need to remember things, like a counter that knows it is on 7. So next to each table sits a , a one-bit memory that updates once per tick of the clock.
A little switch, set when the design loads, picks what leaves the block: the fresh answer, or the one stored in the flip-flop. A table plus a flip-flop is the basic building brick.1
Bricks come in bundles called . Each company names them differently, but the idea is the same: a handful of tables and flip-flops, with short wires between them.415
Adding numbers needs one more trick. When you add by hand, you carry the 1 to the next column. An FPGA has a special fast wire, a , that passes the carry straight to the next block. Without it, every carry would take the slow way round.5
Pair each LUT with a and a configuration-controlled mux that selects either the registered or the direct output, and you have a basic logic element. Dedicated flip-flops are now universal in commercial FPGAs; they are edge-triggered and usually offer set, reset and clock-enable controls.1
Basic logic elements are grouped into clusters, the , with fast local wiring inside. Signals that stay inside a cluster avoid the slower general routing, and a cluster needs fewer external inputs than the sum of its LUT inputs, because many signals feed several LUTs at once.1 Every vendor has its own names for the same layers:
| Vendor and family | Logic block | Inside it |
|---|---|---|
| AMD 7 series | CLB of 2 slices | Per slice: four 6-input LUTs, eight flip-flops, carry logic4 |
| Altera (Stratix 10, Agilex) | LAB of ALMs | Per ALM: an 8-input fracturable LUT, two full adders, two or four registers10 |
| Lattice iCE40 | Logic tile | 8 logic cells, each a 4-input LUT, a carry unit and a flip-flop15 |
| Lattice ECP5 | Logic tile of 4 slices | Per slice: two LUTs, two flip-flops, fast carry19 |
| Microchip PolarFire | Logic cluster | 12 logic elements, each a 4-input LUT with carry chain and a D flip-flop23 |
Carry chains
Arithmetic is common enough to get its own hardware. An adder computes each bit from the two operand bits and the carry from the bit below. Built purely from LUTs, every carry would cross the general routing. Instead, nearly every modern FPGA has dedicated carry logic and a fast path between neighboring elements: a .1 In the AMD 7 series it runs upward through four bits per slice and cascades from slice to slice; adder delay still grows linearly with width, but each step is short.5 PolarFire’s logic elements chain through a 3-bit carry-lookahead circuit, and each Altera ALM contains two dedicated full adders.2310
The basic logic element (BLE) is a K-LUT, a D flip-flop and an output select; a cluster groups N BLEs behind a local crossbar with I external inputs.1 Clustering grows the logic block quadratically (the crossbar) instead of exponentially (a bigger LUT), and shared inputs keep I well below , which Under the hood quantifies. The crossbar need not be full: one study removed at least half of the cluster-input-to-LUT connections and 50–75% of the feedback connections with no loss in delay or cluster count, about a 10% area saving, at the cost of a router that must search inside clusters.1
Control signals are where clusters differ most. In iCE40 the clock, clock enable and set/reset are shared by all eight cells of a tile, as is the clock-polarity bit, while use of the flip-flop and its set/reset behavior are per cell; the cell’s configuration is 20 bits.15 Some families force these controls to be shared cluster-wide while others let each element select its own; sharing saves routing but means flip-flops with different enables or resets can’t share a block.1
Carry logic in detail
With a dedicated chain, each adder bit uses a LUT plus a carry mux and an XOR, with the carry on a hard wire between neighbors. In the 7 series each bit has a carry mux (MUXCY) and a dedicated XOR, four bits per slice, and CO3 of one slice drives the carry-in of the next to form wider adders.5 iCE40’s carry unit computes the majority of two LUT inputs and the incoming carry, and the chain enters a tile from the tile below.15 Vendors have added carry-lookahead and carry-skip structures to shorten the ripple.1 Because the chain runs up a column, a wide adder occupies a vertical run of slices.5
Registered: the output comes from the flip-flop, updated once per clock edge. Tap a part.
The blocks sit in a grid, like islands. Between them run channels full of wires. That is why this design is called island style.1
Two kinds of switch join things up. Where a block’s pin meets a channel, switches pick which wire the pin uses. Where two channels cross, a lets a wire carry straight on or turn the corner.1
Here is the big surprise about FPGAs. The wires and switches, not the lookup tables, take up most of the chip. They also cause most of the delay, because every switch a signal passes slows it down.2
Short wires are flexible, but a signal going far passes many switches. Long wires skip switches, but waste space on short trips. Real chips mix both.2
Most commercial FPGAs use an island-style layout: logic blocks in a two-dimensional grid, with routing channels on all four sides of each block. Each channel holds a fixed number of tracks, W, chosen when the chip is designed.1 Three kinds of programmable part make up the :
- Wire segments. Each track is cut into segments. A short segment spans one logic block; a long one spans several and passes some switch boxes unbroken. Start points are staggered so every block can reach the start of a wire of each length.1
- Connection boxes. Where a block’s pin meets a channel, switches connect the pin to some of the tracks. The fraction of tracks a pin can reach is called Fc.1 Input connection boxes are built as multiplexers, so each pin picks exactly one track; that saves considerable area over separate switches.3
- Switch boxes. At every channel crossing, a lets a segment connect to segments on the other three sides. The number of choices per segment is Fs; 3 is common.1
Why routing dominates
In an FPGA, “the delay of a circuit … is mostly due to routing delays, rather than logic block delays, and most of an FPGA’s area is devoted to programmable routing.”2 Every programmable switch adds resistance and capacitance to the path, and every switch needs a configuration cell to control it. Most of an SRAM FPGA’s configuration cells set the select lines of routing multiplexers; most of the rest hold LUT contents.1 Studies put 60–70% of FPGA power in the interconnect too.1
The best wire length is a balance. Betz and Rose found that length-1 wires, then common, were worse in both delay and routing area, and that segments spanning 4 to 8 logic blocks did best.2 Real fabrics follow that pattern: iCE40 routing is built from span-4 and span-12 wires, and ECP5 general routing uses wires spanning 1, 2 and 6 tiles.1520 Clocks are separate: dedicated low-skew networks reach every flip-flop without using the general routing.1
The detailed routing architecture is set by a handful of parameters: channel width W, the segment-length distribution (and its stagger), input and output connection-box flexibility and , and switch-box flexibility and pattern.1 Hierarchical routing, used by some older Altera families, gave way to flat island style because real designs’ wire-length distributions rarely match a fixed hierarchy and crossing a level costs a large, fixed delay.1
Switch-box patterns
With , a disjoint switch box connects track k only to track k on the other sides, which splits the fabric into W disjoint domains. The Wilton switch box uses the same number of switches but rotates track numbers on turns, so a route can change domain; both were designed for length-1 wires. With longer segments, most designs make turns only at segment ends and connect mid-wire points sparsely, as in the Imran and “shifty” patterns.1
Directional, single-driver wiring
Early fabrics had bidirectional wires driven through tri-state buffers or pass transistors from several places; Betz and Rose’s study, in that era, found 50–80% pass-transistor switches best.2 Modern fabrics use directional single-driver wires: each segment is driven by one multiplexer at its start, which merges the switch box and output connection box into one routing driver block and gives area and delay gains.1 Both open-documented Lattice fabrics show it: iCE40 routing consists of directional tristate buffers, and ECP5 general routing is unidirectional.1520
Area, delay and power
Because routing dominates area and delay, routing-architecture choices dominate FPGA quality.2 Leakage is a particular problem: after configuration most interconnect is unused, but it still leaks, and 60–70% of total FPGA power has been attributed to the interconnect.1 Architecture research evaluates all this empirically, by routing benchmarks on candidate fabrics; Under the hood shows how.
Logic blocks sit in a grid; channels of wires run on every side. Tap a part, or Step to trace one connection.
Some jobs come up in almost every design: storing lots of numbers, and multiplying. You could build them from lookup tables, but it takes lots of space and runs slowly.
So FPGAs have ready-made blocks in columns across the chip. each hold thousands of bits. each hold a fast multiplier.67
Many FPGAs also have for special jobs. Some send data to other chips very fast. Some even hold a small processor, like the one in a phone.2314
Ready-made blocks help a lot with size. In one study, a design built only from small blocks took 35 times the space of a custom chip. Designs that also used memory and math blocks took about 18 times.1 But if a design doesn’t need them, those blocks just sit there unused.
Modern FPGAs are heterogeneous. Columns of dedicated blocks sit among the logic tiles, each doing one common job far more efficiently than LUTs could.1
- Block RAM. is dedicated SRAM in blocks of tens of kilobits. AMD 7-series block RAMs hold 36 Kb (or two 18 Kb halves) and are dual-ported.6 Altera’s M20K blocks hold 20,480 bits.11 ECP5 has 18 kbit blocks, and iCE40 4-kilobit ones.1916 PolarFire has 20 Kb large blocks and 768-bit small ones.23 For small memories, some LUTs double as distributed RAM.4
- DSP blocks. A is a hard multiplier with adders and an accumulator, the multiply-accumulate () operation behind filters and neural networks. AMD’s DSP48E1 has a 25 × 18 multiplier, a pre-adder and a 48-bit accumulator.7 Altera’s DSP blocks support multipliers from 9 × 9 to 36 × 36 depending on the family.12 ECP5 has 18 × 18 sysDSP slices, iCE40 UltraPlus 16-bit multiply-accumulate units, and PolarFire 18 × 18 math blocks.191723
- Hard IP. blocks are fixed circuits for demanding standard jobs. PolarFire includes 12.7 Gb/s transceivers and two hard Gen2 controllers.23 Processor subsystems turn the FPGA into a : Arm Cortex-A9 cores with a hard DDR memory controller in AMD’s Zynq 7000, Arm Cortex-A53 cores in Altera’s Agilex 7 SoC, and five RISC-V cores in PolarFire SoC.91423 Even the small iCE40 UltraPlus has hard I2C and SPI controllers.17
What hard blocks buy
Kuon and Rose compared a 90 nm FPGA with a 90 nm standard-cell ASIC on the same benchmarks. Designs using only LUTs and flip-flops took 35× the area; designs that also used DSP blocks took 25×, and those using both DSP and memory blocks 18×, because a hard block is about as dense as the ASIC’s own circuit.1 Delay barely moved (3.4× to 3.0–3.5×), because the rest of each critical path still runs through soft logic and routing.1 Because the gap is a statement about economics too, the comparison continues in Where FPGAs win.
Kuon et al. distinguish soft-fabric heterogeneity (flip-flops and carry logic in every logic block) from tile-based heterogeneity (distinct tiles of hard structures, usually in columns).1 Block RAMs are typically aligned in vertical columns within the tile array, and hard blocks reach the routing much as logic tiles do. ECP5 shows it directly: special functions connect through common interconnect blocks, “effectively a logic tile with the slices removed.”19
Memory blocks
Block RAMs are synchronous, dual-ported, aspect-configurable and cascadable. The 7-series RAMB36E1 runs from 32K × 1 to 1K × 36, or 512 × 72 in simple dual-port mode, with optional FIFO logic and ECC.6 M20K runs from 16K × 1 to 512 × 40.11 PolarFire’s 20 Kb LSRAM is 1,024 × 20 and supports ECC at 33-bit width; its 768-bit µSRAMs fill the small-memory niche that LUT RAM fills elsewhere.23
DSP blocks
The multiplier width is a bet on the workload. The DSP48E1’s 25 × 18 multiplier with a 48-bit ALU and cascade paths targets fixed-point filters.7 Altera’s blocks fracture into several narrow multipliers (up to eight 9 × 9 or four 18 × 18) or combine into wide ones.12 The same multiply-accumulate is the core of the systolic arrays in AI accelerators; on an FPGA, a design builds such an array from DSP blocks and routing.
Hard IP and the area bound
The Kuon–Rose numbers are optimistic for hard blocks: only blocks a benchmark used were counted. Real designs must accept the chip’s fixed ratio of logic to RAM to DSP. With every hard block on a Stratix II in use, the area gap could fall to about 4.7×, a lower bound.1 Hard IP pushes the same logic further for interfaces that soft logic can’t reach at all at speed, such as multi-gigabit and PCIe.23
Area: 35× for soft logic alone, falling to 18× when designs used both DSP and memory blocks, because a hard block is about as dense as the custom chip’s own circuit.
Every lookup table cell and every switch needs to be told what to do. That information lives in thousands or millions of tiny memory cells across the chip, called .
The file that fills them is the . It is just a long list of 1s and 0s, one for every cell and switch, plus some instructions for loading them.18
Most FPGAs use SRAM for these cells. SRAM is fast and can be rewritten forever. But it forgets everything when the power goes off. So each time the chip starts, it copies the bitstream in from a small memory chip beside it.1
Some FPGAs keep their design when off. Flash ones work like a USB stick. Antifuse ones have links that are melted shut once and never change. They are safe from copying, but you only get one try.1
Some FPGAs can even change one part while the rest keeps running. That is called .8
is distributed across the whole fabric: one cell beside every LUT bit, mux select and block option. Three technologies are in use, and the choice shapes the rest of the chip.1
| SRAM | Flash | Antifuse | |
|---|---|---|---|
| Keeps its setting when off? | No | Yes | Yes |
| Reprogrammable? | Yes, without limit | Yes, limited cycles | No, one time |
| Storage element | 6 transistors | 1 transistor | 0 transistors |
| Manufacturing process | Standard CMOS | Flash process | Special antifuse steps |
| Switch on-resistance | ~500–1,000 Ω | ~500–1,000 Ω | 20–100 Ω |
Source: Kuon, Tessier and Rose, Table 3.1.1 SRAM won because it needs only standard CMOS, so SRAM FPGAs move to every new process node first. Its costs are size, the need for an external memory to hold the design while powered off, and exposure of the bitstream as it loads; several families encrypt it.1 Some SRAM parts embed the flash on the same chip to load themselves.1 Altera calls its configuration file an SRAM Object File, for SRAM-based devices.13 Microchip’s PolarFire is a non-volatile FPGA family built on a 28 nm non-volatile process.23 Flash parts work at power-up without loading; antifuse parts are smallest and hardest to copy but can be programmed only once.1 The cells themselves are described in Memory cells.
What a bitstream looks like
A is not just an image of the configuration cells. It is a short program for the chip’s configuration controller: commands that select a region, then the data to write there, then checks. In Lattice’s iCE40 the commands set the width, height, offset and number of a memory bank, then write configuration or block-RAM data, followed by a CRC check.18 In AMD’s 7 series, the configuration is organized into frames of 101 32-bit words, each frame covering a slice of a column of tiles.21
Vendors don’t publish full bitstream formats. Open projects have documented several by experiment: Project IceStorm for iCE40, Project Trellis for ECP5, and Project X-Ray for the AMD 7 series, to enable open-source tools.181921 How the tools produce a bitstream from a design is the subject of From hardware code to bitstream.
Partial reconfiguration
Because configuration is written region by region, some FPGAs can rewrite part of the chip while the rest runs. AMD’s Dynamic Function eXchange loads a partial bitstream into a reconfigurable region; the static logic outside keeps working, unaffected.8 A design can then swap accelerators in and out of one region instead of fitting all of them at once.
In an SRAM FPGA, dedicated circuitry initializes every configuration cell at power-up and then writes the user-supplied configuration into them.1 Because every cell sits next to the element it controls, the configuration address space mirrors the floor plan. Project X-Ray’s description of the 7 series makes it concrete: a 32-bit frame address selects a block type (logic and I/O, or block-RAM content), half, clock row, column and minor frame; each frame carries 101 32-bit words, 100 for the tiles in that column of a clock row and one for the horizontal clock row.21
Volatility, security and upsets
SRAM configuration is volatile, so an external memory must hold it, and the stream crossing the board can be intercepted; modern families encrypt it.1 Flash configuration (a floating-gate switching transistor controlled by a smaller programming transistor) is non-volatile and instant-on, but needs a special process and has finite endurance: 500 programming cycles for one cited family.1 Metal-to-metal antifuses occupy no silicon area and switch with 20–100 Ω, but need large programming transistors, can’t be fully tested before programming (one family quoted 90% programming yield), and trail CMOS nodes by generations.1
SRAM configuration cells can also be flipped by radiation, changing the circuit itself. Routing bits cause nearly 80% of configuration soft errors, and an upset-aware router reduced the number of sensitive bits by 14%; triple modular redundancy is the brute-force fix.1
Partial reconfiguration constraints
Frames are the fundamental unit of configuration, so a reconfigurable region is made of whole frames; the static logic outside it keeps running while the region is replaced.218 AMD lists reduced device size, algorithm flexibility and in-field updates among the uses.8
Power on: an SRAM FPGA loads its bitstream from external flash before it works; flash and antifuse parts are live immediately.
Programmable chips are older than FPGAs. In the late 1970s, a chip called the PAL let engineers set up their own simple logic by blowing tiny fuses inside it.25
Later chips joined several PAL-like blocks with one big switch grid. These are called . They are quick and simple. But big, many-step circuits fit them badly.261
The first FPGA came out in 1985. It had 64 logic blocks. Its inventor bet that transistors would get so cheap that wasting a few didn’t matter. He was right.241
Today’s biggest FPGAs have hundreds of thousands of blocks, plus memory, math blocks, fast links and processors.1
Programmable logic began with arrays of AND and OR gates. Monolithic Memories’ PAL (programmable array logic) of 1978 had a programmable AND plane feeding fixed OR gates and flip-flop macrocells, programmed by blowing fuses.251 Altera followed in 1983 with erasable, reusable EPROM-based devices.25 As designs grew, combined many PAL-like blocks around one programmable interconnect; with only one extra stage of switches their delay stays close to a PAL’s, and their configuration is stored on chip, typically in EEPROM.26
The FPGA took a different path. A memory-configured array was proposed as early as 1967, but memory cells cost too much area to be practical until the mid-1980s.1 Xilinx’s XC2064 arrived in 1985: logic blocks whose connections are configured, and reconfigured, by software.24 It had 64 logic blocks and 58 I/O pins.1 Co-founder Ross Freeman’s bet was that transistors would become so cheap that leaving many unused wouldn’t matter.24
Since then, FPGAs have grown by orders of magnitude and gained block RAM, DSP blocks, transceivers and processors.1 The FPGA vendors themselves have changed hands; the box below has the current picture.
Before FPGAs, programmability came from two-level logic. A fuse PROM with N address inputs implements any N-input function, but area grows exponentially; PLAs added programmable AND and OR planes; PALs kept the AND plane programmable and fixed the OR plane, feeding D-flip-flop macrocells.1 The dates vary slightly by source: the survey gives 1977 for MMI’s PAL, the Computer History Museum 1978.125 PAL-style devices route through a near-full crossbar, which grows as inputs times outputs; CPLDs contain that by grouping macrocells into blocks around a central matrix.26
Wahlstrom’s 1967 proposal already had both logic and interconnect configured from a stream of bits in static memory, with storage in each cell; SRAM’s area per switch kept it uncommercial until transistor cost fell in the mid-1980s.1 The survey and the Computer History Museum date Xilinx’s first architecture to 1984; the XC2064 shipped in 1985.12524 By the survey’s writing, the largest parts had about 330,000 equivalent logic blocks and around 1,100 I/Os, plus hard memory and arithmetic.1 Hard IP has kept growing since, up to whole processor subsystems, a theme taken up in Where FPGAs win.
1985: Xilinx’s XC2064, the first commercial FPGA: a grid of 64 configurable logic blocks with SRAM-programmed connections. Its bet was that transistors would get cheap.
Part 1: fill in a lookup table. Tap the memory cells to set the answers, or pick a ready-made rule. Then flip the inputs and watch which cell gets read.
Part 2: build an adder on a tiny FPGA. Tap the dots to close switches and join wires. Your goal is to add three bits, A + B + Cin, and get the Sum and Carry right for all 8 cases. Stuck? Press “Show an answer.”
Part 1 is a single 4-input LUT: 16 configuration bits read by a 16:1 multiplexer. Edit the bits or choose a function, set the inputs, and watch the select path and the hex value of the configuration. Part 2 is a tiny fabric: input pads, three LUTs and output pads beside a channel of five tracks. Things to try:
- In part 1, make a 3-input majority function. Which inputs does it ignore?
- In part 2, build a full adder (Sum = A ⊕ B ⊕ Cin, Cout = majority) with two LUTs. How many routing switches did you use, and how many LUT bits?
- Connect two outputs to the same track. What happens?
The fabric has 120 routing bits (one per crossing and one per switch-box track) against 48 LUT bits, so routing is 71% of the configuration even in this toy. Input pins are modeled as connection-box muxes (one track per pin), as in VPR; outputs may drive several tracks. Nets are found with union-find; a net with two drivers is a short. Try:
- Build the adder with three LUTs (P = A ⊕ B, Sum = P ⊕ Cin, Cout = majority). Where does the channel run out of tracks, and how does opening a switch box let two nets share one track?
- Implement Cout as A·B + P·Cin instead of majority. Count the tracks needed through LUT3’s row: does it still fit in five?
- Load your part-1 table into a fabric LUT (“From the LUT editor”) and watch the bitstream change.
- Configuration bits in a 6-input LUT
- 64
- Functions a 4-input LUT can hold
- 65,536
- FPGA ÷ ASIC area: soft logic vs with DSP + RAM
- 35× vs 18×
- Share of FPGA power in the interconnect
- 60–70%
What these numbers mean:
- 64: a table with 6 inputs has 64 rows, so it needs 64 memory cells. Add one more input and it needs 128.
- 65,536: a table with just 4 inputs can follow that many different rules. Which one is decided by its 16 cells.
- 35 versus 18: in one careful study, an FPGA needed 35 times the space of a custom chip. Using the ready-made memory and math blocks cut that to 18 times.1
- 60–70%: most of an FPGA’s power goes into its wires and switches, not its tables.1
Two readings of the table. First, both LUT sizes are current: the Lattice and Microchip families here use 4-input LUTs, while AMD and Altera moved their large families to 6-input or fracturable LUTs, the trend the architecture studies predicted. Second, every vendor pairs its soft logic with the same kinds of hard block: RAMs of a few to a few tens of kilobits and multipliers of 16 to 25 bits. The area ratios behind the KeyNumbers come from a 2007-era 90 nm study, so treat them as the shape of the gap, not a current measurement.1
Comparing logic capacity across vendors by LUT count is unreliable: a fracturable 6-input LUT, an ALM and a 4-input LUT hold different amounts of logic. Count what a design needs (LUTs at a given K, flip-flops, carry bits, RAM bits by block geometry, multipliers by width) and map it to each architecture. The interconnect share of power and the gap figures come from the survey’s summary of 90 nm-era studies.1 The tiny fabric in the sim spends 71% of its configuration bits on routing, a toy with one-hot switches; real fabrics encode muxes, but routing still holds most of the bits.1
An FPGA pays for being rewirable. It is bigger, slower and uses more power than a chip built for one job. Most of that cost is the wires and switches.1
Every design choice inside it is a balance. Bigger tables do more each, but need many more cells. Longer wires are faster over distance, but waste space on short trips. Ready-made blocks are small and fast, but sit unused if a design doesn’t need them.12
And memory that forgets when off is the price of being able to rewrite the chip forever.1
| You get | You give up |
|---|---|
| A circuit you can change in under a second | Roughly 20–35× the area and 3–4× the delay of an ASIC, mostly in routing |
| Bigger LUTs: fewer levels of logic | Bits that double with every input, and more pins to route |
| Longer wire segments: fewer switches per connection | Wasted track on short connections, less flexibility |
| Hard RAM, DSP and IP: near-ASIC efficiency for those jobs | Dead area when a design doesn’t use them; a fixed mix |
| SRAM configuration: newest process, unlimited rewrites | Volatile; needs a boot memory; larger cells; radiation upsets |
Common ways an architecture choice goes wrong
- Routing too thin. Fewer tracks save area, but designs that can’t route waste the logic around them; architects size channels by routing real benchmarks.3
- The wrong hard-block mix. Measured gains assume the blocks get used; a design that needs more RAM than the chip offers falls back to slow LUT memory.1
- Shared controls. Flip-flops that must share clock, enable and reset signals within a block can strand logic that has the wrong control set.1
- Granularity. Larger LUTs and clusters shorten critical paths with diminishing returns (significant up to K = 6 and N ≈ 3–4) while area per block grows exponentially in K and quadratically in N.1
- Segmentation and drivers. Length 4–8 segments beat length 1 in delay and area; directional single-driver wiring beats bidirectional, and pairs with switch patterns designed for multi-length wires.21
- Heterogeneity. Hard blocks reduce area and (with both DSP and RAM) power, but barely reduce delay, and their fixed ratio is a bet on the workload mix; the full-use bound of 4.7× is far from typical.1
- Programming technology. SRAM tracks CMOS scaling but is volatile, large and upset-prone; flash and antifuse fix volatility and (for antifuse) security and switch resistance, at the cost of process and reprogrammability.1
- Leakage. Most interconnect is idle after configuration but still leaks, so the routing that buys flexibility also sets a large part of static power.1
How these costs compare with building an ASIC, and when the flexibility is worth them, is the subject of Where FPGAs win.
Larger LUTs cut logic depth and routing hops but cost exponentially more bits and more input pins; studies found gains in speed up to about six inputs.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
1. The LUT as a mux tree
A K-input LUT with mask computes
built from two-input multiplexers in K levels. In the sim, bit i is the output for inputs reading i in binary. Two counts drive the architecture: storage bits, and the number of distinct functions . The tree’s delay is K pass-transistor stages plus buffering regardless of the function, which is why the mapper, not the LUT, decides how fast logic is: it chooses how many LUT levels sit on the critical path.1
2. Sizing a cluster
For a cluster of N basic logic elements with K-input LUTs, the number of external inputs needed to use nearly all of them is not KN, because signals are shared. Betz and Rose found for 4-input LUTs; Ahmed and Rose generalized it to
1 For , , rather than 60. That matters because every external input needs an input connection-box mux over a fraction of the channel’s W tracks, and every LUT input needs a local-crossbar mux over the signals inside the cluster.
A rough per-tile count of configuration bits makes the routing share concrete. With encoded muxes:
With illustrative values , , , and : , , and , before any switch-box or wire-driver bits. Even counted compactly, the routing around a cluster needs about as many bits as its LUTs, and the switch boxes add more; that is the survey’s observation that most SRAM cells set routing mux selects.1
3. Describing and evaluating a routing architecture
Architecture research is empirical: map benchmarks to the candidate fabric, place and route them, and measure area and delay with transistor-level models. VPR made that practical. Its architecture file describes logic-block pins, their sides and equivalences, the switch-box style and the Fc values; VPR turns the description into a routing-resource graph in which every track and pin is a node and every allowed switch an edge.3 A standard figure of merit for routing is the minimum number of tracks per channel at which a set of benchmarks routes.3 A simplified description in the style of today’s VPR format:
<architecture>
<complexblocklist>
<pb_type name="clb">
<input name="I" num_pins="33" equivalent="full"/>
<output name="O" num_pins="10"/>
<clock name="clk" num_pins="1"/>
<pb_type name="ble" num_pb="10">
<pb_type name="lut6" blif_model=".names" num_pb="1"/>
<pb_type name="ff" blif_model=".latch" num_pb="1"/>
</pb_type>
<interconnect> <complete input="clb.I ble.out" output="ble.in"/> </interconnect>
<fc in_type="frac" in_val="0.15" out_type="frac" out_val="0.10"/>
</pb_type>
</complexblocklist>
<device> <switch_block type="wilton" fs="3"/> </device>
<segmentlist>
<segment freq="1.0" length="4" type="unidir"/>
</segmentlist>
</architecture>- 1L4I = K/2·(N+1) = 33 for K = 6, N = 10; ‘full’ marks the inputs logically equivalent.
- 2L7Ten basic logic elements, each a 6-LUT and a flip-flop.
- 3L11A full local crossbar; real parts depopulate it.
- 4L12Fc: each input pin reaches 15% of the channel’s tracks, each output 10%.
- 5L15Switch-box pattern and Fs.
- 6L17Directional length-4 wires: inside Betz and Rose’s 4–8 sweet spot.
The tool flow that VPR represents (packing, simulated-annealing placement, negotiated-congestion routing) is covered in From hardware code to bitstream and, for ASICs, in Routing.
4. From routing choices to bits
Once routed, every resource decision becomes configuration: LUT masks, mux select codes on every wire driver and pin, flip-flop and carry modes, hard-block options. The bitstream is those bits arranged in the device’s configuration address space plus the commands to write it. In iCE40 the address space is a set of banks of configuration RAM and block RAM, written by commands that set a bank’s width, height and offset before streaming data, and closed with a CRC check.18
preamble
set bank width (opcode 6) # 16-bit value
set bank height (opcode 7)
set bank offset (opcode 8)
set bank number (opcode 1)
write CRAM data (opcode 0) # width*height/8 bytes of configuration cells
... repeat for each CRAM bank ...
write BRAM data (opcode 0) # initial block-RAM contents
CRC check (opcode 2)
wakeup- 1L2Each write first describes the shape of the region it fills.
- 2L6Configuration RAM: every LUT bit and routing switch in the bank.
- 3L8Block RAM contents travel in the same stream, under separate commands.
- 4L9A checksum guards against a corrupted load before the device starts.
AMD’s 7-series stream is organized around frames instead: a 32-bit frame address (block type, half, clock row, column, minor index) and 101 words per frame: 100 for the 50 tiles of a column within a clock row and one for the horizontal clock row.21 Project X-Ray reconstructed this mapping black-box: scripts it calls fuzzers ask the vendor tools for large numbers of designs, and the resulting bitstreams are cross-correlated to discover what each bit does.22
5. Configuration reliability
Because SRAM configuration cells define the circuit, a radiation-induced bit flip changes the hardware, not just a value. Routing bits account for nearly 80% of such configuration soft errors. One defensive router counted the sensitive configuration bits of each route alongside delay and cut them by 14%; triple modular redundancy masks errors at about three times the logic.1 User flip-flops, which aren’t minimum-size, are less vulnerable than configuration cells, while block-RAM bits are vulnerable and can be protected with error correction.1
You have seen what is inside an FPGA: tables, one-bit memories, wires and switches, all set by one file. But how do engineers make that file? They describe the circuit in code, and software turns it into a bitstream in minutes. The next chapter, From hardware code to bitstream, shows how.
This chapter described the target: LUTs, logic blocks, routing, hard blocks and configuration memory. The next chapter, From hardware code to bitstream, follows a design onto it: synthesis and mapping to LUTs, into logic blocks, placement and routing on the fixed fabric, timing, and the open-source tools that do it. After that, Where FPGAs win asks when all this flexibility is worth its cost. The same steps for a custom chip are in the Design Flow guide, starting at Logic synthesis.
Every architectural parameter here becomes a constraint there. K and fracturability set what technology mapping targets; cluster inputs, shared control sets and carry chains set packing; segment lengths, Fc and Fs shape the routing-resource graph the router searches; and frames and partial regions shape what a bitstream generator writes. From hardware code to bitstream covers those tools, including the open flow built on the documentation projects above, and Where FPGAs win weighs the FPGA–ASIC gap against non-recurring cost.
From code to configured fabric. This chapter explained the right-hand end; the next chapter covers the steps that produce the bitstream. Tap a step.
Q1How many configuration bits does a 6-input LUT store?
Q2Why do FPGAs include dedicated carry chains?
Q3In an island-style FPGA, what does a switch box do?
Q4One study found a soft-logic design took 35× the area of a custom chip, but designs that also used DSP and memory blocks took about 18×. Why?
Sources
Show Hide 28 sources
- FPGA Architecture: Survey and Challenges (Foundations and Trends in Electronic Design Automation 2(2))FPGAs configure in under a second; 20–35× area, 3–4× delay, ~10× dynamic power vs standard cells, mostly from routing; SRAM cells mostly set routing mux selects, the rest LUT contents; SRAM vs flash vs antifuse (Table 3.1); LUT-size and cluster studies (significant returns up to 6 inputs; I = K/2·(N+1)); island-style routing, Fc, Fs, segment lengths; hard blocks and the FPGA:ASIC gap (Table 7.1); 60–70% of power in interconnect; routing causes ~80% of configuration soft errors; Wahlstrom 1967; PAL; first Xilinx FPGA with 64 logic blocks.
- FPGA Routing Architecture: Segmentation and Buffering to Optimize Speed and DensityCircuit delay in an FPGA is mostly routing delay and most of the area is programmable routing; length-1 wires are worse in delay and area; segments of 4 to 8 logic blocks are best; 50–80% of switches should be pass transistors.
- VPR: A New Packing, Placement and Routing Tool for FPGA ResearchArchitecture description file (pins, Fc, switch-block style); every track and pin becomes a node of a routing-resource graph; input connection boxes built as multiplexers save area; architectures compared by the minimum tracks per channel needed to route.
- 7 Series FPGAs Configurable Logic Block User Guide (UG474): CLB OverviewTwo slices form a CLB; four 6-input LUTs, eight flip-flops, multiplexers and carry logic form a slice; each LUT is one 6-input LUT or two 5-input LUTs with common inputs; about a third of slices (SLICEM) can use LUTs as 64-bit distributed RAM or 32-bit shift registers.
- 7 Series FPGAs Configurable Logic Block User Guide (UG474): Carry LogicThe carry chain runs upward, four bits per slice, with a carry mux and dedicated XOR per bit; chains cascade across slices; adder delay grows linearly with operand width.
- Vivado 7 Series Libraries Guide (UG953): RAMB36E136 Kb (or 18 Kb) block RAMs usable as true or simple dual-port RAM, FIFOs or ECC RAM; widths from 32K × 1 to 1K × 36 (512 × 72 simple dual-port); cascadable.
- Vivado 7 Series Libraries Guide (UG953): DSP48E1A dedicated block for compact, high-speed arithmetic: 25 × 18 multiplier, pre-adder, 48-bit ALU and accumulator, cascading between slices, pattern detection.
- Vivado Design Suite User Guide: Dynamic Function eXchange (UG909)Modifying an operating FPGA design by loading a partial BIT file; static logic keeps running while reconfigurable regions are replaced.
- Zynq 7000 SoC Technical Reference Manual (UG585): DDR Memory ControllerA processing system with single or dual Arm Cortex-A9 cores and a hard DDR memory controller beside the programmable logic.
- Adaptive Logic Module (ALM) Definition (Quartus Prime Help glossary)Up to eight inputs; two combinational cells and two or four registers; two dedicated full adders, a carry chain and a 64-bit LUT mask; Stratix 10 has 4 registers per 8-input fracturable LUT.
- M20K Memory Block Definition (Quartus Prime Help glossary)A synchronous true dual-port block of 20,480 bits, 16K × 1 to 512 × 40.
- DSP Block Definition (Quartus Prime Help glossary)DSP blocks implement multiply, multiply-add and multiply-accumulate; multiplier sizes from 9 × 9 to 36 × 36 depending on the family.
- SRAM Object File (.sof) Definition (Quartus Prime Help glossary)The binary configuration file for SRAM-based Altera devices, written by the Assembler.
- Agilex 7 Hard Processor System Technical Reference ManualThe Agilex 7 SoC FPGA provides an Arm Cortex-A53 hard processor system with a variety of hard IP, dedicated I/O and direct external memory access.
- Project IceStorm: LOGIC Tile DocumentationiCE40 logic tile: 8 logic cells, each a 4-input LUT, a carry unit and a flip-flop; 20 configuration bits per cell; span-4 and span-12 wires; all routing is directional tristate buffers; local tracks feed LUT inputs through 16-way choices.
- Project IceStorm: RAM Tile DocumentationEach pair of RAM tiles implements one SB_RAM40_4K block RAM; read/write widths set by configuration bits.
- Project IceStorm: UltraPlus Features DocumentationiCE40 UltraPlus adds DSP units with 16-bit multiply and 32-bit accumulate, 1 Mbit of single-ported RAM, and hard I2C and SPI cores.
- Project IceStorm: Bitstream File Format DocumentationThe iCE40 bitstream is a command stream: set bank width, height, offset and number, then write configuration RAM (CRAM) or block RAM data, with a CRC check.
- Project Trellis: TilesECP5 logic tiles hold 4 slices; each slice has 2 LUTs, 2 flip-flops and fast carry, and can act as distributed RAM; 18 kbit EBRs and 18 × 18 sysDSP slices in columns.
- Project Trellis: General RoutingECP5 general routing is unidirectional; LUT inputs A–D; wires spanning 1, 2 and 6 tiles; general routing spans up to 12 tiles.
- Project X-Ray: Configuration (Xilinx 7-Series Architecture)Project X-Ray documents the 7-series bitstream to enable open-source tools; configuration is organized by clock row, column and frame; each frame is 101 32-bit words, the fundamental unit of configuration.
- Project X-Ray: Database Development ProcessBlack-box method: fuzzers have the vendor tools generate many designs, and the resulting bitstreams are cross-correlated to discover what each bit does.
- PolarFire Family Fabric User Guide (DS60001725)Non-volatile 28 nm FPGAs; logic clusters of 12 logic elements, each a 4-input LUT with carry chain (3-bit lookahead) and a D flip-flop; 20 Kb LSRAM and 768-bit µSRAM blocks; 18 × 18 math blocks; 12.7 Gbps transceivers and PCIe Gen2; PolarFire SoC adds five 64-bit RISC-V cores.
- Chip Hall of Fame: Xilinx XC2064 FPGAThe XC2064 (1985), the first FPGA: logic blocks whose connections are configured and reconfigured by software; Ross Freeman’s bet that transistors would become cheap.
- 1978: PAL User-Programmable Logic Devices IntroducedMMI’s PAL (1978) with fuse programming; Altera’s reusable EPROM-based devices (1983); Xilinx FPGA architectures (1984).
- ELEC 464 Lecture 6: Programmable Logic DevicesPALs compute sum-of-products with macrocells; CPLDs combine many sum-of-product macrocells with one extra programmable interconnect, keep delay close to a PAL’s, and use internal (typically EEPROM) configuration memory.
- AMD Completes Acquisition of Xilinx (press release)Acquisition completed February 14, 2022.
- Intel Corporation Form 8-K (September 12, 2025): sale of 51% of AlteraClosing on September 12, 2025: the purchaser (a Silver Lake affiliate) acquired 51% of Altera’s equity; Intel retained 49%.