Architectures · Chapter 16 of 16 · FPGAs

Where FPGAs win

An FPGA is slower and uses more power than a chip built for one job, but it can be changed. That makes it good for testing new chips, for products made in small numbers, for jobs that change often, and for jobs where every microsecond counts.

FPGAs trade efficiency for flexibility. They win where a custom chip would never pay back its design cost, where the design must change after it ships, when prototyping chips before tapeout, and in networking and trading jobs that need low, predictable latency. Hybrids mix FPGA fabric with fixed processors and engines.

The FPGA–ASIC gap in area, speed and power and how hard blocks narrow it, NRE and break-even volume, field upgrades, prototyping economics, SmartNICs and low-latency trading, datacenter and AI-inference deployments, radiation and obsolescence, and adaptive SoCs, eFPGAs and structured ASICs.

An is a chip you can rewire after it is made. You load a file, and its switches join up into the circuit you want. The two chapters before this one showed how the fabric works and how the file is made.

The other choice is a custom chip, called an . It is built for one job and can never change. Because nothing is wasted on switches, it is smaller, faster and uses less power.

So why would anyone pick an FPGA? Three reasons. A custom chip costs millions of dollars before the first one exists. It takes a year or more to make. And once it ships, it is frozen.

That last point matters more than it seems. Rules for sending data change. Bugs turn up. A product with an FPGA can get a new file over the internet, called a . A custom chip has to be made again.

An (field-programmable gate array) implements a circuit by configuring a prefabricated grid of , memories, multipliers and programmable wiring; Lookup tables and the FPGA fabric shows how, and From hardware code to bitstream shows how a design becomes the configuration file, the bitstream. This chapter is about when that is the right way to build something.

The alternative is an (application-specific integrated circuit): the same design laid out as fixed transistors and wires and manufactured for you. It wins on everything you can measure on a datasheet. Kuon and Rose built the same circuits both ways on the same 90 nm process and found the FPGA versions about 35 times larger, 3 to 4 times slower and about 14 times more power-hungry. That ratio is the .

The FPGA wins anyway in four kinds of situation:

  • Low volume. An ASIC carries a large one-time cost, its (NRE). A custom chip can cost up to $100 million and more than two years to design. Spread over a few thousand units, that is far more than the FPGA’s premium.
  • Change. Requirements move. Microsoft’s Azure networking team wrote that an ASIC took 1–2 years from specification to silicon, during which the requirements kept changing, so new silicon arrived already behind. An FPGA can take a : a new loaded after deployment.
  • Speed to market. A prebuilt FPGA can carry a complete system within weeks, skipping physical design, fabrication and much of the verification an ASIC needs.
  • Prototyping. Before an ASIC exists, FPGAs are the fastest way to run its design.

Those reasons explain where FPGAs show up: networking, wireless base stations, chip prototyping, high-frequency trading, datacenters, and space and defense. The rest of the chapter takes them in turn.

You know an FPGA is a fabric of LUTs, and programmable routing configured by a (see Lookup tables and the FPGA fabric) and compiled by a synthesis, place and route flow (From hardware code to bitstream). This chapter is about the build-or-buy decision and the markets it creates. Three quantities set it:

  • The efficiency gap. Per unit of function, an FPGA spends 35× the area, 3.4–4.6× the delay and 14× the dynamic power on plain logic at 90 nm; hard multipliers and memories cut area to 18× and power to 7.1× but leave delay near 3×. A later study of neural-network accelerators, as reported in a 2021 survey, found the area gap smaller but still about 9× for full-featured designs that use RAM and DSP blocks heavily.
  • NRE versus unit cost. Mask sets alone rose from $65K at 250 nm to $5.7M at 16 nm in Moonwalk’s model, and labor, CAD tools and licensed IP add more. That fixed cost sets a below which the FPGA is cheaper.
  • The value of change. An ASIC is specified once for its whole service life. Azure’s networking team had to spec its SmartNIC logic for 7 years ahead (1–2 years to silicon plus a 5-year server life) and judged that impossible. A turns that risk into a software release.

The rest follows from those: prototyping (time-to-silicon is the cost), latency-critical streaming (where a spatial pipeline beats a processor whatever the gap), datacenter offload (scale amortizes NRE but services change too fast for silicon), and space, defense and long-lived products (low volume, obsolescence, and radiation behavior). The last section closes the guide by placing FPGAs beside CPUs, GPUs and AI accelerators.

ASICfixed at the fabFPGAreconfigurabledesigndesignmasks + fabcompileshipshiprespin + recallbitstreamtime → (not to scale)
1 / 4

1 · Design: both paths start from the same hardware description and the same verification.

The same design as a custom chip (ASIC) and as an FPGA. Step through build, ship and a change after shipping. Schematic, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

Every circuit in an FPGA is built from rewirable parts. A custom chip just uses plain, fixed wires and gates. So what does the rewiring cost?

Two researchers in Toronto measured it in 2006. They built the same circuits both ways, using the same factory process. The FPGA versions needed about 35 times more space on the chip. They ran 3 to 4 times slower. And they used about 14 times more power.

Why so much? Most of an FPGA is not logic at all. It is switches, wires and tiny memory cells that remember how the switches are set. A signal passes through many switches on its way. Each one adds a little delay and wastes a little power.

FPGA makers fight back with ready-made blocks for common jobs, like multiplying and storing numbers. These parts are fixed, just like in a custom chip. With them, the space gap fell by half, and so did the power gap. The speed gap barely changed.

Kuon and Rose ran a set of benchmark circuits through two flows on matched 90 nm processes: an Altera Stratix II FPGA, and a standard-cell ASIC flow. They compared the core logic only, leaving out the input/output circuits.

FPGA ÷ ASICLogic onlyWith hard multipliers and memories
Silicon area35×18× (about 4.7× if every were used)
Critical-path delay3.4× (fastest speed grade), 4.6× (slowest)3.0×
Dynamic power14×7.1×

Where does it go? A general-purpose lookup table (a tiny memory that stores a truth table) replaces a few gates, and a network of programmable switches replaces dedicated wires. Each switch is controlled by a configuration memory cell, so most of the area is routing and configuration, not computation. Signals cross many switches, which adds delay, and the long programmable wires have more capacitance to charge, which adds power.

How hard blocks narrow it

A hard block, such as a multiplier (DSP) block or a block of memory (), is fixed silicon inside the FPGA, so it is nearly as dense as the same function in an ASIC. Designs that leaned on them closed the area gap from 35× to 18×, and the authors estimated about 4.7× for a chip whose hard blocks were fully used. Power followed area. Speed did not: paths still cross programmable routing to reach the fixed block columns.

The gap for a whole product depends on its mix. Azure’s networking team estimated generic FPGA logic at 10–20× an ASIC’s size, but argued their whole SmartNIC was only 2–3× larger, because most of a network card is memory, transceivers and protocol blocks that are hard logic on an FPGA too. For neural-network accelerators, a later study reported in Boutros and Betz’s survey still found about 9×.

Kuon and Rose implemented identical RTL on a 90 nm Stratix II and in a 90 nm standard-cell flow, measuring FPGA area as every used LAB, memory and DSP tile including its share of routing. Logic-only averages: 35× area, 3.4× critical-path delay against the fastest speed grade and 4.6× against the slowest, 14× dynamic power. Static power came out at 87× (typical silicon, 25 °C) or 5.4× (worst case), which the authors declined to interpret: leakage is too process- and screening-dependent to compare across two foundries. It did correlate with area (r ≈ 0.8).

Hard blocks behaved asymmetrically. Area fell to 25× with , 33× with memories, 18× with both; dynamic power to 12×, 14× and 7.1×. Delay barely moved (3.5×, 3.5×, 3.0×), and some circuits got slower with the hard multipliers: 5×5 multiplies gained nothing from 9×9 hardware, 32-bit multiplies wasted 36×36 blocks, and routing to the fixed columns ate any remaining advantage. Boutros and Betz report the same pattern in later FIR-filter studies: hard DSPs gave large area gains and about 2× frequency, still well short of 35×, because each DSP tile carries its own programmable routing interface.

Two caveats on the headline numbers. First, they are core-logic comparisons on one process generation; I/O, hard transceivers and memory controllers are near-parity, which is why whole-system gaps (2–3× in Azure’s SmartNIC estimate) can be far smaller than logic gaps. Second, the area and speed gaps compound if you try to buy back speed with parallelism: to match ASIC throughput by replication you need about 35 × 3.4 ≈ 119× the area, Kuon and Rose’s “at least 119”. Under the hood works through what that implies for hard-block-heavy devices.

FPGA ÷ ASIC, same process1×2×5×10×20×50×1× = ASICArea35×Delay3.4×Dynamic power14×
FPGA resources

FPGA ÷ ASIC for the same circuits on the same 90 nm process, logic only. Tap a bar.

The FPGA–ASIC gap measured on the same 90 nm process (Kuon & Rose, 2007). Log scale; 1× would mean no gap.Share freely with credit: ‘Figure from chipfieldguide.com’

Making a custom chip has two kinds of cost. First is a huge cost paid once, before any chips exist. Then there is a small cost for each chip made.

The one-time cost covers the engineers, the design software and parts of the design bought from other companies. It also covers the , the stencils used to print the chip. For an older factory process, a set of masks cost about $65,000. For a newer one, it was nearly $6 million. A whole custom chip can cost up to $100 million before the first one works.

An FPGA is the opposite. Someone else already paid to design it. You just buy it, at a higher price for each chip.

So it comes down to how many you will make. Make a few thousand, and the one-time cost of a custom chip is a big share of each one. Make millions, and it nearly disappears. The number where the two choices cost the same is the point.

Engineers call the one-time cost , non-recurring engineering. For an ASIC it includes design and verification labor, EDA tool licenses, licensed blocks such as memory controllers and PCIe interfaces, package design, and the . Moonwalk, a UC San Diego study of custom chips for datacenters, modeled each part.

  • Masks rose from about $65K at 250 nm to $700K at 65 nm, $2.25M at 28 nm and $5.7M at 16 nm.
  • Labor, tools and IP can dominate at older nodes (up to 95% of NRE there), while masks can reach 90% at advanced ones.

The FPGA vendor has already paid the NRE for its chip and spreads it over every customer. You still pay your own design work and board, but the per-unit price is higher, and so is the power each unit draws. That gives a simple comparison:

Total cost=NRE+N×(unit price+lifetime energy per unit)\text{Total cost} = \text{NRE} + N \times (\text{unit price} + \text{lifetime energy per unit})

Plot it against volume NN and the two lines cross. The crossing is the :

N∗=NREASIC−NREFPGAper-unit costFPGA−per-unit costASICN^{*} = \frac{\text{NRE}_{\text{ASIC}} - \text{NRE}_{\text{FPGA}}}{\text{per-unit cost}_{\text{FPGA}} - \text{per-unit cost}_{\text{ASIC}}}

For example (made-up but plausible numbers): if the ASIC’s NRE is $11 million more and each FPGA unit costs $220 more over its life, break-even is 11,000,000 ÷ 220 = 50,000 units. Three things move it a lot:

  • The process node. A newer node makes each ASIC cheaper and lower-power but raises the NRE. Moonwalk argues older nodes like 65 nm are often the better choice for modest volumes.
  • Power and service life. Each FPGA unit’s extra watts, paid for every hour of every year in service, are added to its per-unit cost.
  • Change. If the design must change after it ships, the ASIC pays a and possibly a hardware swap, and the FPGA pays for a new bitstream.

The cost of a whole fleet of chips, power and cooling included, is covered in Comparing chips; the cost structure of a new chip project is in Specification.

Model each option as a fixed cost FF plus a variable cost vv per unit, where vv includes the unit price and lifetime energy:

C(N)=F+Nv,v=p+W⋅8,760⋅L1,000⋅eC(N) = F + N v, \qquad v = p + \frac{W \cdot 8{,}760 \cdot L}{1{,}000} \cdot e

with WW in watts, LL years in service and ee dollars per kWh. Then N∗=(FA−FF)/(vF−vA)N^{*} = (F_{\mathrm{A}} - F_{\mathrm{F}}) / (v_{\mathrm{F}} - v_{\mathrm{A}}). The model is linear in everything, so sensitivities are easy: N∗N^{*} scales with the NRE difference and inversely with the per-unit difference. Moonwalk’s NRE model for ASIC Clouds breaks FAF_{\mathrm{A}} into mask cost (a step function of node), backend labor proportional to gate count, frontend labor and CAD (roughly node-independent), IP licensing (rising quickly for DDR and PCIe at newer nodes), package and system design. Its headline conclusion is that node selection is the main NRE lever: for many datacenter workloads an advanced node like 16 nm gave worse total cost of ownership, and older nodes like 65 nm made more accelerators viable, because the NRE saved outweighed the energy and silicon given up.

Two things the linear model hides. First, NRE is paid up front and units over time, so a discount rate raises the effective NRE. Second, schedule: an FPGA product ships a year or more earlier, which for a market with a short window can be worth more than any unit-cost difference. Microsoft’s Catapult team set the FPGA’s budget as a TCO cap instead of a break-even: at most 30% added to the cost of ownership of a server, including at most 10% added power. Total cost of ownership for accelerators is worked through in Comparing chips.

Mask-set cost$100k$1M$10M250 nm$65k180 nm$105k130 nm$290k90 nm$560k65 nm$700k40 nm$1.25M28 nm$2.25M16 nm$5.7M
Show

A full mask set grows about 90× from 250 nm to 16 nm. Tap a bar.

Mask-set cost by process node (Khazraee et al., 2017), and the same cost spread over a production run. Log scale. Masks only; design work costs more.Share freely with credit: ‘Figure from chipfieldguide.com’

This calculator compares an FPGA with a custom chip. The lines show what each one costs per unit, for any number of units. Where they cross is the break-even point.

Slide the number of units from a hundred up to ten million. Watch the custom chip get cheaper as its big first cost is shared out. Then turn on “The design must change halfway” and see what happens.

The calculator plots cost per unit against volume (both on log scales) for an FPGA, an ASIC and a structured ASIC, a middle option covered later in the chapter. Mask costs come from Moonwalk; every other number is illustrative. Things to try:

  • At the defaults, where is the FPGA–ASIC break-even? Which option wins just above it?
  • Switch the ASIC process from 28 nm to 16 nm, then to 65 nm. Which way does break-even move?
  • Turn on “Design changes mid-life” at 50,000 units. Who wins now, and why?
  • Raise FPGA power to 30 W and years in service to 10. How much of the FPGA’s per-unit cost is now electricity?

The Expert view adds electricity price and the FPGA’s own NRE, and prints the break-even arithmetic. The model is Ci(N)=Fi+NviC_i(N) = F_i + N v_i for each option. The structured ASIC is illustrative: 80% of the ASIC’s design cost, a quarter of its mask set, twice its unit cost and power. A mid-life change costs the FPGA 20% of its NRE; a fixed chip pays its custom masks plus 20% of its design cost and replaces the half of its units already shipped. Try:

  • Find the volume band where the structured ASIC wins at the defaults (about 37,000 to 124,000 units). How does it shift with the change switch on?
  • Set electricity to $0.40/kWh and 15 years. How low does the FPGA price have to go to keep break-even above 100,000?
  • With the change on, push the ASIC design cost to $3M. Is the ASIC’s advantage now the masks, the unit cost or the power?
Loading simulation…

Before a custom chip is made, its design has to be tested. The first way is a program on an ordinary computer that pretends to be the chip. It shows everything happening inside. But it is slow: for a big chip, it runs thousands of times a second, while the real chip will tick billions of times a second.

At that speed, just starting up a phone’s software would take months. So teams also load the design onto a board of FPGAs. There it runs millions of times a second. That is fast enough to start real software, plug in real cameras and screens, and let programmers start work long before the chip exists.

Every bug found this way is cheap to fix. A bug found after the chip is made can mean paying for a new set of masks.

Running a design before is covered as a verification method in Verification. Here the question is economic: what does an hour of pre-silicon run time buy, and on what platform? Three platforms trade speed against visibility and cost.

PlatformSpeed on a large designCost and catch
Software RTL simulationkilohertz rangeCheap, compiles fast, shows every signal; far too slow for long software runs
(processor-based)up to about 4 MHzSeveral million dollars; full visibility
about 10–50 MHz on one FPGA; below 1 MHz when split across manyModest cost; slow compiles; poor visibility

The catch with FPGAs is capacity. An FPGA holds far less logic than an ASIC of the same generation (that is the area gap again), so a large chip must be split across many FPGAs, and the wires between them slow the whole system down. That is one reason FPGA vendors build their largest parts from several dies on an interposer: emulation and prototyping platforms are among the customers that need the most logic in one device.

Cloud FPGAs changed the economics for smaller teams. Berkeley’s open-source FireSim runs designs on rented FPGAs: individual simulated nodes boot Linux at tens to hundreds of megahertz, and a 1,024-node simulated cluster used $12.8 million worth of FPGAs for about $100 per simulation hour.

The economic case is time to the bug. A platform’s useful speed is cycles per second of wall clock once compile time is included, and the cost of a bug rises steeply after tapeout, when a fix needs a . Bachrach et al. set out the trade-off: software simulators (kHz on large designs) compile fast with full visibility; processor-based emulators (up to 4 MHz, several million dollars) keep visibility at high cost; FPGA boards (10–50 MHz on one device, sub-MHz across many) are cheap and fast but compile slowly and need re-synthesis to observe new signals. The general platform trade-offs and the testbench problem (non-synthesizable stimulus must stay on a host) are in Verification.

Two structural points matter for the FPGA business. First, capacity: large prototypes are built from the biggest devices available, which is part of why vendors adopted multi-die interposer parts early, and only a fraction of routing tracks cross die boundaries (23% of vertical tracks on Virtex-7 interposer parts), which shapes how designs must be partitioned. Second, elasticity: on cloud FPGAs, FireSim simulated 1,024 quad-core nodes at 3.4 MHz (under 1,000× slower than real time) and runs single nodes at tens to hundreds of megahertz, turning a capital purchase into an hourly cost of about $100 for the full cluster.

Time to finish the job1 ms1 min1 day100 yrRTL simulation120 daysEmulator42 minFPGA prototype5.6 minSilicon5 s
Job

Wall-clock time = cycles ÷ clock rate. Log scale: each grid line is 100× longer. Tap a platform.

Pre-silicon platforms compared on one job. Speeds from Bachrach et al. (2017); the job sizes and the 2 GHz chip are illustrative. Log scale.Share freely with credit: ‘Figure from chipfieldguide.com’

Some jobs need an answer right now, every time. A stock exchange sends out a new price. Trading firms race to react first, and a millionth of a second can decide who wins.

A normal program on a computer takes tens of millionths of a second to react. Worse, the time jumps around from one message to the next. An FPGA works like an assembly line built for that one job. Each message moves through the same steps and takes exactly the same time.

In 2024, an FPGA system set a record: it read the price and started sending an order in just 13.9 billionths of a second. Light travels only about 4 meters in that time.

Why not a custom chip? The way exchanges format their messages keeps changing, and a custom chip can’t keep up. You can read more on the fintech and trading page.

A processor handles a network message through layers of software, caches and interrupts, so its latency is long and variable. An FPGA can process the message as it streams off the wire, in a pipeline that takes a fixed number of clock cycles: .

Trading. Dvořák and Kořenek describe software trading systems as having latency of tens of microseconds that is nondeterministic, and FPGA cards as the usual answer; ASICs are unsuitable because exchange message formats change often. Their FPGA design updated an order book in 253 ns. The STAC-T0 benchmark measures the network I/O part of latency, from the last bit of an incoming message to the first bit of the outgoing order; a June 2024 result on an AMD FPGA card reached 13.9 ns. The fintech industry page covers the rest of that world.

Networking. The same properties matter in a cloud. A is a network card that applies the cloud’s networking rules itself, so the server’s processors are left for customers. Azure compared three ways to build one:

  • An ASIC had the best performance but took 1–2 years to arrive while the rules kept changing.
  • A many-core chip was easy to program but added 10 µs or more per packet, with more variation, at 40 Gb/s and above.
  • An FPGA gave hardware speed with software-like updates. Azure chose it.

Since late 2015, every new Azure server has shipped with one, in a fleet of over a million hosts, giving customers under 15 µs between virtual machines.

Latency-critical streaming is where the FPGA–ASIC gap matters least and the processor–FPGA gap most. A cut-through pipeline starts work on the first bytes of a frame, holds state in on-chip SRAM, and has no instruction fetch, cache miss or scheduler in its path, so its latency is a fixed cycle count.

  • Trading. Book handling in 253 ns with cuckoo hashing into external QDR SRAM, against tens of microseconds, nondeterministic, in software. STAC-T0 measures only network I/O (last inbound bit needed for the decision to first outbound bit) with essentially no trading logic; the June 2024 minimum was 13.9 ns on an AMD Alveo UL3524 with Exegy’s framework.
  • SmartNICs. AccelNet offloads Azure’s software-defined-networking flow tables to an FPGA placed as a bump in the wire between the host NIC and the top-of-rack switch. The paper rejects ASIC NICs (1–2 years spec-to-silicon against a 5-year server life, and firmware-limited embedded cores) and multicore SoC NICs (often 10 µs or more and high variability at 40 GbE and above, and single flows pinned to one core). Reported results: <15 μs<15\,\mu\mathrm{s} VM-to-VM TCP latency and 32 Gb/s, on over a million hosts, built by a team averaging fewer than five FPGA developers.

The design method mattered as much as the hardware: Azure deployed iteratively, measured real traffic, and changed its hashing and caching strategies several times after first release, which an ASIC schedule would not have allowed.

Latency per message10 ns100 ns1 µs10 µs100 µsSoftwaretens of µsMany-core NIC≥ 10 µsFPGA order book253 nsFPGA I/O record13.9 ns
Show

Reaction time per message, log scale: each grid line is 10× longer. Tap a row.

Latency per message for software, many-core and FPGA paths (Dvořák & Kořenek 2014; Firestone et al. 2018; STAC 2024). Log scale; spreads are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Big tech companies run millions of computers. At that scale, even a small saving is worth a lot. In 2014, Microsoft put an FPGA in each of 1,632 servers to help rank web search results. The servers did 95% more work, and used only 10% more power.

Search rules change all the time, so a custom chip would soon be out of date. The FPGAs could be updated along with the software.

Later, the same kind of FPGA was used for AI. Most AI chips work best on many questions at once. But a person typing a search wants an answer to their own question right now. Microsoft’s FPGA design answered one question at a time, very quickly.

Microsoft published three generations of datacenter FPGA work. The figure below steps through them.

  • Catapult (2014). One FPGA board per server, 1,632 servers, with the FPGAs in each half-rack of 48 servers cabled into their own network. Offloading part of Bing’s ranking raised throughput per server by 95% at the same latency, for 10% more power. The authors’ reason for FPGAs over ASICs: datacenter services evolve too quickly for fixed hardware.
  • AccelNet (from 2015). The FPGA moved between the server and the network, becoming the described in the previous section.
  • Brainwave (2018). A neural-network processor built in FPGA logic for real-time . It served requests one at a time (a batch size of 1), where GPUs are least efficient, and reported more than ten times better latency and throughput than GPUs of the time on large recurrent networks.

Why could an FPGA compete on AI at all, given the gap? Because it can be shaped to the model. Brainwave used a narrow number format, a block floating point with mantissas of only 2–5 bits, which maps cheaply onto lookup tables and , and set its vector widths to the model when compiled. (Number formats are covered in Number formats.) FPGA vendors have since added hardware for this: DSP blocks with native 8-bit modes and AI tensor blocks that pack more low-precision multipliers into the same space.

Even so, a 2020 survey put FPGAs at roughly 10–100× a CPU’s inference efficiency, against 100–1,000× for AI ASICs, and noted that FPGAs are rarely used for training. FPGAs win in AI where the model changes faster than silicon, where batch-1 latency matters, or where the AI is one small part of a stream the FPGA is already processing.

Catapult placed one Stratix V on a PCIe daughtercard per server, under 25 W so the PCIe slot could power it, and wired 48 FPGAs per half-rack into a 6 × 8 torus. Bing’s ranking was split across pipelines of eight FPGAs. Results: +95% ranking throughput per server at a fixed latency distribution, or 29% lower tail latency at equal throughput, for +10% power and within a 30% TCO increase. Catapult kept the FPGAs off the datacenter network on purpose, which did not suit networking offload. AccelNet’s converged design instead puts the FPGA between the NIC and the top-of-rack switch, with two PCIe links to the CPUs, so one board serves as both SmartNIC and accelerator and large models can span FPGAs over the network.

Brainwave is the clearest case for FPGA inference. Its NPU is a single-threaded SIMD instruction set over a distributed microarchitecture that dispatches over 7 million operations from one instruction, with up to 96,000 MACs on a Stratix 10 280, and reached 10 to over 35 TFLOPS on large RNNs with no batching. Two FPGA-specific levers made that possible: synthesis specialization (native vector dimension, lane count and matrix tiles chosen per model at compile time) and narrow block floating point (a shared 5-bit exponent per 128 values, 2–5-bit mantissas after brief fine-tuning), which removes per-MAC shifters and packs 2–3-bit multiplies into DSP blocks.

Vendors moved the same lever into silicon. Intel’s Stratix 10 NX replaced DSP blocks with AI tensor blocks (30 int8 or 60 int4 MACs per block at about the same size); Agilex DSPs added int9, fp16 and bfloat16; Versal’s DSP58 supports int8; Achronix’s Speedster7t added machine-learning processing blocks from int16 down to int3 plus fp24, fp16 and bfloat16. Each step trades some generality in the hard block for density on the dominant workload. The ranges CSET reports (FPGA inference efficiency ~10–100× a CPU, ASIC ~100–1,000×) are broad survey estimates, not measurements. Measured benchmarks and TCO methods for AI chips are in Comparing chips.

Catapult (2014): an accelerator cardserverCPUsNICFPGAToR switchnetworkto 47 more FPGAsPCIe
1 / 3

Catapult: one FPGA card per server over PCIe, 48 FPGAs per half-rack cabled into their own network. Bing ranking ran 95% more throughput per server for 10% more power.

Microsoft’s published FPGA deployments: Catapult (2014), AccelNet (2018 paper) and Brainwave (2018). Schematic, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

Satellites, fighter jets and medical implants have a few things in common. Few of them are made. They must work for many years. And they are very hard to fix once they are out there.

FPGAs fit the first point well: there is no huge one-time cost for a small number of units. A NASA report says FPGAs are used in every space application.

Space brings a special problem: fast particles. In most FPGAs, the circuit is set by tiny memory cells. If a particle hits one and flips it, the circuit itself changes. So space FPGAs either keep rewriting their setup, called , or store it in tougher cells that particles can’t flip.

And in an implant, even a small change to the circuit needs government approval first. Being able to rewire a chip only helps if you are allowed to.

Space, defense and long-lived industrial or medical products share low volumes, long service lives and high costs of failure. Low volume favors FPGAs on NRE alone; DARPA notes a custom chip can cost up to $100 million, which is why defense systems often settle for general-purpose parts. The National Academies report FPGAs in the F-35’s radar, communication and navigation systems.

Radiation. In an SRAM-configured FPGA, a in a configuration cell changes the circuit, not just a value, and it stays changed until the cell is rewritten. Designs answer with and with of the configuration. Radiation-tolerant parts harden the device itself. The RTG4, for example, keeps its configuration in total-dose-hardened flash cells and hardens its flip-flops with self-correcting TMR. NASA tested it as a Microsemi part; Microchip bought Microsemi in 2018. More on radiation effects is on the space industry page.

Obsolescence. A 2017 Defense Science Board finding, quoted by the National Academies, is that about 70% of the electronics in a weapons system are obsolete or out of production before it is fielded. The usual fixes are lifetime buys and redesigns, a discipline called . A design kept as portable hardware code can be recompiled for a newer FPGA; Azure reports moving its FPGA code between vendors without much difficulty. See the defense industry page.

Regulation. In regulated products, a field upgrade is still a design change. For the highest-risk U.S. medical devices, changes to circuits or components that affect safety or effectiveness need an approved FDA supplement before they are made. Cars bring their own qualification and safety rules, covered on the automotive and medical pages.

In an SRAM-based FPGA, the configuration memory is a large target whose upsets are persistent functional faults, because its bits define the logic and the routing. NASA Goddard’s error model for FPGAs adds three terms: configuration upsets, functional-logic upsets (flip-flops, ), and single-event functional interrupts in control logic. Two architectural answers exist:

  • Harden the configuration. Flash- or antifuse-configured parts remove configuration upsets. RTG4 combines hardened flash configuration cells with self-correcting TMR flip-flops, SET filters at flip-flop inputs, and hardened SRAM, multipliers and PLLs; the tested RT4G150 has 158,214 4-input LUTs.
  • Mitigate around SRAM configuration. Commercial and radiation-tolerant SRAM parts rely on scrubbing plus design-level redundancy. In a cyclotron test, distributed TMR (triplicating everything but global clocks, resets and enables) beat flip-flop-only TMR. The payoff is reconfiguration in orbit; AMD’s space-grade XQR Versal parts advertise unlimited on-orbit reprogramming with a configuration-memory scrubber built in.

For obsolescence, the FPGA’s advantage is portability of the design, not of the part: the part goes out of production like any other, but RTL can be retargeted to a successor device with a new compile and requalification, where an ASIC would need a redesign on a new process. The cost is that the bitstream and the FPGA supply chain become assets to protect; the National Academies criticize published contract notices that named the exact FPGAs bought for the F-35. In regulated markets, a field upgrade is a regulatory event, so the option value of reprogrammability is discounted by the cost of re-approval.

Configuration bitsLUT: 00 01 10 11routing S0 S1 S2 S300010100SRAM: upsets possibleResulting circuitABLUTpinA Bwantget0 0000 1001 0001 111
Configuration memory

Eight configuration bits: four make the lookup table an AND gate, four set routing switches. Fire a particle at them.

A slice of FPGA fabric and its configuration bits. Inject upsets, scrub, and compare SRAM with flash configuration.Share freely with credit: ‘Figure from chipfieldguide.com’

Today, many FPGAs are not just FPGAs. One chip can hold normal processor cores, a rewirable FPGA part, and special engines for AI math. These are called . The processors run normal programs. The FPGA part handles jobs that need a custom circuit.

It also works the other way around. A custom chip can include a small rewirable corner, called an . Most of the chip is fixed and efficient. Only the part that might change is rewirable.

There is even a half-made chip, the . Its bottom layers are made in advance and shared by many customers. Only the top layers of wires are made for your design. That saves most of the cost of the masks.

The FPGA–ASIC gap is small inside hard blocks and large in the programmable fabric, so vendors keep moving more of the chip into hard blocks. Four hybrid designs result.

  • Adaptive SoCs put hard processor cores, FPGA fabric and arrays of small vector engines on one , joined by a . AMD’s Versal family combines programmable logic, Arm application and real-time cores, and AI Engines, which are small vector processors for inference and signal processing. Altera’s Agilex 5 SoCs pair Arm Cortex-A76 and Cortex-A55 cores with the fabric.
  • CPU + FPGA packages put a server processor and an FPGA side by side in one package, so the FPGA can share the CPU’s memory over a short, coherent link. In UCLA’s measurements, Intel’s co-packaged Xeon+FPGA had lower CPU-to-FPGA latency than a PCIe card, which matters only for jobs that exchange small pieces of data often. Packaging several dies together is covered in Packaging and chiplets.
  • Embedded FPGAs (eFPGAs) are fabric licensed as a block inside an ASIC. SLAC built eFPGAs with the open-source FABulous generator into detector-readout test chips at 130 nm and 28 nm, so the data-reduction algorithm could change after fabrication. FABulous uses the open-source Yosys and nextpnr tools for the user’s flow.
  • Structured ASICs build each design on a prefabricated base, customizing only one or a few metal and via layers. With two custom metal layers, Ahmed et al. estimated their die cost at 2–10× lower than a standard-cell chip at 45 nm: about 10× for small circuits, about 2× for large designs. They also use less power than an FPGA because there are no programmable switches. Adoption has been slower than expected.

The hybrids are different answers to one question: where should the reconfigurable boundary sit?

  • Adaptive SoC. Hard CPUs for control, fabric for bit-level and streaming work, a tiled array of SIMD VLIW vector processors (Versal’s AI Engines) for dense arithmetic, hard memory controllers, and a hard NoC. A hard NoC uses far less area and power than one built in soft logic, runs at a fixed frequency decoupled from the fabric, and makes good use of the few wires that cross interposer boundaries; Versal and Speedster7t both include one. Altera’s Agilex 5 HPS combines Cortex-A76 and Cortex-A55 cores.
  • Co-packaged CPU + FPGA. Intel’s Xeon+FPGA v2 placed an Arria 10 die in the Xeon package with one coherent QPI/UPI and two PCIe links. Choi et al. found QPI-attached shared memory gave the lowest latency, but only 64 KB of coherent cache on the FPGA (under 5% of its on-chip memory), and that link quality matters only for kernels with a low computation-to-communication ratio.
  • eFPGA. Fabric at FPGA density, but only for the volatile function, with an FPGA tool flow. SLAC’s FABulous-based eFPGAs at 130 nm and 28 nm ran a machine-learning classifier for particle-detector data reduction. FABulous reports more than a dozen tapeouts across five processes, including SkyWater 130 nm.
  • Structured ASIC. Metal-programmable bases amortize the expensive lower-layer masks across designs. Ahmed et al. found cost minimized with the fewest custom layers, while a few more layers improve delay and power, and attributed slow adoption partly to immature CAD. Reconfigurability is lost after fabrication; only the NRE changes.
Die (top view)CPU coresProgrammable logicAI EnginesNetwork on chipMemory, I/O
Hybrid

Adaptive SoC: processors, programmable logic and AI engines on one chip, joined by a network on chip. Tap a part.

Four hybrids of fixed and reconfigurable hardware. Dashed outlines mark the parts that stay programmable after manufacturing. Schematic, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’
FPGA ÷ ASIC area, plain logic (90 nm)
35×
FPGA ÷ ASIC delay, plain logic
3.4–4.6×
Mask set, 250 nm → 16 nm
$65K → $5.7M
FPGA tick-to-trade I/O record (2024)
13.9 ns

What these numbers mean:

  • 35×: built from rewirable parts, a circuit took 35 times more space than on a custom chip.
  • 3.4 to 4.6 times slower, because signals pass through many switches.
  • $65,000 to $5.7 million: what one set of stencils cost, from an old factory process to a newer one. That is a big reason small products use FPGAs.
  • 13.9 billionths of a second: how fast an FPGA read a price and started an order.
FigureValueSource
Area, delay, dynamic power gap, plain logic, 90 nm35×, 3.4–4.6×, 14×
Same, with hard multipliers and memories18×, 3.0×, 7.1×
Whole SmartNIC, FPGA versus ASIC (Azure’s estimate)2–3× the area
Mask set at 65 / 28 / 16 nm$0.7M / $2.25M / $5.7M
Upper end of custom chip design costup to $100M, over 2 years
FPGA prototype speed on one deviceabout 10–50 MHz
Catapult: Bing ranking throughput per server, power added+95%, +10%
Azure SmartNIC fleetover 1 million hosts

Read the gap numbers as a 2006-era, 90 nm, core-logic measurement. Architecture has moved since (6-input fracturable LUTs, larger logic blocks, hard NoCs, AI tensor blocks), and, as Boutros and Betz report from a later CNN-inference study, full-featured designs that lean on RAMs and DSPs measured about 9× in area rather than 35×. The direction is stable: hard blocks close area and power, programmable routing keeps delay at a few times an ASIC’s. For the economics, Moonwalk’s mask table is the anchor; NRE components other than masks vary more by project than by node.

An FPGA gives you freedom: start now, change later, and skip the huge first cost. You pay for it with a bigger, slower chip that uses more power, and a higher price for each one.

A custom chip is the opposite. It is cheap and efficient once you make lots of them. But it is slow to get, costly to start, and frozen once it is made.

Many products start on FPGAs. If they sell well and the design settles down, they move to custom chips later. Bitcoin mining machines went that way.

You get (FPGA)You give up
No chip NRE; a product in weeksHigher unit price, so an ASIC wins above some break-even volume (about 48,000 units at the sim’s illustrative defaults)
Field upgrades for bugs and new standardsSeveral times the power and area of an ASIC for the same function
Deterministic, low-latency streamingClock rates a few times lower than an ASIC’s
One part that can become many productsPaying for fabric and hard blocks your design doesn’t use

Common ways the decision goes wrong

  • Ignoring power over the service life. Each FPGA unit’s extra watts are paid every hour it runs; for always-on equipment this can rival the price difference.
  • Assuming the design is frozen. Azure’s requirements changed during the 1–2 years an ASIC would have taken.
  • Underestimating FPGA engineering. FPGAs still need hardware designers and long compiles; Catapult’s authors named programmability the main long-term challenge.
  • Moving to an ASIC too early. When a mining algorithm or trading protocol stops changing, the ASIC wins; the fintech page tells how Bitcoin mining moved from FPGAs to ASICs once its algorithm was fixed.
  • Efficiency versus option value. The FPGA premium buys the right to change the function later. Its value scales with the probability of change times the cost of a respin plus field replacement; it vanishes in regulated markets where every change needs re-approval.
  • Hard-block mix. The gap you pay depends on how much of your design maps to hard IP: Azure’s 2–3× for a NIC dominated by SRAM, transceivers and MACs, against 35× for pure logic. Unused hard blocks are pure overhead, which is why vendors keep specializing families (communications DSPs, AI tensor blocks).
  • Clock rate and timing closure. Programmable routing keeps fabric clocks a few times below an ASIC’s, so FPGAs scale by width and pipelining, not frequency. Brainwave ran at 250 MHz and made up for it with 96,000 MACs; the five platforms in Choi et al.’s study clocked their FPGA logic at 200–400 MHz.
  • Lock-in and supply. Vendor-specific primitives (memory widths, DSP modes) tie RTL to a family; Azure found porting manageable when planned for. The part itself is a commercial product with its own end of life and supply chain risk.
  • The middle options have their own costs. eFPGAs add an FPGA tool flow and fabric area to an ASIC project; structured ASICs give up reconfigurability and have had limited adoption.
Where the answers pointFPGAMiddle groundASIC← flexibilityefficiency →illustrative weights
Volume
Changes after shipping
Power and size
Time to market

Lean FPGA: low volume, expected change and schedule pressure outweigh its higher unit cost and power.

Four questions that push the build-or-buy choice one way or the other. Illustrative weights; the sim is the cost model.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

1. Break-even with change and risk

For options i∈{F,A}i \in \{\mathrm{F}, \mathrm{A}\} with fixed cost FiF_i and per-unit cost viv_i (unit price plus lifetime energy), add a design change that happens with probability qq partway through production. A change costs a fixed RiR_i (a respin for the ASIC, engineering for the FPGA) and replacement of a fraction rir_i of shipped units at unit price pip_i:

E[Ci(N)]=Fi+qRi+N (vi+q ri pi)\mathbb{E}[C_i(N)] = F_i + q R_i + N\,(v_i + q\, r_i\, p_i)
N∗=(FA+qRA)−(FF+qRF)(vF+qrFpF)−(vA+qrApA)N^{*} = \frac{(F_{\mathrm{A}} + q R_{\mathrm{A}}) - (F_{\mathrm{F}} + q R_{\mathrm{F}})}{(v_{\mathrm{F}} + q r_{\mathrm{F}} p_{\mathrm{F}}) - (v_{\mathrm{A}} + q r_{\mathrm{A}} p_{\mathrm{A}})}

The sim sets q=1q = 1 when the change switch is on, rA=0.5r_{\mathrm{A}} = 0.5 (half the units are already in the field), rF=0r_{\mathrm{F}} = 0, RA=masks+0.2DR_{\mathrm{A}} = \text{masks} + 0.2D and RF=0.2FFR_{\mathrm{F}} = 0.2 F_{\mathrm{F}}. At its defaults ($10M design, 28 nm masks at $2.25M, $1M FPGA NRE, $200 versus $20 units, 10 W versus 1.5 W, 5 years at $0.15/kWh) the break-even moves from about 47,700 units to about 67,700. Mask costs are Moonwalk’s; the rest is illustrative. In practice qq is the number to argue about: Azure’s SmartNIC case is q≈1q \approx 1 over a 7-year horizon.

break_even.py (illustrative)text
HOURS = 8760
def per_unit(price, watts, years, usd_kwh):
    return price + watts * HOURS * years / 1000 * usd_kwh

F_fpga, F_asic = 1.0e6, 10e6 + 2.25e6      # NRE: FPGA board/eng; ASIC design + 28 nm masks
v_fpga = per_unit(200, 10.0, 5, 0.15)       # 200 + 65.70  = 265.70
v_asic = per_unit(20,  1.5, 5, 0.15)        #  20 +  9.855 =  29.855

n_star = (F_asic - F_fpga) / (v_fpga - v_asic)   # 11.25e6 / 235.845 ≈ 47,700 units

# mid-life change, q = 1: ASIC respins and replaces half the fleet
F_asic_c = F_asic + 2.25e6 + 0.2 * 10e6     # + masks + 20% of design
v_asic_c = v_asic + 0.5 * 20                # replacements spread per unit
F_fpga_c = F_fpga + 0.2 * F_fpga            # new bitstream engineering
n_star_c = (F_asic_c - F_fpga_c) / (v_fpga - v_asic_c)   # ≈ 67,700 units
  1. 1L5Mask cost is the only cited input (Moonwalk Table 1, 28 nm); design cost is illustrative.
  2. 2L7Lifetime energy is a per-unit cost: 10 W for 5 years at $0.15/kWh is $65.70.
  3. 3L9ΔNRE ÷ Δ(per-unit cost). Doubling the FPGA’s power moves N* down by about a fifth.
  4. 4L12Replacing shipped units raises the ASIC’s per-unit cost, not just its fixed cost.

2. Why the gap compounds: area × delay

If an FPGA implementation is gag_a times larger and gdg_d times slower, matching an ASIC’s throughput by replicating the datapath needs gdg_d copies, so the silicon ratio at equal throughput is

Gthroughput=ga⋅gd≈35×3.4≈119,G_{\text{throughput}} = g_a \cdot g_d \approx 35 \times 3.4 \approx 119,

which is Kuon and Rose’s “area gap of at least 119”, assuming ideal parallelization. This is the right metric for throughput workloads (inference, packet processing), and it explains why FPGA accelerators compete only where their specialization (custom precision, no instruction overhead, batch-1 pipelines) recovers an order of magnitude or more against a programmable processor.

3. How much must map to hard blocks?

A simple mixing model (our arithmetic, not the paper’s): let a fraction ϕ\phi of the ASIC’s area be functions that map onto hard blocks with gap ghg_h, and the rest onto soft logic with gap gsg_s. Then the area gap is

G=ϕ gh+(1−ϕ) gs,ϕ=gs−Ggs−gh.G = \phi\, g_h + (1 - \phi)\, g_s, \qquad \phi = \frac{g_s - G}{g_s - g_h}.

With gs=35g_s = 35 and gh≈1.5g_h \approx 1.5 (Kuon and Rose assumed 1 for large memories and 2 for other hard blocks), reaching their fully-utilized estimate of G≈4.7G \approx 4.7 needs ϕ≈0.9\phi \approx 0.9: about nine-tenths of the work must land in hard blocks. That is the economic logic of adaptive SoCs, AI tensor blocks and hard NoCs: each moves a large, common share of designs into hard silicon so the fabric is reserved for the part that needs it. It is also why a SmartNIC (mostly memory and I/O) sees 2–3× and a logic-heavy design sees 20× or more.

That’s the end of the Architectures guide. Look back at the four kinds of chips. A CPU can run any program, one step at a time. A GPU runs thousands of the same step at once. An AI chip is built for neural networks and does the most work per unit of energy. An FPGA sits in between: you build the exact circuit you need, and you can change it later.

The more jobs a chip can do, the less work it gets out of each unit of energy. That trade runs through the whole guide.

No chip works alone, though. Thousands of them are packed into computers, wired together and cooled in giant buildings. The Systems guide picks up from here.

This guide has followed one trade-off from the most flexible chip to the least.

Every one of these chips ends up in a larger machine: a package, a server, racks joined by networks, and a building that powers and cools them. Those are the subject of the Systems guide, starting with Packaging and chiplets and the server, where FPGA SmartNICs and accelerator cards plug in.

Across the guide, each architecture removes a different overhead and gives up a matching freedom. CPUs keep full generality and spend area on single-thread latency (speculation, caches, coherence). GPUs amortize instruction overhead across warps and matrix instructions. AI accelerators hardwire dataflow and memory hierarchy to tensor workloads. FPGAs remove instructions entirely and pay the programmable-routing tax instead. CSET’s survey ranges for inference relative to a CPU at the same node put GPUs at about 1–10× in efficiency, FPGAs at 10–100× and ASICs at 100–1,000×, with generality falling in the same order. Chapters: 1–4 (CPUs), 5 (GPUs), 6–13 (AI accelerators), 14–16 (FPGAs).

The Systems guide takes the chip as given and covers what surrounds it: Packaging and chiplets, Board and server (PCIe, CXL and the cards SmartNICs sit on), Scale-up fabrics, Scale-out networking, Optics, parallelism and Power and cooling.

Inference efficiency vs CPUgenerality1×10×100×1,000×CPUGPUFPGAAI acceleratorCPU = 1×, same process
Compare

Flexibility trades against efficiency. Ranges are CSET’s estimates for AI inference relative to a CPU. Tap a row.

Flexibility versus efficiency across the guide. CSET (Khan & Mann, 2020) estimates for AI inference versus a CPU on the same process. Log scale.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 4 correct
  1. Q1Kuon and Rose found plain-logic FPGA circuits about 35× larger than ASICs. Adding hard multipliers and memories mainly narrowed which gaps?

  2. Q2An ASIC’s one-time cost is $12 million more than an FPGA design’s. Each FPGA unit costs $240 more over its life, power included. Where is break-even?

  3. Q3Why did Microsoft build Azure’s SmartNICs from FPGAs rather than ASICs?

  4. Q4What is an eFPGA?

Sources

Show Hide 27 sources
  1. Measuring the Gap Between FPGAs and ASICsIan Kuon and Jonathan Rose · IEEE Transactions on Computer-Aided Design 26(2) (author copy, University of Toronto) · 200790 nm Stratix II versus 90 nm standard cells: logic-only area 35×, delay 3.4× (fastest speed grade) to 4.6× (slowest), dynamic power 14×; with hard multipliers and memories area 18×, delay 3.0×, power 7.1×; about 4.7× area if hard blocks were fully used; static power inconclusive (5.4–87×); at least 119× area to match ASIC performance.
  2. FPGA Architecture: Principles and ProgressionAndrew Boutros and Vaughn Betz · IEEE Circuits and Systems Magazine 21(2) (author copy, University of Toronto) · 2021Lower NRE and time to market; a system in weeks; field upgrades by loading a new bitstream; uses (wireless, networking, ASIC prototyping, trading, datacenters); citing a later study of CNN inference (Boutros et al., TRETS 2018), full-featured designs that make heavy use of RAMs and DSPs show a reduced area gap that is ‘still ≈9×’; hard DSP blocks; low-precision DSP and AI tensor blocks (Stratix 10 NX, Agilex, Versal DSP58, Speedster7t MLP); hard NoCs in Versal and Speedster7t; interposer-based FPGAs for capacity and emulation.
  3. Moonwalk: NRE Optimization in ASIC CloudsMoein Khazraee, Lu Zhang, Luis Vega, Michael Bedford Taylor · ASPLOS 2017 (author copy) · 2017Mask-set cost by node (Table 1): $65K at 250 nm, $700K at 65 nm, $2.25M at 28 nm, $5.7M at 16 nm; NRE components (labor, CAD, IP, packaging, masks); masks up to 90% of NRE at advanced nodes, non-mask NRE up to 95% at old nodes.
  4. Circuit Realization at Faster Timescales (CRAFT)DARPACustom ICs can cost up to $100 million and take more than two years to design; the Defense Department often uses general-purpose circuits instead, at a cost in power.
  5. AI Chips: What They Are and Why They MatterSaif M. Khan, Alexander Mann · Center for Security and Emerging Technology, Georgetown University · 2020Table 2: inference efficiency and speed versus a CPU at the same node (GPU ~1–10× / ~1–100×, FPGA ~10–100× / ~10–100×, ASIC ~100–1,000× / ~10–1,000×) and generality; ASICs grow obsolete as algorithms change; FPGAs rarely used for training.
  6. Azure Accelerated Networking: SmartNICs in the Public CloudDaniel Firestone et al. (Microsoft) · USENIX NSDI 2018 (open access) · 2018FPGA SmartNICs on all new Azure servers since late 2015, over 1M hosts; <15 µs VM-to-VM TCP latency, 32 Gb/s; ASICs lacked programmability (1–2 years from spec to silicon, 5-year server life); multicore SoC NICs ≥10 µs with more variability; generic FPGA logic 10–20× larger than ASIC but 2–3× for the whole NIC; fewer than 5 FPGA developers; code ported between vendors.
  7. A Reconfigurable Fabric for Accelerating Large-Scale Datacenter ServicesAndrew Putnam et al. (Microsoft) · ISCA 2014 (author copy, University of Washington) · 20141,632 servers, one Stratix V FPGA each, 6×8 torus per 48 servers; Bing ranking +95% throughput at fixed latency or −29% tail latency; +10% power within a 30% TCO limit; services evolve too fast for non-programmable hardware.
  8. A Configurable Cloud-Scale DNN Processor for Real-Time AIJeremy Fowers et al. (Microsoft) · ISCA 2018 (Microsoft Research copy) · 2018Brainwave NPU: more than 10× better latency and throughput than GPUs on large RNNs at batch 1; 10 to over 35 TFLOPS on a Stratix 10 280 without batching; up to 96,000 MACs; narrow block floating point; synthesis specialization of datapath and precision per model.
  9. Low Latency Book Handling in FPGA for High Frequency TradingMilan Dvořák, Jan Kořenek · IEEE DDECS 2014 (author copy, Brno University of Technology) · 2014Software trading latency tens of µs and nondeterministic; FPGA cards typical; ASICs unsuitable because message formats change often; 253 ns order-book update.
  10. STAC Report: New STAC-T0 results with an Exegy/AMD FPGA solutionSTAC (Strategic Technology Analysis Center) · STAC Research · 2024June 2024: 13.9 ns minimum actionable tick-to-trade network I/O latency on an AMD Alveo UL3524 FPGA with Exegy nxFramework; STAC-T0 measures from the last inbound bit needed for the decision to the first outbound bit.
  11. Cyclist: Accelerating Hardware DevelopmentJonathan Bachrach, Albert Magyar, Palmer Dabbelt, Patrick Li, Richard Lin, Krste Asanović · ICCAD 2017; author copy archived by the Wayback Machine (DOI 10.1109/ICCAD.2017.8203892) · 2017Software RTL simulation in the kHz range; FPGA boards at about 10–50 MHz, sub-1 MHz when split across many FPGAs; FPGAs have much less capacity than ASICs of the same generation; poor visibility; processor-based emulators up to 4 MHz, priced in the several millions.
  12. FireSim: FPGA-Accelerated Cycle-Exact Scale-Out System Simulation in the Public CloudSagar Karandikar et al. · ISCA 2018 (author copy) · 2018Open-source FPGA-accelerated simulation on Amazon EC2 F1 cloud FPGAs; nodes boot Linux at 10s to 100s of MHz; a 1,024-node cluster at 3.4 MHz using $12.8M of FPGAs for about $100 per simulation hour.
  13. The Growing Threat to Air Force Mission-Critical Electronics: Lethality at Risk: Unclassified Summary (Discussion of Selected Topics)National Academies of Sciences, Engineering, and Medicine · The National Academies Press · 2019FPGAs in the F-35’s radar, communication and navigation systems; a published contract for 83,169 Xilinx FPGAs; about 70% of a weapons system’s electronics obsolete before fielding; lifetime buys.
  14. PMA Supplements and AmendmentsU.S. Food and Drug Administration · FDAChanges to design specifications, circuits, components or physical layout that affect safety or effectiveness need an approved PMA supplement before they are made.
  15. Localized Triple Modular Redundancy vs. Distributed Triple Modular Redundancy on a ProASIC3E Reprogrammable FPGAAlex McGuffey, Melanie Berg and Jonathan Pellish · NASA Technical Reports Server (NASA USRP internship final report) · 2010FPGAs are used in every space application; most are radiation-hardened and very expensive; cheaper commercial reprogrammable FPGAs need designer mitigation such as TMR.
  16. NEPP Independent Single Event Upset Testing of the Microsemi RTG4: Preliminary DataMelanie D. Berg, Kenneth LaBel, Jonathan Pellish (NASA GSFC) · NASA Electronic Parts and Packaging Program (NTRS) · 2016RTG4 radiation-mitigated architecture: total-dose-hardened flash configuration cells, self-correcting TMR flip-flops with SET filters, hardened SRAM, multipliers and PLLs; RT4G150 with 158,214 4-input LUTs; FPGA error categories (configuration, functional logic, SEFI).
  17. AMD XQR Versal Adaptive SoCs Enable Next-Generation Signal Processing and AI in Space (SEFUW 2025, non-NDA)AMD · ESA Space FPGA Users Workshop (indico.esa.int) · 2025First 7 nm adaptive SoC for space; on-orbit reconfiguration with unlimited programming cycles; configuration-memory SEU mitigation (XilSEM, internal scrubber); no SEL in testing; MIL-PRF-38535 Class B, up to 7-year missions; XQRVC1902 shipping, XQRVE2302 qualified.
  18. Versal Adaptive SoC AI Engine Architecture Manual (AM009): Introduction to Versal Adaptive SoCsAMD · AMD technical documentationVersal adaptive SoCs combine programmable logic, a processing system with Arm application and real-time processors, and AI Engines (SIMD VLIW vector processors), joined by a programmable network on chip.
  19. Hard Processor System Technical Reference Manual: Agilex 5 SoC FPGAAltera · Altera technical documentationAgilex 5 SoC FPGAs: an Arm Cortex-A76 and Cortex-A55 hard processor system with hard IP, dedicated I/O and external memory access, beside the FPGA fabric.
  20. iCE40 UltraPlus Family Data Sheet (FPGA-DS-02008-2.4)Lattice Semiconductor · Lattice Semiconductor · 2025Ultra-low-power FPGA and sensor manager for mobile applications; 2,800 or 5,280 logic cells; standby current as low as 100 µA typical.
  21. In-Depth Analysis on Microarchitectures of Modern Heterogeneous CPU-FPGA PlatformsYoung-Kyu Choi, Jason Cong, Zhenman Fang, Yuchen Hao, Glenn Reinman, Peng Wei · ACM Transactions on Reconfigurable Technology and Systems 12(1) (NSF Public Access Repository) · 2019PCIe cards, IBM CAPI and Intel Xeon+FPGA v1/v2 compared; v2 co-packages the CPU and an Arria 10 FPGA with one coherent QPI/UPI and two PCIe links; QPI gives lower latency; only 64 KB coherent cache; communication matters only for low compute-to-communication kernels.
  22. Embedded FPGA Developments in 130nm and 28nm CMOS for Machine Learning in Particle Detector ReadoutJulia Gonski et al. (SLAC National Accelerator Laboratory) · arXiv:2404.17701 · 2024eFPGA puts reconfigurable logic inside an ASIC; open-source FABulous fabrics fabricated and tested at 130 nm and 28 nm; a machine-learning classifier for detector data reduction validated on the eFPGA.
  23. FABulous: an Embedded FPGA Framework (documentation)FABulous project (University of Manchester and contributors) · Read the Docs (open-source project docs)Open-source eFPGA fabric generator using Yosys and nextpnr for the user flow; Apache 2.0; 12+ tapeouts on TSMC 180 nm, SkyWater 130 nm, IHP SG13G2, GF180MCU and 28 nm.
  24. Performance and Cost Tradeoffs in Metal-Programmable Structured ASICs (MPSAs)Usman Ahmed, Guy G. F. Lemieux, Steven J. E. Wilton · IEEE Transactions on VLSI Systems (author copy, University of British Columbia) · 2010A structured ASIC is partially fabricated with generic masks and customized with one or more metal/via layers; lower power than an FPGA (no programmable switches); fewest custom layers minimize cost, a few more improve delay and power; abstract: ‘With two custom metal layers, MPSAs can be 2×–10× cheaper than cell-based ICs (CBICs)’, a 45 nm die-cost comparison: about 10× for small circuits, 2× for large designs with embedded macro blocks; slower adoption than expected.
  25. AMD Completes Acquisition of XilinxAMD · AMD press release · 2022February 14, 2022: AMD completed its all-stock acquisition of Xilinx.
  26. Microchip Technology Announces Completion of Microsemi Acquisition (Exhibit 99.1 to Form 8-K)Microchip Technology · Microchip press release, filed with the U.S. Securities and Exchange Commission (EDGAR) · 2018Microchip completed its acquisition of Microsemi Corporation; filed with Form 8-K on May 29, 2018, the merger date (the release’s dateline misprints the year as 2016).
  27. Intel Corporation Form 10-Q for the quarter ended September 27, 2025 (Note 9: Divestitures)Intel Corporation · U.S. Securities and Exchange Commission (EDGAR) · 2025On September 12, 2025, Intel completed the sale of 51% of Altera’s common stock to an affiliate of Silver Lake Partners and retained a 49% minority investment.