Extras · Industries · Fintech and trading

Fintech and trading

Nanosecond trading on FPGAs, certified secure chips in payment cards and HSMs, and single-function mining ASICs.

STAC-T0 tick-to-trade I/O record, June 2024 (FPGA)
13.9 ns minimum
EU clock rule for high-frequency traders
Within 100 µs of UTC, 1 µs stamps
First ASIC Bitcoin mining rig delivered
January 2013
16 nm Bitcoin miners vs first 130 nm ASICs
~80× more energy efficient

At a glance

Where the flow bends

  1. 01Specification

    For trading machines, the main goal is a time limit far shorter than a blink. For bank card chips, the goal is passing a security test run by outside experts.

    For trading hardware, the specification sets a time limit: how long from a price update arriving to an order leaving (the tick-to-trade latency), and how much that time may vary. For payment and banking security chips, it names the security certificate the chip must earn, such as a Common Criteria level, EMVCo approval for payment cards, or the U.S. FIPS 140-3 standard for cryptographic modules.

    Specify latency as a distribution, not one number, at a defined measurement point (STAC-T0, for example, measures from the last inbound bit needed for the decision to the first outbound bit), with limits on both the minimum and the slow tail. For a secure element, the protection profile’s threats (leakage, probing, malfunction, physical manipulation, abuse of functionality) become design requirements from the start.

  2. 02Architecture

    Trading chips start working on a message before all of it has arrived. Mining chips copy one small circuit many times.

    Trading designs use cut-through processing: they start reading a message while its bytes are still arriving instead of waiting for all of it, keep the table of current prices (the order book) in fast memory, and decide in a fixed number of clock ticks. The main platform choice is a reprogrammable FPGA or a custom ASIC.

    Keep queues, arbitration and shared resources off the critical path, because each adds delay that varies from message to message. The usual budget items are off-chip memory lookups for large order books and crossings between clock domains.

  3. 03RTL design

    The design works like an assembly line. Every step takes exactly the same time.

    The circuit is written as a deep assembly line (a pipeline) in which every stage takes exactly one clock tick and nothing ever waits. The parsers that read the exchange’s market-data messages and write orders are built directly in hardware.

    Fixed-latency pipelines, parsers that handle several message types in parallel, and narrow datapaths that start work on the first bytes of a field. Mining RTL is the extreme case: the 64 rounds of SHA-256 laid out as 64 pipeline stages, finishing one hash per clock.

  4. 04Verification

    Tests run by themselves every time the design changes. They also check the safety limits that stop bad trades.

    Tests run automatically on every change, using open-source simulators such as Verilator and the Python library cocotb. The built-in safety limits on trading (risk checks) are verified as carefully as the trading logic itself.

    Turnaround matters more than methodology: some firms skip UVM, the standard verification framework, for C++ and Python testbenches that reuse their software code. Secure chips add a separate track of penetration testing and side-channel evaluation by accredited labs.

  5. 06Design for test

    Factories need a way to look inside each chip to test it. A thief could use that same way in, so on security chips it gets locked after testing.

    Chips include test features, such as scan chains that let the factory read out every stored bit, so each chip can be checked. On a security chip those same features could let an attacker read out secrets, so they are locked or disabled after factory testing.

    The security IC protection profile requires preventing abuse of functionality across the chip’s life cycle, and test modes are the obvious case. Test access is designed together with fuses, life-cycle states and tamper sensors, so that it closes once the chip leaves the factory.

  6. 10Clock tree synthesis

    Parts of a chip can tick at different speeds. Passing a signal between them costs a tiny wait, so trading designs keep these handoffs rare.

    Different parts of a chip can run on different clocks. Each time a signal passes from one clock to another it must wait to be captured safely, so trading designs keep these crossings, from the network receiver into the core logic, to a minimum. The clock used for timestamps is kept in step with world time (UTC), often with the Precision Time Protocol (PTP).

    Each asynchronous crossing costs synchronizer cycles and adds variation. Timestamp logic needs a clock traceable to UTC, and the EU rule asks firms to identify the exact point in the system where a timestamp is applied and show that it stays consistent.

  7. 13GDS & tapeout

    Most trading designs never become custom chips. They run on chips that can be rewired overnight by loading a new file.

    Most trading logic never becomes a custom chip: it ships as a configuration file (a bitstream) loaded onto an FPGA, which can be replaced overnight. A custom ASIC makes sense only when the logic is stable enough to justify the cost of its masks. Bitcoin miners, by contrast, raced to each new manufacturing process.

    ASIC NRE and months of fab time conflict with fast-changing strategies. In mining, the efficiency gain of each new process node made the previous generation unprofitable, so tapeout timing decided which companies survived.

Banks and traders use special chips for two very different jobs: speed and trust.

Speed. On a stock exchange, prices change all the time, and the first order to arrive wins. So some trading firms build their trading rules right into hardware that sits on the network cable. It can react to a price change in a few billionths of a second. In that time, light crosses only a small room.

Most of this hardware is an , a chip that can be rewired by loading a new file.

Trust. The chip on a bank card holds secret codes, called keys. It must keep them safe even if a thief takes the card to a lab full of tools. Outside experts test and approve these chips before banks can use them.

Trading firms keep their hardware secret. This page uses only facts that have been made public.

Electronic trading rewards speed. The key measure is latency: the time from a price update (a “tick”) arriving from the exchange to an order going back out. Software on an ordinary processor takes on the order of tens of microseconds (millionths of a second), and the time varies from message to message. That is why trading systems widely use cards, chips that can be configured into custom circuits, to process market data.

Public benchmarks show how far hardware has pushed this. STAC-T0 measures from the last bit of incoming data needed for a decision to the first bit of the outgoing order. In a 2024 test that AMD asked STAC to run, an FPGA system posted a minimum of 13.9 nanoseconds (billionths of a second), beating the previous record. That benchmark covers only getting data in and out of the network, with essentially no trading logic in between, so a real strategy adds its own time.

Consistency matters as much as raw speed. A circuit built for one job gives : the same message takes the same number of clock ticks every time.

The second job is trust. Payment cards and the machines banks use to guard their keys contain chips designed to keep secrets even when an attacker holds them, and these chips must pass independent security certification.

Why FPGAs dominate trading. Exchange message formats change, and strategies change faster. A 2014 FPGA trading paper states plainly that ASIC technology is not suitable for decoding exchange messages, because the format of incoming messages often changes. Hudson River Trading, writing about how it verifies its own trading hardware, notes that FPGAs are increasingly used in trading and that trading is particularly time-to-market sensitive.

When ASICs make sense. An trades (the one-time cost of design and masks) and months of fab time for lower latency, power and unit cost. That pays off for functions that change rarely, such as network interfaces or stable compute kernels. Startups have publicly pursued trading ASICs: in 2021 Bloomberg reported on one building an AI chip for high-frequency trading, with a foundry slated to manufacture it in 2022. Those were the company’s plans, not shipped results.

Software (CPU)FPGA circuitnetwork I/O only2020: 24.2 ns1 ns10 ns100 ns1 µs10 µs100 µslight goes ≈4 mlog scale →

Send a burst of price ticks to both systems.

Tick-to-trade on a log scale. Software times are an illustrative spread of “tens of microseconds”; 13.9 ns is the 2024 STAC-T0 minimum, which counts network in and out only.Share freely with credit: ‘Figure from chipfieldguide.com’
  • Being first. A few billionths of a second can decide who gets the deal.
  • Being steady. A machine that is usually fast but sometimes slow is risky. So designers remove anything that causes random delays.
  • Knowing the time. The law says trading firms must record when each trade happened. In Europe, the fastest traders must keep their clocks within a ten-thousandth of a second of world time.
  • Beating thieves. Bank card chips must keep their keys secret, even when an attacker holds the chip.

The network is part of the race. A trading message travels over Ethernet, so the circuitry that receives and sends network data sits on the critical path. The fastest results come from specialized network circuit blocks: a 2020 STAC-T0 result reached a 24.2 ns minimum using dedicated blocks for the network protocol (TCP) and for the lowest Ethernet layers on an FPGA.

Clocks and timestamps. Regulators need to reconstruct who did what, when. In the EU, trading firms using high-frequency algorithmic trading must keep their business clocks within 100 µs of UTC (world time) and record timestamps in steps of 1 µs or finer. The (IEEE 1588) synchronizes clocks across a network to better than a microsecond, and to better than a nanosecond in a properly designed network.

Certified security. The chip in a payment card is a , a small processor built to keep keys secret. Before it can be used, it goes through EMVCo’s approval process: a recognized laboratory evaluates the chip, and EMVCo issues a certificate if the evaluation report is complete, including vulnerability analysis and penetration testing. Banks keep their master keys in , which are validated to FIPS 140-3 under the Cryptographic Module Validation Program, run jointly by the U.S. NIST and the Canadian Centre for Cyber Security; independent accredited laboratories do the testing.

Where the nanoseconds go. A tick-to-trade path has a fixed sequence of steps:

  1. Receive bits from the cable and recover the Ethernet frame (the transceiver, PCS and MAC layers).
  2. Parse the exchange’s message to find the instrument, price and quantity.
  3. Look up and update that instrument in the .
  4. Decide, then build the outgoing order message.
  5. Send it back through the MAC, PCS and transceiver.

Memory is often the bottleneck in step 3. A published FPGA design needs about 105 Mbit to hold 100,000 instruments, too much for on-chip memory, so the book lives in external QDR SRAM. To find an instrument’s record quickly it uses cuckoo hashing: each key may live at only a few addresses, one per hash function (three in this design), so a lookup reads those few addresses and compares keys, with no long search. The table itself is built in software: an insertion that finds every candidate address taken evicts an occupant and moves it to one of its other addresses, repeating until everything fits. The hardware only does lookups. The design updates a book of 119,275 instruments in 253 ns on average using 144 Mbit of QDR SRAM. Every clock-domain crossing and off-chip access shows up in the slow tail of the latency distribution.

Timestamp placement. The EU rule requires firms to document a system of traceability to UTC, identify the exact point in the system where a timestamp is applied, show that this point stays consistent, and review compliance at least once a year. In hardware, that means time-stamping at a fixed pipeline stage near the network interface.

Secure elements. Smart-card chips are certified under against a protection profile, a standard list of threats and requirements. The one for security ICs, written by Infineon, NXP, STMicroelectronics and Inside Secure, claims EAL4 augmented by AVA_VAN.5 (advanced methodical vulnerability analysis) and ALC_DVS.2 (security of the development environment). Its requirements protect data against malfunction, leakage, physical manipulation and probing, and prevent abuse of functionality. Leakage covers : power measurements can reveal keys, even when the cryptography is a small fraction of the chip’s total power. These countermeasures live in the RTL, the layout and analog sensors, so they can’t be added late.

FPGAinoutReceivePHY·PCS·MACParsemessageOrder booklookupDecide+ buildSendMAC·PCS·PHYSRAMoff-chipUTC via PTPtimestamp
1 / 6

A tick arrives. It is timestamped at a fixed point by the network port, from a clock synced to UTC by PTP. EU rule for the fastest traders: within 100 µs of UTC, 1 µs steps.

The tick-to-trade path from the list above, inside one FPGA, with the timestamp point and the external order-book memory. Step a tick through.Share freely with credit: ‘Figure from chipfieldguide.com’
  1. Trading: change often. Trading designs run on FPGAs. So they can be updated whenever the plan or the exchange’s rules change.
  2. Trading: test all the time. Tests run by themselves after every change. They check the safety limits on trading, too.
  3. Bank card chips: invite attackers. Before a chip is approved, outside labs try hard to break it.
  4. Mining chips: copy one block. The chip is one small circuit, copied over and over.

An FPGA flow. A custom chip is designed by writing code that describes the circuit, testing it in simulation, turning it into logic gates, laying those out on silicon and sending the layout to a factory. An FPGA design shares the first half. The second half is replaced by the FPGA maker’s tools, which fit the design onto the chip’s existing blocks and produce a configuration file instead of a layout for a factory.

Testing for fast change. Hudson River Trading describes running its designs in simulation with two open-source tools: Verilator, which turns a hardware design into a fast C++ program, and cocotb, which lets tests written in Python drive and inspect the simulated signals. An automation server reruns the tests on every change to catch regressions and try new random test inputs. Verification checks both that the hardware behaves as intended and that it obeys the firm’s risk checks, the built-in limits on what it may trade.

Security chips. The evaluation report behind an EMVCo chip certificate includes vulnerability analysis and penetration testing by a recognized lab, and the development and production sites may need audits too. Under Common Criteria, the security of the development environment is assessed as well.

Timing closure for latency. Latency through a pipeline is the number of stages times the clock period. For example (illustrative numbers), eight stages at 400 MHz, a 2.5 ns period, take 8×2.5 ns=20 ns8 \times 2.5\,\mathrm{ns} = 20\,\mathrm{ns}. Teams close timing at a fixed, high clock and then fight to remove stages, since each one saved is a whole clock period. Cut-through designs act on fields before the frame’s checksum (which arrives last) has been checked, so they need a way to cancel work on a corrupted frame. Expect custom network IP, as in the published STAC-T0 stacks, along with as few clock-domain crossings as possible and placement constraints that pin the critical path next to the transceivers.

Methodology choices. HRT reports skipping UVM, the standard SystemVerilog verification methodology, in favor of C++ and Python testbenches, for code reuse with the languages it already uses. The trade is less standard methodology for faster turnaround.

How a power attack works. Differential power analysis (DPA) needs no access inside the chip:

  1. Record the chip’s power consumption (a trace) during many encryptions, along with the data going in or out.
  2. Guess one small part of the key, such as one byte. For each trace, use the guess to predict one intermediate bit inside the computation.
  3. Split the traces into two groups by that predicted bit, and subtract the group averages point by point.
  4. With a wrong guess, the groups are random and the difference flattens toward zero. With the right guess, a spike appears where the chip handles that bit. Repeat for each key byte.

Masking against it. Masking stores a secret value XX as two random shares: a random byte RR and A=X⊕RA = X \oplus R (⊕\oplus is bitwise XOR). Neither share alone reveals XX, and refreshing the random share regularly stops statistical tests like first-order DPA from accumulating information about the secret. Implementation tools must not recombine those shares through logic optimization or sharing, so masked blocks often get special synthesis constraints. Test and debug access must be closed by life-cycle state after manufacturing test, since the protection profile requires preventing abuse of functionality, and resistance to malfunction and physical manipulation is part of the same requirement set. AVA_VAN.5 is decided by the lab’s attacks on real silicon.

power traces (3 of 200)difference of group averages, guess 0x4bit handled herepeak per guess0x00xF
Traces

Guess 0x4: peak 0.35. Tallest of all 16 guesses: 0xB.

A toy power attack on synthetic traces: guess a 4-bit piece of the key, split the traces by the bit your guess predicts, and subtract the group averages. Illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
RTLSimon every changeshared front halfFitBitstreamLoadovernightLayoutMasksFabLab testmonthsattack, certifyFPGAASIC
Product

Trading: same front half (RTL, simulation on every change), then the FPGA tools fit the design and produce a bitstream instead of a layout.

The FPGA flow shares the front half of the chip flow and swaps the back half for a bitstream. Pick a product and push a change through.Share freely with credit: ‘Figure from chipfieldguide.com’

Bitcoin mining is a guessing race. A miner scrambles a block of data into a short code, a bit like a fingerprint. If the code is below a target number, the miner wins a prize. If not, it changes one number and tries again.

More guesses per second means more chances to win. But every guess uses electricity, and the power bill is the biggest cost.

Miners first used ordinary computers, then graphics cards, then FPGAs. Then came chips built only for mining. The first of those machines reached a buyer in January 2013.

By 2017, a mining machine made about 80 times more guesses per unit of electricity than the first mining chips, and about 3,000 times more than graphics cards.

Mining is the clearest example of a chip built for a single function, and Michael Bedford Taylor documented its history in IEEE Computer. A miner searches for a number (the nonce) that makes a double SHA-256 hash of a block header fall below a target. The hash can’t be run backwards, so the only approach is brute force: try nonce after nonce. The network adjusts the target so that, across all miners, a block is found about every 10 minutes. The hardware story, with hash rates in millions (MH/s), billions (GH/s) and trillions (TH/s) of hashes per second:

  • One assembly line, copied. SHA-256 runs 64 rounds of simple operations, each depending on the last. FPGA miners laid out all 64 rounds as separate hardware, with storage between them, so a new guess entered every clock tick and one finished every tick. Early ASICs copied that design.
  • ASICMiner (130 nm). Each USB-stick miner’s chip hashed at 330 MH/s using 2.5 W, about 40 times more energy efficient than a 28 nm graphics card.
  • Butterfly Labs (65 nm). Each chip held 16 double SHA-256 pipelines on a 7.5 × 7.5 mm die. Taylor writes that preorder revenue presumably covered the $500,000 of one-time mask costs.
  • Bitmain Antminer S9 (16 nm). 189 chips in one shoebox-size machine delivered 13.5 TH/s using 1,323 W.

The mining story compresses ASIC economics into a few years.

  • Power estimation failures are fatal. Butterfly Labs’ chip drew four to eight times more power than expected. Every system had to be redesigned, and clearing the order backlog took nearly a year.
  • Time to market beats cost efficiency. HashFast and CoinTerra built cost-efficient 28 nm chips, but at more than 1.1 W per GH/s they were less energy efficient than BitFury’s 55 nm parts, which had shipped months earlier. That contributed to both companies going out of business.
  • Not every design is a long pipeline. BitFury’s chips used “rolled” hashes that iterate in place rather than unrolled pipelines, and connected chips’ power pins in series, removing the DC–DC converters that make up 20 to 40 percent of a mining server’s cost.
  • Low voltage is the endgame. The 16 nm leaders run at ultralow voltages. Mining revenue per GH/s falls as total network hash rate rises, and when it drops below a machine’s energy cost per GH/s the machine should be switched off; lowering the voltage buys a few extra months of profitable life.
  • Vertical integration. Leading miners co-design the ASIC, the machine and the datacenter, which removes the need to support varied customer environments.

The same pattern appears in trading. When the function is fixed and the payoff per unit of latency or energy is clear, custom silicon wins. When the function keeps changing, the FPGA’s flexibility is worth more than the ASIC’s efficiency.

0.0010.010.1110100GH per joule (log) →28 nmGPU×1130 nmASICMiner40×28 nmHashFast, CoinTerra<275×16 nmAntminer S93,092×
Ratio relative to

Hashes per joule of electricity, on a log scale: each gridline is 10× more. Tap a bar.

Energy efficiency computed from the specs above (log scale). The S9 figure is for the whole machine; HashFast/CoinTerra is an upper bound.Share freely with credit: ‘Figure from chipfieldguide.com’

Sources

Show Hide 12 sources
  1. Low Latency Book Handling in FPGA for High Frequency TradingMilan Dvořák, Jan Kořenek · IEEE DDECS 2014 (author copy, Brno University of Technology) · 2014FPGA cards widely used; software latency tens of µs and nondeterministic; ASICs unsuitable as message formats change; cuckoo hashing in external QDR SRAM; 253 ns book update.
  2. STAC Report: New STAC-T0 results with an Exegy/AMD FPGA solutionSTAC (Strategic Technology Analysis Center) · STAC Research · 2024Free STAC news summary of an AMD-requested test (full report for subscribers): 13.9 ns minimum actionable tick-to-trade network I/O latency on an Exegy/AMD FPGA stack, below the previous record.
  3. How We Verify Custom HardwareTodd Strader (Hudson River Trading) · HRT Beat · 2021FPGAs in trading; time-to-market sensitivity; cocotb, Verilator, CI; verifying risk checks; C++/Python testbenches instead of UVM.
  4. This Startup Is Building a Chip to Save Traders Vital MicrosecondsHooyeon Kim and Whanwoong Choi (Bloomberg) · Data Center Knowledge (syndicated from Bloomberg, free to read) · 2021Kept as a fallback (no open primary source from the company found). Rebellions, a Seoul startup, developing an AI ASIC for high-frequency trading, with TSMC to start making it in 2022.
  5. STAC Report: New LDA/Xilinx solution under STAC-T0 (tick-to-trade network I/O)STAC (Strategic Technology Analysis Center) · STAC Research · 2020Free STAC news summary: STAC-T0 measures network I/O with essentially no trading logic; 24.2 ns minimum actionable latency using LDA TCP and 16-bit MAC/PCS IP cores on a Xilinx FPGA.
  6. Commission Delegated Regulation (EU) 2017/574 (RTS 25) on the level of accuracy of business clocks, as adopted by the EUEuropean Commission · legislation.gov.uk (The National Archives) · 2016Original EU text. Annex Table 2: high-frequency algorithmic trading needs 100 µs maximum divergence from UTC and 1 µs timestamp granularity; Article 4: traceability to UTC, the exact timestamp point, annual review.
  7. IEEE 1588-2019: IEEE Standard for a Precision Clock Synchronization Protocol for Networked Measurement and Control SystemsIEEE Standards Association · IEEE · 2019PTP: sub-microsecond synchronization, sub-nanosecond under optimal network design.
  8. Chip & Platform Approval ProcessEMVCo · EMVCoSecurity evaluation certificates for payment ICs and platforms through recognized laboratories; reports must include vulnerability analysis and penetration testing; site audits.
  9. Cryptographic Module Validation ProgramNIST Computer Security Resource Center · NISTCMVP, run by NIST and the Canadian Centre for Cyber Security, validates cryptographic modules to FIPS 140-3 via accredited testing laboratories.
  10. Certification Report BSI-CC-PP-0084-2014: Security IC Platform Protection Profile with Augmentation PackagesFederal Office for Information Security (BSI) · Common Criteria Portal · 2014Smart-card IC protection profile: EAL4 augmented by AVA_VAN.5 and ALC_DVS.2; threats of malfunction, leakage, manipulation, probing.
  11. Introduction to differential power analysisPaul Kocher, Joshua Jaffe, Benjamin Jun, Pankaj Rohatgi · Journal of Cryptographic Engineering (open access, Springer) · 2011Power measurements leak secret keys; attacks are practical and non-invasive; the DPA difference-of-means procedure; masking.
  12. The Evolution of Bitcoin HardwareMichael Bedford Taylor · IEEE Computer (author copy, UC San Diego) · 2017How mining works; CPU to GPU to FPGA to ASIC miners; unrolled SHA-256 pipelines; early ASIC case histories; efficiency by node; mining economics.