Extras · Industries · Space

Space

Chips that keep working through radiation and temperature extremes, years from the nearest repair.

GR740 total-dose tolerance
300 krad(Si)
GR740: development start to QML-V tests done
2009 → 2020
HPSC compute vs current space processors
Over 100×
U.S. export-control threshold for rad-hard ICs (total dose)
5 × 10³ Gy(Si)

At a glance

Where the flow bends

  1. 01Specification

    The plan lists how much radiation the chip must survive, how hot and cold it gets, and how many years it must run with no repairs.

    The written requirements gain a radiation section: how much radiation the chip must absorb over its life without failing (the total dose, measured in krad), how often a stray particle may flip a stored bit, and immunity to particle-triggered short circuits (latch-up). They also set the temperature range, the mission length in years, and the formal approval the parts must earn, such as QML-V, the U.S. space grade.

    Write the environment as numbers the chip can be tested against: a total-dose target (for example 300 krad(Si)); a latch-up threshold, stated as the LET (energy a particle deposits per unit of path length) up to which no latch-up is allowed; an upset rate per device per day for the intended orbit; the junction temperature range; and the qualification flow (MIL-PRF-38535 class V, or ESCC in Europe). Check the export-control line as well: U.S. rules list ICs rated for 5,000 Gy(Si), which is 500 krad(Si), under ECCN 3A001.a.1.

  2. 02Architecture

    The chip keeps spare copies and memory that checks itself. So one particle hit doesn’t cause a failure.

    Memories store extra check bits so the chip can spot and fix a flipped bit (error detection and correction, or EDAC). A background task re-reads memory and repairs errors before they pile up (scrubbing). A watchdog timer restarts the chip if it stops responding, and some blocks are built in three copies.

    Decide which blocks to harden and which to leave standard; hardening everything costs too much area and power. Pick an error-correcting code per memory (a SECDED Hamming code corrects one flipped bit per word and detects two; Reed-Solomon corrects bursts of errors), set the scrub interval against the expected upset rate, and design recovery from a SEFI, a particle-induced freeze: watchdog, separately resettable regions, power cycling.

  3. 03RTL design

    Important data is kept in three copies. A vote picks the answer that two of them agree on.

    In the code that describes the chip, each important stored bit is kept three times, and a small voting circuit passes on whatever at least two copies say. This is triple modular redundancy (TMR). The code also adds the circuits that write and check the memories’ extra check bits.

    Choose between flip-flop-level TMR (three flip-flops and one voter per bit), full TMR (the voters and logic are tripled too, so a glitch on one shared wire can’t reach all three copies), or a hardened flip-flop cell such as DICE. Self-correcting TMR writes the voted value back into all three copies, so a single upset is repaired instead of sitting there until a second one arrives; this matters where clock gating stops flip-flops from being refreshed.

  4. 04Verification

    Engineers fake particle hits on a computer. Then they fire real particle beams at the finished chips.

    In simulation, engineers deliberately flip bits and inject glitches (fault injection) to check that the copies and check bits catch them. After the chips are made, samples go to particle accelerators, where beams of heavy ions and protons measure how often real upsets happen.

    Inject upsets into flip-flops and short glitches into logic, in the RTL and in the gate-level netlist, including the clock, reset and test logic, not only the data path. Plan beam testing early and give the test setup a view of the chip’s internal state: an upset in test logic has been mistaken for latch-up in a beam test, because both showed up as a jump in supply current.

  5. 05Logic synthesis

    The design is built from special toughened building blocks.

    The tool that turns the code into logic gates draws on a library of radiation-hardened building blocks. It also has to be told to keep all three copies: an ordinary optimizer sees identical logic as waste and merges it.

    Map registers to the hardened library’s DICE or TMR flip-flops, and use cells with transient filters where needed. Mark the triplicated registers and voters so the tool preserves them and doesn’t share logic between copies; a gate shared by all three copies is a single point of failure.

  6. 09Placement

    The three copies are placed apart, so one particle can’t hit all of them.

    When the building blocks are arranged on the chip, the three copies are kept apart. One particle can disturb several neighboring cells at once (a multi-bit upset), so copies placed side by side could all flip together.

    Add placement constraints that separate redundant flip-flops and the paired storage nodes of DICE cells. As transistors shrink, the charge from one strike reaches more neighbors, which is why plain DICE loses effectiveness at advanced nodes.

  7. 12Signoff

    Speed is checked from far below freezing to hotter than boiling water, a wider range than a phone chip sees.

    The final timing check, which confirms every signal arrives before the next tick of the clock, is run across the whole mission temperature range. It uses timing data for the hardened building blocks, which are larger and slower than ordinary ones.

    Run static timing analysis across the qualified junction range (−40 °C to +125 °C for the case-study part), leave margin for aging and for the threshold-voltage drift that radiation dose causes, and include the delay of voters and transient filters on the slowest paths.

  8. 13GDS & tapeout

    After the chip is made, it goes through months or years of testing before it is allowed to fly.

    After manufacture, every part is screened and samples go through qualification: electrical, mechanical, life and radiation tests, for example to the U.S. QML-V space level or Europe’s ESCC system. Qualification can take years.

    Plan for the QML-V qualification groups (A electrical, B mechanical and environmental, C life, D package, E radiation), lot acceptance, and the package work space parts need, such as hermetic ceramic packages and column grid arrays. The case-study part went from validated flight silicon in 2018 to completed QML-V tests in 2020.

Space is full of tiny pieces of atoms, flying much faster than any bullet. They come from the Sun and from deep space.

When one zips through a chip, it can turn a stored 0 into a 1, or cause a quick glitch. Rarely, it can start a short circuit that burns the chip out.

Over the years, all those hits also wear the chip down slowly. And once a spacecraft launches, nobody can open it up to swap a part.

So space chips are built to take hits and keep going. They keep important data in three copies and go with the majority. They check their memory all the time. And before they fly, engineers fire real particle beams at them to test them.

A chip is built from billions of transistors, tiny electrical switches, and it stores information as 0s and 1s in memory cells and in , small circuits that each hold one bit. Radiation in space harms these in two ways.

The first is slow. Over the years a chip absorbs a growing (TID), measured in krad, a unit of absorbed radiation energy. This causes “parametric or functional degradation”: the transistors gradually leak more and switch at the wrong voltages until the chip stops meeting its specification.

The second is sudden. A happens “when a single radiation particle strike deposits enough charge to cause an effect.” The main kinds:

  • : a stored bit flips. Nothing is damaged, but the data is wrong.
  • : a brief false pulse in the chip’s logic. It becomes an error only if a flip-flop stores it.
  • : the strike switches on an unintended path through the chip’s silicon, causing a high current that can destroy it.
  • : the chip freezes or misbehaves, a functional “hiccup.” A soft one clears with a restart; a hard one needs the power switched off and on.

There are two broad ways to make a chip tougher, called hardening. One is to build it in a factory process that is itself tolerant of radiation. The other is : keep an ordinary commercial process but change “the transistors, gates, and physical layouts inside the silicon,” using hardened building blocks and radiation-aware layout.

The two kinds of damage have different physics. Total dose comes from charge that builds up in the thin insulating oxide under each transistor’s gate; it raises leakage current and shifts the threshold voltage at which the transistor turns on, so the chip slowly degrades. Single-event effects come from one energetic particle (a proton, neutron or heavy ion) leaving a trail of charge through a transistor that is switched off. Some are “soft” and lose data; some are “hard” and can cause permanent failure.

Smaller transistors changed the balance. Their very thin gate oxides make modern processes more tolerant of total dose: many standard CMOS processes withstand up to 300 krad, while many space missions see below 100 krad. Latch-up was a serious problem in older processes, and silicon-on-insulator processes, which sit the transistors on an insulating layer, are inherently immune to it. Once a process with enough tolerance to both has been chosen, upsets and transients remain the main reliability problem.

The design response comes in layers:

  1. Cells. Use hardened such as , and layout measures: guard rings (rings of doped silicon that collect stray charge), arrays of well contacts, and extra spacing between critical transistors.
  2. Logic. Triple important registers with and protect memories with codes.
  3. System. Add , watchdog timers, selective power cycling and software-based fault isolation.

Every layer costs area and power, so high-performance processors increasingly harden only the most vulnerable blocks and add architectural redundancy only where needed.

Reading the units. A gray, Gy(Si), is one joule of radiation energy absorbed per kilogram of silicon; 5×1035 \times 10^3 Gy(Si) equals 5×1055 \times 10^5 rad(Si), or 500 krad(Si). That number matters legally. U.S. export rules list ICs designed or rated to withstand a total dose of 5×1035 \times 10^3 Gy(Si) or more, a dose-rate upset of 5×1065 \times 10^6 Gy(Si)/s, or a neutron fluence of 5×10135 \times 10^{13} n/cm² under ECCN 3A001.a.1. So a total-dose target written into the spec can decide where the finished part may be sold.

flip-flops / memory01101001controlOKFFclksubstrateI
Effect

Choose an effect, then fire one particle.

Four single-event effects from one particle strike: stored bits, a logic wire into a flip-flop, the control block and the supply current. Schematic.Share freely with credit: ‘Figure from chipfieldguide.com’
0100200300400500600total dose (krad)leakagespec limitmany missionsrating (e.g. GR740)

80 krad absorbed. Many missions stay below 100 krad over their life; leakage has barely moved.

Total ionizing dose: as krad accumulate, transistors leak more until the chip leaves its spec. Markers from the page; the curve’s shape is illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
  • Radiation, which flips bits and slowly wears chips out.
  • Heat and cold: a spacecraft bakes in sunlight and freezes in shadow.
  • No repairs: the chip must last the whole mission, often many years.
  • Cost versus risk: a big science mission pays for proven, toughened chips. A small, cheap satellite may use ordinary chips, like those in phones, and accept more risk.

Ordinary parts or hardened ones. (commercial off-the-shelf) parts, the kind made for phones and laptops, “offer superior performance, energy efficiency, and affordability,” but “tend to be highly susceptible to radiation.” Small satellites have traditionally used COTS processors and memory, surrounded by radiation-hardened supporting parts: error-correcting memory, watchdog timers that restart a stuck processor, scrubbing and spare copies. For harsher environments or longer missions, a part’s flight history and measured radiation performance become the main reasons to choose it.

Formal approval. Parts for military and space use go through a long approval process called qualification. In the U.S., the standard MIL-PRF-38535 has offered hermetic (airtight, sealed) parts in class Q for military use and class V for space, under a . For space parts, the Defense Logistics Agency, NASA’s Jet Propulsion Laboratory and the Aerospace Corporation act as the approving body. In Europe, the European Space Components Coordination (ESCC) runs the specifications and its own Qualified Parts List, with ESA as the qualification authority. NASA’s NEPP program publishes how parts perform, how they fail and how to test them.

More computing on board. Messages to and from Earth take time to arrive, so spacecraft beyond Earth orbit must run autonomy, AI, image processing and fault recovery on board, without waiting for ground control, and NASA notes that “radiation and extreme temperatures can affect electronic components.”

FPGAs. An is a chip whose logic is configured after manufacture. A NASA Goddard report notes that FPGAs are used in every space application. Radiation-hardened FPGAs come with hardened elements built in, at high cost; commercial reprogrammable FPGAs are cheaper but must be protected by the designer. The report compared two ways of applying in the design loaded into one:

  • Localized TMR triples only the flip-flops.
  • Distributed TMR triples everything except the global wires that carry clocks, resets and enables.

In a cyclotron beam test, distributed TMR reduced single-event errors more than localized TMR. Tripling the logic and voters too means a glitch in shared logic can’t corrupt all three copies, but it costs more power and area. Errors on the untripled global wires remain, and tripling those as well (global TMR) brings timing difficulties and still more overhead.

How a beam test is read. Particles are described by their LET (linear energy transfer), the energy they deposit per unit of path length. A test counts errors at one LET and divides by the fluence, the number of particles that crossed each square centimeter; the result is the cross-section, σ\sigma, in cm². For example (illustrative numbers): 20 errors after a fluence of 10710^7 ions/cm² gives σ=20/107=2×10−6 cm2\sigma = 20 / 10^7 = 2 \times 10^{-6}\,\mathrm{cm^2}. Repeating this at several LETs gives a curve, which is combined with the particle population of the intended orbit to predict errors per device per day.

The data needs care. LaBel and Berg describe an FPGA beam test in which supply current stepped up and the device configuration changed. That pattern is usually read as latch-up, but the cause was an upset in the JTAG test access port controller, a block of built-in test logic. Their conclusion: the line between a SEFI and latch-up “can be very blurry,” and classifying an event needs both an understanding of the device and as much visibility as the test setup can give. Test logic is part of the area a particle can hit, and the test setup should be able to see what state the chip is in.

A1011B1011C1011voter2 / 3output1011correct

Three copies feed a majority voter. Hit a copy to flip one of its bits.

Triple modular redundancy: three copies of a register and a majority voter. Scrubbing rewrites every copy from the vote.Share freely with credit: ‘Figure from chipfieldguide.com’

First, the plan says how much radiation, heat and cold the chip must survive.

Next, designers add memory that checks itself and three copies of important parts. They spread the copies apart on the chip, so one particle can’t hit them all.

Then they fake particle hits on a computer, to make sure the protections work.

Finally, real chips go to a particle accelerator, a machine that fires beams of fast particles. After that come long tests before the chip is allowed to fly.

A chip is designed in stages: a written specification, an architecture (the major blocks and how they connect), code that describes the circuit, conversion of that code into logic gates, arrangement of the gates on the silicon, and a series of checks before the design is sent to the factory. Space changes most of them:

StageWhat space adds
SpecificationTotal dose, latch-up immunity, acceptable upset rate, temperature range, mission life, approval class
ArchitectureSelf-correcting memory, scrubbing, watchdog timers, spare copies
Code and gatesThree copies of important bits, or hardened flip-flops; the tools must keep every copy
LayoutCopies placed apart so one particle can’t hit them all
CheckingSimulated particle hits; real beam tests once chips exist
Timing checkThe full mission temperature range
After manufactureScreening and qualification (QML, ESCC)

The simplest form of triple modular redundancy replaces one flip-flop with three and a voting circuit, which hides an error in any one of the three. One of the most widely used hardened flip-flops is the cell, which instead keeps two linked copies of the bit inside a single cell.

Choosing a hardened flip-flop. Each option fails in a different way, so it helps to walk through them:

  1. DICE keeps two interlocked copies inside one cell. It gives very good hardness in older processes but is less effective at highly scaled ones (for example 45 nm), and the standard version is not robust to multi-bit upsets. Variants such as T-DICE, F-DICE and LEAP-DICE add transistors or layout rules to fix that.
  2. Basic TMR uses three ordinary flip-flops and a voter. It masks an upset in any one flip-flop, but if a glitch on the data path arrives exactly at the clock edge, all three flip-flops capture it together and the vote is wrong.
  3. Full TMR also triples the voter and the logic in front of it. It handles the glitch case but costs a large increase in silicon area and complicates the scan chains used for manufacturing test.
  4. Filtered TMR puts a filter at each flip-flop input that swallows glitches shorter than a chosen width. The chosen pulse width directly sets the delay the filter adds.

Schrape et al. built five TMR variants (baseline, latch-based, TSPC, scannable and self-correcting) from a commercial non-hardened cell library. They handled multi-bit upsets with placement constraints and glitches with custom filters and carefully sized clock and reset buffers. Their test chips tolerated particles with LETs from 32.4 to 62.5 MeV·cm²/mg before errors appeared, depending on the variant. The self-correcting variant matters when clock gating (turning a block’s clock off to save power) is used: a gated flip-flop isn’t refreshed, so upsets can accumulate until a second one defeats the vote.

Keeping the copies through the flow. Synthesis and physical optimization see three identical registers as waste and will merge them or share logic between them unless told otherwise, so the flow marks them to be preserved. Placement rules keep redundant cells apart, and timing signoff includes the delay of voters and filters on the slowest paths.

Package and qualification. Space parts often use hermetic ceramic packages and column grid arrays (the chip package stands on rows of small solder columns rather than balls), which bring their own bonding and board-reliability qualification, as the case study shows.

placed cellsABCcopies upset: 3/3 · vote fails
Copy placement

A multi-cell upset reached 3 copies at once: the majority is wrong. Placement rules must keep copies apart.

A row of cells holding three TMR copies. One strike can upset neighbouring cells; compare copies placed together and spread apart. Strike size illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Europe’s space agency, ESA, wanted a faster brain for its spacecraft. The result is the GR740, a chip with four processors built to handle radiation.

Work started in 2009, and the full space tests finished in 2020. That is more than ten years from start to a chip that can fly.

NASA is now building a new space processor called HPSC. It aims to be more than 100 times as powerful as today’s space processors.

The GR740 is a processor for spacecraft, developed for ESA. Its main features, in plain terms:

  • Four processor cores (LEON4, using the SPARC V8 instruction set), each built to tolerate faults, running at 250 MHz, that is 250 million clock ticks a second.
  • A 2 MiB on-chip cache, a fast local memory that holds recently used data.
  • Error correction (Reed-Solomon ) on the connection to external memory, and EDAC on the connection to its startup memory.
  • A router for SpaceWire, a data-network standard used on spacecraft.
  • Built in a 65 nm process for space chips, with a dedicated library of space building blocks.

The goal was a European alternative to the RAD750 processor, with more performance than earlier European space processors. Its published results:

  • Tolerates a total dose of 300 krad(Si).
  • No latch-up up to an LET of 125 MeV·cm²/mg (LET measures how much energy a particle deposits along its path), tested above 85 °C at maximum supply voltage.
  • Fewer than one single-event error per 100,000 days per chip in geostationary orbit (below 1×10−51 \times 10^{-5} events per device per day).
  • Works from −40 °C to +125 °C at the chip itself, in ceramic packages, using under 2 W at 40 °C.

Where the time went. The timeline shows where space projects spend their years:

  • 2009: design in VHDL (a hardware description language) and verification by simulation and on FPGAs.
  • 2014: implementation in the 65 nm space process.
  • 2016: engineering models evaluated.
  • 2018: flight silicon made and validated, including radiation.
  • 2020: all QML-V qualification tests complete.
  • Early 2021: complementary ESCC tests (a “delta” evaluation) complete.

Qualification followed the QML-V flow, in groups: A (electrical), B (mechanical and environmental), C (life), D (package) and E (radiation).

Physical hurdles. Many of the problems were not in the logic. Flip-chip packaging, in which the die is soldered face-down onto the package, was not available to the project, so the team developed 625-pin ceramic land grid and column grid array packages with wire bonds instead. The space cell library required gold ball bonding, and the chip’s complexity required gold wires only 20 µm in diameter, stacked on four bond decks. Because gold wire meets aluminum pads on the die, an extra over-pad metallisation layer was added at the interface.

A second die revision. Revision 1 fixed the level-2 cache’s fault tolerance, extended error logging, corrected a ring oscillator and the on-chip temperature sensor, and added pipelining (extra register stages) to ease the back-end layout work. The team’s lessons: release prototypes early, for functional validation and radiation characterization, and keep software compatibility with earlier LEON parts.

The next generation. NASA’s HPSC is a system-on-chip built on an open-source instruction set, with fault tolerance, radiation tolerance, a security suite, fine-grained power control and Ethernet connectivity, aimed at mission needs through 2040 and beyond.

2009VHDL design201465 nm layout2016eng. models2018flight silicon2020QML-V2021ESCC deltayears from design start0
1 / 6

2009: design in VHDL, verified by simulation and on FPGAs.

The GR740’s path from design start to a qualified flight part, from ESA’s final report.Share freely with credit: ‘Figure from chipfieldguide.com’

Sources

Show Hide 11 sources
  1. 8.0 Small Spacecraft Avionics (State-of-the-Art of Small Spacecraft Technology)NASA Small Spacecraft Systems Virtual Institute · NASATID, SEE, SEU and SEL definitions; COTS trade-offs and the COTS-first pattern; RHBD and selective hardening; watchdogs and power cycling.
  2. NASA Electronic Parts and Packaging (NEPP) ProgramNASAWhat NEPP does; home of GSFC radiation test data tools.
  3. Complex Parts, Complex Data: Why You Need to Understand What Radiation Single Event Testing Data Does and Doesn’t Show and the Implications ThereofKenneth A. LaBel and Melanie D. Berg · NASA Technical Reports Server (GOMACTech 2015) · 2015SEFI definitions (JEDEC and Swift); an FPGA beam test where a JTAG TAP controller upset looked like latch-up.
  4. Design and Evaluation of Radiation-Hardened Standard Cell Flip-FlopsOliver Schrape, Marko Andjelkovic, Anselm Breitenreiter, Steffen Zeidler, Alexey Balashov and Miloš Krstić · IEEE Transactions on Circuits and Systems I (open access via TIB) · 2021TID mechanism; SEL and SOI; SET definition; DICE limits at scaled nodes; TMR flip-flop variants; MBU placement constraints; SET filters; selective hardening.
  5. Localized Triple Modular Redundancy vs. Distributed Triple Modular Redundancy on a ProASIC3E Reprogrammable FPGAAlex McGuffey, Melanie Berg and Jonathan Pellish · NASA Technical Reports Server (NASA USRP internship final report) · 2010FPGAs in space; RH vs COTS FPGAs; LET and cross-section; DTMR beat LTMR in a cyclotron test.
  6. MIL-PRF-38535 Standard Microcircuits: Hermetic and Non-hermeticShri G. Agarwal (NASA JPL) · NASA Electronic Parts and Packaging Program, Electronic Technology Workshop · 2021Hermetic classes Q and V (class levels B and S); the qualifying activity for space microcircuits.
  7. European Space Components CoordinationEuropean Space AgencyESCC specification system, Qualified Parts List, ESA as qualification authority.
  8. GR740 next-generation microprocessorEuropean Space Agency · 2016Quad LEON4, LEON lineage from ESA’s LEON2-FT.
  9. GR740 Next Generation Microprocessor Flight Models (TEC-ED & TEC-SW Final Presentation Day)Cobham Gaisler, for ESA · ESA (nebula.esa.int study deliverable) · 2021GR740 timeline, features, EDAC, radiation and qualification results, package work, design changes, lessons learned.
  10. High Performance Spaceflight Computing (HPSC)NASA Game Changing Development Program · NASAWhy space computing is hard; HPSC capabilities, over 100× compute; project status (tape-out mid-2025, testing as of March 2026).
  11. 15 CFR Part 774, Supplement No. 1: The Commerce Control List (ECCN 3A001.a.1)Electronic Code of Federal Regulations (eCFR)Radiation-hardened IC thresholds for export control; Gy(Si) to rad(Si) conversion.