Extras · Industries · Automotive

Automotive

Chips that must detect their own faults, survive under the hood and ship with near-zero defects.

AEC-Q100 Grade 0 ambient range
−40 °C to +150 °C
ASIL D single-point fault metric target
≥ 99%
AEC-Q100 HTOL sample
77 parts × 3 lots, 0 fails
Tesla FSD chip (2019)
6 billion transistors, 260 mm²

At a glance

Where the flow bends

  1. 01Specification

    The plan says how much harm a failure could cause, how hot the chip will get, and how it must keep hackers out.

    The specification adds safety goals, each with an ASIL rating (A to D) from an analysis of what could go wrong in the car, a temperature grade, how many years and hours the chip must work, numeric targets for catching faults, and security goals from an analysis of likely attacks.

    Chip-level safety requirements are derived from the vehicle function’s safety goals: target SPFM, LFM and PMHF for the ASIL, the time allowed between a fault and reaching a safe state, the safe states themselves, and assumptions of use if the SoC is developed before its final application is known. ISO/SAE 21434 cybersecurity goals come from the TARA.

  2. 02Architecture

    The chip gets built-in guards. Two processors check each other’s work, memory fixes its own small errors, and a timer notices if the software freezes.

    The architecture adds safety mechanisms: pairs of processors that run the same program and compare results, memory that detects and corrects its own errors, watchdog timers that notice when software stops responding, and often a separate “safety island” that supervises the big compute blocks.

    ASIL-D control is partitioned onto lockstep cores and a safety island with independent clock and power; diagnostic coverage is decided per block; and the hardware is partitioned so that less-critical functions cannot interfere with safety-critical ones.

  3. 03RTL design

    Engineers write the checking circuits right alongside the normal ones.

    The design code includes the checking circuits: error-correcting encoders and checkers for memories, comparators for paired processors, wiring that reports every error to a central fault-handling block, and hooks that let the chip test itself.

    Every safety mechanism needs an error output that reaches a fault-handling path, plus a way to test the mechanism itself, since an untested checker that has failed silently is a latent fault.

  4. 04Verification

    Engineers break the design on purpose, thousands of times, in a computer copy. Each time, they check that the guards catch the problem.

    Fault-injection campaigns simulate the design thousands of times, each with one deliberate fault (a wire stuck at 0 or 1, a stored bit flipped), and classify each fault as harmless, caught or dangerous. The results feed the chip’s safety metrics.

    Simulation-based fault campaigns are run per safety mechanism, formal tools prove unobservable faults safe, and the classified results roll into the FMEDA to compute SPFM and LFM against the ASIL targets.

  5. 06Design for test

    Test circuits are built in so the factory can catch nearly every bad chip. The chip can also test itself later, in the car.

    Very thorough factory testing, aimed at near-zero defects, depends on extra test circuitry that can reach almost every part of the chip. Built-in self-test circuits for logic and memory also let the chip test itself in the car, to find faults that appear over its life.

    Scan coverage targets for stuck-at and transition faults are higher than for consumer parts, with at-speed patterns, outlier screening (Part Average Testing), and in-field LBIST and MBIST whose run time must fit the diagnostic test interval and power budget.

  6. 12Signoff

    The final checks allow for more heat and more years of use than for a phone chip.

    The final checks of timing and reliability are run up to the chip’s temperature grade (as hot as 150 °C for Grade 0), and the expected years of use set how much margin is left for wires and transistors to wear.

    Electromigration, IR drop and aging are signed off against the automotive mission profile (the expected temperature and usage over the vehicle’s life), and the safety mechanisms must meet timing in every corner, since a checker that misses timing is itself a dangerous fault.

  7. 13GDS & tapeout

    After the chips are made, samples go through months of tough tests, like six weeks of running in a hot oven. Carmakers use them only if they pass.

    After manufacturing, parts are qualified to AEC-Q100, including 1,000 hours of powered operation at high temperature on samples from three manufacturing lots, with zero failures allowed.

    Teams plan the AEC-Q100 test groups (environmental, lifetime, package, electrical verification), zero-defect practices from AEC-Q004 in production, and a safety case with FMEDA and fault-injection evidence for the customer’s assessment.

A chip in a car may help steer, brake or even drive. If it fails at the wrong moment, people can get hurt. So car chips must catch their own faults, survive years of heat and shaking, and almost never leave the factory broken.

The rules for that first part are called . Each possible danger gets an , a safety grade from A to D. The grade depends on how bad the harm would be, how often the situation comes up, and whether a driver could still control the car. The higher the grade, the more checking the chip needs.

Every chip will eventually suffer random faults: a transistor wears out, a stray particle flips a stored bit. In a phone that means a crash and a restart. In a brake or steering controller it could hurt someone. is the discipline of making sure such faults are caught and handled before they cause harm, and for cars it is defined by the standard ISO 26262. Its 2018 second edition covers all road vehicles except mopeds and added Part 11, guidance on applying the standard to chips.

It starts with the car, not the chip. Engineers list each hazard (say, unintended braking) and rate it for severity (how badly people could be hurt), exposure (how often the driving situation occurs) and controllability (whether a typical driver could still cope). A table combines the three ratings into an , from A (least strict) to D (most strict), for each safety goal.

For a chip, the ASIL turns into three numeric targets:

MetricWhat it measuresASIL BASIL CASIL D
Share of faults that are harmless alone or caught by a checker≥90%\ge 90\%≥97%\ge 97\%≥99%\ge 99\%
Share of hidden faults that are eventually found≥60%\ge 60\%≥80%\ge 80\%≥90%\ge 90\%
Rate of dangerous failures, in failures per billion hoursunder 100under 100under 10

Under 10 failures per billion hours means, on average, fewer than one dangerous failure in 100 million hours of driving. Separately, chips sold into cars are qualified with a set of stress tests called , which show they survive heat, cold and age.

For a chip team, functional safety produces three deliverables. First, : error-correcting codes on memories, lockstep comparators, watchdogs, each with a claimed diagnostic coverage (the share of the faults in its block that it catches). Second, an , a table that lists each block’s failure modes and failure rates, applies the coverage of its mechanisms, and sums the result into the SPFM, LFM and PMHF figures the ASIL demands. Third, fault-injection evidence that the coverage claims hold. An analysis of RISC-V automotive safety argues that the dominant costs are these engineering activities: FMEDA generation, diagnostic coverage analysis, tool qualification, safety-case documentation and fault-injection campaigns.

A small, illustrative FMEDA row shows how the numbers work. Suppose a memory block has a failure rate of 100 FIT (failures per 10910^9 hours) and every one of its failures could violate a safety goal. Error-correcting code with 99% diagnostic coverage leaves 1 FIT of residual faults that escape it. If that block were the whole chip, the SPFM would be 1−1/100=99%1 - 1/100 = 99\%, just meeting ASIL D. Every block’s row adds to the totals, so a single block with weak coverage can sink the whole chip’s metric.

ISO 26262 covers malfunction. Two other standards sit beside it: ISO 21448 (SOTIF, safety of the intended functionality) addresses perception and decision failures outside the random-hardware-fault model, which matters for autonomy, and ISO/SAE 21434, released in August 2021, sets cybersecurity engineering and risk-analysis principles.

memory block100 FITECC99.0%detected 99.0escapes 1.0FIT = failures / 10⁹ hSPFM85%≥99residual rate (PMHF)0.1100 FIT10✓ ASIL D
ASIL

Coverage 99.0%: SPFM 99.0% (needs ≥ 99%), 1.0 FIT escape (needs < 10). Meets ASIL D. LFM must also reach 90%.

One illustrative FMEDA row: a 100 FIT memory block protected by ECC, treated as the whole chip. Targets from the page’s table.Share freely with credit: ‘Figure from chipfieldguide.com’
  • Safety: the chip must catch its own faults.
  • Heat and cold: next to an engine, a chip may need to work from −40 °C, colder than a freezer, up to 150 °C, hotter than boiling water.
  • Zero defects: carmakers count bad chips per million shipped, and the goal is none.
  • Hackers: cars now connect to the internet, so their chips must resist attack.
  • Big brains: self-driving needs the power of a huge computer chip and the reliability of a brake.

Temperature. AEC-Q100 sorts chips into four grades by the air temperature around them: Grade 0, −40 °C to +150 °C (next to the engine); Grade 1, −40 °C to +125 °C; Grade 2, −40 °C to +105 °C; Grade 3, −40 °C to +85 °C. One of its tests, high-temperature operating life (HTOL), runs chips powered for 1,000 hours at the grade’s maximum temperature: 77 chips from each of three separate manufacturing lots, with zero failures allowed.

Zero defects. Car makers track , and the goal is zero. The automotive industry’s zero-defects framework is a menu of practices across process design, product design, production and improvement. For chip designers, its main ask is design for test: extra circuitry so that “as many nodes as possible” can be tested in a reasonable time on the factory tester. That includes scan testing for wires stuck at 0 or 1 and for transitions that are too slow, and built-in self-test, where the chip tests itself. In production, Part Average Testing removes chips whose measurements are statistical outliers from the population before they ship.

Cybersecurity. A connected car can be attacked remotely. The car-cybersecurity standard ISO/SAE 21434 sets high-level principles for threat analysis and risk assessment () but does not prescribe specific use cases. For a chip, the results become requirements such as secure boot (only signed software can start), protected key storage, and locking down the debug port.

Autonomy compute. Large multicore processors and GPU-style accelerators are increasingly the only way to reach the performance autonomous cars need, but the safety support they offer is uneven. A complements them with watchdogs, monitoring, test orchestration and diverse redundant execution.

Lockstep. ASIL-D components are generally deployed on dual-core CPUs: the same software runs on two identical cores, one a few cycles behind the other. A disturbance that hits both at once, such as a supply glitch, catches them at different points in the computation, so their outputs differ and a comparator flags the error. Watchdogs, which expect a regular “still alive” signal, should be independent of what they monitor, with their own clock and power supply where possible, or a single failure could stop both.

12345678cycleCore ACore BCompare3172235445861768315223944586delay=≠error!=≠error!==
Checker core delay
Fault

Mismatch on step 2 (cycle 4): the comparator flags the fault before the wrong result is used.

Dual-core lockstep, cycle by cycle (values illustrative). The checker core runs a set number of cycles behind; a comparator checks each step.Share freely with credit: ‘Figure from chipfieldguide.com’

Safety shows up at almost every step. The design gets checkers, like twin processors that compare answers.

Then engineers break the design on purpose, thousands of times, in a computer copy. Each time, they check that a checker catches the problem.

Last, sample chips spend months in tough tests before a carmaker says yes.

Designing a chip runs through a fixed sequence of steps, from the written specification to the file sent to the factory. For a car chip, almost every step gains a safety task:

StepWhat it normally doesWhat automotive adds
SpecificationSays what the chip must doSafety goals and ASIL, temperature grade, years of use, security goals
ArchitectureChooses the main blocksTwin processors in lockstep, error-correcting memory, watchdogs, a safety island
Design code (RTL)Describes the circuit in a hardware languageCheckers whose error signals all lead to one fault-handling block
VerificationChecks the design worksFault injection, to measure how many faults the checkers catch
Design for testAdds circuits for the factory testerVery high test coverage, plus self-tests the chip runs in the car
SignoffFinal checks of timing and reliabilityChecks up to the temperature grade, over the car’s whole lifetime
TapeoutSends the design to the factoryAEC-Q100 qualification and a written safety case

The distinctive new step is , which ISO 26262 highly recommends during chip development to judge how well the design handles random hardware failures. Engineers simulate the design thousands of times, each time with one deliberate fault, such as a wire stuck at 0 or a stored bit flipped. For each run they ask two questions: did the fault change the chip’s output, and did a checker raise an error? A fault that changes the output with no error raised is dangerous; the others are counted as safe or caught.

Fault classification. A fault campaign defines observation points (the functional outputs) and detection points (the safety mechanism’s error signals), then sorts every injected fault into four bins:

  • Observed and detected: the mechanism did its job.
  • Not observed but detected: harmless, and a diagnosis fired anyway.
  • Observed but not detected: a dangerous, undetected fault that counts against SPFM.
  • Neither: harmless now, but a potential latent fault that needs another diagnostic, such as .

Simulation-based campaigns aren’t exhaustive, because a fault may simply not be exercised by the tests, so unobserved faults traditionally need manual analysis. Formal tools can prove many of them untestable by design. In one published campaign on an error-correcting (SECDED) memory block, formal analysis of 1,726 faults found 512 that corrupted output data without raising an error.

Campaign cost. Campaigns for transient faults (bit flips that last one cycle) grow with design size times simulation length, since each fault can be injected at any cycle. Pruning the fault list helps: dynamic HDL slicing, which skips injection times when the affected register’s value isn’t used, cut injections by up to 10% on an industrial core.

DFT doubles as safety. The same scan and built-in self-test hardware that drives DPPM down at the factory provides diagnostics in the car. The zero-defects framework lists what BIST costs: chip area, supply requirements and test time, weighed against the faults it covers. In the field, LBIST overwrites the logic’s working state, so it runs in windows when the function can be paused, and the run must fit the diagnostic interval set by the safety concept.

inputsdesign blockoutputsafety checkerror flagObserved + detected0Detected only0Observed, missed0Neither (latent?)0

A block with an output (observation point) and a checker (detection point). Inject faults one at a time.

A fault-injection campaign on an illustrative block: each fault lands in one of four bins.Share freely with credit: ‘Figure from chipfieldguide.com’

In 2019, Tesla showed its own self-driving chip at a big chip conference. Each car’s computer holds two of these chips and two power supplies, so one can back up the other.

Tesla says the design took just 14 months. To move fast, it used proven parts from other companies for most of the chip. The part Tesla built itself does the heavy math to make sense of what the cameras see.

In 2019 Tesla presented its self-driving chip at Hot Chips, an open engineering conference. The published figures:

  • Made in a 14 nm process, 260 mm² of silicon with 6 billion transistors, qualified to AEC-Q100.
  • 12 general-purpose processor cores, a graphics processor, an image processor for the cameras, a video encoder, safety and security blocks, and two neural-network accelerators that Tesla designed itself.
  • Each accelerator is a 96 × 96 grid of multiply-and-add units (9,216 of them) running at over 2 GHz, with 32 MB of on-chip memory.
  • Goals: over 50 trillion operations per second, under 40 W per chip, and under 100 W for the whole computer.

Safety came from the board as much as the chip: the computer carries two of these chips, redundant power supplies, and cameras whose views overlap, with redundant paths. The neural-network engines themselves are covered as a design in the Architectures guide.

Two design choices stand out. First, redundancy is at system level: Tesla listed “lower part costs to enable redundancy architectures” as a platform goal and designed the chip to be “modular to enable various platform redundancy uses.” This matches the broader pattern for autonomy, where high-performance compute is paired with separate safety supervision because the safety support built into such compute is uneven.

Second, schedule risk was cut by limiting custom work. Standard functions (CPUs, GPU, image processor, video encoder, memory controller, PHYs, on-chip interconnect) came from proven IP, and the team chose simpler clock and power distribution. The custom accelerator uses a single clock domain, state-machine control and DVFS-enabled power and clock distribution, and keeps programs resident in SRAM to avoid DRAM traffic. Choices like these remove work from later flow stages: a single clock domain, for example, means no signals inside the block cross between unrelated clocks, a classic source of hard-to-verify bugs.

What the slides do not include is equally typical: no ASIL rating, FMEDA or fault-injection results are published. Automotive safety cases go to customers and assessors, so public case studies show the architecture and keep the safety evidence private.

computer boardcamerasAFSD chipPSUBFSD chipPSUone chip · block diagram12 CPUsGPUNNA 96×9632 MBNNA 96×9632 MBISPvideosafety
Failure

Two chips, two power supplies, overlapping cameras with redundant paths. Tap a block of the chip.

Tesla’s 2019 self-driving computer as presented at Hot Chips. Board and chip drawn as block diagrams, not layouts; three cameras stand for many.Share freely with credit: ‘Figure from chipfieldguide.com’

Sources

Show Hide 10 sources
  1. Assessment of Safety Standards for Automotive Electronic Control Systems (DOT HS 812 285)Qi D. Van Eikema Hommes · National Highway Traffic Safety Administration · 2016Section 3.6: ISO 26262 assesses an ASIL for each safety goal from a tabular combination of severity, exposure and controllability.
  2. ISO 26262WikipediaThe 2018 second edition extended the scope to all road vehicles except mopeds and added Part 11 on semiconductors; ASIL D is the highest level.
  3. Formal Assisted Fault Campaign for ISO26262 CertificationNitin Ahuja, Mayank Agarwal and Sandeep Jana · DVCon Europe proceedings · 2019SPFM/LFM/PMHF targets by ASIL; observed/detected fault classification; formal analysis of a SECDED ECC block finding 512 of 1,726 faults observed but not detected.
  4. Accelerating Transient Fault Injection Campaigns by using Dynamic HDL SlicingAhmet Cagri Bagbaba, Maksim Jenihhin, Jaan Raik and Christian Sauer · arXiv · 2020ISO 26262 highly recommends fault injection during IC development; dynamic HDL slicing cuts injections by up to 10% on an industrial core.
  5. AEC-Q100 Rev-J1: Failure Mechanism Based Stress Test Qualification for Integrated Circuits in Automotive ApplicationsComponent Technical Committee · Automotive Electronics Council · 2026Table 1 temperature grades 0–3; HTOL (test B1) for 1,000 hours at the grade’s maximum ambient on 77 parts × 3 lots with 0 fails.
  6. AEC-Q004: Automotive Zero Defects FrameworkComponent Technical Committee · Automotive Electronics Council · 2020Zero-defect practices across process design, product design, production and improvement: BIST and its trade-offs, DfT reaching as many nodes as possible (scan stuck-at and transition coverage), Part Average Testing to remove outliers.
  7. Envisioning a Safety Island to Enable HPC Devices in Safety-Critical DomainsJaume Abella, Francisco J. Cazorla, Sergi Alcaide, Michael Paulitsch, Yang Peng and Inês Pinto Gouveia · arXiv · 2023HPC devices increasingly the only way to reach autonomy performance, with heterogeneous safety support; dual-core lockstep with staggering for ASIL D; watchdogs independent in clock and power; safety islands.
  8. Security Risk Analysis Methodologies for Automotive SystemsMohamed Abouelnaga and Christine Jakobs · arXiv · 2023ISO/SAE 21434 released in August 2021; gives high-level risk-analysis steps but does not specify use cases.
  9. Compute and Redundancy Solution for the Full Self-Driving Computer (Hot Chips 31 slides)Pete Bannon, Ganesh Venkataramanan, Debjit Das Sarma, Emil Talpes, Bill McGee and team (Tesla) · Hot Chips 31 · 2019FSD computer with dual redundant SoCs and power supplies; chip goals (over 50 TOPS, under 40 W); 14 nm, 260 mm², 6 billion transistors, AEC-Q100; proven IP for standard functions; 96×96 MAC NNAs with 32 MB SRAM at 2 GHz+; 14 months architecture to tapeout; single clock domain.
  10. RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted CertificationNick Andreasyan, Mikhail Struve, Alexey Popov, Maksim Nikolaev and Vadim Vashkelis · arXiv · 2026Dominant certification costs are engineering activities (FMEDA, diagnostic coverage analysis, tool qualification, safety-case documentation, fault injection); ISO 21448 SOTIF covers perception and decision failures outside the random-hardware-fault model.