Extras

After tapeout

Sending the design to the factory is only halfway to a working chip. This page follows the file as it becomes glass stencils, then chips on a silicon disc, then tested and packaged chips that ship.

At tapeout the design team sends the factory one file: an exact drawing of every layer of the chip. Months later the team holds finished, packaged chips. In between come stencil making, a few hundred manufacturing steps, testing and packaging, then the team’s own work: switching the first chips on, finding and fixing problems, and proving the chip will last.

After tapeout, yield, test, debug and reliability decide whether a design becomes a product. This page walks through each: what limits lithography, how wafer sort and known-good die work, how a team brings up and characterizes first silicon, when a fix can be made in the wiring alone and when it needs a full respin, and the statistics behind qualification.

At tapeout the chip is still only a drawing. Turning it into something you can hold takes months and several different kinds of factories.

First, each layer of the drawing becomes a glass stencil. A chip factory, called a fab, uses the stencils to build the chip on a , a thin disc of pure silicon about the size of a dinner plate. Then every chip is tested, cut out and sealed in a protective case.

When the first chips come back, engineers check that they work hot and cold, fast and slow. If they find a bug, they choose between a cheap fix and making the chip again. Only after stress tests show the chip will last for years does the factory start making it by the millions.

A chip design ends as one file, usually in a format called GDSII: an exact drawing of every shape on every layer of the chip. Sending that file to the factory is called tapeout, and the GDS and tapeout chapter covers how it is prepared. The factory, called a fab or foundry, builds many copies of the chip side by side on a thin, round slice of silicon called a wafer. Each copy is a die; after testing, the wafer is cut apart and each good die is sealed in a package with pins or solder balls.

This page follows the design from that handoff to a product, in seven steps:

  • Making the masks: a stencil for each layer.
  • Processing the wafer in the fab.
  • Testing every die while it is still on the wafer.
  • Packaging the good dies and testing them again.
  • Switching on in the lab and measuring it.
  • Fixing what is wrong.
  • Proving the chip will last, then ramping up production.

The fab’s work has two halves. First it builds the transistors, the tiny switches, in the silicon surface; then it builds the stack of metal wiring that connects them. A modern chip takes more than 300 processing steps in a fixed order and can have eleven or more layers of metal wiring, and a wafer for an advanced chip spends 11–13 weeks in the fab on average, up to 15. The Transistors guide’s Making them chapter covers the fab in depth, and the Systems guide’s Packaging and chiplets chapter does the same for packaging. This page keeps to what the design team sees and does.

Tapeout hands the foundry the finished layout. From then on the design team works through four streams of evidence at once:

  • the fab’s process and yield data, which say how well the wafers were made;
  • the production test program (wafer sort, final test and burn-in), which decides which parts ship;
  • lab bring-up and characterization, which show how the design actually behaves; and
  • reliability qualification, which shows how long it will keep working.

They meet at one decision per chip revision: ship it, fix it by changing only the wiring layers, or respin with a new full mask set. Hunting design bugs in silicon (post-silicon validation) follows four steps: detect a problem, localize it, find the root cause, then fix or bypass it by patching, by editing the circuit, or, as a last resort, with a new mask.

This is where choices made before tapeout get judged. Design for test (DFT), the extra logic that lets a tester reach inside each chip, decides how fast bad dies are found and diagnosed. Spare cells, unused gates scattered through the layout, decide whether a bug can be fixed in metal. Debug access and on-chip monitors decide how long a bug hunt takes. The fab and packaging steps are covered in depth in Making them and Packaging and chiplets; this page follows the design team’s side of them.

new metal masksTapeoutMasksFabWafer sortPackageBring-upQualifyProduction
Bring-up finds

Tap a station to read about it, then choose what bring-up finds.

From tapeout to production in eight stations. The bring-up result decides whether the design moves on or goes back for new masks.Share freely with credit: ‘Figure from chipfieldguide.com’
Sequenced processing steps, modern chip
300+
Fab time at advanced nodes, industry average
11–13 weeks
AEC-Q100 Grade 1 HTOL
1,000 h at 125 °C

Sources: Wikipedia on semiconductor fabrication and the AEC-Q100 qualification standard.

Each layer of the chip needs its own stencil, a glass plate called a . A machine draws the pattern onto it with a fine beam of electrons or a laser. Then the plate is checked, cleaned and covered with a . That is a thin, clear film stretched on a frame a short distance above the pattern. Dust that lands on the film is blurry and doesn’t print.

A full set for a modern chip can need 60 or more plates. Each must be nearly perfect, because the same plate prints every chip on every wafer. That is why a set for an advanced chip costs millions of dollars.

A is a plate of very pure quartz glass coated with a thin film of chrome. The pattern for one layer is cut into the chrome, so light passes through where the chip needs a shape and is blocked everywhere else. A mask shop makes each one in about a dozen steps:

  1. Write the pattern into a light-sensitive coating (resist) on the plate with an electron beam or laser.
  2. Develop the resist, etch the chrome where it is exposed, and strip the resist.
  3. Measure the critical dimensions (the widths of the smallest shapes) and check that every feature sits in the right place.
  4. Clean the plate, inspect it for defects such as leftover chrome or pinholes, and repair them.
  5. Clean again, mount the , and run a final check.

The pellicle is a transparent film on a frame glued over the mask. Dust that lands on it is too far out of focus to print. Masks for the newest machines, which use extreme ultraviolet light, work differently: they reflect the light from a stack of alternating molybdenum and silicon layers instead of letting it through.

Masks are expensive because there are many of them and each must be close to perfect. A Older processes needed a few dozen masks. A 16 nm set has about 60, and sets for 7 nm-class processes were expected to need more than 70, because the densest layers are printed in two or more passes, each with its own mask (this is called multiple patterning). Finer features also take longer to write and inspect. A flaw on a mask repeats in every chip it prints, so masks are inspected far more strictly than any single wafer.

Mask-set cost climbs steeply with each process generation. One estimate, quoted in a UC Davis course handout from AMD’s chief technology officer in 2016, put a full set at about $2 million for a 28 nm process, $4 million at 14/16 nm and $8–10 million at 7 nm. Two drivers dominate: finer features, which take longer to write and inspect, and the extra masks that multiple patterning needs for a single layer.

Mask inspection measures feature placement as well as size, because a correctly sized shape in the wrong place turns into an alignment (overlay) error on every wafer. For the design team, mask cost shapes the fix strategy later. The masks for the transistor layers at the bottom of the stack (the base layers) are the most expensive. A fix that changes only the metal wiring reuses them and buys only new metal and via masks. A full respin needs a new set of base-layer masks before any wafer can start.

UV lightquartzchromepelliclelenswafer + resistdiesall clean
Pellicle

Light passes the clear quartz and is blocked by the chrome. The lens focuses the chrome pattern onto the resist on the wafer.

A photomask with its pellicle, the lens and the wafer, side view, not to scale. A speck on the pellicle is out of focus; one on the chrome prints.Share freely with credit: ‘Figure from chipfieldguide.com’

The fab first builds the transistors, the tiny switches, in the surface of the wafer. Then it stacks layers of metal wiring on top, often ten or more.

Each layer is made the same way. The wafer gets a coating that changes where light hits it. Light shines through a mask onto the coating. The changed parts wash away, and the factory adds or removes material where the coating is gone. It works like a stencil, but with light instead of paint. This is called .

The smallest shapes are thousands of times thinner than a hair, too small for ordinary light to print sharply. The newest machines use a special ultraviolet light, made by blasting tiny drops of melted tin with a powerful laser. Air would swallow this light, so it travels through empty space inside the machine and bounces off mirrors.

All in all, a wafer spends about three months in the fab.

Transistors first, then wiring

Wafer processing has two halves. The builds the transistors directly in the silicon surface. The back end of line (BEOL) builds the above them: layers of wires separated by insulator and joined by vertical plugs called vias. After each wiring layer the wafer is polished flat by so the next layer starts on a level surface. Polishing works evenly only if the metal is spread evenly, which is why the layout has to meet metal-density rules before tapeout.

Printing each layer

Every layer that has a mask goes through the same loop. The wafer is coated with photoresist, a light-sensitive film. A machine called a scanner shines light through the mask and a lens that shrinks the image onto the resist. The exposed resist is washed away (or kept, depending on the resist type), and the wafer is then etched, or has atoms implanted, through the openings.

How small a shape can be printed depends on two things: the wavelength of the light, and how wide a cone of light the lens can collect, its . The smallest feature scales as wavelength divided by numerical aperture, so shorter light or a wider lens both help. Deep-ultraviolet machines use lasers at 248 nm and 193 nm. Even with water between lens and wafer to widen the cone, one exposure at 193 nm can print parallel lines no closer than about 76 nm center to center (that spacing is called the pitch). Denser layers use : split the shapes across two masks and print them in turn (called LELE, for litho-etch-litho-etch), or grow thin walls (spacers) on the sides of a first pattern and use those walls as the lines, doubling or quadrupling their density (SADP and SAQP, self-aligned double and quadruple patterning).

cuts the wavelength to 13.5 nm. Current EUV scanners have a numerical aperture of 0.33; the newest high-NA tools raise it to 0.55. Even EUV may still need a second exposure on the densest layers, for example to print long lines and then cut them into pieces.

Keeping it on target

With hundreds of steps in a row, a small error anywhere ruins the result, so the fab measures constantly. This is . Film thickness is measured optically, by how the film reflects light. On every critical layer the fab checks the width of the smallest shapes and the , how well the layer lines up with the one below; even a misalignment of less than a nanometer can make a chip fail. Inspections between steps catch processing mistakes as they happen, and just before wafers leave the fab an electrical test checks special test patterns printed beside the chips.

The design team sees the fab mostly through data. Each wafer carries small test structures between the dies, built to measure one process parameter each: transistor turn-on threshold, and the resistance of polysilicon and contacts. Their results (the parametric data) show where each wafer landed within the process’s allowed range. Yield losses come in two kinds. Line yield is whole wafers lost to damage or processing mistakes before they reach test; die yield is the share of dies on a finished wafer that pass, which is what wafer sort measures. Lining up a lot’s parametric data (a lot is a batch of wafers processed together) with its test results is the first step in explaining a bad wafer: a lot with slow transistors and a lot with a particle problem fail in different ways.

Lithography choices also reach back into the design. A layer printed with two masks needs every shape assigned to one of them, a step called coloring, and neighboring shapes on the same mask must stay far enough apart to print; the router enforced those rules before tapeout. High-NA EUV uses anamorphic optics, which magnify differently in the two directions and halve the area exposed in one shot from 26 × 33 mm to 26 × 16.5 mm. That halves the largest die one exposure can print, and pushes very large designs toward stitching two exposures together or splitting into several dies.

Why the 76 nm figure, and how far EUV goes, is worked through step by step in Under the hood below. The process steps themselves (deposition, etching, implantation, polishing) are in Making them.

silicon wafer✓measure: flatness, particles
1 / 6

Bare wafer: A polished silicon wafer, before any processing.

The two halves of wafer processing in cross-section, not to scale: transistors (FEOL), then the metal stack (BEOL), each layer measured. Four metal layers stand in for ten or more.Share freely with credit: ‘Figure from chipfieldguide.com’

Before the wafer is cut up, a machine presses a card of tiny needles onto each chip and runs tests. Chips that fail are marked and never get packaged. The share of chips that pass is called the . If 90 out of 100 chips pass, the yield is 90 percent.

Passing chips are not all equal. Some run faster than others. Some have one small broken part that can be switched off. Makers sort them into grades, like eggs by size. The fastest become pricey models, and the rest become cheaper ones. This is called .

The first test happens while the dies are still on the wafer. It is called . A computer-controlled tester (, automatic test equipment) connects to a machine called a prober, which steps the wafer under a probe card: a custom circuit board carrying fine needles that touch each die’s connection pads. The tests come in two kinds:

  • Electrical (parametric) tests check that every pad makes contact, that the die doesn’t draw too much current, and that it doesn’t leak.
  • Logic tests use test circuitry the designers built in (see Design for test). In test mode, the chip’s thousands of one-bit storage cells (flip-flops) are linked into long chains so the tester can shift a pattern in, run the logic for one clock tick, and shift the result out. Small self-test engines check each on-chip memory.

The tester also measures small test structures printed between the dies, which report process values such as how much voltage it takes to switch a transistor on. Each die’s result is recorded at its position on a , and failing dies are marked so they are never packaged.

Why big chips yield worse

Most dies that fail were hit by a defect: a speck of dust or a flaw in one step that breaks or shorts a wire. If such killer defects are scattered at random with an average of D0D_0 per square centimeter, the chance that a die of area AA escapes all of them is Y=e−D0AY = e^{-D_0 A}. Two examples at D0=0.2D_0 = 0.2 defects per cm²:

  • A 1 cm² die: e−0.2≈0.82e^{-0.2} \approx 0.82, so about 82% of dies work.
  • A 4 cm² die: e−0.8≈0.45e^{-0.8} \approx 0.45, so only about 45% work.

That is one reason large designs are split into several smaller dies in one package. Then every die has to work, so each is tested before assembly; a die that passes is called a . Testing first keeps one bad die from ruining a whole stack of good ones.

Each tested die also goes into a bin. Fail bins record which test failed; pass bins record speed or power grade. Parts that miss the top speed, or that have one defective block, can be sold at a lower clock speed or with that block switched off.

Tester time is expensive, so production tests are kept short. They give a go/no-go answer with no diagnosis, and testers save time by testing several dies at once. That means yield work needs a separate mode that logs the full failure data (which pattern failed, at which output) for a sample of dies, so the failures can be traced to a location later.

Wafer maps carry signatures that the yield model has to separate.

  • Edge loss. Deposited films are often well controlled across the middle of the wafer but not near the edge, so dies near the edge fail together. Parametric tests and in-line inspection usually skip edge dies, so this loss shows up as die yield loss even though random defects don’t cause it.
  • Clustering. Real defects bunch together, so some dies collect several while their neighbors get none. That makes the random (Poisson) model pessimistic for large dies; the negative binomial model adds a cluster parameter α\alpha to account for it.

For products built from several dies the arithmetic is unforgiving. If four dies are assembled untested and each is good with probability 0.9, the package works only if all four do: 0.94≈66%0.9^4 \approx 66\%, before any loss in assembly itself. Testing each die first and assembling only known-good die removes that multiplication. It also means the test logic has to reach every die, including dies buried in the middle of a stack.

Wafer mapprobed: 0 / 616good: 0fail: 0killer defect
Die area
Color by

616 whole dies, 1 cm² each. The dots are killer defects at 0.2 per cm². Probe the wafer to fill in the map.

Wafer sort builds a wafer map. Defects are a random scatter at the page’s example density of 0.2 per cm²; bin splits are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Next the wafer is sawn into separate chips. Each good chip goes into a , the hard case that protects it and connects it to the rest of a device. Inside, the chip is joined to the package by wires much thinner than a hair, or it sits face-down on tiny balls of solder.

The biggest chips, such as the ones that run AI, put several chips in one package. They sit side by side or stack on top of each other. Every packaged chip is tested again. Some are also left running in a hot oven for hours. This makes weak chips fail in the factory instead of in your phone.

After sort the wafer is ground thinner and sawn apart, and only the unmarked dies go on. There are two main ways to connect a die to its package.

  • A uses fine wires of aluminum, copper, silver or gold, each welded from a pad on the die to a pad on the package. It is the cheapest and most flexible method and is used for the vast majority of packages.
  • puts a small solder ball (a bump) on each pad, turns the die face-down onto the package, and melts the solder to make every connection at once. The gap is then filled with an insulating glue, the underfill. The short connections let signals switch faster and carry heat away better, and the underfill spreads the stress that builds because the die and the board expand by different amounts when they warm up.

Either way, a , a small multilayer circuit board, carries the connections out to the balls that solder to the main board.

Big chips are increasingly built from several dies in one package, each called a . Splitting a design lets a product be bigger than the largest die a lithography machine can print, lets each part use the process that suits it, and gives smaller dies that yield better. How densely the dies can be wired together depends on the bump pitch, the spacing between neighboring solder bumps:

  • In a standard package, dies sit side by side on the package substrate, with bumps about 100–130 µm apart.
  • In an advanced (“2.5D”) package, the dies sit on a , a slab of silicon used only for wiring, or are joined by a small silicon bridge. Bumps can then be 25–55 µm apart.
  • True 3D stacks put dies on top of one another and connect them through , metal-filled holes straight through the silicon. replaces the solder bumps, which sit tens of micrometers apart, with direct copper-to-copper bonds about 9 µm apart in production.

The Packaging and chiplets chapter covers these options in depth, and the AI accelerators industry page shows where they matter most.

then runs the full test program on each packaged part at the speed the maker guarantees. Some products also get burn-in: parts run at high temperature and above-normal voltage while being tested, so that weak parts fail in the factory instead of in a customer’s hands. Short burn-in lasts 10–30 hours; long burn-in 100–1,000 hours.

Every assembly step can damage a good die, so testing moves earlier and repeats: each die at wafer sort, sometimes each partial stack, and the finished package. Chiplet designs pay for this in area, because every chiplet needs its own clocking, power management, test and debug circuitry that a single large die would have shared. Stacks add two more problems. Test signals have to pass through the dies between the tester and the target, and heat builds up: in one simulation study, folding a 7 nm processor into two stacked tiers raised its peak temperature by up to 12 °C over the flat version, and putting memory on one tier and logic on the other halved that rise.

Die-to-die links also need a way to survive a bad connection. The UCIe standard for chiplet links takes different approaches in its two package types: the advanced-package version provides spare lanes that can stand in for a failed one, while the standard-package version instead runs the link at reduced width. Burn-in is very expensive and has to be balanced against the reliability target, so long burn-in is kept for products whose reliability requirements justify the oven time.

Stacking order is a yield decision. Wafer-to-wafer bonding, which bonds two whole wafers before cutting, is easier to align and reaches the finest pitches, 400 nm in research; chip-on-wafer bonding, which places individual dies onto a wafer, reached about 2 µm in the same work. But a wafer-to-wafer stack is defective if any one of its dies is, whereas stacking individual dies lets each be tested first and even matched by bin. For the design team, that means the DFT plan has to cover each die alone at wafer sort, each partial stack, and the finished package.

substratedie (face up)bond wires
Connection

Tap a part to read about it, then compare the four ways to connect dies.

Four ways to join dies to a package, in cross-section, not to scale. Pitches are the page’s cited values. Toggle final test to see the part in a test socket.Share freely with credit: ‘Figure from chipfieldguide.com’

When the first chips arrive, engineers put one on a test board and switch it on slowly. They watch how much electricity it pulls, because too much can mean a short circuit inside. In 2023 the Tiny Tapeout project streamed this live online. The chip pulled a normal amount, and the first simple designs worked. One counter counted from 4 to 9 instead of 0 to 9, maybe because a signal arrived a little too late.

Then the team tests the chip again and again. They change how fast it runs, how hard the electricity pushes (the voltage) and how hot it gets. A chip usually needs a stronger push to run faster. The results go into a , a grid that shows where the chip works and where it fails.

is the first lab work on a new chip. It goes in an order where each step relies only on what earlier steps proved works:

  1. Power. Apply the supply voltages through a lab supply set to a current limit, and check the current drawn against the design’s estimate.
  2. Clocks. Feed in a steady reference clock and check that the on-chip PLL (phase-locked loop, a circuit that multiplies the reference up to the chip’s working speed) locks onto it.
  3. Debug port. Talk to the chip through JTAG, a standard four- or five-wire test port. The first thing to read is the IDCODE, a fixed number built into every chip of the design; reading the right one proves the port works.
  4. Built-in tests. Run the scan chains and the memory self-tests that wafer sort uses (see the previous sections).
  5. First program. Let the chip run its boot ROM, a small program stored in the chip itself, then load real software.

Small things can stall it. The Tiny Tapeout team found its external scan chain worked only after one pin was grounded by hand, likely because a pull-down resistor (which holds an unused input at 0 V) was missing. A log from a typical session might look like this.

bringup_board07.loglog
[00:00] PSU  VDD_CORE 0.80 V, limit 0.50 A; VDDIO 1.80 V   (illustrative)
[00:01] Power on: I(VDD_CORE) = 0.042 A, I(VDDIO) = 0.006 A       OK
[00:02] REFCLK 25.000 MHz present on CLK_IN                       OK
[00:03] JTAG: IR length 5, IDCODE 0x1A2B3C4D (expected)           OK
[00:04] PLL lock, core 800 MHz (monitor pin 12.5 MHz = /64)       OK
[00:06] Scan integrity flush: 400/400 chains pass                 OK
[00:09] MBIST: 37/37 memories pass                                OK
[00:12] Boot ROM: UART banner "boot v1.0"                         OK
[00:15] DDR training: FAIL at 3200 MT/s, PASS at 2400 MT/s        INVESTIGATE
  1. 1L1Two supplies: VDD_CORE feeds the logic, VDDIO the input/output pins. The current limit protects the only chips the lab has; a short circuit shows up as a supply stuck at its limit.
  2. 2L2The current drawn while idle, compared with the power estimate. Far too high suggests a short, a stuck pin or a block that never left reset.
  3. 3L4The IDCODE read back correctly: the first proof of life from inside the die. Every later debug step goes through this port.
  4. 4L5The core clock is too fast to probe on the board, so the chip divides it by 64 and sends it to a pin to be measured.
  5. 5L6Shifting a known pattern through every scan chain checks the test logic itself before trusting any scan result.
  6. 6L7MBIST is memory built-in self-test: small engines on the chip that write and read every memory and report pass or fail.
  7. 7L8The boot program printed its greeting over a serial port (UART): the processor is running code.
  8. 8L9DDR is the link to external memory, and training tunes its timing. It works at a lower rate but not the target: a margin problem. Next step: sweep data rate against voltage.

Testing before tapeout ran in simulation, which can see every signal but runs very slowly. Silicon runs orders of magnitude faster, so it reaches situations simulation never did, but engineers can watch only a few internal signals. The bugs found now are of two kinds: logic bugs, mistakes in the design that slipped past verification; and electrical bugs, where interference between wires, noise on the power supply, heat or manufacturing variation make correct logic misbehave. Finding them relies on access planned at design time: the JTAG port, the scan chains (which can dump the contents of every storage cell), and that record chosen signals until a trigger fires. This hunt for design bugs is called .

measures the real operating limits. Engineers choose a test that some chips pass and others fail, take a statistically meaningful sample of chips, repeat the test over combinations of conditions such as supply voltage, temperature and clock speed, and plot each result as a . The same work produces the shorter production test that every chip will get.

shmoo_core_25C.txttext
SHMOO  core_func  T = 25 C  part #07   * = pass  . = fail  (illustrative)
VDD(V)  f(MHz):  600  700  800  900  1000  1100  1200
 0.95             *    *    *    *    *     *     .
 0.90             *    *    *    *    *     .     .
 0.85             *    *    *    *    .     .     .
 0.80             *    *    *    *    .     .     .
 0.75             *    *    *    .    .     .     .
 0.70             *    *    .    .    .     .     .
 0.65             *    .    .    .    .     .     .
  1. 1L1One test, one chip, one temperature. Characterization repeats this across many chips and temperatures.
  2. 2L2Rows are supply voltages, columns clock frequencies. Each * or . is one run of the test.
  3. 3L6The normal supply. The target is 800 MHz, and 900 MHz also passes: one 100 MHz step of frequency margin.
  4. 4L8At 0.70 V the chip still runs at 700 MHz. The diagonal edge is the speed limit falling with voltage, because transistors switch more slowly at lower voltage. That is normal.

Detection is cheap, because silicon runs real programs at full speed and a crash is obvious. Localizing the bug, narrowing it to one block and a short stretch of time, dominates the effort and cost, and one bug can hide others behind it. Electrical bugs are the hardest: a logic bug typically takes hours to days to localize, an electrical bug days to weeks, with more expensive equipment.

The tools, and their limits:

  • Trace buffers. How much a single run can capture is set by on-chip trace-buffer capacity. Compressing the data adds 20–30%, and sharing buffers between blocks on the fly, instead of fixing each block’s share at design time, doubled the data captured in one case study.
  • Speed-path debug. A speed path is the slowest chain of logic, the one that limits the clock frequency. To find it, teams collect results at several clock frequencies, using clock-tuning circuits built into the design that stretch or shrink the clock for selected paths only.
  • Scan dumps. Stopping the clock and shifting out every storage cell gives a complete snapshot, but it helps only if the failure repeats exactly. Many electrical bugs depend on noise and temperature and don’t repeat, which is why the shmoo sweep comes first: it turns a vague failure into a reproducible point in voltage, frequency and temperature.
  • Assertion monitors. Assertions are rules written during verification about how a block’s inputs and outputs must behave. Built into the chip as hardware checkers, a failing one narrows a crash to one block and one moment.

The methods used to diagnose manufacturing defects transfer only partly. They rely on scan test turning the chip’s logic into a simple input-to-output circuit for one clock tick, and a functional failure many millions of cycles into a program can’t be replayed that way. The practical lesson for the next design is to budget debug hardware like any other feature: trace-buffer depth, trigger logic, monitor coverage and clock control all have to exist before the first part comes back.

A bug in the first chips leaves three choices. The first is to work around it in software, if that’s possible.

The second is to change only the top wiring layers. Designers scatter small, unused pieces of logic, called , across the chip just for this. A fix can wire them in with new masks for only a few layers.

The third is to remake the chip with all new masks. This is called a , and it costs the most.

Some bugs show up only after chips are sold. In 1994 a math professor found that some Pentium computer chips got certain division problems slightly wrong. Replacing them cost Intel $475 million.

A late change to a finished design is called an engineering change order, or ECO. The cheapest kind in silicon is a . During layout, designers scatter extra, unconnected logic gates across the chip: these are not connected to the working circuit. A fix rewires some of them by changing only the metal wiring layers. A change confined to metal costs much less than one that needs new masks for every layer, and a spare is useful only if one sits close to the logic being fixed.

Before ordering new metal masks, a team can prove the fix on a few chips with a . A focused ion beam (a finely aimed beam of charged gallium atoms) drills through the layers to cut a wire, and lays down metal to add one.

Some processor bugs can be patched without new silicon. Processors run part of their instruction handling from microcode, internal programs that can be updated in the field. Some designs also add hardware that watches for the sequence of states known to trigger a bug and, when it appears, switches to a slower but safe way of executing instructions. A full respin needs new base-layer masks and another full trip through the fab, so it is the last resort. In the Pentium FDIV bug, five entries in a lookup table inside the divider held zero instead of the right value. Intel replaced affected processors on request, and later versions of the chip were unaffected.

The decision weighs how severe the bug is, whether software can avoid it, how many other bugs are still open, and whether a metal fix is even possible. A respin costs enough that it should carry every fix known at the time.

In practice the fix path depends on what kind of problem was found:

FindingUsual pathWhy
Logic bug with spare cells nearbyFIB edit to prove the fix, then a metal-only ECOOnly the wiring changes; base-layer masks are reused
Logic bug with a software or microcode workaroundErratum now, fix in the next revisionNo schedule hit for the current parts
Speed path a few percent too slowSell faster parts in a separate bin, raise the voltage, or a metal ECO with spare buffersTransistor sizes are fixed in the base layers
Analog, device or I/O circuit problemFull respinNeeds transistor or implant changes
Systematic yield loss traced to a layout patternTune the process now, fix the layout at the next revisionFab changes first; design changes when masks are cut anyway

A metal ECO is not free. Every one reruns the signoff checks on the changed design: equivalence checking (proving the edited netlist does what the fix intended), extraction and static timing analysis (recomputing wire delays and rechecking every timing path), design-rule and layout-versus-schematic checks, and local metal refill. It avoids a full new mask set, which costs far more than a few metal masks, but new metal and via masks and the back-end processing steps still take time.

Spare-cell density is a bet placed before tapeout on where bugs will be. In one study, pre-placing spare cells in 70% of the layout’s unused space let a spare-cell rewiring method repair 70% of injected functional errors. A FIB edit that works on real silicon is the strongest evidence a metal fix will too.

Mask setmetal, viabase0 new masksLogic, top viewA✗bugFFwrongspare cellin
Fix path

A logic bug in one gate, with an unused spare cell nearby. Choose a fix.

Fix paths after first silicon. Left, the mask set (new masks highlighted; the 10 + 14 split is illustrative). Right, the logic from above with a spare cell placed before tapeout.Share freely with credit: ‘Figure from chipfieldguide.com’

Before a chip is sold by the millions, it must pass stress tests that age it quickly. This is called . In one common test, hundreds of chips run in a hot oven for 1,000 hours, about six weeks, and none may fail.

Over a product’s life, failures follow a . Many fail early, few fail for most of the product’s life, and more fail again as parts wear out. Drawn as a graph, it looks like the side view of a bathtub. Burn-in removes many of the early failures.

runs a fixed set of stress tests on chips from several production lots (a lot is a batch of wafers processed together). The automotive standard is freely readable and gives a concrete example. Grade 1 is its rating for parts that must work in surroundings up to 125 °C.

TestGrade 1 conditionSample
, running powered at high temperature (JESD22-A108)125 °C surroundings, 1,000 hours, highest operating voltage77 parts × 3 lots, 0 fails
Early life failure rate (ELFR)Per AEC-Q100-008800 parts × 3 lots, 0 fails
Temperature cycling (JESD22-A104)−55 °C to +150 °C, 1,000 cycles77 parts × 3 lots, 0 fails

Parts are tested before and after each stress, and a later change to the device can require qualifying it again. How often parts fail in use is quoted in : failures per billion (10910^{9}) device-hours. Running parts hot works as a shortcut because many failure processes are chemical or atomic and speed up with heat. The Arrhenius model turns a temperature difference into an , how many hours of normal use each hour of stress stands for.

Ramping up production is . The share of good chips is often low when a new process or design starts, and the team raises it by finding and removing the causes of loss, using inspection in the fab, test structures, and diagnosis of where failing chips are broken. Characterization continues through production to improve the design and the process.

Qualification targets the wear-out mechanisms, each a slow physical change that eventually breaks a part:

  • Electromigration (EM): current gradually pushes metal atoms along a wire until it thins and opens.
  • Hot-carrier degradation: fast-moving electrons damage the transistor near its drain, slowing it.
  • Time-dependent dielectric breakdown (TDDB): the thin insulator under the gate wears through.
  • Negative bias temperature instability (NBTI): PMOS transistors held on at high temperature drift, needing more voltage to switch.

HTOL is usually the final qualification step: typically about 100 parts for 1,000 hours at raised voltage and temperature. Stress levels are chosen to guarantee zero failures, and a test with zero failures yields little statistical data. Worse, one stress condition speeds up each mechanism by its own factor, so a single acceleration factor gives the wrong answer when several mechanisms compete.

AEC-Q100 allows the HTOL temperature to be set by the chip’s own (junction) temperature rather than the surroundings. In that case the team must show the run is equivalent to 1,000 hours at the rated ambient temperature, using an activation energy of 0.7 eV or another justified value. It also asks for drift analysis of key electrical parameters after stress, to confirm the guard bands, the margins between measured performance and the data sheet, are large enough. This ties reliability back to characterization: the margin a part shows when new has to cover how far it drifts over its life.

The sample sizes reflect what each test is for. HTOL uses 231 parts to show the wear-out mechanisms stay outside the product’s life; ELFR uses 2,400 parts, 800 from each of three lots, because early-life failures are rare and need a large population to show up at all. Neither proves a low failure rate in the field on its own. Under the hood, below, turns a zero-failure result into a FIT bound step by step and shows how much that bound depends on the assumed activation energy.

Infant mortalityUseful lifeWear-out↑ failure rate1 h1 week1 year10 yearstime in use (log scale) →burn-inshipped here

Tap a region of the curve, then add burn-in.

The bathtub curve of failure rate over time, and the part that burn-in removes. Shape and time axis are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

The yield-learning loop

Leachman’s course notes split die yield into a part that depends on die area and a part that doesn’t:

DY=Ys e−AD0\mathrm{DY} = Y_{\mathrm{s}}\, e^{-A D_0}

Here AA is the die area, D0D_0 the density of random killer defects, and YsY_{\mathrm{s}} the systematic-limited yield, the losses that hit a die whatever its size. Taking logarithms gives ln⁡DY=ln⁡Ys−AD0\ln \mathrm{DY} = \ln Y_{\mathrm{s}} - A D_0, a straight line in AA with slope −D0-D_0 and intercept ln⁡Ys\ln Y_{\mathrm{s}}.

One product, with one die size, gives only one point on that line. The windowing technique makes more from the same wafer maps:

  1. Compute the ordinary die yield: one point at area AA.
  2. Group neighboring dies into pairs and treat each pair as one die of area 2A2A, which works only if both halves work. Its yield is a second point.
  3. Repeat with groups of three, four and so on, for points at 3A3A, 4A4A, …
  4. Fit a line to ln⁡(yield)\ln(\text{yield}) against area: the slope gives D0D_0 and the intercept gives YsY_{\mathrm{s}}.

The split is by area dependence, not by cause, and a mechanism can land on either side. Edge loss, for instance, grows with die size and so ends up in D0D_0 even though random defects don’t cause it. For large dies, where defects cluster, the negative binomial model Y=(1+AD0/α)−αY = (1 + A D_0/\alpha)^{-\alpha} fits better; as the cluster parameter α\alpha grows past about 10 it becomes the same as the Poisson model.

defectdead zone0-0.5-1-1.5-201.534.56window area, cm²ln(yield)
1 / 4

Step 1: ordinary die yield, 62.0% of 392 dies at A = 1.5 cm². One point fixes neither D₀ nor Y_s.

Windowing on one simulated wafer: 1.5 cm² dies, random defects at D₀ = 0.25 cm⁻² plus a dead zone (systematic loss, Y_s ≈ 0.91). Step through window sizes; the fit uses the points shown.Share freely with credit: ‘Figure from chipfieldguide.com’

Fitted numbers measure the loss; finding its cause takes . From a failing die’s tester response, a diagnosis tool lists candidate defect locations and behaviors. It is judged on resolution (how few candidates it reports) and accuracy (whether one of them is the real defect). Volume diagnosis then compares results across many failing dies to find a shared root cause, and selected dies go to , which gives indisputable confirmation but is destructive and slow, so only a few dies get it. An inaccurate diagnosis sends that analysis to the wrong place on the wrong die.

The other data sources have limits that make diagnosis central. Optical inspection in the fab finds defects less well as features shrink, and test structures such as comb drives and ring oscillators don’t reflect the variety of layout patterns in a real product, so failing product chips have themselves become the main vehicle for yield learning. Diagnosis is still improving: LearnX, which adds machine learning on top of conventional diagnosis, returned a single correct candidate for 73.2% of circuits in 30,000 simulated fault experiments, 86.6% more often than the commercial tool it was compared with.

Physical failure analysis removes layers down to the suspect spot and images it. A focused ion beam can cut a cross-section through the exact location diagnosis points to, and prepare an electron-microscope sample of a single transistor out of millions. The design team’s contribution comes earlier: test logic that makes diagnosis precise, fault models that know the layout, and full failure logs kept from production test.

Characterization methodology

Delft’s outline of characterization reads like a procedure:

  1. Set out to find the exact operating limits, with a worst-case test.
  2. Choose a test that some chips pass and others fail.
  3. Select a statistically significant sample of chips.
  4. Repeat the test for every combination of two or more environmental variables, measuring DC and AC (steady-state and timing) parameters.
  5. Plot the results as a shmoo.
  6. Diagnose and correct design errors.
  7. Develop the production test program.

It then continues for the production life of the chip.

Reading a shmoo is reasoning about physics. Gate delay falls as supply voltage rises, so a path that limits speed (a setup-limited path, where data must arrive before the next clock edge) gives a pass/fail boundary that slopes up and to the right. Hold violations, where data changes too soon after a clock edge, don’t depend on the clock period, so they show up as failures at every frequency, worst at high voltage on fast silicon. Holes inside the pass region suggest noise, a data-dependent electrical bug or a problem with the test itself. The sample has to include fast and slow material and every temperature corner before the team trusts any edge. Guard bands then go between the worst measured edge and the data-sheet limit, and AEC-Q100 asks for drift analysis after stress to confirm they are big enough.

Lithography resolution

Rayleigh’s criterion says two points are just resolved when the central peak of one point’s image falls on the first dark ring of its neighbor’s; for a microscope this gives R=0.61 λ/NAR = 0.61\,\lambda/\mathrm{NA}. The is NA=nsin⁡α\mathrm{NA} = n \sin\alpha, where nn is the refractive index of the medium and α\alpha the largest half-angle of light the lens accepts. Lithography generalizes the criterion to R=k1λ/NAR = k_1 \lambda / \mathrm{NA}, where k1k_1 collects everything else that sets resolution: the resist, the mask and the illumination. Depth of focus falls faster: in Rayleigh’s low-NA approximation, DOF=k2λ/NA2\mathrm{DOF} = k_2 \lambda / \mathrm{NA}^2. As an illustration, a 248 nm (KrF) tool with NA 0.6 working at k1=0.75k_1 = 0.75 resolves 0.75×248 nm/0.6≈0.31 μm0.75 \times 248\,\mathrm{nm} / 0.6 \approx 0.31\,\mu\mathrm{m}.

Resolution enhancement, meaning pre-distorted mask shapes (OPC), shaped illumination and phase-shifting masks, drives k1k_1 down, but a dense pattern of lines and spaces has a hard floor: k1k_1 cannot go below 0.25. Since k1k_1 here describes the half-pitch (one line or one space), the smallest pitch one exposure can print is 2×0.25 λ/NA=0.5 λ/NA2 \times 0.25\,\lambda/\mathrm{NA} = 0.5\,\lambda/\mathrm{NA}. The numbers follow directly:

Toolλ\lambdaNAMinimum single-exposure pitch
ArF immersion193 nmabove 1 (water)≈76 nm, the expected practical limit
EUV13.5 nm0.330.5×13.5/0.33≈20 nm0.5 \times 13.5 / 0.33 \approx 20\,\mathrm{nm}
High-NA EUV13.5 nm0.550.5×13.5/0.55≈12 nm0.5 \times 13.5 / 0.55 \approx 12\,\mathrm{nm}

The EUV wavelength and NA values come from IEEE Spectrum, the immersion figure from the multiple-patterning article, and the EUV pitches are arithmetic. The depth-of-focus term explains why high NA costs more than the tool itself: the focus budget shrinks as 1/NA21/\mathrm{NA}^2, so going from NA 0.33 to 0.55 cuts it to about (0.33/0.55)2≈36%(0.33/0.55)^2 \approx 36\% of what it was, and wafer flatness, resist thickness and the polishing uniformity that metal fill protects all get tighter.

Reliability models

The has an early-failure period with a high but rapidly falling rate, an intrinsic period with a roughly constant rate, and a wear-out period in which degradation failures rise. Burn-in addresses the first; FIT ratings describe the second; qualification and design rules push the third past the product’s lifetime.

The acceleration factor is the ratio of time to failure in use to time to failure under stress, tuse=AF×tstresst_{\mathrm{use}} = \mathrm{AF} \times t_{\mathrm{stress}}, for the same failure mechanism. For temperature the Arrhenius model gives

AF=exp⁡ ⁣[Eak(1Tuse−1Tstress)]\mathrm{AF} = \exp\!\left[ \frac{E_{\mathrm{a}}}{k} \left( \frac{1}{T_{\mathrm{use}}} - \frac{1}{T_{\mathrm{stress}}} \right) \right]

with k=8.617×10−5 eV/Kk = 8.617 \times 10^{-5}\,\mathrm{eV/K}, temperatures in kelvin, and EaE_{\mathrm{a}}, the mechanism’s activation energy, typically from 0.3 or 0.4 eV up to 1.5 eV or more. Worked through for Ea=0.7 eVE_{\mathrm{a}} = 0.7\,\mathrm{eV}, the value AEC-Q100 names, stress at 125 °C against use at 55 °C:

  1. Convert to kelvin: 125 °C = 398.15 K and 55 °C = 328.15 K.
  2. Ea/k=0.7/(8.617×10−5)≈8124 KE_{\mathrm{a}}/k = 0.7 / (8.617 \times 10^{-5}) \approx 8124\,\mathrm{K}.
  3. 1/328.15−1/398.15≈0.003047−0.002512=0.0005361/328.15 - 1/398.15 \approx 0.003047 - 0.002512 = 0.000536 per K.
  4. AF=exp⁡(8124×0.000536)=exp⁡(4.35)≈78\mathrm{AF} = \exp(8124 \times 0.000536) = \exp(4.35) \approx 78.

So each hour at 125 °C stands for about 78 hours at 55 °C, for mechanisms that really have Ea=0.7 eVE_{\mathrm{a}} = 0.7\,\mathrm{eV}.

Zero-failure results only bound the failure rate. If failures happen at a constant rate λ\lambda, the chance of seeing none in TT equivalent device-hours is e−λTe^{-\lambda T}. Requiring that chance to be at most 1−C1 - C, for a confidence level CC, gives λ≤−ln⁡(1−C)/T\lambda \le -\ln(1 - C)/T. For an AEC-style HTOL:

  1. Parts × hours: 3×77=2313 \times 77 = 231 parts for 1,000 hours = 231,000 device-hours.
  2. Times the acceleration factor: T≈231,000×78≈1.8×107T \approx 231{,}000 \times 78 \approx 1.8 \times 10^{7} equivalent device-hours.
  3. At 60% confidence, −ln⁡(0.4)≈0.916-\ln(0.4) \approx 0.916, so λ≤0.916/(1.8×107)≈5.1×10−8\lambda \le 0.916 / (1.8 \times 10^{7}) \approx 5.1 \times 10^{-8} per hour.
  4. FIT counts failures per 10910^{9} device-hours, so that is about 51 FIT. At 90% confidence (−ln⁡0.1≈2.30-\ln 0.1 \approx 2.30) it is about 128 FIT.

The bound is only as good as the AF, and the AF is only right for mechanisms that share the assumed activation energy. Redo step 2 with Ea=0.4 eVE_{\mathrm{a}} = 0.4\,\mathrm{eV} and the AF falls to about 12, raising the 60% bound more than sixfold, to about 330 FIT.

AFuse h per stress h1101001,00010,000FIT boundfails per 10⁹ h1101001,00010,000E_a/k = 8123 K · 1/328.15 − 1/398.15 = 0.000536 /KT = 231,000 × 78 = 1.79 × 10⁷ hλ ≤ −ln(0.4) / T = 5.11 × 10⁻⁸ /h = 51 FIT
Confidence

AF = 78: each stress hour stands for 78 h at 55 °C. 17.9 M equivalent device-hours, 0 fails → λ ≤ 51 FIT at 60% confidence.

Zero-fail HTOL (231 parts × 1,000 h at 125 °C) to a FIT bound via the Arrhenius model. The bound is only as good as the assumed activation energy; slide E_a to see how much it moves.Share freely with credit: ‘Figure from chipfieldguide.com’

Wear-out mechanisms get their own models. For , Black’s equation gives the median time to failure of a metal line as t50=AJ−nexp⁡(Ea/kT)t_{50} = A J^{-n} \exp(E_{\mathrm{a}}/kT), with JJ the current density and AA a constant for the material and process. The original form uses n=2n = 2; failures dominated by void growth, and narrow lines, fit n=1n = 1. Higher current density and higher temperature both shorten life, which is what the electromigration limits checked at signoff protect against.

Novice · 0 of 5 correct
  1. Q1A metal-only ECO fixes a bug by changing only the chip’s wiring. Which part of wafer processing does it leave untouched?

  2. Q2Why are the densest layers of an advanced chip often printed in two or more passes, each with its own mask?

  3. Q3Defect density is 0.2 per cm². Using the Poisson model Y=e−D0AY = e^{-D_0 A}, what is the die yield for a 1 cm² die?

  4. Q4During bring-up, why do engineers read the chip’s ID number through its JTAG test port before trying to run software?

  5. Q5Why does a multi-die package need known-good die?

Sources

Show Hide 34 sources
  1. Post-Silicon Validation: Opportunities, Challenges and Recent AdvancesSubhasish Mitra, Sanjit A. Seshia, Nicola Nicolici · Design Automation Conference; author copy archived by the Wayback Machine (DOI 10.1145/1837274.1837280) · 2010Four steps (detect, localize, root-cause, fix by patching, circuit editing or a respin); comparison with pre-silicon verification and manufacturing test; electrical vs. logic bugs, hours to days vs. days to weeks to localize; trace buffers; speedpath debug with clock stretching; microcode and field-repairable patches; spare-cell metal fixes.
  2. Semiconductor device fabricationWikipediaOver 300 sequenced steps and eleven or more metal levels; FEOL and BEOL; 300 mm wafers since 2000; thin-film metrology; 11–13 weeks average and up to 15 weeks in the fab at advanced nodes; probing, backgrinding and dicing.
  3. AEC-Q100 Rev-J1: Failure Mechanism Based Stress Test Qualification for Integrated Circuits in Automotive ApplicationsAutomotive Electronics Council · Automotive Electronics Council · 2026HTOL per JESD22-A108 at 125 °C for 1,000 hours for Grade 1, 77 parts from each of 3 lots, zero fails; 0.7 eV for Tj-based HTOL; ELFR with 800 parts per lot; temperature cycling; requalification of changed devices.
  4. EEC 116 lecture handout: PhotomasksBevan Baas · University of California, DavisTwelve mask-making steps from e-beam or laser write to pellicle and audit; approximate mask-set costs of $2 million at 28 nm, $4 million at 14/16 nm and $8–10 million at 7 nm (quoting AMD’s CTO, 2016), driven by finer features and extra masks for multiple patterning.
  5. PhotomaskWikipediaPellicle as a film held out of focus above the pattern; EUV masks reflect with Mo/Si multilayers; high-NA anamorphic optics shrink the field so large designs need stitched exposures.
  6. 7nm Fab ChallengesMark LaPedus · Semiconductor Engineering (April 21, 2016)Citing an eBeam Initiative survey: 60 masks per mask set at 16 nm, expected to rise to 77 below 11 nm; 7 nm logic can take more than 80 photo passes with 193 nm immersion and multiple patterning.
  7. Focused ion beamWikipediaFIB cuts connections and deposits conductors on ICs; used for circuit modification, defect analysis, photomask repair and site-specific TEM sample preparation.
  8. Lecture 40: Lithography: Imaging Tools (CHE323/CHE384)Chris A. Mack · University of Texas at Austin course notes, on the author’s site (lithoguru.com) · 2013g-line 436 nm and i-line 365 nm lamps; KrF 248 nm and ArF 193 nm excimer lasers; steppers from g-line NA 0.28 to i-line NA 0.65; deep-UV steppers and scanners from 1988, ArF scanners from 1998, immersion up to NA 1.35; step-and-scan through a slit, used by all state-of-the-art tools.
  9. Lecture 43: Lithography: Projection Imaging, part 1 (CHE323/CHE384)Chris A. Mack · University of Texas at Austin course notes, on the author’s site (lithoguru.com) · 2013Numerical aperture NA = n sin α, with n the refractive index of the medium and α the maximum half-angle of light making it through the lens.
  10. Lecture 48: Lithography: Resolution and Immersion (CHE323/CHE384)Chris A. Mack · University of Texas at Austin course notes, on the author’s site (lithoguru.com) · 2013Rayleigh resolution equation R = k1·λ/NA; k1 collects everything else that improves resolution (better resist, phase-shifting masks, off-axis illumination), with a physical limit of 0.25.
  11. Lecture 46: Lithography: Defocus and DOF (CHE323/CHE384)Chris A. Mack · University of Texas at Austin course notes, on the author’s site (lithoguru.com) · 2013Rayleigh depth of focus DOF = k2·λ/NA² in the paraxial (low-NA) approximation; smaller pitches lose more image to defocus.
  12. Angular resolutionWikipediaRayleigh criterion: two point sources are just resolved when the central maximum of one Airy disk falls on the first minimum of the other; for a microscope with equal objective and condenser NA, R = 0.61λ/NA.
  13. This Machine Could Keep Moore’s Law on TrackJan van Schoot · IEEE Spectrum · 202313.5 nm EUV from tin droplets hit by a CO₂ laser; mirrors in vacuum; CD proportional to λ/NA; k1 has a physical lower limit of 0.25; NA 0.33 today, 0.55 for high-NA; anamorphic optics halve the field to 26 × 16.5 mm.
  14. Multiple patterningWikipediaPitch below 0.5λ/NA not resolvable in one exposure; about 76 nm minimum pitch for single immersion exposure; LELE pitch splitting; spacer (SADP/SAQP) patterning; EUV may also need line-and-cut.
  15. Overlay Metrology Using Physics and AI-Based Scanning Electron MicroscopyNational Institute of Standards and Technology · NISTOverlay and CD measurements are essential for process control; even a sub-nanometer misalignment can render a chip non-functional.
  16. Yield Modeling and Analysis (IEOR 130 course notes)Robert C. Leachman · University of California, Berkeley, IEOR 130 course page (Internet Archive copy) · 2017Line yield vs. die yield; in-line inspection and parametric test of test patterns before wafer probe; edge loss; Poisson and negative binomial models; DY = Ys·e^(−A·D0); wafer maps and the windowing technique.
  17. VLSI Test Technology and Reliability, Module 2: VLSI Test Process and ATESaid Hamdioui · TU Delft OpenCourseWare · 2010Characterization over combinations of conditions plotted as shmoo plots; production test is go/no-go; burn-in 10–30 hours for infant mortality, 100–1,000 hours long-term; wafer sort with test-site characterization; probe cards and pin electronics.
  18. Product binningWikipediaSorting tested products by characteristics into market tiers; lower clocks or disabled components sold at lower prices.
  19. Three-dimensional integrated circuitWikipediaTSVs; 2.5D interposers vs. true 3D; wafer-to-wafer stacks fail if any one die is bad; die-level stacking lets each die be tested first; heat in stacks.
  20. Wire bondingWikipediaMost cost-effective and flexible interconnect, used for the vast majority of packages; Al, Cu, Ag and Au wire; ball and wedge bonding.
  21. Flip chipWikipediaC4 solder bumps, die flipped face-down and reflowed, underfill; shorter connections with lower inductance and better heat removal; thermal-expansion mismatch.
  22. The UCIe 1.1 Specification: Future Applications of ChipletsDebendra Das Sharma · UCIe Consortium · 2023Chiplets exceed the reticle size, mix process nodes and yield better as smaller dies; standard packages (100–130 µm bump pitch) vs. advanced 2.5D packages with interposers or bridges (25–55 µm).
  23. Hybrid Bonding Plays Starring Role in 3D ChipsSamuel K. Moore · IEEE Spectrum · 2024Solder microbumps at tens of micrometers; production hybrid bonds about 9 µm apart; 400 nm wafer-on-wafer pitch in research; chip-on-wafer vs. wafer-on-wafer.
  24. TT02 Premiere: Silicon Bring-upTiny Tapeout project · tinytapeout.com · 2023Live open-source bring-up: 33.9 mA at 3.3 V on first power-up; scan chain needed pin 8 grounded (likely a missing pull-down); clock tests to 20 MHz; one design counted 4 to 9, possibly setup/hold.
  25. Resource-Aware Functional ECO Patch GenerationAn-Che Cheng, Iris Hui-Ru Jiang, Jing-Yang Jou · DATE 2016 (open proceedings archive) · 2016Metal-only ECO modifies only metal layers after placement is frozen; spare cells are spread over the design during placement, unconnected, and rewired when a fix is needed; too few spares near the patch mean long wires, timing violations and congestion.
  26. Engineering change orderWikipedia contributors · WikipediaAfter masks are made, a change confined to a few (typically metal) layers costs much less than a rebuild needing new masks for all layers; designers sprinkle unused gates for such fixes.
  27. Pentium FDIV bugWikipediaFive lookup-table entries in the SRT divider came out as zero; found by Thomas Nicely in 1994; replacement offered; $475 million pretax charge; later steppings unaffected.
  28. NIST/SEMATECH e-Handbook of Statistical Methods, 8.1.2.4: “Bathtub” curveNIST/SEMATECH · NISTEarly failure period with a high, falling rate; a stable intrinsic period; wearout with a rising rate.
  29. NIST/SEMATECH e-Handbook of Statistical Methods, 8.1.5.1: ArrheniusNIST/SEMATECH · NISTAF = exp[(ΔH/k)(1/T1 − 1/T2)]; k = 8.617 × 10⁻⁵ eV/K; ΔH typically 0.3 or 0.4 up to 1.5 eV or more; for chemical, diffusion and migration mechanisms.
  30. Microelectronics Reliability: Physics-of-Failure Based Modeling and Lifetime Evaluation (JPL Publication 08-5)Mark White, Joseph B. Bernstein · NASA Jet Propulsion Laboratory · 2008Wearout mechanisms EM, HCI, TDDB and NBTI; FIT as failures per 10⁹ part-hours; HTOL of about 100 parts for 1,000 hours yields little statistical data at zero failures; competing mechanisms; Black’s equation for EM.
  31. LearnX: A Hybrid Deterministic-Statistical Defect Diagnosis MethodologySoumya Mittal, R. D. (Shawn) Blanton · IEEE European Test Symposium 2019 (NSF Public Access Repository) · 2019Yield is low for new processes and designs; yield learning; inline inspection and test structures; diagnosis resolution and accuracy; volume diagnosis steers physical failure analysis, which is destructive and slow.
  32. Methods and systems of performing device failure analysis, electrical characterization and physical characterization (US Patent 7,842,920)Theodore R. Lundquist · US Patent and Trademark Office, via Google Patents · 2010Background: focused ion beam systems edit circuits to validate design changes without a trip through fabrication, and cut cross-section trenches at a failing site for SEM imaging in failure analysis.
  33. Understanding Chiplets Today to Anticipate Future Integration Opportunities and LimitsGabriel H. Loh, Samuel Naffziger and Kevin Lepak (AMD) · Design, Automation and Test in Europe (DATE) 2021, proceedings archive · 2021Each chiplet needs its own clocking, power management, test and debug circuitry; chiplets that pass testing are known good die, assembled into a package.
  34. Thermal Analysis of a 3D Stacked High-Performance Commercial Microprocessor using Face-to-Face Wafer Bonding TechnologyRahul Mathur, Chien-Ju Chao, Rossana Liu, Nikhil Tadepalli, Pranavi Chandupatla, Shawn Hung, Xiaoqing Xu, Saurabh Sinha and Jaydeep Kulkarni · arXiv (ECTC 2020) · 2020Overlapping hotspots in 3D stacks; a 7 nm CPU folded into two tiers runs up to 12 °C hotter than 2D, about 6 °C with logic-over-memory partitioning.