Design Flow · Stage 1 of 13 · Front end

Specification

Write down what the chip must do, how fast it must be, how much power it may use and what it may cost. Every later step is checked against this list.

The team writes down measurable targets for speed, power, size and cost, along with the standard connections the chip must support and the factory process it will use. Every later stage, including testing, is checked against these numbers.

Performance, power, area and cost targets; interface standards and how compliance will be shown; power and thermal limits; package and I/O; process choice; manufacturing-test, safety and security requirements; and schedule. Mistakes here cost the most later, so good specs are versioned and every requirement links to the test that proves it.

Builds The targets

Every chip starts as a problem someone wants solved. Maybe a phone battery should last longer, or a car should brake by itself. Before anyone designs a single circuit, a team writes down what the chip must do and the limits it must stay within. That document is the , or spec.

The spec answers three questions with numbers: how fast, how much power, and how big. Engineers call this trio , for power, performance and area. The spec also sets the price, the deadline, and which other devices the chip must talk to.

Getting it right matters. Once a chip is made, a fix means new factory tools and a fresh run through the factory. That can take months and cost millions of dollars.

A chip is a small slab of silicon carrying billions of microscopic electronic switches, called transistors, wired into circuits that compute, store data and communicate. Getting from an idea to a manufactured chip takes a long chain of stages, often called the design flow. It ends with a file that a chip factory (a foundry) uses to make the chip. Specification is the first stage. Its job is to write down, in measurable terms, what the chip must do and the limits it must meet, before anyone starts designing.

The central limits are usually summed up as :

  • Performance: how much work the chip does per second, such as video frames processed or images recognized.
  • Power: how much electrical energy it uses per second, in watts. Nearly all of it ends up as heat.
  • Area: how much silicon it takes up, in square millimeters. Area largely sets what each chip costs to make.

The spec adds cost, a schedule, the standard connections (interfaces) the chip must support, the conditions it must survive, such as heat and cold, and any test, safety and security rules it must meet.

A spec is a stack of documents. It usually starts with a (PRD), which says what the product must do without saying how. It is often written by the company’s marketing department. The engineers then break it down into a detailed technical specification. For a chip, that means numbered . Each one is a single sentence using the word “shall” that can be checked by test, analysis, inspection or demonstration.

The stage produces three things later stages depend on:

  • A go or no-go decision: can this chip be built at a price the market will pay?
  • Budgets that later stages inherit: power for each part of the chip, chip size, speed targets and the number of connections to the outside.
  • A plan that links every requirement to the check that will prove the chip meets it.

Next, architecture decides how the chip is organized into blocks. Engineers then describe each block in a programming-like design language (the RTL stage), and verification proves that description does what the spec says.

You know the outline of the flow: the chip is described in code, the code is turned into logic gates, and the gates are laid out as the physical pattern the factory prints. The spec comes before all of it, and it is where most of the decisions that can’t be undone later are made:

  • the , meaning which factory and which manufacturing generation will build the chip;
  • the , the case that connects the chip to a circuit board and carries its heat away;
  • the die-size class, which largely sets what each chip costs;
  • the interface standards it must support, such as PCIe, USB or a DDR memory type;
  • which blocks to license from other companies instead of designing them; and
  • the level of safety and security assurance the market demands.

Every later stage trades inside those limits. Few of them can move the limits.

The economics punish late discovery. A NASA study of large aerospace hardware and software projects put the relative cost of fixing a requirements error at 1 during requirements, 3–8 in design, 7–16 in build, 21–78 in integration and test, and from 29 to more than 1,500 in operations. That study did not look at chips, but chips have the same structure, with a sharp step at , when the finished layout goes to the factory. After that, a fix needs a new set of (the stencils used to print each layer) and another trip through the factory, which takes months.

Relative cost to fix a requirements error found in integration and test (NASA)
21–78×
…found in operations
29 to 1,500+×
Modeled foundry price per 300 mm wafer, 90 nm vs. 5 nm (CSET, 2020)
$1,650 vs. $16,988

The wafer prices come from a 2020 cost model by Georgetown’s Center for Security and Emerging Technology (CSET). A wafer is the 300 mm silicon disc that hundreds of chips are printed on at once.

1×10×100×1,000×Requirements1×Design3–8×Build7–16×Integration and test21–78×Operations29–1,500+×
Error found during

Found during requirements: the fix is an edit to a document. This is the baseline cost, 1.

Relative cost of fixing a requirements error, by the phase in which it is found (NASA study of large aerospace projects, log scale). Pick a phase.Share freely with credit: ‘Figure from chipfieldguide.com’

The spec stage starts with three things. There is a need: what customers want. There is a market: how many chips, at what price, by when. And there are rules, such as laws for chips in cars.

It ends with the spec: a numbered list of musts. Each one has a target number and a way to check it. The spec also shares out the power and space among the chip’s parts, like splitting an allowance.

Later stages of the flow pass precise computer files from one design tool to the next. The spec stage is different: its inputs and outputs are mostly documents, spreadsheets and models. The table lists the usual ones.

DirectionWhatTypical form
InCustomer and market needs: uses, competitors, target price, sales volume, launch dateMarket and product requirements documents (MRD, PRD), slides, spreadsheets
InInterface standards: the agreed rules for connecting to memory (DDR, LPDDR), expansion cards and other chips (PCIe), USB, Ethernet, and cameras and displays (MIPI)Standards documents from the bodies that own them (JEDEC, PCI-SIG, USB-IF, IEEE)
InQuality, safety and regulatory rulesISO 26262 (car safety), AEC-Q100 (car-grade chip testing), DO-254 (aircraft electronics), export rules, customer quality targets
InTechnology data: candidate manufacturing processes, wafer cost, defect rates, catalogs of ready-made blocks, package optionsData from the factory and block suppliers, usually under non-disclosure agreements; internal cost models
InData from the previous chip: measured power, performance and area, failures in the fieldSilicon reports, spreadsheets
OutChip requirements: numbered “shall” statements with targets, conditions and a way to check eachRequirements database or controlled document
OutArchitecture spec (started here, finished in the architecture stage)Document, block diagrams, memory map
OutBudgets: power per block and mode, area, number of external connections, clock targetsSpreadsheets
OutExecutable models: cost and power model, reference model, performance modelSpreadsheet, C/C++, Python, SystemC
OutVerification plan and requirements verification matrixTestplan (a structured text file), matrix
OutList of blocks to design or buySpreadsheet

There is no standard file format for a spec. What matters is that every requirement has a unique ID and a recorded source. Then a verification matrix, a table with one row per requirement, can show how each one will be checked.

The outputs that matter most downstream are often the least glamorous: the per-mode and the I/O list.

  • The power budget later becomes the chip’s file, a standard format that tells the design tools which regions run at which voltage and which can be switched off. It also sets targets for the power grid, the mesh of metal wires that carries supply current across the chip.
  • The I/O list names every external signal, the standard it follows, its voltage, and whether it needs a (the analog circuit that drives a high-speed link). It drives the package choice, the number of solder balls under the package and the pad ring: the row of connection pads around the edge of the die.

Pads have a minimum size and spacing, so a die with many I/Os and little logic ends up : its size is set by the ring of pads, not by the circuits inside.

inputsoutputsMarket needsStandardsSafety, qualityTechnology dataLast chip’s dataRequirementsArchitectureBudgetsSoftware modelsVerification planMake-or-buy listSpecstage

Inputs on the left, outputs on the right. Tap or hover any of them.

The spec stage’s usual inputs and outputs, as in the table. They are mostly documents, spreadsheets and models, not design files.Share freely with credit: ‘Figure from chipfieldguide.com’

From a wish to a list of musts

The business team talks to customers and writes down their wishes. Engineers turn each wish into a must with numbers in it, and give it an ID number so nobody loses track of it.

A watt measures how fast something uses energy. A bright LED light bulb uses about 10 watts.

The three-way tug of war

Speed, power and size pull against each other. A faster chip uses more power and gets hotter. More processors get more work done, but they take up more space, and space on a chip costs money. A phone chip picks low power, since a phone has no fan and a small battery. A chip for training AI picks speed and accepts a huge power bill.

Why small chips are cheaper

Factories make hundreds of chips at once on a thin, round slice of silicon about the size of a dinner plate. It is called a . Think of cutting cookies from a sheet of dough: smaller cookies means more cookies per sheet.

Tiny specks of dust also ruin any chip they land on. A small chip is less likely to catch a speck, so more of them work. That is why the spec sets a size limit.

Talking to other devices

A chip has to connect to memory, other chips and cables. It follows shared rules called standards, such as USB. Making a USB part that works with every gadget in the world is hard. So most companies buy that part ready-made and tested from experts.

Checking it later

Every line of the spec gets a matching test. Later, testers tick off each test as it passes. This link from promise to test is called . If a line has no test, nobody can prove the chip keeps that promise.

Where requirements come from

Requirements flow down a chain of documents, each more detailed than the last:

  1. Customer needs, gathered by talking to customers and studying competitors.
  2. The PRD, which states what the product must do, not how. When marketing writes it, it may be called a market requirements document (MRD).
  3. The chip requirements: the numbered “shall” statements.
  4. The , which describes the chip as a set of large blocks (processors, memories, special-purpose accelerators, connections to the outside) and how they connect.
  5. The , one per block, which say in detail how the block works, down to what happens on each tick of the chip’s clock. The clock is a signal that ticks at a steady rate, often a billion or more times a second (one gigahertz, GHz), and every step inside the chip happens in time with it.

Chip architects and the program lead usually own the requirements and the architecture, and the engineers who design each block write its microarchitecture spec. Reviewers come from software, verification, physical layout, manufacturing test and packaging, and, in regulated markets, from safety and security assessment.

A finished spec is a gate that work must pass. OpenTitan, an open-source design for a security chip, is a good public example. Before a block can reach OpenTitan’s first design milestone, D1, its “feature set [must be] finalized, spec complete”. Before it can reach the first verification milestone, V1, its test plan must be written and reviewed.

Writing a requirement

Each requirement is one “shall” sentence with a unique ID, a measurable target, the conditions it applies under, a source and a way to check it. NASA’s guidance bans words such as fast, small, adequate, robust and user-friendly, because nobody can verify them.

  • Weak: “The chip shall have low standby power.”
  • Better: “PWR-005: In standby, with only the always-on domain powered, the chip shall draw at most 2 mW at 25 °C on typical silicon. Verify by analysis before silicon and by test on silicon.”

The better version names its conditions. The always-on domain is the small part of the chip that stays powered while the rest sleeps, such as the circuit that watches for a button press. Typical silicon means an average chip: every chip comes out of the factory slightly faster or slower than the next, and slow or fast ones can draw quite different power. And 2 mW is two thousandths of a watt.

PPA targets, plus cost, schedule and risk

  • Performance is stated as work done: frames per second, images recognized per second, response time, or how many bytes per second can move to and from memory. Clock speed is only one way to get there, and the architecture chooses it.
  • Power is stated for each operating mode: peak, sustained, idle and standby. The sustained limit is the (thermal design power): the power the chip can turn into heat without its transistors going above their maximum safe temperature.
  • Area sets the cost of each chip and has to fit the package, the protective case the chip is mounted in.
  • Cost has a one-time part, called (non-recurring engineering: design work, design-software licenses and the factory’s printing masks), and a per-chip part (silicon, packaging and testing). NRE is spread over every chip sold, so a chip that sells in small numbers must keep it small.
  • Schedule is set by the market: when the product has to be on shelves.
  • Risk grows with every first: a new manufacturing process, a new interface standard, a new architecture. Many teams limit how many firsts one chip can carry.

Executable specs

Prose can be read two ways; a program gives one answer. So teams also write an : software that behaves the way the chip should. It comes in layers: a spreadsheet that estimates power, performance and area; a reference model that computes the right answers; and a performance model that estimates how fast the chip will run. SystemC, a C++ library standardized as IEEE 1666, is widely used to explore which jobs belong in hardware and which in software. It is also used to build virtual platforms, software copies of the whole chip made of , so programmers can start writing the chip’s software before its design exists. Open-source simulators such as gem5 model processor cores, caches (small, fast on-chip memories) and the main memory for architecture studies. The reference model is used again later, as the “golden model” that verification compares the real design against.

Traceability

Every requirement links up to the need it serves and down to the tests that prove it. NASA asks for these links in both directions. OpenTitan’s test plans show one practical form. Each planned test (a “testpoint”) maps to a feature of the design and lists the tests that cover it, and a tool maps the test results back onto the plan as a report. Because the is written from the spec, verification can start at the same time as design, not after it.

Why late changes cost more

In the NASA study, a requirements error that costs 1 unit to fix during requirements costs 21–78 units if it is found during integration and test. On a chip, a change during the spec stage is an edit to a document. Once the design is being written, it means rework and retesting. After the design has gone to the factory, it means new masks and another manufacturing run. So teams declare a and send any later change through a formal change request.

The document stack and its owners

The PRD says what the product must do, not how, and is often written by marketing (sometimes as a separate MRD). The chip requirements spec, sometimes called the product or target spec, is usually owned by the chip architect and the program lead. Below it sit:

  • The : how the chip is split into blocks; the memory map (which address reaches which block); the on-chip interconnect; the clock and power domains (regions with their own clock or their own on/off switch); the boot sequence; and the security architecture.
  • Per-block : pipeline stages, state machines, interfaces with their timing cycle by cycle, and the register map. Registers here are the numbered control and status locations software reads and writes to drive the block.

Keep the register map in one machine-readable file and generate the design code, the verification models and the software’s header files from it, so the three can’t disagree. OpenTitan’s D1 checklist requires exactly that generated collateral alongside a complete spec.

Requirements that survive signoff

Signoff is the final round of checks before the design goes to the factory. A requirement survives it only if it can be checked, and that takes conditions. Transistors vary from chip to chip and slow down when hot or at low voltage, so engineers check designs at named combinations of process, voltage and temperature called . “1.5 GHz” means nothing without the corner it applies at. Likewise, “≤ 5 W” means nothing without a workload.

The verification method decides who owns closing each requirement:

  • Test: running the design in simulation, or measuring real chips.
  • Analysis: calculation by a tool. Examples are , which computes the worst-case delay through every path in the circuit without simulating it, power estimation and thermal simulation.
  • Inspection: looking at the design files, layout or documents.
  • Demonstration: showing the function working.

Requirements that can only be measured on finished chips, such as standby power, need a before-silicon analysis that stands in for the measurement, with explicit margin. Requirements must also be attainable. A clock target that no available cell library (the set of ready-made logic gates for a process) can meet at the chosen voltage is a defect in the spec.

Budgets and margin

Power and area budgets flow down to blocks. At the spec stage the numbers are estimates, scaled from the previous chip or from early models. So teams hold margin back at chip level and release it as better estimates arrive: first from the design code, then from the gate-level netlist, the list of logic gates and their connections that synthesis produces. A block that goes over its allocation either finds savings or negotiates margin from the pool. A pool that is empty before the design code is frozen is an early warning worth acting on.

Traceability in regulated flows

In safety-critical work, traceability is audited. FAA guidance for airborne electronic hardware calls a correct and complete set of requirements the cornerstone of development assurance. At the two most critical design assurance levels (DAL A and B), it asks for a review of traceability between the detailed design and the device requirements. It also asks for a review of the reports from synthesis and place-and-route, the tool steps that turn the design into gates and then into a layout. And it asks for a check that each verification case suits the requirements it traces to, and that those requirements are completely covered.

In cars, ISO 26262 starts from hazard analysis. It sets top-level safety goals, gives each an (automotive safety integrity level, from A to the strictest, D), breaks the goals into safety requirements allocated to hardware and software components, and requires traceability between dependent work products.

Change control

Late changes cost tens to hundreds of times more than early ones. A is really a series of freezes: requirements, then architecture, then the features in the design code. A change request after each one carries an impact analysis across design code, verification, physical layout, software, manufacturing test and schedule. The commonest failure is not a bad decision but a silent one: the design code changes behavior, the spec is never updated, and the test plan keeps checking the old behavior.

requirements specThe chip shall have low standby power.✓one “shall”·a number·conditions·ID + source·how to verifyNot verifiable
1 / 5

“Low” can’t be verified: low compared with what? NASA’s guidance bans words like this.

Building a requirement from a weak one, following NASA’s rules: one “shall”, a number, conditions, an ID with its source, and a way to verify. IDs are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’
Customer needscustomers, studied by marketingPRDproduct marketingChip requirementschip architect, program leadArchitecture specchip architectsMicroarchitecture specseach block’s design engineersone need, followed down:“Record sharp 4K video all day.”
1 / 5

Customer needs: Gathered by talking to customers and studying competitors.

Requirements flow down a chain of documents, each more detailed than the last. Step down to follow one need; the example is illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Think of the chip’s power like a monthly allowance. The total is set by the battery, or by how much heat the device can get rid of. The spec splits that total among the chip’s parts. It keeps a little back for surprises.

Size has a budget too. A chip twice as big costs more than twice as much, because more of the big ones are ruined by specks of dust.

Power and heat

Most of a chip’s power goes into switching. Every time a signal flips between 0 and 1, a tiny amount of electric charge flows into or out of the wires and transistor inputs it drives. That charge-holding ability is called capacitance, and the energy spent moving the charge ends up as heat. The standard formula for this is:

P=αCV2fP = \alpha C V^2 f
  • (activity factor): how often, per clock tick, a signal charges up (goes from 0 to 1). The clock does this every tick (α=1\alpha = 1); a signal that flips every tick charges up only every other tick (α=1/2\alpha = 1/2); typical logic averages around 0.1.
  • CC: the capacitance being switched. Bigger circuits and longer wires mean more.
  • VV: the supply voltage. It is squared, so it matters most.
  • ff: the clock frequency, in ticks per second.

On top of that comes : current that trickles through transistors that are supposed to be off, whether or not anything switches. Leakage grows as the chip heats up.

Faster switching needs a higher voltage, so power grows faster than speed. Here is a worked example. A 20% faster clock that needs 10% more voltage costs 1.2 (for f)×1.1×1.1 (for V2)≈1.451.2\ (\text{for } f) \times 1.1 \times 1.1\ (\text{for } V^2) \approx 1.45 times the switching power: 45% more power for 20% more speed.

The chip’s case and cooling set the sustained limit, the . The spec then sets a for each mode: peak, sustained, idle and standby. Battery devices care most about standby. Data-center chips care most about sustained power and cooling.

Area and the cost of each chip

Chips are made on a , a thin disc of silicon, today usually 300 mm across. Hundreds of copies of a design are printed side by side, and the wafer is then cut into individual chips, each called a die. Tiny flaws, such as a speck of dust, land at random during manufacturing, and one in the wrong place ruins the die it lands on. Three formulas turn these facts into a cost:

  • =wafer cost÷(dies per wafer×yield)= \text{wafer cost} \div (\text{dies per wafer} \times \text{yield}).
  • ≈π(d/2)2/A−πd/2A\approx \pi (d/2)^2/A - \pi d/\sqrt{2A}, for wafer diameter dd and die area AA. The first term is the wafer’s area divided by the die’s area. The second subtracts the partial dies wasted around the round edge.
  • , the share of dies with no fatal flaw, is in the simplest (Poisson) model Y=e−D0AY = e^{-D_0 A}. Here is the average number of fatal defects per cm², so D0AD_0 A is the expected number of defects per die.

A worked example on a 300 mm wafer with D0=0.1D_0 = 0.1 defects per cm²: a 100 mm² die (1 cm²) gives about 640 candidates, of which 90.5% work, so about 579 good dies. Double the area to 200 mm² and you get about 306 candidates at 81.9% yield, or about 251 good dies. At an illustrative $10,000 per wafer, the silicon cost of each good die rises from about $17 to about $40: 2.3 times the cost for 2 times the area.

Choosing a manufacturing process

Chip factories offer several generations of manufacturing process, called and named with a number such as “5 nm”. A newer node packs more transistors into each square millimeter and spends less energy per operation, but each wafer costs more. In one 2020 cost model, the factory’s price for a 300 mm wafer rose from about $1,650 at 90 nm to about $16,988 at 5 nm, and design costs per node grew about 24% a year. The one-time cost also includes the , the stencils used to print each layer. For products made in small numbers, an FPGA avoids that one-time cost altogether; see Where FPGAs win. Some bought-in blocks, such as the analog circuits that drive high-speed connections, come as finished layout for one factory’s process and can’t easily be moved to another. So what can be bought for a node can decide which node is used. Older nodes remain the right answer for many analog, automotive and cost-sensitive chips.

Package and connections

The list of interfaces sets how many signals must leave the chip. Add the connections for power and ground and you have the number of solder balls (or pins) the package needs. The is the case that holds the die, wires it to the circuit board and carries heat away. Its main choices are how the die connects to it (fine wires bonded to the die’s edge, or , with the die face-down on tiny solder bumps), how many balls it has, how well it conducts heat, and what it costs. They are made now because they limit the die size, where the connection pads go, and the TDP. A small die with many connections is : its size is set by the ring of connection pads around its edge, not by the logic inside.

Power budgets that hold up

A useful spec-stage power budget states four things: the operating mode, the workload, the junction temperature (the temperature of the transistors themselves) and the process corner (whether it assumes average, fast or slow transistors). It splits each block’s power into switching and leakage.

Leakage rises steeply with temperature, so the budget at the hot corner is the one that sets TDP. A design with a lot of leakage can even run away thermally if the package is marginal: more heat means more leakage, which means more heat.

Peak current is a separate requirement. When a large block wakes up, the current it draws jumps within nanoseconds. The package’s power connections have inductance, which resists sudden changes in current, so the supply voltage on the die briefly sags. The fixes are more power bumps between die and package and more on the die, small charge reservoirs that cover the sudden demand. All of this has to be planned, and it can bite at mode changes even when average power is fine.

Battery products add a standby budget. It usually dictates the (the part that never sleeps) and the architecture (switches that cut the supply to idle blocks so they stop leaking). Those decisions later become the chip’s power-intent file.

Cost beyond the die

Unit cost is the die cost, plus testing and packaging, plus the one-time costs (NRE) divided by the number of chips sold. Large dies run into two walls: low yield, and the , the largest area the lithography machine can print in one exposure. That is why big AI accelerators are split into , several smaller dies in one package. Chiplets trade die yield for the cost of die-to-die interfaces, advanced packaging, and testing each die before assembly to make sure it is a .

Node economics

The price per wafer is only half the story; what matters is the price per transistor. CSET priced a chip with a fixed number of transistors at each node. The factory’s price fell from $2,433 at 90 nm to $233 at 7 nm, then edged up to $238 at 5 nm, because at the newest node the wafer price rose about as fast as density. Cost per transistor no longer falls automatically, so the case for a new node rests on energy per operation, performance, or the sheer transistor count the product needs. Schedule risk adds to the cost: early in a node’s life the defect density is higher, and the cell libraries, process design kits and bought-in blocks are still changing.

per block (dark: leakage)CPU cluster1.70NPU1.15ISP + video0.63SRAM + wiring0.42LPDDR5 + PHY0.45PCIe, USB, MIPI0.20Always-on0.05Totalbudget 5.00 W4.60 W + 0.40 W margin (dashed)
Estimate

Spec-stage estimate: 4.60 W across seven blocks, with 0.40 W (8%) held back as margin. Tap a block.

The illustrative VX-200 power budget from “In practice”, split by block, with margin held at chip level. Switch to the later gate-level estimate.Share freely with credit: ‘Figure from chipfieldguide.com’
Dies per wafer≈ 640Yield90.5%Good dies≈ 579Cost per good die$17drawn: 584/648whole dies clean
Defect density D₀

100 mm²: ≈ 640 dies per wafer × 90.5% yield ≈ 579 good dies, $17 each (1.00× the 100 mm² die for 1.0× the area).

A 300 mm wafer cut into dies, with fatal defects scattered at random. Readouts use the formulas above (Poisson yield); the $10,000 wafer price is illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Some musts have nothing to do with speed. A car chip must work in freezing winters and next to a hot engine. A payment chip must guard its secrets. And every chip must be testable, so the factory can catch broken ones. These go into the spec at the start, because adding them later is like adding a basement to a finished house.

Interfaces and standards

An interface is an agreed set of rules for connecting two devices: the wires, voltages, timing and message formats. The spec lists every external interface with its standard, version, width and speed. Common ones are DDR or LPDDR memory (standards from JEDEC), PCIe for expansion cards and other chips (PCI-SIG), USB (USB-IF), Ethernet (IEEE 802.3), and links to cameras and displays. The spec also says how the chip will prove it follows each standard. USB-IF, for example, adds products that pass its certification tests to a public Integrators List, and only certified products whose makers hold a trademark license may use the USB logo. That makes a line in the schedule.

Most teams license these blocks rather than design them. An is a ready-made, already-tested design block that shortens the design cycle. Interface IP typically pairs a digital controller with a , the analog circuit that drives and receives the signals on the wires. The PHY is delivered as finished layout for one factory’s process. The list of which blocks to make and which to buy is one of the spec’s outputs.

Manufacturing test requirements

Every chip is tested on a machine before it ships, to catch the ones with manufacturing defects. No test catches everything, so the spec sets an outgoing-quality target, usually in (DPPM), shipped chips that are faulty. That target implies a target. Coverage is measured by imagining simple faults, such as one wire stuck at 0, and counting how many of them the tests would detect. The number of bad chips that escape depends on both the yield and the coverage, and more than 500 per million is considered unacceptable for commercial chips.

The coverage target in turn sets the test time, the number of pins reserved for testing, and the built-in test circuits. The main one is the . A chip stores its working data in millions of , tiny circuits that each hold one bit. In test mode, scan links them into long chains so a tester can load and read them directly. Others are a JTAG test port (a standard four- or five-wire test and debug connection) and self-test circuits for the on-chip memories. The design for test stage builds them.

Safety

For a car, follows the standard ISO 26262. It starts with hazard analysis: what could go wrong, and how badly. Each hazard is rated for severity, for how often drivers are exposed to the situation, and for how well a driver could control it. The ratings give an (automotive safety integrity level), from A to the strictest, D. Safety requirements are then allocated to hardware and software components. On a chip, that means requirements such as error-correcting codes on memories (extra bits that let the chip detect and fix a flipped bit), pairs of processors that run the same program and compare results, or self-tests that run while the car is in use. Automotive chips also name a temperature grade. AEC-Q100 grade 0 covers an air temperature around the part of −40 °C to +150 °C, and grade 3 covers −40 °C to +85 °C.

Security

Security requirements come from a . It names the assets to protect (secret keys, the chip’s software, unique device IDs), the attackers, and the ways in. OpenTitan’s public threat model counts the test and debug interfaces as a way in. It also lists physical attacks such as , which reads secrets from the chip’s power use, timing or radio emissions, and fault injection, which glitches the voltage, the clock or a laser onto the chip to make it misbehave.

Where industry requirements enter

Each market’s rules become chip requirements here: standby power for battery devices, memory bandwidth and TDP for AI accelerators, ASIL and temperature grade for cars, ultra-low power for medical implants, radiation tolerance for space, trusted supply chains for defense, predictable response time for financial trading. The industry pages cover each.

Interfaces drive the floorplan

Interface choices fix bandwidth and pin budgets early. Take a 32-bit memory interface at 6,400 MT/s (million transfers per second). Each transfer moves 32 bits, or 4 bytes, so the peak is 4×6,400 million=25.6 GB/s4 \times 6{,}400\,\text{million} = 25.6\,\mathrm{GB/s}. The spec must also say what fraction of that peak the workload needs to sustain, because that sets how many requests the memory controller must queue and how much buffering the chip needs.

Each PHY sits on the die edge and brings its own power supplies, electrostatic-discharge (ESD) protection and assignment of package balls. So the number of PHYs constrains the die’s shape and the package from day one. Compliance programs also require test modes, such as loopbacks that send a link’s output straight back into its input, which must be in the design. Pick IP with a compliance record on your node, because hard PHYs do not move between processes.

Test, safety and security interact

Manufacturing-test targets follow from the quality target. Defect level, the fraction of bad chips that pass test, is a function of yield and fault coverage. So a DPPM goal translates into coverage numbers for two fault models: , where a node is stuck at 0 or 1, and , where a node switches too slowly.

Safety adds test in the field. (logic built-in self-test) puts a pattern generator and a checker on the chip so the logic can test itself while the car runs. Safety also adds metrics with hard targets that depend on the ASIL. Diagnostic coverage measures how many faults a safety mechanism catches. The single-point fault metric must reach at least 99% for ASIL D.

Security pulls the other way. Scan chains and JTAG give direct access to internal registers, and MITRE’s catalog of hardware weaknesses has a whole category for debug and test problems, next to power, clock, thermal and reset concerns. The usual answer is a set of life-cycle states. In OpenTitan, for example, debug and scan-based test are enabled during manufacturing and disabled in production, moving between states requires a token, and the design ensures that states with debug access cannot expose the device’s secrets in normal use. A spec that wants high coverage, self-test in the field and protected keys must define these states, how an unlock is authorized and how secrets are handled during test, before architecture starts.

dieCPUCPU copySRAM + ECCscan chainKeyssealedJTAGDDR controllerPCIe controllerDDR PHYPCIe PHYTester+150 °C−40 °CAEC grade 0
Life-cycle state

Production state: debug and scan are disabled. Unlocking needs an authorized token. Tap any part.

Where interface, test, safety and security requirements land on a generic chip (a sketch, not to scale). Tap a part; switch the life-cycle state.Share freely with credit: ‘Figure from chipfieldguide.com’

Play product planner. Sliders set the chip’s speed, its number of cores (processors), and its power allowance. The starting chip just fits a 5-watt allowance. Turn up the speed and watch the bar shoot past the line. Then try more cores at a lower speed. You get the same work done for less power, but the chip gets bigger.

You are setting the targets for a small processor chip. Set the clock frequency (0.5–3.0 GHz), the number of cores, meaning processors (1–8), and the power budget (1–15 W). The readouts show the supply voltage the chosen clock needs, and the power as a stacked bar against the budget. The bar splits power into dynamic (switching) power, leakage and “uncore”: the shared parts outside the cores, which stay the same whatever you choose. There is also the throughput in billions of simple instructions per second (each core does one per clock tick), the die area, how many dies fit on a 300 mm wafer, and a verdict.

The default, 4 cores at 1.5 GHz, uses 4.93 W and fits the 5 W budget, with a throughput of 6.0 and an area of 14 mm². Push the clock to 2.0 GHz and power reaches 7.50 W, over budget. Now try 6 cores at 1.0 GHz: 4.29 W for the same 6.0 throughput. Watch the area and dies-per-wafer readouts as you add cores.

The model is illustrative but exact, and the Expert view shows each formula with the current numbers substituted:

  • supply voltage V=0.6+0.2fV = 0.6 + 0.2 f (ff in GHz);
  • switching power per core 0.8 V2f0.8\,V^2 f, and leakage per core 0.15 V0.15\,V;
  • a fixed 0.5 W ;
  • area 2.0 mm² per core plus 6.0 mm²; and
  • dies per wafer ≈π⋅1502/A−π⋅300/2A\approx \pi \cdot 150^2/A - \pi \cdot 300/\sqrt{2A}.

The straight-line link between voltage and frequency is a stand-in for the alpha-power law (see “Under the hood”). In that law, a gate’s delay is roughly the charge it must move, CVC V, divided by the current its transistors can drive. That current grows as (V−VT)α(V - V_{\mathrm{T}})^{\alpha}, where VTV_{\mathrm{T}} is the voltage at which the transistors start to conduct.

Compare the default with 6 cores at 1.0 GHz. The supply drops from 0.9 V to 0.8 V, and that is where the saving comes from. Then find the highest throughput that fits in 5 W. Note which term dominates there (switching, leakage or uncore), and what it costs in dies per wafer.

Loading simulation…

At a spec review, engineers walk through three things.

  1. The list of musts. Each row has an ID, the promise, a number and how it will be checked. Reviewers hunt for rows that are vague, impossible or missing.
  2. The power budget. A table splits the total power among the chip’s parts. If the parts already add up to more than the total, the plan must change now.
  3. The test report. Later on, a report lists each must next to its tests and whether they pass. A must with no test is a red flag.

Below is an illustrative excerpt from the requirements for a made-up smart-camera chip, the VX-200. It is a system on a chip (): one chip that holds the processors, an AI accelerator (the NPU), the video circuits and the memory controllers. Each line has an ID, a single “shall”, a target with its conditions, a way to check it, and the item in the product requirements that it came from. The notes explain the abbreviations.

The four blocks below follow one made-up system on a chip, the VX-200 smart-camera chip, from requirements to evidence: the requirements table, a spec-stage power budget, an excerpt from the test plan that ties tests to requirement IDs, and a requirements trace report from late in the project. All numbers are illustrative.

requirementcovered bystatusPERF-001CPU ≥ 1.5 GHzPERF-0044K at 30 frames/sPWR-002≤ 5.0 W on W1IF-003PCIe Gen4 ×2DFT-001≥ 99% stuck-atSEC-004JTAG lockedENV-001−40 to 105 °C
1 / 3

Spec stage: requirements with IDs and targets. Nothing is linked to them yet. Tap a row.

Traceability over an illustrative project: requirements, then tests, then results. Step through; tap a row.Share freely with credit: ‘Figure from chipfieldguide.com’
vx200_requirements.txt (excerpt, illustrative)text
# VX-200 smart-camera SoC: chip requirements specification, rev 0.7 (illustrative excerpt)
# Method: T = test (simulation or silicon), A = analysis, I = inspection, D = demonstration

ID        Requirement: the SoC shall ...                  Target and conditions                        Method  Source
--------  ----------------------------------------------  -------------------------------------------  ------  ------
PERF-001  run the CPU cluster at its target clock         >= 1.5 GHz at slow corner, 0.81 V, 105 °C    A, T    PRD-12
PERF-004  encode 4K video in real time                    >= 30 frames/s, 3840 × 2160                  T       PRD-03
PWR-002   stay inside the package thermal envelope        <= 5.0 W on workload W1, Tj <= 105 °C        A, T    PRD-20
PWR-005   draw little power in standby                    <= 2 mW, always-on domain only, 25 °C        A, T    PRD-21
AREA-001  meet the die-cost target                        <= 25 mm² die, <= 400 package balls          I       BIZ-04
IF-003    provide a PCIe root port                        Gen4, 2 lanes, passes standard compliance    T       PRD-15
IF-007    support external DRAM                           LPDDR5, 32-bit, 6400 MT/s                    T, A    PRD-16
DFT-001   reach manufacturing test coverage               >= 99% stuck-at, >= 90% transition           A       QA-02
SEC-004   keep debug locked in production                 JTAG closed unless authenticated unlock      T, I    TM-07
SAF-002   correct single-bit errors in on-chip SRAM       SECDED ECC on every SRAM over 4 KB           I, T    SAF-01
ENV-001   operate across the junction temperature range   -40 °C to 105 °C                             T       PRD-30
  1. 1L2The four classic ways to verify a requirement. Every requirement must name at least one, or nobody owns proving it.
  2. 2L6A clock target needs conditions: slow-corner transistors, the lowest allowed supply (0.81 V) and the hottest temperature. Before silicon it is checked by analysis (static timing analysis, which computes worst-case delays); after silicon, by testing real chips.
  3. 3L7Performance stated as work done: 30 frames per second of 4K video. How many cores and what clock to use are the architecture’s job, not the spec’s.
  4. 4L8A power limit needs a named workload (W1) and a temperature. Tj is the junction temperature, the temperature of the transistors themselves. The per-block split lives in the power budget (next block).
  5. 5L9Standby power shapes how the chip is split into power regions: a small always-on region, and switches that cut power to everything else.
  6. 6L10Die area and the number of package balls (the solder connections to the board) come from the business case (BIZ-04), not from engineering preference. Checked by inspecting the floorplan and the package drawing.
  7. 7L11PCIe is a standard high-speed link to other chips; a root port is the end that runs the link. Passing the standard’s compliance program pushes the team to license a controller and PHY that have already passed.
  8. 8L12LPDDR5 is a memory standard. 32 bits at 6,400 million transfers per second is 25.6 GB/s at peak. The memory PHY also sets ball count and how much of the die edge is used.
  9. 9L13The share of modeled manufacturing faults the factory test must detect: wires stuck at 0 or 1, and signals that switch too slowly (transition faults). The targets come from the outgoing-quality goal and shape the test circuitry.
  10. 10L14JTAG is the standard test and debug port. This line traces to item 7 of the threat model, not to a customer request: security requirements have their own source.
  11. 11L15SRAM is on-chip memory. SECDED ECC (single-error-correct, double-error-detect error-correcting code) adds check bits to each word. A safety mechanism from the safety analysis, checked by inspecting the design and by deliberately injecting errors.

The power budget turns PWR-002 into a limit for each block. It is a spec-stage estimate, scaled from a previous chip, so it keeps some margin in reserve. “Dynamic” is switching power; “leakage” is the power drawn even when nothing switches.

vx200_power_budget.txt (illustrative)text
# VX-200 power budget, workload W1 (4K30 encode + NPU object detection), Tj = 105 °C, typical silicon
# Spec-stage estimate, rev 0.7 (illustrative)

Block                          Dynamic (W)   Leakage (W)   Total (W)   Share
CPU cluster (4 cores)                 1.45          0.25        1.70     37%
NPU                                   1.00          0.15        1.15     25%
ISP + video encoder                   0.55          0.08        0.63     14%
Interconnect + on-chip SRAM           0.30          0.12        0.42      9%
LPDDR5 controller + PHY               0.40          0.05        0.45     10%
PCIe, USB, MIPI PHYs + I/O            0.16          0.04        0.20      4%
Always-on, PLLs, sensors              0.03          0.02        0.05      1%
---------------------------------------------------------------------------
Estimated total                       3.89          0.71        4.60
Margin held at chip level                                       0.40   (8% of budget)
Budget (PWR-002)                                                5.00
  1. 1L1Workload, temperature and kind of silicon are all stated. Leakage rises steeply with temperature, so the hot case is the one that sets the thermal limit.
  2. 2L5The biggest consumer gets the most scrutiny. Its number came from the previous chip, scaled for the new process and clock, which is the usual spec-stage method.
  3. 3L6The NPU is the AI accelerator. It does one kind of math very efficiently.
  4. 4L7The ISP (image signal processor) cleans up raw camera data before the encoder compresses it.
  5. 5L8On-chip memory leaks a lot: 0.12 W of 0.42 W. Large memories are a common place to use low-leakage transistors or modes that keep data with most of the power off.
  6. 6L9PHY power is set mostly by the interface standard and the bought-in block, and shrinks little with the logic. Treat it as fixed.
  7. 7L11PLLs generate the chip’s clocks. The always-on blocks are tiny, but they set standby power.
  8. 8L13Leakage is about 15% of the total at the hot corner. At room temperature the split looks very different.
  9. 9L14Margin is released as better power estimates arrive from the design code and then from the gate-level netlist. An empty margin before the design is frozen is an early warning.

Traceability in practice: the test plan names the requirement each planned test covers. This excerpt follows the shape of OpenTitan’s test plans, written in Hjson (a human-friendly form of the JSON data format). Each testpoint has a name, a description of its goal, stimulus and check, a stage by which it must be done, and a list of tests.

vx200_soc_testplan.hjson (excerpt, illustrative)text
// Modeled on OpenTitan's testplan format; requirement IDs added to each description
{
  name: vx200_soc
  testpoints: [
    {
      name: pcie_rootport_link_gen4
      desc: '''Covers IF-003. Goal: link trains to Gen4, 2 lanes, from reset.
            Stimulus: PCIe endpoint verification IP; every supported width and speed.
            Check: link reaches L0 at 16 GT/s on 2 lanes; config reads return expected IDs.'''
      stage: V2
      tests: ["pcie_link_gen4_x2", "pcie_link_downtrain"]
    }
    {
      name: debug_jtag_locked_prod
      desc: '''Covers SEC-004. Goal: JTAG stays closed in the production life-cycle state.
            Stimulus: drive the TAP in PROD state with and without a valid unlock token.
            Check: no TAP access without the token; access granted with it.'''
      stage: V2
      tests: ["jtag_lock_prod", "jtag_unlock_token"]
    }
    {
      name: sram_ecc_single_bit
      desc: '''Covers SAF-002. Inject single- and double-bit errors into every SRAM.
            Check: single-bit errors corrected and logged; double-bit errors flagged.'''
      stage: V2
      tests: ["sram_ecc_inject"]
    }
  ]
}
  1. 1L1The requirement ID in each description is what lets a script build the trace report.
  2. 2L6Testpoint names follow feature_subfeature naming, so the plan reads like a feature list.
  3. 3L8Verification IP is a ready-made model of the device at the other end of the link, bought or built so the test can talk PCIe to the design.
  4. 4L9Checks are concrete and observable: L0 is PCIe’s normal working state, and 16 GT/s (billion transfers per second) is the Gen4 speed. “PCIe works” is not a check.
  5. 5L10Stage says when this testpoint must be done (V2: testing complete).
  6. 6L11One testpoint can need several tests, and one test can serve several testpoints.
  7. 7L16A security requirement tested like any other, including the negative case: access must be refused without the token. The TAP is the logic behind the JTAG port.

A script joins the requirements, the test plans, the results of the automated test runs, and the timing and power reports into a trace report. This is what a program review looks at. In airborne hardware it is the evidence reviewers use to confirm that every requirement is completely covered by verification. Two terms in the report: WNS (worst negative slack) is how far the slowest path misses its timing target, and ATPG coverage is the share of faults detected by the patterns from automatic test-pattern generation, the tool that creates the factory tests.

vx200_trace_report.log (illustrative)log
Requirements trace report, VX-200 spec rev 0.7, regression 2026-09-28 (illustrative)
REQ       Method  Covered by                          Evidence          Status
PERF-001  A, T    sta_signoff_ss_0p81v_105c           WNS -12 ps        OPEN
PERF-004  T       video_enc_4k30 (emulation)          31.2 frames/s     PASS
PWR-002   A, T    power_est_w1 (gate-level)           5.21 W            FAIL  budget 5.00 W
PWR-005   A, T    power_est_standby                   1.6 mW            PASS  silicon test pending
AREA-001  I       floorplan_area                      24.1 mm²          PASS
IF-003    T       pcie_rootport_link_gen4             2/2 tests         PASS
IF-007    T, A    ddr_training, ddr_bandwidth         2/2 tests         PASS
DFT-001   A       atpg_coverage                       98.6% stuck-at    FAIL  target 99%
SEC-004   T, I    debug_jtag_locked_prod              2/2 tests         PASS
SAF-002   I, T    sram_ecc_single_bit                 1/1 tests         PASS
ENV-001   T       (none)                              -                 UNTRACED
Summary: 11 requirements: 7 pass, 1 open, 2 fail, 1 untraced
  1. 1L1A regression is the full automated test suite, rerun regularly (often nightly) so that new changes can’t silently break old features.
  2. 2L3Closed by analysis: static timing analysis at the spec’s slow corner (ss), 0.81 V and 105 °C. The slowest path is still 12 picoseconds too slow, so the requirement stays open, not failed, while the team fixes it.
  3. 3L4Measured on emulation, special hardware that runs the design fast enough to process real video. Silicon will confirm.
  4. 4L5Gate-level power is over budget. Options: spend the 0.40 W margin (already gone), lower the voltage or clock through a spec change, add power gating, or revisit the workload definition. Each is a change request.
  5. 5L6Analysis passes, but the requirement also names test, so it stays partly open until real chips can be measured.
  6. 6L10Test coverage short of target. Usually fixed by adding test points (extra logic that makes hard-to-reach nodes controllable or visible) or new test modes, which costs area and schedule.
  7. 7L13No testpoint at all: nobody has planned how to prove the temperature range. Add analysis at the temperature corners now, and a characterization test on real chips.
  8. 8L14The summary goes to the program review. Two fails and an untraced line block tapeout signoff.
  • Vague wishes. Nobody can test “fast” or “low power,” so nobody can say whether the chip meets them.
  • Late changes. A feature added halfway through ripples through everything. It can delay the chip by months.
  • Forgotten musts. Leave out factory testing or security, and the gap shows up at the end. That is when it is hardest to fix.
  • Wishful numbers. A power or cost target that only works if everything goes right usually fails once the real chip is made.
  • Requirements nobody can check. “Low latency” (fast response), or “≤ 5 W” with no workload. The fix is a review checklist built from NASA’s rules: one “shall”, a number, conditions, and a way to check it.
  • Saying how instead of what. A PRD that dictates the design leaves the engineers no room to find the best solution. Requirements should state the need and leave the solution open.
  • Requirements with no test. They reach the factory unproven. A trace report that maps every test result back to the plan, produced with every run of the test suite, catches them.
  • Optimistic power. Budgets that assume an average chip at room temperature miss the extra leakage of a hot chip. State the temperature and the kind of chip in the budget.
  • A bought-in block that doesn’t exist for the chosen process. Analog interface blocks are built for one factory’s process. Finding out, after the process is chosen, that the PCIe block isn’t available for it forces a redesign or a change of process.
  • Spec drift. The design changes behavior and the document is never updated, so verification keeps checking the old behavior. A frozen baseline with change control prevents it.
  • Corner mismatch. The spec states a frequency at one voltage and temperature, and signoff checks timing at another, so passing signoff doesn’t prove the requirement. Write the signoff corners into the requirement, so that timing closure at those corners is the evidence.
  • TDP without a peak-current requirement. Average power fits the package, but the sudden jump in current when blocks switch modes makes the supply voltage sag (in engineering shorthand, a di/dt problem: current changing fast). Peak current needs its own requirement, which feeds the number of power bumps and the amount of on-die decoupling capacitance.
  • A yield model too kind for a big die. Using the defect density of a mature process for a young one, or the wrong yield model, changes cost per good die by tens of percent at large area. Run both the Poisson and negative binomial models (see “Under the hood”), and use the factory’s numbers where you have them.
  • Security bolted on. A debug port left open in production, or secrets readable through scan, are catalogued hardware weaknesses. Test and debug interfaces belong in the threat model from the start.
  • Safety metrics found late. The automotive hardware metrics have hard targets, such as a single-point fault metric of at least 99% for ASIL D: at most 1% of faults that would directly break a safety goal may go uncaught by a safety mechanism. If the fault campaigns (simulations that inject faults and count how many are caught) run only after the design is frozen and fall short, adding safety mechanisms then costs area, power and re-verification.
  • The executable model drifts. The reference model and the design code drift apart. The checker that compares them then starts treating design bugs as expected behavior. Version the model with the spec and review changes to both together.
  • Too many firsts. A new process, a new interface generation and a new architecture on one chip multiply schedule risk. Late requirement changes then land on a project with no slack, which is where the NASA cost multipliers come from.
✗ Can’t be checked“The chip shall consume ≤ 5 W.”✗ How, not what“The chip shall use 4 CPU coresat 2 GHz.”✗ No testENV-001 −40 to 105 °Ccovered by: (none)✗ Optimistic powerPower budget: 5 W at 25 °C, assumingan average chip.
Show

Four spec lines, each with a common flaw. Tap one to see what is wrong, or show the fixes.

A spec review, line by line: four common flaws and their fixes. Lines are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

This section covers the models behind the numbers in a spec: why power, not transistor count, limits modern chips, how voltage and frequency trade, and how die size turns into cost. A transistor here is just a voltage-controlled switch. A small voltage on its control input, the gate, turns it on or off, and the threshold voltage, VTV_{\mathrm{T}}, is roughly where it flips.

Dennard scaling

comes from a 1974 paper. Dennard and colleagues showed what happens if you shrink a transistor’s dimensions and its supply voltage by the same factor κ\kappa (and raise the doping, the impurities that set its electrical behavior, by κ\kappa). Its current and capacitance both scale by 1/κ1/\kappa. The delay of each circuit scales by 1/κ1/\kappa, so it gets faster. The power of each circuit scales by 1/κ21/\kappa^2. The area of each circuit also falls by 1/κ21/\kappa^2, so power per square millimeter stays constant. In their words: “even if many more circuits are placed on a given integrated circuit chip, the cooling problem is essentially unchanged.” For decades, Moore’s law plus Dennard scaling gave exponential performance gains. Each new process brought more transistors, faster transistors and more efficient transistors at once.

Why it broke, and dark silicon

Dennard scaling requires the to fall along with the supply. But a transistor below threshold is not fully off. A small subthreshold current still flows, and it grows about tenfold for every 100 mV the threshold drops (at a subthreshold slope near 100 mV per decade). Harris’s 65 nm example gives off-currents of 1, 10 and 100 nA per µm of transistor width at VTV_{\mathrm{T}} of 0.5, 0.4 and 0.3 V. Multiply that by billions of transistors and the leakage becomes unaffordable. So the threshold stopped falling, supply voltage scaling slowed, and energy per transistor stopped falling as fast as transistor counts rose. Chips became limited by power.

Esmaeilzadeh and colleagues modeled what this means for multicore chips. They projected that 21% of a fixed-size chip must be powered off at 22 nm, and more than 50% at 8 nm, with only a 7.9× average speedup on parallel workloads through 2024. That powered-off fraction is . For a spec writer, it means the binding constraint on a large digital chip is usually the power budget, not the transistor budget. That is why modern SoCs fill their area with specialized accelerators that sit idle most of the time.

Switching power and leakage

Where P=αCV2fP = \alpha C V^2 f comes from: each time a node charges up and back down, it draws CV2C V^2 of energy from the supply. A node that does this fsw=αff_{\mathrm{sw}} = \alpha f times per second therefore dissipates αCV2f\alpha C V^2 f. A clock has α=1\alpha = 1, a signal that switches once every cycle has α=1/2\alpha = 1/2, and typical logic averages around 0.1. Short-circuit current, which flows for a moment while both transistors of a gate are partly on, usually adds less than 10% of switching power. Static power is the sum of four leakage currents times the supply: (Isub+Igate+Ijunction+Icontention) V(I_{\mathrm{sub}} + I_{\mathrm{gate}} + I_{\mathrm{junction}} + I_{\mathrm{contention}})\,V. These are subthreshold, gate-insulator, junction and contention currents.

At the spec stage, each input comes from somewhere different. CC comes from the switched capacitance per mm² of a previous design, scaled for the new process. α\alpha comes from traces of the target workloads. Leakage comes from the cell library’s data at the hot corner. The least certain term is α\alpha, which is why budgets state their workload.

The voltage–frequency relationship

Sakurai and Newton’s alpha-power law (their Berkeley report writes the exponent as n) models a transistor’s saturated current as proportional to (VGS−VT)α(V_{\mathrm{GS}} - V_{\mathrm{T}})^{\alpha}. Here VGSV_{\mathrm{GS}} is the gate voltage and VTV_{\mathrm{T}} the threshold. The exponent α\alpha falls from 2 toward 1 as velocity saturation grows: in short transistors, the electrons hit a speed limit. The law gives a gate’s delay in terms of CLVDDC_{\mathrm{L}} V_{\mathrm{DD}} (the charge to move) divided by that current. Maximum frequency therefore scales roughly as (V−VT)α/V(V - V_{\mathrm{T}})^{\alpha} / V. Well above threshold this is close to linear in VV, which is why the simulator can use V=0.6+0.2fV = 0.6 + 0.2 f. Near threshold, it collapses.

Combine that with P∝V2fP \propto V^2 f, step by step:

  1. If frequency is roughly proportional to VV over the operating range, then VV rises with ff.
  2. So switching power, V2fV^2 f, rises roughly with the cube of frequency.
  3. Energy per operation, V2V^2, rises with the square.
  4. Throughput, meanwhile, grows linearly with the number of cores.

So more cores at a lower frequency and voltage can match the throughput of fewer, faster cores for less power: “wide and slow.” In the simulator, 4 cores at 1.5 GHz (0.9 V) need 4.93 W, while 6 cores at 1.0 GHz (0.8 V) deliver the same 6.0 throughput for 4.29 W. The costs are more area (18 mm² against 14 mm²), more leakage from more transistors, and workloads that must actually split across cores.

The trade has a floor. Dreslinski and colleagues report that lowering the supply from a nominal 1.1 V into the near-threshold range of 400–500 mV gives up to 10× better energy efficiency, at about 10× lower performance. In their 45 nm example, a standard delay benchmark (the fanout-of-four inverter delay) was 10× slower at 400 mV. Below threshold, delay grows exponentially, so leakage energy per operation rises, and total energy reaches a shallow minimum. Commercial parts typically keep the supply above about 70% of nominal because of robustness and performance concerns. designs therefore need massive parallelism to recover throughput, and they suit power-capped products such as sensors. The same physics underlies (dynamic voltage and frequency scaling), which lowers voltage and clock together whenever full speed isn’t needed.

Dies per wafer

The standard estimate is:

DPW≈π(d/2)2A−πd2A\mathrm{DPW} \approx \frac{\pi (d/2)^2}{A} - \frac{\pi d}{\sqrt{2A}}

The first term is the wafer’s area divided by the die’s area. The second approximates the loss at the edge. The wafer’s circumference, πd\pi d, divided by the diagonal of a square die, 2A\sqrt{2A}, roughly counts the partial dies that straddle the edge. For a 300 mm wafer, a 100 mm² die gives about 640 and a 600 mm² die about 91. Production estimates also subtract the scribe lanes where the wafer is sawn, an unusable rim at the edge, and test sites.

Yield models

With mean defect density D0D_0 and die area AA, the expected number of fatal defects per die is D0AD_0 A. The models below differ in how they treat the fact that defects are not spread evenly: they vary from wafer to wafer and bunch together. Leachman’s notes give the standard family:

  • Poisson: Y=e−D0AY = e^{-D_0 A}. Assumes defects land independently, like raindrops. Accurate for small dies (about 0.25 cm² or less) and D0AD_0 A below 1. Because the per-layer yields multiply, each layer’s defect density can be separated.
  • Murphy: Y=(1−e−D0AD0A)2Y = \left(\frac{1 - e^{-D_0 A}}{D_0 A}\right)^2. Lets the defect density itself vary across wafers, with a triangular distribution.
  • Seeds: Y=1/(1+D0A)Y = 1/(1 + D_0 A). Assumes the defect density follows an exponential distribution, which means strong bunching.
  • Negative binomial: Y=(1+D0A/α)−αY = (1 + D_0 A/\alpha)^{-\alpha}. The defect density follows a gamma distribution with a cluster parameter α\alpha (a different α\alpha from the activity factor). With α\alpha of 10 or more it is essentially Poisson, α=5\alpha = 5 closely approximates Murphy, and α=1\alpha = 1 gives Seeds.

Bunching puts several defects on the same dies and leaves others clean, so Poisson underestimates the yield of large dies. Baas’s course notes use the negative binomial with α≈3\alpha \approx 3, and 0.5–1 defects per cm² as typical values for CMOS processes. The models agree for small dies and diverge for large ones. At D0=0.2D_0 = 0.2 defects per cm², a 100 mm² die yields 82–83% under every model. A 600 mm² die (D0A=1.2D_0 A = 1.2) yields about 30% under Poisson, 34% under Murphy, 36% under the negative binomial with α=3\alpha = 3, and 45% under Seeds.

0%25%50%75%100%0200400600800 mm²Poisson30% $366Murphy34% $325Neg. bin. α=336% $303Seeds45% $243yield and cost per good die at the dashed line

D₀A = 1.20 defects per die, 91 dies per wafer. Yield 30–45% across the models: cost per good die differs by 51% between them.

Die yield against area for the four models in the text, negative binomial at α = 3. They agree for small dies and spread apart for large ones; Poisson is the pessimist. Cost per good die at an illustrative $10,000 per 300 mm wafer.Share freely with credit: ‘Figure from chipfieldguide.com’

Cost per good die

Cost per good die is:

wafer costDPW×Y\frac{\text{wafer cost}}{\mathrm{DPW} \times Y}

Dies per wafer falls slightly faster than 1/A1/A, and yield falls exponentially in AA, so cost per good die grows faster than area. Use the numbers above with an illustrative $10,000 wafer and D0=0.2D_0 = 0.2 per cm²:

  • A 100 mm² die costs about $19.
  • A 600 mm² die costs about $300 under the negative binomial (α=3\alpha = 3), and about $370 under Poisson.

That is roughly 16–19 times the cost for 6 times the area. Full unit cost adds testing and packaging, plus the one-time costs divided by volume. These sensitivities are why the spec fixes a die-size class early, and why very large designs move to chiplets.

Novice · 0 of 5 correct
  1. Q1A 300 mm wafer has 0.1 fatal defects per cm² on average. Growing a chip from 100 mm² to 200 mm² changes the number of working chips per wafer from about 579 to about…

  2. Q2Switching power follows P=αCV2fP = \alpha C V^2 f. A block’s clock is made 20% faster, and its supply voltage must rise 10% so its circuits can keep up. By roughly how much does its switching power grow?

  3. Q3What does a chip’s TDP (thermal design power) describe?

  4. Q4A requirement reads “PWR-002: The chip shall consume ≤ 5.0 W.” What is the most important thing missing?

  5. Q5A traceability report shows a requirement with no linked test. What does that mean?

Sources

Show Hide 29 sources
  1. Product requirements documentWikipedia contributors · WikipediaA PRD says what a product should do, not how; often written by a marketing department (then also called an MRD); the maker breaks it down into a functional specification.
  2. NASA Systems Engineering Handbook, Appendix C: How to Write a Good RequirementNASA · National Aeronautics and Space Administration“Shall” marks a requirement; one thought per requirement; requirements must be verifiable by test, demonstration, inspection or analysis; bidirectional traceability; list of unverifiable words to avoid.
  3. NASA Systems Engineering Handbook, Appendix D: Requirements Verification MatrixNASA · National Aeronautics and Space AdministrationA matrix that defines how every “shall” is verified, with a unique identifier and source document for each.
  4. Error Cost Escalation Through the Project Life CycleJonette M. Stecklein, Jim Dabney, Brandon Dick, Bill Haskins, Randy Lovell, Gregory Moroney · NASA Technical Reports Server (JSC-CN-8435) · 2004Relative cost of fixing a requirements error: 1 in requirements, 3–8 in design, 7–16 in build, 21–78 in integration and test, 29 to over 1,500 in operations.
  5. Hardware Development StageslowRISC and the OpenTitan contributors · OpenTitan documentationD1 requires “feature set finalized, spec complete” and generated register collateral; V1 requires a reviewed testplan; V2 requires 90% code and functional coverage; V3 100% with waivers.
  6. Testplanner tool and testplan formatlowRISC and the OpenTitan contributors · OpenTitan documentationHjson testpoints with name, stage, desc and tests; each testpoint maps to a design feature; simulation results are mapped back to testpoints in a report.
  7. Lightweight Threat ModellowRISC and the OpenTitan contributors · OpenTitan documentationAssets, attacker profiles and attack surfaces (including test and debug interfaces) used to produce security requirements; side-channel analysis and fault injection as physical attacks.
  8. Device Life CyclelowRISC and the OpenTitan contributors · OpenTitan documentationLife-cycle states enable or disable debug and DFT (scan) functions; production states disable both; transitions need tokens; states with invasive debug cannot expose secrets in mission mode.
  9. CWE VIEW-1194: Hardware DesignMITRE · Common Weakness EnumerationHardware weakness categories, including debug and test problems (JTAG and scan chains, improper access control on debug and test interfaces) and power, clock, thermal and reset concerns.
  10. AC 20-152A: Development Assurance for Airborne Electronic HardwareFederal Aviation Administration, AIR-622 · U.S. Department of Transportation · 2022Correct and complete requirements are the cornerstone; traceability reviews for DAL A/B; review of synthesis and place-and-route reports; each verification case must fit the requirements it traces to and cover them completely.
  11. ISO 26262Wikipedia contributors · WikipediaASIL from severity, exposure and controllability; safety goals and requirements allocated to hardware and software components; traceability between work products.
  12. Assessment of Safety Standards for Automotive Electronic Control Systems (DOT HS 812 285)Qi D. Van Eikema Hommes · National Highway Traffic Safety Administration · 2016Section 3.6: ISO 26262 sets top-level safety goals and assesses an ASIL for each from severity, exposure and controllability.
  13. Formal Assisted Fault Campaign for ISO26262 CertificationNitin Ahuja, Mayank Agarwal, Sandeep Jana · DVCon Europe proceedings (Accellera) · 2019Diagnostic coverage measures a safety mechanism; SPFM, LFM and PMHF targets by ASIL (SPFM ≥ 99% for ASIL D); LBIST as a diagnostic for latent faults.
  14. AEC-Q100 Rev-J1: Failure Mechanism Based Stress Test Qualification for Integrated Circuits in Automotive ApplicationsComponent Technical Committee · Automotive Electronics CouncilTable 1: operating temperature grades 0 (−40 °C to +150 °C) to 3 (−40 °C to +85 °C).
  15. Test Economics (CMPE 418 lecture notes)Chintan Patel · University of Maryland, Baltimore County · 2004Defect level is the fraction of bad chips that pass test, in DPM; it depends on yield and fault coverage; above 500 ppm is unacceptable for commercial VLSI.
  16. Semiconductor intellectual property coreWikipedia contributors · WikipediaLicensed, pre-verified blocks shorten design cycles; interface IP (PCIe, Ethernet, USB, SDRAM) needs digital and analog parts; hard cores and PHYs are tied to one foundry process.
  17. ComplianceUSB Implementers Forum · USB-IFCertified products are added to the Integrators List; logo use needs certification plus a trademark license; workshops and independent test labs.
  18. SystemC overviewAccellera Systems Initiative · systemc.orgC++ class library for system-level modeling; hardware/software partitioning; early software development before RTL; TLM model exchange and virtual prototyping; IEEE Std 1666-2023.
  19. About gem5The gem5 project · gem5.orgOpen-source, modular simulator for system and microarchitecture research: CPU models, caches and DRAM controller models.
  20. Cost (EEC 116 lecture handout)Bevan Baas · University of California, DavisCost per chip = fixed NRE (design, masks, CAD) + variable cost (silicon, packaging, test); dies-per-wafer formula with edge loss; yield (1 + D·A/α)^−α with α ≈ 3; die cost = wafer cost ÷ (dies per wafer × yield).
  21. Yield Modeling and Analysis (IEOR 130 course notes)Robert C. Leachman · University of California, Berkeley, IEOR 130 course page (Internet Archive copy) · 2017Poisson, Murphy, Seeds and negative binomial die-yield models; clustering makes Poisson pessimistic for large dies; cluster parameter α.
  22. AI Chips: What They Are and Why They MatterSaif M. Khan, Alexander Mann · Center for Security and Emerging Technology, Georgetown University · 2020Modeled foundry sale price per 300 mm wafer from $1,650 (90 nm) to $16,988 (5 nm); price per chip at a fixed transistor count; IBS design-cost estimates rising about 24% a year.
  23. Pad Ring and Floor Planning (ELEC6231 lecture notes)Iain McNally · University of SouthamptonPad-limited (small core, many pads) vs. core-limited (large core, few pads) dies.
  24. Lecture 7: Power (E158, CMOS VLSI Design 4th ed. slides)David Harris · Harvey Mudd CollegeP_switching = α·C·V²·f; activity factor (clock α = 1, typical α ≈ 0.1); short-circuit current under 10%; static power components; 65 nm off-current of 100, 10 and 1 nA/µm at Vt of 0.3, 0.4 and 0.5 V.
  25. Lecture 4: Nonideal Transistor Theory (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd CollegeRising temperature lowers on-current and raises off-current (leakage); process, voltage and temperature corners.
  26. Design of Ion-Implanted MOSFET’s with Very Small Physical DimensionsRobert H. Dennard, Fritz H. Gaensslen, Hwa-Nien Yu, V. Leo Rideout, Ernest Bassous, Andre R. LeBlanc · IEEE Journal of Solid-State Circuits SC-9(5), reprinted as Appendix D of “The Future of Computing Performance” (National Academies Press, 2011) · 1974Table I: scaling dimensions and voltage by 1/κ scales delay by 1/κ and power per circuit by 1/κ², so power density stays constant.
  27. Dark Silicon and the End of Multicore ScalingHadi Esmaeilzadeh, Emily Blem, Renée St. Amant, Karthikeyan Sankaralingam, Doug Burger · ISCA 2011 (author copy, University of Wisconsin–Madison) · 2011Failure of Dennard scaling slowed supply-voltage scaling and left designs power-limited; TDP defined as the chip power budget; 21% dark at 22 nm and over 50% at 8 nm; 7.9× average speedup through 2024.
  28. A Simple MOSFET Model for Circuit Analysis and Its Application to CMOS Gate Delay Analysis and Series-Connected MOSFET Structure (Memorandum UCB/ERL M90/19)Takayasu Sakurai, A. Richard Newton · University of California, Berkeley, Electronics Research Laboratory · 1990Drain current ∝ (V_GS − V_TH)^α with α falling from 2 toward 1 as velocity saturation grows; inverter delay in terms of C_L·V_DD ÷ I_D0.
  29. Near-Threshold Computing: Reclaiming Moore’s Law Through Energy Efficient Integrated CircuitsRonald G. Dreslinski, Michael Wieckowski, David Blaauw, Dennis Sylvester, Trevor Mudge · Proceedings of the IEEE (author copy, David Blaauw’s group, University of Michigan) · 2010Near-threshold operation gives about 10× lower energy for about 10× lower performance; energy minimum below threshold; commercial Vdd floor about 70% of nominal.