896 terms

Glossary

Definitions follow your detail level. Open any term to compare all three.

896 of 896 · showing Novice definitions

2D mesh

An on-chip network layout where cores sit on a grid and each is linked only to its four nearest neighbors. Data going farther takes several steps (hops) through the cores in between.

All three levels

Beginner. A grid of roads connecting the cores, like city blocks: each core connects to the ones directly above, below, left and right.

Novice. An on-chip network layout where cores sit on a grid and each is linked only to its four nearest neighbors. Data going farther takes several steps (hops) through the cores in between.

Expert. A k×kk \times k mesh has node degree 4, diameter 2(k−1)2(k-1) and half the bisection bandwidth of the equivalent torus. Its short, equal-length wires lay out well on a die, which is why it dominates large on-chip networks.

3Cs (compulsory, capacity, conflict misses)

The three causes of cache misses. Compulsory: the first time a piece of data is ever used. Capacity: the cache is too small for everything the program is working on. Conflict: several pieces of data compete for the same few places in the cache.

All three levels

Beginner. Three reasons memory lookups fail: it’s the first time, the memory is too small, or two things fight over the same spot.

Novice. The three causes of cache misses. Compulsory: the first time a piece of data is ever used. Capacity: the cache is too small for everything the program is working on. Conflict: several pieces of data compete for the same few places in the cache.

Expert. Defined by simulation: compulsory misses happen even in an infinite cache; capacity misses are the extra misses of a fully associative cache of the same size; the rest are conflict misses. More capacity cuts capacity misses, more associativity cuts conflict misses, and bigger lines cut compulsory misses when nearby data is used together.

3D stacking

Bonding dies face to face or face to back so that connections run vertically across their overlap, using through-silicon vias and microbumps or hybrid bonds.

All three levels

Beginner. Putting chips right on top of each other in one package, wired straight up and down.

Novice. Bonding dies face to face or face to back so that connections run vertically across their overlap, using through-silicon vias and microbumps or hybrid bonds.

Expert. Trades the edge-limited bandwidth of side-by-side links for area-scaled bandwidth, at the cost of heat that must cross the upper die, power that must pass through the lower one, and test access for buried dies.

Abstract (layout abstract)

A stand-in for a cell or block that shows only its outline, its connection points (pins) and the areas wires must avoid. Placement and routing work with abstracts, usually in the LEF format, because they don’t need the inside.

All three levels

Beginner. A simple outline of a chip part: its size and where its connections are, with the inside left out.

Novice. A stand-in for a cell or block that shows only its outline, its connection points (pins) and the areas wires must avoid. Placement and routing work with abstracts, usually in the LEF format, because they don’t need the inside.

Expert. The LEF view of a cell or macro: SIZE, PIN shapes by layer and OBS (obstructions). At tapeout every abstract must be replaced by the real layout of the same name; one that isn’t becomes an empty hole on the masks.

Abutted floorplan

A floorplan in which blocks touch edge to edge with no channels between them. It saves area, but every connection between blocks must be planned through the blocks’ own pins.

All three levels

Beginner. Blocks packed edge to edge, with no gaps between them.

Novice. A floorplan in which blocks touch edge to edge with no channels between them. It saves area, but every connection between blocks must be planned through the blocks’ own pins.

Expert. No top-level logic or channels: signals that cross the chip, including the clock, route through the blocks, so they must be planned into each block up front.

Acceleration factor

How many times faster a stress ages a part than normal use: time to fail in use divided by time to fail under stress. For temperature, the Arrhenius model gives AF=exp⁡ ⁣[(Ea/k)(1/Tuse−1/Tstress)]\mathrm{AF} = \exp\!\left[(E_{\mathrm{a}}/k)(1/T_{\mathrm{use}} - 1/T_{\mathrm{stress}})\right], with temperatures in kelvin, EaE_{\mathrm{a}} the failure mechanism’s activation energy and kk Boltzmann’s constant.

All three levels

Beginner. How many times faster a stress test ages a chip compared with normal use.

Novice. How many times faster a stress ages a part than normal use: time to fail in use divided by time to fail under stress. For temperature, the Arrhenius model gives AF=exp⁡ ⁣[(Ea/k)(1/Tuse−1/Tstress)]\mathrm{AF} = \exp\!\left[(E_{\mathrm{a}}/k)(1/T_{\mathrm{use}} - 1/T_{\mathrm{stress}})\right], with temperatures in kelvin, EaE_{\mathrm{a}} the failure mechanism’s activation energy and kk Boltzmann’s constant.

Expert. Each failure mechanism has its own activation energy and voltage dependence, so one AF for a whole product is an approximation. Moving EaE_{\mathrm{a}} across its typical 0.3–1.5 eV range changes the AF by orders of magnitude.

Accelerator

A block built for one kind of work, such as graphics (a GPU), neural networks (an NPU), video or encryption, that the main processor hands jobs to. It saves energy by skipping the work of fetching and decoding general-purpose instructions, using only as many bits as the job needs, doing many operations at once and keeping data in nearby memory.

All three levels

Beginner. A part of a chip built to do one kind of job, such as video or AI, much faster and with much less energy than a general-purpose processor.

Novice. A block built for one kind of work, such as graphics (a GPU), neural networks (an NPU), video or encryption, that the main processor hands jobs to. It saves energy by skipping the work of fetching and decoding general-purpose instructions, using only as many bits as the job needs, doing many operations at once and keeping data in nearby memory.

Expert. Its benefit is limited three ways: by the share of the work it can take on (Amdahl’s law), by the cost of handing work over and collecting results, and by memory bandwidth. Once arithmetic is cheap, memory usually takes most of its area and power.

Accumulator

The register that holds the running sum in a multiply-accumulate unit. It is usually much wider than the inputs (for example 32 bits for 8-bit inputs) so that adding thousands of products doesn’t overflow or lose small contributions.

All three levels

Beginner. The running total inside a math unit, where many multiplied pairs are added up.

Novice. The register that holds the running sum in a multiply-accumulate unit. It is usually much wider than the inputs (for example 32 bits for 8-bit inputs) so that adding thousands of products doesn’t overflow or lose small contributions.

Expert. Fixed-point (Kulisch-style, exact) for integer inputs, typically INT32 for INT8; floating point (FP32, sometimes FP16 or a reduced-width internal format) for float inputs. Its width and the length of the reduction set the accumulation error; long reductions are often split into chunks summed at higher precision.

Activation

A value produced by one layer of a neural network and consumed by the next. In a matrix multiplication the activations are one operand and the weights the other.

All three levels

Beginner. The numbers flowing through a neural network: the input, and every in-between result one layer hands to the next.

Novice. A value produced by one layer of a neural network and consumed by the next. In a matrix multiplication the activations are one operand and the weights the other.

Expert. Activations change with every input while weights are fixed during inference, which is why weight-stationary designs pin the weights and stream activations. Their total size grows with batch and sequence length.

Activation recomputation (checkpointing)

A memory-saving trick: keep only some layer inputs during the forward pass and recompute the rest just before the backward pass needs them. It costs extra computation.

All three levels

Beginner. Throwing away some in-between results to save memory, then working them out again when they are needed.

Novice. A memory-saving trick: keep only some layer inputs during the forward pass and recompute the rest just before the backward pass needs them. It costs extra computation.

Expert. Full recomputation stores only each layer’s input (≈2sbh\approx 2sbh bytes) and costs one extra forward pass (about 33% more FLOPs). Selective recomputation re-does only the cheap attention core, removing the 5as/h5as/h term for a few percent of compute.

Activations (training memory)

The intermediate outputs of every layer during the forward pass. Training keeps them until the backward pass uses them, so their memory grows with batch size, sequence length and layer count.

All three levels

Beginner. The in-between results a model produces while working on an example, which it has to keep until it has learned from that example.

Novice. The intermediate outputs of every layer during the forward pass. Training keeps them until the backward pass uses them, so their memory grows with batch size, sequence length and layer count.

Expert. Per transformer layer about sbh (34+5as/h)sbh\,(34 + 5as/h) bytes in 16-bit training (Korthikanti et al.); the attention term grows with s2s^2. Reduced by tensor and sequence parallelism, and traded for compute by recomputation.

Active electrical cable (AEC)

A copper cable with a retimer chip in each plug. The retimers clean up the signal, which lets copper reach a few meters further than a plain cable, at the cost of some power.

All three levels

Beginner. A copper cable with small booster chips in its plugs so it can reach farther.

Novice. A copper cable with a retimer chip in each plug. The retimers clean up the signal, which lets copper reach a few meters further than a plain cable, at the cost of some power.

Expert. Twinax with a retimer in each connector shell (a non-retimed active copper cable uses a linear redriver instead). Reach of several meters, beyond passive copper; the retimers add several watts per end.

Activity factor (α)

How often, per clock tick, a signal charges up (goes from 0 to 1). The clock does this every tick (α=1\alpha = 1); a signal that flips every tick does it every other tick (α=1/2\alpha = 1/2); typical logic signals average around 0.1.

All three levels

Beginner. How often, on average, a part of the chip switches each clock tick.

Novice. How often, per clock tick, a signal charges up (goes from 0 to 1). The clock does this every tick (α=1\alpha = 1); a signal that flips every tick does it every other tick (α=1/2\alpha = 1/2); typical logic signals average around 0.1.

Expert. Depends on the workload, and is the least certain input at the spec stage. Early estimates come from the previous chip’s measurements, or from simulating representative workloads on the design code and feeding the recorded switching to a power-estimation tool.

Adaptive routing

Choosing each packet’s or flowlet’s output port from the current load on the candidate ports, so traffic moves to less busy paths as conditions change.

All three levels

Beginner. Switches that steer traffic away from busy routes as they see them, instead of following a fixed rule.

Novice. Choosing each packet’s or flowlet’s output port from the current load on the candidate ports, so traffic moves to less busy paths as conditions change.

Expert. Load-aware path selection in the switch (local queue depth, or remote congestion state), per packet or per flowlet. Common in InfiniBand and HPC fabrics and offered on recent Ethernet switch chips; like spraying, per-packet variants need reordering tolerance at the receiver.

Adaptive SoC

A system on a chip that combines hard processor cores, FPGA fabric and, increasingly, arrays of small vector or AI engines, joined by an on-chip network. AMD uses the name for its Versal family; Altera and Microchip build similar SoC FPGAs.

All three levels

Beginner. One chip that holds normal processor cores, a rewirable FPGA part, and often special engines for AI math.

Novice. A system on a chip that combines hard processor cores, FPGA fabric and, increasingly, arrays of small vector or AI engines, joined by an on-chip network. AMD uses the name for its Versal family; Altera and Microchip build similar SoC FPGAs.

Expert. Hard CPUs for control, fabric for bit-level and streaming logic, a vector-engine array for dense MACs, hard memory controllers and a hard NoC. The goal is to move most of the area into hard blocks, where the FPGA–ASIC gap is small, while keeping fabric for what changes.

ADC (analog-to-digital converter)

Converts a voltage or current into a digital number with a fixed number of bits. An 8-bit ADC distinguishes 256 levels. More bits mean finer steps but more energy, area and time.

All three levels

Beginner. A circuit that measures an electrical signal and turns it into a number.

Novice. Converts a voltage or current into a digital number with a fixed number of bits. An 8-bit ADC distinguishes 256 levels. More bits mean finer steps but more energy, area and time.

Expert. In analog CIM, the column ADC usually dominates power and area. Energy per conversion grows roughly linearly with bits at low resolution and about 4× per bit once thermal noise sets the limit; ENOB, not nominal bits, measures what you actually get.

AEC-Q100

The Automotive Electronics Council’s set of stress tests that a chip must pass to be used in cars, with temperature grades 0 (hottest) to 3.

All three levels

Beginner. The set of tough tests a chip must pass before carmakers will use it.

Novice. The Automotive Electronics Council’s set of stress tests that a chip must pass to be used in cars, with temperature grades 0 (hottest) to 3.

Expert. Defines test groups and sample sizes (for example HTOL on 77 parts from each of 3 lots, zero fails). It qualifies reliability; functional safety is covered separately by ISO 26262.

AI accelerator (GPU, NPU)

A separate chip, usually on its own board or module, with many parallel math units and its own fast memory (often HBM). The host CPU sends it data and instructions over a link such as PCIe.

All three levels

Beginner. A chip built to do the huge piles of multiplication in AI much faster than a normal processor.

Novice. A separate chip, usually on its own board or module, with many parallel math units and its own fast memory (often HBM). The host CPU sends it data and instructions over a link such as PCIe.

Expert. A device (GPU or other) with its own memory hierarchy and usually a dedicated scale-up fabric to its peers. To the host it is a PCIe endpoint, sometimes a CXL Type 2 device; its host link carries data staging, kernel launches and synchronization, not the bulk of training traffic.

All-gather

A collective in which each participant contributes a block and every participant ends up with all the blocks, concatenated.

All three levels

Beginner. Each chip has one piece of a puzzle; at the end, every chip has the whole puzzle.

Novice. A collective in which each participant contributes a block and every participant ends up with all the blocks, concatenated.

Expert. Moves (p−1)/p(p-1)/p of the final buffer per participant at minimum; no arithmetic. The second half of a bandwidth-optimal all-reduce.

All-reduce

A collective that combines (usually sums) a buffer from every participant element by element and leaves the full result on every participant. Training uses it to add up gradients across chips.

All three levels

Beginner. Every chip starts with its own list of numbers; at the end, every chip holds the sum of all the lists.

Novice. A collective that combines (usually sums) a buffer from every participant element by element and leaves the full result on every participant. Training uses it to add up gradients across chips.

Expert. Semantically reduce + broadcast; efficiently implemented as reduce-scatter followed by all-gather, which moves 2(p−1)/p2(p-1)/p of the buffer per participant, the bandwidth lower bound for algorithms that don’t reduce inside the network.

All-to-all

A collective in which each participant sends a distinct block to each of the others: a full personalized exchange. Mixture-of-experts models use it to route tokens to the chips that hold each expert.

All three levels

Beginner. Every chip sends a different message to every other chip at the same time.

Novice. A collective in which each participant sends a distinct block to each of the others: a full personalized exchange. Mixture-of-experts models use it to route tokens to the chips that hold each expert.

Expert. Each participant sends and receives (p−1)/p(p-1)/p of its buffer, but every pair exchanges distinct data, so it stresses bisection bandwidth and incast; it cannot be shortened by pipelining around a ring the way reductions can.

Alpha-beta (α-β) model

A cost model in which sending nn bytes takes α+n/B\alpha + n/B: α\alpha is a fixed per-message latency and BB is the link bandwidth (β=1/B\beta = 1/B is the time per byte). A collective’s time is the sum over its steps.

All three levels

Beginner. A simple rule for how long sending data takes. There is a fixed delay to get started, plus a time that grows with the amount of data.

Novice. A cost model in which sending nn bytes takes α+n/B\alpha + n/B: α\alpha is a fixed per-message latency and BB is the link bandwidth (β=1/B\beta = 1/B is the time per byte). A collective’s time is the sum over its steps.

Expert. Hockney-style linear model T=α+nβT = \alpha + n\beta (+ nγ+\, n\gamma for reduction arithmetic). Ignores contention, protocol switching, pipelining granularity and congestion, but captures the latency-bound vs bandwidth-bound trade-off that drives algorithm selection.

Alpha-power law

A MOSFET model in which saturation current grows as (VGS−Vt)(V_{\mathrm{GS}} - V_{\mathrm{t}}) to a power α\alpha between 1 and 2 instead of exactly 2. α=2\alpha = 2 gives back the square law; α\alpha near 1 means strong velocity saturation.

All three levels

Beginner. A formula that fits real, tiny transistors better than the textbook rule, because their current grows more gently.

Novice. A MOSFET model in which saturation current grows as (VGS−Vt)(V_{\mathrm{GS}} - V_{\mathrm{t}}) to a power α\alpha between 1 and 2 instead of exactly 2. α=2\alpha = 2 gives back the square law; α\alpha near 1 means strong velocity saturation.

Expert. IDsat=Pc(β/2)(VGS−Vt)αI_{\mathrm{Dsat}} = P_{\mathrm{c}} (\beta/2)(V_{\mathrm{GS}} - V_{\mathrm{t}})^{\alpha} and VDSAT=Pv(VGS−Vt)α/2V_{\mathrm{DSAT}} = P_{\mathrm{v}} (V_{\mathrm{GS}} - V_{\mathrm{t}})^{\alpha/2}, with α\alpha, PcP_{\mathrm{c}} and PvP_{\mathrm{v}} fitted to measured or simulated curves. Sakurai and Newton’s nnth-power-law form also fits the linear region; α≈1.3\alpha \approx 1.3 is a typical value for a 65 nm process.

always-on domain

The part of a chip that stays powered in every mode. It holds the wake-up logic, timers and the controller that turns the other parts on and off.

All three levels

Beginner. The small part of the chip that never sleeps. It notices a button press or a message and wakes up the rest.

Novice. The part of a chip that stays powered in every mode. It holds the wake-up logic, timers and the controller that turns the other parts on and off.

Expert. Kept small and built from low-leakage cells. Its signals cross switched domains through always-on buffers, and its leakage sets the device’s floor current in deep sleep.

AMAT (average memory access time)

Average memory access time=hit time+miss rate×miss penalty\text{Average memory access time} = \text{hit time} + \text{miss rate} \times \text{miss penalty}. Example: a 1-cycle cache that misses 10% of the time, with 20 cycles to fetch a missed line, averages 1+0.1×20=31 + 0.1 \times 20 = 3 cycles. With several cache levels, the miss penalty is itself the average access time of the next level.

All three levels

Beginner. On average, how long the processor waits each time it asks memory for something.

Novice. Average memory access time=hit time+miss rate×miss penalty\text{Average memory access time} = \text{hit time} + \text{miss rate} \times \text{miss penalty}. Example: a 1-cycle cache that misses 10% of the time, with 20 cycles to fetch a missed line, averages 1+0.1×20=31 + 0.1 \times 20 = 3 cycles. With several cache levels, the miss penalty is itself the average access time of the next level.

Expert. A first-order latency model. Processors that keep working past a miss (out-of-order cores), or that have several misses outstanding at once, hide part of the penalty, so the real slowdown is smaller than AMAT suggests for those cores and close to it for simple in-order ones.

Amdahl’s law

Speedup=1/((1−f)+f/s)\text{Speedup} = 1 / \bigl((1 - f) + f/s\bigr), where ff is the fraction of the original run time that gets faster and ss is how much faster it gets. If f=0.8f = 0.8, the other 20% still takes its full time, so the speedup can never exceed 5×, however large ss is.

All three levels

Beginner. The rule that speeding up part of a job can only help so much, because the part you didn’t speed up still takes its time.

Novice. Speedup=1/((1−f)+f/s)\text{Speedup} = 1 / \bigl((1 - f) + f/s\bigr), where ff is the fraction of the original run time that gets faster and ss is how much faster it gets. If f=0.8f = 0.8, the other 20% still takes its full time, so the speedup can never exceed 5×, however large ss is.

Expert. Applies to accelerators (ff is the share of time that can be offloaded, and the cost of handing work over belongs in the part that isn’t sped up) and to multicore chips (ss is the number of cores). It assumes a fixed job; Gustafson’s law covers jobs that grow with the machine.

Analog compute

Representing numbers as physical quantities (voltage, current, charge, time) and letting physics do the arithmetic: currents add where wires meet, charge spreads across capacitors. Cheap at low precision, but noise and mismatch limit accuracy.

All three levels

Beginner. Doing math with electricity that can be at any level, not just on or off. For example, flows of electricity add up in a wire.

Novice. Representing numbers as physical quantities (voltage, current, charge, time) and letting physics do the arithmetic: currents add where wires meet, charge spreads across capacitors. Cheap at low precision, but noise and mismatch limit accuracy.

Expert. Energy efficient below roughly 8 bits of required SNR, where structural simplicity beats digital logic; above that, kT/CkT/C noise makes each extra bit cost about 4× the energy. Needs converters at the digital boundary.

analog front end (AFE)

The amplifiers, filters and analog-to-digital converter that pick up the body’s electrical signals, such as heartbeats or nerve activity, which range from millionths to thousandths of a volt, and turn them into numbers.

All three levels

Beginner. The part of a chip that picks up the body’s tiny electric signals and makes them big and clear enough to measure.

Novice. The amplifiers, filters and analog-to-digital converter that pick up the body’s electrical signals, such as heartbeats or nerve activity, which range from millionths to thousandths of a volt, and turn them into numbers.

Expert. Designed for input-referred noise of a few µVrms at µA currents, rejection of large electrode offsets through AC coupling, and low 1/f1/f noise through large input devices or chopping. Noise efficiency factor is the usual figure of merit.

And-Inverter Graph (AIG)

A way to represent any logic using only two-input AND gates, with some connections marked as inverted (NOT). Its uniform shape makes it easy for tools to simplify.

All three levels

Beginner. A way of drawing any logic using only one kind of block, AND, plus markers that flip yes to no.

Novice. A way to represent any logic using only two-input AND gates, with some connections marked as inverted (NOT). Its uniform shape makes it easy for tools to simplify.

Expert. A directed acyclic graph of two-input AND nodes whose edges may be inverted. Structural hashing merges duplicate nodes as it is built. Node count and depth roughly track mapped area and delay, and rewriting, balancing, mapping and SAT-based equivalence checking all work on it. It is the core data structure in ABC.

Antenna effect

Damage that can happen during manufacturing. While metal layers are etched with electrically charged gas (plasma), a long wire connected only to a transistor’s input collects charge that can break through the transistor’s very thin insulating layer. Antenna rules limit how much wire can connect to one input before a protection measure is needed.

All three levels

Beginner. Static charge that builds up on long wires while the chip is made. It can zap the chip’s tiny switches.

Novice. Damage that can happen during manufacturing. While metal layers are etched with electrically charged gas (plasma), a long wire connected only to a transistor’s input collects charge that can break through the transistor’s very thin insulating layer. Antenna rules limit how much wire can connect to one input before a protection measure is needed.

Expert. Antenna rules limit the ratio of metal area connected to a gate to the gate area, per layer and cumulatively up the stack. Fixes: a protection diode, jumping to a higher metal layer near the gate (so the long segment isn’t connected until later), or splitting the net. Routers repair most violations; signoff re-checks the final layout with the foundry’s rules.

AOCV (advanced on-chip variation)

A way of setting timing safety margins that depends on the path: long chains of gates get a smaller margin per gate, because random differences between gates partly cancel out, and paths spread far across the chip get a larger one.

All three levels

Beginner. A safety margin that shrinks on paths with more steps, because small random errors tend to cancel out.

Novice. A way of setting timing safety margins that depends on the path: long chains of gates get a smaller margin per gate, because random differences between gates partly cancel out, and paths spread far across the chip get a larger one.

Expert. Derates looked up from tables by path depth (number of stages) and distance, instead of one flat percentage. Depth models random variation, which averages out over many stages; distance models systematic variation across the die. At advanced nodes it has largely given way to POCV with LVF.

Architecture and microarchitecture

Architecture is the chip-level plan: which big blocks the chip has (processors, special-purpose units, memories, the connections between them) and how they work together. Microarchitecture is how one block is organized inside, for example how many steps a processor splits each instruction into and how big its caches are.

All three levels

Beginner. The overall plan of a chip: which big parts it has, what each one does, and how they connect.

Novice. Architecture is the chip-level plan: which big blocks the chip has (processors, special-purpose units, memories, the connections between them) and how they work together. Microarchitecture is how one block is organized inside, for example how many steps a processor splits each instruction into and how big its caches are.

Expert. The word is used two ways. In processor design, “architecture” often means the instruction set, the list of instructions software can use, and “microarchitecture” is the hardware that carries them out. In chip (SoC) design, “architecture” means the plan the RTL team implements: the block partition, memory system, interconnect, clock and power plan, and per-block budgets.

Architecture specification

The document that describes the chip as a set of large blocks (processors, memories, accelerators, connections to the outside) and how they talk to each other. It answers, block by block, how the chip will meet its requirements.

All three levels

Beginner. The document that sketches how the chip will be built: which big blocks it has and how they connect.

Novice. The document that describes the chip as a set of large blocks (processors, memories, accelerators, connections to the outside) and how they talk to each other. It answers, block by block, how the chip will meet its requirements.

Expert. Owned by the chip architects. It fixes how the chip is split into blocks, the interfaces between them, the memory map (which addresses reach which block), the clock and power domains (regions that run at their own speed or can be switched off), and the power, area and performance budget each block gets. The block-level microarchitecture specs are written from it.

Arithmetic (operational) intensity

How many operations a task does per byte it moves to or from memory, usually main memory. Multiplying two large matrices reuses each number many times, so its intensity is high; adding two long lists of numbers uses each number once, so it is low.

All three levels

Beginner. How much calculating a job does for each piece of data it fetches from memory.

Novice. How many operations a task does per byte it moves to or from memory, usually main memory. Multiplying two large matrices reuses each number many times, so its intensity is high; adding two long lists of numbers uses each number once, so it is low.

Expert. Depends on the algorithm, the number format, how the work is split into tiles that fit on chip, and how much on-chip memory there is. A larger scratchpad or better tiling raises intensity and moves a task to the right on the roofline.

Array utilization

Useful multiply-accumulates done divided by the most the array could have done in the same time (number of PEs × cycles). Not to be confused with placement utilization in chip layout.

All three levels

Beginner. The share of the time the calculator cells are doing useful work. 100% would mean every cell is busy on every tick.

Novice. Useful multiply-accumulates done divided by the most the array could have done in the same time (number of PEs × cycles). Not to be confused with placement utilization in chip layout.

Expert. U=MKN/(PEs×cycles)U = MKN / (\text{PEs} \times \text{cycles}). It falls with fill/drain, matrix shapes that don’t fill the array (fragmentation), stalls waiting for weights or memory, and small batches. TPU v1 averaged 23% useful MACs across six production apps.

ASIC

Application-specific integrated circuit: a custom chip manufactured from your own design. High up-front cost, but the best speed and efficiency.

All three levels

Beginner. A chip built for one job. Once it is made, it can’t be changed.

Novice. Application-specific integrated circuit: a custom chip manufactured from your own design. High up-front cost, but the best speed and efficiency.

Expert. Fixed function after tapeout, with mask NRE and months of fab and bring-up time. It pays off when the function is stable and either volume or performance margin justifies the cost.

ASIL

Automotive Safety Integrity Level. Each hazard is rated for how severe the harm would be, how often the driving situation occurs, and how well a driver could still control the car; a table combines the three into QM (ordinary quality, no safety requirement) or ASIL A to D, with D the strictest.

All three levels

Beginner. A safety grade from A to D for how dangerous a failure would be. D is the strictest.

Novice. Automotive Safety Integrity Level. Each hazard is rated for how severe the harm would be, how often the driving situation occurs, and how well a driver could still control the car; a table combines the three into QM (ordinary quality, no safety requirement) or ASIL A to D, with D the strictest.

Expert. Sets the hardware metric targets (SPFM, LFM, PMHF), the required methods and how independent the confirmation reviews must be. ASIL decomposition can split one requirement across redundant, independent elements at lower ASILs.

Aspect ratio

Height divided by width (OpenROAD’s convention). A ratio of 1 is a square. Squarish shapes keep wires short in both directions.

All three levels

Beginner. How tall a chip is compared with how wide it is. A ratio of 1 means a square.

Novice. Height divided by width (OpenROAD’s convention). A ratio of 1 is a square. Squarish shapes keep wires short in both directions.

Expert. Free to choose at the chip level, within limits set by the package and by how many pads or bumps must fit; for a block inside a chip it is negotiated with the top-level plan. A long, thin core lengthens wires along its long side and puts more wiring demand in one direction than the other.

Assertion (SVA)

A rule written into the design or testbench and checked continuously while the design runs, such as “a reply always comes within four clock cycles.” If it is ever broken, the simulator reports where and when. SystemVerilog Assertions (SVA) is the standard notation.

All three levels

Beginner. A rule written into a design or a test. It sounds an alarm the moment the rule is broken.

Novice. A rule written into the design or testbench and checked continuously while the design runs, such as “a reply always comes within four clock cycles.” If it is ever broken, the simulator reports where and when. SystemVerilog Assertions (SVA) is the standard notation.

Expert. Immediate assertions check a condition at one moment; concurrent ones describe sequences over clock cycles and are understood by both simulators and formal tools. Engineers write assume for rules about inputs, assert for what must hold, and cover to show that a scenario can happen at all.

Asynchronous FIFO

A small first-in, first-out queue between two clock domains. One side writes items in at its pace and the other reads them out at its own. Only the two position counters (where to write next, where to read next) cross between the clocks, through synchronizers.

All three levels

Beginner. A waiting line between two parts of the chip with different clocks: one side adds items, the other removes them at its own pace.

Novice. A small first-in, first-out queue between two clock domains. One side writes items in at its pace and the other reads them out at its own. Only the two position counters (where to write next, where to read next) cross between the clocks, through synchronizers.

Expert. The standard way to move a stream of multi-bit data between clocks. The pointers cross in Gray code, so a pointer caught mid-change reads as either its old or its new value. The full and empty flags are therefore a cycle or two late, but never wrong.

At-speed test

A scan test in which a change is started and its result caught exactly one normal clock period apart, so a path that is even slightly too slow gives the wrong answer. Loading and unloading the chains can still run slowly.

All three levels

Beginner. Testing the chip at its real working speed, to catch flaws that only make signals a bit late.

Novice. A scan test in which a change is started and its result caught exactly one normal clock period apart, so a path that is even slightly too slow gives the wrong answer. Loading and unloading the chains can still run slowly.

Expert. Targets transition, path delay and small-delay faults. Needs closely spaced launch and capture pulses (usually from on-chip clock control), timing signoff in test mode, and care with capture power and supply droop.

ATE (automatic test equipment)

The factory tester. Under the control of a test program, it applies stored patterns to a chip through its pins, compares the responses with the expected values and marks the chip good or bad, either on the wafer or after packaging.

All three levels

Beginner. The big, expensive machine in the factory that plugs into each chip, runs the tests and sorts good from bad.

Novice. The factory tester. Under the control of a test program, it applies stored patterns to a chip through its pins, compares the responses with the expected values and marks the chip good or bad, either on the wafer or after packaging.

Expert. Cost scales with channel count, speed, pattern memory and instruments, so DFT aims to cut the channels needed (compression), pattern memory and test time, and to let one tester test several chips at once (multi-site).

Atomic layer deposition (ALD)

Two chemicals are pulsed over the wafer in turn. Each reacts only with the surface until the surface is used up, so each cycle adds a fixed sliver, about a tenth of a nanometer.

All three levels

Beginner. A way of building a film one layer of atoms at a time, so it can be made extremely thin and even.

Novice. Two chemicals are pulsed over the wafer in turn. Each reacts only with the surface until the surface is used up, so each cycle adds a fixed sliver, about a tenth of a nanometer.

Expert. Self-limiting surface reactions give thickness control by cycle count and near-perfect conformality in deep, narrow features. Used for high-κ gate dielectrics, work-function metals and the inner-spacer fill in nanosheet transistors.

ATPG (automatic test pattern generation)

Software that works out input patterns that make a list of modeled faults visible at the outputs. With scan, it can treat every flip-flop as an extra input and output, which turns a hard problem into a much easier one.

All three levels

Beginner. Software that works out which test patterns to feed the chip, so that as many pretend flaws as possible would show up.

Novice. Software that works out input patterns that make a list of modeled faults visible at the outputs. With scan, it can treat every flip-flop as an extra input and output, which turns a hard problem into a much easier one.

Expert. Random patterns first for easy faults, then a deterministic search (PODEM- or FAN-style, or SAT-based) per remaining fault, with fault simulation to drop faults already detected and compaction to shrink the set. Outputs the patterns plus a class for every fault: detected, undetectable, ATPG-untestable, aborted.

Attention

A layer that scores how relevant each position is to each other position (a matrix multiply of queries by keys), turns the scores into weights with softmax, and takes a weighted mix of the values (another matrix multiply).

All three levels

Beginner. The step where a model compares every word with every other word to decide which ones to pay attention to.

Novice. A layer that scores how relevant each position is to each other position (a matrix multiply of queries by keys), turns the scores into weights with softmax, and takes a weighted mix of the values (another matrix multiply).

Expert. softmax(QKT/d) V\mathrm{softmax}(QK^T/\sqrt{d})\,V per head. Its FLOPs and score matrix grow as the square of sequence length; during decode it reads every cached key and value per new token, which makes it memory-bound.

Available, Preview and RDI

Available systems can be bought or rented today. Preview systems must become available within about the next benchmark round. RDI (research, development or internal) covers everything else, such as prototypes and in-house systems.

All three levels

Beginner. Labels that say whether the tested machine is something you can buy now, something coming soon, or a lab prototype.

Novice. Available systems can be bought or rented today. Preview systems must become available within about the next benchmark round. RDI (research, development or internal) covers everything else, such as prototypes and in-house systems.

Expert. Available needs public or on-request pricing and shipment to at least one third party. A Preview result must be resubmitted as Available after 140 days or by the next round (whichever is later) with performance within 2%, or it is marked invalid.

backside power delivery

Putting the chip’s power wiring on the back of the silicon instead of in the stack of wiring layers above the transistors. The front layers are then left for signal wires, and power reaches the transistors through tiny vertical connections from below.

All three levels

Beginner. Bringing power into the chip from underneath. That leaves the wires on top free to carry signals.

Novice. Putting the chip’s power wiring on the back of the silicon instead of in the stack of wiring layers above the transistors. The front layers are then left for signal wires, and power reaches the transistors through tiny vertical connections from below.

Expert. Separates the power delivery network from the signal metal stack, cutting IR drop and freeing routing tracks. It changes cell design, the heat path and failure analysis, because the substrate side is no longer open for probing.

Band gap

The range of energies with no allowed electron states, between the top of the valence band (EvE_{\mathrm{v}}) and the bottom of the conduction band (EcE_{\mathrm{c}}). In silicon Eg=1.12 eVE_{\mathrm{g}} = 1.12\,\mathrm{eV} at room temperature.

All three levels

Beginner. The jump in energy an electron in silicon must make to break free and move around. Until it makes the jump, it is stuck.

Novice. The range of energies with no allowed electron states, between the top of the valence band (EvE_{\mathrm{v}}) and the bottom of the conduction band (EcE_{\mathrm{c}}). In silicon Eg=1.12 eVE_{\mathrm{g}} = 1.12\,\mathrm{eV} at room temperature.

Expert. EgE_{\mathrm{g}} sets nin_{\mathrm{i}} through ni2=NcNv e−Eg/kTn_{\mathrm{i}}^2 = N_{\mathrm{c}} N_{\mathrm{v}}\, e^{-E_{\mathrm{g}}/kT}, so it controls intrinsic carriers, junction leakage and the temperature at which doping stops mattering. Silicon’s gap is indirect, which is why it is a poor light emitter.

Bandwidth density

Bandwidth per unit of die edge (GB/s per mm of shoreline) or per unit area (GB/s per mm²). It is the number of connections in that space times the data rate of each.

All three levels

Beginner. How much data can cross between chips for each millimeter of chip edge you use for connections.

Novice. Bandwidth per unit of die edge (GB/s per mm of shoreline) or per unit area (GB/s per mm²). It is the number of connections in that space times the data rate of each.

Expert. Shoreline density = signal bumps per mm² × bump-field depth × lane rate. Areal density matters for 3D, where the whole overlap can carry bonds. UCIe’s targets span about 22–125 GB/s/mm² (standard), 188–1,350 (advanced) and about 4,000 at 9 µm hybrid bonding.

Batch (batch size)

The number of inputs processed together by the same weights. In a matrix multiplication it adds rows to the activation matrix, which gives a weight-stationary array more work per weight load.

All three levels

Beginner. How many requests a chip works on together. Bigger batches keep the hardware busier but make each request wait longer.

Novice. The number of inputs processed together by the same weights. In a matrix multiplication it adds rows to the activation matrix, which gives a weight-stationary array more work per weight load.

Expert. Raises weight reuse and utilization but adds queueing latency; inference services cap it with response-time limits (7 ms at the 99th percentile for one TPU v1 workload).

Batching

Running several inputs through a model in one pass. The weights are fetched once and used for every input in the batch, which raises the work done per byte.

All three levels

Beginner. Serving many requests together so each number fetched from memory gets used for all of them at once.

Novice. Running several inputs through a model in one pass. The weights are fetched once and used for every input in the batch, which raises the work done per byte.

Expert. Raises weight reuse linearly with batch size, at the cost of per-request latency and KV-cache capacity. It does nothing for per-sequence traffic such as the KV cache.

Bathtub curve

Failure rate plotted against age, in three parts: early failures (high and falling), useful life (low and roughly constant) and wear-out (rising).

All three levels

Beginner. The shape of failure rate over a product’s life: many early failures, then a long quiet period, then rising failures as parts wear out.

Novice. Failure rate plotted against age, in three parts: early failures (high and falling), useful life (low and roughly constant) and wear-out (rising).

Expert. Burn-in and screening attack the first region; design rules for electromigration, oxide breakdown and transistor aging push the third beyond the product’s lifetime. FIT numbers describe the middle.

Bayesian optimization

A way to search for good settings when each try is slow. It fits a surrogate model to the tries so far, then picks the next try where the model predicts a good result or is most unsure.

All three levels

Beginner. A smart way to search for good settings when each try is slow. It keeps a running guess about which settings are good, and picks the next try to learn the most.

Novice. A way to search for good settings when each try is slow. It fits a surrogate model to the tries so far, then picks the next try where the model predicts a good result or is most unsure.

Expert. Usually a Gaussian process as the surrogate plus an acquisition function, such as expected improvement, that scores where to sample next. Works best below about 20 continuous settings with slow, noisy evaluations. TPE, a related method, models where good and poor samples fall instead.

BDD (binary decision diagram)

Binary decision diagram: a compact graph that represents a logic function as a chain of yes/no decisions on its inputs, so tools can reason about very large numbers of input combinations at once.

All three levels

Beginner. A compact diagram that stands for a yes-or-no rule. It lets tools reason about huge numbers of cases at once.

Novice. Binary decision diagram: a compact graph that represents a logic function as a chain of yes/no decisions on its inputs, so tools can reason about very large numbers of input combinations at once.

Expert. For a fixed order of the input variables the graph is unique, so checking two functions for equality is a pointer comparison. Its size depends heavily on that order, and for some functions, multipliers among them, it is exponential whatever the order.

BF16 (bfloat16)

“Brain floating point”: 1 sign bit, 8 exponent bits and 7 mantissa bits. It is the top half of an FP32 number, so it has the same range (about 10−3810^{-38} to 103810^{38}) with roughly 2–3 significant decimal digits.

All three levels

Beginner. A 16-bit number format that reaches numbers just as big and as tiny as the classic 32-bit format, but keeps fewer digits.

Novice. “Brain floating point”: 1 sign bit, 8 exponent bits and 7 mantissa bits. It is the top half of an FP32 number, so it has the same range (about 10−3810^{-38} to 103810^{38}) with roughly 2–3 significant decimal digits.

Expert. FP32 with the low 16 mantissa bits dropped (round-to-nearest-even on conversion). Same exponent and bias (127) as FP32, so no loss scaling is needed; precision p=8p = 8, relative step 2−72^{-7}. The standard training format on most accelerators, usually with FP32 accumulation.

Bilinear interpolation

Interpolating in two directions at once: first along one axis between two pairs of table entries, then along the other axis between the two results. Timing tools use it to read NLDM tables between their grid points.

All three levels

Beginner. Guessing a value that falls between four measured ones by blending them, trusting the closest ones most.

Novice. Interpolating in two directions at once: first along one axis between two pairs of table entries, then along the other axis between the two results. Timing tools use it to read NLDM tables between their grid points.

Expert. f=(1−u)(1−v) y00+u(1−v) y10+uv y11+(1−u)v y01f = (1-u)(1-v)\,y_{00} + u(1-v)\,y_{10} + uv\,y_{11} + (1-u)v\,y_{01}, with u,vu, v the fractional positions in the bracketing interval. Linear along each axis, quadratic along a diagonal. Outside the grid, the same formula with uu or vv beyond [0,1][0, 1] extrapolates linearly.

Binning

Putting each tested chip into a category, or bin: failing chips by which test they failed, passing chips by speed, power or how many of their parts work. Faster chips can sell as higher-priced models, and chips with one broken block switched off as cheaper ones.

All three levels

Beginner. Sorting tested chips into grades, like eggs by size, and selling each grade as a different product.

Novice. Putting each tested chip into a category, or bin: failing chips by which test they failed, passing chips by speed, power or how many of their parts work. Faster chips can sell as higher-priced models, and chips with one broken block switched off as cheaper ones.

Expert. Bin limits are written into the test program and come from characterization. Bin counts per wafer and per lot are among the earliest yield signals, and built-in redundancy (spare cores, spare memory rows) turns some failures into lower-grade passes.

biocompatibility

Whether the materials that touch the body cause harm, such as irritation, toxicity or an immune reaction. It is evaluated under the standard ISO 10993-1, as part of a risk analysis.

All three levels

Beginner. Making sure the materials that touch the body don’t harm it, for example by causing swelling or poisoning.

Novice. Whether the materials that touch the body cause harm, such as irritation, toxicity or an immune reaction. It is evaluated under the standard ISO 10993-1, as part of a risk analysis.

Expert. Drives package materials, coatings and electrode metals, and any material change can reopen the biological evaluation.

Bisection bandwidth

The total bandwidth of the links crossing the worst-case cut that splits the chips into two equal halves. It limits traffic where every chip talks to chips far away, such as all-to-all.

All three levels

Beginner. Cut a team of chips in half. This is how much data per second can cross between the two halves.

Novice. The total bandwidth of the links crossing the worst-case cut that splits the chips into two equal halves. It limits traffic where every chip talks to chips far away, such as all-to-all.

Expert. Minimum over balanced cuts of the cut capacity. Full-bisection switched fabrics give pB/2pB/2; a ring gives a constant (2 links); a k×kk \times k 2D torus gives 2k2k links per direction, twice a mesh without wraparound.

Bit line

A vertical wire shared by every cell in one column of a memory array. Data is written by driving it and read by sensing small changes on it. SRAM uses a pair per column (BL and its complement BLB); DRAM uses a single bit line per column, shared by hundreds of cells.

All three levels

Beginner. The wire that carries a bit into or out of a memory cell. Every cell in a column shares it.

Novice. A vertical wire shared by every cell in one column of a memory array. Data is written by driving it and read by sensing small changes on it. SRAM uses a pair per column (BL and its complement BLB); DRAM uses a single bit line per column, shared by hundreds of cells.

Expert. A heavily loaded line: each attached cell adds junction and wire capacitance, so bit-line capacitance grows with cells per column and dominates read delay and energy. Designs trade cells per bit line (density) against swing, speed and sense-amplifier count.

Bitcell

The circuit that stores one bit in a memory array, repeated millions of times in a grid of rows and columns. Its area sets most of the array’s area.

All three levels

Beginner. The smallest unit of a memory: the little circuit that stores one bit.

Novice. The circuit that stores one bit in a memory array, repeated millions of times in a grid of rows and columns. Its area sets most of the array’s area.

Expert. The repeated unit of an array, often drawn by the foundry with special ‘core’ design rules tighter than ordinary logic rules. Quoted in µm²; periphery (decoders, sense amplifiers, write drivers) is extra and dominates small macros.

Bitstream

The configuration file for an FPGA: every LUT bit, switch setting and block mode, plus commands for loading them. Loading a different bitstream turns the same chip into a different circuit.

All three levels

Beginner. The file that sets up an FPGA. It holds a 1 or 0 for every table and switch.

Novice. The configuration file for an FPGA: every LUT bit, switch setting and block mode, plus commands for loading them. Loading a different bitstream turns the same chip into a different circuit.

Expert. A command stream that writes configuration frames (on AMD 7 series, 101 32-bit words each) and block-RAM contents, usually with CRC checks and often encryption. Formats are vendor-specific and mostly undocumented.

Bitstream generation

The final FPGA step: turning the placed and routed design into the configuration file, by looking up which bits control each lookup table, flip-flop, switch and pin, and setting them.

All three levels

Beginner. The last step of an FPGA build: writing the file that sets every tiny memory and switch on the chip.

Novice. The final FPGA step: turning the placed and routed design into the configuration file, by looking up which bits control each lookup table, flip-flop, switch and pin, and setting them.

Expert. Serializing the configured state into the device’s configuration frames (addresses, data, CRC, commands) from the device database’s resource-to-bit map, for example 101 32-bit words per frame on Xilinx 7-series.

Block RAM (BRAM)

Dedicated blocks of SRAM, tens of kilobits each, laid out in columns across an FPGA. A design uses them for buffers, tables and FIFOs instead of building memory from LUTs.

All three levels

Beginner. Ready-made memory blocks built into an FPGA, for storing lots of numbers on the chip.

Novice. Dedicated blocks of SRAM, tens of kilobits each, laid out in columns across an FPGA. A design uses them for buffers, tables and FIFOs instead of building memory from LUTs.

Expert. Hard synchronous, usually dual-port SRAM (36 Kb on AMD 7 series, 20 Kb M20K on Altera, 18 kbit EBR on ECP5) with configurable aspect ratio, cascading and optional ECC.

Blocking assignment (=)

The = sign inside a SystemVerilog block. It updates the variable at once, so the next line already sees the new value, as in ordinary software. Used for combinational logic.

All three levels

Beginner. An instruction in a chip-design language that takes effect immediately, before the next line runs.

Novice. The = sign inside a SystemVerilog block. It updates the variable at once, so the next line already sees the new value, as in ordinary software. Used for combinational logic.

Expert. Executes immediately, in the simulator’s Active region, before the next statement. Correct inside always_comb. In a clocked block the result depends on statement order, and across blocks on which block the simulator happens to run first.

BMC (baseboard management controller)

A small, always-on processor on the server board with its own network port. It reads sensors, controls fans and power, keeps logs and offers a remote screen and keyboard, independent of the main CPUs and their operating system.

All three levels

Beginner. A tiny extra computer inside the server. It watches temperatures and power. Staff can use it to fix the machine from far away, even when it is off.

Novice. A small, always-on processor on the server board with its own network port. It reads sensors, controls fans and power, keeps logs and offers a remote screen and keyboard, independent of the main CPUs and their operating system.

Expert. An out-of-band management SoC (often Arm-based, running OpenBMC or vendor firmware) on standby power. It exposes Redfish and IPMI, mediates firmware updates, and with a hardware root of trust verifies firmware. DC-SCM moves it onto a separate module.

Body (substrate or well)

The fourth terminal: the doped silicon under the channel. NMOS bodies are usually tied to ground and PMOS bodies (an n-type well) to the supply voltage.

All three levels

Beginner. The silicon a transistor is built in. It is usually just held at one fixed voltage.

Novice. The fourth terminal: the doped silicon under the channel. NMOS bodies are usually tied to ground and PMOS bodies (an n-type well) to the supply voltage.

Expert. Its voltage relative to the source shifts the threshold (the body effect). Body ties come from well taps in the layout; deliberate body biasing tunes threshold and leakage in some processes.

Body effect

The threshold voltage rises when the source is at a higher voltage than the body, which happens in transistors stacked in series.

All three levels

Beginner. A switch gets a little harder to turn on when it sits higher up in a stack of switches.

Novice. The threshold voltage rises when the source is at a higher voltage than the body, which happens in transistors stacked in series.

Expert. VT=VT0+γ(ϕs+VSB−ϕs)V_{\mathrm{T}} = V_{\mathrm{T0}} + \gamma\left(\sqrt{\phi_{\mathrm{s}} + V_{\mathrm{SB}}} - \sqrt{\phi_{\mathrm{s}}}\right). It weakens upper transistors in series stacks and is exploited deliberately by body biasing.

Boost (turbo) clock

A clock rate above the guaranteed base rate that a processor runs at when power, current and temperature allow, for example when only a few cores are busy. Intel’s name is Turbo Boost.

All three levels

Beginner. A faster clock a chip can use for a short time, or when only a few cores are busy, because it has power and heat to spare.

Novice. A clock rate above the guaranteed base rate that a processor runs at when power, current and temperature allow, for example when only a few cores are busy. Intel’s name is Turbo Boost.

Expert. Opportunistic DVFS within the TDP: with idle cores power-gated, the busy ones get the headroom. Base clock is what all cores sustain at full load; boost is what one or a few can reach.

Bounce buffer

A staging area in the CPU’s main memory. Data read from a drive is copied there first and then copied again to the accelerator, so it crosses the wires twice.

All three levels

Beginner. A stop in the processor’s memory where data waits before moving on to where it is really going.

Novice. A staging area in the CPU’s main memory. Data read from a drive is copied there first and then copied again to the accelerator, so it crosses the wires twice.

Expert. Host-memory staging between two DMA transfers. It doubles host-memory traffic, adds latency, burns CPU cycles and makes the root complex (and the inter-socket link) part of the path. Peer-to-peer DMA such as GPUDirect Storage removes it.

Boundary scan (IEEE 1149.1, JTAG)

A standard, IEEE 1149.1, usually called JTAG. It puts a small test cell behind every pin and adds a four- or five-pin test port, so a board tester can set and read every pin of every chip one bit at a time and check the board’s wiring.

All three levels

Beginner. A small, standard test plug on a chip. It is used to check that every pin is soldered properly to the circuit board.

Novice. A standard, IEEE 1149.1, usually called JTAG. It puts a small test cell behind every pin and adds a four- or five-pin test port, so a board tester can set and read every pin of every chip one bit at a time and check the board’s wiring.

Expert. A 16-state TAP controller, an instruction register of at least 2 bits, a 1-bit bypass register and an optional 32-bit IDCODE. EXTEST, SAMPLE, PRELOAD and BYPASS are mandatory; INTEST and RUNBIST are optional. The pin map is published in BSDL.

Bounded model checking (BMC)

A formal check of the first kk clock cycles after start-up: the tool asks whether any input sequence of that length can break a rule. It finds bugs quickly but says nothing about cycle k+1k + 1 and beyond.

All three levels

Beginner. Checking every possible input for the first few steps after the chip starts up, to see if anything goes wrong.

Novice. A formal check of the first kk clock cycles after start-up: the tool asks whether any input sequence of that length can break a rule. It finds bugs quickly but says nothing about cycle k+1k + 1 and beyond.

Expert. Copies the design’s logic kk times, one copy per cycle, into one large formula and hands it to a SAT solver. Returns the shortest counterexample and needs no variable ordering. A pass is a bounded result only.

Branch divergence

When threads in one warp take different sides of an if/else, the warp runs each side in turn with the other threads switched off, wasting part of the hardware.

All three levels

Beginner. When workers on the same team need to do different things, they have to take turns, and the team slows down.

Novice. When threads in one warp take different sides of an if/else, the warp runs each side in turn with the other threads switched off, wasting part of the hardware.

Expert. Handled with an active mask and, classically, a reconvergence stack that rejoins paths at the immediate post-dominator. Lane efficiency falls to the fraction of active lanes; Volta-style per-thread PCs allow interleaving but still execute one path per issue.

Branch prediction

A branch is an instruction that chooses which instruction runs next, like an if-statement. A branch predictor guesses the outcome so the pipeline can keep fetching instructions before the branch is resolved. A wrong guess means throwing away the instructions fetched down the wrong path.

All three levels

Beginner. The processor’s guess about which way the program will go next, so the assembly line can keep moving before it knows for sure.

Novice. A branch is an instruction that chooses which instruction runs next, like an if-statement. A branch predictor guesses the outcome so the pipeline can keep fetching instructions before the branch is resolved. A wrong guess means throwing away the instructions fetched down the wrong path.

Expert. The cost of a wrong guess is roughly the number of stages from fetch to where the branch resolves, so it grows with pipeline depth. Better accuracy lowers the expected cost per stage and so moves the best pipeline depth deeper.

Branch target buffer (BTB)

A small cache that maps the address of a branch instruction to where it jumped last time. Fetch looks it up every cycle, so it can follow a taken branch before the instruction is even decoded.

All three levels

Beginner. A list that remembers where each jump in a program went last time, so the processor can follow it before it has even read the jump.

Novice. A small cache that maps the address of a branch instruction to where it jumped last time. Fetch looks it up every cycle, so it can follow a taken branch before the instruction is even decoded.

Expert. Tagged by fetch address and consulted in the first fetch stage; modern cores use multi-level BTBs with thousands of entries. Indirect branches need their own predictors, and returns use a return address stack.

Break-even volume

The production volume where two options cost the same in total. Below it, the option with the lower one-time cost (often an FPGA) is cheaper; above it, the option with the lower cost per unit (often an ASIC) wins.

All three levels

Beginner. The number of units at which making your own chip becomes cheaper than buying ready-made ones.

Novice. The production volume where two options cost the same in total. Below it, the option with the lower one-time cost (often an FPGA) is cheaper; above it, the option with the lower cost per unit (often an ASIC) wins.

Expert. N* = ΔNRE ÷ Δ(variable cost per unit), where the variable cost includes the unit price plus lifetime energy and any per-unit change costs. Uncertainty about future changes and schedule shifts it, often by more than the unit-price estimate.

Bridging fault

A fault model for a short circuit between two wires. What the shorted wires then carry depends on which one’s driver is stronger.

All three levels

Beginner. A pretend flaw where two neighboring wires accidentally touch.

Novice. A fault model for a short circuit between two wires. What the shorted wires then carry depends on which one’s driver is stronger.

Expert. Only practical when candidate pairs come from the layout, since checking every pair of nets is quadratic. Extracted from wires that run physically close, often alongside cell-aware faults for shorts inside cells.

Bring-up

The first lab work on a new chip: apply power carefully, get its clocks running, connect to its built-in test port, run its self-tests, then run its first program and finally real software. Each step uses only what the steps before it proved works.

All three levels

Beginner. The first time engineers switch on a new chip in the lab and get its parts working, one by one.

Novice. The first lab work on a new chip: apply power carefully, get its clocks running, connect to its built-in test port, run its self-tests, then run its first program and finally real software. Each step uses only what the steps before it proved works.

Expert. Ordered so each step relies only on what already works. Planned before tapeout with a board, debug access and a test plan, because something as small as a missing pull-down resistor or an unreachable debug port can stall the lab.

Buffer

A cell whose output simply copies its input with fresh strength. Inserting buffers along a long wire, or in front of many inputs, keeps the signal fast and its edges sharp.

All three levels

Beginner. A small booster part that repeats a signal at full strength so it can travel farther.

Novice. A cell whose output simply copies its input with fresh strength. Inserting buffers along a long wire, or in front of many inputs, keeps the signal fast and its edges sharp.

Expert. A non-inverting repeater (BUF_X1 … BUF_X16, or an inverter pair). Inserted by repair_design for slew, capacitance, fanout and long-wire RC limits, and by timing repair for setup; each one needs a free legal site near the right point on the net.

Built-in potential (φbi)

The voltage across a p-n junction at equilibrium, set up by the fixed charges in the depletion region. In silicon it is typically 0.6–1 V depending on doping.

All three levels

Beginner. The natural “hill” at a p-n junction that stops electrons and holes from mixing any further.

Novice. The voltage across a p-n junction at equilibrium, set up by the fixed charges in the depletion region. In silicon it is typically 0.6–1 V depending on doping.

Expert. ϕbi=(kT/q)ln⁡(NaNd/ni2)\phi_{\mathrm{bi}} = (kT/q)\ln(N_{\mathrm{a}} N_{\mathrm{d}}/n_{\mathrm{i}}^2). It cannot be read with a voltmeter, because the metal-semiconductor contact potentials around the loop cancel it exactly.

Bulk synchronous parallel (BSP)

A parallel programming model that runs in supersteps: each processor computes on its local data, then all processors exchange data, then a barrier makes everyone wait until the exchange is finished.

All three levels

Beginner. A way of organizing teamwork in rounds: everyone works alone, then everyone swaps results, then everyone waits until all are done before the next round.

Novice. A parallel programming model that runs in supersteps: each processor computes on its local data, then all processors exchange data, then a barrier makes everyone wait until the exchange is finished.

Expert. Separates compute from communication in time, so the exchange can be scheduled at compile time with no contention surprises; the cost is that compute and communication do not overlap within a superstep and every barrier waits for the slowest participant.

Bump pitch

The center-to-center distance between neighboring connection points (solder bumps, microbumps or copper bond pads) on a die. Connections per unit area grow as one over the pitch squared.

All three levels

Beginner. The spacing between the tiny metal dots that connect a chip to what it sits on. Closer dots mean more connections fit.

Novice. The center-to-center distance between neighboring connection points (solder bumps, microbumps or copper bond pads) on a die. Connections per unit area grow as one over the pitch squared.

Expert. The first-order knob for die-to-die bandwidth density. Roughly 100–130 µm on organic substrates, 25–55 µm on interposers and bridges, under 10 µm with hybrid bonding. Pitch also sets bump current capacity, so power and ground take a share of the array.

Burn-in

A stress test: chips run at raised temperature and voltage for hours so that weak ones, which would otherwise fail early in a customer’s product, fail in the factory instead.

All three levels

Beginner. Running chips hot and hard for hours, so weak ones fail in the factory instead of in a customer’s hands.

Novice. A stress test: chips run at raised temperature and voltage for hours so that weak ones, which would otherwise fail early in a customer’s product, fail in the factory instead.

Expert. Short burn-in (10–30 hours in the Delft course’s example) targets infant mortality; long burn-in runs hundreds of hours. It is expensive, so its length is balanced against reliability targets.

Bus

A shared set of wires that connects several blocks. Only one transfer can use it at a time, so it is simple and cheap, but all the blocks share its bandwidth.

All three levels

Beginner. A single shared road that all the parts of a chip take turns using.

Novice. A shared set of wires that connects several blocks. Only one transfer can use it at a time, so it is simple and cheap, but all the blocks share its bandwidth.

Expert. Every block sees every transfer, which made the bus the natural base for snooping cache coherence. As more blocks attach, waiting for a turn (arbitration) and the long, heavily loaded wires limit its speed and bandwidth.

Busbar

A solid copper bar that distributes DC power along the rack. Servers clip onto it instead of using individual power cords.

All three levels

Beginner. A thick metal bar running up the back of a rack that carries electricity to every server.

Novice. A solid copper bar that distributes DC power along the rack. Servers clip onto it instead of using individual power cords.

Expert. In Open Rack, a rear 48 V-class DC bar fed by rack-level power shelves. Distributing at ~50 V instead of 12 V cuts current fourfold and resistive loss sixteenfold for the same conductor.

Butterfly curve

One inverter’s transfer curve plotted together with the other inverter’s curve flipped across the diagonal. The two curves cross three times: two stable states and one balance point between them. The open ‘wings’ show stability.

All three levels

Beginner. A graph engineers draw of a memory cell’s two halves. It looks like a pair of wings, and bigger wings mean the cell holds its bit more firmly.

Novice. One inverter’s transfer curve plotted together with the other inverter’s curve flipped across the diagonal. The two curves cross three times: two stable states and one balance point between them. The open ‘wings’ show stability.

Expert. Plot of VTC1\mathrm{VTC}_1 and the inverse of VTC2\mathrm{VTC}_2 in the (Q, QB) plane. Each lobe corresponds to one stored state; its largest inscribed square gives that state’s SNM, and the smaller lobe sets the cell’s SNM. Under mismatch the lobes become unequal; when one closes the cell is monostable.

Cache

A small, fast memory that automatically keeps copies of recently used data. Data moves in fixed chunks called lines, often 64 bytes, and each line has a tag recording which memory address it came from. A hit (the data is there) takes a few clock cycles; a miss means fetching the line from the next, slower level.

All three levels

Beginner. A small, fast memory that keeps copies of recently used data close to the processor so it doesn’t have to fetch it from far away again.

Novice. A small, fast memory that automatically keeps copies of recently used data. Data moves in fixed chunks called lines, often 64 bytes, and each line has a tag recording which memory address it came from. A hit (the data is there) takes a few clock cycles; a miss means fetching the line from the next, slower level.

Expert. Designed by capacity, line size, associativity (how many places a given line may sit), replacement policy, write policy, and number of banks and ports. Larger, more associative caches miss less but take longer and more energy per access, which is why the first-level cache is kept small enough to answer in one or two cycles.

Cache coherence

Keeping copies of the same data in different caches in agreement. When one processor writes to an address, hardware invalidates or updates the other copies so no one reads a stale value. Each cached line carries a state, such as Modified, Shared or Invalid (the MSI protocol).

All three levels

Beginner. Keeping copies of the same data in different places in agreement, so no part of the chip reads a stale value.

Novice. Keeping copies of the same data in different caches in agreement. When one processor writes to an address, hardware invalidates or updates the other copies so no one reads a stale value. Each cached line carries a state, such as Modified, Shared or Invalid (the MSI protocol).

Expert. Snooping broadcasts every request to all caches, which doesn’t scale to many cores. A directory records which caches may hold each line and messages only those, which scales but costs storage and adds a hop through the directory. Accelerators are often attached without coherence, with software flushing caches before handing over data.

Cache line (block)

The fixed-size unit a cache stores and moves, usually 64 bytes. Each line has a tag saying which memory address it came from. Fetching a whole line at once is what makes spatial locality pay.

All three levels

Beginner. The chunk a cache moves in one go, often 64 bytes. Asking for one number brings its neighbors along too.

Novice. The fixed-size unit a cache stores and moves, usually 64 bytes. Each line has a tag saying which memory address it came from. Fetching a whole line at once is what makes spatial locality pay.

Expert. The unit of allocation, transfer and coherence. The address splits into tag, index and offset; offset bits = log2(line size). Larger lines cut compulsory misses and tag overhead but waste bandwidth on unused bytes and raise false sharing.

Cache miss (and hit rate)

A miss is an access the cache cannot answer; the line is fetched from the next level. The hit rate is hits ÷ accesses, and the miss rate is one minus that. Misses are sorted into compulsory, capacity and conflict (the three Cs).

All three levels

Beginner. When the cache doesn’t have what the core asked for, so it must be fetched from slower memory. The hit rate is how often it does have it.

Novice. A miss is an access the cache cannot answer; the line is fetched from the next level. The hit rate is hits ÷ accesses, and the miss rate is one minus that. Misses are sorted into compulsory, capacity and conflict (the three Cs).

Expert. Local miss rate counts misses per access to that cache; global miss rate per CPU access; misses per kilo-instruction (MPKI) normalizes by work. Out-of-order cores overlap misses, so stall time depends on memory-level parallelism, not just count.

Capex and opex

Capital expenditure is the up-front cost of equipment and buildings, spread over their useful life (depreciation). Operating expenditure is the ongoing cost: electricity, maintenance, staff, rent.

All three levels

Beginner. Capex is money spent once to buy something. Opex is money spent over and over to keep it running, like the power bill.

Novice. Capital expenditure is the up-front cost of equipment and buildings, spread over their useful life (depreciation). Operating expenditure is the ongoing cost: electricity, maintenance, staff, rent.

Expert. For accelerators, capex per hour falls with a longer depreciation life, but only if the hardware stays useful; opex scales with measured power, PUE and price. Idle hardware still accrues capex, so low duty cycle raises cost per unit of work.

Carrier mobility

How fast electrons or holes drift for a given electric field. In silicon, electrons (which carry NMOS current) are more mobile than holes (PMOS), so a PMOS has to be wider to carry the same current.

All three levels

Beginner. How easily electricity moves through the material of a switch. Higher mobility means a stronger, faster switch.

Novice. How fast electrons or holes drift for a given electric field. In silicon, electrons (which carry NMOS current) are more mobile than holes (PMOS), so a PMOS has to be wider to carry the same current.

Expert. μn/μp\mu_{\mathrm{n}}/\mu_{\mathrm{p}} sets the PMOS width needed to match NMOS drive: roughly 2–3 in classic planar processes. The ratio is process-dependent, so the preferred P/N width ratio differs from kit to kit.

Carry chain

Dedicated logic and wiring between neighboring logic elements that passes an adder’s carry bit along without going through the general routing, which makes adders and counters much faster and smaller.

All three levels

Beginner. A special fast wire that passes the ‘carry the one’ from each bit of an adder to the next.

Novice. Dedicated logic and wiring between neighboring logic elements that passes an adder’s carry bit along without going through the general routing, which makes adders and counters much faster and smaller.

Expert. A hard ripple (sometimes with short lookahead) path through each logic block, cascading to the next block in a column. Delay stays linear in width, but each hop is far shorter than a LUT plus routing hop.

CDU (coolant distribution unit)

The unit that separates the clean liquid flowing through servers from the building’s water. It pumps that liquid, controls its temperature, pressure and cleanliness, and moves the heat across a heat exchanger.

All three levels

Beginner. A box with pumps that sends cool, clean liquid to the computers and passes their heat on to the building’s water pipes.

Novice. The unit that separates the clean liquid flowing through servers from the building’s water. It pumps that liquid, controls its temperature, pressure and cleanliness, and moves the heat across a heat exchanger.

Expert. Five functions: temperature control, flow control, pressure control (including negative-pressure designs to limit leaks), fluid treatment, and heat exchange with isolation. Liquid-to-air CDUs reject heat into the room; liquid-to-liquid CDUs into facility water. Rack-mounted units cover tens of kW; floor-mounted units hundreds of kW to over 1 MW.

Cell characterization

Running circuit simulations of every standard cell over a grid of input transition times and output loads, at each process, voltage and temperature corner, and storing the measured delays, transitions, power and timing checks in Liberty files.

All three levels

Beginner. Measuring each building block in a library, by careful computer simulation, so the design tools know how fast it is and how much power it uses.

Novice. Running circuit simulations of every standard cell over a grid of input transition times and output loads, at each process, voltage and temperature corner, and storing the measured delays, transitions, power and timing checks in Liberty files.

Expert. Automated SPICE runs on the extracted cell netlist: function recognition, arc enumeration, delay and slew at fixed thresholds, internal power, leakage per state, pin capacitance, and setup/hold/min-pulse-width bisection. Grid choice and thresholds decide how accurate the interpolated tables are; libraries are re-characterized whenever models or layouts change.

Cell ratio (β ratio)

In a 6T SRAM cell, the strength (width) of a pull-down transistor divided by that of an access transistor. A larger ratio keeps the stored 0 from being pushed up during a read.

All three levels

Beginner. How much stronger a memory cell’s grip on its 0 is than the door that connects it to the outside wire.

Novice. In a 6T SRAM cell, the strength (width) of a pull-down transistor divided by that of an access transistor. A larger ratio keeps the stored 0 from being pushed up during a read.

Expert. CR=(W/L)pd/(W/L)acc\mathrm{CR} = (W/L)_{\mathrm{pd}} / (W/L)_{\mathrm{acc}}, commonly 1.5–2.5. It sets the read bump VREADV_{\mathrm{READ}} against the opposite inverter’s trip point, and so the read SNM.

Cell-aware test

A way of testing for defects inside the library’s building blocks (cells). Each cell’s internal layout is checked for likely shorts and breaks, each one is simulated in detail, and the result becomes a condition on the cell’s pins that test software can aim at.

All three levels

Beginner. Testing for tiny flaws inside the chip’s smallest building blocks, not only on the wires between them.

Novice. A way of testing for defects inside the library’s building blocks (cells). Each cell’s internal layout is checked for likely shorts and breaks, each one is simulated in detail, and the result becomes a condition on the cell’s pins that test software can aim at.

Expert. Each likely defect in a cell layout is SPICE-simulated, and the input conditions that expose it become 1-time-frame (static) or 2-time-frame (delay) faults at the cell pins. It catches parts that pass stuck-at and transition tests.

CFET (complementary FET)

A future transistor in which an nMOS and a pMOS device are built as one stacked structure instead of side by side. That could nearly double transistor density.

All three levels

Beginner. Stacking the two kinds of transistor every logic gate needs, one directly on top of the other, to save space.

Novice. A future transistor in which an nMOS and a pMOS device are built as one stacked structure instead of side by side. That could nearly double transistor density.

Expert. n- and p-type nanosheet stacks built vertically, either monolithically (one process, one wafer) or sequentially (bonded layers). It removes n-p spacing from cell height, enabling roughly 4-track cells, but needs contacts and power from the backside for the bottom device. Demonstrated by Intel, Samsung and TSMC at IEDM 2023.

CGRA (coarse-grained reconfigurable array)

A reconfigurable chip whose building blocks work on whole numbers (words) rather than single bits, connected by a configurable network. More efficient than an FPGA, more flexible than a fixed accelerator.

All three levels

Beginner. A chip made of a grid of small math blocks whose connections can be rewired for each job, a middle ground between a fixed chip and a fully rewirable one.

Novice. A reconfigurable chip whose building blocks work on whole numbers (words) rather than single bits, connected by a configurable network. More efficient than an FPGA, more flexible than a fixed accelerator.

Expert. Word-level functional units and memories in a static or packet-switched interconnect, configured per kernel or per program. Avoids the bit-level routing overhead of FPGAs, at the cost of supporting only the operations and data widths its units provide.

Channel (inversion layer)

A thin layer at the silicon surface under the gate where the gate’s field has gathered so many carriers of the opposite type to the body that the surface behaves like the source and drain: n-type in an NMOS, p-type in a PMOS.

All three levels

Beginner. The thin layer of electrons that a transistor’s gate pulls together under itself. It is the bridge that lets electricity cross the switch.

Novice. A thin layer at the silicon surface under the gate where the gate’s field has gathered so many carriers of the opposite type to the body that the surface behaves like the source and drain: n-type in an NMOS, p-type in a PMOS.

Expert. Forms in strong inversion once VGSV_{\mathrm{GS}} passes VTV_{\mathrm{T}}. Its charge per area is about Cox(VGS−VT)C_{\mathrm{ox}}(V_{\mathrm{GS}} - V_{\mathrm{T}}), which is what the square-law model integrates. Why the surface inverts is covered in the optional physics chapter.

Channel (link) load

The amount of traffic that must cross one link of the network, for example in bytes per image processed. The most heavily loaded link sets how fast the whole network can go.

All three levels

Beginner. How much data has to squeeze through one particular road between two cores.

Novice. The amount of traffic that must cross one link of the network, for example in bytes per image processed. The most heavily loaded link sets how fast the whole network can go.

Expert. Defined for a traffic pattern and routing function. Saturation throughput is link bandwidth divided by the maximum channel load, so placement and routing that lower the worst link raise throughput even if total traffic is unchanged.

Channel (macro channel)

The gap between two macros, or between a macro and the edge of the core. Wires to the macros’ pins have to pass through it, so it must be wide enough for them.

All three levels

Beginner. The gap between two big blocks, like an aisle, that wires must squeeze through.

Novice. The gap between two macros, or between a macro and the edge of the core. Wires to the macros’ pins have to pass through it, so it must be wide enough for them.

Expert. Sized from routing demand: pins on the facing edges, wires passing through, power straps that must cross, and any cells you allow there. Narrow channels are often blocked for placement so cells aren’t stranded without room to wire, or the macros are pushed together (abutted) to remove the channel.

Channel release

Removing the silicon-germanium layers with an etch that attacks them and barely touches silicon, so the silicon nanosheets are left suspended between source and drain, ready for the gate to wrap around them.

All three levels

Beginner. The step that eats away the in-between layers of a nanosheet stack, leaving the silicon ribbons hanging like bridges.

Novice. Removing the silicon-germanium layers with an etch that attacks them and barely touches silicon, so the silicon nanosheets are left suspended between source and drain, ready for the gate to wrap around them.

Expert. Highly selective, isotropic SiGe-to-Si etch after dummy-gate removal. Wet etches risk sheet collapse by capillary forces; hot HCl adds thermal budget (600–760 °C); remote-plasma dry etches avoid both. The gate stack is then deposited around every sheet by ALD.

Channel-length modulation (CLM)

As VDSV_{\mathrm{DS}} rises past saturation, the depleted region at the drain widens and the effective channel gets shorter, so a little more current flows. It is modeled as a factor (1+λVDS)(1 + \lambda V_{\mathrm{DS}}), where λ\lambda is fitted to measurements.

All three levels

Beginner. After the current levels off, it still creeps up a little as the push across the transistor grows.

Novice. As VDSV_{\mathrm{DS}} rises past saturation, the depleted region at the drain widens and the effective channel gets shorter, so a little more current flows. It is modeled as a factor (1+λVDS)(1 + \lambda V_{\mathrm{DS}}), where λ\lambda is fitted to measurements.

Expert. Leff=L−Ld(VDS)L_{\mathrm{eff}} = L - L_{\mathrm{d}}(V_{\mathrm{DS}}). Modeled as ID=(β/2)VGT2(1+λVDS)I_{\mathrm{D}} = (\beta/2)V_{\mathrm{GT}}^2(1 + \lambda V_{\mathrm{DS}}), λ=1/VA\lambda = 1/V_{\mathrm{A}} by analogy with the BJT Early voltage. It sets the finite output resistance ro≈1/(λID)r_{\mathrm{o}} \approx 1/(\lambda I_{\mathrm{D}}) and gets worse as L shrinks.

Characterization

Detailed measurements of timing, current and voltage levels on a statistically meaningful sample of chips, across combinations of conditions, to find the real operating limits and to build the shorter test that every production chip will get.

All three levels

Beginner. Testing a new chip at every speed, voltage and temperature it might face, to find the limits it can safely be sold with.

Novice. Detailed measurements of timing, current and voltage levels on a statistically meaningful sample of chips, across combinations of conditions, to find the real operating limits and to build the shorter test that every production chip will get.

Expert. Its outputs are data-sheet limits with a safety margin (guard band) added, the bin definitions, and a production test program cheap enough to run on every part. It continues through production to track drift and improve yield.

Charge sharing

What happens when a DRAM cell’s capacitor is connected to its bit line: the charge spreads over both capacitances, so the bit line moves only a small fraction of the cell’s voltage.

All three levels

Beginner. When a tiny bucket of charge is joined to a much bigger wire, the charge spreads out and changes the wire only a little.

Novice. What happens when a DRAM cell’s capacitor is connected to its bit line: the charge spreads over both capacitances, so the bit line moves only a small fraction of the cell’s voltage.

Expert. ΔV=(Vcell−Vpre)⋅CS/(CS+CBL)\Delta V = (V_{\mathrm{cell}} - V_{\mathrm{pre}}) \cdot C_{\mathrm{S}}/(C_{\mathrm{S}} + C_{\mathrm{BL}}). With CBLC_{\mathrm{BL}} several to tens of times CSC_{\mathrm{S}}, the transfer ratio is a few percent, so the signal is a small fraction of the cell voltage, and the sense amplifier must resolve it against its own offset and coupling noise.

Chemical-mechanical polishing (CMP)

A step that polishes the wafer flat with a rotating pad and a chemical slurry. How fast it removes material depends on how densely the layer underneath is patterned, so an uneven layout polishes unevenly.

All three levels

Beginner. Polishing the chip flat after each layer is built, a bit like sanding, so the next layer goes on a smooth surface.

Novice. A step that polishes the wafer flat with a rotating pad and a chemical slurry. How fast it removes material depends on how densely the layer underneath is patterned, so an uneven layout polishes unevenly.

Expert. Thickness after polishing follows an effective local density, a weighted average over a neighborhood set by how much the pad bends. Uneven density leaves thickness variation (dishing, where wide metal is scooped out, and erosion, where dense areas wear down), which changes wire resistance and capacitance.

Chiller

A refrigeration machine that cools water using a compressor, then rejects the heat outdoors through a cooling tower or air-cooled condenser. It is usually the biggest energy user in a cooling plant.

All three levels

Beginner. A giant air conditioner that makes cold water for the datacenter.

Novice. A refrigeration machine that cools water using a compressor, then rejects the heat outdoors through a cooling tower or air-cooled condenser. It is usually the biggest energy user in a cooling plant.

Expert. Efficiency is expressed as a coefficient of performance (heat moved ÷ electrical input) or kW/ton; it improves as the chilled-water supply temperature rises, which is why warm-water liquid cooling and economizers reduce chiller energy.

chiplet

A die (one piece of silicon cut from the wafer) designed to be assembled with other dies in a single package and connected to them by short, dense links, so together they behave like one larger chip.

All three levels

Beginner. A small chip made to be joined with others in one package, so together they act like one big chip.

Novice. A die (one piece of silicon cut from the wafer) designed to be assembled with other dies in a single package and connected to them by short, dense links, so together they behave like one larger chip.

Expert. A die whose interfaces, test and power delivery assume multi-die integration. Splitting a design into chiplets lets each piece use the process that suits it and improves yield, at the cost of die-to-die power, latency and package cost.

Clock

A signal that switches between 0 and 1 at a steady rate, for example a billion times a second (1 GHz). Each rise from 0 to 1 is a clock edge, the moment flip-flops take in new values. The time between edges is the clock period.

All three levels

Beginner. A steady tick signal that tells every memory cell on the chip when to update, like a metronome.

Novice. A signal that switches between 0 and 1 at a steady rate, for example a billion times a second (1 GHz). Each rise from 0 to 1 is a clock edge, the moment flip-flops take in new values. The time between edges is the clock period.

Expert. The time reference of a synchronous design. Every path between flip-flops must settle within one period, minus margins, so the slowest path sets the highest usable clock frequency. Chips usually have several clocks, which divides them into clock domains.

Clock buffer

A buffer is a small amplifier cell that passes its input to its output at full strength, so it can drive long wires and many inputs. Clock buffers (and clock inverters, which also flip the signal) are versions made for clock trees, with strong drive and equal rise and fall delays, such as CLKBUF_X8 or CLKINV_X4.

All three levels

Beginner. A small amplifier that re-strengthens the clock signal so it can travel further and feed more parts.

Novice. A buffer is a small amplifier cell that passes its input to its output at full strength, so it can drive long wires and many inputs. Clock buffers (and clock inverters, which also flip the signal) are versions made for clock trees, with strong drive and equal rise and fall delays, such as CLKBUF_X8 or CLKINV_X4.

Expert. Clock cells trade area and input capacitance for matched rise/fall delay, a tight delay spread under variation, and outputs that tolerate electromigration. Trees built from inverters flip the edge at every level, which keeps rise/fall mismatches from piling up into duty-cycle drift.

Clock domain

All the flip-flops (one-bit storage elements that update on each tick of a clock signal) driven by the same clock, or by clocks locked in a fixed relationship. Separate domains let blocks run at different speeds or stop independently.

All three levels

Beginner. A group of parts on a chip that all march to the same clock beat.

Novice. All the flip-flops (one-bit storage elements that update on each tick of a clock signal) driven by the same clock, or by clocks locked in a fixed relationship. Separate domains let blocks run at different speeds or stop independently.

Expert. Each unrelated (asynchronous) domain needs its own clock network and timing constraints, and every signal that crosses between unrelated domains needs a synchronizer, a handshake or an asynchronous FIFO. The domain plan is set at architecture and drives clock-tree and crossing checks later.

Clock domain crossing (CDC)

A signal sent from a flip-flop in one clock domain to one in another, unrelated domain. The receiving flip-flop can catch the signal mid-change and hang between 0 and 1 for a while (metastability), so crossings need a synchronizer circuit.

All three levels

Beginner. A signal passing between parts of the chip that run on different clocks, which needs special care so it isn’t misread.

Novice. A signal sent from a flip-flop in one clock domain to one in another, unrelated domain. The receiving flip-flop can catch the signal mid-change and hang between 0 and 1 for a while (metastability), so crossings need a synchronizer circuit.

Expert. Single bits pass through two flip-flops in a row as a synchronizer; multi-bit values use handshakes or asynchronous FIFOs with Gray-coded pointers. Crossings are checked with dedicated structural and formal CDC tools, because RTL simulation and static timing analysis don’t catch these bugs.

Clock gating

Stopping the clock signal to a group of flip-flops whenever they would just keep their current value. With no clock ticks they don’t switch, which saves power.

All three levels

Beginner. Pausing parts of the chip that have nothing to do, like turning off the lights in empty rooms.

Novice. Stopping the clock signal to a group of flip-flops whenever they would just keep their current value. With no clock ticks they don’t switch, which saves power.

Expert. Synthesis replaces the feedback multiplexers of flip-flops that share an enable with one integrated clock-gating (ICG) cell, which contains a latch so the gated clock can’t glitch. Teams set a minimum group size, check that the enable reaches the ICG in time, and connect the ICG’s test-enable pin for scan testing.

Clock mesh

A grid of clock wires, all connected together and driven by many buffers at once, with short wires down to the flip-flops. Arrival times come out very even, at a high cost in power.

All three levels

Beginner. A grid of clock wires all tied together, so every area of the chip gets the tick from several directions at once.

Novice. A grid of clock wires, all connected together and driven by many buffers at once, with short wires down to the flip-flops. Arrival times come out very even, at a high cost in power.

Expert. The grid puts branch resistances in parallel and the drivers short together, which averages out variation, so skew at the grid is near zero. Costs: wire capacitance, current wasted between drivers that switch at slightly different times, and analysis that needs circuit simulation (SPICE) because the network is not a tree.

Clock sink

An endpoint of the clock network: the clock input of a flip-flop, a latch, a memory block or a clock-gating cell. A block of a chip can have tens of thousands of them.

All three levels

Beginner. Any part of the chip that the clock has to reach.

Novice. An endpoint of the clock network: the clock input of a flip-flop, a latch, a memory block or a clock-gating cell. A block of a chip can have tens of thousands of them.

Expert. CTS balances arrival times at sink pins. Some pins are declared excluded or ignored (not balanced), and memory blocks and other macros can carry their own internal clock delay that CTS must account for.

Clock skew

The difference between the times the same clock tick reaches two flip-flops. On this page, skew = arrival at the receiving (capture) flip-flop minus arrival at the sending (launch) one, so positive skew gives the data traveling between them extra time.

All three levels

Beginner. The gap in time between the same clock tick reaching two different parts of the chip.

Novice. The difference between the times the same clock tick reaches two flip-flops. On this page, skew = arrival at the receiving (capture) flip-flop minus arrival at the sending (launch) one, so positive skew gives the data traveling between them extra time.

Expert. Global skew is the latest minus the earliest arrival over all sinks. Local skew is measured only between flip-flops joined by a data path, and only local skew decides whether timing passes. Textbooks and tools disagree on the sign, so check which arrival is subtracted.

Clock transition (slew)

How long the clock signal takes to rise from low to high (or fall back), usually measured between 10% and 90% (or 20% and 80%) of the voltage swing. CTS keeps it under a maximum.

All three levels

Beginner. How sharp the clock tick is: a crisp click or a slow fade.

Novice. How long the clock signal takes to rise from low to high (or fall back), usually measured between 10% and 90% (or 20% and 80%) of the voltage swing. CTS keeps it under a maximum.

Expert. A slow clock edge makes flip-flops respond later and raises their setup and hold requirements, wastes power while a gate’s pull-up and pull-down transistors briefly conduct together, and makes the edge’s timing more sensitive to noise. Clock nets usually get a tighter limit than data nets.

Clock tree synthesis (CTS)

The layout step, run after the logic cells have been placed, that builds the network carrying the clock to every flip-flop. It inserts and places small amplifier cells (clock buffers) and connects them so the clock tick arrives everywhere at nearly the same moment, with a sharp edge.

All three levels

Beginner. The step that builds a branching network of wires and small amplifiers to carry the clock’s tick to every memory cell on the chip.

Novice. The layout step, run after the logic cells have been placed, that builds the network carrying the clock to every flip-flop. It inserts and places small amplifier cells (clock buffers) and connects them so the clock tick arrives everywhere at nearly the same moment, with a sharp edge.

Expert. Builds each clock network (cluster the endpoints, choose a branching structure, insert and size buffers, balance branch delays) against targets for skew, latency and edge rate, then re-times the design with the real clock delays and repairs setup and hold. Commercial tools: Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler; open source: OpenROAD’s TritonCTS.

Clock uncertainty

A safety margin, set in the timing constraints, that is taken away from the time available for setup and added to the hold requirement. It covers jitter (the tick wobbling slightly from one cycle to the next) and, before the clock tree exists, the skew the tree is expected to add.

All three levels

Beginner. A safety margin for the clock tick wobbling a little from cycle to cycle.

Novice. A safety margin, set in the timing constraints, that is taken away from the time available for setup and added to the hold requirement. It covers jitter (the tick wobbling slightly from one cycle to the next) and, before the clock tree exists, the skew the tree is expected to add.

Expert. Before CTS: jitter + skew estimate + margin. After CTS the skew part is removed, because the timer now computes real skew. Setup and hold values are set separately, and inter-clock uncertainty covers paths between two clocks.

Clock-to-Q delay

The delay from a clock tick at a flip-flop to its output (called Q) showing the newly stored value. It is the first piece of delay on every path that starts at a flip-flop.

All three levels

Beginner. The short moment a memory cell needs after the clock tick before its new value appears at its output.

Novice. The delay from a clock tick at a flip-flop to its output (called Q) showing the newly stored value. It is the first piece of delay on every path that starts at a flip-flop.

Expert. Characterized in the Liberty library as a function of the clock edge rate and the output load, so a slow clock edge lengthens it. Timing reports show it as the first line of the data path.

Clos network

A multi-stage switching network, described by Charles Clos in 1953 for telephone exchanges, in which every switch of one stage connects to every switch of the next. With enough switches in the middle it can connect any free input to any free output.

All three levels

Beginner. A way to build one huge switch out of many small ones arranged in stages.

Novice. A multi-stage switching network, described by Charles Clos in 1953 for telephone exchanges, in which every switch of one stage connects to every switch of the next. With enough switches in the middle it can connect any free input to any free output.

Expert. Three-stage (or more) network of small crossbars. With nn inputs per ingress switch and mm middle switches it is strictly non-blocking for m≥2n−1m \ge 2n - 1 and rearrangeably non-blocking for m≥nm \ge n. Data-center fabrics “fold” it so ingress and egress stages are the same leaf switches.

Closed and Open divisions

In the Closed division the model and its processing must match the reference, so differences come from hardware and software. The Open division allows a different model or retraining, as long as the task and quality target are met.

All three levels

Beginner. Two kinds of MLPerf results. Closed means everyone ran the same AI model, so it compares the machines. Open lets people change the model too.

Novice. In the Closed division the model and its processing must match the reference, so differences come from hardware and software. The Open division allows a different model or retraining, as long as the task and quality target are met.

Expert. Closed inference allows post-training quantization calibrated only on the supplied calibration set, with a publicly described method, but no retraining. Inference also has a Network division for systems served over a network. A public comparison of an Open result with a Closed one must say how the Open result differs.

CMOS

Complementary metal–oxide–semiconductor: logic built from NMOS and PMOS transistors working in pairs, so that in a steady state there is no direct path from the supply to ground and almost no power is drawn.

All three levels

Beginner. The way nearly all chips are built: the two kinds of transistor work in pairs, so one is off whenever the other is on.

Novice. Complementary metal–oxide–semiconductor: logic built from NMOS and PMOS transistors working in pairs, so that in a steady state there is no direct path from the supply to ground and almost no power is drawn.

Expert. Static CMOS uses a PMOS pull-up network and an NMOS pull-down network that are logical complements. It replaced nMOS-only logic because it draws essentially only leakage when idle.

Co-packaged optics (CPO)

Optical engines mounted on the same package substrate as the switch chip, millimeters away from it. The electrical path is so short that the switch chip can use small, low-power SerDes and no DSP. Fibers run from the package to the front panel.

All three levels

Beginner. Putting the parts that send and catch light right next to the switch chip, instead of in plug-in modules at the front of the switch.

Novice. Optical engines mounted on the same package substrate as the switch chip, millimeters away from it. The electrical path is so short that the switch chip can use small, low-power SerDes and no DSP. Fibers run from the package to the front panel.

Expert. Optical engines (PIC + EIC) on the ASIC substrate, typically 3.2–6.4 Tb/s each, fed by external laser sources. Energy targets of 5–10 pJ/bit for the optics. Trades serviceability, test and supply-chain flexibility for power and density.

cocotb

An open-source framework for writing testbenches in Python, a general-purpose programming language, instead of a hardware language. The design runs in an ordinary simulator, and Python drives its inputs and checks its outputs.

All three levels

Beginner. A free toolkit for writing chip tests in Python, a popular programming language.

Novice. An open-source framework for writing testbenches in Python, a general-purpose programming language, instead of a hardware language. The design runs in an ordinary simulator, and Python drives its inputs and checks its outputs.

Expert. Connects to the simulator through its standard programming interfaces (VPI, VHPI or FLI). Tests are Python coroutines: functions that pause until the simulator reaches a given time or signal change. Python’s libraries make reference models, file handling and data analysis easy.

Code coverage

Coverage measured automatically from the design’s code: which lines ran, which way each if-statement went, which signals switched between 0 and 1, and which states of each state machine were visited.

All three levels

Beginner. A count of how much of the design’s written description the tests have actually used.

Novice. Coverage measured automatically from the design’s code: which lines ran, which way each if-statement went, which signals switched between 0 and 1, and which states of each state machine were visited.

Expert. Line, branch, condition, toggle, and FSM state and transition coverage. It measures activation only: a line holding a bug counts as covered even if the wrong value never reaches a checker. Required for sign-off, never sufficient on its own.

Cold plate

A liquid-cooled heat exchanger mounted on a chip or module. Water or coolant flows through channels in the plate and carries away far more heat than air can.

All three levels

Beginner. A metal plate with water flowing through it, pressed against a hot chip to carry the heat away.

Novice. A liquid-cooled heat exchanger mounted on a chip or module. Water or coolant flows through channels in the plate and carries away far more heat than air can.

Expert. Sized by heat flux (W/cm²), coolant flow and allowed temperature rise. It must stay in contact across the whole surface, which gets harder as the area and the thermal-expansion mismatch grow.

Collective communication

Communication operations that involve a whole group of processors at once, such as all-reduce, all-gather, reduce-scatter and all-to-all, as opposed to one sender talking to one receiver.

All three levels

Beginner. Group conversations between chips, where many chips swap data in one planned move.

Novice. Communication operations that involve a whole group of processors at once, such as all-reduce, all-gather, reduce-scatter and all-to-all, as opposed to one sender talking to one receiver.

Expert. Implemented by libraries such as NCCL as ring, tree or hierarchical schedules over the topology. Performance is usually reported as bus bandwidth, which normalizes the algorithm’s data volume to compare with link speed.

Collective operation

A communication step in which all members of a group participate together: broadcast, reduce, all-reduce, all-gather, reduce-scatter and all-to-all. Libraries such as MPI and NCCL provide them.

All three levels

Beginner. A team job where every chip takes part at once, like everyone adding their numbers together and all getting the total.

Novice. A communication step in which all members of a group participate together: broadcast, reduce, all-reduce, all-gather, reduce-scatter and all-to-all. Libraries such as MPI and NCCL provide them.

Expert. A group-wide communication pattern with a defined result (MPI and NCCL semantics). Its cost depends on the algorithm (ring, tree, recursive halving/doubling, direct) mapped onto the physical topology; libraries switch algorithms by message size and group size.

Combinational logic

Logic with no memory: its output depends only on its inputs at that moment. Adders, comparators and multiplexers (circuits that pick one of several inputs) are examples. In SystemVerilog it is written with always_comb or assign.

All three levels

Beginner. Circuitry whose output depends only on its inputs right now, with no memory.

Novice. Logic with no memory: its output depends only on its inputs at that moment. Adders, comparators and multiplexers (circuits that pick one of several inputs) are examples. In SystemVerilog it is written with always_comb or assign.

Expert. Everything between flip-flops. Signals take time to ripple through it, and the slowest path between two flops sets the shortest usable clock period. How the RTL is written (chains of if/else, operator order, bit widths) shapes what synthesis can build from it.

Common Criteria

An international standard (ISO/IEC 15408) for security evaluation. An accredited lab checks a product against a stated set of security claims and assigns an Evaluation Assurance Level (EAL) from 1 to 7.

All three levels

Beginner. A worldwide system where outside labs test how hard a product is to break into.

Novice. An international standard (ISO/IEC 15408) for security evaluation. An accredited lab checks a product against a stated set of security claims and assigns an Evaluation Assurance Level (EAL) from 1 to 7.

Expert. Smart-card chips are evaluated against a protection profile, typically EAL4 augmented with AVA_VAN.5 (resistance to attackers with high attack potential) and ALC_DVS.2 (security of the development sites).

Compact model

A set of equations with fitted parameters that a circuit simulator such as SPICE uses to compute a transistor’s currents and charges from its terminal voltages. The foundry fits the parameters to measured devices and ships them in the process design kit.

All three levels

Beginner. A set of formulas that lets a computer predict how a transistor will behave, so engineers can test circuits before building them.

Novice. A set of equations with fitted parameters that a circuit simulator such as SPICE uses to compute a transistor’s currents and charges from its terminal voltages. The foundry fits the parameters to measured devices and ships them in the process design kit.

Expert. Physics-based, all-region equations (BSIM4 for planar bulk, BSIM-CMG for FinFETs and multi-gate devices, PSP and others) with hundreds of parameters, fitted per process and corner. Continuity of currents and their derivatives across regions is essential for the simulator’s Newton iterations.

Complex gate (AOI/OAI)

A CMOS gate whose pull-down network mixes series and parallel transistors, such as AND-OR-INVERT (AOI) or OR-AND-INVERT (OAI). It computes functions like NOT((A AND B) OR C) in one stage.

All three levels

Beginner. A single gate that does a small combination of AND, OR and NOT in one step.

Novice. A CMOS gate whose pull-down network mixes series and parallel transistors, such as AND-OR-INVERT (AOI) or OR-AND-INVERT (OAI). It computes functions like NOT((A AND B) OR C) in one stage.

Expert. Also called compound gates. Any inverting function expressible as a series-parallel network can be one stage; AOI21, AOI22, OAI21 and OAI22 are library staples because they replace two or three simple gates with less area and delay. Stack depth limits how large they get.

Compliance testing

Testing a product against a standard body’s official test suite, at authorized test labs or at events where companies test their products together. Passing can earn a listing and the right to use the standard’s logo.

All three levels

Beginner. Official tests that prove a chip follows a standard, such as USB, so it works with everyone else’s devices.

Novice. Testing a product against a standard body’s official test suite, at authorized test labs or at events where companies test their products together. Passing can earn a listing and the right to use the standard’s logo.

Expert. Plan it at the spec stage: pick IP that has already passed, budget for test chips or boards, and add the test modes the standard requires (such as loopbacks) to the requirements.

Compute-bound

A workload whose run time is set by how many operations it needs, because it does more work per byte than the chip’s ratio of compute to bandwidth. More math hardware would speed it up.

All three levels

Beginner. A job is limited by math when the math parts of the chip are the slow part. They’re busy all the time, and numbers arrive fast enough.

Novice. A workload whose run time is set by how many operations it needs, because it does more work per byte than the chip’s ratio of compute to bandwidth. More math hardware would speed it up.

Expert. Arithmetic intensity above the ridge point. Its achieved rate is then capped by peak compute or by a lower ceiling (missing parallelism, the wrong instruction mix, tile or wave quantization).

Compute-in-memory (CIM)

Also called in-memory computing (IMC). The memory array stores the weights and also performs the multiply-accumulate, so weights never travel. Built from SRAM, flash, or resistive memories such as RRAM and PCM, in analog or digital form.

All three levels

Beginner. Doing math inside the memory itself. The numbers don’t have to travel to the math parts and back.

Novice. Also called in-memory computing (IMC). The memory array stores the weights and also performs the multiply-accumulate, so weights never travel. Built from SRAM, flash, or resistive memories such as RRAM and PCM, in analog or digital form.

Expert. Analog CIM sums currents or charges on bit lines and needs DACs and ADCs at the edges; digital CIM puts bitwise multipliers and adder trees beside the bit cells. Both are weight-stationary to the extreme, which limits flexibility and capacity.

Conductance drift

A slow change in a programmed resistive memory cell’s conductance after it is written. If every cell drifts in the same way it can be corrected; differences between cells cause errors.

All three levels

Beginner. Some memory cells slowly change their stored value over time, like ink fading.

Novice. A slow change in a programmed resistive memory cell’s conductance after it is written. If every cell drifts in the same way it can be corrected; differences between cells cause errors.

Expert. In PCM, G(t)=G(t0) (t/t0)−νG(t) = G(t_0)\,(t/t_0)^{-\nu} from structural relaxation of the amorphous phase; RRAM shows a fast early “relaxation.” Global per-core rescaling fixes the mean; the spread in ν\nu remains as error that grows with time.

Conduction band

The band of allowed electron energies above the band gap. In a semiconductor it is nearly empty, so the few electrons in it can move freely. Its bottom edge is written EcE_{\mathrm{c}}.

All three levels

Beginner. The upper energy level in silicon. Electrons that reach it are free to move and carry electricity.

Novice. The band of allowed electron energies above the band gap. In a semiconductor it is nearly empty, so the few electrons in it can move freely. Its bottom edge is written EcE_{\mathrm{c}}.

Expert. Electron density n=Nc e−(Ec−EF)/kTn = N_{\mathrm{c}}\, e^{-(E_{\mathrm{c}} - E_{\mathrm{F}})/kT} in the Boltzmann limit, with Nc≈2.8×1019 cm−3N_{\mathrm{c}} \approx 2.8 \times 10^{19}\,\mathrm{cm^{-3}} in Si at 300 K. Band bending (a change in EcE_{\mathrm{c}} with position) is the electrostatic potential drawn as energy.

Configuration memory

The memory cells spread across an FPGA that hold every LUT bit and switch setting. In most FPGAs they are SRAM, which forgets when the power goes off; some FPGAs use flash or one-time antifuses instead.

All three levels

Beginner. The many tiny memory cells inside an FPGA that hold its design: what each table says and which switches are on.

Novice. The memory cells spread across an FPGA that hold every LUT bit and switch setting. In most FPGAs they are SRAM, which forgets when the power goes off; some FPGAs use flash or one-time antifuses instead.

Expert. Distributed SRAM (or flash, or antifuse) cells whose outputs drive LUT masks, mux selects and block modes. Most bits control routing. SRAM configuration can be upset by radiation, mostly in routing bits.

Congestion

More wires needing to pass through an area than the metal layers there have room for. Tools divide the chip into tiles, estimate how many wires each tile needs, compare that with how many fit, and show over-full tiles as hot spots on a heat map.

All three levels

Beginner. A traffic jam of wires: a spot where more wires need to pass than there is room for.

Novice. More wires needing to pass through an area than the metal layers there have room for. Tools divide the chip into tiles, estimate how many wires each tile needs, compare that with how many fit, and show over-full tiles as hot spots on a heat map.

Expert. Routing demand above routing capacity, counted per routing tile (GCell) and per layer and reported as overflow. During placement it is estimated (RUDY or a quick global route). Typical causes: high local cell or pin density, macro corners and narrow channels, and clusters of high-fanout or many-pin cells.

Constrained-random testing

Tests whose inputs the computer picks at random, within rules (constraints) that keep them legal, such as “the length is between 1 and 64.” A starting number called the seed fixes the random sequence, so rerunning the same seed replays a failure exactly.

All three levels

Beginner. Letting the computer make up tests by rolling dice, with rules so every test still makes sense.

Novice. Tests whose inputs the computer picks at random, within rules (constraints) that keep them legal, such as “the length is between 1 and 64.” A starting number called the seed fixes the random sequence, so rerunning the same seed replays a failure exactly.

Expert. A constraint solver picks legal values. Weights and ordering controls in the test language decide how often each kind of case comes up, and coverage reports show engineers where to adjust them.

Contacted poly pitch (CPP, gate pitch)

The center-to-center distance between neighboring gates, with room for a contact between them. Together with the metal pitch it sets how small a logic cell can be.

All three levels

Beginner. The spacing between neighboring transistor gates on a chip, from the middle of one to the middle of the next.

Novice. The center-to-center distance between neighboring gates, with room for a contact between them. Together with the metal pitch it sets how small a logic cell can be.

Expert. Gate length + 2 spacers + source/drain contact width. It scaled about 0.85–0.9× per node in the FinFET era (ASAP7 assumes 54 nm) and, with M1/M2 pitch and track count, sets cell area. The ‘G’ in IRDS’s G-M-T density notation.

Context parallelism

Dividing the sequence of a very long input across accelerators for every layer, including attention. Each accelerator still needs keys and values from the others, which it receives by passing them around a ring or gathering them.

All three levels

Beginner. Splitting a very long input, such as a whole book, across chips so no single chip has to hold all of it.

Novice. Dividing the sequence of a very long input across accelerators for every layer, including attention. Each accelerator still needs keys and values from the others, which it receives by passing them around a ring or gathering them.

Expert. Partitions the full sequence across CP ranks; attention needs remote KK and VV, exchanged by ring passing overlapped with blockwise compute (Ring Attention) or by all-gather (Llama 3). Used for 100K-token-scale contexts.

Continuous batching

A serving method that rebuilds the batch at every decode step: finished requests return immediately and waiting requests join, which keeps the batch full.

All three levels

Beginner. Letting users get on and off the shared bus at every step, so nobody waits for the slowest answer to finish.

Novice. A serving method that rebuilds the batch at every decode step: finished requests return immediately and waiting requests join, which keeps the batch full.

Expert. Iteration-level scheduling, introduced in Orca. It needs a KV-cache allocator that tolerates sequences of different, growing lengths, which is what paged KV-cache managers provide.

Controllability and observability

Controllability is how easily a test can force a point inside the circuit to 0 or 1 from the chip’s pins. Observability is how easily the value at that point can be seen at an output. Scan makes every flip-flop fully controllable and observable.

All three levels

Beginner. How easy it is to set a spot inside the chip to a value, and how easy it is to see what value it holds.

Novice. Controllability is how easily a test can force a point inside the circuit to 0 or 1 from the chip’s pins. Observability is how easily the value at that point can be seen at an output. Scan makes every flip-flop fully controllable and observable.

Expert. Quantified by testability measures such as SCOAP, which ATPG uses to steer its search and which guide test-point insertion. Deep sequential state and logic that random patterns rarely exercise score poorly on both.

Convolution

A layer that applies the same small filter at every position of an image (or other grid). Image networks are built mostly from convolutions, and libraries turn each one into a matrix multiplication.

All three levels

Beginner. Sliding a small pattern detector across an image, checking each spot for that pattern.

Novice. A layer that applies the same small filter at every position of an image (or other grid). Image networks are built mostly from convolutions, and libraries turn each one into a matrix multiplication.

Expert. KK filters of size C×R×SC \times R \times S over an input of CC channels; lowering (im2col) turns it into a K×CRSK \times CRS by CRS×NPQCRS \times NPQ GEMM at the cost of duplicating inputs, or it is computed directly with the same reuse.

Convolutional neural network (CNN)

A neural network that slides small learned filters across an image, so it picks up local patterns wherever they appear. A chip layout can be cut into a grid of tiles and treated like an image.

All three levels

Beginner. A kind of learning program built for pictures, good at spotting patterns in a grid of pixels.

Novice. A neural network that slides small learned filters across an image, so it picks up local patterns wherever they appear. A chip layout can be cut into a grid of tiles and treated like an image.

Expert. Fully convolutional versions output a map the same shape as the input, such as one congestion value per layout tile, which fits layout-to-hotspot prediction. The filter sizes and depth limit how far away a feature can influence a prediction.

Cost per good die

The cost of one wafer divided by the number of working chips it gives: wafer cost÷(dies per wafer×yield)\text{wafer cost} \div (\text{dies per wafer} \times \text{yield}).

All three levels

Beginner. What one working chip costs to make, counting the ones that came out broken.

Novice. The cost of one wafer divided by the number of working chips it gives: wafer cost÷(dies per wafer×yield)\text{wafer cost} \div (\text{dies per wafer} \times \text{yield}).

Expert. The silicon part of unit cost. Add testing and packaging, then spread the one-time costs (NRE) over the number of chips sold to get unit cost. Because both dies per wafer and yield fall as area grows, cost rises faster than area.

COTS

Commercial off-the-shelf: ordinary parts made for the mass market, not designed or tested for radiation.

All three levels

Beginner. Ordinary parts that anyone can buy, like the chips in phones and laptops.

Novice. Commercial off-the-shelf: ordinary parts made for the mass market, not designed or tested for radiation.

Expert. Better performance, power and cost than radiation-hardened parts, but highly susceptible to radiation. Small satellites often use them surrounded by hardened supervisors, error correction, watchdogs and scrubbing.

Counterexample (trace)

What a formal tool returns when a rule can be broken: a specific sequence of inputs, clock cycle by clock cycle, that leads from start-up to the failure. Engineers replay it like a failing test.

All three levels

Beginner. A step-by-step recipe, found by a proof tool, that makes the design break a rule.

Novice. What a formal tool returns when a rule can be broken: a specific sequence of inputs, clock cycle by clock cycle, that leads from start-up to the failure. Engineers replay it like a failing test.

Expert. Usually written out as a waveform or a small testbench. Bounded model checking returns the shortest one. A “counterexample” from the step of an induction proof that does not start from a reachable state is not a bug; it means the proof needs strengthening.

Counterfeit part

A part that is recycled, re-marked, copied or otherwise misrepresented. It may work at first, then fail early or out of specification.

All three levels

Beginner. A fake or used chip sold as if it were new and real.

Novice. A part that is recycled, re-marked, copied or otherwise misrepresented. It may work at first, then fail early or out of specification.

Expert. Most common for parts no longer in production, bought from independent brokers. Countermeasures include buying from authorized sources, inspection and testing, and authentication built into the package, such as DARPA’s SHIELD dielet.

Coupling capacitance

Two wires next to each other, separated by insulator, form a small capacitor: a pair of conductors that store charge between them. The longer they run together and the closer they are, the larger it is, and the more a voltage change on one pushes on the other.

All three levels

Beginner. The invisible electrical link between two wires running side by side, which lets a signal on one nudge the other.

Novice. Two wires next to each other, separated by insulator, form a small capacitor: a pair of conductors that store charge between them. The longer they run together and the closer they are, the larger it is, and the more a voltage change on one pushes on the other.

Expert. Grows with parallel run length and falls with spacing. At tight pitches, same-layer coupling to the adjacent wires can exceed a wire’s capacitance to ground, which makes neighbor choice a main timing and noise lever. Extraction reports it per pair of nets in SPEF.

Coverage closure

The late phase of verification in which engineers look at every coverage item no test has reached yet, and either write a test for it, adjust the random rules, or exclude it with a reviewed reason (for example, because it can never happen).

All three levels

Beginner. The long last stretch of testing, when engineers chase down every unticked item on the checklist.

Novice. The late phase of verification in which engineers look at every coverage item no test has reached yet, and either write a test for it, adjust the random rules, or exclude it with a reviewed reason (for example, because it can never happen).

Expert. Iterative: merge coverage across every seed, rank tests by what they add, and sort each hole into missing stimulus, missing checker, unreachable, or unsupported. Every exclusion needs a written reason and a reviewer.

CPI (cycles per instruction)

Cycles per instruction: the average number of clock ticks a processor takes per instruction. Performance=clock frequency÷CPI\text{Performance} = \text{clock frequency} \div \mathrm{CPI}. Anything that makes the processor wait, such as a wrong guess about a branch or a cache miss, raises CPI. Its inverse, instructions per cycle, is called IPC.

All three levels

Beginner. On average, how many clock ticks the processor needs to finish one command.

Novice. Cycles per instruction: the average number of clock ticks a processor takes per instruction. Performance=clock frequency÷CPI\text{Performance} = \text{clock frequency} \div \mathrm{CPI}. Anything that makes the processor wait, such as a wrong guess about a branch or a cache miss, raises CPI. Its inverse, instructions per cycle, is called IPC.

Expert. Usually broken into a base CPI plus a stall term for each cause (branch mispredictions, misses at each memory level, busy units). Deeper pipelines raise frequency and CPI together, so compare designs on frequency÷CPI\text{frequency} \div \mathrm{CPI}, never on either alone.

CPLD (complex programmable logic device)

A programmable chip made of several PAL-like sum-of-products blocks joined by one programmable interconnect matrix. It keeps its setup when powered off and has short, predictable delay, but limited size.

All three levels

Beginner. An older kind of programmable chip: a few small blocks of AND and OR gates, joined by one big switch grid.

Novice. A programmable chip made of several PAL-like sum-of-products blocks joined by one programmable interconnect matrix. It keeps its setup when powered off and has short, predictable delay, but limited size.

Expert. Macrocell blocks with programmable AND arrays around a central switch matrix, usually EEPROM configured; one extra interconnect stage keeps timing near a PAL’s. Two-level logic grows costly for large multi-level circuits.

CPPR / CRPR

Common path pessimism removal (also called clock reconvergence pessimism removal). When a timing check assumes the launching clock is slow and the capturing clock fast, the stretch of clock network the two share can’t be both at once. CPPR gives back that impossible difference.

All three levels

Beginner. Taking back a safety margin that was counted twice, on the stretch of clock wiring that two memory cells share.

Novice. Common path pessimism removal (also called clock reconvergence pessimism removal). When a timing check assumes the launching clock is slow and the capturing clock fast, the stretch of clock network the two share can’t be both at once. CPPR gives back that impossible difference.

Expert. Under OCV derating, the shared clock segment is treated as late for launch and early for capture in the same check. CPPR adds the late-minus-early delay of that segment back to slack. Trees that give critical flip-flop pairs a long shared path therefore gain slack.

CPU (central processing unit)

The general-purpose processor that runs the operating system and ordinary programs. It fetches instructions from memory and carries them out; a modern chip holds several CPU cores.

All three levels

Beginner. The main chip in a computer or phone that follows a program’s commands one by one, very fast.

Novice. The general-purpose processor that runs the operating system and ordinary programs. It fetches instructions from memory and carries them out; a modern chip holds several CPU cores.

Expert. A general-purpose core optimized for single-thread latency: pipelining, branch prediction, out-of-order execution and deep caches. One die usually carries many cores sharing a last-level cache.

Credit-based flow control

The receiving end of each link tells the sender how much buffer space it has (credits). The sender spends credits as it sends and stops when they run out, so nothing is ever dropped for lack of space.

All three levels

Beginner. A rule that a sender may only send as much as the receiver has said it has room for, like tickets for parking spaces.

Novice. The receiving end of each link tells the sender how much buffer space it has (credits). The sender spends credits as it sends and stops when they run out, so nothing is ever dropped for lack of space.

Expert. Hop-by-hop, per-virtual-channel buffer accounting. Lossless like PFC but insensitive to cable delay misestimates (a long cable only lowers throughput) and it gives the sender per-VC occupancy for scheduling. Native to InfiniBand; Ultra Ethernet defines an Ethernet version (CBFC).

Critical dimension (CD)

The width of the key features on a layer, measured after lithography and after etch. Gate length is the classic CD because it sets a transistor’s speed and leakage.

All three levels

Beginner. The size of the smallest, most important shapes on a layer, such as the width of a gate.

Novice. The width of the key features on a layer, measured after lithography and after etch. Gate length is the classic CD because it sets a transistor’s speed and leakage.

Expert. Controlled through dose and focus in the scanner and through resist and etch bias, and measured on dedicated targets. CD and overlay are the two numbers each critical layer is held to.

Critical path

The timing path with the least slack, usually from one flip-flop through a long chain of gates to the next. Its delay sets the fastest clock the design can use.

All three levels

Beginner. The slowest route a signal takes through the circuit. It sets how fast the whole chip can run.

Novice. The timing path with the least slack, usually from one flip-flop through a long chain of gates to the next. Its delay sets the fastest clock the design can use.

Expert. The path with the worst slack in a path group. Optimization spends most of its effort there and recovers area everywhere else. Fixing one usually exposes the next-worst, so look at how many paths fail and by how much, not only at the worst one.

Crossbar

A switch that connects any of nn inputs directly to any of mm outputs in one step, so transfers that don’t want the same output can happen at the same time.

All three levels

Beginner. A switchboard that can connect any part of the chip directly to any other part at the same time.

Novice. A switch that connects any of nn inputs directly to any of mm outputs in one step, so transfers that don’t want the same output can happen at the same time.

Expert. Its area and wiring grow roughly with n×mn \times m connection points, so crossbars stay small: inside network routers, or joining a handful of high-bandwidth blocks. Larger systems use a network of smaller switches.

Crossbar array (for computing)

Rows and columns of wires with a programmable resistor or memory cell at each crossing. Voltages on the rows times the cells’ conductances give currents (Ohm’s law), and each column wire adds them up (Kirchhoff’s law): a matrix-vector multiply in one step.

All three levels

Beginner. A grid of wires with a memory cell where each pair crosses. Signals go in along the rows. They come out of the columns already multiplied and added up.

Novice. Rows and columns of wires with a programmable resistor or memory cell at each crossing. Voltages on the rows times the cells’ conductances give currents (Ohm’s law), and each column wire adds them up (Kirchhoff’s law): a matrix-vector multiply in one step.

Expert. Array size is bounded by IR drop, wire capacitance, sneak paths (hence 1T1R cells with a selector transistor) and the ADC’s dynamic range. Signed weights use a differential pair, w∝G+−G−w \propto G^+ - G^-. Not to be confused with an interconnect crossbar switch.

Crosstalk (signal integrity)

Interference between neighboring wires. Two wires running side by side act like a small capacitor, so when one switches it pushes on the other: it can delay or speed up the neighbor’s signal, or put a brief false blip (a glitch) on a wire that should be quiet.

All three levels

Beginner. When a busy wire nudges the wire next to it. The nudge can speed a signal up, slow it down or cause a blip.

Novice. Interference between neighboring wires. Two wires running side by side act like a small capacitor, so when one switches it pushes on the other: it can delay or speed up the neighbor’s signal, or put a brief false blip (a glitch) on a wire that should be quiet.

Expert. Coupling capacitance lets a switching aggressor net change a victim net’s delay (delta delay) or inject a glitch (noise). The effect depends on whether the two nets’ switching windows overlap and how they align, and the windows themselves depend on delays, so signal-integrity-aware STA iterates until they converge. Noise analysis checks whether a glitch is large enough to propagate or be captured.

CUDA

A programming model and toolkit, introduced in 2007, that extends C++ so programmers write a function for one thread and launch it across thousands of GPU threads. AMD’s HIP is a closely matching, portable alternative.

All three levels

Beginner. NVIDIA’s programming system for writing general-purpose programs that run on its GPUs.

Novice. A programming model and toolkit, introduced in 2007, that extends C++ so programmers write a function for one thread and launch it across thousands of GPU threads. AMD’s HIP is a closely matching, portable alternative.

Expert. Grid → block → warp → thread hierarchy, explicit shared memory and barriers, plus libraries (cuBLAS, cuDNN, CUTLASS). Its install base and libraries are a large part of why GPUs dominate AI software.

Current-source model (CCS)

A Liberty timing model that stores the cell’s output current over time instead of only a delay and a transition. Tools can rebuild the real output waveform from it, which is more accurate when wires are long and resistive.

All three levels

Beginner. A more detailed way of describing a building block’s speed, which records the shape of its output signal instead of just two numbers.

Novice. A Liberty timing model that stores the cell’s output current over time instead of only a delay and a transition. Tools can rebuild the real output waveform from it, which is more accurate when wires are long and resistive.

Expert. CCS (and the similar ECSM) store current or voltage waveforms per slew/load point, so the receiving net’s resistance and the Miller effect can be handled. It costs larger libraries and slower analysis; NLDM stays the default for synthesis and early placement.

Cutoff

The region where VGSV_{\mathrm{GS}} is below the threshold voltage. No channel forms, so the drain current is close to zero. In a simple model it is exactly zero; in a real device a small subthreshold current still flows.

All three levels

Beginner. The off state: the gate voltage is too low to open the switch, so only a tiny trickle gets through.

Novice. The region where VGSV_{\mathrm{GS}} is below the threshold voltage. No channel forms, so the drain current is close to zero. In a simple model it is exactly zero; in a real device a small subthreshold current still flows.

Expert. VGS<VtV_{\mathrm{GS}} < V_{\mathrm{t}}. The square law sets ID=0I_{\mathrm{D}} = 0 here, but the real current falls exponentially with VGSV_{\mathrm{GS}} at the subthreshold slope, which sets the device’s off-current.

CXL (Compute Express Link)

An open standard that runs over the PCIe physical wires but adds two new protocols: one lets a device cache the CPU’s memory, the other lets the CPU use a device’s memory as ordinary memory.

All three levels

Beginner. A standard that uses the same wires as PCIe to give the processor extra memory on a plug-in card.

Novice. An open standard that runs over the PCIe physical wires but adds two new protocols: one lets a device cache the CPU’s memory, the other lets the CPU use a device’s memory as ordinary memory.

Expert. Three dynamically multiplexed protocols on the PCIe PHY: CXL.io (PCIe semantics), CXL.cache (device caches host memory coherently) and CXL.mem (host load/store to device memory). Device Types 1, 2 and 3 use different subsets. 1.1 is single-host; 2.0 adds a switch level and pooling; 3.0 moves to 64 GT/s and fabrics.

CXL Type 3 device (memory expander)

A CXL device that holds memory and lets the CPU read and write it like normal memory. It adds capacity and bandwidth without adding memory channels to the CPU, at the cost of some extra delay.

All three levels

Beginner. A card full of extra memory that plugs into the fast PCIe wires instead of a memory slot.

Novice. A CXL device that holds memory and lets the CPU read and write it like normal memory. It adds capacity and bandwidth without adding memory channels to the CPU, at the cost of some extra delay.

Expert. Implements CXL.io and CXL.mem only. Appears to software as a CPU-less NUMA node. Load latency is roughly a remote-socket access or more, depending on the controller; useful for capacity tiers and pooling.

Cycle-based simulation

A faster way to simulate: compute each register’s next value once per clock tick, in an order worked out in advance, instead of tracking every small change between ticks. It cannot show delays or brief glitches.

All three levels

Beginner. A faster way to simulate a chip by computing everything once per clock tick instead of tracking every tiny change.

Novice. A faster way to simulate: compute each register’s next value once per clock tick, in an order worked out in advance, instead of tracking every small change between ticks. It cannot show delays or brief glitches.

Expert. Sorting the logic once, at compile time, in data-flow order removes most of the run-time bookkeeping of an event queue. Verilator compiles RTL into ordered C++ this way, with two-state values (0 and 1 only), evaluating logic only when its inputs may have changed.

D-algorithm

The first systematic method for finding a test for a stuck-at fault, published by J. Paul Roth in 1966. It tracks the good and the faulty circuit at once using a symbol D, pushes the fault’s effect toward an output, and works backward to find the inputs needed.

All three levels

Beginner. An early method for working out a test pattern that makes a flaw show up.

Novice. The first systematic method for finding a test for a stuck-at fault, published by J. Paul Roth in 1966. It tracks the good and the faulty circuit at once using a symbol D, pushes the fault’s effect toward an output, and works backward to find the inputs needed.

Expert. Five values (0,1,X,D,D‾0, 1, X, D, \overline{D}), where DD means good 1 / faulty 0. It searches over values on internal lines using the D-frontier and the lines still to justify, with implication and backtracking. Complete, but slow on circuits where signals split and rejoin, which PODEM and FAN addressed.

DAC (digital-to-analog converter)

Converts a digital number into an analog voltage, current or pulse length. In a compute-in-memory array, DACs turn each input into a row voltage or pulse.

All three levels

Beginner. A circuit that turns a number into an electrical signal. A bigger number gives a stronger signal.

Novice. Converts a digital number into an analog voltage, current or pulse length. In a compute-in-memory array, DACs turn each input into a row voltage or pulse.

Expert. Often avoided by feeding inputs bit-serially (a 1-bit DAC is just a driver) or as pulse widths, trading cycles for area and power. Every row needs one, so multi-bit DACs add up quickly.

Damascene process

Copper can’t be patterned by plasma etching, so the pattern is etched as trenches in the insulator instead. A barrier layer, copper electroplating and chemical-mechanical polishing leave copper only in the trenches.

All three levels

Beginner. Making copper wires by cutting grooves in an insulating layer, filling them with copper and polishing off the excess, like inlaying metal in wood.

Novice. Copper can’t be patterned by plasma etching, so the pattern is etched as trenches in the insulator instead. A barrier layer, copper electroplating and chemical-mechanical polishing leave copper only in the trenches.

Expert. Dual damascene etches the via holes and the line trenches, then fills both in one plating step. Ta/TaN barriers keep copper out of the dielectric; CMP dishing and erosion in wide or dense regions are why metal density rules exist.

Dark silicon

The share of a chip that cannot run at full speed at the same time without breaking its power budget. It appears because the number of transistors per chip has grown faster than the energy each one uses has fallen.

All three levels

Beginner. Parts of a chip that must stay switched off at any moment, because turning everything on at once would make it too hot.

Novice. The share of a chip that cannot run at full speed at the same time without breaking its power budget. It appears because the number of transistors per chip has grown faster than the energy each one uses has fallen.

Expert. Projected in 2011 at 21% of a fixed-size multicore chip at 22 nm and over 50% at 8 nm. The industry’s response has been specialization: accelerators that sit idle most of the time but do their one job very efficiently when used.

Data contamination

Test problems, or their answers, ending up in a model’s training data. The model may then score well by recall, which inflates the result.

All three levels

Beginner. When the test questions, or their answers, were already in what the AI studied, so a high score may mean it remembered instead of reasoned.

Novice. Test problems, or their answers, ending up in a model’s training data. The model may then score well by recall, which inflates the result.

Expert. Hard to rule out when the training data is undisclosed and the benchmark came from a public website. Detection uses statistical signals such as Min-K% Prob or how repetitive the model’s samples are; mitigations include held-out or freshly written problems.

Data parallelism (DP)

Each accelerator, or group of them, holds a full copy of the model and trains on a different part of each batch. At the end of every step the copies average their gradients so they stay identical.

All three levels

Beginner. Giving every chip its own copy of the model and a different slice of the examples, then having them share what they learned.

Novice. Each accelerator, or group of them, holds a full copy of the model and trains on a different part of each batch. At the end of every step the copies average their gradients so they stay identical.

Expert. Replicas of degree dd exchange gradients once per step with an all-reduce, 2(d−1)/d2(d-1)/d of the gradient buffer per rank. Communication can overlap with the backward pass; memory is not reduced unless state is sharded (ZeRO/FSDP).

DCQCN

Data Center Quantized Congestion Notification, the common congestion-control scheme for RoCEv2. Switches mark packets with ECN; the receiving NIC sends a congestion notification packet back; the sending NIC cuts its rate, then raises it again in steps.

All three levels

Beginner. A set of rules network cards follow to slow down when switches say they’re crowded and speed up again when they’re not.

Novice. Data Center Quantized Congestion Notification, the common congestion-control scheme for RoCEv2. Switches mark packets with ECN; the receiving NIC sends a congestion notification packet back; the sending NIC cuts its rate, then raises it again in steps.

Expert. Zhu et al. 2015. CP: ECN marking via RED; NP: at most one CNP per flow per 50 µs; RP: RC←RC(1−α/2)R_{\mathrm{C}} \leftarrow R_{\mathrm{C}}(1 - \alpha/2), α←(1−g)α+g\alpha \leftarrow (1 - g)\alpha + g on CNP, α\alpha decays without CNPs, then fast recovery toward RTR_{\mathrm{T}} and additive increase. Rate-based, NIC-resident, and sensitive to its many parameters.

Deadlock (network)

A state in which a set of packets each hold a buffer the next one needs, forming a cycle, so none can advance. Routing rules are designed so such cycles cannot form.

All three levels

Beginner. A jam where several pieces of data each wait for space held by another, in a circle, so none can ever move.

Novice. A state in which a set of packets each hold a buffer the next one needs, forming a cycle, so none can advance. Routing rules are designed so such cycles cannot form.

Expert. Arises from cyclic channel dependencies. Avoided by restricting turns (dimension-order or turn-model routing), by virtual channels with an escape path, or, for message-level protocol deadlock, by separating request and response classes.

Decap cell

A cell with no logic function that is simply a capacitor, a tiny store of electric charge, connected between supply and ground. When nearby logic suddenly needs current, the decap supplies it for the moment before the rest of the supply can catch up.

All three levels

Beginner. A tiny part that stores a little electricity right next to busy parts, like a water tower for the neighborhood.

Novice. A cell with no logic function that is simply a capacitor, a tiny store of electric charge, connected between supply and ground. When nearby logic suddenly needs current, the decap supplies it for the moment before the rest of the supply can catch up.

Expert. Decap answers the first and fastest dip in supply voltage, and at high frequency only helps logic within a small radius, so it goes near busy logic. It costs area and leakage, and adds to the chip capacitance that resonates with the package’s inductance.

Decode (generation)

The inference phase that produces output tokens one after another. Each step must read all the model’s weights to produce just one token per user, so it is limited by memory speed.

All three levels

Beginner. The second part of answering: the model writes its reply one word at a time, and each word depends on the ones before.

Novice. The inference phase that produces output tokens one after another. Each step must read all the model’s weights to produce just one token per user, so it is limited by memory speed.

Expert. A sequential loop of forward passes, one token per sequence each. Weight GEMMs shrink to MM = batch size, so intensity is about batch size (in FLOP/byte at 16-bit) until KV-cache reads take over.

Deemed export

Under ITAR, handing controlled technical information to a foreign person inside the U.S. counts as exporting it to every country where that person holds citizenship or permanent residency.

All three levels

Beginner. Showing secret tech to someone from another country counts as sending it abroad. It counts even if nobody leaves the U.S.

Novice. Under ITAR, handing controlled technical information to a foreign person inside the U.S. counts as exporting it to every country where that person holds citizenship or permanent residency.

Expert. It drives access control on design repositories, EDA compute servers and design reviews. Teams of mixed nationality need licenses, or partitioned data, before anyone touches controlled RTL, netlists or layout.

DEF

Design Exchange Format: a text file that records a design’s physical state: the chip outline, the rows, where each part sits and which way it faces, the pins, the blocked areas and, later, every wire.

All three levels

Beginner. A computer file that records where everything has been put on the chip.

Novice. Design Exchange Format: a text file that records a design’s physical state: the chip outline, the rows, where each part sits and which way it faces, the pins, the blocked areas and, later, every wire.

Expert. Each component carries a placement status: FIXED (automatic tools may not move it), PLACED (they may), COVER (part of a cover macro; nothing may move it) or UNPLACED. A floorplan DEF is the hand-off to placement or physical synthesis, so its units, site names and orientations must match the LEF.

Defect clustering

Killer defects are not spread perfectly at random: they gather in some regions, wafers or lots. For a given average defect density, clustering leaves more dies defect-free than a purely random spread would.

All three levels

Beginner. Flaws on a wafer tend to bunch together instead of landing evenly, so some chips get several and others none.

Novice. Killer defects are not spread perfectly at random: they gather in some regions, wafers or lots. For a given average defect density, clustering leaves more dies defect-free than a purely random spread would.

Expert. Modeled by letting the defect density itself vary (compound Poisson). A gamma distribution gives the negative binomial model with cluster parameter α\alpha; small α\alpha means strong clustering, and α\alpha of about 10 or more behaves like Poisson.

Defect density (D₀)

The average number of fatal manufacturing flaws per square centimeter of wafer. Multiply it by a chip’s area to get the expected number of flaws per chip.

All three levels

Beginner. How many tiny flaws, like specks of dust, land on each patch of a wafer, on average.

Novice. The average number of fatal manufacturing flaws per square centimeter of wafer. Multiply it by a chip’s area to get the expected number of flaws per chip.

Expert. Falls as a manufacturing process matures, so a new process starts with a higher D0D_0 than an established one. Factories track it layer by layer; in the Poisson model the per-layer yields multiply, so each layer’s contribution can be separated.

Dennard scaling

The 1974 observation that shrinking a transistor’s dimensions and its voltage by the same factor keeps the power per square millimeter constant while making circuits faster. It held until the mid-2000s.

All three levels

Beginner. The old rule that as transistors got smaller they also used proportionally less power, so chips could get faster without getting hotter.

Novice. The 1974 observation that shrinking a transistor’s dimensions and its voltage by the same factor keeps the power per square millimeter constant while making circuits faster. It held until the mid-2000s.

Expert. Scaling dimensions, voltage and current by 1/κ1/\kappa gives delay 1/κ1/\kappa, power per circuit 1/κ21/\kappa^2 and constant power density. It broke when threshold and supply voltage stopped falling, because of leakage, so energy per transistor no longer falls as fast as transistor count rises.

Density overflow

During global placement, the share of cell area that sits in over-full bins (the small squares the chip is divided into for counting). It starts near 1, with every cell piled up, and the placer stops when it falls to a target such as 0.1.

All three levels

Beginner. A measure of how much the parts are still piled on top of each other during the rough first pass.

Novice. During global placement, the share of cell area that sits in over-full bins (the small squares the chip is divided into for counting). It starts near 1, with every cell piled up, and the placer stops when it falls to a target such as 0.1.

Expert. Per bin, cell area minus capacity; summed over over-full bins and divided by total movable area. The convergence metric for Nesterov placers, and the trigger for timing-driven reweighting and routability inflation passes.

Density window

A square of fixed size over which a density rule is measured, such as “metal must cover 20% to 80% of every 50 µm window.” The checker steps the window across the whole chip, often in overlapping steps.

All three levels

Beginner. A square patch of the chip in which the factory checks how much of the area is covered with metal.

Novice. A square of fixed size over which a density rule is measured, such as “metal must cover 20% to 80% of every 50 µm window.” The checker steps the window across the whole chip, often in overlapping steps.

Expert. Checked over a fixed dissection: windows of size ww stepped by w/rw/r, so each window covers r×rr \times r tiles. A larger rr comes closer to checking every possible window position, at more runtime.

Depletion region

A region swept clear of mobile carriers, leaving the fixed, charged dopant atoms behind. Their charge creates the electric field across a p-n junction and under a MOS gate. It acts as an insulator, so it behaves like the gap in a capacitor.

All three levels

Beginner. A thin empty zone with almost no free electrons or holes, where they have been pushed away or used up.

Novice. A region swept clear of mobile carriers, leaving the fixed, charged dopant atoms behind. Their charge creates the electric field across a p-n junction and under a MOS gate. It acts as an insulator, so it behaves like the gap in a capacitor.

Expert. Width W=2εsV/(qN)W = \sqrt{2\varepsilon_{\mathrm{s}} V/(qN)} for a one-sided junction, with VV the total potential across it and NN the lighter doping. Junction capacitance is εsA/W\varepsilon_{\mathrm{s}} A/W; under a MOS gate WW saturates at WdmaxW_{\mathrm{dmax}} once the surface inverts.

Deposition

Growing or laying down a film: chemical vapor deposition (CVD) reacts gases on the wafer; physical vapor deposition (sputtering) knocks atoms off a target, mostly for metals; electroplating fills copper; atomic layer deposition builds a film one atomic layer at a time.

All three levels

Beginner. Adding a very thin, even layer of material over the whole wafer, like a coat of paint only a few atoms to a few hundred atoms thick.

Novice. Growing or laying down a film: chemical vapor deposition (CVD) reacts gases on the wafer; physical vapor deposition (sputtering) knocks atoms off a target, mostly for metals; electroplating fills copper; atomic layer deposition builds a film one atomic layer at a time.

Expert. Chosen by conformality, temperature and film quality. CVD and ALD coat sidewalls; sputtering is more line-of-sight. Spacers, liners and gate dielectrics depend on conformal films of controlled thickness.

Design for test (DFT)

Design techniques that make it cheap to create, run and judge manufacturing tests. The main ones link the chip’s internal memory cells into test chains, give memories their own built-in testers, and add a standard test port.

All three levels

Beginner. Extra parts built into a chip so a test machine can quickly check each copy for flaws.

Novice. Design techniques that make it cheap to create, run and judge manufacturing tests. The main ones link the chip’s internal memory cells into test chains, give memories their own built-in testers, and add a standard test port.

Expert. Test structures added mostly to the gate-level netlist after synthesis: scan chains and compression for logic, memory BIST, logic BIST where needed, and JTAG/IJTAG for access. Judged by fault coverage, pattern count, test time, and what the structures cost in area, timing and power.

Design rule checking (DRC)

Running the factory’s rule book over the layout to find shapes the factory cannot reliably make: wires too thin, gaps too narrow, connections not covered by enough metal, areas with too little metal, and so on.

All three levels

Beginner. Checking that every shape in the drawing follows the factory’s rules. Shapes can’t be too small or too close.

Novice. Running the factory’s rule book over the layout to find shapes the factory cannot reliably make: wires too thin, gaps too narrow, connections not covered by enough metal, areas with too little metal, and so on.

Expert. Running the foundry’s rule deck over the final merged layout, including metal fill and all third-party blocks. Modern decks are dominated by contextual rules, such as spacing that depends on a neighbor’s width or on whether a wire ends there. The result is a list of rules with violation counts and markers to step through.

Design under test (DUT)

The piece of the design being tested, placed inside a testbench. It can be one small block or the whole chip. Also called the DUV, design under verification.

All three levels

Beginner. The part of the chip that a test is checking.

Novice. The piece of the design being tested, placed inside a testbench. It can be one small block or the whole chip. Also called the DUV, design under verification.

Expert. Its boundary decides how easily a test can steer it (controllability) and see the effects (observability). Small DUTs are easier to drive into rare cases and to check, which is why most random and formal effort happens on individual blocks.

Design-space exploration (DSE)

Searching through combinations of settings, such as the target clock speed, how tightly to pack the gates and tool options, to find the ones that give the best power, speed and area.

All three levels

Beginner. Trying many combinations of choices to find the best one, like testing different recipes.

Novice. Searching through combinations of settings, such as the target clock speed, how tightly to pack the gates and tool options, to find the ones that give the best power, speed and area.

Expert. Each sample is a tool run that takes minutes to hours and gives slightly different results from run to run, so the budget is tens to hundreds of samples. Methods range from grid sweeps to Bayesian optimization, TPE, evolutionary search and RL. Results depend on how the goals are weighted and on run-to-run noise.

Detailed placement

Small improvements after legalization: swapping two cells, reordering a few neighbors in a row, or flipping a cell left to right, keeping only the changes that shorten wires and stay legal.

All three levels

Beginner. The final polish: small swaps between nearby parts that make the wires a little shorter.

Novice. Small improvements after legalization: swapping two cells, reordering a few neighbors in a row, or flipping a cell left to right, keeping only the changes that shorten wires and stay legal.

Expert. Moves such as global swap, vertical swap, local reordering of small windows, mirroring and single-segment clustering. Advanced-node detailed placers also fix cell-adjacency rule violations and pin-access conflicts.

Detailed routing

The second, exact pass of routing. It draws every wire and via as a real shape on the tracks, connects to the actual pin shapes of each cell, and repairs every broken manufacturing rule. Its output is the finished wiring, written as a DEF file.

All three levels

Beginner. The exact pass: drawing every wire in its real lane, with every connection and every manufacturing rule checked.

Novice. The second, exact pass of routing. It draws every wire and via as a real shape on the tracks, connects to the actual pin shapes of each cell, and repairs every broken manufacturing rule. Its output is the finished wiring, written as a DEF file.

Expert. Typically A* search on the track grid inside small regions, followed by many search-and-repair iterations that rip up nets near rule-violation markers and reroute them with rising costs. Getting the violation count to zero is its main runtime and schedule risk.

Deterministic execution

Execution in which the timing of every operation is known in advance. The compiler, not the hardware, decides when each piece of data moves and each instruction runs, so there are no caches or traffic jams to make the timing vary.

All three levels

Beginner. A chip that does exactly the same steps, in exactly the same time, every time you run the same job.

Novice. Execution in which the timing of every operation is known in advance. The compiler, not the hardware, decides when each piece of data moves and each instruction runs, so there are no caches or traffic jams to make the timing vary.

Expert. Achieved by removing reactive hardware (caches, arbiters, dynamic routing, replay) and exposing fixed latencies to a compiler that schedules every cycle. Multi-chip versions need aligned clocks or counters and paced, error-corrected links.

deterministic latency

Delay that doesn’t vary with load or with what happened just before, because every operation takes a fixed number of clock ticks.

All three levels

Beginner. Taking exactly the same time, every time.

Novice. Delay that doesn’t vary with load or with what happened just before, because every operation takes a fixed number of clock ticks.

Expert. Achieved by removing queues, caches, arbitration and shared buses from the critical path. Residual variation comes from clock-domain crossings, serial-link alignment and off-chip memory.

DICE flip-flop

Dual Interlocked storage Cell: a flip-flop design that holds each bit on two linked pairs of internal nodes, so if a particle disturbs one node, the others restore it.

All three levels

Beginner. A memory cell that keeps a backup copy inside itself, so one hit can’t flip it.

Novice. Dual Interlocked storage Cell: a flip-flop design that holds each bit on two linked pairs of internal nodes, so if a particle disturbs one node, the others restore it.

Expert. Cheaper than TMR in area, but one strike that disturbs two nodes at once can flip it. Effectiveness drops at highly scaled nodes, where charge spreads further, which led to variants (T-DICE, F-DICE, LEAP-DICE) and node-spacing rules.

Die

One copy of a chip design on a wafer. A wafer holds many dies side by side; they are tested, sawn apart along thin lanes, and the good ones are packaged.

All three levels

Beginner. One chip, as it sits on the wafer before it is cut out and packaged.

Novice. One copy of a chip design on a wafer. A wafer holds many dies side by side; they are tested, sawn apart along thin lanes, and the good ones are packaged.

Expert. The unit of manufacture and test. Its area sets die yield and dies per wafer; its edge length sets how many I/O signals and bumps it can carry. Plural: dies (or dice).

Die and core

The die is the whole rectangle of silicon cut from the wafer. The core is the inner rectangle where the logic goes. The border between them holds the chip’s connection pads, power wiring around the edge, or simply clearance.

All three levels

Beginner. The die is the whole tiny slice of silicon that becomes one chip. The core is the inner area that holds the working parts, with a border around it.

Novice. The die is the whole rectangle of silicon cut from the wafer. The core is the inner rectangle where the logic goes. The border between them holds the chip’s connection pads, power wiring around the edge, or simply clearance.

Expert. Set directly as rectangles (initialize_floorplan -die_area/-core_area, or DIE_AREA/CORE_AREA in OpenROAD-flow-scripts) or derived from a target utilization, aspect ratio and margin. The die may also be a right-angled polygon, such as an L shape. Core edges should line up with the placement-site and routing-track grids so rows and pins line up with them.

Die yield

The fraction of chips on a wafer with no fatal defect. It falls as chip area grows, because a bigger chip is more likely to catch one. The simplest model is Y=e−D0AY = e^{-D_0 A}, where D0D_0 is defects per cm² and AA is the chip’s area.

All three levels

Beginner. The share of chips on a wafer that come out working.

Novice. The fraction of chips on a wafer with no fatal defect. It falls as chip area grows, because a bigger chip is more likely to catch one. The simplest model is Y=e−D0AY = e^{-D_0 A}, where D0D_0 is defects per cm² and AA is the chip’s area.

Expert. Modeled with Poisson, Murphy, Seeds or negative binomial statistics, which differ in how they treat defects that bunch together. Bunching leaves more chips clean than random scattering would, so the Poisson model is pessimistic for large chips; the negative binomial’s cluster parameter α\alpha sets how strongly defects bunch.

Dies per wafer

Roughly the area of the round wafer divided by the area of one chip, minus the partial chips lost around the curved edge. Smaller chips fit more copies and waste less at the edge.

All three levels

Beginner. How many copies of a chip fit on one wafer, the thin round slice of silicon that chips are made on.

Novice. Roughly the area of the round wafer divided by the area of one chip, minus the partial chips lost around the curved edge. Smaller chips fit more copies and waste less at the edge.

Expert. A common estimate is π(d/2)2/A−πd/2A\pi (d/2)^2/A - \pi d/\sqrt{2A} for wafer diameter dd and die area AA. Real counts also subtract the scribe lanes (the gaps where the wafer is sawn), an unusable rim at the edge, and test structures.

Dimension-order (XY) routing

A routing rule for a mesh in which a packet travels along the X dimension until it is in the destination’s column, then along Y. Every packet between the same two cores takes the same path.

All three levels

Beginner. A simple rule for the roads between cores: go left or right first until you reach the right column, then go up or down.

Novice. A routing rule for a mesh in which a packet travels along the X dimension until it is in the destination’s column, then along Y. Every packet between the same two cores takes the same path.

Expert. Deterministic and minimal. Because a packet never turns from Y back to X, no cycle of channel dependencies can form, so it is deadlock-free without extra virtual channels. The price is no path diversity: it cannot steer around a hot link.

Direct-attach copper (DAC)

A passive cable of twin-axial copper pairs with a module-shaped plug on each end. The switch chip’s own SerDes drives the signal straight through it. Cheapest and lowest power, but only a meter or two at today’s speeds.

All three levels

Beginner. A plain copper cable with a plug on each end and no chips inside.

Novice. A passive cable of twin-axial copper pairs with a module-shaped plug on each end. The switch chip’s own SerDes drives the signal straight through it. Cheapest and lowest power, but only a meter or two at today’s speeds.

Expert. Passive twinax assembly (for example 802.3ck CR or 802.3dj CR PHYs). The host SerDes must equalize host traces, connectors and cable together; reach falls as the lane rate rises (about 2 m at 100G per lane, about 1 m at 200G).

Direct-to-chip liquid cooling (cold plates)

Metal cold plates with tiny internal channels sit on the hottest chips in place of air heat sinks. Water or a water-glycol mix pumps through them and carries the heat out of the rack. Other parts still need some air cooling.

All three levels

Beginner. Cooling a chip with water that flows through a metal block pressed right on top of it.

Novice. Metal cold plates with tiny internal channels sit on the hottest chips in place of air heat sinks. Water or a water-glycol mix pumps through them and carries the heat out of the rack. Other parts still need some air cooling.

Expert. Single-phase water/propylene-glycol loops (the technology cooling system) fed by a CDU; two-phase refrigerant versions exist. Captures most but not all rack heat, so residual air cooling is still needed. Enables warm supply water (ASHRAE W32–W45) and dry coolers, but needs tight control of chemistry, filtration and flow.

Directed test

A test whose inputs and expected results an engineer picks by hand, to check one known feature or tricky case.

All three levels

Beginner. A test written by hand to check one situation an engineer thought of.

Novice. A test whose inputs and expected results an engineer picks by hand, to check one known feature or tricky case.

Expert. Cheap to debug and good for first bring-up, known corner cases and checking that a fixed bug stays fixed. It covers only what its author imagined, and every spec change means editing it.

DMSMS

Diminishing Manufacturing Sources and Material Shortages: the loss, or coming loss, of the companies or materials needed to keep making a part.

All three levels

Beginner. When the companies that make a part stop making it, but the machine that needs it is still in use.

Novice. Diminishing Manufacturing Sources and Material Shortages: the loss, or coming loss, of the companies or materials needed to keep making a part.

Expert. A program-level discipline in the U.S. Defense Department. For chips it means lifetime buys, redesign onto a newer process, emulation, or portable RTL and documentation that let a part be rebuilt decades later.

Domain-adaptive pretraining (DAPT)

Continuing to train a general language model on one field’s documents and code, so it learns that field’s vocabulary and facts, before teaching it to follow instructions.

All three levels

Beginner. Giving a general AI model extra study on one field’s documents, such as chip-design manuals and code, so it knows that field better.

Novice. Continuing to train a general language model on one field’s documents and code, so it learns that field’s vocabulary and facts, before teaching it to follow instructions.

Expert. Costs a small fraction of the original pretraining compute but needs a large, clean domain corpus, which in chip design is proprietary. Applied to a model already tuned for chat, it can wreck the model’s instruction following, so the instruction tuning is done afterward.

Doping

Introducing impurity atoms into silicon. Donors such as phosphorus or arsenic add free electrons (n-type); acceptors such as boron add holes (p-type). Heavily doped regions are marked n+ or p+.

All three levels

Beginner. Mixing a tiny pinch of another element into silicon so that it carries electricity in a controlled way.

Novice. Introducing impurity atoms into silicon. Donors such as phosphorus or arsenic add free electrons (n-type); acceptors such as boron add holes (p-type). Heavily doped regions are marked n+ or p+.

Expert. Doping sets the body’s carrier concentration and hence the threshold voltage; separate threshold-adjust implants produce the low-, standard- and high-VTV_{\mathrm{T}} device flavors of one process.

DPPM

Defective parts per million: how many faulty chips reach customers for every million shipped. Car makers track it closely, with a goal of zero.

All three levels

Beginner. How many bad chips slip through out of every million shipped.

Novice. Defective parts per million: how many faulty chips reach customers for every million shipped. Car makers track it closely, with a goal of zero.

Expert. Driven down by high test coverage, outlier screening such as Part Average Testing, burn-in or stress screens, and design for test and analysis.

Drain current (ID)

The current flowing from drain to source through the channel, written IDI_{\mathrm{D}} or IDSI_{\mathrm{DS}}. It depends on the gate-to-source voltage VGSV_{\mathrm{GS}} and the drain-to-source voltage VDSV_{\mathrm{DS}}.

All three levels

Beginner. How much electricity flows through a transistor, from one end of the switch to the other.

Novice. The current flowing from drain to source through the channel, written IDI_{\mathrm{D}} or IDSI_{\mathrm{DS}}. It depends on the gate-to-source voltage VGSV_{\mathrm{GS}} and the drain-to-source voltage VDSV_{\mathrm{DS}}.

Expert. ID(VGS,VDS,VBS)I_{\mathrm{D}}(V_{\mathrm{GS}}, V_{\mathrm{DS}}, V_{\mathrm{BS}}), usually quoted per micrometer of width because it scales linearly with WW. The function is what every MOSFET model, from the square law to BSIM, tries to reproduce.

Drain-induced barrier lowering (DIBL)

In a short channel the drain is close enough to the source that its voltage lowers the energy barrier the gate is supposed to control. The threshold voltage falls as VDSV_{\mathrm{DS}} rises, raising both on-current and leakage.

All three levels

Beginner. In very tiny transistors, the voltage at the far end helps open the switch a little, so it leaks more when off.

Novice. In a short channel the drain is close enough to the source that its voltage lowers the energy barrier the gate is supposed to control. The threshold voltage falls as VDSV_{\mathrm{DS}} rises, raising both on-current and leakage.

Expert. Modeled as Vt=Vt0−ηVDSV_{\mathrm{t}} = V_{\mathrm{t0}} - \eta V_{\mathrm{DS}}, with η\eta (DIBL) quoted in mV/V. On a log IDI_{\mathrm{D}}–VGSV_{\mathrm{GS}} plot it shows as a sideways shift between low-VDSV_{\mathrm{DS}} and high-VDSV_{\mathrm{DS}} curves, and it multiplies off-current by 10ηΔVDS/S10^{\eta \Delta V_{\mathrm{DS}}/S}.

DRAM

Dynamic RAM, a computer’s main memory. It stores each bit as charge on a tiny capacitor, which makes it dense and cheap, and it sits on separate chips or is stacked in the same package. Reaching it takes tens to over a hundred nanoseconds, many processor clock cycles. Its peak bandwidth is the width of the connection times its data rate.

All three levels

Beginner. The chip’s big main memory, usually on separate chips. It holds a lot but takes a long time to reach.

Novice. Dynamic RAM, a computer’s main memory. It stores each bit as charge on a tiny capacitor, which makes it dense and cheap, and it sits on separate chips or is stacked in the same package. Reaching it takes tens to over a hundred nanoseconds, many processor clock cycles. Its peak bandwidth is the width of the connection times its data rate.

Expert. For the architect it is a budget of bandwidth and latency. Sustained bandwidth is well below the datasheet peak because the memory pauses to refresh its cells, requests collide on the same internal bank, and the bus loses time switching between reads and writes. The controller and its PHY (the circuits that drive the off-chip wires) also take die edge and power.

Drive strength

The variants of one logic function built with wider transistors (in SKY130: inv_1, inv_2, inv_4 and so on). A stronger cell drives more load capacitance with less delay, but takes more area and presents more input capacitance to whatever drives it.

All three levels

Beginner. How strong a building block is at pushing its signal: a stronger one switches a big load faster, but it is bigger and uses more power.

Novice. The variants of one logic function built with wider transistors (in SKY130: inv_1, inv_2, inv_4 and so on). A stronger cell drives more load capacitance with less delay, but takes more area and presents more input capacitance to whatever drives it.

Expert. Wider (or more fingered) devices lower the cell’s effective output resistance, which is the slope of delay against load in its NLDM table. Families share a Liberty cell_footprint so optimizers can swap sizes in place. Upsizing one cell slows its driver, which is why sizing is a path-level problem.

DSP (in an optical module)

A digital signal processor that retimes and equalizes the signal twice: once for the electrical link from the switch chip, and once for the optical link. It makes each half of the link independent, but it is the hungriest part of the module.

All three levels

Beginner. The chip inside an optical module that cleans up the signal on both sides, like a translator who also fixes typos.

Novice. A digital signal processor that retimes and equalizes the signal twice: once for the electrical link from the switch chip, and once for the optical link. It makes each half of the link independent, but it is the hungriest part of the module.

Expert. A full retimer per lane on the host side plus PAM4 transmit/receive DSP (ADC, FFE, sometimes MLSE) on the line side, with gearboxing and FEC monitoring. Often around half of module power. Removing it is the point of LPO.

DSP block

A hard arithmetic block with a multiplier, adders and an accumulator, used for filters, signal processing and neural-network math.

All three levels

Beginner. A built-in multiplying machine in an FPGA, much faster and smaller than one made from small logic blocks.

Novice. A hard arithmetic block with a multiplier, adders and an accumulator, used for filters, signal processing and neural-network math.

Expert. A hard multiply-accumulate slice (for example 25 × 18 with a 48-bit ALU, or 18 × 18), with input and pipeline registers and cascade paths down its column.

DTCO (design-technology co-optimization)

Choosing process rules (pitches, fin counts, metal layers) by laying out real cells and memories and measuring the result, instead of shrinking every dimension and hoping the cells still fit.

All three levels

Beginner. Designing the factory process and the circuit building blocks together, so each choice helps the other.

Novice. Choosing process rules (pitches, fin counts, metal layers) by laying out real cells and memories and measuring the result, instead of shrinking every dimension and hoping the cells still fit.

Expert. Iterating process assumptions against standard-cell, SRAM and routing results: fin-to-metal gear ratio, track height, diffusion breaks, contact schemes. ASAP7’s rules came from DTCO on a 7.5-track cell and SRAM bitcells. Much of each ‘node’s’ density gain now comes from DTCO, not from a smaller gate.

DVFS

Dynamic voltage and frequency scaling: lowering the chip’s supply voltage and clock speed while it runs, whenever the work allows. Switching power grows with the square of voltage, so a small voltage cut saves a lot.

All three levels

Beginner. Slowing the chip down, and turning down the electrical push that runs it, when there is little work. It saves battery.

Novice. Dynamic voltage and frequency scaling: lowering the chip’s supply voltage and clock speed while it runs, whenever the work allows. Switching power grows with the square of voltage, so a small voltage cut saves a lot.

Expert. Needs characterized voltage–frequency operating points, regulators that can change voltage on command, clock switching or PLL relock, and timing signoff at every supported pair, often separately per domain.

Dynamic power

Power spent each time signals switch, as tiny amounts of charge fill and drain wires and transistor inputs. P=αCV2fP = \alpha C V^2 f: how often signals switch (α\alpha), how much capacitance they charge (CC), the supply voltage squared (V2V^2) and the clock rate (ff).

All three levels

Beginner. The power a chip uses each time its tiny switches flip. More switching, faster, means more power.

Novice. Power spent each time signals switch, as tiny amounts of charge fill and drain wires and transistor inputs. P=αCV2fP = \alpha C V^2 f: how often signals switch (α\alpha), how much capacitance they charge (CC), the supply voltage squared (V2V^2) and the clock rate (ff).

Expert. The V2V^2 term makes voltage the strongest lever. Short-circuit current, which flows while both transistors of a gate are briefly on, adds a little more. Clock gating (stopping the clock to idle blocks) cuts α\alpha, smaller cells and shorter wires cut CC, and DVFS lowers VV and ff together when full speed isn’t needed.

Dynamic range

The ratio between the largest and smallest non-zero values a format can represent, often counted in powers of two (binades). FP32 and BF16 cover about 2−1262^{-126} to 21282^{128}; FP16 only about 2−142^{-14} to 2162^{16} (down to 2−242^{-24} with subnormals).

All three levels

Beginner. The span from the smallest to the largest number a format can store.

Novice. The ratio between the largest and smallest non-zero values a format can represent, often counted in powers of two (binades). FP32 and BF16 cover about 2−1262^{-126} to 21282^{128}; FP16 only about 2−142^{-14} to 2162^{16} (down to 2−242^{-24} with subnormals).

Expert. Usually quoted in binades (octaves) from the smallest subnormal to the largest normal: 18 for OCP E4M3 and 32 for E5M2. It decides whether a tensor fits without a scale factor; a scale only shifts the window, it doesn’t widen it.

EAR

Export Administration Regulations: U.S. Commerce Department rules for items with both civilian and military uses. Each item gets a classification code (an ECCN) on the Commerce Control List.

All three levels

Beginner. U.S. rules for sending abroad things that have both everyday and military uses.

Novice. Export Administration Regulations: U.S. Commerce Department rules for items with both civilian and military uses. Each item gets a classification code (an ECCN) on the Commerce Control List.

Expert. Also covers the less sensitive “600 series” military items; ASICs programmed for them fall under ECCN 3A611.f. Radiation-hardened ICs above set thresholds are controlled under 3A001.a.1.

ECMP (equal-cost multipath)

Equal-cost multipath routing: a switch with several equally good next hops hashes fields of each packet’s header (addresses and ports) to pick one. All packets of a flow hash the same way, so they stay in order on one path.

All three levels

Beginner. When there are several equally short routes, the switch picks one for each conversation by a fixed rule. That spreads the traffic out.

Novice. Equal-cost multipath routing: a switch with several equally good next hops hashes fields of each packet’s header (addresses and ports) to pick one. All packets of a flow hash the same way, so they stay in order on one path.

Expert. Static per-flow hashing over the 5-tuple (or a dedicated entropy field). Balanced only in expectation: with few large flows, collisions leave some links doubled up and others idle, and repeated hashing at each tier can correlate badly (hash polarization).

ECN (explicit congestion notification)

Explicit congestion notification: switches set a bit in the IP header of packets that pass through a queue above a threshold. The receiver reports the marks back, and the sender lowers its rate.

All three levels

Beginner. A mark a busy switch puts on passing data to say “I’m getting crowded.” The sender then slows down before anything is lost.

Novice. Explicit congestion notification: switches set a bit in the IP header of packets that pass through a queue above a threshold. The receiver reports the marks back, and the sender lowers its rate.

Expert. Two IP-header bits (RFC 3168). Data-center schemes mark on instantaneous queue depth (RED with Kmin/Kmax, or a single threshold); DCQCN relays marks as CNPs, and Ultra Ethernet asks switches to mark at dequeue rather than enqueue for a fresher signal.

ECO (engineering change order)

A small, targeted change to a nearly finished design, such as swapping one gate for a stronger version or adding a delay, made without rerunning the whole layout process.

All three levels

Beginner. A small fix to a nearly finished design, instead of redoing everything.

Novice. A small, targeted change to a nearly finished design, such as swapping one gate for a stronger version or adding a delay, made without rerunning the whole layout process.

Expert. Timing ECOs size, buffer or swap cells; functional ECOs patch logic; metal-only ECOs rewire spare cells placed earlier, so only metal and via masks change. Every ECO is followed by re-extraction, STA, physical verification and equivalence checking, because a fix for one check can break another.

Economizer (free cooling)

A cooling mode that rejects heat without a refrigeration compressor: either by bringing in filtered outside air (airside) or by cooling water in outdoor towers or dry coolers (waterside). The warmer the IT equipment can run, the more hours a year it works.

All three levels

Beginner. Using cool outside air or water to cool the building, instead of big air conditioners, whenever the weather allows.

Novice. A cooling mode that rejects heat without a refrigeration compressor: either by bringing in filtered outside air (airside) or by cooling water in outdoor towers or dry coolers (waterside). The warmer the IT equipment can run, the more hours a year it works.

Expert. Airside and waterside economizers, often with evaporative (adiabatic) assist, cut chiller hours and PUE. Warm-water liquid cooling extends economizer hours, but evaporative assist trades energy for water use (higher WUE).

EDA (electronic design automation)

Software for designing and checking electronic systems. In chip design, a chain of EDA programs takes the design from a description written in code to the final layout drawing, checking it at every step.

All three levels

Beginner. The software engineers use to design chips. It turns a written plan into a drawing a factory can build. It also checks the work.

Novice. Software for designing and checking electronic systems. In chip design, a chain of EDA programs takes the design from a description written in code to the final layout drawing, checking it at every step.

Expert. Commercial EDA comes mainly from three vendors whose suites cover most steps and share data internally; open-source EDA centers on Yosys, OpenROAD, OpenSTA, KLayout and Magic. Which tools a team can use depends as much on what the chip factory supports and trusts as on how good the algorithms are.

EDA license

Paid permission to run a commercial tool. A company buys a number of licenses, and each running copy of the tool checks one out from a license server, so the count limits how many jobs can run at the same time. Licenses are often the largest tool cost.

All three levels

Beginner. Paid permission to use a design program. Companies buy a set number, so only that many copies can run at once.

Novice. Paid permission to run a commercial tool. A company buys a number of licenses, and each running copy of the tool checks one out from a license server, so the count limits how many jobs can run at the same time. Licenses are often the largest tool cost.

Expert. License counts shape methodology: how many corners (combinations of manufacturing variation, voltage and temperature) can be analyzed in parallel, how many exploratory runs fit overnight, and whether renting extra cloud machines at busy times helps. Open-source tools remove this cap, so compute becomes the limit.

EDAC

Error detection and correction: extra check bits stored alongside each word of memory, computed so the chip can tell when a bit has flipped and put it right.

All three levels

Beginner. Extra check bits stored with data, so the chip can spot and fix mistakes.

Novice. Error detection and correction: extra check bits stored alongside each word of memory, computed so the chip can tell when a bit has flipped and put it right.

Expert. Common codes are Hamming SECDED (single-error correct, double-error detect) and Reed-Solomon, which corrects bursts. Pair it with scrubbing so a single error is fixed before a second hit in the same word makes it uncorrectable.

Effective width (W_eff)

The total width of channel surface the gate controls, which sets how much current a transistor can carry. For a flat transistor it is the drawn width; for a fin it is two sides plus the top; for a nanosheet it is the whole outline of every sheet.

All three levels

Beginner. How much of the channel’s edge the gate touches, added up. More of it lets more electricity flow.

Novice. The total width of channel surface the gate controls, which sets how much current a transistor can carry. For a flat transistor it is the drawn width; for a fin it is two sides plus the top; for a nanosheet it is the whole outline of every sheet.

Expert. Gated perimeter per device: WW (planar), nfin(2Hfin+Wfin)n_{\mathrm{fin}}(2H_{\mathrm{fin}} + W_{\mathrm{fin}}) (tri-gate FinFET), nsh⋅2(w+t)n_{\mathrm{sh}} \cdot 2(w + t) (nanosheet). Drive current and intrinsic gate capacitance both scale with it, so the useful figure is WeffW_{\mathrm{eff}} per footprint.

eFPGA (embedded FPGA)

FPGA fabric licensed as a block (IP) and placed inside an ASIC or SoC. Most of the chip stays fixed and efficient; the eFPGA holds the part of the job that may change after tapeout.

All three levels

Beginner. A small rewirable FPGA part built inside a custom chip, so that one corner of the chip can still change.

Novice. FPGA fabric licensed as a block (IP) and placed inside an ASIC or SoC. Most of the chip stays fixed and efficient; the eFPGA holds the part of the job that may change after tapeout.

Expert. Delivered as hard or soft IP with its own configuration interface and an FPGA tool flow. Costs fabric area at FPGA density for the volatile function only. Open generators such as FABulous use Yosys and nextpnr for the user flow.

Elaboration

The first step of synthesis. The tool reads the design code, fills in settings such as bus widths, connects the modules, and builds a generic circuit of adders, comparators, selectors and storage elements that isn’t tied to any library yet.

All three levels

Beginner. The step where the tool reads the design description and builds its own picture of it.

Novice. The first step of synthesis. The tool reads the design code, fills in settings such as bus widths, connects the modules, and builds a generic circuit of adders, comparators, selectors and storage elements that isn’t tied to any library yet.

Expert. Resolving parameters and generate blocks, building the module hierarchy, and turning each always block into word-level operators, multiplexers and inferred flip-flops, latches or memories. The result is a technology-independent netlist. Unintended latches and width mismatches show up here first, so elaboration warnings get read before anything else.

Electrical rule check (ERC)

Checks for connections that make no electrical sense, such as a gate input connected to nothing (it would drift unpredictably), two outputs wired together, or a piece of the silicon that isn’t tied to the supply as it must be.

All three levels

Beginner. Checking for wiring mistakes that make no sense, like an input joined to nothing.

Novice. Checks for connections that make no electrical sense, such as a gate input connected to nothing (it would drift unpredictably), two outputs wired together, or a piece of the silicon that isn’t tied to the supply as it must be.

Expert. Checks on extracted connectivity: floating gates, shorted outputs, wells without taps, and gates tied directly to a supply rather than through a tie cell. Usually run with the LVS deck. Designs with several supplies produce false reports unless the deck knows which supply each region uses.

Electromigration (EM)

Slow wear-out of a wire caused by the current itself: over years, flowing electrons push metal atoms along, leaving thin spots or gaps that can eventually break the wire. Checks make sure no wire carries more current than its width allows.

All three levels

Beginner. Wires slowly wearing out, because the flowing electricity pushes metal atoms along.

Novice. Slow wear-out of a wire caused by the current itself: over years, flowing electrons push metal atoms along, leaving thin spots or gaps that can eventually break the wire. Checks make sure no wire carries more current than its width allows.

Expert. Current-driven migration of metal atoms that forms voids (and eventually opens) or extrusions. Checked against per-layer current-density limits (average, RMS and peak). Rule-based checks build on Black’s equation and the Blech short-length effect. Power wires carry mostly one-way current; signal and clock wires carry two-way current with partial recovery, so their limits differ. Temperature matters strongly.

Electrostatic (analytical) placement

Global placement that treats cells as electric charges that push each other apart, balanced against a smooth wirelength score that pulls connected cells together. All cells move a little at a time in the direction that improves the combined score.

All three levels

Beginner. A way of spreading parts out by having them push each other away, like magnets, while wires pull connected parts together.

Novice. Global placement that treats cells as electric charges that push each other apart, balanced against a smooth wirelength score that pulls connected cells together. All cells move a little at a time in the direction that improves the combined score.

Expert. ePlace’s eDensity: charge = cell area, density penalty = potential energy, field from Poisson’s equation solved by FFT, objective W+λNW + \lambda N minimized with Nesterov’s method. Used in RePlAce, OpenROAD gpl and DREAMPlace.

Elementwise operation

An operation that produces each output value from the matching input value(s) alone, such as an activation function or adding two tensors. It does about one operation per value read.

All three levels

Beginner. A simple step done to each number on its own, such as turning negatives into zero.

Novice. An operation that produces each output value from the matching input value(s) alone, such as an activation function or adding two tensors. It does about one operation per value read.

Expert. Intensity well under 1 FLOP/byte, so always memory-bound unless fused into a neighboring kernel. Reductions (softmax, layer norm) are only slightly better.

Elephant flow

A flow that carries a large amount of data for a long time, often at the full speed of the link. AI training produces a few elephant flows per NIC rather than thousands of small ones.

All three levels

Beginner. One very large, long transfer of data, as opposed to many tiny ones.

Novice. A flow that carries a large amount of data for a long time, often at the full speed of the link. AI training produces a few elephant flows per NIC rather than thousands of small ones.

Expert. Long-lived, high-rate flows. Hash-based balancing assumes many small flows that average out; a handful of elephants per link turns every hash collision into a lasting bottleneck.

Elmore delay

A quick estimate of signal delay through a network of wires: for each piece of wire, multiply its resistance by all the capacitance beyond it, and add up the results. Because both grow with length, an unbuffered wire’s delay grows roughly with the square of its length.

All three levels

Beginner. A quick way to estimate how long a signal takes to travel along a wire. Delay grows quickly with length: a wire twice as long takes about four times as long.

Novice. A quick estimate of signal delay through a network of wires: for each piece of wire, multiply its resistance by all the capacitance beyond it, and add up the results. Because both grow with length, an unbuffered wire’s delay grows roughly with the square of its length.

Expert. The first moment of an RC tree’s impulse response: cheap, additive along a path and “faithful” (lowering it almost always lowers the true delay), so good for comparing options, but not accurate enough for signoff. It ignores the input edge shape, driver nonlinearity and inductance.

Emulation

Running the design on a dedicated machine, built from many FPGAs or from arrays of custom processors, that executes it in hardware at up to a few million clock cycles per second. That is fast enough to boot an operating system before the chip exists.

All three levels

Beginner. Running a chip design on a big, special machine. It is fast enough to start real software, but slower than the finished chip.

Novice. Running the design on a dedicated machine, built from many FPGAs or from arrays of custom processors, that executes it in hardware at up to a few million clock cycles per second. That is fast enough to boot an operating system before the chip exists.

Expert. Trades compile time, visibility and price against speed. Processor-based emulators compile quickly and can record every signal; FPGA-based ones compile slowly and need a recompile to watch different signals. Everything mapped onto the machine must be synthesizable.

Endcap (boundary) cell

A small cell with no logic placed at both ends of every row and along the edges of large blocks, to finish off the rows.

All three levels

Beginner. A cap placed at the start and end of every row, like bookends, to finish the row properly.

Novice. A small cell with no logic placed at both ends of every row and along the edges of large blocks, to finish off the rows.

Expert. LEF CLASS ENDCAP PRE/POST. Inserted with the taps before placement and FIXED. Advanced nodes add corner and edge variants around macros and row ends.

Energy per bit (pJ/bit)

The energy a link spends per transferred bit, in picojoules (10−1210^{-12} joules). Power in watts = bits per second × joules per bit, so 1 Tb/s at 1 pJ/bit costs 1 W.

All three levels

Beginner. How much energy it takes to send one bit, a single 1 or 0, across a link. Less energy per bit means a cooler, cheaper link.

Novice. The energy a link spends per transferred bit, in picojoules (10−1210^{-12} joules). Power in watts = bits per second × joules per bit, so 1 Tb/s at 1 pJ/bit costs 1 W.

Expert. Covers transmitter, receiver, clocking and the wire’s capacitance; shorter, finer wires need less. Targets fall from about 0.5 pJ/b on an organic substrate to 0.25 on an interposer and under 0.05 with hybrid bonding; SerDes built to cross a circuit board spend far more.

ENOB (effective number of bits)

The resolution a converter actually achieves given its noise and distortion. A converter with 10 output bits might deliver only 8 effective bits.

All three levels

Beginner. How many bits of real detail a measurement has once noise is counted. It is often fewer than the label says.

Novice. The resolution a converter actually achieves given its noise and distortion. A converter with 10 output bits might deliver only 8 effective bits.

Expert. ENOB=(SINAD−1.76 dB)/6.02 dB\mathrm{ENOB} = (\mathrm{SINAD} - 1.76\,\mathrm{dB}) / 6.02\,\mathrm{dB}. In analog MACs it is the right way to compare the whole signal chain, devices included, to a digital datapath of a given width.

EOT (equivalent oxide thickness)

The thickness of plain silicon dioxide that would give the same gate capacitance as the real gate insulator. High-k insulators let the physical layer stay thicker (less leakage) while the EOT keeps shrinking.

All three levels

Beginner. How thin the insulating layer under the gate effectively is. Thinner gives the gate a stronger grip.

Novice. The thickness of plain silicon dioxide that would give the same gate capacitance as the real gate insulator. High-k insulators let the physical layer stay thicker (less leakage) while the EOT keeps shrinking.

Expert. tphys⋅εSiO2/εhigh-kt_{\mathrm{phys}} \cdot \varepsilon_{\mathrm{SiO_2}} / \varepsilon_{\text{high-}k} plus interfacial layer. It sets CoxC_{\mathrm{ox}}, so it enters both the body factor and λ\lambda. Around 1 nm in modern high-k/metal-gate stacks; going lower raises gate tunneling and mobility loss, which is part of why architecture (NN) took over as the scaling knob.

Equalization

Circuits that compensate for a channel’s loss. Because high frequencies fade more, an equalizer boosts them (or subtracts the smeared-out echoes of earlier bits) so the receiver can tell the levels apart again.

All three levels

Beginner. Undoing the blurring a long wire does to a signal, by boosting the parts that faded the most.

Novice. Circuits that compensate for a channel’s loss. Because high frequencies fade more, an equalizer boosts them (or subtracts the smeared-out echoes of earlier bits) so the receiver can tell the levels apart again.

Expert. Transmit FFE, receive CTLE, and DFE or ADC-plus-DSP FFE/MLSE. More loss to equalize means more taps, faster ADCs and more power; noise is boosted along with the signal, so loss budgets cannot grow indefinitely.

Errata

A document listing the known bugs in a chip version that has shipped, what triggers each one, and how to avoid it, usually with a change in software or firmware.

All three levels

Beginner. A published list of known mistakes in a chip, with tips for working around them.

Novice. A document listing the known bugs in a chip version that has shipped, what triggers each one, and how to avoid it, usually with a change in software or firmware.

Expert. Each erratum records a decision not to fix the silicon yet. Workarounds range from driver and compiler changes to microcode or configuration patches; a later revision of the chip may fix them.

essential performance

In the medical-equipment safety standard IEC 60601-1, the functions whose loss or degradation would put a patient at unacceptable risk. Tests are judged against them.

All three levels

Beginner. The jobs a medical device must keep doing right so that patients are not hurt.

Novice. In the medical-equipment safety standard IEC 60601-1, the functions whose loss or degradation would put a patient at unacceptable risk. Tests are judged against them.

Expert. Identified through risk management and used to set test pass criteria, for example that electromagnetic interference must not cause unintended or unsafe stimulation.

EUV lithography

Lithography with extreme ultraviolet light of 13.5 nm wavelength, about fourteen times shorter than the 193 nm deep-ultraviolet light used before it. The light comes from tin droplets hit by a laser, and because air and glass absorb it, it travels in a vacuum and is focused with mirrors.

All three levels

Beginner. Printing chip patterns with a special ultraviolet light. Its waves are much shorter than those of ordinary light, so it can print smaller shapes.

Novice. Lithography with extreme ultraviolet light of 13.5 nm wavelength, about fourteen times shorter than the 193 nm deep-ultraviolet light used before it. The light comes from tin droplets hit by a laser, and because air and glass absorb it, it travels in a vacuum and is focused with mirrors.

Expert. Uses reflective masks made of alternating molybdenum and silicon layers and all-mirror optics. Current tools have a numerical aperture of 0.33; high-NA tools reach 0.55 by using anamorphic optics (different magnification in the two directions), which halves the exposure field to 26 × 16.5 mm.

Event-driven simulation

A way of simulating a circuit that keeps a to-do list of signal changes in time order, and re-runs only the parts of the design affected by each change.

All three levels

Beginner. A way of simulating a chip that only does work when some signal actually changes.

Novice. A way of simulating a circuit that keeps a to-do list of signal changes in time order, and re-runs only the parts of the design affected by each change.

Expert. The method SystemVerilog’s meaning is defined by. Each instant of simulated time is processed in ordered phases called regions. Within a region the standard does not fix the order in which processes run, and code whose result depends on that order has a race.

Executable specification

A program, in a spreadsheet, C/C++, Python or SystemC, that captures what the chip must compute or how fast it must do it. Unlike prose, it gives one unambiguous answer for any input.

All three levels

Beginner. A computer program that behaves the way the chip is supposed to, so people can try out the idea before the chip exists.

Novice. A program, in a spreadsheet, C/C++, Python or SystemC, that captures what the chip must compute or how fast it must do it. Unlike prose, it gives one unambiguous answer for any input.

Expert. Comes in layers: power-performance-area spreadsheets, a reference model of the algorithms, a cycle-approximate performance model, and a transaction-level virtual platform that runs real software. The reference model later becomes the golden model that verification compares the design against.

Expert parallelism (EP)

Placing the experts of a mixture-of-experts layer on different accelerators. Before each expert layer, every accelerator sends its tokens to wherever their chosen experts live, and the results are sent back afterwards.

All three levels

Beginner. Putting different specialists of a model on different chips and sending each word to the chip that holds its specialist.

Novice. Placing the experts of a mixture-of-experts layer on different accelerators. Before each expert layer, every accelerator sends its tokens to wherever their chosen experts live, and the results are sent back afterwards.

Expert. Experts sharded over an EP group, usually carved out of the data-parallel dimension. Each MoE layer needs a dispatch and a combine all-to-all forward and again backward, moving about k bsh (EP−1)/EPk\,bsh\,(\mathrm{EP}-1)/\mathrm{EP} elements per rank each time.

Exponent (floating point)

The bits of a floating-point number that pick the power of two it is multiplied by. It is stored with a bias (for example 127 in FP32) so that both large and small powers fit in an unsigned field. More exponent bits mean a wider range.

All three levels

Beginner. The part of a stored number that says how big it is, like the “× 1,000” in 3 × 1,000.

Novice. The bits of a floating-point number that pick the power of two it is multiplied by. It is stored with a bias (for example 127 in FP32) so that both large and small powers fit in an unsigned field. More exponent bits mean a wider range.

Expert. A biased unsigned field of ee bits. The all-zeros code marks zero and subnormals; IEEE formats reserve all-ones for infinity and NaN, while OCP E4M3 and the MX FP6/FP4 types reclaim most or all of those codes for finite values. Range in binades ≈2e\approx 2^e minus the reserved codes, plus the subnormal binades.

External laser source (ELS)

A pluggable module that holds only lasers. It sends steady, unmodulated light over fiber to the co-packaged optical engines, which imprint the data onto it. Lasers are the part most likely to fail and they dislike heat, so keeping them replaceable and cool helps.

All three levels

Beginner. Keeping the lasers in their own plug-in box, away from the hot switch chip, and piping their light in through fiber.

Novice. A pluggable module that holds only lasers. It sends steady, unmodulated light over fiber to the co-packaged optical engines, which imprint the data onto it. Lasers are the part most likely to fail and they dislike heat, so keeping them replaceable and cool helps.

Expert. Continuous-wave lasers in a front-panel form factor (for example the OIF ELSFP). Costs optical loss in the fiber and connectors, made up with higher laser power, but separates the laser from ASIC heat and keeps it field-replaceable.

Failure domain

The set of things that go down together when one component fails. A failed pluggable module takes down one port; a failed co-packaged engine can take down many ports, or a whole switch until it is repaired.

All three levels

Beginner. How much stops working when one part breaks.

Novice. The set of things that go down together when one component fails. A failed pluggable module takes down one port; a failed co-packaged engine can take down many ports, or a whole switch until it is repaired.

Expert. Scope of impact × time to repair. Pluggables have a small domain and minutes-scale repair; CPO enlarges the domain and lengthens repair to a board or chassis swap, so it needs much lower FIT rates, redundancy or graceful lane degradation.

False path

An SDC command (set_false_path) telling the tools to ignore a path for timing, for example a signal crossing between two unrelated clocks through a synchronizer circuit built for that purpose.

All three levels

Beginner. A route the tool is told to ignore when checking speed, because no signal ever really needs to race along it.

Novice. An SDC command (set_false_path) telling the tools to ignore a path for timing, for example a signal crossing between two unrelated clocks through a synchronizer circuit built for that purpose.

Expert. An exception that removes a path from both optimization and checking. Use it only where a synchronizer or protocol makes timing irrelevant, scope it narrowly and review it: a wrong false path hides a real timing bug until the chip comes back.

False sharing

Two threads write different variables that share a cache line. Coherence works per line, so each write invalidates the other core’s copy and the line ping-pongs, even though no data is actually shared.

All three levels

Beginner. When two cores use different numbers that happen to sit in the same block. The block bounces between them as if they were sharing.

Novice. Two threads write different variables that share a cache line. Coherence works per line, so each write invalidates the other core’s copy and the line ping-pongs, even though no data is actually shared.

Expert. A coherence-miss pathology: each write needs ownership of the whole line. Fixed by padding or aligning per-thread data to the line size (64 bytes, or more where lines or adjacent-line prefetch pair lines), or by per-thread copies merged later.

FAN

A 1983 refinement of PODEM by Fujiwara and Shimono. It fills in every value that is forced, stops tracing backward at parts of the circuit that can always be satisfied later, and pursues several goals at once.

All three levels

Beginner. An improved search for test patterns that takes shortcuts where wires split.

Novice. A 1983 refinement of PODEM by Fujiwara and Shimono. It fills in every value that is forced, stops tracing backward at parts of the circuit that can always be satisfied later, and pursues several goals at once.

Expert. Adds immediate implication of uniquely forced values, stops backtrace at headlines (the outputs of fanout-free cones) and propagates through all branches of a fanout, cutting backtracks compared with PODEM.

Fan-in

The number of inputs to a gate. A NAND2 has fan-in 2. In CMOS, higher fan-in means taller transistor stacks, so gates are usually kept to four inputs or fewer.

All three levels

Beginner. How many inputs a gate has.

Novice. The number of inputs to a gate. A NAND2 has fan-in 2. In CMOS, higher fan-in means taller transistor stacks, so gates are usually kept to four inputs or fewer.

Expert. Input count of a gate. Each extra input adds a series device to one network, raising logical effort (NAND: (n+2)/3(n+2)/3) and parasitic delay (≈n\approx n by logical effort; growing faster, roughly as n2n^2, by Elmore). Libraries stop around four inputs; wider functions are built as trees.

Fat tree

A tree-shaped network, proposed by Charles Leiserson in 1985, in which the links get more numerous (fatter) closer to the root. In datacenters it is built from many identical switches in layers so every layer has as much total bandwidth as the one below.

All three levels

Beginner. A network shaped like a tree whose branches get thicker toward the top, so the top doesn’t become a traffic jam.

Novice. A tree-shaped network, proposed by Charles Leiserson in 1985, in which the links get more numerous (fatter) closer to the root. In datacenters it is built from many identical switches in layers so every layer has as much total bandwidth as the one below.

Expert. Leiserson’s universal network; in practice the kk-ary fat-tree of Al-Fares et al.: kk pods of k/2k/2 edge and k/2k/2 aggregation switches, (k/2)2(k/2)^2 cores, k3/4k^3/4 hosts, all from identical kk-port switches. A folded Clos by another name.

Fault collapsing

Merging faults that are equivalent, meaning every test that catches one also catches the other, so the software only has to handle one fault from each group.

All three levels

Beginner. Removing duplicate pretend flaws that any test would catch together, so the list is shorter.

Novice. Merging faults that are equivalent, meaning every test that catches one also catches the other, so the software only has to handle one fault from each group.

Expert. Equivalence collapsing keeps one fault per class; dominance collapsing goes further for test generation. Coverage on collapsed and uncollapsed lists differs, so say which one a number refers to.

Fault coverage

The number of modeled faults the test patterns detect, divided by the total number of modeled faults. A related number, fault efficiency, leaves out faults that are proven impossible to test, so they don’t count against the patterns.

All three levels

Beginner. The share of pretend flaws that the tests would catch. 99% means they catch 99 out of every 100.

Novice. The number of modeled faults the test patterns detect, divided by the total number of modeled faults. A related number, fault efficiency, leaves out faults that are proven impossible to test, so they don’t count against the patterns.

Expert. Always quote it with the fault model, the denominator, and whether the fault list is collapsed. Tools differ on names (fault coverage, test coverage, fault efficiency) and on how they credit possibly-detected faults, so read the report’s definitions.

Fault injection

Deliberately inserting faults, such as a wire stuck at 0 or a stored bit flipped, into a simulation of the chip, and checking whether each one reaches an output and whether a safety mechanism detects it.

All three levels

Beginner. Breaking a design on purpose, in a computer copy, to check that its safety checks work.

Novice. Deliberately inserting faults, such as a wire stuck at 0 or a stored bit flipped, into a simulation of the chip, and checking whether each one reaches an output and whether a safety mechanism detects it.

Expert. Classifies faults as observed or not at functional outputs and detected or not at error signals; observed-but-undetected faults are dangerous and reduce SPFM. Campaigns are expensive and are accelerated with fault-list pruning and formal analysis.

Fault model

A simplified, logical description of what a physical defect does, such as “this wire is always 0.” It turns countless possible defects into a finite list that software can count, target and check. Common models are stuck-at, transition, path delay and bridging.

All three levels

Beginner. A simple pretend version of a flaw, such as “this wire is stuck at 0,” that software can work with.

Novice. A simplified, logical description of what a physical defect does, such as “this wire is always 0.” It turns countless possible defects into a finite list that software can count, target and check. Common models are stuck-at, transition, path delay and bridging.

Expert. Chosen because it is tractable and correlates with real defects. Each model has its own fault list, coverage number and pattern set, and programs run several models to push shipped defect levels down.

Fault simulation

Simulating the circuit once without faults and again with each modeled fault inserted, and comparing the outputs, to find out which patterns detect which faults and what the coverage is.

All three levels

Beginner. Running a computer model of the chip with each pretend flaw added, to see which tests would catch it.

Novice. Simulating the circuit once without faults and again with each modeled fault inserted, and comparing the outputs, to find out which patterns detect which faults and what the coverage is.

Expert. Serial, parallel (bit-parallel over faults or patterns), deductive and concurrent algorithms exist. PPSFP (parallel-pattern single-fault propagation) dominates for full-scan logic; concurrent simulation handles general timing and sequential models. Fault dropping speeds it up but is turned off for diagnosis.

FDA device class

The U.S. Food and Drug Administration’s risk rating for a medical device: Class I (lowest risk), Class II, and Class III (highest risk). Class III devices, which include most implants, generally need premarket approval (PMA) before they can be sold.

All three levels

Beginner. How risky the U.S. government rates a medical device. The riskier it is, the more proof it needs before it can be sold.

Novice. The U.S. Food and Drug Administration’s risk rating for a medical device: Class I (lowest risk), Class II, and Class III (highest risk). Class III devices, which include most implants, generally need premarket approval (PMA) before they can be sold.

Expert. Class sets the controls and pathway: general controls, special controls, and 510(k) clearance for most non-exempt Class I and II devices; PMA for Class III. Changes after approval then follow that pathway’s change rules.

FEOL and BEOL

Front end of line (FEOL) is the first part of wafer processing, which builds the transistors in the silicon surface. Back end of line (BEOL) is the second part, which builds the stack of metal wiring layers above them, separated by insulating layers and joined by vertical plugs called vias.

All three levels

Beginner. The two halves of building a chip: first the tiny switches in the silicon, then the layers of wiring on top.

Novice. Front end of line (FEOL) is the first part of wafer processing, which builds the transistors in the silicon surface. Back end of line (BEOL) is the second part, which builds the stack of metal wiring layers above them, separated by insulating layers and joined by vertical plugs called vias.

Expert. FEOL masks are the base layers and the most expensive to change. BEOL masks are the metal and via layers. A late fix that changes only wiring (a metal-only ECO) replaces just those, which is why designers scatter unused spare cells across the chip at layout time: they sit in the FEOL waiting to be wired in.

Fermi level (EF)

The energy at which an available electron state has a 50% chance of being occupied. In pure silicon it sits near the middle of the band gap; n-type doping raises it toward EcE_{\mathrm{c}}, p-type lowers it toward EvE_{\mathrm{v}}. At equilibrium it is flat across a whole device.

All three levels

Beginner. A “fill line” for electrons: energy levels below it are mostly full and levels above it are mostly empty.

Novice. The energy at which an available electron state has a 50% chance of being occupied. In pure silicon it sits near the middle of the band gap; n-type doping raises it toward EcE_{\mathrm{c}}, p-type lowers it toward EvE_{\mathrm{v}}. At equilibrium it is flat across a whole device.

Expert. Carrier densities follow from Ec−EFE_{\mathrm{c}} - E_{\mathrm{F}} and EF−EvE_{\mathrm{F}} - E_{\mathrm{v}}. Under bias the single EFE_{\mathrm{F}} splits into quasi-Fermi levels for electrons and holes; their separation is the applied voltage. A voltage between two terminals is a difference in their Fermi levels divided by qq.

FIB circuit edit

Changing the wiring of an individual finished chip with a focused ion beam (FIB): a finely aimed beam of charged atoms that can drill through layers to cut a wire, or lay down metal to add one.

All three levels

Beginner. Cutting and adding tiny wires on a finished chip with a very fine beam, to try out a fix before new stencils are made.

Novice. Changing the wiring of an individual finished chip with a focused ion beam (FIB): a finely aimed beam of charged atoms that can drill through layers to cut a wire, or lay down metal to add one.

Expert. Proves a wiring fix on a handful of chips in days, before the team pays for new metal masks. Edits are slow and only a few fit on each chip, so a FIB edit validates a fix and never replaces one.

Field upgrade (reconfigurability)

Updating the hardware function of a deployed product by loading a new FPGA configuration, often over a network. It can fix bugs, add features or follow a changed standard without replacing the board.

All three levels

Beginner. Changing what a chip’s hardware does after it is already in people’s hands, by sending it a new setup file.

Novice. Updating the hardware function of a deployed product by loading a new FPGA configuration, often over a network. It can fix bugs, add features or follow a changed standard without replacing the board.

Expert. Requires secure, authenticated bitstream delivery, a fallback image if an update fails, and, in regulated products, the same approval as any design change. Partial reconfiguration narrows the update to one region while the rest keeps running.

Fill and drain

The cycles at the start of a pass while data is still reaching the far PEs (fill) and at the end while the last results are still leaving (drain). During them only part of the array is busy.

All three levels

Beginner. The start and end of a job on a grid, when the numbers are still marching in or the last answers are marching out, so some cells sit idle.

Novice. The cycles at the start of a pass while data is still reaching the far PEs (fill) and at the end while the last results are still leaving (drain). During them only part of the array is busy.

Expert. For an R×CR \times C array a pass of TT streamed steps takes about 2R+C+T−22R + C + T - 2 cycles, so the overhead is fixed per pass and shrinks relative to work as TT grows. Overlapping passes (double buffering) hides part of it.

Final test (package test)

Production test of each chip after it is sealed in its package, run in a socket on an automated handler, usually with the full set of patterns plus electrical measurements.

All three levels

Beginner. Testing each chip again after it is sealed in its protective case, before it ships.

Novice. Production test of each chip after it is sealed in its package, run in a socket on an automated handler, usually with the full set of patterns plus electrical measurements.

Expert. Catches assembly damage and defects wafer sort missed, and runs tests that probe cards can’t deliver. Results feed speed binning and the shipped defect level.

Fine-tuning

Training an already trained model further on a smaller, targeted set of examples, such as instruction-and-answer pairs or hardware-coding problems.

All three levels

Beginner. Giving an already trained AI model extra practice on a narrower task.

Novice. Training an already trained model further on a smaller, targeted set of examples, such as instruction-and-answer pairs or hardware-coding problems.

Expert. Supervised fine-tuning (SFT) on instruction data turns a base model into an assistant. Too many passes over a narrow set overfit: the model grows more confident and less varied, which can raise pass@1 while lowering pass@10.

FinFET

A transistor built on a thin vertical fin of silicon. The gate covers the two sides and the top of the fin, which gives it much better control than a flat (planar) transistor. Its width comes in whole fins, so designers use one, two or three fins per transistor.

All three levels

Beginner. A transistor whose channel stands up like a thin fin, with the gate draped over it on three sides, so it grips the channel better than a flat one.

Novice. A transistor built on a thin vertical fin of silicon. The gate covers the two sides and the top of the fin, which gives it much better control than a flat (planar) transistor. Its width comes in whole fins, so designers use one, two or three fins per transistor.

Expert. A multigate MOSFET (double- or tri-gate) on a fin a few nanometers thick, invented at UC Berkeley in 1998 and in volume production from 2011. Channel width is quantized to 2Hfin+Wfin2H_{\mathrm{fin}} + W_{\mathrm{fin}} per fin; the sub-fin under the gated region needs a punch-through stopper. ASAP7 models 32 nm × 6.5 nm fins on a 27 nm pitch.

Finite state machine (FSM)

Control logic that is always in one of a fixed list of states and moves between them according to its inputs. It is built from a register holding the current state plus logic that computes the next state and the outputs.

All three levels

Beginner. A circuit that moves step by step through a fixed set of modes, like a washing machine going from fill to wash to spin.

Novice. Control logic that is always in one of a fixed list of states and moves between them according to its inputs. It is built from a register holding the current state plus logic that computes the next state and the outputs.

Expert. Coded as a state register (always_ff) plus next-state and output logic (always_comb). The bit pattern chosen for each state (binary, one-hot, Gray) trades flip-flop count against logic depth and switching, and synthesis tools may re-encode the machine.

First silicon

The first chips made from a new design, used to switch the design on for the first time, find problems and measure its real limits before volume production.

All three levels

Beginner. The first chips that come back from the factory, before anyone knows if they work.

Novice. The first chips made from a new design, used to switch the design on for the first time, find problems and measure its real limits before volume production.

Expert. Usually a small engineering lot. Its arrival starts the clock on the decision to ship, fix in metal or respin, so the test boards, lab plan and debug tools have to be ready before it lands.

FIT (failures in time)

Failures per billion (10910^{9}) device-hours. A part rated at 10 FIT is expected to fail once in 100 million hours of operation, so a fleet of 100,000 such parts would see about one failure every 1,000 hours, roughly six weeks.

All three levels

Beginner. A way to count rare failures: one FIT means one failure in a billion hours of chips running.

Novice. Failures per billion (10910^{9}) device-hours. A part rated at 10 FIT is expected to fail once in 100 million hours of operation, so a fleet of 100,000 such parts would see about one failure every 1,000 hours, roughly six weeks.

Expert. Assumes a constant failure rate, the flat floor of the bathtub curve. A system’s FIT is the sum of its components’ FITs. Qualification data with zero failures gives only an upper bound, at a chosen statistical confidence.

Flash memory

Non-volatile memory where each cell is a transistor with an extra, insulated charge-storage layer. Trapped electrons shift the transistor’s threshold voltage, and the threshold encodes the data. NAND flash chains cells in series for density.

All three levels

Beginner. Memory that keeps its data with the power off, used in phones, memory cards and SSDs. It stores bits by trapping electrons.

Novice. Non-volatile memory where each cell is a transistor with an extra, insulated charge-storage layer. Trapped electrons shift the transistor’s threshold voltage, and the threshold encodes the data. NAND flash chains cells in series for density.

Expert. Floating-gate or charge-trap transistors programmed and erased by Fowler–Nordheim tunneling through the tunnel oxide. Programmed page by page, erased block by block, with endurance of thousands of program/erase cycles. Multi-level cells store 2–4 bits by dividing the threshold window.

Flat-band voltage (Vfb)

The gate voltage at which the energy bands in the silicon are flat right up to the oxide: no carriers are attracted or repelled. It is set by the difference between the gate material and the silicon, plus any charge trapped in the oxide.

All three levels

Beginner. The gate voltage that leaves the silicon under the gate just as it would be with no gate at all.

Novice. The gate voltage at which the energy bands in the silicon are flat right up to the oxide: no carriers are attracted or repelled. It is set by the difference between the gate material and the silicon, plus any charge trapped in the oxide.

Expert. Vfb=ψg−ψs−Qox/CoxV_{\mathrm{fb}} = \psi_{\mathrm{g}} - \psi_{\mathrm{s}} - Q_{\mathrm{ox}}/C_{\mathrm{ox}}. For an n+\mathrm{n^+} poly gate (ψg=χSi=4.05 V\psi_{\mathrm{g}} = \chi_{\mathrm{Si}} = 4.05\,\mathrm{V}) on p-type silicon, Vfb=−(Ec−EF)/qV_{\mathrm{fb}} = -(E_{\mathrm{c}} - E_{\mathrm{F}})/q, roughly −1 V. It is the first term in the threshold equation.

Flip-chip bumps

A packaging method in which small solder bumps are placed on the chip’s top surface and the chip is mounted upside down, so the bumps connect directly to the package. The connections are short, and the bumps can sit in a grid across the chip rather than only along its edge.

All three levels

Beginner. Tiny solder dots spread across the face of a chip. The chip is flipped over and stuck down onto its base by those dots.

Novice. A packaging method in which small solder bumps are placed on the chip’s top surface and the chip is mounted upside down, so the bumps connect directly to the package. The connections are short, and the bumps can sit in a grid across the chip rather than only along its edge.

Expert. Bumps form a grid at a fixed pitch. Each signal or supply net is assigned to a bump and connected to its I/O cell through a redistribution layer (RDL), an extra metal layer on top of the chip. Bump planning therefore couples to I/O cell placement, the power grid and macro placement.

Flip-flop (register)

A circuit that stores one bit (a 0 or a 1). It copies its input to its output only at the instant the clock rises, and holds that value until the next rise. A group of flip-flops that hold a number together is a register.

All three levels

Beginner. A tiny memory that holds one bit and only changes when the clock ticks.

Novice. A circuit that stores one bit (a 0 or a 1). It copies its input to its output only at the instant the clock rises, and holds that value until the next rise. A group of flip-flops that hold a number together is a register.

Expert. The storage element of a clocked design. Its input must be steady for a short window around the clock edge (the setup time before, the hold time after); a change inside that window can leave it metastable. RTL implies a flop wherever a clocked block assigns a value, and cell libraries offer versions with reset, enable and test (scan) inputs.

Flit

A fixed-size block (256 bytes in PCIe 6.0) that carries packets plus error-checking bytes. Fixed sizes make it possible to add error correction.

All three levels

Beginner. A small chunk of data, always the same size, sent over a link.

Novice. A fixed-size block (256 bytes in PCIe 6.0) that carries packets plus error-checking bytes. Fixed sizes make it possible to add error correction.

Expert. Flow-control unit. PCIe 6.0’s 256-byte flit holds 236 B of transaction-layer packets, 6 B of data-link payload, 8 B of CRC and 6 B of FEC; CXL 1.x/2.0 uses a 68-byte flit. Flit mode drops per-packet framing and per-packet CRC.

Floating gate / charge trap

A flash cell’s storage layer. A floating gate is a conductor completely surrounded by oxide; a charge trap is an insulating layer (silicon nitride) that holds electrons in traps. Either way, stored charge raises the transistor’s threshold voltage.

All three levels

Beginner. An island inside a flash transistor, walled off by insulation, where electrons can be parked for years.

Novice. A flash cell’s storage layer. A floating gate is a conductor completely surrounded by oxide; a charge trap is an insulating layer (silicon nitride) that holds electrons in traps. Either way, stored charge raises the transistor’s threshold voltage.

Expert. Floating gates (planar NAND) store charge on a conductor, so one leak path drains the whole gate; charge-trap layers (most 3D NAND) store it in discrete traps, tolerate oxide defects better and endure more cycles, but lose charge faster soon after programming, including sideways between cells.

Floating point

A number format that stores a sign, an exponent and a mantissa. The value is ±(1.mantissa)×2exponent−bias\pm(1.\text{mantissa}) \times 2^{\text{exponent} - \text{bias}}. Because the exponent moves the binary point, the format covers a huge range with a fixed number of significant bits.

All three levels

Beginner. A way of storing numbers like scientific notation: a few digits plus a note saying where the decimal point goes, so the same few digits can describe tiny or huge numbers.

Novice. A number format that stores a sign, an exponent and a mantissa. The value is ±(1.mantissa)×2exponent−bias\pm(1.\text{mantissa}) \times 2^{\text{exponent} - \text{bias}}. Because the exponent moves the binary point, the format covers a huge range with a fixed number of significant bits.

Expert. Binary floating point per IEEE 754 and its descendants: significand precision pp (mantissa bits + 1 hidden bit), exponent width setting the range in binades, round-to-nearest-even, subnormals for gradual underflow. Relative rounding error is bounded by 2−p2^{-p} regardless of magnitude.

Floorplan

The plan of a chip’s physical layout, made before any of its millions of small parts are placed: how big the chip is, where its connections to the outside world sit, where the large ready-made blocks go, and which areas are reserved or kept clear. Tools save it as a file, usually DEF, that later steps start from.

All three levels

Beginner. The plan for a chip’s space: how big it is and where the big blocks go, made before the small parts are placed.

Novice. The plan of a chip’s physical layout, made before any of its millions of small parts are placed: how big the chip is, where its connections to the outside world sit, where the large ready-made blocks go, and which areas are reserved or kept clear. Tools save it as a file, usually DEF, that later steps start from.

Expert. The set of decisions every later layout step inherits: the outline, the rows that standard cells will sit in, I/O pin or bump positions, each macro’s location and orientation with a keep-out halo, blockages, power-domain regions and, when the chip is split into blocks, each block’s outline, pins and timing budget. Macro positions are rarely changed once fixed, so their quality limits timing, congestion and the power grid downstream.

FLOP and FLOP/s

A floating-point operation: one add or one multiply. FLOP/s (FLOPS) counts them per second; AI chips are rated in teraFLOP/s (101210^{12} per second) or more.

All three levels

Beginner. One math step on a number with a decimal point, like one multiplication. Chips are rated by how many they can do each second.

Novice. A floating-point operation: one add or one multiply. FLOP/s (FLOPS) counts them per second; AI chips are rated in teraFLOP/s (101210^{12} per second) or more.

Expert. A multiply-accumulate counts as 2 FLOPs. Vendor peaks are usually dense matrix rates at a stated precision, sometimes doubled for structured sparsity, so always check which. Integer rates are quoted as OPS or TOPS.

Flow control (credit-based)

The mechanism that decides when a packet may move to the next router. With credits, a sender keeps a count of free buffer slots downstream and only sends when that count is above zero.

All three levels

Beginner. A rule that stops a core from sending data until the next one has room to take it, so nothing is dropped.

Novice. The mechanism that decides when a packet may move to the next router. With credits, a sender keeps a count of free buffer slots downstream and only sends when that count is above zero.

Expert. Credit-based, hop-by-hop backpressure is standard on chip: it never drops packets, but a stalled destination backs traffic up through the network, which is how one hot link can slow unrelated streams.

FMEDA

Failure Modes, Effects and Diagnostic Analysis: a table listing, for each block of the chip, the ways it can fail, how often, what each failure would do to a safety goal, and what share a safety mechanism would catch.

All three levels

Beginner. A big table of every way each part of a chip could fail, and whether something would catch it.

Novice. Failure Modes, Effects and Diagnostic Analysis: a table listing, for each block of the chip, the ways it can fail, how often, what each failure would do to a safety goal, and what share a safety mechanism would catch.

Expert. The quantitative backbone of the hardware safety case: summing its rows produces SPFM, LFM and PMHF. Its diagnostic-coverage figures should be validated by fault injection rather than taken from tables.

FO4 delay

The delay of one inverter (the simplest logic gate) driving four copies of itself. Measuring the logic in each pipeline stage in FO4 delays gives a number that stays roughly the same from one manufacturing process to the next.

All three levels

Beginner. A standard yardstick for how fast a chip’s tiny switches are, so chips made in different factories can be compared fairly.

Novice. The delay of one inverter (the simplest logic gate) driving four copies of itself. Measuring the logic in each pipeline stage in FO4 delays gives a number that stays roughly the same from one manufacturing process to the next.

Expert. A unit for architecture-level budgets: logic per stage plus register and clock overhead, both in FO4, sets the cycle time. The overhead doesn’t shrink when the logic does, so in very short stages it takes a large share of the cycle.

Forksheet

A nanosheet layout in which the nMOS and pMOS sheets sit on either side of a thin insulating wall, so they can be packed closer together than separate nanosheet devices.

All three levels

Beginner. A variation of the stacked-ribbon transistor where a thin wall lets the two kinds of transistor sit closer together.

Novice. A nanosheet layout in which the nMOS and pMOS sheets sit on either side of a thin insulating wall, so they can be packed closer together than separate nanosheet devices.

Expert. A dielectric wall between n and p nanosheet stacks shrinks n-p spacing in the cell, which lets the sheets be wider or the cell shorter. imec proposed it to extend the nanosheet roadmap (to around its A10 node) before CFET takes over.

Formal verification

Checking a design by mathematical proof instead of by running tests. A tool considers every possible input at once and either proves that a rule always holds or produces a counterexample: one specific input sequence that breaks it.

All three levels

Beginner. Using math to prove a design follows a rule for every possible input, without running tests one by one.

Novice. Checking a design by mathematical proof instead of by running tests. A tool considers every possible input at once and either proves that a rule always holds or produces a counterexample: one specific input sequence that breaks it.

Expert. Exhaustive within its assumptions and any depth limit. The number of possible states doubles with every storage bit added (state explosion), so it is applied block by block, with assumptions about inputs and simplifications of the design.

Forward error correction (FEC)

Adding redundant bits to the data stream so the receiver can detect and correct errors without asking for a resend. It lets links run with a raw error rate far worse than what the application could accept.

All three levels

Beginner. Extra check bits sent with the data so the receiver can fix a few wrong bits on its own.

Novice. Adding redundant bits to the data stream so the receiver can detect and correct errors without asking for a resend. It lets links run with a raw error rate far worse than what the application could accept.

Expert. Ethernet at 100G per lane and above uses RS(544,514) “KP4” FEC, which lets links run at pre-FEC bit-error ratios in the 10−410^{-4} range. Decoding adds latency, which is why scale-up links weigh lighter FEC against retries.

Forwarding (bypassing)

Extra wires and selectors that send a result from a later pipeline stage straight back to the input of the execute stage, so a dependent instruction doesn’t have to wait for the result to reach the register file.

All three levels

Beginner. A shortcut that passes an answer straight to the next step that needs it, instead of waiting for it to be written down first.

Novice. Extra wires and selectors that send a result from a later pipeline stage straight back to the input of the execute stage, so a dependent instruction doesn’t have to wait for the result to reach the register file.

Expert. Comparators match each source register in X against the destinations held in the X/M and M/W pipeline registers (newest wins, x0 excluded) and steer a mux in front of each ALU input. It removes ALU-to-ALU stalls but sits on the X-stage critical path.

Foundry

A company that manufactures chips to other companies’ designs. It publishes the rules a layout must obey, receives the final layout file at tapeout, makes the masks and processes the wafers.

All three levels

Beginner. A factory that makes chips designed by other companies.

Novice. A company that manufactures chips to other companies’ designs. It publishes the rules a layout must obey, receives the final layout file at tapeout, makes the masks and processes the wafers.

Expert. Supplies the process design kit (device models, design rules, rule files for the checking tools, layer map) and runs its own intake checks and mask data preparation on every tapeout.

Foundry reference flow

A recommended recipe a chip factory publishes for one of its processes: which design tools and versions to use at each step, plus scripts that pass data from one tool to the next.

All three levels

Beginner. The factory’s list of recommended design programs and settings. Following it lowers the risk that the chip has problems.

Novice. A recommended recipe a chip factory publishes for one of its processes: which design tools and versions to use at each step, plus scripts that pass data from one tool to the next.

Expert. Built on tools whose results the foundry has checked against its own process, especially the final-check (signoff) tools: design-rule and layout-versus-schematic checking, wire extraction, timing, and power-grid analysis. The rule files ship in those tools’ own formats, which is the practical lock-in behind tool qualification.

FP16 (half precision)

IEEE half precision: 1 sign bit, 5 exponent bits, 10 mantissa bits. Its largest value is 65,504 and the smallest non-zero value is about 6×10−86 \times 10^{-8}, so small gradients can vanish without help.

All three levels

Beginner. The standard 16-bit number format. It keeps more digits than BF16 but can’t hold very big or very tiny numbers.

Novice. IEEE half precision: 1 sign bit, 5 exponent bits, 10 mantissa bits. Its largest value is 65,504 and the smallest non-zero value is about 6×10−86 \times 10^{-8}, so small gradients can vanish without help.

Expert. IEEE binary16: bias 15, p=11p = 11, max 65,504, min normal 2−142^{-14}, min subnormal 2−242^{-24}. Three more mantissa bits than BF16 but far less range (about 30 binades of normal range against 254); training with it needs loss scaling and FP32 master weights.

FP4 (E2M1)

A 4-bit float with 1 sign, 2 exponent and 1 mantissa bit. Its non-negative values are 0, 0.5, 1, 1.5, 2, 3, 4 and 6. It is always used with block scale factors, as in MXFP4 or NVFP4.

All three levels

Beginner. A tiny 4-bit number format with only 15 different values. It works for AI only with a shared multiplier for each small group.

Novice. A 4-bit float with 1 sign, 2 exponent and 1 mantissa bit. Its non-negative values are 0, 0.5, 1, 1.5, 2, 3, 4 and 6. It is always used with block scale factors, as in MXFP4 or NVFP4.

Expert. E2M1, bias 1, no Inf or NaN, max 6, one subnormal (0.5). Fifteen distinct values. Used as MXFP4 (32-element blocks, E8M0 scale, 4.25 bits/value) or NVFP4 (16-element blocks, E4M3 scale plus FP32 tensor scale, 4.5 bits/value).

FP8 (E4M3 and E5M2)

8-bit floating point, standardized by the Open Compute Project. E4M3 has 4 exponent and 3 mantissa bits (largest value 448); E5M2 has 5 and 2 (largest 57,344). Each tensor usually gets its own scale factor so its values land in range.

All three levels

Beginner. Two 8-bit number formats used for AI: one keeps a little more detail, the other a little more range.

Novice. 8-bit floating point, standardized by the Open Compute Project. E4M3 has 4 exponent and 3 mantissa bits (largest value 448); E5M2 has 5 and 2 (largest 57,344). Each tensor usually gets its own scale factor so its values land in range.

Expert. OCP OFP8: E4M3 (bias 7, no infinities, single NaN pattern, 18 binades) for weights and activations; E5M2 (bias 15, IEEE-style specials, 32 binades) for gradients. Used with per-tensor or finer scaling, saturating conversion and higher-precision accumulation.

FPGA

Field-programmable gate array: a chip made of many small configurable logic blocks and programmable wiring. A configuration file (the bitstream) sets what it does, so the function can change without making a new chip.

All three levels

Beginner. A chip that can be rewired with new instructions after it is built.

Novice. Field-programmable gate array: a chip made of many small configurable logic blocks and programmable wiring. A configuration file (the bitstream) sets what it does, so the function can change without making a new chip.

Expert. Avoids the up-front cost of masks and allows updates in the field, at a cost in power, area and speed against a custom ASIC. It is still a commercial part with its own supply chain, and the bitstream becomes an asset to protect.

FPGA prototyping

Loading the design onto FPGAs, chips whose circuits can be reconfigured after manufacture, so it runs fast enough to run real software and connect to real devices.

All three levels

Beginner. Loading a chip design onto chips that can be rewired by software, so it runs fast enough to use with real devices.

Novice. Loading the design onto FPGAs, chips whose circuits can be reconfigured after manufacture, so it runs fast enough to run real software and connect to real devices.

Expert. Needs FPGA-specific changes (clocking, memories, splitting the design across several devices), so it is not cycle-identical to the final chip. Compiles are slow and visibility is limited; watching a new signal usually means recompiling.

FPGA–ASIC gap

The ratios of area, delay and power between a circuit built in an FPGA and the same circuit as a custom chip on the same process. A classic 90 nm study measured about 35× the area, 3–4× the delay and 14× the dynamic power for plain logic.

All three levels

Beginner. How much bigger, slower and more power-hungry an FPGA is than a custom chip doing the same job. It is the price of being able to rewire it.

Novice. The ratios of area, delay and power between a circuit built in an FPGA and the same circuit as a custom chip on the same process. A classic 90 nm study measured about 35× the area, 3–4× the delay and 14× the dynamic power for plain logic.

Expert. Measured by implementing identical RTL both ways on matched processes (Kuon & Rose: 35× area, 3.4–4.6× delay, 14× dynamic power, logic only). Hard blocks narrow area and power far more than delay; the gap for a whole product depends on how much of it maps to hard IP.

Full mesh

A topology where each of pp chips has a direct link to each of the other p−1p - 1. No switch is needed, and any two chips are one hop apart.

All three levels

Beginner. Every chip has its own direct wire to every other chip.

Novice. A topology where each of pp chips has a direct link to each of the other p−1p - 1. No switch is needed, and any two chips are one hop apart.

Expert. Degree p−1p - 1, so the chip’s fixed lane budget is split p−1p - 1 ways: per-pair bandwidth falls as 1/(p−1)1/(p-1) and the domain size is capped by the port count (8 chips with seven links is typical).

Functional coverage

A list of situations that the specification cares about, written by engineers, which the simulator ticks off as tests reach them; for example, “the buffer is full when a reset arrives.” The total shows how much of the plan the tests have actually exercised.

All three levels

Beginner. A checklist of situations the tests should reach. Each one gets ticked off when a test actually reaches it.

Novice. A list of situations that the specification cares about, written by engineers, which the simulator ticks off as tests reach them; for example, “the buffer is full when a reset arrives.” The total shows how much of the plan the tests have actually exercised.

Expert. Written as SystemVerilog covergroups (coverpoints that sort values into bins, and crosses that combine coverpoints) or as cover properties. Only as good as the list: a feature missing from it reads 100% while untested. Crosses multiply bin counts, so engineers choose them from the spec’s risky areas.

Functional safety

The engineering discipline of making sure a malfunctioning electronic system doesn’t create an unreasonable risk of harm. For cars it is defined by the standard ISO 26262.

All three levels

Beginner. Making sure that when electronics break, they break in a way that doesn’t hurt anyone.

Novice. The engineering discipline of making sure a malfunctioning electronic system doesn’t create an unreasonable risk of harm. For cars it is defined by the standard ISO 26262.

Expert. Covers random hardware faults (quantified by architectural metrics and PMHF) and systematic faults such as design bugs (controlled by process). Failures of perception and of the intended function, with no fault present, fall under ISO 21448 (SOTIF).

Gate (of a transistor)

The terminal that sits on top of the thin insulating oxide over the channel. It is not electrically connected to the silicon; it acts through its electric field, like one plate of a capacitor.

All three levels

Beginner. The control connection of a transistor. Its voltage, an electric push, decides whether the switch is on or off.

Novice. The terminal that sits on top of the thin insulating oxide over the channel. It is not electrically connected to the silicon; it acts through its electric field, like one plate of a capacitor.

Expert. Polysilicon in older processes such as SKY130, a metal stack over a high-k dielectric in advanced ones. Its capacitance (CoxWLC_{\mathrm{ox}} W L plus overlap and fringe) is the load every driving gate must charge. Not to be confused with a logic gate.

Gate capacitance

The capacitance between the gate and the channel, about Cox×width×lengthC_{\mathrm{ox}} \times \text{width} \times \text{length}. Every time a signal changes, the circuit driving it must charge or discharge the gates it connects to.

All three levels

Beginner. How much electric charge a transistor’s gate must be filled with before the switch flips. More charge takes more time and energy.

Novice. The capacitance between the gate and the channel, about Cox×width×lengthC_{\mathrm{ox}} \times \text{width} \times \text{length}. Every time a signal changes, the circuit driving it must charge or discharge the gates it connects to.

Expert. CoxWLC_{\mathrm{ox}} W L plus overlap and fringe terms, typically 1–2 fF per µm of width. It dominates the input load of a logic gate and, with on-resistance, sets the RC delay unit for the whole technology.

Gate delay (propagation delay)

The time from an input crossing half the supply voltage to the output crossing half the supply voltage. It is mostly the time a transistor needs to charge or discharge the capacitance the gate drives.

All three levels

Beginner. How long a logic gate takes to change its output after its input changes. A few trillionths of a second on a modern chip.

Novice. The time from an input crossing half the supply voltage to the output crossing half the supply voltage. It is mostly the time a transistor needs to charge or discharge the capacitance the gate drives.

Expert. Measured 50% to 50% and characterized per cell, per input slew and per output load in the library’s timing tables. To first order tpd≈CLVDD/(2Ion)t_{\mathrm{pd}} \approx C_{\mathrm{L}} V_{\mathrm{DD}} / (2 I_{\mathrm{on}}); the slowest chain of gates between two flip-flops sets the clock period.

Gate leakage

Current that tunnels through the gate insulator when it is only a few atoms thick. It flows whenever a voltage sits across the gate, whether or not the transistor is switching.

All three levels

Beginner. Electricity that sneaks straight through the super-thin insulating layer inside a transistor, a layer only a few atoms thick.

Novice. Current that tunnels through the gate insulator when it is only a few atoms thick. It flows whenever a voltage sits across the gate, whether or not the transistor is switching.

Expert. Direct tunneling current, exponential in oxide thickness. It became a major leakage term as SiO2\mathrm{SiO_2} gates thinned toward roughly a nanometer, and is the reason high-k gate dielectrics replaced SiO2\mathrm{SiO_2}: a thicker high-k layer gives the same capacitance with far less tunneling.

Gate length (L, L_g)

The distance from source to drain under the gate, along the direction current flows. Shorter gates switch faster and fit more transistors in the same space, but the gate loses control when they get too short.

All three levels

Beginner. How long the switch’s control strip is, measured in the direction current flows. Shorter means a smaller, faster switch.

Novice. The distance from source to drain under the gate, along the direction current flows. Shorter gates switch faster and fit more transistors in the same space, but the gate loses control when they get too short.

Expert. Physical gate length, usually a little longer than the effective channel length set by the source/drain junctions. It stopped tracking node names in the mid-1990s and has scaled more slowly than gate pitch since; FinFET-era values are around 20 nm (ASAP7 assumes 21 nm).

Gate overdrive (VGS − Vt)

The amount by which the gate-to-source voltage exceeds the threshold voltage, written VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}} or VGTV_{\mathrm{GT}}. Above-threshold current depends on this difference, not on VGSV_{\mathrm{GS}} alone.

All three levels

Beginner. How far the gate voltage is past the tipping point. The farther past, the more current flows.

Novice. The amount by which the gate-to-source voltage exceeds the threshold voltage, written VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}} or VGTV_{\mathrm{GT}}. Above-threshold current depends on this difference, not on VGSV_{\mathrm{GS}} alone.

Expert. VGT=VGS−VtV_{\mathrm{GT}} = V_{\mathrm{GS}} - V_{\mathrm{t}} sets the channel charge per unit area (CoxVGTC_{\mathrm{ox}} V_{\mathrm{GT}}) at the source end. Drive current scales as VGT2V_{\mathrm{GT}}^2 in the square law and roughly VGTαV_{\mathrm{GT}}^{\alpha} (1<α<21 < \alpha < 2) in partly velocity-saturated devices.

Gate oxide

The insulating film, originally silicon dioxide, between the gate and the channel. Thinner oxide gives the gate a stronger grip on the channel but lets more current tunnel through.

All three levels

Beginner. A layer of glass, far thinner than a soap bubble, between a transistor’s gate and the silicon. It stops electricity from flowing into the gate.

Novice. The insulating film, originally silicon dioxide, between the gate and the channel. Thinner oxide gives the gate a stronger grip on the channel but lets more current tunnel through.

Expert. About 4.1 nm (electrical) for SKY130’s 1.8 V NMOS. At around 1 nm, SiO2\mathrm{SiO_2} leaks too much by tunneling, so advanced processes use high-k dielectrics that give the same capacitance with a physically thicker film.

Gate-all-around (GAA) nanosheet

The successor to the FinFET. The channel is a stack of thin horizontal silicon sheets, and the gate wraps all the way around each one. That gives even better control, and the sheets can be made wider or narrower to tune the current. Companies call it MBCFET or RibbonFET.

All three levels

Beginner. A transistor whose channel is a few thin ribbons stacked on top of each other, each completely surrounded by the gate.

Novice. The successor to the FinFET. The channel is a stack of thin horizontal silicon sheets, and the gate wraps all the way around each one. That gives even better control, and the sheets can be made wider or narrower to tune the current. Companies call it MBCFET or RibbonFET.

Expert. Stacked nanosheets (typically three) released from a Si/SiGe superlattice by selectively etching the SiGe, then wrapped with a high-k/metal gate by atomic layer deposition. Sheet thickness, not fin width, sets λ\lambda, and the sub-channel leakage path of the fin disappears. Width is continuous within a cell but limited by n-p spacing and sheet stack height.

Gate-level simulation (GLS)

Rerunning some tests on the netlist, the version of the design after it has been turned into actual logic gates, sometimes with each gate’s real delay included. It catches problems that the higher-level description hid.

All three levels

Beginner. Rerunning some tests on a later, more detailed version of the design, to catch problems the earlier version hid.

Novice. Rerunning some tests on the netlist, the version of the design after it has been turned into actual logic gates, sometimes with each gate’s real delay included. It catches problems that the higher-level description hid.

Expert. Runs either with zero delay or with delays from an SDF file (Standard Delay Format, written by the timing tools). Catches unknown-value bugs masked in RTL simulation, mismatches between how synthesis and simulation read the code, reset problems and test-mode issues. Slow, so teams run a short, targeted list.

GBA vs. PBA

Two modes of timing analysis. Graph-based analysis is fast and checks everything, but it is cautious: at each point it assumes the worst case from any path passing through. Path-based analysis re-checks one path at a time with that path’s own conditions, which is more accurate but much slower.

All three levels

Beginner. Two ways to check timing. One is fast and careful for every path. The other is slow and exact, for the paths that matter.

Novice. Two modes of timing analysis. Graph-based analysis is fast and checks everything, but it is cautious: at each point it assumes the worst case from any path passing through. Path-based analysis re-checks one path at a time with that path’s own conditions, which is more accurate but much slower.

Expert. Graph-based analysis stores one worst arrival and slew per pin; at a gate with several inputs, every path through the output inherits the worst input slew. Path-based analysis re-times a specific path with the slews that actually occur along it, so its slack is never worse. It costs several times the runtime, so flows run GBA everywhere and PBA on the paths that fail.

GCell (global routing cell)

One tile of the coarse grid the global router lays over the chip, usually a dozen or more tracks on a side (OpenROAD’s default is 15 track spacings of layer M3). The router only tracks how many wires cross each border between tiles, on each layer, and compares that with how many tracks cross it.

All three levels

Beginner. One square of the coarse grid that the planner lays over the chip. The planner counts how many wires must pass through each square.

Novice. One tile of the coarse grid the global router lays over the chip, usually a dozen or more tracks on a side (OpenROAD’s default is 15 track spacings of layer M3). The router only tracks how many wires cross each border between tiles, on each layer, and compares that with how many tracks cross it.

Expert. Bigger GCells make global routing faster but hide local trouble such as cells with many pins close together and the edges of blockages. A border’s capacity is its track count minus blockages, pre-routed wires and any derate you set. Demand counts wires and, in some routers, vias and pins.

GDSII

The standard file format for a chip’s layout, from Calma (1978). It is binary (written for programs, not people) and hierarchical: named cells hold shapes and references to other cells, and every shape carries a layer number and a datatype.

All three levels

Beginner. The standard file format that holds the drawing of every layer of a chip.

Novice. The standard file format for a chip’s layout, from Calma (1978). It is binary (written for programs, not people) and hierarchical: named cells hold shapes and references to other cells, and every shape carries a layer number and a datatype.

Expert. A stream of records, each with a 2-byte length, a record type and a data type. Coordinates are 4-byte integers in database units set by the UNITS record; reals use an excess-64 base-16 format, not IEEE 754. The original spec limits layers and datatypes to 0–255 and boundaries to 200 points.

Global placement

The first placement pass. It finds a rough position for every cell that keeps connected cells close, while keeping every small area of the chip below a target fullness. Cells may still overlap a little and are not yet in row slots.

All three levels

Beginner. The first, rough pass that decides about where every part should go, even if a few still overlap.

Novice. The first placement pass. It finds a rough position for every cell that keeps connected cells close, while keeping every small area of the chip below a target fullness. Cells may still overlap a little and are not yet in row slots.

Expert. A continuous optimization over all cell (x,y)(x, y) coordinates, with the per-bin density limit turned into a penalty. Modern engines are analytical: quadratic (SimPL, ComPLx) or nonlinear electrostatic (ePlace, RePlAce, DREAMPlace). It stops when density overflow falls below a target such as 10%.

Global routing

The first, planning pass of routing. It gives every net a rough path through a coarse grid of tiles (GCells), usually with a layer for each piece, while staying within how many wires each tile border can hold and keeping total wire length and via count low. It hands this plan (route guides) and a map of crowded areas to detailed routing.

All three levels

Beginner. The planning pass of routing: choosing a rough path for every connection across a coarse grid, the way you plan a road trip city by city before opening street maps.

Novice. The first, planning pass of routing. It gives every net a rough path through a coarse grid of tiles (GCells), usually with a layer for each piece, while staying within how many wires each tile border can hold and keeping total wire length and via count low. It hands this plan (route guides) and a map of crowded areas to detailed routing.

Expert. A typical modern global router splits multi-pin nets into two-pin pieces with Steiner trees, tries cheap L- and Z-shaped routes, maze-routes what still overflows, rips up and reroutes under rising congestion costs, then assigns layers. Overflow left at the end can turn into detours or rule violations later.

Golden (signoff-quality) tool

A checking tool whose answer is accepted as final, for example on timing or on the factory’s drawing rules. Earlier tools make estimates while they build the layout; a golden tool gives the verdict before the design goes to the factory.

All three levels

Beginner. A checking program whose answer the factory and the design team trust enough to bet a chip on.

Novice. A checking tool whose answer is accepted as final, for example on timing or on the factory’s drawing rules. Earlier tools make estimates while they build the layout; a golden tool gives the verdict before the design goes to the factory.

Expert. Golden status rests on agreement with measured silicon and a long track record, which can only be earned after chips are made. A flow without golden tools has to guardband: add safety margin to its estimates so the design stays safe, at some cost in speed, power or area (PPA).

GPU (graphics processing unit)

A processor made of many simple cores that run the same program on huge amounts of data in parallel. Originally fixed-function graphics hardware, it became programmable in the 2000s and now runs most AI training.

All three levels

Beginner. A chip with thousands of small math units that all work at once. It was built to draw video-game graphics and turned out to be great at AI math too.

Novice. A processor made of many simple cores that run the same program on huge amounts of data in parallel. Originally fixed-function graphics hardware, it became programmable in the 2000s and now runs most AI training.

Expert. A throughput processor: dozens to a few hundred multithreaded cores (SMs or CUs), each keeping thousands of threads resident to hide memory latency, plus matrix units for dense linear algebra and a high-bandwidth memory system.

Gradient

For each weight, the rate at which the training error would change if that weight changed. The backward pass computes one gradient per weight, so gradients take as much memory as the weights themselves.

All three levels

Beginner. A list of nudges, one for every number in the model, saying which way to change it and by how much.

Novice. For each weight, the rate at which the training error would change if that weight changed. The backward pass computes one gradient per weight, so gradients take as much memory as the weights themselves.

Expert. ∂loss/∂w\partial \mathrm{loss} / \partial w for every parameter, produced layer by layer in reverse during the backward pass. In data parallelism, replicas average their gradients with an all-reduce (or reduce-scatter) before the optimizer step.

Graph neural network (GNN)

A neural network that works on a network of connected items. It keeps a short list of numbers for each item and updates it from its neighbors’ lists, round after round. A circuit fits naturally: gates are the items and wires are the connections.

All three levels

Beginner. A kind of learning program that works on networks of connected things, like the parts and wires of a circuit.

Novice. A neural network that works on a network of connected items. It keeps a short list of numbers for each item and updates it from its neighbors’ lists, round after round. A circuit fits naturally: gates are the items and wires are the connections.

Expert. Each layer passes messages along the edges, so after L layers a node has information from L hops away. Timing models often update nodes in signal order. Errors can build up along long paths, and graphs with millions of gates have to be split up to fit in GPU memory.

Gray code

A way of counting in binary where only one bit changes from each number to the next (000, 001, 011, 010, 110…). A count caught in the middle of changing is then either the old number or the new one, never garbage.

All three levels

Beginner. A way of counting where only one digit changes at each step.

Novice. A way of counting in binary where only one bit changes from each number to the next (000, 001, 011, 010, 110…). A count caught in the middle of changing is then either the old number or the new one, never garbage.

Expert. Guarantees that a value sampled mid-transition is the old or the new value, which makes Gray counters safe to pass bit by bit through synchronizers. Logic that converts binary to Gray can briefly produce out-of-sequence codes, so the Gray value must come straight from a register.

Grouped-query attention (GQA)

An attention design in which groups of query heads share one key/value head, so the KV cache shrinks by the group size. Multi-query attention (MQA) is the extreme case with a single shared key/value head.

All three levels

Beginner. A way to build an AI model so the notes it keeps about the chat are much smaller.

Novice. An attention design in which groups of query heads share one key/value head, so the KV cache shrinks by the group size. Multi-query attention (MQA) is the extreme case with a single shared key/value head.

Expert. With hh query heads and gg KV heads the KV cache shrinks h/gh/g times, and the arithmetic intensity of attention in decode rises to about 2h/g2h/g ÷ bytes per element FLOPs per byte. Llama 3 uses 8 KV heads at every size.

Gustafson’s law

Scaled speedup=s+pN\text{Scaled speedup} = s + pN, where ss and pp are the fractions of time spent on serial and parallel work on the NN-processor machine. It assumes the job grows with the machine: more processors are used to solve a bigger problem in the same time.

All three levels

Beginner. The observation that when you get more helpers, you usually take on a bigger job, so the helpers stay useful.

Novice. Scaled speedup=s+pN\text{Scaled speedup} = s + pN, where ss and pp are the fractions of time spent on serial and parallel work on the NN-processor machine. It assumes the job grows with the machine: more processors are used to solve a bigger problem in the same time.

Expert. Fits throughput work whose size grows with the hardware (bigger batches, higher resolution). Amdahl’s law fits a fixed job with a latency target, such as one frame or one query. Pick the one that matches what the spec fixes.

H-tree

A clock layout that branches in nested H shapes: each wire splits into two equal-length wires, again and again, so every endpoint is the same wire distance from the center.

All three levels

Beginner. A symmetric branching pattern shaped like nested letter H’s, so every endpoint is the same distance from the center.

Novice. A clock layout that branches in nested H shapes: each wire splits into two equal-length wires, again and again, so every endpoint is the same wire distance from the center.

Expert. Zero skew by symmetry when the loads are uniform, so it is used for the top levels of distribution. Blockages and uneven sink placement break the symmetry, and it uses more wire, and so more power, than a tree shaped around the actual sinks.

Hallucination

Fluent, confident model output that is wrong or made up: a tool option that doesn’t exist, a wrong signal name, an invented detail of a specification.

All three levels

Beginner. When an AI model states something false with full confidence, such as a tool option that does not exist.

Novice. Fluent, confident model output that is wrong or made up: a tool option that doesn’t exist, a wrong signal name, an invented detail of a specification.

Expert. Common defenses are pasting trusted documents into the prompt (retrieval-augmented generation, RAG), running the output through tools that fail loudly on invalid commands, and checking against a reference. None removes it, so outputs are treated as unchecked drafts.

Halo (keep-out margin)

A border around a large block (a macro, such as a memory) where no small cells may be placed. It is attached to the block and moves with it.

All three levels

Beginner. A clear border kept around a big block so small parts don’t crowd right up against it.

Novice. A border around a large block (a macro, such as a memory) where no small cells may be placed. It is attached to the block and moves with it.

Expert. DEF COMPONENTS + HALO [SOFT] left bottom right top: a placement blockage around a macro’s edges. A soft halo is honored only during initial placement, so later buffering and clock tree synthesis may use it.

Hard IP (FPGA)

Fixed circuits built into an FPGA for common, demanding jobs: high-speed transceivers, PCIe interfaces, memory controllers and processor cores. They are smaller, faster and lower power than the same thing in programmable logic.

All three levels

Beginner. Parts of an FPGA that come already built, like fast links or a small processor, instead of being made from the general blocks.

Novice. Fixed circuits built into an FPGA for common, demanding jobs: high-speed transceivers, PCIe interfaces, memory controllers and processor cores. They are smaller, faster and lower power than the same thing in programmable logic.

Expert. Fixed-function silicon placed in the fabric (SerDes, PCIe, DDR controllers, Arm or RISC-V subsystems). It narrows the gap to an ASIC for the functions it covers and is dead area for designs that don’t use it.

Hardware description language (HDL)

A language for describing circuits rather than programs, such as SystemVerilog or VHDL. The same code can be run on a computer to test it (simulation) or turned into gates (synthesis).

All three levels

Beginner. A programming-like language used to describe circuits instead of software.

Novice. A language for describing circuits rather than programs, such as SystemVerilog or VHDL. The same code can be run on a computer to test it (simulation) or turned into gates (synthesis).

Expert. SystemVerilog (IEEE 1800) and VHDL (IEEE 1076) dominate chip design. Both began as languages for simulating circuits, so the official meaning of the code is what a simulator does with it. Synthesis tools accept only a subset whose hardware meaning they agree on.

Hardware thread (logical processor)

A set of program state (registers and program counter) that a core holds so it can run one instruction stream. A core with two hardware threads appears to the operating system as two processors. Intel calls them logical processors; RISC-V calls them harts.

All three levels

Beginner. One job that a core keeps track of. A core with two hardware threads can juggle two jobs at once, so the computer sees it as two.

Novice. A set of program state (registers and program counter) that a core holds so it can run one instruction stream. A core with two hardware threads appears to the operating system as two processors. Intel calls them logical processors; RISC-V calls them harts.

Expert. The architectural state the hardware duplicates per context: registers, PC, some control state and interrupt controller. Everything else (pipeline, caches, predictors, execution units) is shared or partitioned among the core’s threads.

Hardware Trojan

A deliberately hidden change to a chip’s circuit. It usually has two parts: a trigger that waits for a rare condition, and a payload that then leaks data, changes what the chip does or disables it.

All three levels

Beginner. A hidden, harmful change to a chip. It stays quiet until something wakes it up.

Novice. A deliberately hidden change to a chip’s circuit. It usually has two parts: a trigger that waits for a rare condition, and a payload that then leaks data, changes what the chip does or disables it.

Expert. Can enter through the RTL, third-party IP, design tools, the factory or packaging. A small trigger rarely fires during ordinary testing, so detection combines tests aimed at rare conditions, side-channel measurements (power, timing, heat) and physical reverse engineering against a known-good reference.

Hardware/software co-design

Deciding what goes in hardware and what stays in software at the same time, and adapting the algorithm, number formats and data layout to suit the hardware.

All three levels

Beginner. Designing the chip and the programs that will run on it together, so each fits the other.

Novice. Deciding what goes in hardware and what stays in software at the same time, and adapting the algorithm, number formats and data layout to suit the hardware.

Expert. For accelerators it often means trading extra (cheap) computation for less (expensive) memory traffic, choosing number formats, and settling how the compiler sees on-chip memory early enough to shape scratchpad sizes.

HBM (High Bandwidth Memory)

A tower of DRAM memory chips stacked on top of each other and connected by vertical metal holes through the silicon. It sits beside the processor in the same package and talks to it over a very wide connection, about a thousand wires or more.

All three levels

Beginner. Memory chips stacked into a tower right next to the main chip. Thousands of tiny wires link them, so data flows very fast.

Novice. A tower of DRAM memory chips stacked on top of each other and connected by vertical metal holes through the silicon. It sits beside the processor in the same package and talks to it over a very wide connection, about a thousand wires or more.

Expert. JEDEC-standard stacked DRAM with a 1,024-bit (HBM3) or 2,048-bit (HBM4) interface per stack, wired through an interposer or bridge. The PHY on the processor die, the interposer routing and heat flowing between logic and DRAM all constrain the floorplan.

HBM base die

The logic die at the bottom of an HBM stack. Through-silicon vias bring data down from the DRAM dies above, and the base die drives the wide interface to the processor through its microbumps.

All three levels

Beginner. The bottom chip in a memory stack. It doesn’t store data. It passes data between the memory chips above it and the processor.

Novice. The logic die at the bottom of an HBM stack. Through-silicon vias bring data down from the DRAM dies above, and the base die drives the wide interface to the processor through its microbumps.

Expert. Holds the PHYs that drive the stack’s interface. That interface sits within a few millimeters of the processor’s PHY because the unterminated interposer wires lose bandwidth quickly with length.

HDE (Humanitarian Device Exemption)

An FDA approval route for devices that treat rare conditions. The device must be shown to be safe and to offer a probable benefit, but it is exempt from proving effectiveness to the usual standard.

All three levels

Beginner. A special U.S. approval for devices that help small groups of people with rare illnesses.

Novice. An FDA approval route for devices that treat rare conditions. The device must be shown to be safe and to offer a probable benefit, but it is exempt from proving effectiveness to the usual standard.

Expert. Approval rests on reasonable assurance of safety and probable benefit, documented in a public Summary of Safety and Probable Benefit.

Head-of-line blocking

A packet that can’t move blocks the packets queued behind it, even if those are headed to an idle destination. Pausing a whole traffic class on a link, as PFC does, causes it.

All three levels

Beginner. When one stuck item at the front of a line holds up everything behind it, even items going somewhere else.

Novice. A packet that can’t move blocks the packets queued behind it, even if those are headed to an idle destination. Pausing a whole traffic class on a link, as PFC does, causes it.

Expert. Victim flows that share a queue or a paused priority with a congested flow lose throughput. PFC works per priority, not per flow, so a single hot spot can slow unrelated traffic several hops away.

hermetic package

A sealed case, often titanium or ceramic, that keeps moisture and body fluids away from an implant’s electronics. Wires leave it through sealed pass-throughs called feedthroughs.

All three levels

Beginner. A sealed metal or ceramic case that keeps body fluids away from the electronics.

Novice. A sealed case, often titanium or ceramic, that keeps moisture and body fluids away from an implant’s electronics. Wires leave it through sealed pass-throughs called feedthroughs.

Expert. Qualified by leak and internal water-vapor tests, corrosion soak tests and environmental stress. The number of feedthroughs limits how many electrodes the chip inside can drive.

Heterogeneous (big/little) cores

A multicore design mixing core types that run the same instructions but sit at different speed and power points. The operating system places each task on a core that fits it. Arm calls the approach big.LITTLE; Intel calls its two kinds P-cores and E-cores.

All three levels

Beginner. A chip with two kinds of cores: big fast ones for urgent work and small thrifty ones for background work.

Novice. A multicore design mixing core types that run the same instructions but sit at different speed and power points. The operating system places each task on a core that fits it. Arm calls the approach big.LITTLE; Intel calls its two kinds P-cores and E-cores.

Expert. Single-ISA asymmetry: big out-of-order cores for latency-critical or serial phases, small cores for throughput per watt and per mm². Needs capacity- and energy-aware scheduling and ISA parity across core types, including vector width.

Heterogeneous integration

Combining dies from different process nodes, factories or technologies (logic, memory, analog, photonics) in one package so each part uses the process that suits it.

All three levels

Beginner. Building one product from chips made in different factories, each picked for what it does best.

Novice. Combining dies from different process nodes, factories or technologies (logic, memory, analog, photonics) in one package so each part uses the process that suits it.

Expert. The architectural payoff of chiplets beyond yield: dense logic on the newest node, analog PHYs and SRAM-heavy I/O dies on cheaper nodes, DRAM from a memory process, all joined by die-to-die links.

Hierarchical design (partitioning)

Splitting a large chip into blocks that are laid out separately, often by different teams, each with its own outline, pins and timing targets, and then assembling them at the top level.

All three levels

Beginner. Splitting a huge chip into big pieces that different teams finish on their own, then fitting the pieces together.

Novice. Splitting a large chip into blocks that are laid out separately, often by different teams, each with its own outline, pins and timing targets, and then assembling them at the top level.

Expert. Top-down planning sets block outlines, pin positions and timing budgets from a full-chip view; bottom-up assembly integrates finished blocks. Styles range from channeled (top-level wiring and buffers in channels between blocks) to abutted (no channels). Fixed-outline floorplanning is the matching optimization problem.

High-k/metal gate (HKMG)

A gate stack that replaces silicon dioxide with a material of higher dielectric constant (‘high-k’, such as hafnium oxide) and the silicon gate electrode with metal. It allows a thinner equivalent insulator without a big jump in gate leakage.

All three levels

Beginner. A newer recipe for the gate: a special insulator and a metal electrode instead of glass and silicon, so the gate keeps control without leaking.

Novice. A gate stack that replaces silicon dioxide with a material of higher dielectric constant (‘high-k’, such as hafnium oxide) and the silicon gate electrode with metal. It allows a thinner equivalent insulator without a big jump in gate leakage.

Expert. Introduced at the 45 nm generation. High-k raises CoxC_{\mathrm{ox}} at a physical thickness that keeps direct tunneling down; the metal removes poly depletion and lets work-function metals set VTV_{\mathrm{T}} in undoped multigate channels.

High-level synthesis (HLS)

A tool that writes RTL from an algorithm in C, C++ or SystemC. The input says nothing about clock ticks; the tool decides which operations happen in which tick and how much hardware to build for them.

All three levels

Beginner. A tool that writes the detailed chip description for you from a program written in a language like C++.

Novice. A tool that writes RTL from an algorithm in C, C++ or SystemC. The input says nothing about clock ticks; the tool decides which operations happen in which tick and how much hardware to build for them.

Expert. Works in three steps: scheduling assigns operations to clock cycles, allocation chooses how many hardware units to build, and binding maps each operation onto a unit. Directives from the designer steer the trade between fewer cycles and less hardware.

Hold time

The minimum time data must stay steady after the capturing tick. If new data races in too soon, it overwrites the value being captured. The hold check: launch arrival + data delay must be no earlier than capture arrival + hold time. The clock period does not appear in it.

All three levels

Beginner. How long the data must stay unchanged after the clock tick so the memory cell isn’t confused.

Novice. The minimum time data must stay steady after the capturing tick. If new data races in too soon, it overwrites the value being captured. The hold check: launch arrival + data delay must be no earlier than capture arrival + hold time. The clock period does not appear in it.

Expert. The hold check pairs the fastest (early) data with the latest capture clock. Because the clock frequency does not appear in it, a hold failure in silicon breaks the chip at any speed. Most hold-fixing delay cells are added right after CTS, once real skew is known.

Hole

The absence of an electron in the valence band, treated as a particle with positive charge. Holes are the majority carriers in p-type silicon and carry current just as electrons do.

All three levels

Beginner. An empty spot where an electron is missing. Nearby electrons hop into it, so the hole moves around like a bubble and acts like a positive charge.

Novice. The absence of an electron in the valence band, treated as a particle with positive charge. Holes are the majority carriers in p-type silicon and carry current just as electrons do.

Expert. Modeled with its own effective mass and mobility (Si: about 470 cm²/V·s in lightly doped material, vs 1400 for electrons). PMOS transistors conduct with holes, which is one reason they are drawn wider than NMOS for the same current.

Host CPU

The central processing unit that boots the machine, runs the operating system and the program’s control logic, and talks to every other device. In an AI server it prepares data and launches work on the accelerators.

All three levels

Beginner. The server’s main, all-purpose processor. It runs the machine and hands big math jobs to the accelerators.

Novice. The central processing unit that boots the machine, runs the operating system and the program’s control logic, and talks to every other device. In an AI server it prepares data and launches work on the accelerators.

Expert. The CPU socket(s) whose root complex owns the PCIe hierarchy, the main DRAM channels and the I/O memory map. Its lane count, memory bandwidth and core count bound how many devices it can feed and how fast it can stage data for them.

HPWL (half-perimeter wirelength)

A quick estimate of how much wire one connection needs: draw the smallest rectangle around all the pins it joins and add the rectangle’s width and height. Added up over every connection, it gives one number for the placer to shrink.

All three levels

Beginner. A quick guess of how long a wire will be: draw the smallest box around everything it connects, then add the box’s width and height.

Novice. A quick estimate of how much wire one connection needs: draw the smallest rectangle around all the pins it joins and add the rectangle’s width and height. Added up over every connection, it gives one number for the placer to shrink.

Expert. Width plus height of a net’s pin bounding box. It equals the shortest rectilinear route for a two-pin net and is a lower bound for nets with more pins, since any tree joining the pins spans the box in both directions. It is built from max and min, so it has no smooth gradient; analytical placers optimize smooth stand-ins (log-sum-exp, weighted-average) and report true HPWL.

HSM

Hardware security module: a tamper-resistant device that creates, stores and uses encryption keys without ever revealing them to the computer it is attached to.

All three levels

Beginner. A locked, tamper-proof box that banks use to keep and use their most secret codes.

Novice. Hardware security module: a tamper-resistant device that creates, stores and uses encryption keys without ever revealing them to the computer it is attached to.

Expert. Usually validated to FIPS 140-3. Built from secure processors, crypto accelerators and tamper detection that erases keys on intrusion.

HTOL (high temperature operating life)

A stress test that runs chips at high temperature and their highest allowed voltage, typically for 1,000 hours (about six weeks), testing them before and after. The JEDEC standard JESD22-A108 describes the method.

All three levels

Beginner. A test that runs chips hot and powered for about 1,000 hours to show they won’t wear out too soon.

Novice. A stress test that runs chips at high temperature and their highest allowed voltage, typically for 1,000 hours (about six weeks), testing them before and after. The JEDEC standard JESD22-A108 describes the method.

Expert. Heat and voltage speed up wear-out, so 1,000 hours stands in for years of use. Zero failures gives only an upper bound on the failure rate, and one stress condition speeds up mechanisms with different activation energies by different amounts.

hybrid bonding

A way of stacking chips by pressing their polished surfaces together so copper pads fuse to copper pads, with no solder. Because there is no solder bump, the connections can be far closer together than in older methods.

All three levels

Beginner. A way to join two chips face to face, with copper pads pressed onto copper pads. No solder is needed, so the links can be packed very close.

Novice. A way of stacking chips by pressing their polished surfaces together so copper pads fuse to copper pads, with no solder. Because there is no solder bump, the connections can be far closer together than in older methods.

Expert. Bumpless copper-and-dielectric bonding at pitches from several micrometers down to sub-micrometer in research. It needs extremely flat, clean surfaces, and it makes testing each die before stacking, and getting heat out of the stack, harder.

I/O ring (pad ring)

A ring of special cells around the edge of the chip, each with a pad where a bond wire attaches. Signal cells contain drivers strong enough for the outside world and protection against static electricity; others bring in power and ground. Corner and filler cells close the ring.

All three levels

Beginner. The ring of connection points around the edge of a chip, where it plugs into the outside world.

Novice. A ring of special cells around the edge of the chip, each with a pad where a bond wire attaches. Signal cells contain drivers strong enough for the outside world and protection against static electricity; others bring in power and ground. Corner and filler cells close the ring.

Expert. Built from I/O cells in dedicated I/O rows, joined end to end so their supply rails connect, with corner and filler cells closing gaps. The ring is then wire-bonded or connected to flip-chip bumps through a redistribution layer. Multiple I/O voltage groups, multi-row rings and pad spacing drive die size when the design is pad-limited.

IC3 / PDR

IC3, also called property directed reachability (PDR), is a formal proof method. Instead of unrolling many cycles, it learns small facts that rule out states leading to a failure, until those facts add up to a full proof.

All three levels

Beginner. A modern proof method that builds up small facts about a design, one at a time, until they add up to a full proof.

Novice. IC3, also called property directed reachability (PDR), is a formal proof method. Instead of unrolling many cycles, it learns small facts that rule out states leading to a failure, until those facts add up to a full proof.

Expert. Keeps a series of frames, each a formula covering at least every state reachable in ii steps. It blocks each state that could lead to a violation with a small learned clause, and stops when two neighboring frames become identical. Runs many small, incremental SAT queries.

Ideal clock

A stand-in clock used before the clock network exists: the timing tool pretends it reaches every flip-flop instantly, or after a delay the engineer estimates. Synthesis and placement time the design against it.

All three levels

Beginner. A pretend clock that reaches everything at once. Engineers use it before the real clock wiring exists.

Novice. A stand-in clock used before the clock network exists: the timing tool pretends it reaches every flip-flop instantly, or after a delay the engineer estimates. Synthesis and placement time the design against it.

Expert. In ideal mode, set_clock_latency, set_clock_transition and an enlarged set_clock_uncertainty describe the tree that doesn’t exist yet. Forgetting to remove them after CTS counts the tree twice.

IJTAG (IEEE 1687)

IEEE 1687 (internal JTAG) defines a network of internal test chains behind the JTAG port that can be reconfigured to reach the chip’s built-in test and monitoring circuits, such as memory testers and temperature sensors.

All three levels

Beginner. A standard way to reach the many small test circuits hidden inside a chip through its test plug.

Novice. IEEE 1687 (internal JTAG) defines a network of internal test chains behind the JTAG port that can be reconfigured to reach the chip’s built-in test and monitoring circuits, such as memory testers and temperature sensors.

Expert. Segment insertion bits (SIBs) open or close network segments so only the instruments in use sit in the scan path. ICL describes the network; PDL describes instrument operations, which tools retarget to chip-level patterns.

ILP wall

Diminishing returns from instruction-level parallelism (ILP), the number of instructions in one program that can run at the same time. Dependences between instructions, branches and cache misses cap it, so making one core wider or deeper gains less and less.

All three levels

Beginner. The limit on how many steps of one program a core can do at once. Many steps must wait for the step before, so a wider core stops helping.

Novice. Diminishing returns from instruction-level parallelism (ILP), the number of instructions in one program that can run at the same time. Dependences between instructions, branches and cache misses cap it, so making one core wider or deeper gains less and less.

Expert. Even an ideal window is bounded by the critical path of the dataflow graph; real cores also pay for mispredictions and misses. Tullsen et al. measured under 1.5 IPC on an 8-issue model; the cost of wider issue (rename, wakeup, bypass) grows faster than the IPC it buys.

im2col (lowering)

“Image to column”: copying every input patch a filter touches into one row (or column) of a matrix, so a convolution becomes a single matrix multiplication with the filters.

All three levels

Beginner. Rearranging an image into a big table, copying overlapping patches, so a picture-filtering job becomes an ordinary grid multiplication.

Novice. “Image to column”: copying every input patch a filter touches into one row (or column) of a matrix, so a convolution becomes a single matrix multiplication with the filters.

Expert. Produces a CRS×NPQCRS \times NPQ data matrix in which each input value appears up to R⋅SR \cdot S times. Materializing it costs memory and traffic, so libraries and accelerators (Gemmini, for example) generate it on the fly at the array edge.

Immersion cooling

Servers sit in a tank of dielectric (electrically insulating) fluid. In single-phase immersion an oil-like fluid is pumped to a heat exchanger; in two-phase immersion a fluid boils on the hot parts and condenses on a coil above.

All three levels

Beginner. Dunking whole computers in a special liquid that doesn’t carry electricity. The liquid soaks up all their heat.

Novice. Servers sit in a tank of dielectric (electrically insulating) fluid. In single-phase immersion an oil-like fluid is pumped to a heat exchanger; in two-phase immersion a fluid boils on the hot parts and condenses on a coil above.

Expert. Captures nearly all heat with no server fans and tolerates high temperatures, but complicates service (lifting boards out of fluid), material compatibility and warranties. Many two-phase fluids are fluorinated chemicals, whose supply is shrinking under PFAS pressure.

Immersion lithography

193 nm lithography with ultrapure water (refractive index 1.44 at that wavelength) filling the gap under the last lens. That raises the numerical aperture above 1, up to about 1.35.

All three levels

Beginner. Printing chip patterns with a thin layer of water between the lens and the wafer, which lets the machine print finer detail.

Novice. 193 nm lithography with ultrapure water (refractive index 1.44 at that wavelength) filling the gap under the last lens. That raises the numerical aperture above 1, up to about 1.35.

Expert. Single exposure reaches roughly 75–80 nm pitch; tighter layers need multiple patterning. Still the workhorse for non-critical layers and for spacer-patterned gratings even where EUV is used.

implantable device

An electronic medical device placed inside the body by surgery and left there for years, such as a pacemaker or a nerve stimulator. It runs on a battery that can’t easily be replaced, a rechargeable battery, or power sent wirelessly through the skin.

All three levels

Beginner. A medical device put inside the body, like a pacemaker, and left there for years.

Novice. An electronic medical device placed inside the body by surgery and left there for years, such as a pacemaker or a nerve stimulator. It runs on a battery that can’t easily be replaced, a rechargeable battery, or power sent wirelessly through the skin.

Expert. A device whose electronics sit in a hermetic package exposed to body temperature and fluids for their whole life. Replacement means surgery, so battery life and reliability dominate the design.

In-network reduction

A switch feature in which the switch performs the additions of a reduction as data flows through it and sends the result back to every chip.

All three levels

Beginner. The switch itself adds up the numbers passing through it. So the chips don’t have to send everything back and forth twice.

Novice. A switch feature in which the switch performs the additions of a reduction as data flows through it and sends the result back to every chip.

Expert. Each endpoint sends its buffer once and receives the result once, so the per-endpoint traffic is about nn instead of 2(p−1)/p⋅n2(p-1)/p \cdot n, roughly halving all-reduce time when bandwidth-bound. NVIDIA calls its version SHARP; the switch needs arithmetic units and buffer space.

Incast

A many-to-one traffic pattern in which several senders transmit to one receiver at once. Their combined rate exceeds the receiver’s link, so the last switch’s buffer fills and packets are delayed, dropped or paused.

All three levels

Beginner. Many senders all sending to the same receiver at the same moment, overflowing its connection.

Novice. A many-to-one traffic pattern in which several senders transmit to one receiver at once. Their combined rate exceeds the receiver’s link, so the last switch’s buffer fills and packets are delayed, dropped or paused.

Expert. Synchronized fan-in that oversubscribes the last-hop link. Common in collectives and storage; switch buffers are small on-chip memories shared across all ports, so line-rate incast from even a handful of senders fills them in microseconds.

Inference

Running a trained model forward to get an output. It happens every time someone uses the model, so its total cost over a model’s life can exceed training’s, and it usually has a response-time limit.

All three levels

Beginner. Using a trained neural network to answer a question, recognize a picture or write text.

Novice. Running a trained model forward to get an output. It happens every time someone uses the model, so its total cost over a model’s life can exceed training’s, and it usually has a response-time limit.

Expert. Forward pass only (2N2N FLOPs per token), with latency targets per request. Small batches and the sequential decode loop make it far more often memory-bound than training.

Inference scenario

MLPerf Inference has four: Single stream (one query at a time, latency measured), Multistream (groups of 8), Server (random arrivals, highest rate that still meets a latency limit) and Offline (everything at once, raw throughput).

All three levels

Beginner. The pattern in which questions arrive during an MLPerf test: one at a time, in a steady stream from many users, or all at once in a big pile.

Novice. MLPerf Inference has four: Single stream (one query at a time, latency measured), Multistream (groups of 8), Server (random arrivals, highest rate that still meets a latency limit) and Offline (everything at once, raw throughput).

Expert. Server arrivals follow a Poisson process and the metric is the highest arrival rate at which the tail latency still meets the benchmark’s bound (for LLMs, time to first token and time per output token). Offline needs at least 24,576 samples.

Inferred latch

A storage element a tool adds by accident. If code meant to have no memory leaves a signal unassigned in some case, the signal must keep its old value, and keeping a value needs memory: a latch, which passes its input through while enabled and freezes when not.

All three levels

Beginner. An accidental memory the tool builds when the design forgets to say what a signal should be in some situation.

Novice. A storage element a tool adds by accident. If code meant to have no memory leaves a signal unassigned in some case, the signal must keep its old value, and keeping a value needs memory: a latch, which passes its input through while enabled and freezes when not.

Expert. Usually a bug: unintended state that complicates timing analysis and test. Default assignments at the top of always_comb, a default branch in every case, and lint (Verilator’s LATCH warning) catch it. A latch you do want goes in always_latch.

InfiniBand

A networking standard designed for high-performance computing with RDMA built in. Switches only send data when the next hop has room for it (credit-based flow control), so packets are not dropped when the network is busy.

All three levels

Beginner. A kind of network built for supercomputers. Its switches never throw data away.

Novice. A networking standard designed for high-performance computing with RDMA built in. Switches only send data when the next hop has room for it (credit-based flow control), so packets are not dropped when the network is busy.

Expert. An end-to-end architecture from the InfiniBand Trade Association: RDMA transport over a lossless, in-order link layer with per-virtual-lane credits, centrally computed forwarding tables from a subnet manager, and features such as adaptive routing and in-network reduction in recent switches.

Inner spacer

Small pieces of insulator placed at the ends of the silicon-germanium layers, between neighboring nanosheets. They keep the gate metal that later fills those gaps away from the source and drain.

All three levels

Beginner. A tiny insulating plug between the gate and the source and drain, tucked between the stacked ribbons.

Novice. Small pieces of insulator placed at the ends of the silicon-germanium layers, between neighboring nanosheets. They keep the gate metal that later fills those gaps away from the source and drain.

Expert. Formed by a lateral cavity etch of the SiGe ends, dielectric deposition and etch-back. It cuts gate-to-S/D capacitance and shields the S/D epitaxy during channel release. Too shallow leaves capacitance; too deep adds extension resistance; published work puts the optimum depth under about 5 nm.

Insertion delay (clock latency)

How long the clock tick takes to travel from where it is defined (usually the chip’s clock pin) to a flip-flop. Delay before the chip pin is source latency; delay through the on-chip network is network latency.

All three levels

Beginner. How long the clock tick takes to travel from where it enters the chip to a memory cell.

Novice. How long the clock tick takes to travel from where it is defined (usually the chip’s clock pin) to a flip-flop. Delay before the chip pin is source latency; delay through the on-chip network is network latency.

Expert. Latency matters beyond skew. The part of two flip-flops’ clock paths that is not shared gets variation margins applied to it, so longer latency costs slack; it also collects more jitter and burns more power. Before CTS it is estimated with set_clock_latency; afterwards the timer computes it.

Insertion loss

The fraction of signal power lost along a channel, given in decibels (dB). 10 dB means one tenth of the power arrives; 20 dB means one hundredth. Copper loss grows with both length and frequency.

All three levels

Beginner. How much weaker a signal gets by the time it reaches the far end of a wire.

Novice. The fraction of signal power lost along a channel, given in decibels (dB). 10 dB means one tenth of the power arrives; 20 dB means one hundredth. Copper loss grows with both length and frequency.

Expert. |S21| in dB, quoted at the Nyquist frequency. For twinax and PCB traces it is roughly afa\sqrt{f} (skin effect) + bfb f (dielectric loss) per unit length, so it scales linearly with length and faster than linearly with lane rate. Standards specify end-to-end budgets, such as 28 dB at 26.56 GHz for 100 Gb/s lanes.

Instruction set architecture (ISA)

The contract between software and hardware: the instructions, registers, data types and memory behavior a program may rely on. Any chip that implements the same ISA runs the same programs. x86, Arm and RISC-V are ISAs.

All three levels

Beginner. The list of commands a kind of chip understands, and exactly what each one does. Programs are written for it, and chips are built to follow it.

Novice. The contract between software and hardware: the instructions, registers, data types and memory behavior a program may rely on. Any chip that implements the same ISA runs the same programs. x86, Arm and RISC-V are ISAs.

Expert. The architectural state and the semantics of every instruction, including exceptions and memory ordering. Timing is deliberately left out, so microarchitectures can differ. The IBM System/360 (1964) first separated the two.

Instruction window

The set of instructions an out-of-order core has fetched but not yet retired, from which it picks ready ones to run. A bigger window can find more independent work, for example while a load waits on memory.

All three levels

Beginner. How many steps of the program a processor can hold at once while it looks for ones that are ready to run.

Novice. The set of instructions an out-of-order core has fetched but not yet retired, from which it picks ready ones to run. A bigger window can find more independent work, for example while a load waits on memory.

Expert. Bounded by the ROB and by every other in-flight resource (physical registers, issue-queue entries, load and store queue entries). To keep issuing W per cycle through a latency L takes about W·L entries.

Integer quantization (INT8, INT4)

Representing values as signed integers times a scale: x≈s×qx \approx s \times q with qq from −127 to 127 for INT8 or −7 to 7 for INT4. The steps are evenly spaced, so the format is precise for values near the top of the range and coarse for small ones relative to their size.

All three levels

Beginner. Storing numbers as small whole numbers, like −127 to 127, plus one shared multiplier.

Novice. Representing values as signed integers times a scale: x≈s×qx \approx s \times q with qq from −127 to 127 for INT8 or −7 to 7 for INT4. The steps are evenly spaced, so the format is precise for values near the top of the range and coarse for small ones relative to their size.

Expert. Symmetric (z=0z = 0) or affine (zero-point zz) uniform quantizers. Products of INT8 operands are exact 16-bit integers and accumulate exactly in INT32, which makes the hardware small. Accuracy depends on granularity and calibration, and on handling outliers that stretch the range.

Integrated logic analyzer

Debug logic added to an FPGA design. It samples chosen internal signals every clock into on-chip memory, stops after a trigger condition, and sends the samples to a computer to be viewed as waveforms. FPGA vendors ship their own versions.

All three levels

Beginner. A tiny recorder built into an FPGA design. It watches chosen signals and saves what happens around an event.

Novice. Debug logic added to an FPGA design. It samples chosen internal signals every clock into on-chip memory, stops after a trigger condition, and sends the samples to a computer to be viewed as waveforms. FPGA vendors ship their own versions.

Expert. An embedded capture core: probe ports, trigger comparators, a circular buffer in block RAM and readout over JTAG. It costs RAM, logic and a recompile, and depth × width limits what it sees. The custom-chip equivalent is a trace buffer.

Intrinsic carrier concentration (ni)

The density of electrons (equal to the density of holes) in undoped silicon, about 101010^{10} per cm³ at room temperature. For comparison there are 5×10225 \times 10^{22} silicon atoms per cm³.

All three levels

Beginner. How many electrons heat alone knocks free in perfectly pure silicon. It is very few.

Novice. The density of electrons (equal to the density of holes) in undoped silicon, about 101010^{10} per cm³ at room temperature. For comparison there are 5×10225 \times 10^{22} silicon atoms per cm³.

Expert. ni=NcNv e−Eg/2kTn_{\mathrm{i}} = \sqrt{N_{\mathrm{c}} N_{\mathrm{v}}}\, e^{-E_{\mathrm{g}}/2kT}; about 9.65×109 cm−39.65 \times 10^9\,\mathrm{cm^{-3}} at 300 K by the modern measurement. It rises steeply with temperature, and so do junction leakage (∝ni2\propto n_{\mathrm{i}}^2) and the drift of VTV_{\mathrm{T}}.

Inverse lithography technology (ILT)

Treating mask design as an inverse problem: change the mask, often pixel by pixel, until a simulation of the printed result matches the shape you want.

All three levels

Beginner. Letting a computer work backwards from the shape you want on the wafer to the best possible stencil.

Novice. Treating mask design as an inverse problem: change the mask, often pixel by pixel, until a simulation of the printed result matches the shape you want.

Expert. Gradient-based optimization of mask pixels through a differentiable lithography model. More precise than edge-based OPC, but the curved results must be broken into shapes a mask writer can draw, which raises cost, so it is used selectively.

Inversion layer

A layer at the silicon surface where the minority carrier has become the majority, for example electrons at the surface of p-type silicon. It is only 1–2 nm thick and connects the source to the drain.

All three levels

Beginner. A thin sheet of electrons that forms just under the gate when the gate voltage is high enough. It is the transistor’s bridge.

Novice. A layer at the silicon surface where the minority carrier has become the majority, for example electrons at the surface of p-type silicon. It is only 1–2 nm thick and connects the source to the drain.

Expert. Defined as starting when surface electron density equals NaN_{\mathrm{a}} (ϕs=2ϕB\phi_{\mathrm{s}} = 2\phi_{\mathrm{B}}). Above threshold its sheet charge is Qinv≈−Cox(Vg−VT)Q_{\mathrm{inv}} \approx -C_{\mathrm{ox}}(V_{\mathrm{g}} - V_{\mathrm{T}}) because the depletion charge barely changes. Its finite thickness (TinvT_{\mathrm{inv}}) adds about Tinv/3T_{\mathrm{inv}}/3 to the electrical oxide thickness.

Inverter (NOT gate)

A gate whose output is 1 when the input is 0 and 0 when the input is 1. In CMOS it is one PMOS from the supply to the output and one NMOS from the output to ground, with both gates tied to the input.

All three levels

Beginner. The simplest logic gate: it outputs the opposite of its input.

Novice. A gate whose output is 1 when the input is 0 and 0 when the input is 1. In CMOS it is one PMOS from the supply to the output and one NMOS from the output to ground, with both gates tied to the input.

Expert. One PMOS, one NMOS, gates common. The reference gate for delay (FO4) and logical effort (g=1g = 1), and the building block of buffers, latches and SRAM cells. Its DC transfer curve defines switching threshold and noise margins.

Ion implantation

Dopant atoms such as boron, phosphorus or arsenic are ionized, accelerated (typically 5–200 keV) and shot into the wafer through openings in a mask. A heat treatment (anneal) then repairs the crystal and activates the dopant.

All three levels

Beginner. Firing atoms of another element into the silicon at high speed. This changes how well that patch carries electricity.

Novice. Dopant atoms such as boron, phosphorus or arsenic are ionized, accelerated (typically 5–200 keV) and shot into the wafer through openings in a mask. A heat treatment (anneal) then repairs the crystal and activates the dopant.

Expert. Dose (ions/cm²) and energy set the amount and depth; the profile peaks below the surface. Channeling along crystal axes is avoided by tilting. Every later hot step moves the profile, so the thermal budget is planned across the whole flow.

IP block (silicon IP)

A ready-made design block licensed from a supplier, reused from an earlier chip or taken from open source: processor cores, interconnect, memory controllers, PHYs (the mostly analog circuits that drive signals on and off the chip), memory generators. It comes as hardware-description code (soft IP) or as a finished layout (hard IP).

All three levels

Beginner. A ready-made piece of chip design, bought or reused, so a team doesn’t have to design every part from scratch.

Novice. A ready-made design block licensed from a supplier, reused from an earlier chip or taken from open source: processor cores, interconnect, memory controllers, PHYs (the mostly analog circuits that drive signals on and off the chip), memory generators. It comes as hardware-description code (soft IP) or as a finished layout (hard IP).

Expert. Make-or-buy weighs differentiation, schedule, license fees and royalties, verification quality, configurability and integration risk: protocol versions, clocking, test hooks and power intent must all match the rest of the chip. Hard IP also fixes its shape and pin positions in the floorplan.

IPC (instructions per cycle)

The average number of instructions completed per clock cycle, the inverse of CPI. A simple pipeline reaches at most 1; chips that work on several instructions at once can exceed it.

All three levels

Beginner. How many commands a chip finishes on each tick of its clock, on average.

Novice. The average number of instructions completed per clock cycle, the inverse of CPI. A simple pipeline reaches at most 1; chips that work on several instructions at once can exceed it.

Expert. IPC=1/CPI\mathrm{IPC} = 1/\mathrm{CPI}. Superscalar out-of-order cores sustain IPC above 1 on code with enough independent instructions; misses and mispredictions pull it down.

IR drop

The voltage lost across the resistance of the power wiring: drop=current×resistance\text{drop} = \text{current} \times \text{resistance} (V=IRV = IR, hence “IR”). Cells far from the supply or in busy areas get less than the full supply voltage, and lower voltage makes them slower.

All three levels

Beginner. The push of electricity (the voltage) that gets lost along a chip’s thin power wires, like water pressure dropping along a long pipe.

Novice. The voltage lost across the resistance of the power wiring: drop=current×resistance\text{drop} = \text{current} \times \text{resistance} (V=IRV = IR, hence “IR”). Cells far from the supply or in busy areas get less than the full supply voltage, and lower voltage makes them slower.

Expert. Static IR drop uses each cell’s average current and is a pure resistance problem. Dynamic IR drop follows current over time as many cells switch together, so capacitance and package inductance matter too. Both are budgeted against the supply voltage the timing analysis assumes.

Iron law of processor performance

Run time = instructions in the program × average cycles per instruction (CPI) × time per cycle. Every design idea changes one or more of the three, often making one better and another worse.

All three levels

Beginner. The rule that a program’s time is: how many steps it has, times the ticks each step takes, times how long one tick lasts.

Novice. Run time = instructions in the program × average cycles per instruction (CPI) × time per cycle. Every design idea changes one or more of the three, often making one better and another worse.

Expert. T=Ninstr⋅CPI⋅tcycleT = N_{\mathrm{instr}} \cdot \mathrm{CPI} \cdot t_{\mathrm{cycle}}. Instruction count depends on the ISA and compiler, CPI on the ISA and microarchitecture, cycle time on the microarchitecture and circuits. Credited to Emer and Clark’s VAX-11/780 study.

isolation cell

A small gate placed on a wire leaving a block that can be switched off. While the block is off, its outputs drift to random values; the isolation cell holds the wire at a fixed 0 or 1 so those values can’t confuse the parts still running.

All three levels

Beginner. A tiny gate that holds a wire steady while the part of the chip feeding it is switched off.

Novice. A small gate placed on a wire leaving a block that can be switched off. While the block is off, its outputs drift to random values; the isolation cell holds the wire at a fixed 0 or 1 so those values can’t confuse the parts still running.

Expert. Inserted per UPF on domain crossings and powered from the receiving or always-on supply. Its enable must be asserted before power-down and released only after power-up completes.

ITAR

International Traffic in Arms Regulations: U.S. State Department rules for items on the U.S. Munitions List. Exporting those items, or the technical information that describes them, needs government permission.

All three levels

Beginner. U.S. rules on who can get weapons tech. They cover the designs for some chips too.

Novice. International Traffic in Arms Regulations: U.S. State Department rules for items on the U.S. Munitions List. Exporting those items, or the technical information that describes them, needs government permission.

Expert. USML Category XI(c)(1) covers ASICs and programmable logic devices programmed for defense articles, and Category XI(d) covers the technical data directly related to them, so the design database is controlled. Releasing it to a foreign person, even inside the U.S., counts as an export.

ITUE (IT-power usage effectiveness)

The total power into the IT equipment divided by the power used by the computing parts (processors, memory, storage). The rest goes to fans, power supplies and voltage regulators.

All three levels

Beginner. A score like PUE, but for inside each computer: how much of the power going in actually reaches the chips.

Novice. The total power into the IT equipment divided by the power used by the computing parts (processors, memory, storage). The rest goes to fans, power supplies and voltage regulators.

Expert. Proposed by the Energy Efficient HPC Working Group as “PUE for the server”: ITUE=total IT energy÷compute-component energy\mathrm{ITUE} = \text{total IT energy} \div \text{compute-component energy}. Fans, PSUs, voltage regulators and management controllers are overhead; it needs component-level power telemetry to measure well.

K-feasible cut

For one gate, a set of at most K signals that every path from the circuit’s inputs to that gate must pass through. The logic between the cut and the gate depends only on those K signals, so one K-input lookup table can replace all of it.

All three levels

Beginner. A way to cut out a piece of a circuit that needs only a few input signals, few enough to fit in one lookup table.

Novice. For one gate, a set of at most K signals that every path from the circuit’s inputs to that gate must pass through. The logic between the cut and the gate depends only on those K signals, so one K-input lookup table can replace all of it.

Expert. A cut C of node n with |C| ≤ K; the cone between C and n is a function of at most K inputs. Cuts are enumerated bottom-up by merging the fanins’ cut sets and dropping dominated cuts; LUT mapping picks one cut per mapped node.

k-induction

A two-part proof, like showing that a row of dominoes will all fall: first show the rule holds for the first kk cycles, then show that any kk good cycles in a row are always followed by another good one.

All three levels

Beginner. A proof that works like falling dominoes. If a rule holds for a run of steps, it must hold for the next step too.

Novice. A two-part proof, like showing that a row of dominoes will all fall: first show the rule holds for the first kk cycles, then show that any kk good cycles in a row are always followed by another good one.

Expert. The second part starts from any state at all, including states the design can never reach, so it can fail on a harmless state. Remedies: a larger kk, requiring the states on the path to differ, or helper assertions. SymbiYosys’s smtbmc engine uses it in prove mode.

Kernel (GPU)

A function launched on the GPU across a grid of threads. Each thread runs the same code and uses its thread number to pick which data to work on.

All three levels

Beginner. A small program sent to the GPU that runs many thousands of times at once, once for each piece of data.

Novice. A function launched on the GPU across a grid of threads. Each thread runs the same code and uses its thread number to pick which data to work on.

Expert. A device function launched over a grid of thread blocks. Its per-thread register count, per-block shared memory and block size set occupancy; one layer of a neural network is often one or a few kernels.

Kernel fusion

Combining several operations into one program (kernel) so intermediate results stay in on-chip memory. It cuts memory traffic and raises arithmetic intensity.

All three levels

Beginner. Doing several steps on data while it’s already on the chip, instead of sending it back to memory between each step.

Novice. Combining several operations into one program (kernel) so intermediate results stay in on-chip memory. It cuts memory traffic and raises arithmetic intensity.

Expert. The standard fix for chains of memory-bound ops. FlashAttention fuses QKTQK^T, softmax and PVPV with tiling so the S×SS \times S score matrix never reaches HBM.

known-good die

A bare chip tested thoroughly while still on the wafer, before it is assembled with other chips, so it is very unlikely to make an expensive multi-chip package fail.

All three levels

Beginner. A chip that has been fully tested before it is packed in with other chips.

Novice. A bare chip tested thoroughly while still on the wafer, before it is assembled with other chips, so it is very unlikely to make an expensive multi-chip package fail.

Expert. Wafer-level test coverage high enough to keep package-level yield acceptable. It pushes at-speed test and die-to-die interface test forward to wafer sort, where probing is harder than on a packaged part.

KV cache

The keys and values that attention computed for every earlier token, stored in memory so each new token only needs its own. It grows with the length of the conversation and the number of users served at once.

All three levels

Beginner. The model’s notes on everything it has read so far, kept so it doesn’t have to reread the whole conversation for every new word.

Novice. The keys and values that attention computed for every earlier token, stored in memory so each new token only needs its own. It grows with the length of the conversation and the number of users served at once.

Expert. 2 × layers × KV heads × head dimension × bytes per token per sequence. It is read in full on every decode step, so at long contexts or large batches it, not the weights, dominates memory traffic.

Lane

One SerDes channel: a pair of wires (or one fiber, or one color of light) carrying one serial stream, such as 100 Gb/s. A port bundles several lanes: an 800 Gb/s port is often 8 lanes of 100 Gb/s.

All three levels

Beginner. One high-speed path for data, like one lane on a highway.

Novice. One SerDes channel: a pair of wires (or one fiber, or one color of light) carrying one serial stream, such as 100 Gb/s. A port bundles several lanes: an 800 Gb/s port is often 8 lanes of 100 Gb/s.

Expert. The unit of electrical signaling. Port speed = lanes × lane rate; doubling the lane rate (50 → 100 → 200 Gb/s) halves the lanes per port and the pins per Tb/s, at the cost of reach and equalization power.

Large language model (LLM)

A very large neural network trained on huge amounts of text and code to predict the next word piece (token), then tuned to follow instructions. Chatbots are built on them. In chip design they write hardware code, test code and tool scripts on request.

All three levels

Beginner. A program trained on huge amounts of text and code. It can answer questions and write text or code from a request typed in plain language.

Novice. A very large neural network trained on huge amounts of text and code to predict the next word piece (token), then tuned to follow instructions. Chatbots are built on them. In chip design they write hardware code, test code and tool scripts on request.

Expert. Output is sampled one token at a time, so the same prompt can give different answers, and settings such as temperature, the prompt and the examples in it change the quality. Nothing in the model checks that hardware code behaves correctly, so every output needs an outside check: simulation, a formal tool or human review.

Last-level cache (LLC)

The final, largest level of on-chip cache, shared by all the cores (often called L3). It is usually split into slices spread across the chip, with each memory address assigned to one slice.

All three levels

Beginner. The biggest memory on the chip, shared by all the cores. It is the last stop before the slow trip to main memory.

Novice. The final, largest level of on-chip cache, shared by all the cores (often called L3). It is usually split into slices spread across the chip, with each memory address assigned to one slice.

Expert. A banked, shared cache sliced across the interconnect, with an address hash choosing the slice; it often hosts the coherence directory or snoop filter. Inclusive or non-inclusive of the private caches; its slice distance is part of the NUMA picture on big dies.

Latency hiding

Overlapping slow operations, such as memory reads that take hundreds of clock cycles, with useful work from other threads, so the processor is rarely idle.

All three levels

Beginner. Keeping busy with other work while you wait for something slow, so the waiting doesn’t cost time.

Novice. Overlapping slow operations, such as memory reads that take hundreds of clock cycles, with useful work from other threads, so the processor is rarely idle.

Expert. On GPUs, thread-level parallelism (many resident warps) plus instruction-level parallelism (independent instructions per warp). Little’s law gives the concurrency needed: latency × throughput.

Latent fault metric (LFM)

How well a chip finds hidden faults: ones that cause no harm on their own and go unnoticed, but could combine with a second fault later. Targets: at least 60% for ASIL B, 80% for C and 90% for D.

All three levels

Beginner. How well a chip finds hidden faults. Alone they do no harm, but they could team up with a second fault.

Novice. How well a chip finds hidden faults: ones that cause no harm on their own and go unnoticed, but could combine with a second fault later. Targets: at least 60% for ASIL B, 80% for C and 90% for D.

Expert. Coverage of multiple-point faults that would otherwise stay latent. These often sit in the safety mechanisms themselves; logic and memory BIST and self-checks on the mechanisms raise LFM.

Launch-on-capture (LOC)

Also called broadside. Load the first pattern, switch scan enable off, then give two fast normal clock ticks. The first tick makes the flip-flops change to what the logic computed (the second pattern); the second tick catches the result of that change.

All three levels

Beginner. A way to start a speed test with two quick normal steps after the pattern is loaded.

Novice. Also called broadside. Load the first pattern, switch scan enable off, then give two fast normal clock ticks. The first tick makes the flip-flops change to what the logic computed (the second pattern); the second tick catches the result of that change.

Expert. Scan enable has time to settle, so it is easy to implement and close to normal operation. Because the second pattern must be the logic’s own response to the first, some transitions can’t be launched, which costs coverage compared with LOS.

Launch-on-shift (LOS)

Also called skewed-load. The last shift tick of loading creates the change (the second pattern is the first one moved along by one bit). Scan enable then switches off and one fast tick catches the result.

All three levels

Beginner. A way to start a speed test using the very last step of loading the pattern.

Novice. Also called skewed-load. The last shift tick of loading creates the change (the second pattern is the first one moved along by one bit). Scan enable then switches off and one fast tick catches the result.

Expert. Higher transition coverage, but scan enable must fall within one functional clock period chip-wide, and launch states that normal operation never reaches can fail good parts.

Layer and datatype

Two numbers on every shape in a layout file. The layer number says which physical layer the shape belongs to, such as the first metal layer (M1). The datatype separates kinds of shapes on that layer, such as real wires, dummy fill and slots.

All three levels

Beginner. Two numbers on every shape in a chip file. They say which layer it belongs to and what kind of shape it is.

Novice. Two numbers on every shape in a layout file. The layer number says which physical layer the shape belongs to, such as the first metal layer (M1). The datatype separates kinds of shapes on that layer, such as real wires, dummy fill and slots.

Expert. The numbers have no meaning in the file itself; the foundry’s layer map defines them. Mask data preparation usually merges the drawing and fill datatypes into one mask and subtracts the slot datatype, so a shape on the wrong datatype silently drops or adds metal.

Layer map

A table, set by the foundry for its process, from layer names and purposes (M1 wires, M1 fill, M1 pin labels) to the layer and datatype numbers used in the file, such as 15/0 and 15/1.

All three levels

Beginner. A lookup table that says which number in the chip file means which layer.

Novice. A table, set by the foundry for its process, from layer names and purposes (M1 wires, M1 fill, M1 pin labels) to the layer and datatype numbers used in the file, such as 15/0 and 15/1.

Expert. Every tool that writes or reads the layout must use the same map: the place-and-route export, IP deliveries, the fill tool, and the DRC and LVS rule files. A mismatch is a classic cause of a broken merge.

Layer pipelining (streaming dataflow)

Giving each layer of a model its own group of cores and streaming data from group to group. Every layer works at the same time on a different piece of data, and the slowest group sets the pace.

All three levels

Beginner. Running a model like an assembly line: each group of cores does one step and passes its result to the next group, so all steps work at once on different inputs.

Novice. Giving each layer of a model its own group of cores and streaming data from group to group. Every layer works at the same time on a different piece of data, and the slowest group sets the pace.

Expert. Spatial execution of a dataflow graph as a coarse-grained pipeline with stage buffers between operators. Throughput is set by the slowest stage or link; latency by the sum of stage delays. Requires the whole pipeline’s weights and buffers to fit on chip at once.

Layout versus schematic (LVS)

Rebuilding the circuit from the layout drawing (which transistors exist and which wires connect them) and comparing it with the circuit the designers intended. It reports shorts (wires touching that shouldn’t), opens (breaks in a wire) and missing or wrong devices.

All three levels

Beginner. Checking that the drawing is exactly the circuit the designers meant.

Novice. Rebuilding the circuit from the layout drawing (which transistors exist and which wires connect them) and comparing it with the circuit the designers intended. It reports shorts (wires touching that shouldn’t), opens (breaks in a wire) and missing or wrong devices.

Expert. Device and connectivity extraction from the layout, then comparison with a reference transistor netlist (SPICE/CDL) built from the final gate-level netlist and each cell’s transistor netlist. Comparison is graph matching by iterative partition refinement, usually hierarchical. Power pins, well ties and missing labels cause many first-run failures.

LBIST

Logic built-in self-test: circuitry on the chip that generates test patterns, runs them through the logic, and compresses the results into a signature that is compared with the expected value. The chip can test itself without outside equipment.

All three levels

Beginner. A self-test the chip can run on itself, without any outside test machines.

Novice. Logic built-in self-test: circuitry on the chip that generates test patterns, runs them through the logic, and compresses the results into a signature that is compared with the expected value. The chip can test itself without outside equipment.

Expert. Patterns from an on-chip generator are shifted through scan chains and the responses compacted into a signature. Used in the field to find latent faults; the run must fit the diagnostic test interval and power budget, and functional state must be saved or restored around it.

L·di/dt noise

A voltage dip caused by inductance, the property of any current path that resists a change in current. The dip equals the inductance LL times how fast the current changes (di/dtdi/dt), so a sudden jump in demand through the package’s wiring makes the supply sag.

All three levels

Beginner. A dip in the push of electricity caused by the chip’s need for electricity changing very fast.

Novice. A voltage dip caused by inductance, the property of any current path that resists a change in current. The dip equals the inductance LL times how fast the current changes (di/dtdi/dt), so a sudden jump in demand through the package’s wiring makes the supply sag.

Expert. Dominates dynamic droop near the chip–package resonance. Reduced by on-chip decap, lower-inductance packages and more power bumps, and by limiting how fast large blocks wake up or ramp up activity.

LDO regulator

Low-dropout regulator: a linear voltage regulator that outputs a stable voltage slightly below its input. Simple and fast, but it wastes the voltage difference as heat.

All three levels

Beginner. A small circuit that takes in electricity with a wobbly push (voltage) and turns it into a steady supply for the chip.

Novice. Low-dropout regulator: a linear voltage regulator that outputs a stable voltage slightly below its input. Simple and fast, but it wastes the voltage difference as heat.

Expert. Efficiency is roughly Vout/VinV_{\mathrm{out}}/V_{\mathrm{in}}, so a large input range costs power. Used on-die to clean up supply noise or to absorb a supply that varies with location.

Leaf–spine network

A two-tier network. Leaf switches (also called top-of-rack switches) connect to servers; spine switches connect only to leaves; every leaf links to every spine. Any server reaches any other in at most three switch hops.

All three levels

Beginner. Two layers of switches: the lower ones connect to servers, and every lower switch has a cable to every upper switch.

Novice. A two-tier network. Leaf switches (also called top-of-rack switches) connect to servers; spine switches connect only to leaves; every leaf links to every spine. Any server reaches any other in at most three switch hops.

Expert. A two-stage folded Clos. With radix kk and no oversubscription: kk leaves, k/2k/2 spines, k2/2k^2/2 endpoints, and k/2k/2 equal-cost paths between any two leaves.

Leakage (static) power

Power drawn by current that trickles through transistors that are supposed to be off. It is paid whether or not anything switches, and it grows with temperature.

All three levels

Beginner. Power a chip wastes just by being switched on, even when it is doing nothing.

Novice. Power drawn by current that trickles through transistors that are supposed to be off. It is paid whether or not anything switches, and it grows with temperature.

Expert. Mostly subthreshold leakage, plus leakage through the gate insulator and the junctions. Subthreshold current rises about tenfold for each ~100 mV drop in threshold voltage and rises with temperature, so the hot case sets the leakage budget. It is controlled with slower, low-leakage (high-threshold) cells, power gating and body bias.

LEF

Library Exchange Format: text files that describe the building blocks without their insides. A technology LEF lists the manufacturing layers and their rules; cell and macro LEFs give each block’s size, the shapes of its pins, and the areas it blocks.

All three levels

Beginner. A computer file that describes the outline and plug points of each building block, but not what’s inside.

Novice. Library Exchange Format: text files that describe the building blocks without their insides. A technology LEF lists the manufacturing layers and their rules; cell and macro LEFs give each block’s size, the shapes of its pins, and the areas it blocks.

Expert. An abstract view: SIZE, SYMMETRY, CLASS (CORE, BLOCK, PAD, ENDCAP…), PIN shapes by layer, and OBS (obstructions). Floorplanning uses macro LEF for footprints, pin sides and blocked layers, and tech LEF for placement sites, track pitches and layer directions.

Legalization

The step after global placement that moves each cell to a free slot (site) in a row, so nothing overlaps and nothing sits in a forbidden area, while moving each cell as little as possible.

All three levels

Beginner. Snapping every part into a proper slot in a row so nothing overlaps, while moving each part as little as possible.

Novice. The step after global placement that moves each cell to a free slot (site) in a row, so nothing overlaps and nothing sits in a forbidden area, while moving each cell as little as possible.

Expert. Greedy (Tetris), row-based dynamic programming (Abacus), flow-based, diffusion-based and negotiation-based methods. It must also respect fences, padding, edge-spacing rules and multi-height rail alignment. Large displacement here undoes the wirelength and timing that global placement achieved.

Level shifter

A small cell placed on a signal that crosses from logic running at one supply voltage to logic running at another. Without it, a low-voltage signal may not switch a higher-voltage receiver cleanly.

All three levels

Beginner. A small part that acts like a translator for a signal passing between two areas that run at different power levels.

Novice. A small cell placed on a signal that crosses from logic running at one supply voltage to logic running at another. Without it, a low-voltage signal may not switch a higher-voltage receiver cleanly.

Expert. Required where a low-voltage domain drives a high-voltage one; otherwise the receiver can suffer short-circuit current and leakage. Specified in the UPF, placed near the domain boundary and powered from one or both supplies, so the floorplan reserves room for them.

LFSR

Linear feedback shift register: a short chain of flip-flops whose input is the exclusive-OR of a few of its own bits. With the right feedback an nn-bit LFSR steps through all 2n−12^n - 1 nonzero values in a fixed, random-looking order.

All three levels

Beginner. A small circuit that churns out a long string of random-looking test patterns.

Novice. Linear feedback shift register: a short chain of flip-flops whose input is the exclusive-OR of a few of its own bits. With the right feedback an nn-bit LFSR steps through all 2n−12^n - 1 nonzero values in a fixed, random-looking order.

Expert. Described by a characteristic polynomial; a primitive polynomial gives maximum length. Used as the pattern generator in LBIST, as the expander in compression, and, with extra inputs, as a signature register. The all-zero state is stuck forever, so it must never be loaded.

Liberty (.lib)

The text file that describes a cell library to the design tools: each cell’s logic function, its area, the electrical load of each input, tables of how its delay depends on conditions, and how much power it uses.

All three levels

Beginner. The catalog that tells the tool each building block’s size, speed and power use.

Novice. The text file that describes a cell library to the design tools: each cell’s logic function, its area, the electrical load of each input, tables of how its delay depends on conditions, and how much power it uses.

Expert. One file per corner (a combination of manufacturing process, supply voltage and temperature), often one per threshold-voltage flavor too. It holds lookup tables for delay, output transition and internal energy, timing checks such as flip-flop setup and hold, and limits such as max_transition and max_capacitance. Older flows also kept wire-load models here.

Linear (triode) region

The region where VGSV_{\mathrm{GS}} is above threshold and VDSV_{\mathrm{DS}} is small (below VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}}). The channel stretches all the way from source to drain, and current rises with VDSV_{\mathrm{DS}} almost in proportion, like a resistor whose value the gate controls.

All three levels

Beginner. An on state where more push across the transistor gives more flow, like water down a steeper hose.

Novice. The region where VGSV_{\mathrm{GS}} is above threshold and VDSV_{\mathrm{DS}} is small (below VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}}). The channel stretches all the way from source to drain, and current rises with VDSV_{\mathrm{DS}} almost in proportion, like a resistor whose value the gate controls.

Expert. VGS>VtV_{\mathrm{GS}} > V_{\mathrm{t}} and VDS<VDSATV_{\mathrm{DS}} < V_{\mathrm{DSAT}}. ID=β(VGS−Vt−VDS/2)VDSI_{\mathrm{D}} = \beta(V_{\mathrm{GS}} - V_{\mathrm{t}} - V_{\mathrm{DS}}/2)V_{\mathrm{DS}} in the square law. A switching transistor ends each transition here, and its resistance at small VDSV_{\mathrm{DS}} sets how fast it finishes pulling a node to the rail.

Linear-drive pluggable optics (LPO)

An optical module without a DSP. Its driver and receiver amplifiers pass the signal through linearly, and the switch chip’s own SerDes equalizes the whole path, electrical and optical, end to end. Saves power and delay, but works only with strong host SerDes and is harder to mix and match.

All three levels

Beginner. A plug-in light module with the signal-cleaning chip removed, so it uses less power. The switch chip does that work instead.

Novice. An optical module without a DSP. Its driver and receiver amplifiers pass the signal through linearly, and the switch chip’s own SerDes equalizes the whole path, electrical and optical, end to end. Saves power and delay, but works only with strong host SerDes and is harder to mix and match.

Expert. Linear driver and linear TIA only; no CDR. The link budget spans host SerDes → host channel → module → fiber → module → host channel → host SerDes, so interoperability must be defined end to end (LPO MSA, OIF linear interfaces). Half-retimed (LRO) variants keep a DSP on transmit only.

Lint

A program that reads design code and flags likely mistakes without running it: accidental latches, numbers of mismatched size, unused signals, missing cases. Verilator and Verible are open-source linters.

All three levels

Beginner. An automatic proofreader for design code that flags suspicious patterns before anything is tested.

Novice. A program that reads design code and flags likely mistakes without running it: accidental latches, numbers of mismatched size, unused signals, missing cases. Verilator and Verible are open-source linters.

Expert. Fast enough to run on every code change. Teams keep a fixed set of rules and a reviewed list of exceptions (waivers). A growing waiver file, or a rule switched off for the whole project, is a warning sign, because it also hides every future case of the same problem.

Lithography hotspot

A spot in a chip layout that passes the design rules but is likely to print badly, for example two lines that may merge or a line that may pinch off.

All three levels

Beginner. A spot in a chip’s design whose tiny shapes are likely to print badly when the chip is made, even though they follow the rules.

Novice. A spot in a chip layout that passes the design rules but is likely to print badly, for example two lines that may merge or a line that may pinch off.

Expert. Found traditionally by pattern matching or slow lithography simulation. ML classifiers trained on small layout snippets flag candidates faster, trading false alarms against misses.

Little’s law

L=λWL = \lambda W: the average number of items in a system equals the rate at which they arrive times the average time each one stays. Ten requests per microsecond, each taking half a microsecond, means five in progress on average.

All three levels

Beginner. A simple rule: how many things are waiting equals how fast they arrive times how long each one stays.

Novice. L=λWL = \lambda W: the average number of items in a system equals the rate at which they arrive times the average time each one stays. Ten requests per microsecond, each taking half a microsecond, means five in progress on average.

Expert. Sizes queues and outstanding-request limits: to sustain bandwidth BB when each request takes latency WW, a requester must keep B⋅WB \cdot W bytes in flight. That sets the number of miss-handling registers, DMA tags and interconnect buffers. It holds whatever the arrival pattern.

LLM agent (agentic flow)

A setup in which a language model works in a loop: it plans a step, runs a tool (a simulator, a checker, a design tool), reads the output and decides what to do next.

All three levels

Beginner. An AI model that does more than answer: it runs tools, looks at the results and decides what to do next, step by step.

Novice. A setup in which a language model works in a loop: it plans a step, runs a tool (a simulator, a checker, a design tool), reads the output and decides what to do next.

Expert. Tool feedback grounds the model in real results, and invalid tool calls fail visibly. The risks are mistakes that compound over long loops, cost per task, and success criteria that only check that a script ran, leaving the design goal unchecked.

Load capacitance

The capacitance a gate’s output must charge or discharge: the input (gate) capacitance of every transistor it drives, plus the wire, plus its own output junctions. Measured in femtofarads (fF).

All three levels

Beginner. How much electric charge a gate’s output wire must fill up with before the next gates see the new value. Like the size of a bucket.

Novice. The capacitance a gate’s output must charge or discharge: the input (gate) capacitance of every transistor it drives, plus the wire, plus its own output junctions. Measured in femtofarads (fF).

Expert. CL=Cwire+∑Cin(fanout)+CselfC_{\mathrm{L}} = C_{\mathrm{wire}} + \sum C_{\mathrm{in}}(\text{fanout}) + C_{\mathrm{self}}. Delay and switching energy are both linear in it, which is why wire length and fanout appear in every timing and power estimate.

Load-use hazard

A load followed directly by an instruction that uses the loaded value. The data only arrive at the end of the memory stage, one cycle too late for the next instruction’s execute stage, so the pipeline stalls for one cycle even with forwarding.

All three levels

Beginner. When a step needs a number that is still on its way from memory. Even with the shortcut, it has to wait a moment.

Novice. A load followed directly by an instruction that uses the loaded value. The data only arrive at the end of the memory stage, one cycle too late for the next instruction’s execute stage, so the pipeline stalls for one cycle even with forwarding.

Expert. Load-to-use latency is two cycles in the five-stage pipeline (one for ALU-to-use), so the hazard unit inserts one bubble at distance 1. Compilers schedule an independent instruction into the slot; deeper pipelines and caches lengthen the distance.

Locality (temporal and spatial)

Temporal locality: data used recently is likely to be used again soon. Spatial locality: data near recently used data is likely to be used soon. Loops and arrays have plenty of both.

All three levels

Beginner. Programs tend to reuse what they just used, and to use things that sit next to it. Caches work because of this habit.

Novice. Temporal locality: data used recently is likely to be used again soon. Spatial locality: data near recently used data is likely to be used soon. Loops and arrays have plenty of both.

Expert. Quantified by reuse distance (distinct lines touched between two uses of a line) and stride. A fully associative LRU cache of C lines hits exactly the accesses with reuse distance below C.

Lockstep

Running the same software on two identical processor cores at the same time and comparing their outputs every cycle. Any mismatch means one of them has a fault.

All three levels

Beginner. Two identical processors run the same program at once and keep comparing answers.

Novice. Running the same software on two identical processor cores at the same time and comparing their outputs every cycle. Any mismatch means one of them has a fault.

Expert. Dual-core lockstep (DCLS) runs the second core a few cycles behind the first, so a common-cause disturbance hits them at different points and produces different errors that the comparator catches. Standard for ASIL D control functions.

Lockup latch

A latch (a simple memory cell that lets data through while its clock is at one level and holds it while the clock is at the other) placed between two scan flip-flops whose clock signals arrive at different times. It holds the bit back so it can’t skip a place in the chain.

All three levels

Beginner. A small holding cell placed where a test chain crosses between two parts of the chip that keep slightly different time, so values don’t race ahead.

Novice. A latch (a simple memory cell that lets data through while its clock is at one level and holds it while the clock is at the other) placed between two scan flip-flops whose clock signals arrive at different times. It holds the bit back so it can’t skip a place in the chain.

Expert. Clocked on the opposite phase of the launching flop, so the shifted bit waits about half a cycle and the receiving flop’s capture edge, even if it arrives late, sees the old value. Lockups also act as break points that scan reordering can’t move flops across, so too many of them lengthen scan wiring and add congestion.

Logic block (CLB, LAB, slice, ALM)

The repeated unit of an FPGA’s general-purpose logic: several LUTs, flip-flops, carry logic and local wiring. AMD calls it a CLB made of slices, Altera a LAB made of ALMs, Lattice a logic tile or PFU, Microchip a logic cluster.

All three levels

Beginner. A small group of lookup tables and one-bit memories that sits in the FPGA’s grid.

Novice. The repeated unit of an FPGA’s general-purpose logic: several LUTs, flip-flops, carry logic and local wiring. AMD calls it a CLB made of slices, Altera a LAB made of ALMs, Lattice a logic tile or PFU, Microchip a logic cluster.

Expert. A cluster of N basic logic elements (K-LUT + flip-flop) with a local crossbar and shared control signals. Cluster size trades fast local connections against crossbar area; inputs scale near K/2·(N+1).

Logic depth

The number of gates a signal passes through, one after another, on the longest route from input to output. It is a first rough estimate of delay.

All three levels

Beginner. How many building blocks a signal passes through, one after another, from input to output.

Novice. The number of gates a signal passes through, one after another, on the longest route from input to output. It is a first rough estimate of delay.

Expert. Levels of logic on a path, counted in AIG nodes or mapped cells. It tracks delay only loosely, since real delay depends on cell type, load, input transition and wires, but it is what balancing and depth-oriented mapping minimize.

Logic gate

A circuit that computes a Boolean function: inputs and output are each 0 (low voltage) or 1 (high voltage). NOT, NAND and NOR are the basic gates in CMOS; AND and OR are built from them.

All three levels

Beginner. A tiny circuit that takes one or more on/off signals and produces an on/off answer by a fixed rule.

Novice. A circuit that computes a Boolean function: inputs and output are each 0 (low voltage) or 1 (high voltage). NOT, NAND and NOR are the basic gates in CMOS; AND and OR are built from them.

Expert. A combinational circuit implementing a Boolean function of its inputs. In static CMOS each single-stage gate is a monotone-decreasing (inverting) function; non-inverting functions need a second stage.

Logic locking

Adding extra gates controlled by a secret key, so the chip only works correctly once the right key is loaded after manufacture.

All three levels

Beginner. Building a secret key into a chip, like a password. Without it, a stolen copy gives wrong answers.

Novice. Adding extra gates controlled by a secret key, so the chip only works correctly once the right key is loaded after manufacture.

Expert. Aimed at overproduction and design theft by an untrusted factory. The simplest form inserts XOR/XNOR gates on internal wires, each fed by a key bit. The SAT attack recovers keys by using a working chip as an oracle; later schemes trade resistance to it against other attacks and overhead.

Logic synthesis

The step that turns RTL code into a netlist: a list of logic gates and flip-flops from a library of ready-made cells, plus the wires that connect them. Yosys is an open-source synthesis tool.

All three levels

Beginner. The step where a tool turns the design code into a list of real circuit building blocks.

Novice. The step that turns RTL code into a netlist: a list of logic gates and flip-flops from a library of ready-made cells, plus the wires that connect them. Yosys is an open-source synthesis tool.

Expert. Reads the RTL, builds generic logic from it, optimizes that logic and maps it onto library cells under timing constraints. It keeps the RTL’s cycle-by-cycle behavior but may restructure everything within a cycle; equivalence checking confirms the function did not change.

Logical effort

The ratio of a gate’s input capacitance to that of an inverter that can deliver the same output current. An inverter scores 1; a NAND2 scores 4/3; a NOR2 scores 5/3.

All three levels

Beginner. A score for how much harder a gate has to work than a plain inverter to do the same job.

Novice. The ratio of a gate’s input capacitance to that of an inverter that can deliver the same output current. An inverter scores 1; a NAND2 scores 4/3; a NOR2 scores 5/3.

Expert. gg in the delay model d=gh+pd = gh + p (Sutherland, Sproull and Harris). It captures the cost of the series stacks a function requires. Path effort and the best stage count follow from it, which makes it a fast way to compare gate topologies.

Logical equivalence checking (LEC)

A mathematical proof, rather than a test, that two versions of a design (for example the RTL code and the synthesized netlist) produce the same outputs for every possible input.

All three levels

Beginner. A math-based check that the finished parts list gives the same answers as the original design, for every possible input.

Novice. A mathematical proof, rather than a test, that two versions of a design (for example the RTL code and the synthesized netlist) produce the same outputs for every possible input.

Expert. LEC first pairs up state points (flip-flops, ports, black boxes) between the reference and the implementation, then proves the logic feeding each pair identical, using AIGs, random simulation, SAT solvers and BDDs. Retiming, state re-encoding and merged or deleted flip-flops break the pairing and need guidance from the synthesis tool or sequential checking.

Loss scaling

A trick for FP16 training: multiply the loss by a constant (say 8 or 1,024) before backpropagation so every gradient is scaled up by the same amount and stays above FP16’s smallest value. The gradients are divided back down before the weight update.

All three levels

Beginner. Multiplying numbers up before a calculation so that tiny ones don’t round to zero, then dividing back afterward.

Novice. A trick for FP16 training: multiply the loss by a constant (say 8 or 1,024) before backpropagation so every gradient is scaled up by the same amount and stays above FP16’s smallest value. The gradients are divided back down before the weight update.

Expert. Shifts gradient distributions into FP16’s representable window. Dynamic loss scaling raises the scale until an overflow (Inf/NaN) appears, then skips the step and backs off. BF16 rarely needs it; FP8 uses per-tensor scale factors instead, because skipping steps on overflow would happen too often.

LUT (lookup table)

The basic logic element of most FPGAs: a K-input truth table stored in 2K2^K configuration bits and read through a multiplexer whose select lines are the inputs. It can hold any function of its K inputs.

All three levels

Beginner. A tiny table of answers inside an FPGA. The inputs pick a row, and the table gives back the answer stored there.

Novice. The basic logic element of most FPGAs: a K-input truth table stored in 2K2^K configuration bits and read through a multiplexer whose select lines are the inputs. It can hold any function of its K inputs.

Expert. A 2K2^K-bit SRAM mask read by a K-level mux tree, so delay is fixed by K, not by the function. Modern parts use K = 4 to 6, often fracturable into two smaller LUTs with shared inputs; some can also act as small RAMs or shift registers.

MAC (multiply-accumulate)

The operation a←a+b⋅ca \leftarrow a + b \cdot c: multiply two numbers and add the result to a running total. A matrix multiplication is made of many of them, so AI chips quote their peak speed in MACs per clock cycle.

All three levels

Beginner. One multiplication added onto a running total. AI needs billions of them every second.

Novice. The operation a←a+b⋅ca \leftarrow a + b \cdot c: multiply two numbers and add the result to a running total. A matrix multiplication is made of many of them, so AI chips quote their peak speed in MACs per clock cycle.

Expert. The basic datapath cell of a matrix engine. Its input precision (INT8, FP8, BF16) sets area and energy per operation, and the width of its accumulator sets how many products can be summed before rounding or overflow hurts accuracy.

Machine learning (ML)

Software whose behavior comes from fitting a model to examples rather than from hand-written rules. Shown thousands of past chip layouts and how each one turned out, it learns to guess how a new layout will turn out.

All three levels

Beginner. Software that learns patterns from many examples, instead of following rules a person wrote down, and then uses those patterns to make guesses about new cases.

Novice. Software whose behavior comes from fitting a model to examples rather than from hand-written rules. Shown thousands of past chip layouts and how each one turned out, it learns to guess how a new layout will turn out.

Expert. Three styles appear in chip design: supervised prediction (learn to map early data to a later result), black-box optimization (decide which tool settings to try next), and reinforcement learning (learn a sequence of decisions from a score). The recurring problems are scarce labeled data, new designs that differ from the training set, and the hours of tool runtime each label costs.

Macro (hard macro)

A large pre-designed block whose layout is already finished, such as a memory, a clock generator or an analog circuit. The floorplan can only choose where it goes and which way it faces. Layout tools see only its outline, its connection points and the areas it blocks.

All three levels

Beginner. A big, ready-made block, such as a block of memory. You can move it and turn it, but you can’t change what’s inside.

Novice. A large pre-designed block whose layout is already finished, such as a memory, a clock generator or an analog circuit. The floorplan can only choose where it goes and which way it faces. Layout tools see only its outline, its connection points and the areas it blocks.

Expert. Placed before the standard cells and usually locked (FIXED). Orientation is normally limited to mirroring and 180° rotation, because the direction of the transistor gates inside can’t change. Its pins usually sit along one or two edges, and its internal metal blocks some routing layers above it. Described to the tools by a LEF abstract.

Majority / minority carrier

In n-type silicon electrons are the majority carriers and holes the minority; in p-type it’s the reverse. Their product stays fixed: np=ni2np = n_{\mathrm{i}}^2.

All three levels

Beginner. In a piece of silicon, whichever moving charge there is more of, electrons or holes. The rarer kind is the minority.

Novice. In n-type silicon electrons are the majority carriers and holes the minority; in p-type it’s the reverse. Their product stays fixed: np=ni2np = n_{\mathrm{i}}^2.

Expert. Minority carriers dominate diode current: forward bias injects them, raising the minority density at the depletion edge by eqV/kTe^{qV/kT}, and they diffuse and recombine over a diffusion length. An inverted MOS surface is a layer where the minority carrier has become the majority.

Mantissa

The fraction bits of a floating-point number (also called the trailing significand). With mm mantissa bits, each power-of-two interval is split into 2m2^m equal steps, so mm sets the precision.

All three levels

Beginner. The part of a stored number that holds its digits. More bits here means a more exact number.

Novice. The fraction bits of a floating-point number (also called the trailing significand). With mm mantissa bits, each power-of-two interval is split into 2m2^m equal steps, so mm sets the precision.

Expert. The stored fraction field; the significand is 1.M1.M for normal numbers and 0.M0.M for subnormals, so precision p=m+1p = m + 1. Relative spacing is 2−m2^{-m} within a binade; the significand multiplier’s area grows roughly as p2p^2.

Manufacturing defect

A physical imperfection introduced while the chip is made: two wires shorted together, a wire broken open, missing or extra material, or a transistor too far off its intended size to work.

All three levels

Beginner. A flaw made while building a chip, like a speck of dust joining two wires or a wire that comes out broken.

Novice. A physical imperfection introduced while the chip is made: two wires shorted together, a wire broken open, missing or extra material, or a transistor too far off its intended size to work.

Expert. What test is really trying to catch. Fault models are logical stand-ins for defects, and coverage of a fault model only correlates with defect coverage. That gap is why teams add transition, bridging and cell-aware models on top of stuck-at.

Mapping

The full plan for running a model on a particular chip: how each layer is split into pieces, which cores get which pieces, where data is stored, and when it moves. A compiler produces it.

All three levels

Beginner. The plan for which worker does which part of the job, and where each number is kept.

Novice. The full plan for running a model on a particular chip: how each layer is split into pieces, which cores get which pieces, where data is stored, and when it moves. A compiler produces it.

Expert. A point in the mapspace: tiling factors, loop order, spatial versus temporal assignment of each dimension, buffer allocation and placement. Its quality varies by large factors in energy and throughput for the same hardware and workload.

March test

A memory test made of steps (march elements). Each step visits every address in rising or falling order and does the same reads and writes at each. Test time grows in proportion to memory size.

All three levels

Beginner. A memory test that walks through every spot in order, writing and reading values in a set pattern.

Novice. A memory test made of steps (march elements). Each step visits every address in rising or falling order and does the same reads and writes at each. Test time grows in proportion to memory size.

Expert. Written in notation like ⇑(r0,w1): ascending order, read expecting 0, then write 1. Different algorithms target different memory fault classes (stuck cells, slow transitions, coupling between cells, address decoder faults). March C− costs 10 operations per cell.

Mask data preparation (MDP)

The steps between tapeout and mask writing: chip finishing (adding a seal ring and fill), placing the chip on the reticle with test patterns and alignment marks, optical proximity correction, and fracturing every shape into the rectangles and trapezoids a mask writer can draw.

All three levels

Beginner. The factory’s work of turning the chip file into instructions for the machines that make the stencils.

Novice. The steps between tapeout and mask writing: chip finishing (adding a seal ring and fill), placing the chip on the reticle with test patterns and alignment marks, optical proximity correction, and fracturing every shape into the rectangles and trapezoids a mask writer can draw.

Expert. Input is GDSII or OASIS; output is in the mask writer’s own format. Data volume explodes along the way, especially after OPC, which is why compact formats and processing that keeps the hierarchy matter.

Mask set

The photomasks, plates that each carry the pattern of one layer, that the factory uses to print a design onto wafers. A design needs at least one per patterned layer, and making the set is a large one-time cost that grows with each newer process.

All three levels

Beginner. The set of stencils used to print a chip’s patterns onto silicon. A new design needs a new set.

Novice. The photomasks, plates that each carry the pattern of one layer, that the factory uses to print a design onto wafers. A design needs at least one per patterned layer, and making the set is a large one-time cost that grows with each newer process.

Expert. Mask count rises with the number of layers and with multiple patterning (printing one layer in several exposures), so masks are a large part of the one-time cost at advanced processes. A fix confined to the upper metal wiring layers replaces only those masks, which is why designs scatter unused spare cells that a later wiring change can connect.

Matrix multiplication (GEMM)

C=ABC = AB, where AA is M×KM \times K and BB is K×NK \times N. Every one of the M×NM \times N outputs is a sum of KK products, so it takes MNKMNK multiply-adds. GEMM (general matrix multiply) is the library name for it.

All three levels

Beginner. Multiplying two grids of numbers to make a third. Each box in the answer is a long row of multiplications added up.

Novice. C=ABC = AB, where AA is M×KM \times K and BB is K×NK \times N. Every one of the M×NM \times N outputs is a sum of KK products, so it takes MNKMNK multiply-adds. GEMM (general matrix multiply) is the library name for it.

Expert. 2MNK2MNK FLOPs over MK+KN+MNMK + KN + MN values, so its intensity grows with the smallest dimension. With M=1M = 1 it degenerates to a matrix-vector product (GEMV), which is always memory-bound.

Maze routing

Finding a path through a grid with obstacles by spreading out from the start one step at a time, labeling each point with its distance, until the target is reached, then tracing back along falling labels. It always finds a path if one exists, but may visit a great many grid points.

All three levels

Beginner. Finding a path through a grid with obstacles by spreading outward step by step from the start until you reach the goal, like water flooding a maze.

Novice. Finding a path through a grid with obstacles by spreading out from the start one step at a time, labeling each point with its distance, until the target is reached, then tracing back along falling labels. It always finds a path if one exists, but may visit a great many grid points.

Expert. Lee’s breadth-first wave expansion, generalized to Dijkstra for weighted costs and to A* with a distance-to-target estimate. It is the core search in both global (GCell graph) and detailed (track graph) routers. Variants trade optimality for speed: Hadlock’s detour numbers, Soukup’s line-then-wave search, bidirectional search, and limiting the search to a box around the pins.

MCMM (multi-mode multi-corner)

Checking timing in every scenario. A scenario combines a mode (a way the chip is used, such as normal operation or test) with a corner (a combination of manufacturing extreme, voltage and temperature). Timing must pass in all of them.

All three levels

Beginner. Checking the chip in every way it will be used, in every place it will be.

Novice. Checking timing in every scenario. A scenario combines a mode (a way the chip is used, such as normal operation or test) with a corner (a combination of manufacturing extreme, voltage and temperature). Timing must pass in all of them.

Expert. Each scenario pairs a mode (its own SDC constraints) with a PVT corner (Liberty) and an RC corner (SPEF). The raw count is modes × PVT corners × RC corners; teams prune scenarios that never set the worst slack anywhere. Optimization and ECO tools load many scenarios at once so a fix in one doesn’t break another.

MedRadio / MICS

The U.S. radio rules for medical devices (Medical Device Radiocommunication Service). Its core 402–405 MHz band began in 1999 as the Medical Implant Communication Service (MICS), with a very low power limit.

All three levels

Beginner. Radio channels saved just for medical implants to talk to equipment outside the body.

Novice. The U.S. radio rules for medical devices (Medical Device Radiocommunication Service). Its core 402–405 MHz band began in 1999 as the Medical Implant Communication Service (MICS), with a very low power limit.

Expert. An ultra-low-power, non-voice service with tight EIRP limits and frequency-monitoring rules. Implant radios must also coexist with inductive links and survive tissue attenuation within a µA-level average current budget.

Memory bandwidth

The rate at which data moves between a chip and its memory, in bytes per second. AI chips with stacked memory reach terabytes per second.

All three levels

Beginner. How fast numbers can be brought from memory to the part of the chip that does math. Think of it as the width of a pipe.

Novice. The rate at which data moves between a chip and its memory, in bytes per second. AI chips with stacked memory reach terabytes per second.

Expert. Quoted as a theoretical peak (bus width × data rate). Sustained bandwidth is lower, and the roofline paper measures it with microbenchmarks rather than trusting the pin rate.

Memory BIST (MBIST)

Memory built-in self-test: a small test controller placed beside each on-chip memory. It writes and reads set patterns through every address at full speed and reports pass or fail, often with the failing addresses so the memory can be repaired.

All three levels

Beginner. A tiny built-in tester next to each memory block. It writes values into every spot and reads them back to check.

Novice. Memory built-in self-test: a small test controller placed beside each on-chip memory. It writes and reads set patterns through every address at full speed and reports pass or fail, often with the failing addresses so the memory can be repaired.

Expert. Generated per memory instance: a controller, address and data generators and a comparator. Runs March-style algorithms, feeds repair with spare rows and columns, and is reached through JTAG or IJTAG.

Memory coalescing

When the 32 threads of a warp read neighboring addresses, the hardware combines their requests into a few large memory transactions. Scattered addresses need many transactions and waste bandwidth.

All three levels

Beginner. Fetching data for a whole team in one trip instead of one trip per worker.

Novice. When the 32 threads of a warp read neighboring addresses, the hardware combines their requests into a few large memory transactions. Scattered addresses need many transactions and waste bandwidth.

Expert. Global loads are served in 32-byte sectors; a warp touching 32 consecutive 4-byte words needs 4 sectors, while 32 scattered words can need 32, cutting effective bandwidth up to 8×.

Memory consistency model

The contract between hardware and software about the order in which memory operations from different cores can be observed. Sequential consistency is the strictest; x86, Arm and RISC-V allow some reordering and provide fences to stop it.

All three levels

Beginner. The rules that say in what order one core’s changes to memory can appear to other cores.

Novice. The contract between hardware and software about the order in which memory operations from different cores can be observed. Sequential consistency is the strictest; x86, Arm and RISC-V allow some reordering and provide fences to stop it.

Expert. Coherence orders accesses to one location; consistency orders accesses to different locations. x86 is TSO (a load may pass an earlier store to another address); Armv8 and RISC-V RVWMO are weaker, with barriers and acquire/release accesses.

Memory hierarchy

The layers of memory in a computer, from fastest and smallest to slowest and largest: registers inside the processor, then on-chip caches or SRAM (fast memory built from transistors on the chip itself), then DRAM main memory on separate chips. Each level is bigger, slower and cheaper per bit than the one above. It works because programs tend to reuse the same and nearby data.

All three levels

Beginner. Layers of memory from tiny and very fast, right next to the processor, to huge and slow, far away.

Novice. The layers of memory in a computer, from fastest and smallest to slowest and largest: registers inside the processor, then on-chip caches or SRAM (fast memory built from transistors on the chip itself), then DRAM main memory on separate chips. Each level is bigger, slower and cheaper per bit than the one above. It works because programs tend to reuse the same and nearby data.

Expert. Each level is sized by capacity, access time, bandwidth and energy per access, plus rules for which copies are kept (inclusion, replacement) and how copies stay consistent (coherence). Each on-chip level becomes blocks of SRAM in the floorplan, so these choices set much of the die area.

Memory semantics

A fabric property: accelerators issue ordinary loads, stores and atomic operations to addresses that live on another accelerator. Hardware turns them into read, write and atomic requests across the link.

All three levels

Beginner. One chip can read or write another chip’s memory directly, as if it were its own. It doesn’t have to pack data into messages.

Novice. A fabric property: accelerators issue ordinary loads, stores and atomic operations to addresses that live on another accelerator. Hardware turns them into read, write and atomic requests across the link.

Expert. Load/store/atomic transactions with a defined ordering model, carried as small requests (64–256 B) rather than NIC-driven messages. Enables one-sided, PGAS-style communication and fine-grained synchronization, but requires many outstanding requests to fill the link and careful handling of ordering and errors.

Memory wall

The widening gap between how fast processors compute and how fast memory can deliver data. Wulf and McKee named it in 1995; in AI chips it shows up as math units idling while weights stream in from DRAM.

All three levels

Beginner. The problem that a chip can do math much faster than it can fetch the numbers to do math on, so its math parts sit waiting.

Novice. The widening gap between how fast processors compute and how fast memory can deliver data. Wulf and McKee named it in 1995; in AI chips it shows up as math units idling while weights stream in from DRAM.

Expert. Compute throughput has grown faster than DRAM bandwidth, capacity and latency for decades (peak FLOPS about 3.0× per two years against 1.6× for DRAM bandwidth), so the ridge point of each new chip moves right and more kernels fall on the bandwidth slope of the roofline.

Memory-bound

A workload whose run time is set by how many bytes it moves, because it does too little work per byte. More bandwidth or more reuse of each byte would speed it up; more math hardware would not.

All three levels

Beginner. A job is waiting on memory when the math parts of the chip sit idle, because the numbers they need can’t arrive fast enough.

Novice. A workload whose run time is set by how many bytes it moves, because it does too little work per byte. More bandwidth or more reuse of each byte would speed it up; more math hardware would not.

Expert. Arithmetic intensity below the ridge point, so attainable = bandwidth × intensity. Fixes raise the intensity (batching, fusion, tiling, fewer bits per value) or the bandwidth.

MESI protocol

A cache-coherence protocol with four states per line: Modified (only copy, changed), Exclusive (only copy, clean), Shared (one of several clean copies) and Invalid. A core must hold a line in M or E to write it, which invalidates other copies.

All three levels

Beginner. The rules cores use to keep their copies of shared data in agreement. Each copy is marked changed, only copy, shared, or empty.

Novice. A cache-coherence protocol with four states per line: Modified (only copy, changed), Exclusive (only copy, clean), Shared (one of several clean copies) and Invalid. A core must hold a line in M or E to write it, which invalidates other copies.

Expert. Write-invalidate protocol enforcing single-writer/multiple-reader per line. E lets private data be written without a bus transaction; MOESI adds Owned so a dirty line can be shared without writing it back. Real implementations add many transient states.

Metal fill (density fill)

Small unconnected metal shapes added to empty areas of the layout. The factory polishes each layer flat, and areas with too little metal polish unevenly, so every region must have roughly the right amount of metal.

All three levels

Beginner. Extra metal added to empty spots, so the factory can polish each layer flat and even.

Novice. Small unconnected metal shapes added to empty areas of the layout. The factory polishes each layer flat, and areas with too little metal polish unevenly, so every region must have roughly the right amount of metal.

Expert. Shapes inserted so every density window meets minimum and maximum density rules for chemical-mechanical polishing (CMP). Fill adds coupling capacitance, so signoff extraction and timing must include it or model it; timing-aware fill keeps it away from critical nets. Fill also changes DRC results, so physical verification runs after fill.

Metal pitch

The center-to-center distance of the tightest wires, on the first few metal layers. Cell height is usually given as a number of these ‘tracks’.

All three levels

Beginner. The spacing between neighboring wires on the lowest wiring layers of a chip.

Novice. The center-to-center distance of the tightest wires, on the first few metal layers. Cell height is usually given as a number of these ‘tracks’.

Expert. Minimum line plus space on the tightest metal layers (M1–M3 in ASAP7, 36 nm with EUV). Cell height = tracks × metal pitch, and the ratio to fin pitch (the ‘gear ratio’) must come out to whole or half numbers. The ‘M’ in G-M-T.

Metal stack

The layers of metal wiring built on top of the transistors, numbered from M1 at the bottom to M10 or more at the top, separated by insulating glass and joined by vias. Lower layers are thin and tightly packed for short local connections; upper layers are thicker and wider, so they resist current less and suit long routes, clocks and power.

All three levels

Beginner. The layers of copper wiring stacked above the chip’s tiny switches, like floors in a parking garage. Low floors have thin, tightly packed wires; high floors have thick, roomy ones.

Novice. The layers of metal wiring built on top of the transistors, numbered from M1 at the bottom to M10 or more at the top, separated by insulating glass and joined by vias. Lower layers are thin and tightly packed for short local connections; upper layers are thicker and wider, so they resist current less and suit long routes, clocks and power.

Expert. Pitch (the center-to-center spacing of wires) and thickness grow up the stack, so resistance per µm falls steeply from the lower layers to the upper ones. Promoting critical nets to upper layers, giving them wider wires and resistance-aware layer assignment all exploit that. At advanced nodes the lowest layers are printed with several masks and must run strictly in one direction, and power may move below the transistors or to the back of the wafer.

Metastability

An undecided state between 0 and 1 that a flip-flop can enter when its input changes at almost the same instant as the clock edge. It settles to 0 or 1 eventually, usually within a fraction of a nanosecond, but with no guaranteed limit.

All three levels

Beginner. A moment when a memory cell can’t decide between 0 and 1 because its input changed just as the clock ticked.

Novice. An undecided state between 0 and 1 that a flip-flop can enter when its input changes at almost the same instant as the clock edge. It settles to 0 or 1 eventually, usually within a fraction of a nanosecond, but with no guaranteed limit.

Expert. The time to settle is exponentially distributed with a time constant τ\tau (tau): the chance a flop is still undecided after waiting SS is e−S/τe^{-S/\tau}. Each extra τ\tau of waiting divides the failure rate by ee (about 2.7), and synchronizer design rests on that.

Metrology

Measurement inside the fab: the thickness of each deposited film, the width of the smallest printed shapes (the critical dimension), and how well each layer lines up with the one below (overlay). Inspections for defects run alongside. The results keep each step within its limits.

All three levels

Beginner. Measuring the wafer during manufacturing to make sure each layer came out the right size and in the right place.

Novice. Measurement inside the fab: the thickness of each deposited film, the width of the smallest printed shapes (the critical dimension), and how well each layer lines up with the one below (overlay). Inspections for defects run alongside. The results keep each step within its limits.

Expert. The data feed statistical process control (charts that flag a tool drifting outside its normal range) and run-to-run correction (adjusting a tool’s settings for the next batch from the last one’s measurements). Measuring every wafer would cost tool time and delay, so fabs sample, trading cost against how quickly they catch a problem.

MFU (model FLOPs utilization)

Observed tokens per second × the operations each token needs (about 6 per model parameter for training) ÷ the system’s peak operations per second.

All three levels

Beginner. A fair score for how well a training job uses its chips. Count only the math the model truly needs, and compare it with the chips’ top speed.

Novice. Observed tokens per second × the operations each token needs (about 6 per model parameter for training) ÷ the system’s peak operations per second.

Expert. Introduced with PaLM. Unlike hardware FLOPs utilization (HFU), it excludes recomputation (rematerialization) and other implementation-specific work, so it can’t be inflated by doing extra arithmetic. PaLM 540B reported 46.2% MFU and 57.8% HFU.

Microarchitecture specification

A detailed design document for one block: the steps data passes through, the control logic that sequences them, the storage the block needs, and what happens on each tick of the chip’s clock. Engineers write the block’s design code from it.

All three levels

Beginner. A detailed plan for one block of the chip, written by the engineers who will build it.

Novice. A detailed design document for one block: the steps data passes through, the control logic that sequences them, the storage the block needs, and what happens on each tick of the chip’s clock. Engineers write the block’s design code from it.

Expert. Written by the block owner and reviewed by architecture, verification and physical design. It covers pipeline stages, state machines, interfaces with cycle-by-cycle timing, and the register map: the numbered control and status registers software uses to drive the block. Ideally that map lives in one machine-readable file that generates the design code, the verification models and the software header files.

Microbatch

A part of a training batch processed in one forward and backward pass. Gradients from several microbatches are added up before the weights are updated, and pipelines keep several microbatches in flight at once.

All three levels

Beginner. A small slice of a batch of examples, sent through the model by itself.

Novice. A part of a training batch processed in one forward and backward pass. Gradients from several microbatches are added up before the weights are updated, and pipelines keep several microbatches in flight at once.

Expert. With global batch BB, data-parallel degree dd and microbatch size bb, each pipeline sees m=B/(db)m = B/(db) microbatches per step. mm sets the pipeline bubble, and bb sets kernel efficiency and activation memory.

Microbump

A small solder-capped copper bump at a pitch of tens of micrometers, used in 2.5D packages and memory stacks where ordinary flip-chip bumps are too coarse.

All three levels

Beginner. A very small dot of solder, a metal that melts easily, used to join chips to a silicon slab or to each other.

Novice. A small solder-capped copper bump at a pitch of tens of micrometers, used in 2.5D packages and memory stacks where ordinary flip-chip bumps are too coarse.

Expert. Typically 25–55 µm pitch in UCIe’s advanced-package range. Below roughly 10 µm, solder becomes impractical and bumpless copper-to-copper hybrid bonding takes over.

Microscaling (MX)

Block-scaled formats standardized by the Open Compute Project: every 32 consecutive values share one 8-bit power-of-two scale, and each value is stored in FP8, FP6, FP4 or INT8. MXFP4, for example, costs 4.25 bits per value.

All three levels

Beginner. Giving each small group of numbers its own shared multiplier, so one huge number only affects its own group.

Novice. Block-scaled formats standardized by the Open Compute Project: every 32 consecutive values share one 8-bit power-of-two scale, and each value is stored in FP8, FP6, FP4 or INT8. MXFP4, for example, costs 4.25 bits per value.

Expert. OCP MX v1.0: block size k=32k = 32, scale XX in E8M0 (2−1272^{-127} to 21272^{127}, one NaN code), elements PiP_i in E4M3/E5M2, E3M2/E2M3, E2M1 or INT8. X=2⌊log⁡2max⁡∣v∣⌋−emax⁡,elemX = 2^{\lfloor \log_2 \max|v| \rfloor - e_{\max,\mathrm{elem}}}; Pi=round(vi/X)P_i = \mathrm{round}(v_i/X) with saturation. A dot product computes XAXB∑PAPBX_A X_B \sum P_A P_B, with internal precision implementation-defined.

Middle of line (MOL)

The contacts and local connections that join transistor sources, drains and gates to the first metal layer. Once a few simple steps, it became its own module as transistors shrank.

All three levels

Beginner. The small set of steps that connect the finished switches to the wiring above them.

Novice. The contacts and local connections that join transistor sources, drains and gates to the first metal layer. Once a few simple steps, it became its own module as transistors shrank.

Expert. In ASAP7 it is three layers: gate and source/drain local interconnect (LIG, LISD) and V0 up to M1, all printed with EUV.

Minimum energy point

The supply voltage that minimizes energy per operation. Below it, each operation takes so long that the leakage energy collected during it outweighs the switching energy saved.

All three levels

Beginner. The power-supply level where a circuit does each job with the least energy. Go lower and the chip gets so slow that leaking wastes more than you save.

Novice. The supply voltage that minimizes energy per operation. Below it, each operation takes so long that the leakage energy collected during it outweighs the switching energy saved.

Expert. The minimum of E(V)=αCV2+VIleaktop(V)E(V) = \alpha C V^2 + V I_{\mathrm{leak}} t_{\mathrm{op}}(V). Switching energy falls as V2V^2, but topt_{\mathrm{op}} rises exponentially below VTV_{\mathrm{T}}, so the minimum usually sits near or just below threshold. It moves up when activity is low or leakage is high, and down when activity is high.

Misprediction penalty

The clock cycles lost when a branch prediction turns out wrong: the wrong-path work is discarded and the pipeline refills from the right instruction. On recent desktop and server cores it is roughly 10 to 25 cycles.

All three levels

Beginner. The time lost when the processor guesses an “if” wrong and has to throw away work and start again.

Novice. The clock cycles lost when a branch prediction turns out wrong: the wrong-path work is discarded and the pipeline refills from the right instruction. On recent desktop and server cores it is roughly 10 to 25 cycles.

Expert. Roughly the fetch-to-resolve depth, plus however long the branch waited for its operands. Added CPI ≈ branches per instruction × mispredict rate × penalty, and each lost cycle costs W issue slots on a W-wide core.

MISR

Multiple-input signature register: an LFSR that also mixes in many output bits every clock, folding a whole test’s worth of results into one short value (a signature) that is compared once at the end with the expected value.

All three levels

Beginner. A small circuit that boils a huge stream of test results down to one short fingerprint.

Novice. Multiple-input signature register: an LFSR that also mixes in many output bits every clock, folding a whole test’s worth of results into one short value (a signature) that is compared once at the end with the expected value.

Expert. Linear, so the final signature is the XOR of each input stream’s polynomial remainder. A single unknown (X) input makes the signature unknown, so MISR-based BIST and compaction need X-free responses or masking.

Mixed-precision training

A training recipe where matrix multiplications use 16-bit (or 8-bit) inputs, but results are added up in 32-bit and a 32-bit “master” copy of the weights receives the updates.

All three levels

Beginner. Training a model with small, fast number formats for most of the math and a bigger format where accuracy really matters.

Novice. A training recipe where matrix multiplications use 16-bit (or 8-bit) inputs, but results are added up in 32-bit and a 32-bit “master” copy of the weights receives the updates.

Expert. Low-precision GEMM inputs, FP32 accumulation, an FP32 master copy of weights and optimizer state, plus loss scaling for FP16 and per-tensor scaling for FP8. Precision is chosen per operation: reductions, normalizations and softmax typically stay in FP32.

Mixed-size placement

Placing a design that contains both large macros and millions of small standard cells. The macros may be placed first and locked, or placed together with the cells.

All three levels

Beginner. Putting a few big blocks and millions of tiny parts in place at the same time.

Novice. Placing a design that contains both large macros and millions of small standard cells. The macros may be placed first and locked, or placed together with the cells.

Expert. Sequential flows place the cells, cluster them, floorplan the macros inside a fixed outline, lock them and re-place the cells. Concurrent flows move macros and cells in one analytical placer and then legalize the macros, snapping them to legal, non-overlapping positions.

Mixture of experts (MoE)

A model whose feed-forward layers are replaced by many parallel ‘experts’. A small router sends each token to a few of them, so the model can have many more parameters without doing more work per token.

All three levels

Beginner. A model built from many specialist parts, where each word is handled by only a few of the specialists.

Novice. A model whose feed-forward layers are replaced by many parallel ‘experts’. A small router sends each token to a few of them, so the model can have many more parameters without doing more work per token.

Expert. Sparse layers with EE experts and top-kk routing; total parameters grow with EE while FLOPs per token track kk. Introduces load balancing, token dropping or capacity limits, and all-to-all traffic when experts live on different devices.

MLPerf

Benchmark suites from MLCommons, an industry and academic consortium, that time complete AI tasks (training a model to a target accuracy, or answering queries under latency limits). Submissions are reviewed by the other submitters before publication.

All three levels

Beginner. A shared set of AI speed tests, run under the same rules by many companies, so their results can be compared fairly.

Novice. Benchmark suites from MLCommons, an industry and academic consortium, that time complete AI tasks (training a model to a target accuracy, or answering queries under latency limits). Submissions are reviewed by the other submitters before publication.

Expert. Separate suites for training, datacenter and edge inference, mobile, tiny devices and more, each released in numbered versions. Each result is tied to a division, an availability category and a scenario; only results with the same benchmark and scenario from compatible versions should be compared.

Model checking

A formal method that treats the design as a machine moving from state to state, one step per clock cycle, and checks whether any state it can reach breaks a rule.

All three levels

Beginner. A proof that looks at every situation a design can ever get into, to see if a rule can be broken.

Novice. A formal method that treats the design as a machine moving from state to state, one step per clock cycle, and checks whether any state it can reach breaks a rule.

Expert. Symbolic engines (BDD reachability, SAT-based bounded model checking, k-induction, interpolation, IC3/PDR) represent huge sets of states as formulas instead of listing them one by one. Production tools run several engines at once and take whichever answers first.

MOS capacitor

A gate electrode separated from a doped silicon body by a thin insulator (originally silicon dioxide). Depending on the gate voltage the silicon surface is in accumulation, depletion or inversion. Add a source and a drain and it becomes a MOSFET.

All three levels

Beginner. A metal plate, a very thin layer of glass and a slab of silicon, stacked up. It is the heart of every transistor in a chip.

Novice. A gate electrode separated from a doped silicon body by a thin insulator (originally silicon dioxide). Depending on the gate voltage the silicon surface is in accumulation, depletion or inversion. Add a source and a drain and it becomes a MOSFET.

Expert. Analyzed with Vg=Vfb+ϕs+VoxV_{\mathrm{g}} = V_{\mathrm{fb}} + \phi_{\mathrm{s}} + V_{\mathrm{ox}} and Gauss’s law at the interface. Its C–V curve, measured slowly (quasi-static) or at high frequency, is how ToxT_{\mathrm{ox}}, VfbV_{\mathrm{fb}}, VTV_{\mathrm{T}}, doping and interface charge are extracted in practice.

MOSFET

Metal–oxide–semiconductor field-effect transistor. The gate sits on a thin insulator above the silicon. Its electric field (no current flows into the gate) controls whether a conducting channel connects the source to the drain.

All three levels

Beginner. The kind of transistor almost every chip is built from. Its name lists its layers: a metal gate, a thin layer of glass, and silicon underneath.

Novice. Metal–oxide–semiconductor field-effect transistor. The gate sits on a thin insulator above the silicon. Its electric field (no current flows into the gate) controls whether a conducting channel connects the source to the drain.

Expert. A four-terminal device (gate, source, drain, body). Enhancement-mode NMOS and PMOS are the building blocks of CMOS. Planar devices have given way to FinFETs and gate-all-around nanosheets at advanced nodes, but the switch model is the same.

MTBF (synchronizer)

Mean time between failures: the average time between errors. A well-designed synchronizer has an MTBF far longer than the product will ever be used, often longer than the age of the universe.

All three levels

Beginner. The average time between failures. For a good synchronizer it is far longer than the product’s life.

Novice. Mean time between failures: the average time between errors. A well-designed synchronizer has an MTBF far longer than the product will ever be used, often longer than the age of the universe.

Expert. For a synchronizer, MTBF=eS/τ/(TWFCFD)\mathrm{MTBF} = e^{S/\tau} / (T_{\mathrm{W}} F_{\mathrm{C}} F_{\mathrm{D}}): SS is the settling time allowed, τ\tau and TWT_{\mathrm{W}} are properties of the flip-flop, FCF_{\mathrm{C}} is the clock rate and FDF_{\mathrm{D}} the rate at which the incoming data changes. It grows exponentially with SS and falls in proportion to the clock rate, the data rate and the number of synchronizers.

Multi-level cell (MLC, TLC, QLC)

A flash cell programmed to one of 4 (MLC, 2 bits), 8 (TLC, 3 bits) or 16 (QLC, 4 bits) threshold-voltage levels instead of 2 (SLC). More bits per cell means more capacity, but narrower windows between levels and more errors.

All three levels

Beginner. A flash cell that stores more than one bit by using several in-between levels instead of just ‘full’ and ‘empty’.

Novice. A flash cell programmed to one of 4 (MLC, 2 bits), 8 (TLC, 3 bits) or 16 (QLC, 4 bits) threshold-voltage levels instead of 2 (SLC). More bits per cell means more capacity, but narrower windows between levels and more errors.

Expert. Each extra bit halves the voltage window per state, so programming needs more ISPP steps and verify passes, reads need more reference voltages, and retention, disturb and wear errors rise. Endurance drops with each step (SLC≫MLC≫TLC\mathrm{SLC} \gg \mathrm{MLC} \gg \mathrm{TLC}), and controllers rely on strong ECC and read-retry.

Multi-project wafer (MPW)

A shared manufacturing run, also called a shuttle, in which many designs are placed side by side on the same masks and wafers. Each customer pays for its area and gets a small number of chips back.

All three levels

Beginner. Many teams or students sharing one set of stencils and one batch of chips, like carpooling, so each pays only a share.

Novice. A shared manufacturing run, also called a shuttle, in which many designs are placed side by side on the same masks and wafers. Each customer pays for its area and gets a small number of chips back.

Expert. Trades cost for flexibility: fixed submission dates, area sold in set sizes, few parts, the provider’s rules and tools, and usually no volume production from the same masks. Ideal for prototypes, test chips, research and teaching.

multi-Vt

Building the same logic gates in versions whose transistors have different threshold voltages (the voltage at which they switch on). Low-threshold gates are fast but leak more current; high-threshold gates are slower but leak much less. Designers mix them gate by gate.

All three levels

Beginner. Using a mix of fast switches that leak a bit of power and slow ones that leak less, picking the right kind for each spot.

Novice. Building the same logic gates in versions whose transistors have different threshold voltages (the voltage at which they switch on). Low-threshold gates are fast but leak more current; high-threshold gates are slower but leak much less. Designers mix them gate by gate.

Expert. Synthesis and leakage recovery assign a VTV_{\mathrm{T}} flavor per cell under timing constraints: low-VTV_{\mathrm{T}} only on critical paths, high-VTV_{\mathrm{T}} everywhere with slack. The best mix shifts with corner and temperature, so leakage is signed off hot.

Multibit flip-flop

A library cell that holds two, four or more flip-flops side by side and shares one clock input and its internal clock circuitry among them, which saves clock power.

All three levels

Beginner. Several one-bit memories packed into one block so they can share wiring.

Novice. A library cell that holds two, four or more flip-flops side by side and shares one clock input and its internal clock circuitry among them, which saves clock power.

Expert. Banking single-bit flip-flops into multibit cells shares the internal clock buffers, lowering clock-pin capacitance and clock power and sometimes area. It happens in synthesis when the library has the cells and the bits are related, or at placement, where the tool knows which flip-flops sit close together.

Multicore processor

A processor chip with two or more cores, each able to run its own stream of instructions, sharing the chip’s caches, memory connections and power budget. Since the mid-2000s nearly every CPU is multicore.

All three levels

Beginner. A chip with several complete processors, called cores, side by side. Each core can run its own program at the same time as the others.

Novice. A processor chip with two or more cores, each able to run its own stream of instructions, sharing the chip’s caches, memory connections and power budget. Since the mid-2000s nearly every CPU is multicore.

Expert. Thread-level parallelism on one die: N cores with private L1/L2, a shared last-level cache and interconnect, kept coherent by hardware. It replaced frequency scaling once power, not transistors, became the binding limit.

Multicycle path

An SDC command (set_multicycle_path) that gives a path more than one clock cycle, used when the design only reads the result every few cycles.

All three levels

Beginner. A route that is allowed more than one beat of the chip’s clock to finish.

Novice. An SDC command (set_multicycle_path) that gives a path more than one clock cycle, used when the design only reads the result every few cycles.

Expert. Moves the setup check NN cycles later. In standard SDC the hold check moves with it, so a setup multiplier of NN is normally paired with a hold multiplier of N−1N - 1 to put the hold check back where it was.

Multigate transistor

Any MOSFET whose gate controls the channel from two or more sides: double-gate and tri-gate (fin) devices, and gate-all-around wires and sheets.

All three levels

Beginner. A transistor where the gate touches the channel from more than one side, so it has a firmer grip.

Novice. Any MOSFET whose gate controls the channel from two or more sides: double-gate and tri-gate (fin) devices, and gate-all-around wires and sheets.

Expert. Family name for double-gate, tri-gate (FinFET) and gate-all-around MOSFETs. More gated faces shorten the natural length, so the channel can stay undoped and the gate can be shorter for the same SS and DIBL.

Multiphase buck converter

A step-down converter made of several identical stages (“phases”) wired in parallel, each switching at a different moment. Splitting the current spreads the heat and smooths the output.

All three levels

Beginner. Several small power converters taking turns, sharing the job of feeding one chip.

Novice. A step-down converter made of several identical stages (“phases”) wired in parallel, each switching at a different moment. Splitting the current spreads the heat and smooths the output.

Expert. nn buck phases interleaved at 360∘/n360^\circ/n, sharing input and output capacitors. Ripple cancellation reduces capacitance, effective output inductance during transients drops by nn, and controllers shed phases at light load. Per-phase current is typically kept to tens of amps.

Multiple patterning

Printing one very dense layer with two or more masks (the stencils light is shone through), because its lines are too close together to print in one exposure. Shapes that are too close must go on different masks, so the router labels each shape with a “color” for its mask.

All three levels

Beginner. Printing one very dense wiring layer in two or more steps because its lines are too fine to print at once.

Novice. Printing one very dense layer with two or more masks (the stencils light is shone through), because its lines are too close together to print in one exposure. Shapes that are too close must go on different masks, so the router labels each shape with a “color” for its mask.

Expert. Litho-etch-litho-etch (LELE) needs a valid two-coloring of the conflict graph; an odd cycle cannot be two-colored, so a shape must be split or a wire rerouted. Spacer-based patterning (SADP, SAQP) makes regular rows of lines that favor on-track, one-direction wiring, with line ends set by separate cut masks.

Multiplexer (mux)

A selector circuit: a control input decides which of several data inputs is passed to the output. An if-else in RTL usually becomes a multiplexer.

All three levels

Beginner. A switch that picks one of several inputs and passes it along.

Novice. A selector circuit: a control input decides which of several data inputs is passed to the output. An if-else in RTL usually becomes a multiplexer.

Expert. The hardware form of a conditional. Synthesis builds mux trees from if and case statements; resource sharing and flip-flop enables add muxes, and clock gating removes the enable muxes again.

n-type

Silicon doped with donors (P, As, Sb). Electrons are the majority carriers and the Fermi level sits closer to the conduction band.

All three levels

Beginner. Silicon with extra free electrons, made by mixing in atoms that each have one electron to spare.

Novice. Silicon doped with donors (P, As, Sb). Electrons are the majority carriers and the Fermi level sits closer to the conduction band.

Expert. n≈Ndn \approx N_{\mathrm{d}} when Nd≫niN_{\mathrm{d}} \gg n_{\mathrm{i}}, p=ni2/Ndp = n_{\mathrm{i}}^2/N_{\mathrm{d}}, and EF−Ei=kTln⁡(Nd/ni)E_{\mathrm{F}} - E_{\mathrm{i}} = kT \ln(N_{\mathrm{d}}/n_{\mathrm{i}}). Above roughly 1019 cm−310^{19}\,\mathrm{cm^{-3}} the material becomes degenerate and Boltzmann statistics no longer apply.

NAND gate

NOT-AND. In CMOS, an n-input NAND has n NMOS in series to ground and n PMOS in parallel to the supply. The output falls only when every input is high.

All three levels

Beginner. A gate whose output is 0 only when all of its inputs are 1. It means “not AND.”

Novice. NOT-AND. In CMOS, an n-input NAND has n NMOS in series to ground and n PMOS in parallel to the supply. The output falls only when every input is high.

Expert. Series NMOS stack, parallel PMOS. Logical effort (n+2)/3(n+2)/3 and parasitic delay ≈n\approx n. Preferred over NOR in planar CMOS because its series devices are the faster NMOS.

Nanowire transistor

A gate-all-around transistor whose channels are thin round or square wires, a few nanometers across. The gate controls them very well, but each wire carries little current, and adding wires adds capacitance.

All three levels

Beginner. A transistor whose channel is a very thin wire with the gate wrapped all around it. It turns off well but carries little current.

Novice. A gate-all-around transistor whose channels are thin round or square wires, a few nanometers across. The gate controls them very well, but each wire carries little current, and adding wires adds capacitance.

Expert. GAA with channels of similar width and thickness. Best electrostatics per channel but small WeffW_{\mathrm{eff}} per footprint, which is why industry moved to wider, flat nanosheets.

Natural length (scale length, λ)

A length set by the shape and materials of a transistor that says how far the electric field of the source and drain spreads under the gate. A device behaves well when its gate is several natural lengths long.

All three levels

Beginner. How far the drain’s influence reaches into the channel. The gate stays in charge only if the channel is several times longer than this.

Novice. A length set by the shape and materials of a transistor that says how far the electric field of the source and drain spreads under the gate. A device behaves well when its gate is several natural lengths long.

Expert. The decay length of the 2D Poisson solution along the channel, λ≈εSi tch tox/(Nεox)\lambda \approx \sqrt{\varepsilon_{\mathrm{Si}}\, t_{\mathrm{ch}}\, t_{\mathrm{ox}} / (N \varepsilon_{\mathrm{ox}})} in the simple model, where NN grows with the number of gated sides. Short-channel effects scale as exp⁡(−L/2λ)\exp(-L/2\lambda), so minimum LL is a multiple of λ\lambda (about 3λ3\lambda for planar bulk in published scaling theory).

near-threshold computing

Running a chip with its supply voltage only slightly above the transistors’ switch-on voltage, roughly 400–500 mV in one 45 nm study, instead of about 1.1 V. Each operation takes far less energy, but the circuits run much slower.

All three levels

Beginner. Running a chip on just barely enough electrical push for its switches to work. It saves lots of energy but runs slower.

Novice. Running a chip with its supply voltage only slightly above the transistors’ switch-on voltage, roughly 400–500 mV in one 45 nm study, instead of about 1.1 V. Each operation takes far less energy, but the circuits run much slower.

Expert. Energy per operation falls steeply as supply approaches VTV_{\mathrm{T}}, but delay, and its sensitivity to process, voltage and temperature variation, rise sharply. It needs variation-aware timing, SRAM assist circuits or alternative bitcells, and robust level shifting.

Negotiated congestion

A form of rip-up and reroute in which a resource’s price rises with how many nets share it now and with how often it has been overused before. Nets with alternatives move away; the net with no other option keeps the resource.

All three levels

Beginner. Letting wires compete for crowded spots, with each spot getting more expensive every time it is fought over, until every wire finds a place.

Novice. A form of rip-up and reroute in which a resource’s price rises with how many nets share it now and with how often it has been overused before. Nets with alternatives move away; the net with no other option keeps the resource.

Expert. From PathFinder (1995): node cost (b+h)⋅p(b + h) \cdot p, with base cost bb, a history term hh that only grows, and a present-sharing penalty pp. Every net is rerouted each iteration, and timing-critical nets can weight delay over congestion. The idea underlies most academic and open global routers.

Net

A set of two or more pins that must be electrically connected: usually one output pin that drives a signal and the input pins that receive it. A chip has one net per signal, and routing turns each net into metal.

All three levels

Beginner. One connection in the chip’s parts list: a group of pins that must all be joined by wire.

Novice. A set of two or more pins that must be electrically connected: usually one output pin that drives a signal and the input pins that receive it. A chip has one net per signal, and routing turns each net into metal.

Expert. The unit the router works on. Its pin count (degree) decides how it is decomposed: about half of all nets have two pins, and multi-pin nets are split into two-pin connections along a Steiner tree. Power and ground are special nets, routed earlier by power planning.

Netlist

A list of every component in a circuit and the wires (nets) that join their pins. After synthesis each component is a standard cell from a library, and the list is written as a plain text file in the Verilog language, one line per cell.

All three levels

Beginner. A list of every part in a circuit and which parts connect to which, like the wiring list for a model kit.

Novice. A list of every component in a circuit and the wires (nets) that join their pins. After synthesis each component is a standard cell from a library, and the list is written as a plain text file in the Verilog language, one line per cell.

Expert. The gate-level Verilog file that carries the design from synthesis to layout: one line per cell instance, naming the cell, the instance and what each pin connects to. Later stages edit it (test logic, buffers, the clock tree), and an equivalence check re-proves it matches the RTL after each big change.

Network-on-chip (NoC)

An on-chip network of small routers joined by links. Blocks send data as packets that hop from router to router, much like a computer network. Common layouts are a ring, a 2D grid (mesh) and a grid whose edges wrap around (torus).

All three levels

Beginner. A small road network on the chip, with intersections, that carries data between the parts in little packets.

Novice. An on-chip network of small routers joined by links. Blocks send data as packets that hop from router to router, much like a computer network. Common layouts are a ring, a 2D grid (mesh) and a grid whose edges wrap around (torus).

Expert. Chosen by average hop count, bisection bandwidth (the total bandwidth across a line that cuts the network in half), router port count and layout cost. Routing, flow control, virtual channels and priority schemes then decide latency under load and whether traffic can deadlock.

Neural network

A model made of layers. Each layer multiplies its input by a grid of learned numbers (the weights), adds them up and applies a simple function. Training adjusts the weights until the outputs are useful.

All three levels

Beginner. A computer program that learns from examples instead of being told the rules. Inside, it’s a huge pile of numbers that get multiplied and added.

Novice. A model made of layers. Each layer multiplies its input by a grid of learned numbers (the weights), adds them up and applies a simple function. Training adjusts the weights until the outputs are useful.

Expert. A composition of parameterized layers (linear/convolutional, attention, normalization, nonlinearities). The linear and attention layers are matrix multiplications, which is why accelerators are built around them.

nextpnr

An open-source, timing-driven place-and-route tool for real FPGAs. It packs, places and routes a netlist from Yosys and, with an open bitstream database, writes the configuration file. It supports several FPGA families, some still experimental.

All three levels

Beginner. A free program that decides where each part goes on an FPGA and how to wire it all up.

Novice. An open-source, timing-driven place-and-route tool for real FPGAs. It packs, places and routes a netlist from Yosys and, with an open bitstream database, writes the configuration file. It supports several FPGA families, some still experimental.

Expert. Architecture-neutral P&R in which each architecture implements a device-database API. It has simulated-annealing and HeAP-style analytic placers and a timing-driven A* rip-up-and-reroute router, with backends for iCE40, ECP5, Nexus, Gowin and others.

NIC (network interface card)

A card that turns data in memory into network packets and back. AI servers often have one fast NIC per accelerator so each can talk to other servers.

All three levels

Beginner. The part that connects the server to the network cables.

Novice. A card that turns data in memory into network packets and back. AI servers often have one fast NIC per accelerator so each can talk to other servers.

Expert. A PCIe endpoint with Ethernet or InfiniBand ports, usually RDMA-capable. With GPUDirect-style peer-to-peer it can read and write accelerator memory directly, which is why NIC and accelerator often share a PCIe switch.

NLDM (non-linear delay model)

Liberty’s way of storing a cell’s delay: a small table whose rows are how quickly the input signal changes and whose columns are how much load the output drives. The tool reads between the entries.

All three levels

Beginner. A small table that tells how slow a building block is, depending on what goes into it and how much it has to drive.

Novice. Liberty’s way of storing a cell’s delay: a small table whose rows are how quickly the input signal changes and whose columns are how much load the output drives. The tool reads between the entries.

Expert. Two-dimensional tables of delay and output transition, measured in circuit simulation at a grid of input-transition and load values. Tools interpolate inside the grid and extrapolate, less accurately, outside it. Advanced nodes add current-source models (CCS, ECSM), but synthesis runs well on NLDM.

NMOS

An n-channel MOSFET: n+ source and drain in a p-type body. A positive gate-to-source voltage above the threshold pulls electrons into a channel and turns it on.

All three levels

Beginner. The kind of transistor switch that turns on when its gate voltage is high.

Novice. An n-channel MOSFET: n+ source and drain in a p-type body. A positive gate-to-source voltage above the threshold pulls electrons into a channel and turns it on.

Expert. Electron channel, so roughly 2–3× the drive of a same-size PMOS. Passes a strong 0 and a degraded 1, which is why it forms the pull-down half of a CMOS gate.

Noise margin

The gap between the worst output voltage a gate produces for a logic level and the worst input voltage the next gate still accepts for that level. NML=VIL−VOL\mathrm{NM}_{\mathrm{L}} = V_{\mathrm{IL}} - V_{\mathrm{OL}} and NMH=VOH−VIH\mathrm{NM}_{\mathrm{H}} = V_{\mathrm{OH}} - V_{\mathrm{IH}}.

All three levels

Beginner. How much a signal can be nudged by stray electricity before the next gate reads it wrong.

Novice. The gap between the worst output voltage a gate produces for a logic level and the worst input voltage the next gate still accepts for that level. NML=VIL−VOL\mathrm{NM}_{\mathrm{L}} = V_{\mathrm{IL}} - V_{\mathrm{OL}} and NMH=VOH−VIH\mathrm{NM}_{\mathrm{H}} = V_{\mathrm{OH}} - V_{\mathrm{IH}}.

Expert. Static noise margins taken at the unity-gain (slope −1) points of the VTC. They budget DC disturbances such as supply drop, ground bounce and coupling; dynamic noise (glitch width and energy) needs separate analysis.

Noise-aware training

Adding random perturbations, shaped like the hardware’s real errors, to the weights during training. The network learns settings that are less sensitive to those errors.

All three levels

Beginner. Training an AI model with random errors added on purpose. Then it still works on chips that aren’t perfect.

Novice. Adding random perturbations, shaped like the hardware’s real errors, to the weights during training. The network learns settings that are less sensitive to those errors.

Expert. Hardware-aware training injects noise drawn from measured device statistics, optionally with quantization and ADC models in the loop; chip-in-the-loop fine-tuning then absorbs residual errors layer by layer.

Non-blocking assignment (<=)

The <= sign. The simulator first reads every right-hand side, then updates all the left-hand sides together. That matches real flip-flops, which all take in their inputs at the same clock edge. Used for clocked logic.

All three levels

Beginner. An instruction that waits and updates all the stored values together at the clock tick, like everyone flipping their cards at once.

Novice. The <= sign. The simulator first reads every right-hand side, then updates all the left-hand sides together. That matches real flip-flops, which all take in their inputs at the same clock edge. Used for clocked logic.

Expert. The right-hand side is evaluated in the Active region and the update is queued for the NBA region later in the same time step. Because no register changes until every clocked block has read its inputs, the order in which those blocks run no longer matters.

Non-default routing rule (NDR)

A wiring rule that overrides the normal wire width and spacing for chosen wires. Clock wires often get double spacing, so neighboring wires disturb them less, and sometimes double width, which lowers their resistance.

All three levels

Beginner. Special, roomier wiring rules used for the clock so it stays fast and quiet.

Novice. A wiring rule that overrides the normal wire width and spacing for chosen wires. Clock wires often get double spacing, so neighboring wires disturb them less, and sometimes double width, which lowers their resistance.

Expert. Double spacing cuts coupling capacitance and so crosstalk-induced delay changes; double width cuts resistance and helps electromigration. Usually applied to the trunk and upper levels; leaf nets keep default rules to save routing tracks.

NOR gate

NOT-OR. In CMOS, an n-input NOR has n NMOS in parallel to ground and n PMOS in series to the supply. The output rises only when every input is low.

All three levels

Beginner. A gate whose output is 1 only when all of its inputs are 0. It means “not OR.”

Novice. NOT-OR. In CMOS, an n-input NOR has n NMOS in parallel to ground and n PMOS in series to the supply. The output rises only when every input is low.

Expert. Parallel NMOS, series PMOS stack. Logical effort (2n+1)/3(2n+1)/3 and parasitic delay ≈n\approx n. The series PMOS stack makes it slower and larger than a NAND of the same fan-in when PMOS is the weaker device.

NPU

Neural processing unit: a block on a chip built to run trained AI models efficiently, usually a grid of multiply-and-add units with its own local memory.

All three levels

Beginner. The part of a phone chip built for AI jobs, like knowing faces or speech.

Novice. Neural processing unit: a block on a chip built to run trained AI models efficiently, usually a grid of multiply-and-add units with its own local memory.

Expert. Sized for performance per watt within a tight area budget. It usually has its own power domain and DVFS, and its SRAM is a major leakage contributor.

NRE

Non-recurring engineering: the one-time cost of design work, licensed blocks, tools and the photomasks, paid once however many chips are made.

All three levels

Beginner. The one-time cost of designing a new chip, paid before the first one is made.

Novice. Non-recurring engineering: the one-time cost of design work, licensed blocks, tools and the photomasks, paid once however many chips are made.

Expert. Spread across all units, so it sets the volume at which an ASIC beats an FPGA. Mask cost rises steeply at advanced process nodes.

NRZ

Non-return-to-zero: two signal levels, one bit per symbol. Used for lanes up to 25 Gb/s, after which links moved to PAM4.

All three levels

Beginner. The simplest way to send bits: high means 1, low means 0.

Novice. Non-return-to-zero: two signal levels, one bit per symbol. Used for lanes up to 25 Gb/s, after which links moved to PAM4.

Expert. Two-level signaling (PAM2). Nyquist frequency is half the bit rate, so a 25.78 Gb/s NRZ lane has its Nyquist at 12.9 GHz, about the same as a 53 Gb/s PAM4 lane.

NUMA (non-uniform memory access)

A design where each CPU has its own local memory and can also reach the other CPU’s memory, but more slowly. Software tries to keep data near the CPU that uses it.

All three levels

Beginner. When some memory sits closer to a processor than other memory. The far memory takes longer to reach.

Novice. A design where each CPU has its own local memory and can also reach the other CPU’s memory, but more slowly. Software tries to keep data near the CPU that uses it.

Expert. Memory is partitioned into nodes with different latency and bandwidth from each core. Remote-socket accesses cross the inter-socket link; CXL memory shows up as a further, CPU-less node. Placement and page migration policies decide the performance.

Numerical aperture (NA)

A measure of how wide a cone of light a lens collects: NA=nsin⁡α\mathrm{NA} = n \sin\alpha, where α\alpha is the half-angle of the cone and nn is the refractive index of whatever fills the gap between lens and wafer (1 for air or vacuum). Filling the gap with water lets NA go above 1.

All three levels

Beginner. A number that says how wide a cone of light a lens can gather. A bigger number means sharper detail.

Novice. A measure of how wide a cone of light a lens collects: NA=nsin⁡α\mathrm{NA} = n \sin\alpha, where α\alpha is the half-angle of the cone and nn is the refractive index of whatever fills the gap between lens and wafer (1 for air or vacuum). Filling the gap with water lets NA go above 1.

Expert. Resolution improves as 1/NA1/\mathrm{NA}, but depth of focus shrinks as 1/NA21/\mathrm{NA}^2. Doubling NA therefore halves the smallest printable feature but cuts the focus window to a quarter, which demands flatter wafers and thinner resist.

NVMe drive

NVM Express: the standard way a computer talks to solid-state (flash) drives over PCIe. Data-center NVMe drives usually use four PCIe lanes each.

All three levels

Beginner. A fast storage drive, like the one in a laptop, that plugs straight into the processor’s fast wires.

Novice. NVM Express: the standard way a computer talks to solid-state (flash) drives over PCIe. Data-center NVMe drives usually use four PCIe lanes each.

Expert. A queue-based command protocol (many deep submission/completion queues) over PCIe, RDMA or TCP. Form factors include U.2, M.2 and EDSFF E1.S/E3.S; an E3 1C connector carries x4, a 2C connector x8.

Nyquist frequency

Half the symbol rate of a signal. A lane sending 53 billion symbols per second has its Nyquist frequency at about 26.6 GHz. Losses are usually quoted at this frequency.

All three levels

Beginner. Roughly the fastest wiggle a signal has to make to carry its data.

Novice. Half the symbol rate of a signal. A lane sending 53 billion symbols per second has its Nyquist frequency at about 26.6 GHz. Losses are usually quoted at this frequency.

Expert. fN=baud/2f_{\mathrm{N}} = \mathrm{baud}/2. The reference frequency for channel loss specs: 26.56 GHz for 100G PAM4 (53.125 GBd) and 53.125 GHz for 200G PAM4 (106.25 GBd).

OAM (OCP Accelerator Module)

An Open Compute Project specification for a flat accelerator module that mounts onto a baseboard through two connectors, instead of plugging into a slot like a graphics card.

All three levels

Beginner. A standard size and plug for accelerator boards, so different companies’ chips fit the same servers.

Novice. An Open Compute Project specification for a flat accelerator module that mounts onto a baseboard through two connectors, instead of plugging into a slot like a graphics card.

Expert. A 102 × 165 mm mezzanine with one or two x16 host links and up to seven inter-module links, powered from 48/54 V for high-TDP modules. A universal baseboard (UBB) carries eight of them with the scale-up wiring between them.

OASIS

Open Artwork System Interchange Standard (SEMI P39, version 1.0 in 2004), a successor to GDSII that stores the same cells and shapes in much smaller files.

All three levels

Beginner. A newer, more compact file format for chip drawings.

Novice. Open Artwork System Interchange Standard (SEMI P39, version 1.0 in 2004), a successor to GDSII that stores the same cells and shapes in much smaller files.

Expert. Compacts with variable-length integers, short records for rectangles, modal variables that let a record omit values repeated from the previous one, repetition records for arrays and irregular sets of identical shapes, and optional compressed cells. Fill and post-OPC data benefit most.

Occupancy

The number of warps resident on a GPU core divided by the maximum it supports. More resident warps give the scheduler more choices while others wait on memory.

All three levels

Beginner. How many teams of workers a GPU core is holding at once, compared with how many it could hold.

Novice. The number of warps resident on a GPU core divided by the maximum it supports. More resident warps give the scheduler more choices while others wait on memory.

Expert. Limited by warp slots, registers per thread, shared memory per block and block-count limits. Necessary for latency hiding only up to the Little’s-law requirement; tuned GEMM kernels often run at low occupancy and hide latency with ILP instead.

Off-current (Ioff)

The drain current with the gate at 0 V and the drain at the supply voltage. It is mostly subthreshold current, so it rises steeply as the threshold voltage is lowered or the chip gets hotter.

All three levels

Beginner. The tiny current a transistor leaks when it is off. Billions of tiny leaks add up to wasted power.

Novice. The drain current with the gate at 0 V and the drain at the supply voltage. It is mostly subthreshold current, so it rises steeply as the threshold voltage is lowered or the chip gets hotter.

Expert. IDSI_{\mathrm{DS}} at VGS=0V_{\mathrm{GS}} = 0, VDS=VDDV_{\mathrm{DS}} = V_{\mathrm{DD}}. Dominated by subthreshold conduction (plus gate and junction leakage), raised by DIBL at full VDSV_{\mathrm{DS}} and roughly tenfold for each SS of VtV_{\mathrm{t}} reduction. The ratio Ion/IoffI_{\mathrm{on}}/I_{\mathrm{off}} summarizes a device’s switch quality.

On-chip variation (OCV)

Small manufacturing differences that make identical cells on one chip run at slightly different speeds. Timing tools allow for it by assuming one path runs a few percent slow while the path it is compared with runs a few percent fast (for example, late delays ×1.05 and early delays ×0.95).

All three levels

Beginner. Tiny manufacturing differences that make identical parts on the same chip run at slightly different speeds.

Novice. Small manufacturing differences that make identical cells on one chip run at slightly different speeds. Timing tools allow for it by assuming one path runs a few percent slow while the path it is compared with runs a few percent fast (for example, late delays ×1.05 and early delays ×0.95).

Expert. A flat derate applies one factor everywhere; AOCV scales the factor by path depth and distance; POCV gives each cell its own delay sigma, stored in the library as LVF. Clock paths are derated too, which is why a long unshared clock path costs slack.

On-current (Ion)

The drain current with the gate and drain both at the supply voltage (VGS=VDS=VDDV_{\mathrm{GS}} = V_{\mathrm{DS}} = V_{\mathrm{DD}}), usually quoted per micrometer of width. It sets how fast a gate can charge the next one.

All three levels

Beginner. The most current a transistor passes when it is fully on. More means faster circuits.

Novice. The drain current with the gate and drain both at the supply voltage (VGS=VDS=VDDV_{\mathrm{GS}} = V_{\mathrm{DS}} = V_{\mathrm{DD}}), usually quoted per micrometer of width. It sets how fast a gate can charge the next one.

Expert. IDSI_{\mathrm{DS}} at VGS=VDS=VDDV_{\mathrm{GS}} = V_{\mathrm{DS}} = V_{\mathrm{DD}}, in saturation. Usually quoted in µA/µm at a stated corner and temperature. Gate delay scales roughly as CVDD/IonC V_{\mathrm{DD}}/I_{\mathrm{on}}, so IonI_{\mathrm{on}} is the headline speed figure on a device’s I-V plot.

On-resistance

The resistance between source and drain when the transistor is on. It falls as the gate voltage rises and as the transistor is made wider. Designers quote it as a number times the width, such as kΩ·µm.

All three levels

Beginner. How hard it is for electricity to squeeze through a switch that is on. Lower is better.

Novice. The resistance between source and drain when the transistor is on. It falls as the gate voltage rises and as the transistor is made wider. Designers quote it as a number times the width, such as kΩ·µm.

Expert. Usually an effective resistance averaged over a switching event, chosen so that delay ≈RC\approx RC. A common estimate is about 0.75 VDD/Ion0.75\,V_{\mathrm{DD}}/I_{\mathrm{on}}. It scales as 1/W1/W, which makes RCRC of a transistor driving its own gate width-independent.

One-hot encoding

A way to record which step a state machine is on: one flip-flop per step, with exactly one of them set to 1. It uses more flip-flops than counting in binary, but the logic that chooses the next step is simpler and faster.

All three levels

Beginner. A way to keep track of which step a machine is on by giving each step its own on/off switch.

Novice. A way to record which step a state machine is on: one flip-flop per step, with exactly one of them set to 1. It uses more flip-flops than counting in binary, but the logic that chooses the next step is simpler and faster.

Expert. One of several state encodings (binary, Gray, one-hot) a tool may choose after it extracts a state machine from the RTL. Re-encoding changes the set of flip-flops, so the equivalence checker must be told the new encoding or check sequentially.

Open Compute Project (OCP)

A foundation where large data-center operators and vendors publish open specifications for racks, power, servers and modules, so parts from different suppliers fit together.

All three levels

Beginner. A group of companies that share open designs for data-center hardware.

Novice. A foundation where large data-center operators and vendors publish open specifications for racks, power, servers and modules, so parts from different suppliers fit together.

Expert. Hosts open specs such as Open Rack V3 (48 V busbar), OAM/UBB, OCP NIC 3.0 and DC-SCM. Contributions range from base specifications to full reference designs.

Open-source FPGA toolchain

A chain of open-source tools that goes from hardware code to a bitstream with no vendor software: usually Yosys for synthesis, nextpnr for place and route, and a bit packer built on an open documentation project such as IceStorm or Trellis.

All three levels

Beginner. A set of free programs that take FPGA code all the way to the finished settings file, without the chip maker’s software.

Novice. A chain of open-source tools that goes from hardware code to a bitstream with no vendor software: usually Yosys for synthesis, nextpnr for place and route, and a bit packer built on an open documentation project such as IceStorm or Trellis.

Expert. A free HDL-to-bitstream flow for commercial devices. It depends on reverse-engineered device data (routing graph, timing, bit mapping): Yosys + nextpnr + IceStorm, Trellis, Oxide or Apicula, or F4PGA with VPR or nextpnr.

Operand reuse

How many operations use each value read from a given level of memory. In a matrix multiplication every input value is needed many times, so a design that keeps it nearby saves memory traffic.

All three levels

Beginner. Using a number many times after fetching it once, instead of fetching it again each time.

Novice. How many operations use each value read from a given level of memory. In a matrix multiplication every input value is needed many times, so a design that keeps it nearby saves memory traffic.

Expert. Measured per level of the hierarchy (DRAM, buffer, array, register). In an R×CR \times C weight-stationary array each activation is reused CC times on its way across and each weight once per streamed row. Reuse is bounded by the algorithm and the capacity of each level.

Operator fusion

Combining consecutive operations of a model (say, a matrix multiply followed by an activation) so the intermediate result stays on chip instead of being written to and read back from external memory.

All three levels

Beginner. Doing two or more steps in one go, so the in-between result never has to be written out to slow memory.

Novice. Combining consecutive operations of a model (say, a matrix multiply followed by an activation) so the intermediate result stays on chip instead of being written to and read back from external memory.

Expert. Raises arithmetic intensity by removing intermediate DRAM round trips. On GPUs it usually means hand-written or compiler-generated combined kernels covering a few operators; spatial dataflow chips can pipeline much longer operator chains across the fabric.

Optical engine

A compact transmitter-receiver: a photonic chip with modulators and photodetectors, plus an electronic chip with drivers and amplifiers. In CPO, several engines surround the switch chip.

All three levels

Beginner. A tiny chip set that turns electrical signals into light and back, small enough to sit beside a processor.

Novice. A compact transmitter-receiver: a photonic chip with modulators and photodetectors, plus an electronic chip with drivers and amplifiers. In CPO, several engines surround the switch chip.

Expert. PIC (modulators, photodiodes, waveguides, fiber coupling) stacked or bonded with an EIC (drivers, TIAs, control). Lasers are usually external. Power, bandwidth per mm of package edge and yield of known-good engines are its key metrics.

Optical proximity correction (OPC)

Pre-distorting the shapes on a mask to make up for blurring in lithography: moving edges, widening line ends into serifs or hammerheads, and adding thin assist bars too narrow to print themselves.

All three levels

Beginner. Changing the stencil shapes on purpose so that, after the blurry printing, the shapes on the chip come out right.

Novice. Pre-distorting the shapes on a mask to make up for blurring in lithography: moving edges, widening line ends into serifs or hammerheads, and adding thin assist bars too narrow to print themselves.

Expert. Model-based OPC splits polygon edges into segments, simulates the printed contour with a calibrated optical and resist model, measures the edge placement error at evaluation points, and moves each segment along its normal until the error converges. Full-chip runs need large compute farms.

Optimizer state

Extra values the optimizer keeps per weight. The common Adam optimizer stores two running averages per weight plus a high-precision copy of the weight, so it uses several times more memory than the weights.

All three levels

Beginner. Extra notes the learning method keeps about each number in the model, like which way it has been moving lately.

Novice. Extra values the optimizer keeps per weight. The common Adam optimizer stores two running averages per weight plus a high-precision copy of the weight, so it uses several times more memory than the weights.

Expert. For mixed-precision Adam: fp32 master weights, momentum and variance, 12 bytes per parameter (K=12K = 12 in the ZeRO paper), on top of 2-byte weights and 2-byte gradients: 16 bytes per parameter in total.

order book

For each traded instrument, the list of outstanding buy orders (bids) and sell orders (offers) by price. A trading system keeps its own copy, updated from the exchange’s messages.

All three levels

Beginner. A list of everyone who wants to buy or sell something, and at what price.

Novice. For each traded instrument, the list of outstanding buy orders (bids) and sell orders (offers) by price. A trading system keeps its own copy, updated from the exchange’s messages.

Expert. Often kept in aggregated form, as price levels with total quantity at each. In hardware the challenge is finding an instrument’s record quickly: a large book doesn’t fit on chip, so lookups go to external SRAM through a hash table.

Out-of-order execution

A way of running a program in which the processor executes each instruction as soon as its inputs are ready, rather than strictly in program order, and then commits the results in program order so the program sees no difference.

All three levels

Beginner. When a processor runs the steps of a program in whatever order their inputs are ready, then puts the results back in the right order.

Novice. A way of running a program in which the processor executes each instruction as soon as its inputs are ready, rather than strictly in program order, and then commits the results in program order so the program sees no difference.

Expert. Dynamic scheduling of a window of renamed instructions in dataflow order, with in-order commit through a reorder buffer. It hides latency (cache misses, long operations) by finding independent work, bounded by window size, issue width and branch prediction.

Outliers (in activations)

Values far larger than the typical value in a tensor. Large language models develop a few activation channels with values up to about twenty times larger than the rest, and they matter for accuracy, so they can’t simply be clipped.

All three levels

Beginner. A few numbers in a block of data that are much bigger than all the others.

Novice. Values far larger than the typical value in a tensor. Large language models develop a few activation channels with values up to about twenty times larger than the rest, and they matter for accuracy, so they can’t simply be clipped.

Expert. Systematic large-magnitude features that appear in specific hidden dimensions of transformers (about 0.1% of features, up to ~20× larger, in all layers beyond a few billion parameters). With a per-tensor scale they force a coarse grid on everything else; fixes include mixed-precision decomposition, finer scaling granularity and rotation transforms.

Output resistance (ro)

The inverse of the slope of the IDI_{\mathrm{D}}–VDSV_{\mathrm{DS}} curve in saturation: ro=ΔVDS/ΔIDr_{\mathrm{o}} = \Delta V_{\mathrm{DS}}/\Delta I_{\mathrm{D}}. An ideal current source has infinite ror_{\mathrm{o}}; real transistors have a finite ror_{\mathrm{o}} because of channel-length modulation and DIBL.

All three levels

Beginner. How well a transistor keeps its current steady when the push across it changes. Higher means steadier.

Novice. The inverse of the slope of the IDI_{\mathrm{D}}–VDSV_{\mathrm{DS}} curve in saturation: ro=ΔVDS/ΔIDr_{\mathrm{o}} = \Delta V_{\mathrm{DS}}/\Delta I_{\mathrm{D}}. An ideal current source has infinite ror_{\mathrm{o}}; real transistors have a finite ror_{\mathrm{o}} because of channel-length modulation and DIBL.

Expert. ro=(∂ID/∂VDS)−1≈1/(λID)r_{\mathrm{o}} = (\partial I_{\mathrm{D}}/\partial V_{\mathrm{DS}})^{-1} \approx 1/(\lambda I_{\mathrm{D}}) from CLM alone, lower again with DIBL. Combined with transconductance it gives the intrinsic gain gmrog_{\mathrm{m}} r_{\mathrm{o}}, which falls as channels get shorter.

Output-stationary (OS)

A dataflow in which each processing element owns one output and accumulates it in place while activations and weights stream in. Partial sums never move until the output is finished.

All three levels

Beginner. A way of running a grid of calculators where each cell keeps its own running total and both kinds of input flow past it.

Novice. A dataflow in which each processing element owns one output and accumulates it in place while activations and weights stream in. Partial sums never move until the output is finished.

Expert. Minimizes partial-sum traffic and keeps the wide accumulator local; both narrow operands must be streamed, and a new output tile means re-reading the other operand. Eyeriss distinguishes several variants by which outputs are computed together.

Overlay

The misalignment between a layer and the layer it has to connect to, measured on special alignment marks. Too much overlay error leaves contacts missing their targets or shorts neighboring wires together.

All three levels

Beginner. How well one layer lines up with the layer under it.

Novice. The misalignment between a layer and the layer it has to connect to, measured on special alignment marks. Too much overlay error leaves contacts missing their targets or shorts neighboring wires together.

Expert. A shared error budget: the scanner, the mask, the process and, on layers printed with several masks, the placement of each mask relative to the others all add to it. At advanced nodes even sub-nanometer misalignment can make a chip fail.

Oversubscription

The ratio of the total bandwidth of the links below a switch to the bandwidth of the single link above it. At 4:1, devices that all send toward the CPU at once each get a quarter of their own link speed.

All three levels

Beginner. When more data can pour in than a shared path can carry. If everyone sends at once, everyone slows down.

Novice. The ratio of the total bandwidth of the links below a switch to the bandwidth of the single link above it. At 4:1, devices that all send toward the CPU at once each get a quarter of their own link speed.

Expert. Downstream lanes ÷ upstream lanes (same generation). Harmless for traffic that stays below the switch (peer-to-peer) or that is bursty and not simultaneous; a hard limit for host-bound streaming.

p-n junction

The interface between p-type and n-type regions. Carriers diffuse across, leaving a depletion region and a built-in voltage. Forward bias (p positive) lowers the barrier and current rises exponentially; reverse bias raises it and only a tiny leakage flows. Every source and drain in a MOSFET forms one with the body.

All three levels

Beginner. The border where p-type and n-type silicon meet. It lets electricity through one way and blocks it the other way.

Novice. The interface between p-type and n-type regions. Carriers diffuse across, leaving a depletion region and a built-in voltage. Forward bias (p positive) lowers the barrier and current rises exponentially; reverse bias raises it and only a tiny leakage flows. Every source and drain in a MOSFET forms one with the body.

Expert. Ideal current I=I0(eqV/kT−1)I = I_{0}(e^{qV/kT} - 1) with I0=Aqni2 (Dp/(LpNd)+Dn/(LnNa))I_{0} = Aqn_{\mathrm{i}}^2\,(D_{\mathrm{p}}/(L_{\mathrm{p}} N_{\mathrm{d}}) + D_{\mathrm{n}}/(L_{\mathrm{n}} N_{\mathrm{a}})). Real junctions add generation-recombination current, series resistance, high-injection effects and avalanche or tunneling breakdown.

p-type

Silicon doped with acceptors (usually boron). Holes are the majority carriers and the Fermi level sits closer to the valence band.

All three levels

Beginner. Silicon with extra holes, made by mixing in atoms that each have one electron too few.

Novice. Silicon doped with acceptors (usually boron). Holes are the majority carriers and the Fermi level sits closer to the valence band.

Expert. p≈Nap \approx N_{\mathrm{a}}, n=ni2/Nan = n_{\mathrm{i}}^2/N_{\mathrm{a}}, Ei−EF=kTln⁡(Na/ni)=qϕBE_{\mathrm{i}} - E_{\mathrm{F}} = kT \ln(N_{\mathrm{a}}/n_{\mathrm{i}}) = q\phi_{\mathrm{B}}. The body of an NMOS transistor is p-type; ϕB\phi_{\mathrm{B}} sets the 2ϕB2\phi_{\mathrm{B}} band bending needed for inversion.

Package

The case that holds the bare silicon chip, connects its signals and power to solder balls or pins on the circuit board, and carries heat away. In a wire-bond package fine wires join pads on the chip’s edge to the package; in a flip-chip package the chip sits face-down on tiny solder bumps.

All three levels

Beginner. The protective case around the silicon chip that connects it to the circuit board and helps carry heat away.

Novice. The case that holds the bare silicon chip, connects its signals and power to solder balls or pins on the circuit board, and carries heat away. In a wire-bond package fine wires join pads on the chip’s edge to the package; in a flip-chip package the chip sits face-down on tiny solder bumps.

Expert. Package choice sets the number of balls, the inductance of the power connections, the thermal resistance (and so the TDP), the cost, and whether several dies can share one package. It is chosen at the spec stage because it limits die size, pad placement and the number of signals.

Package substrate

A small multilayer circuit board inside the package. It carries signals and power from the chip’s tiny connections out to the much larger solder balls that attach the package to the main board.

All three levels

Beginner. The small circuit board inside a chip’s case. It spreads the chip’s tiny connections out to bigger ones that solder onto the main board.

Novice. A small multilayer circuit board inside the package. It carries signals and power from the chip’s tiny connections out to the much larger solder balls that attach the package to the main board.

Expert. Its line width, via size and layer count limit how many of the die’s connections can be routed out and how well power is delivered. When it can’t reach the wiring density two dies need between them, a silicon interposer or bridge goes between the dies and the substrate.

Packet spraying

Spreading the packets of a single flow across all equal-cost paths, packet by packet. Load becomes even, but packets can arrive out of order, so the receiver has to put them back together.

All three levels

Beginner. Sending the pieces of one big message down all the available routes at once instead of one route.

Novice. Spreading the packets of a single flow across all equal-cost paths, packet by packet. Load becomes even, but packets can arrive out of order, so the receiver has to put them back together.

Expert. Per-packet (or per-small-batch) multipath. Removes hash collisions; requires a transport that tolerates reordering (selective acknowledgment, out-of-order placement) and handles loss detection without relying on ordering. Ultra Ethernet’s unordered delivery modes are designed around it.

Packing (clustering)

The FPGA step between technology mapping and placement. It groups lookup tables and flip-flops into the chip’s logic blocks, within each block’s limits on inputs, clocks and size. Connections inside a block use fast local wiring instead of the general routing.

All three levels

Beginner. Grouping small parts that work together into one bigger block on the FPGA, so the wires between them stay short.

Novice. The FPGA step between technology mapping and placement. It groups lookup tables and flip-flops into the chip’s logic blocks, within each block’s limits on inputs, clocks and size. Connections inside a block use fast local wiring instead of the general routing.

Expert. Clustering basic logic elements into logic-block sites under legality constraints (input pins, control signals, carry chains). Classic tools seed a cluster and greedily absorb the element with the most shared nets, weighted by timing criticality (VPack, T-VPack).

Pad-limited vs. core-limited

A pad-limited chip is bigger than its logic needs because its edge must be long enough to fit all its pads. A core-limited chip’s size is set by the logic inside, and the pads have room to spare.

All three levels

Beginner. Whether a chip’s size is set by how many connection points must fit around its edge, or by how much fits inside.

Novice. A pad-limited chip is bigger than its logic needs because its edge must be long enough to fit all its pads. A core-limited chip’s size is set by the logic inside, and the pads have room to spare.

Expert. Perimeter needed≈pad count×pad pitch\text{Perimeter needed} \approx \text{pad count} \times \text{pad pitch} (the pitch is the center-to-center spacing). Ways out include staggered or multi-row pads, fewer power pads, or flip-chip bumps, which take I/O off the perimeter.

PAM4

Four-level pulse-amplitude modulation. Each symbol takes one of four voltage (or light) levels, so it carries two bits. It doubles the data rate at the same symbol rate, but the levels are closer together, so the signal is more sensitive to noise.

All three levels

Beginner. A way of sending two bits at once by using four different signal levels instead of just on and off.

Novice. Four-level pulse-amplitude modulation. Each symbol takes one of four voltage (or light) levels, so it carries two bits. It doubles the data rate at the same symbol rate, but the levels are closer together, so the signal is more sensitive to noise.

Expert. Two bits per symbol; for the same peak swing each eye is one third the height of NRZ, about a 9.5 dB SNR penalty. Used from 50 Gb/s lanes upward, which is why links rely on forward error correction (RS(544,514) in Ethernet).

Parameter (weight)

A learned number in a neural network, mostly the weights of its matrix multiplications. Memory needed = number of parameters × bytes per parameter: 2 bytes each in 16-bit formats, 1 byte in 8-bit.

All three levels

Beginner. One of the numbers a trained AI model has learned. Big models have billions of them, and all have to be stored somewhere.

Novice. A learned number in a neural network, mostly the weights of its matrix multiplications. Memory needed = number of parameters × bytes per parameter: 2 bytes each in 16-bit formats, 1 byte in 8-bit.

Expert. Weight footprint is PbP b bytes. Inference also needs activations and a KV cache; training adds gradients and optimizer state, typically several times the weight footprint.

Parasitic (RC) extraction

Calculating, from the shapes of the finished wiring, how much resistance each wire has (how much it resists current) and how much capacitance (how much charge it must be filled with to switch), including the capacitance between neighboring wires. These unwanted but unavoidable properties are the “parasitics” that slow signals down.

All three levels

Beginner. Measuring the hidden drag on every wire in the finished drawing. The drag slows signals down.

Novice. Calculating, from the shapes of the finished wiring, how much resistance each wire has (how much it resists current) and how much capacitance (how much charge it must be filled with to switch), including the capacitance between neighboring wires. These unwanted but unavoidable properties are the “parasitics” that slow signals down.

Expert. Computing each routed net’s resistor network, capacitance to ground and coupling capacitance to neighbors from the layout and a calibrated technology file, once per interconnect corner. Full-chip extractors look up pre-characterized wire configurations rather than solving the physics directly; the result is written as SPEF for timing.

Pareto front

When there are several goals (say, low power and high speed), the options that no other option beats on every goal at once. Each point on the front is a different trade-off.

All three levels

Beginner. The set of best compromises: options where you can’t improve one thing without making another worse.

Novice. When there are several goals (say, low power and high speed), the options that no other option beats on every goal at once. Each point on the front is a different trade-off.

Expert. Multi-objective optimizers return the front instead of one answer and leave the final weighting to the engineer. With noisy objectives, points near the front can swap places on a rerun.

Partial reconfiguration

Loading a partial bitstream that rewrites only one region of the FPGA, while the rest of the design keeps working. AMD calls it Dynamic Function eXchange.

All three levels

Beginner. Changing the design in one part of an FPGA while the rest keeps running.

Novice. Loading a partial bitstream that rewrites only one region of the FPGA, while the rest of the design keeps working. AMD calls it Dynamic Function eXchange.

Expert. Rewriting the configuration frames of a reconfigurable region while the static logic keeps running. Frames are the smallest unit of configuration, so regions are made of whole frames.

Partial sum

An output value part-way through its accumulation: the sum of some, but not all, of the KK products that make it up.

All three levels

Beginner. A running total that isn’t finished yet. More products still have to be added to it.

Novice. An output value part-way through its accumulation: the sum of some, but not all, of the KK products that make it up.

Expert. Often abbreviated psum. Partial sums are wider than the operands (for example 32-bit sums of 8-bit products), so moving them costs more wire and energy than moving inputs; where they live is a key dataflow choice.

pass@k

A score for code-writing models: the chance that at least one of kk attempts passes a problem’s tests, averaged over problems. pass@1 is the first-try success rate.

All three levels

Beginner. A score for code-writing AI: the chance that at least one of several tries passes the tests.

Novice. A score for code-writing models: the chance that at least one of kk attempts passes a problem’s tests, averaged over problems. pass@1 is the first-try success rate.

Expert. Estimated without bias by drawing n≥kn \ge k samples, counting the cc that pass, and averaging 1−(n−ck)/(nk)1 - \binom{n-c}{k} \big/ \binom{n}{k} over problems. It assumes something can pick out the passing attempt, and it is only as strong as the tests behind it.

Path delay fault

A fault model in which the total delay along one specific chain of gates (a path) is longer than one clock period, even though no single gate is badly slow. Each path has a rising and a falling version.

All three levels

Beginner. A pretend flaw where each step along a route is a little slow, and together they make the whole route too slow.

Novice. A fault model in which the total delay along one specific chain of gates (a path) is longer than one clock period, even though no single gate is badly slow. Each path has a rising and a falling version.

Expert. Models delay spread along a path. The number of paths grows exponentially with circuit size, so teams test the critical paths that timing analysis reports. A robust test detects the fault whatever the other paths do; a non-robust one assumes they are fault-free.

PCIe (PCI Express)

A point-to-point serial link standard. Each link is made of one or more lanes; each new generation doubles the data rate per lane while staying compatible with older devices.

All three levels

Beginner. The standard kind of super-fast wiring that joins the parts inside a computer to its processor.

Novice. A point-to-point serial link standard. Each link is made of one or more lanes; each new generation doubles the data rate per lane while staying compatible with older devices.

Expert. A load-store interconnect with a layered protocol (transaction, data link, physical). Generations 1–5 used NRZ signaling with 8b/10b or 128b/130b coding; 6.0 and 7.0 use PAM4 and fixed 256-byte flits with FEC. Links negotiate width and rate down to what both ends support.

PCIe lane

The basic unit of a PCIe link: two differential pairs, one per direction, so a lane sends and receives at the same time. A “x16” link bundles 16 lanes and has 16 times the bandwidth of one.

All three levels

Beginner. One set of PCIe wires. It works like a two-way road, with one side for sending and one for receiving.

Novice. The basic unit of a PCIe link: two differential pairs, one per direction, so a lane sends and receives at the same time. A “x16” link bundles 16 lanes and has 16 times the bandwidth of one.

Expert. One transmit and one receive differential pair, each carrying the full per-lane rate (for example 32 GT/s at PCIe 5.0). A x16 link therefore has 64 signal wires. Lanes are the scarce resource on a CPU package and a board: the lane budget.

PCIe switch

A chip with one upstream port toward the CPU and several downstream ports toward devices. It lets more devices attach than the CPU has lanes for, and lets devices behind it send data straight to each other.

All three levels

Beginner. A chip that lets several devices share one connection to the processor, like a power strip shares one wall socket.

Novice. A chip with one upstream port toward the CPU and several downstream ports toward devices. It lets more devices attach than the CPU has lanes for, and lets devices behind it send data straight to each other.

Expert. A set of logical PCI-PCI bridges. It fans out lanes at the cost of oversubscribing the upstream link and adding some latency, and it routes peer-to-peer traffic (GPU↔NIC, NVMe→GPU) without touching the root complex when access control allows.

PDK (process design kit)

The package a chip factory (foundry) gives designers for one manufacturing process: the rules for which shapes may be drawn, models of how its transistors behave, and usually a library of ready-made logic building blocks, described in files the design software can read.

All three levels

Beginner. The rulebook a chip factory gives designers. It says which shapes are allowed, how its switches behave, and which ready-made parts come with it.

Novice. The package a chip factory (foundry) gives designers for one manufacturing process: the rules for which shapes may be drawn, models of how its transistors behave, and usually a library of ready-made logic building blocks, described in files the design software can read.

Expert. Contains rule files for design-rule and layout-versus-schematic checks written for particular checking tools, SPICE transistor models, technology files, and cell libraries in LEF (shapes) and Liberty (timing and power) form. Commercial PDKs are covered by non-disclosure agreements, and some files are encrypted so only approved tools can read them. Open PDKs (SKY130, GF180MCU, IHP SG13G2) publish everything, currently as preview releases.

PDU (power distribution unit)

The equipment that splits the UPS output into circuits for each rack, often with breakers and meters. Floor-standing PDUs may include a step-down transformer; rack PDUs are the power strips inside the rack.

All three levels

Beginner. A big power strip that shares electricity out to the racks and the computers inside them.

Novice. The equipment that splits the UPS output into circuits for each rack, often with breakers and meters. Floor-standing PDUs may include a step-down transformer; rack PDUs are the power strips inside the rack.

Expert. Losses are mostly in any internal transformer and in cabling; PDU output is one of the standard points for measuring PUE’s IT energy. In ORv3-style racks the rack PDU is replaced by power shelves feeding a DC busbar.

Peak throughput

The theoretical maximum rate of operations, usually in trillions of operations per second (TFLOP/s or TOPS). It is calculated, not measured: number of math units × operations per clock tick × clock rate.

All three levels

Beginner. The fastest a chip could possibly do math, if every part of it were busy every moment. It’s the big number printed on the box.

Novice. The theoretical maximum rate of operations, usually in trillions of operations per second (TFLOP/s or TOPS). It is calculated, not measured: number of math units × operations per clock tick × clock rate.

Expert. Computed at the boost clock for one precision, often with structured sparsity doubling it. It assumes every multiply-accumulate unit issues every cycle, so it ignores memory stalls, data movement, non-matrix work and power or thermal throttling.

Pellicle

A transparent film on a frame glued over one side of a photomask (the stencil for one chip layer). Dust that lands on it sits far enough from the pattern to be out of focus, so it doesn’t print onto the wafer.

All three levels

Beginner. A thin, clear film held just above a stencil’s pattern. Dust lands on the film, where it is blurry and doesn’t print.

Novice. A transparent film on a frame glued over one side of a photomask (the stencil for one chip layer). Dust that lands on it sits far enough from the pattern to be out of focus, so it doesn’t print onto the wafer.

Expert. Mounted after the mask’s final clean and inspection. An EUV pellicle is much harder to build than a deep-UV one: EUV masks reflect rather than transmit, so the 13.5 nm light crosses the pellicle twice, and the film has to let enough of it through while surviving the heat it absorbs.

Performance model

Any model that predicts a chip’s speed before it is built: formulas in a spreadsheet, simulators such as gem5 that run real programs on a software imitation of the chip, transaction-level models in SystemC, and later the hardware description itself, simulated or loaded onto an FPGA (a reprogrammable chip).

All three levels

Beginner. A computer program or math that predicts how fast a chip will be before it is built.

Novice. Any model that predicts a chip’s speed before it is built: formulas in a spreadsheet, simulators such as gem5 that run real programs on a software imitation of the chip, transaction-level models in SystemC, and later the hardware description itself, simulated or loaded onto an FPGA (a reprogrammable chip).

Expert. A ladder trading speed for fidelity. Formulas prune the design space, simulators rank the survivors on real workloads, and RTL measurements calibrate the simulators. Until it is calibrated, a model’s output is a hypothesis; track its error against RTL as blocks arrive.

Performance per watt

Throughput divided by power, for example tokens per second per watt, which equals tokens per joule. It changes a lot depending on whose power you count: just the chip, the whole server, or the building too.

All three levels

Beginner. How much useful work a chip does for each bit of electricity it uses. Higher is better, because electricity and cooling cost money.

Novice. Throughput divided by power, for example tokens per second per watt, which equals tokens per joule. It changes a lot depending on whose power you count: just the chip, the whole server, or the building too.

Expert. Specify the numerator (peak or delivered, which precision) and the denominator boundary (chip TDP, measured chip, wall power of the node, including switches, × PUE). MLPerf Power measures full-system wall power and reports samples per joule.

PFC (priority flow control)

Priority-based flow control: when a switch’s buffer for one traffic class passes a threshold, it sends a pause frame to the device feeding it, which stops sending that class until told to resume. It keeps packets from being dropped, at the cost of pausing traffic that may not be causing the problem.

All three levels

Beginner. A signal that tells the switch sending data to pause for a moment, because the next switch is nearly full.

Novice. Priority-based flow control: when a switch’s buffer for one traffic class passes a threshold, it sends a pause frame to the device feeding it, which stops sending that class until told to resume. It keeps packets from being dropped, at the cost of pausing traffic that may not be causing the problem.

Expert. Per-priority XOFF/XON pause (IEEE 802.1Q). Needs per-port headroom proportional to cable length × rate; pauses propagate hop by hop and cause head-of-line blocking, congestion spreading, pause storms and, with cyclic buffer dependencies, deadlock.

Phase-change memory (PCM)

A non-volatile memory whose cell material switches between an amorphous (high-resistance) and crystalline (low-resistance) state when heated by current pulses. Partial states give intermediate conductances that can store analog weights.

All three levels

Beginner. A memory cell made of a material that can be glassy or crystal. The two forms let electricity through in different amounts.

Novice. A non-volatile memory whose cell material switches between an amorphous (high-resistance) and crystalline (low-resistance) state when heated by current pulses. Partial states give intermediate conductances that can store analog weights.

Expert. Conductance drifts after programming roughly as G(t)=G(t0) (t/t0)−νG(t) = G(t_0)\,(t/t_0)^{-\nu}, with ν\nu varying by state and device, so analog PCM chips compensate globally and use two devices per weight to raise SNR.

Photolithography

The patterning step repeated for every layer of a chip: coat the wafer with photoresist (a light-sensitive coating), shine light through the mask so the pattern lands on the resist, wash away the exposed or unexposed resist, then etch or add material where the resist has gone, and strip what’s left.

All three levels

Beginner. Printing a pattern onto a chip with light. Light shines through a stencil onto a coating that changes where the light hits it.

Novice. The patterning step repeated for every layer of a chip: coat the wafer with photoresist (a light-sensitive coating), shine light through the mask so the pattern lands on the resist, wash away the exposed or unexposed resist, then etch or add material where the resist has gone, and strip what’s left.

Expert. The smallest printable feature scales as k1λ/NAk_1 \lambda / \mathrm{NA}: wavelength λ\lambda divided by the lens’s numerical aperture NA, times a process factor k1k_1. Depth of focus shrinks faster, as λ/NA2\lambda / \mathrm{NA}^2. Production machines (scanners) move mask and wafer together past a slit of light, then step to the next chip-sized field on the wafer.

Photomask

A quartz plate carrying the pattern of one layer, in chrome (or, for extreme-ultraviolet light, as a mirror pattern). A chip needs a full mask set, at least one per patterned layer, used one after another.

All three levels

Beginner. A glass plate with a pattern on it. It works like a stencil to print one layer of a chip.

Novice. A quartz plate carrying the pattern of one layer, in chrome (or, for extreme-ultraviolet light, as a mirror pattern). A chip needs a full mask set, at least one per patterned layer, used one after another.

Expert. Written by electron-beam or laser writers from fractured data, measured for feature size and placement, inspected, repaired and covered with a pellicle. Multiple patterning splits one layer across several masks, a large part of why mask sets cost so much at advanced nodes.

Photoresist

A polymer film spun onto the wafer before each lithography step. After exposure through a mask and development, the remaining resist protects what’s under it from the following etch or implant, then is stripped.

All three levels

Beginner. A light-sensitive coating. Where light hits it, it changes so it can be washed away, leaving a stencil on the wafer.

Novice. A polymer film spun onto the wafer before each lithography step. After exposure through a mask and development, the remaining resist protects what’s under it from the following etch or implant, then is stripped.

Expert. Deep-UV and EUV resists are chemically amplified: exposure turns a photo-acid generator into acid, which during the post-exposure bake catalyzes many deprotection reactions. Acid diffusion and photon counts set a floor on line-edge roughness.

PHY

The physical-layer block of an interface such as PCIe, USB, Ethernet or DDR memory: the analog circuits that drive and receive the electrical signals on the wires, recover their timing and calibrate the link. It comes as finished layout for one manufacturing process.

All three levels

Beginner. The part of a chip that turns digital data into the actual electrical signals sent over a wire, and back.

Novice. The physical-layer block of an interface such as PCIe, USB, Ethernet or DDR memory: the analog circuits that drive and receive the electrical signals on the wires, recover their timing and calibrate the link. It comes as finished layout for one manufacturing process.

Expert. Hard IP with its own power supplies, electrostatic-discharge protection and package-pin requirements. It sits at the die edge, so the number and type of PHYs constrain the chip’s shape, the package and the floorplan from the spec stage onward.

Physical failure analysis (PFA)

Finding and photographing the actual flaw in a failed chip: removing layers one at a time, cutting a cross-section through the suspect spot with a focused ion beam, and imaging it with an electron microscope.

All three levels

Beginner. Opening up a failed chip and looking inside with powerful microscopes to find the actual flaw.

Novice. Finding and photographing the actual flaw in a failed chip: removing layers one at a time, cutting a cross-section through the suspect spot with a focused ion beam, and imaging it with an electron microscope.

Expert. Destructive and slow, so only a few dies get it; scan diagnosis and volume analysis decide which dies, and where to cut. Its confirmation closes the yield-learning loop.

Pin access

Whether the wiring step can connect a wire to each of a cell’s pins (its input and output contact points) without breaking a manufacturing rule. Tightly packed cells can block each other’s pins.

All three levels

Beginner. Whether wires can actually reach each connection point on a part without breaking a rule.

Novice. Whether the wiring step can connect a wire to each of a cell’s pins (its input and output contact points) without breaking a manufacturing rule. Tightly packed cells can block each other’s pins.

Expert. At advanced nodes, tiny low-metal pins, few tracks per cell and complex via rules mean two legal neighbors can still leave a pin unreachable. Fixes include padding, pin-access-aware detailed placement, and placement refinement during routing.

Pin assignment

Choosing exactly where on the edge of a chip or block each signal enters or leaves, and on which metal layer. The usual aim is to put each pin close to the logic that uses it.

All three levels

Beginner. Deciding where on a block’s edge each wire comes in or goes out.

Novice. Choosing exactly where on the edge of a chip or block each signal enters or leaves, and on which metal layer. The usual aim is to put each pin close to the logic that uses it.

Expert. Pins go on the routing-track grid with minimum spacing and keep-outs near corners, often under constraints that tie groups of pins to an edge, part of an edge, or an order. In a design split into blocks, block pins are assigned from the top-level view so neighboring blocks agree.

Pinch-off

When VDSV_{\mathrm{DS}} reaches VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}}, the gate-to-drain voltage falls to the threshold and the channel vanishes at the drain end. Carriers still cross the short depleted gap, but the current no longer grows with VDSV_{\mathrm{DS}}.

All three levels

Beginner. When the path of electrons under the gate gets squeezed to nothing at the far end, so more push stops adding current.

Novice. When VDSV_{\mathrm{DS}} reaches VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}}, the gate-to-drain voltage falls to the threshold and the channel vanishes at the drain end. Carriers still cross the short depleted gap, but the current no longer grows with VDSV_{\mathrm{DS}}.

Expert. Occurs at VGD=VtV_{\mathrm{GD}} = V_{\mathrm{t}}, i.e. VDS=VGS−VtV_{\mathrm{DS}} = V_{\mathrm{GS}} - V_{\mathrm{t}} in the long-channel model. The pinch-off point moves toward the source as VDSV_{\mathrm{DS}} rises, which shortens the effective channel: that is channel-length modulation.

Pipeline bubble

An empty slot moving through a pipeline, where a stage has no useful work. In a CPU it comes from a stall or a flushed instruction; in pipeline-parallel training it is the idle time while the pipeline fills at the start of a step and drains at the end, which shrinks when there are many more microbatches than stages.

All three levels

Beginner. Waiting time on an assembly line: at the start, later stations have nothing to do yet, and at the end, early ones have finished.

Novice. An empty slot moving through a pipeline, where a stage has no useful work. In a CPU it comes from a stall or a flushed instruction; in pipeline-parallel training it is the idle time while the pipeline fills at the start of a step and drains at the end, which shrinks when there are many more microbatches than stages.

Expert. In a CPU pipeline, a no-op inserted by a stall or a squash, costing one cycle of CPI per bubble. In pipeline-parallel training, (p−1)(tf+tb)(p-1)(t_{\mathrm{f}} + t_{\mathrm{b}}) of idle time per step with flushes, a fraction (p−1)/m(p-1)/m of ideal compute time; interleaving vv model chunks per device divides it by vv at vv times the point-to-point traffic.

Pipeline hazard

Something that stops the next instruction from moving down the pipeline: it needs a result that isn’t ready yet (data hazard), it follows a branch whose direction isn’t known yet (control hazard), or it needs a unit another instruction is using (structural hazard).

All three levels

Beginner. A hiccup on the assembly line: a step can’t go ahead because it needs a result that isn’t ready, or the line went the wrong way.

Novice. Something that stops the next instruction from moving down the pipeline: it needs a result that isn’t ready yet (data hazard), it follows a branch whose direction isn’t known yet (control hazard), or it needs a unit another instruction is using (structural hazard).

Expert. A hazard costs roughly the number of stages between where a value is produced and where it is needed, so deeper pipelines lose more cycles per hazard. Forwarding results early, predicting branches and reordering instructions hide part of the cost.

Pipeline parallelism

Placing different layers of a model on different processors. Data flows from one stage to the next, and several inputs can be in flight at once, one per stage.

All three levels

Beginner. Splitting a model into stages, like an assembly line, with each chip (or part of a chip) doing a few steps and passing the work on.

Novice. Placing different layers of a model on different processors. Data flows from one stage to the next, and several inputs can be in flight at once, one per stage.

Expert. Each stage holds only its layers’ weights, so capacity scales with stage count; only activations cross stage boundaries. Costs: pipeline fill and drain, latency per hop, and load imbalance between stages.

Pipelining

Splitting a long piece of logic into shorter stages separated by registers, like an assembly line. Each stage does part of the work per clock tick, so the clock can run faster and a new result can finish every tick, even though each single operation takes several ticks.

All three levels

Beginner. Splitting a job into steps like an assembly line, so several jobs are in progress at once and a new one can finish every tick.

Novice. Splitting a long piece of logic into shorter stages separated by registers, like an assembly line. Each stage does part of the work per clock tick, so the clock can run faster and a new result can finish every tick, even though each single operation takes several ticks.

Expert. Clock period≈logic delay÷number of stages+per-stage overhead\text{Clock period} \approx \text{logic delay} \div \text{number of stages} + \text{per-stage overhead} (the register’s setup and clock-to-output times, plus clock skew and jitter). More stages also mean more cycles lost per hazard and more flip-flops, clock load and power, so performance peaks at a finite depth.

Placement blockage

An area where the placer may not put cells. A hard blockage allows none. A soft blockage is kept empty during the main placement, but cells added later for timing or for the clock may use it. A partial blockage lets cells fill only a set percentage of the area.

All three levels

Beginner. An area marked “no parts here” (or “only a few parts here”) to keep space free for wires.

Novice. An area where the placer may not put cells. A hard blockage allows none. A soft blockage is kept empty during the main placement, but cells added later for timing or for the clock may use it. A partial blockage lets cells fill only a set percentage of the area.

Expert. DEF BLOCKAGES with PLACEMENT (hard), PLACEMENT + SOFT, or PLACEMENT + PARTIAL maxDensity. Typical uses: macro channels, notches, areas over sensitive analog, and partial blockages to thin out predicted congestion hot spots.

Planar transistor

A MOSFET whose channel is a thin layer at the flat top surface of the silicon, with the gate lying on top. Its width is drawn freely in the layout, but the gate only acts from one side.

All three levels

Beginner. The flat kind of transistor used for most of chip history. Its gate sits on top of the channel, like a lid.

Novice. A MOSFET whose channel is a thin layer at the flat top surface of the silicon, with the gate lying on top. Its width is drawn freely in the layout, but the gate only acts from one side.

Expert. Single-gate MOSFET on a bulk (or SOI) wafer. In bulk, the gated depth is set by doping, so short-channel control costs channel doping, mobility and random dopant fluctuation. Planar bulk led until Intel’s 22 nm FinFETs in 2011; planar FD-SOI is still made for low-power parts.

Plasma (dry) etch

Etching with ions and reactive gas fragments in a low-pressure plasma. Unlike a liquid etch, which eats in every direction, ions accelerated toward the wafer etch mostly downward, so the shape in the resist transfers with vertical walls.

All three levels

Beginner. Carving away material with a glowing, electrically charged gas that cuts straight down instead of eating sideways.

Novice. Etching with ions and reactive gas fragments in a low-pressure plasma. Unlike a liquid etch, which eats in every direction, ions accelerated toward the wafer etch mostly downward, so the shape in the resist transfers with vertical walls.

Expert. Reactive ion etching balances physical sputtering (directional, unselective) against chemical attack by radicals (selective, isotropic). Pressure, ion energy and gas mix set anisotropy, selectivity to the layer underneath, and damage.

PLL (phase-locked loop)

A phase-locked loop: a circuit that takes a slower reference clock from outside the chip and generates the faster clock the logic runs on, locked in step with the reference.

All three levels

Beginner. The part that makes the chip’s clock, the steady beat that keeps all the other parts in step.

Novice. A phase-locked loop: a circuit that takes a slower reference clock from outside the chip and generates the faster clock the logic runs on, locked in step with the reference.

Expert. A mixed analog and digital macro, usually bought or reused as IP. Its jitter is sensitive to supply and substrate noise, so floorplans keep it away from busy logic and give it a short, clean path to its reference-clock pin.

Pluggable optical module

A hot-swappable transceiver in a standard form factor (such as QSFP-DD or OSFP) that slots into a cage on the switch’s front panel. It contains the electronics, lasers and light detectors for one port.

All three levels

Beginner. A small box that plugs into the front of a switch and turns electrical signals into light and back.

Novice. A hot-swappable transceiver in a standard form factor (such as QSFP-DD or OSFP) that slots into a cage on the switch’s front panel. It contains the electronics, lasers and light detectors for one port.

Expert. Contains a host-side retimer/DSP, laser drivers, modulators or directly modulated lasers, photodiodes and TIAs, plus management. Standard electrical interface (AUI) to the host, standard optical interface (an IEEE PMD such as DR or FR) to the fiber. A field-replaceable unit.

PMA (premarket approval)

The FDA’s strictest review, used for the highest-risk (Class III) devices. The maker must show evidence that the device is safe and effective before it can be sold.

All three levels

Beginner. The strictest U.S. approval for medical devices. It is used for the riskiest ones.

Novice. The FDA’s strictest review, used for the highest-risk (Class III) devices. The maker must show evidence that the device is safe and effective before it can be sold.

Expert. After approval, any change affecting safety or effectiveness, explicitly including changes to circuits, components, physical layout and manufacturing facility, requires an approved PMA supplement before it is implemented.

PMHF

Probabilistic Metric for random Hardware Failures: the expected rate at which random faults violate a safety goal, in FIT (failures per billion hours of operation). Targets: under 100 FIT for ASIL B and C, under 10 FIT for D.

All three levels

Beginner. A number for how often random breakdowns in a chip could lead to danger.

Novice. Probabilistic Metric for random Hardware Failures: the expected rate at which random faults violate a safety goal, in FIT (failures per billion hours of operation). Targets: under 100 FIT for ASIL B and C, under 10 FIT for D.

Expert. Budgeted across the whole item (the vehicle function), so an SoC typically gets a fraction of it; silicon base failure rates, package failures and transient soft-error rates all count.

PMOS

A p-channel MOSFET: p+ source and drain in an n-type well. It turns on when the gate is pulled sufficiently below the source, which usually sits at the supply voltage, so a low gate turns it on.

All three levels

Beginner. The kind of transistor switch that turns on when its gate voltage is low. It is the mirror image of NMOS.

Novice. A p-channel MOSFET: p+ source and drain in an n-type well. It turns on when the gate is pulled sufficiently below the source, which usually sits at the supply voltage, so a low gate turns it on.

Expert. Hole channel with lower mobility, so it is made wider than the matching NMOS for equal drive. Passes a strong 1 and a degraded 0, so it forms the pull-up half of a CMOS gate.

POCV and LVF

A way of handling manufacturing variation in timing. Each building block gets its own measured spread of delay, and the tool combines the spreads along a path statistically instead of assuming every gate is at its worst at the same time. LVF (Liberty Variation Format) is the format that stores those spreads in the cell library.

All three levels

Beginner. Giving every part its own small wobble in speed. The wobbles are added up as odds, not all at their worst at once.

Novice. A way of handling manufacturing variation in timing. Each building block gets its own measured spread of delay, and the tool combines the spreads along a path statistically instead of assuming every gate is at its worst at the same time. LVF (Liberty Variation Format) is the format that stores those spreads in the cell library.

Expert. Parametric OCV: each cell arc has a mean delay and a standard deviation (sigma). Along a path the means add and the sigmas combine by root-sum-square, and the check is made at a chosen number of sigmas. LVF tables give sigma for delay, output transition and setup/hold constraints as functions of input slew and output load.

PODEM

Path-Oriented Decision Making (Goel, 1981), a test-search method that only ever guesses values for the circuit’s inputs, simulates what follows, and undoes the latest guess when the fault can no longer be made visible.

All three levels

Beginner. A faster way to search for a test pattern, by guessing only the chip’s inputs and checking each guess.

Novice. Path-Oriented Decision Making (Goel, 1981), a test-search method that only ever guesses values for the circuit’s inputs, simulates what follows, and undoes the latest guess when the fault can no longer be made visible.

Expert. The decision tree is over primary inputs, not all lines. Each objective is traced back to one input using controllability estimates, and a check that the fault effect can still reach an output prunes dead branches early. The basis of most later ATPG engines.

Point of load (PoL)

The final voltage conversion stage, placed as close as possible to the chip so the high current at low voltage travels only a short distance.

All three levels

Beginner. The last power converter, sitting right next to the chip it feeds.

Novice. The final voltage conversion stage, placed as close as possible to the chip so the high current at low voltage travels only a short distance.

Expert. The last stage of the conversion chain (for example 12 V→0.8 V). Its distance to the die sets lateral PDN resistance and inductance; vertical power delivery moves it under the package to shorten that path further.

Post-silicon validation

Running manufactured chips under real conditions to find mistakes in the design itself that testing in simulation missed. Manufacturing test is different: it looks for physical flaws in each individual copy.

All three levels

Beginner. Running real programs on the real chip to find design mistakes that earlier tests on a computer missed.

Novice. Running manufactured chips under real conditions to find mistakes in the design itself that testing in simulation missed. Manufacturing test is different: it looks for physical flaws in each individual copy.

Expert. Four steps: detect, localize, find the root cause, fix. Silicon runs orders of magnitude faster than simulation, but engineers can see only a few internal signals, so localizing the bug takes most of the effort and depends on debug hardware built in at design time: trace buffers, triggers, scan dumps and clock control.

Power budget

A table that splits the chip’s total power allowance among its blocks and its operating modes (peak, typical, idle, standby), for a stated workload and temperature, with some held back in reserve.

All three levels

Beginner. A plan that splits the chip’s total power allowance among its parts, like a household budget.

Novice. A table that splits the chip’s total power allowance among its blocks and its operating modes (peak, typical, idle, standby), for a stated workload and temperature, with some held back in reserve.

Expert. Split into switching and leakage power for each block, mode and supply, at a stated temperature and manufacturing corner. It later feeds the power-intent file (UPF) that says which blocks can be switched off, the sizing of the on-chip power wiring, and the package choice. It is re-estimated once the design code, and later the gate-level netlist, exist.

Power conversion chain

The sequence from the utility’s high voltage down to about 1 volt at the chip: transformers, a backup system (UPS), distribution units (PDUs), the server’s power supply (PSU) and the voltage regulators beside each chip. Each stage’s efficiency multiplies with the others.

All three levels

Beginner. The steps that bring electricity from the power lines down to the gentle 1 volt a chip needs. A little is lost as heat at each step.

Novice. The sequence from the utility’s high voltage down to about 1 volt at the chip: transformers, a backup system (UPS), distribution units (PDUs), the server’s power supply (PSU) and the voltage regulators beside each chip. Each stage’s efficiency multiplies with the others.

Expert. Typically medium-voltage AC → low-voltage AC (transformer) → UPS → PDU → PSU (AC to a 48 V-class DC bus in modern racks) → intermediate bus converter → multiphase point-of-load regulators. End-to-end efficiency is the product of stage efficiencies, and losses upstream scale with every watt lost downstream.

Power delivery network (PDN)

Everything that carries the supply from the voltage regulator on the circuit board, through the chip’s package and into the chip, down to every logic cell, plus the capacitors along the way that steady the voltage.

All three levels

Beginner. All the wiring that carries electricity from outside the chip to every tiny switch inside it.

Novice. Everything that carries the supply from the voltage regulator on the circuit board, through the chip’s package and into the chip, down to every logic cell, plus the capacitors along the way that steady the voltage.

Expert. Treated on chip as a large mesh of resistors (with capacitance, and sometimes inductance, added for time-varying analysis), fed by voltage sources at the bumps and loaded by the current each cell draws. It is judged by voltage drop, wire wear-out (electromigration), how steady the supply stays when demand jumps, and how much routing space it takes from signal wires.

Power density

Power divided by area, in watts per square centimeter or per square millimeter. Cooling limits how high it can go.

All three levels

Beginner. How much heat a patch of chip gives off for its size. Too much in one spot and the chip overheats.

Novice. Power divided by area, in watts per square centimeter or per square millimeter. Cooling limits how high it can go.

Expert. Dennard scaling held it constant; once voltage stopped scaling, power density rose with transistor density at a given activity and frequency. It is what a power budget really limits, and why clock rates stalled and dark silicon appeared.

Power domain

A region of the chip with its own power supply that can be switched off (power gating) or run at its own voltage. Signals leaving a region that may be off need isolation cells to hold them at a safe value; signals between different voltages need level shifters.

All three levels

Beginner. A part of the chip that can be powered or slowed down separately from the rest.

Novice. A region of the chip with its own power supply that can be switched off (power gating) or run at its own voltage. Signals leaving a region that may be off need isolation cells to hold them at a safe value; signals between different voltages need level shifters.

Expert. Captured as power intent in a UPF file (IEEE 1801) and laid out as a voltage area. Each domain adds isolation, level shifters, retention registers, power switches and more power states to verify, so the number of domains is an architectural cost.

power gating

Disconnecting an idle block of the chip from its power supply through a transistor switch, so it stops wasting power on leakage. The block loses whatever it was storing unless that is saved first.

All three levels

Beginner. Cutting off power to a part of the chip that isn’t being used, like switching off the lights in an empty room.

Novice. Disconnecting an idle block of the chip from its power supply through a transistor switch, so it stops wasting power on leakage. The block loses whatever it was storing unless that is saved first.

Expert. Header (supply-side) or footer (ground-side) switch networks between the global rail and a block’s virtual rail, controlled by a power-management unit. It requires isolation on the block’s outputs, retention or state reload, control of in-rush current at wake-up, and power-aware verification.

Power ring

A loop of wide supply and ground wires around the edge of the chip’s logic area (the core) or around a large block such as a memory, so power can be fed in from every side.

All three levels

Beginner. A loop of thick power wires around a block or the whole chip, so electricity can come in from every side.

Novice. A loop of wide supply and ground wires around the edge of the chip’s logic area (the core) or around a large block such as a memory, so power can be fed in from every side.

Expert. A VDD/VSS pair on two upper layers, one layer for the horizontal sides and one for the vertical sides. Most useful when power enters from pads on the chip’s edge, and around memories and switched blocks. When bumps cover the whole chip face, they feed the top straps directly. Ring width is set by the current it must carry without wearing out.

Power strap (stripe)

A wide supply or ground wire on one of the chip’s upper wiring layers, repeated at a regular spacing. Straps on neighboring layers run at right angles to each other and are joined where they cross, forming a mesh.

All three levels

Beginner. A wide power wire that runs across the chip on the top layers, like a highway for electricity.

Novice. A wide supply or ground wire on one of the chip’s upper wiring layers, repeated at a regular spacing. Straps on neighboring layers run at right angles to each other and are joined where they cross, forming a mesh.

Expert. Described by layer, width, pitch (the repeat distance), offset and the spacing between the VDD and VSS straps in each pitch. Wider or closer straps lower resistance and wear-out risk but take routing tracks from signals, so the pitch is often varied by region.

Power switch (header/footer)

A transistor used as an on/off switch between the chip’s main supply and a block’s own local supply wire, so the block can be powered off. Many small switch cells are spread across the block.

All three levels

Beginner. A big switch on the chip that can cut the power to one part, like a light switch for a whole room.

Novice. A transistor used as an on/off switch between the chip’s main supply and a block’s own local supply wire, so the block can be powered off. Many small switch cells are spread across the block.

Expert. Either a PMOS “header” on the supply side or an NMOS “footer” on the ground side. Sized for its on-resistance, which adds voltage drop whenever it is on, against its leakage when off and its area. The on/off signal can reach every switch at once (star) or pass from switch to switch (daisy chain), which staggers turn-on to limit rush current.

Power wall

The point where a chip’s power, and so its heat, became the limit on its speed. Once supply voltage stopped falling with each new manufacturing generation, faster clocks meant more heat than packages could remove, so clock rates flattened in the mid-2000s.

All three levels

Beginner. The limit on how much heat a chip can give off. Around 2005 it stopped chips from simply running their clocks faster.

Novice. The point where a chip’s power, and so its heat, became the limit on its speed. Once supply voltage stopped falling with each new manufacturing generation, faster clocks meant more heat than packages could remove, so clock rates flattened in the mid-2000s.

Expert. With V roughly fixed, P ≈ αCV²f per gate stops falling as fast as gate density rises, and in the DVFS range f ∝ V makes per-core power grow roughly as f³. Designs became power-limited (TDP-bound) rather than area-limited.

Power-aware simulation

Simulation that also models parts of the chip being switched off and back on to save power, following the chip’s power plan, so tests can check that sleeping parts don’t disturb the rest and wake up correctly.

All three levels

Beginner. Testing what happens when parts of the chip switch off and back on to save battery.

Novice. Simulation that also models parts of the chip being switched off and back on to save power, following the chip’s power plan, so tests can check that sleeping parts don’t disturb the rest and wake up correctly.

Expert. Reads the UPF power intent, scrambles the state of powered-down regions, and models isolation, retention and level-shifter cells. It has to exercise the whole power sequence and the control logic around it. Slower than plain RTL simulation, and the unknown values it spreads make failures slow to trace.

PPA (power, performance, area)

Power, performance (speed, usually the clock frequency) and area: the three main measures of a design’s quality. Improving one usually costs another.

All three levels

Beginner. The three things every chip trades off: how much power it uses, how fast it runs and how big it is.

Novice. Power, performance (speed, usually the clock frequency) and area: the three main measures of a design’s quality. Improving one usually costs another.

Expert. The joint goal of the whole implementation flow. In synthesis you steer it with the clock target, effort settings, which cells you allow (fast and leaky or slow and frugal) and transformations such as retiming, clock gating and resource sharing. Synthesis PPA is an estimate; the numbers after routing are what ships.

PRD / MRD (product and market requirements)

Business documents written before engineering starts. A product requirements document (PRD) says what the product must do, without saying how. When a company’s marketing department writes it, it may be called a market requirements document (MRD); some companies keep a separate MRD for customers, competitors and price.

All three levels

Beginner. The document where the business side writes down what customers want from the product and why.

Novice. Business documents written before engineering starts. A product requirements document (PRD) says what the product must do, without saying how. When a company’s marketing department writes it, it may be called a market requirements document (MRD); some companies keep a separate MRD for customers, competitors and price.

Expert. The upstream source for chip requirements. Each chip requirement should name the PRD item it serves, so that when the PRD changes, a script can list every chip requirement, test and budget affected.

Precharge

Charging the bit lines to a fixed voltage before an access. SRAM precharges both lines to the supply; DRAM precharges to half the supply. The read then looks at how the line moves away from that level.

All three levels

Beginner. Filling a memory’s wires with charge to a set starting level before each read, like zeroing a kitchen scale.

Novice. Charging the bit lines to a fixed voltage before an access. SRAM precharges both lines to the supply; DRAM precharges to half the supply. The read then looks at how the line moves away from that level.

Expert. Also equalizes the two lines of a differential pair so the sense amplifier sees only the cell’s signal. In DRAM, the time to restore VDD/2V_{\mathrm{DD}}/2 after a row closes is the tRP timing parameter.

Precise exception

An error or interrupt handled as if the processor stopped exactly between two instructions: everything before has completed and nothing after has had any effect. Out-of-order cores get this by retiring in order.

All three levels

Beginner. When something goes wrong in one step, the processor stops so that every earlier step is finished and no later step has changed anything.

Novice. An error or interrupt handled as if the processor stopped exactly between two instructions: everything before has completed and nothing after has had any effect. Out-of-order cores get this by retiring in order.

Expert. Required for demand paging and restartable faults. Implemented by recording exceptions in the ROB and acting on them only at commit, flushing all younger instructions.

Preferred routing direction

The one direction, horizontal or vertical, that wires on a given layer run in. Neighboring layers alternate, so a wire that needs to turn steps up or down a layer through a via. A short sideways piece on a layer (a “wrong-way” step) may be allowed at a high cost or forbidden outright.

All three levels

Beginner. Each wiring layer runs mostly one way, east–west or north–south, and neighboring layers alternate, like stacked one-way streets.

Novice. The one direction, horizontal or vertical, that wires on a given layer run in. Neighboring layers alternate, so a wire that needs to turn steps up or down a layer through a via. A short sideways piece on a layer (a “wrong-way” step) may be allowed at a high cost or forbidden outright.

Expert. Declared per layer in tech LEF (DIRECTION HORIZONTAL or VERTICAL). On the lowest layers of advanced nodes, the patterning process prints rows of parallel lines, so those layers are effectively one-direction only. Jogs and wrong-way stubs there trigger line-end, minimum-area and mask-coloring rules, so routers keep them for reaching pins.

Prefetching

Bringing lines into the cache ahead of demand. Hardware prefetchers spot patterns, such as a stream of consecutive lines or a fixed stride, and fetch the next ones; software can also issue prefetch instructions.

All three levels

Beginner. Fetching data before the program asks for it, by guessing what it will need next.

Novice. Bringing lines into the cache ahead of demand. Hardware prefetchers spot patterns, such as a stream of consecutive lines or a fixed stride, and fetch the next ones; software can also issue prefetch instructions.

Expert. Judged by accuracy (useful ÷ issued), coverage (misses removed ÷ original misses) and timeliness (arriving before use but not so early it is evicted). Distance must cover latency ÷ time per access; useless prefetches cost bandwidth and pollute the cache.

Prefill

The inference phase that runs the prompt through the model. All prompt tokens are processed together, so the matrix multiplications are large and the chip’s math units stay busy.

All three levels

Beginner. The first part of answering: the model reads your whole question at once.

Novice. The inference phase that runs the prompt through the model. All prompt tokens are processed together, so the matrix multiplications are large and the chip’s math units stay busy.

Expert. One forward pass over B×LinputB \times L_{\mathrm{input}} tokens that also fills the KV cache. Weight GEMMs have MM = tokens in the batch, so it is usually compute-bound; it sets the time to first token.

Process node

A chip factory’s named manufacturing generation, such as “7 nm class”. The number no longer matches any physical size on the chip; it labels a generation with its own density, speed, power, cost and design rules.

All three levels

Beginner. The generation of chipmaking technology a chip is built with. Newer generations pack more tiny switches into the same space.

Novice. A chip factory’s named manufacturing generation, such as “7 nm class”. The number no longer matches any physical size on the chip; it labels a generation with its own density, speed, power, cost and design rules.

Expert. Choosing one fixes the building blocks on offer (the library of basic logic cells and the blocks that can be bought in), wafer and mask cost, defect density and factory capacity. Newer is not always better: older processes are cheaper, better understood, easier for analog circuits and long-life parts, and often the only ones where the needed bought-in blocks already exist.

Processing element (PE)

The repeated cell of a spatial or systolic array. In an AI accelerator it is usually one multiply-accumulate unit plus a few registers that hold a stored value and the values passing through.

All three levels

Beginner. One of the small, identical calculator cells in a grid. Each one does a tiny piece of the work and hands numbers to its neighbors.

Novice. The repeated cell of a spatial or systolic array. In an AI accelerator it is usually one multiply-accumulate unit plus a few registers that hold a stored value and the values passing through.

Expert. A MAC datapath with operand and partial-sum registers, sometimes a small register file (0.5 kB per PE in Eyeriss) and local control. Its contents and its links to neighbors define the dataflow; its pipeline registers set the array’s clock rate.

Program counter (PC)

The register that holds the memory address of the instruction being fetched. It normally moves on to the next instruction (4 bytes further for a 32-bit instruction); a branch or jump loads a new address instead.

All three levels

Beginner. A small counter in the chip that remembers where in the program it is, like a bookmark.

Novice. The register that holds the memory address of the instruction being fetched. It normally moves on to the next instruction (4 bytes further for a 32-bit instruction); a branch or jump loads a new address instead.

Expert. Architectural state; in a pipeline the fetch PC runs ahead of the instructions in flight, and every stage carries its own instruction’s PC for branch targets and precise exceptions.

Propagated clock

A clock whose arrival time at each flip-flop is calculated from the real delays of the buffers and wires in the built network, instead of being assumed.

All three levels

Beginner. The clock as it really behaves after the tree is built, with real travel delays.

Novice. A clock whose arrival time at each flip-flop is calculated from the real delays of the buffers and wires in the built network, instead of being assumed.

Expert. Switched on with set_propagated_clock. Each sink then has its own latency and edge rate, variation derates apply to clock cells, and common-path credit (CPPR) becomes meaningful. All timing work after CTS, including signoff, uses propagated clocks.

Proxy metric

A quick estimate optimized in place of the real result, such as estimated wire length instead of the wire length measured after the slow routing step.

All three levels

Beginner. A quick score that stands in for the real result, like judging a cake by its smell before tasting it.

Novice. A quick estimate optimized in place of the real result, such as estimated wire length instead of the wire length measured after the slow routing step.

Expert. A proxy is only useful while it ranks options in the same order as the real metric. That agreement often weakens among the best options, so a better proxy score can fail to give a better chip.

Pruning

Setting selected weights to zero, usually the smallest in magnitude, then fine-tuning the network so the remaining weights compensate. Done after or during training.

All three levels

Beginner. Setting the least useful numbers in a trained AI model to zero. Then the model practices a bit to make up for them.

Novice. Setting selected weights to zero, usually the smallest in magnitude, then fine-tuning the network so the remaining weights compensate. Done after or during training.

Expert. Magnitude pruning is the baseline; criteria using second-order information (e.g. SparseGPT) allow one-shot pruning of very large models. One-shot or iterative, global or per-layer, with or without retraining; the pattern (unstructured, N:M, block, channel) is a separate choice.

PSU (power supply unit)

An AC-to-DC converter for servers. Its efficiency (output ÷ input) depends on how heavily it is loaded; the 80 PLUS program certifies levels from basic 80% up to Titanium and Ruby.

All three levels

Beginner. The box in a computer that changes wall electricity into the steady kind that electronics need.

Novice. An AC-to-DC converter for servers. Its efficiency (output ÷ input) depends on how heavily it is loaded; the 80 PLUS program certifies levels from basic 80% up to Titanium and Ruby.

Expert. Front ends use power-factor correction then an isolated DC/DC stage. 80 PLUS Titanium (230 V internal redundant) requires 90/94/96/91% at 10/20/50/100% load. Redundant PSUs share load and so often sit near 50% or below; modern AI racks centralize them in power shelves feeding a 48 V-class busbar.

PSU redundancy (N+1, N+N)

N is the number of power supplies needed to carry the load. N+1 adds one spare; N+N (or 2N) doubles them, often so each half can be fed from a separate power feed.

All three levels

Beginner. Having more power supplies than you need, so the server keeps running when one breaks.

Novice. N is the number of power supplies needed to carry the load. N+1 adds one spare; N+N (or 2N) doubles them, often so each half can be fed from a separate power feed.

Expert. Usable capacity is (installed−redundant)×rating(\text{installed} - \text{redundant}) \times \text{rating}. N+1 covers a single PSU failure; N+N covers loss of a whole feed. Running supplies near half load also keeps them near their efficiency peak, but capacity reserved for redundancy is paid for and idle.

PTP (IEEE 1588)

The Precision Time Protocol: devices on a network exchange time-stamped messages to work out how far apart their clocks are and how long messages take, then correct their clocks.

All three levels

Beginner. A way for computers on a network to set their clocks to the same time, very exactly.

Novice. The Precision Time Protocol: devices on a network exchange time-stamped messages to work out how far apart their clocks are and how long messages take, then correct their clocks.

Expert. Accuracy depends on time-stamping in hardware close to the cable, symmetric network paths, and switches that correct for their own delay. The standard supports sub-microsecond accuracy, and sub-nanosecond in a properly designed network.

PUE (power usage effectiveness)

Total power entering the datacenter divided by the power reaching the computers (IT equipment). A PUE of 1.4 means 0.4 W of cooling, power conversion and lighting for every 1 W of computing.

All three levels

Beginner. A score for how much extra electricity a datacenter (a building full of computers) spends on cooling and other needs. 1.0 would mean none at all.

Novice. Total power entering the datacenter divided by the power reaching the computers (IT equipment). A PUE of 1.4 means 0.4 W of cooling, power conversion and lighting for every 1 W of computing.

Expert. An annual average for a facility, not a property of a chip. It multiplies IT energy in a TCO model but says nothing about how efficiently the IT power is used, which is why MLPerf Power leaves it out.

Pull-down network

The NMOS transistors between the output and ground (GND) of a CMOS gate. When they form a conducting path, the output is pulled down to 0.

All three levels

Beginner. The group of switches in a gate that can connect the output to ground, making it a 0.

Novice. The NMOS transistors between the output and ground (GND) of a CMOS gate. When they form a conducting path, the output is pulled down to 0.

Expert. The NMOS network of a static CMOS gate. It conducts exactly when the gate’s output should be 0, so its Boolean function is the complement of the gate’s function. Series depth sets NMOS sizing and fall delay.

Pull-up network

The PMOS transistors between the supply (VDDV_{\mathrm{DD}}) and the output of a CMOS gate. When they form a conducting path, the output is pulled up to 1.

All three levels

Beginner. The group of switches in a gate that can connect the output to power, making it a 1.

Novice. The PMOS transistors between the supply (VDDV_{\mathrm{DD}}) and the output of a CMOS gate. When they form a conducting path, the output is pulled up to 1.

Expert. The PMOS network of a static CMOS gate, the series/parallel dual of the pull-down network. Its worst-case series depth sets PMOS sizing and rise delay.

Pull-up ratio

In a 6T SRAM cell, the strength of a pull-up (PMOS) transistor divided by that of an access transistor. It must be small enough that the access transistor can drag a stored 1 down during a write.

All three levels

Beginner. How strong the part of the cell that holds a 1 is compared with the door used to write into it.

Novice. In a 6T SRAM cell, the strength of a pull-up (PMOS) transistor divided by that of an access transistor. It must be small enough that the access transistor can drag a stored 1 down during a write.

Expert. PR=(W/L)pu/(W/L)acc\mathrm{PR} = (W/L)_{\mathrm{pu}} / (W/L)_{\mathrm{acc}}, usually minimum (about 1) and kept below roughly 1.8–3. A higher PR helps read and hold stability slightly but erodes write margin, especially at slow-NMOS/fast-PMOS corners.

PVT corner

One combination of process (whether the factory made the transistors slow, typical or fast), supply voltage and temperature. Each corner has its own timing data for the building blocks. Signals are usually slowest at slow process and low voltage, and fastest at fast process and high voltage.

All three levels

Beginner. One extreme case the chip must still work in, such as a slow chip on a hot day.

Novice. One combination of process (whether the factory made the transistors slow, typical or fast), supply voltage and temperature. Each corner has its own timing data for the building blocks. Signals are usually slowest at slow process and low voltage, and fastest at fast process and high voltage.

Expert. Named by the speed of the NMOS and PMOS transistors (slow-slow, typical, fast-fast, and the skewed slow-fast and fast-slow), with a voltage and temperature, each with its own Liberty library. Every independent supply multiplies the set. At low supply voltages some processes get slower when cold, not hot, so teams check both temperatures at the slow corner.

QML (Qualified Manufacturers List)

The U.S. system, under the standard MIL-PRF-38535, in which manufacturers get their production and testing approved for military and space parts. Class Q is the military level; class V is the space level.

All three levels

Beginner. A U.S. government list of factories and parts approved for military and space use.

Novice. The U.S. system, under the standard MIL-PRF-38535, in which manufacturers get their production and testing approved for military and space parts. Class Q is the military level; class V is the space level.

Expert. Administered by the Defense Logistics Agency (DLA Land and Maritime); for space microcircuits, DLA, NASA/JPL and the Aerospace Corporation form the qualifying activity. Qualification covers electrical, environmental, life, package and radiation test groups.

Qualification

A fixed set of stress tests on chips from several separate production batches: running them hot for long periods, cycling them between extreme temperatures, and more. Passing usually means zero failures. Customers require it before buying in volume.

All three levels

Beginner. A series of stress tests that show a chip will keep working for years, required before it is sold in volume.

Novice. A fixed set of stress tests on chips from several separate production batches: running them hot for long periods, cycling them between extreme temperatures, and more. Passing usually means zero failures. Customers require it before buying in volume.

Expert. Sample sizes, batch counts and conditions come from the standard the customer requires, such as AEC-Q100 for cars. Changes to the design, process or package can trigger requalification.

Quantization

Converting a tensor from a wide format (usually FP32 or BF16) to a narrow one such as INT8, FP8 or FP4, with a scale factor that maps the tensor’s range onto the format’s range. It introduces rounding error and, if values exceed the range, clipping error.

All three levels

Beginner. Rounding numbers onto a small set of allowed values so they take fewer bits to store and are cheaper to compute with.

Novice. Converting a tensor from a wide format (usually FP32 or BF16) to a narrow one such as INT8, FP8 or FP4, with a scale factor that maps the tensor’s range onto the format’s range. It introduces rounding error and, if values exceed the range, clipping error.

Expert. q=clamp(round(x/s)+z)q = \mathrm{clamp}(\mathrm{round}(x/s) + z) and x^=s(q−z)\hat{x} = s(q - z) for integers, or a scaled cast for small floats. Choices: format, granularity (tensor, channel, block), calibration of ss (max, percentile, MSE), and whether it happens after training (PTQ) or is simulated during training (QAT).

Rack power density

The electrical power drawn by all the equipment in one rack, usually in kilowatts (kW). Since nearly all of it becomes heat, it is also the heat the cooling system must remove from that rack.

All three levels

Beginner. How much electricity one cabinet of computers uses. Almost all of it turns into heat, so it is also how much heat must be removed.

Novice. The electrical power drawn by all the equipment in one rack, usually in kilowatts (kW). Since nearly all of it becomes heat, it is also the heat the cooling system must remove from that rack.

Expert. Quoted as design (provisioned) or measured power per rack. Typical enterprise racks run under 10 kW; liquid-cooled AI racks run 100–140 kW, with designs aimed at several hundred kW to 1 MW. It sets busbar current, coolant flow, floor loading and the size of every upstream electrical block.

Radiation hardening by design (RHBD)

Making a chip resistant to radiation through how its circuits and layout are designed, while building it in an ordinary commercial factory process rather than a special radiation-hardened one.

All three levels

Beginner. Making a chip tough against radiation through smart design, while still using a normal chip factory.

Novice. Making a chip resistant to radiation through how its circuits and layout are designed, while building it in an ordinary commercial factory process rather than a special radiation-hardened one.

Expert. Hardened standard-cell libraries (DICE and TMR flip-flops, cells with transient filters), layout measures (guard rings, well contacts, spacing between sensitive transistors) and architectural redundancy. Applied selectively to the most sensitive cells to limit the area and power cost.

Rail-optimized network

In servers with eight accelerators and eight NICs, NIC 0 of every server goes to the same leaf switch (rail 0), NIC 1 to rail 1, and so on. Accelerators with the same position in different servers then share a switch, which suits how training traffic is laid out.

All three levels

Beginner. A wiring plan where chip 1 of every server plugs into one switch, chip 2 into another, and so on. Matching chips are then one switch apart.

Novice. In servers with eight accelerators and eight NICs, NIC 0 of every server goes to the same leaf switch (rail 0), NIC 1 to rail 1, and so on. Accelerators with the same position in different servers then share a switch, which suits how training traffic is laid out.

Expert. A rail is the set of GPUs with the same local rank across scale-up domains. Rails turn the dominant same-rank traffic into one-hop traffic and multiply the servers under one leaf group by the GPUs per server; cross-rail traffic either climbs to the spine or first hops over the scale-up link (as NCCL’s PXN does).

Random dopant fluctuation

Variation in threshold voltage between nominally identical transistors, caused by the random number and placement of dopant atoms in the small volume under each gate.

All three levels

Beginner. In a very tiny transistor, the number of added atoms changes by chance from one to the next. So transistors that should be twins act a little differently.

Novice. Variation in threshold voltage between nominally identical transistors, caused by the random number and placement of dopant atoms in the small volume under each gate.

Expert. Dopant counts follow Poisson statistics, so the relative spread scales as 1/N1/\sqrt{N} and grows as the depletion volume shrinks. A major reason planar transistors stopped scaling and FinFETs and nanosheets use lightly doped channels with VTV_{\mathrm{T}} set by the gate work function.

Random-pattern-resistant fault

A fault that random patterns almost never detect because it needs one rare input combination, for example the output of an 8-input AND gate stuck at 0, which only shows when all eight inputs are 1.

All three levels

Beginner. A pretend flaw that random tests almost never catch, because it needs one very specific mix of inputs.

Novice. A fault that random patterns almost never detect because it needs one rare input combination, for example the output of an 8-input AND gate stuck at 0, which only shows when all eight inputs are 1.

Expert. Caps LBIST and random-first ATPG coverage. Fixed with weighted random patterns, deterministic top-up patterns, or test points that make the hard spot easier to control or observe.

Rayleigh criterion

A rule for when two points of light can just be told apart: the bright center of one falls on the first dark ring around the other. For a lens it gives a smallest resolvable distance of 0.61 λ/NA0.61\,\lambda/\mathrm{NA}. Chipmakers write it as CD=k1λ/NA\mathrm{CD} = k_1 \lambda / \mathrm{NA}, where CD (critical dimension) is the width of the smallest feature.

All three levels

Beginner. A rule for the smallest detail an optical system can print or see.

Novice. A rule for when two points of light can just be told apart: the bright center of one falls on the first dark ring around the other. For a lens it gives a smallest resolvable distance of 0.61 λ/NA0.61\,\lambda/\mathrm{NA}. Chipmakers write it as CD=k1λ/NA\mathrm{CD} = k_1 \lambda / \mathrm{NA}, where CD (critical dimension) is the width of the smallest feature.

Expert. k1k_1 bundles everything except wavelength and NA: the illumination shape, mask type, resist and optical proximity correction (pre-distorting mask shapes so they print as intended). Its physical floor is 0.25 for a dense pattern of lines and spaces, so a single exposure can’t print a pitch below 0.5 λ/NA0.5\,\lambda/\mathrm{NA}. Anything finer needs multiple patterning or a new tool.

RC corner

An extreme of how the factory might make the wires: for example the version with the most capacitance (wires printed wide and close together) or the one with the most resistance times capacitance (thin, narrow wires). Extraction is run once for each.

All three levels

Beginner. One extreme of how thick or thin the factory makes the wires.

Novice. An extreme of how the factory might make the wires: for example the version with the most capacitance (wires printed wide and close together) or the one with the most resistance times capacitance (thin, narrow wires). Extraction is run once for each.

Expert. Common corners are Cworst/Cbest (maximum/minimum capacitance) and RCworst/RCbest (maximum/minimum resistance × capacitance). Short nets, whose delay is mostly the driver charging capacitance, are worst at Cworst; long nets, where wire resistance matters, can be worst at RCworst. RC corners are paired with device corners on purpose rather than run as a full cross product.

RDMA

Remote direct memory access. Applications register memory with the NIC, and the NICs move data between those buffers directly, skipping the operating system’s networking code. That cuts latency and frees the CPU.

All three levels

Beginner. A way for a network card to copy data straight into another computer’s memory. Neither main processor has to help.

Novice. Remote direct memory access. Applications register memory with the NIC, and the NICs move data between those buffers directly, skipping the operating system’s networking code. That cuts latency and frees the CPU.

Expert. Kernel-bypass, zero-copy messaging with one-sided (read/write) and two-sided (send/receive) operations on queue pairs. The transport (ordering, acknowledgment, retransmission) runs on the NIC, which is why RDMA traditionally assumes a lossless fabric.

Rear-door heat exchanger

An air-to-water coil mounted on the back of a rack. The servers still use fans and air, but the air is cooled right as it leaves the rack instead of mixing into the room. Passive doors rely on server fans; active doors add their own fans.

All three levels

Beginner. A radiator on the back door of a rack. Hot air blows through it, and cool water inside carries the heat away.

Novice. An air-to-water coil mounted on the back of a rack. The servers still use fans and air, but the air is cooled right as it leaves the rack instead of mixing into the room. Passive doors rely on server fans; active doors add their own fans.

Expert. Removes heat close to the source with ordinary IT hardware, so it suits mixed or retrofit sites. Capacity depends on water temperature, flow and server airflow; one vendor’s published sizing curves span 20–40 kW per rack, and supply water must stay above the room dew point.

Recovery and removal checks

Timing checks on the moment a reset signal is released. A reset forces flip-flops to a known value; releasing it too close to a clock tick can leave a flip-flop undecided. Recovery checks the release happens early enough before the tick, removal that it happens late enough after it.

All three levels

Beginner. Timing rules for when the chip’s reset button is let go, so memory cells start up cleanly.

Novice. Timing checks on the moment a reset signal is released. A reset forces flip-flops to a known value; releasing it too close to a clock tick can leave a flip-flop undecided. Recovery checks the release happens early enough before the tick, removal that it happens late enough after it.

Expert. The setup-like (recovery) and hold-like (removal) checks on the release edge of an asynchronous set or reset pin, relative to the clock. A violation can leave the flop metastable as it leaves reset. Reset nets fan out like clocks, so reset synchronizers make release synchronous to each clock domain, and STA checks every asynchronous pin in every scenario.

Redfish

A DMTF standard API for managing servers. Software sends ordinary web requests (HTTPS) to the BMC and gets back JSON documents describing sensors, power supplies, fans and more.

All three levels

Beginner. A shared language that software uses to check on any server’s health and give it commands.

Novice. A DMTF standard API for managing servers. Software sends ordinary web requests (HTTPS) to the BMC and gets back JSON documents describing sensors, power supplies, fans and more.

Expert. A RESTful, schema-backed (OData v4) interface over HTTPS, designed as a secure, scalable replacement for IPMI-over-LAN. Resources such as Chassis, Sensors and PowerSubsystem are linked by @odata.id URIs.

Reduce-scatter

A collective that sums a buffer across all participants and leaves each participant with a different 1/p1/p slice of the result.

All three levels

Beginner. Everyone adds up their lists, but each chip ends up holding only its own slice of the total.

Novice. A collective that sums a buffer across all participants and leaves each participant with a different 1/p1/p slice of the result.

Expert. Moves (p−1)/p(p-1)/p of the buffer per participant at minimum. The first half of a bandwidth-optimal all-reduce, and the basis of sharded-optimizer (ZeRO/FSDP-style) training.

Redundant (double) via

Replacing single-cut vias with two or more cuts side by side wherever space allows, usually in a pass after detailed routing. If one cut fails to form in manufacturing, the other still connects, and two cuts in parallel also resist current less.

All three levels

Beginner. Using two vias side by side instead of one, so the connection still works if one is badly made.

Novice. Replacing single-cut vias with two or more cuts side by side wherever space allows, usually in a pass after detailed routing. If one cut fails to form in manufacturing, the other still connects, and two cuts in parallel also resist current less.

Expert. A post-route optimization that tries alternative via shapes at each location without creating new rule violations, tracked as a double-via rate. Some rules require several cuts on wide wires (LEF MINCUTS in non-default rules). Via pillars are a structured relative.

Reference model

A separate, simpler program, written from the specification, that computes what the design should output for the same inputs. For a processor it is often an instruction-set simulator: a program that carries out the same machine instructions in software.

All three levels

Beginner. A simple, separate program that works out the right answers, so a test can check the chip against them.

Novice. A separate, simpler program, written from the specification, that computes what the design should output for the same inputs. For a processor it is often an instruction-set simulator: a program that carries out the same machine instructions in software.

Expert. Usually works on whole operations and ignores exact cycle timing. Its value comes from independence: written from the spec by someone other than the designer, it catches misreadings that the designer would otherwise repeat in both.

Refresh (DRAM)

Because DRAM capacitors leak, every row must be read and written back periodically. Standard DRAM refreshes each row at least every 64 milliseconds.

All three levels

Beginner. Reading and rewriting every bit of a DRAM chip many times a second, before its leaking charge fades away.

Novice. Because DRAM capacitors leak, every row must be read and written back periodically. Standard DRAM refreshes each row at least every 64 milliseconds.

Expert. The controller issues an auto-refresh command every tREFI (7.8 µs below 85 °C); each one blocks the rank for tRFC while a batch of rows is activated and restored. The interval is set by the leakiest cells, so refresh costs bandwidth and energy that grow with density.

Region, guide and fence

An area assigned to a group of cells. A guide is a preference the placer may break if wirelength or timing are better elsewhere. A fence is a strict rule: the group must stay inside, and no other cells may enter.

All three levels

Beginner. A fenced yard on the chip where one group of parts must stay together.

Novice. An area assigned to a group of cells. A guide is a preference the placer may break if wirelength or timing are better elsewhere. A fence is a strict rule: the group must stay inside, and no other cells may enter.

Expert. DEF REGIONS with TYPE GUIDE or FENCE, with cells assigned by COMPONENTS + REGION. Over-tight fences push local density up and starve the optimizer of space.

Register file

The bank of registers inside a processor core. On a GPU it is huge (hundreds of kilobytes per core) because every resident thread keeps its own registers there all the time.

All three levels

Beginner. The tiny, fastest memory right next to the math units, where each worker keeps the numbers it’s using right now.

Novice. The bank of registers inside a processor core. On a GPU it is huge (hundreds of kilobytes per core) because every resident thread keeps its own registers there all the time.

Expert. 65,536 32-bit registers (256 KB) per NVIDIA SM since Kepler-era designs; 256–512 KiB of VGPRs per AMD CU. It’s partitioned among resident warps, so registers per thread trade directly against occupancy.

Register renaming

Mapping the few register names a program uses onto a larger set of hidden physical registers, giving every result a fresh one. It removes false dependences, where two instructions only reuse the same register name.

All three levels

Beginner. Giving each new answer its own fresh storage box, so steps that only share a name stop waiting for each other.

Novice. Mapping the few register names a program uses onto a larger set of hidden physical registers, giving every result a fresh one. It removes false dependences, where two instructions only reuse the same register name.

Expert. A map table (architectural → physical) plus a free list. Eliminates WAR and WAW hazards, leaving only true RAW dependences. Map-table snapshots or ROB walks restore the mapping after a mispredict.

Register spilling

When a program needs more registers than it’s allowed, the compiler stores some values in slower memory instead and fetches them back when needed.

All three levels

Beginner. When a worker runs out of desk space and has to put things on a far-away shelf, which slows them down.

Novice. When a program needs more registers than it’s allowed, the compiler stores some values in slower memory instead and fetches them back when needed.

Expert. On NVIDIA GPUs spills go to ‘local memory’, which lives in device DRAM behind the caches, so a spill costs global-memory latency and bandwidth. Capping registers raises occupancy but risks spills.

Register-transfer level (RTL)

A way of describing a digital circuit as storage slots (registers) plus the calculations between them. On every tick of the clock, every register takes in a newly computed value. Engineers write RTL as code, and a synthesis tool turns it into logic gates.

All three levels

Beginner. A way of describing a chip as places that store values and the rules for computing the next values, updated every tick of a clock.

Novice. A way of describing a digital circuit as storage slots (registers) plus the calculations between them. On every tick of the clock, every register takes in a newly computed value. Engineers write RTL as code, and a synthesis tool turns it into logic gates.

Expert. The level at which most chip logic is written by hand. It fixes what every register holds after every clock tick, but not which gates compute it or where they sit. Synthesis and the later stages may restructure the gates freely, as long as that cycle-by-cycle behavior is preserved.

Regression

The full collection of tests, each run many times with different random seeds, rerun automatically (usually every night) to catch anything a new change broke. Results, coverage and failures are tracked over time.

All three levels

Beginner. The big set of tests run again and again, usually every night, to catch anything a new change broke.

Novice. The full collection of tests, each run many times with different random seeds, rerun automatically (usually every night) to catch anything a new change broke. Results, coverage and failures are tracked over time.

Expert. Run by tooling that schedules thousands of jobs, groups failures by their first error message, merges coverage and ranks tests. Pass rate, new-failure count and coverage trend are the daily health metrics.

Reinforcement learning (RL)

A program (the agent) looks at a situation, takes an action and gets a score (the reward). Over many tries it learns a strategy that earns a high total score, for example where to put each large block on a chip.

All three levels

Beginner. Learning by trial and error: a program tries actions, gets a score, and slowly learns which choices lead to better scores.

Novice. A program (the agent) looks at a situation, takes an action and gets a score (the reward). Over many tries it learns a strategy that earns a high total score, for example where to put each large block on a chip.

Expert. Formally a Markov decision process: states, actions, the rules for moving between states, and a reward. In chip design the true score needs hours of placement and routing, so agents train on a fast estimate instead, and the result is only as good as that estimate’s agreement with the finished layout.

Reorder buffer (ROB)

A queue that holds every in-flight instruction in program order from dispatch until it retires. Instructions finish in any order but leave the reorder buffer only from its head, in order, which makes wrong guesses and errors easy to undo.

All three levels

Beginner. A list that keeps every step in its original order, so results can be made final in that order even if they finished in a jumble.

Novice. A queue that holds every in-flight instruction in program order from dispatch until it retires. Instructions finish in any order but leave the reorder buffer only from its head, in order, which makes wrong guesses and errors easy to undo.

Expert. Circular buffer from dispatch to commit: tracks completion and exceptions, drives in-order commit, and gives precise exceptions. Its size caps how far ahead the core can speculate.

Replacement metal gate

A process in which a sacrificial ‘dummy’ gate defines where the gate goes during the hot source/drain steps, then is removed and replaced by the high-k insulator and metal gate. In nanosheet transistors the swap is also when the sheets are released and wrapped.

All three levels

Beginner. Building a stand-in gate first, then swapping it for the real metal gate near the end, so the real gate isn’t damaged by the hot early steps.

Novice. A process in which a sacrificial ‘dummy’ gate defines where the gate goes during the hot source/drain steps, then is removed and replaced by the high-k insulator and metal gate. In nanosheet transistors the swap is also when the sheets are released and wrapped.

Expert. Gate-last HKMG flow. In GAA the dummy-gate removal opens the way for channel release, and the interfacial oxide, high-k and work-function metals are then deposited conformally by ALD into the gaps between sheets, only a few nanometers wide.

Replacement policy

How a cache chooses the victim when a set is full: least recently used (LRU), an approximation such as tree pseudo-LRU, first-in first-out, or random.

All three levels

Beginner. The rule a cache uses to pick which block to throw out when it needs room.

Novice. How a cache chooses the victim when a set is full: least recently used (LRU), an approximation such as tree pseudo-LRU, first-in first-out, or random.

Expert. True LRU needs log2(N!) bits per N-way set, so hardware uses tree PLRU or not-recently-used bits; last-level caches use scan- and thrash-resistant policies such as RRIP that predict re-reference distance.

Requirement (“shall” statement)

One sentence using “shall”, with its own ID, a number to hit, the conditions it applies under, where it came from, and how it will be checked: by test, analysis, inspection or demonstration.

All three levels

Beginner. One clear, checkable promise the chip must keep, with its own ID number.

Novice. One sentence using “shall”, with its own ID, a number to hit, the conditions it applies under, where it came from, and how it will be checked: by test, analysis, inspection or demonstration.

Expert. Good requirements are single, verifiable, attainable and traceable. A target needs conditions: “≤ 5 W” means nothing without the workload, temperature and manufacturing corner it applies at. Words like fast, small or robust are banned because nobody can verify them.

Requirements traceability

Links from each requirement up to the customer need it serves and down to the tests that prove it. The links expose requirements that nothing tests, and tests that check something nobody asked for.

All three levels

Beginner. Keeping a link from every promise in the spec to the test that proves the chip keeps it.

Novice. Links from each requirement up to the customer need it serves and down to the tests that prove it. The links expose requirements that nothing tests, and tests that check something nobody asked for.

Expert. Kept in both directions and maintained by tools: requirement IDs appear in test plans, coverage reports and the final signoff checklist, so a script can report the status of every requirement. It is mandatory in regulated work such as aircraft (DO-254) and automotive (ISO 26262) electronics, where reviewers check that verification covers each requirement completely.

Reservation station

A buffer in front of an execution unit where an instruction waits for its operands. Each operand is either a value or a tag naming which instruction will produce it; when that result is broadcast, the station grabs it.

All three levels

Beginner. A waiting seat where a step sits until all the numbers it needs have arrived.

Novice. A buffer in front of an execution unit where an instruction waits for its operands. Each operand is either a value or a tag naming which instruction will produce it; when that result is broadcast, the station grabs it.

Expert. From Tomasulo’s IBM 360/91 floating-point unit: tag-based operand capture from a common data bus. Modern cores use unified or distributed issue queues with the same wakeup-and-select function.

Reset domain crossing (RDC)

A path from a flip-flop on one reset signal to a flip-flop on a different reset, or on none. If the first is reset while the second keeps running, the second sees its input jump at an arbitrary moment and can be caught mid-change, even with a single clock.

All three levels

Beginner. A spot where resetting one part of the chip can scramble a neighboring part that wasn’t reset.

Novice. A path from a flip-flop on one reset signal to a flip-flop on a different reset, or on none. If the first is reset while the second keeps running, the second sees its input jump at an arbitrary moment and can be caught mid-change, even with a single clock.

Expert. Asserting an asynchronous reset is untimed relative to the receiving flop’s clock, so it behaves like a clock crossing. The fixes: order the resets so the receiver enters reset first, isolate the receiver with an enable while the source is in reset, or put both on the same reset.

Reset synchronizer

A pair of flip-flops that lets reset (the signal that puts flip-flops into a known starting state) take effect immediately, but ends it only in step with the clock, so every flip-flop leaves reset on the same tick.

All three levels

Beginner. A small circuit that makes sure the chip comes out of reset in step with its clock.

Novice. A pair of flip-flops that lets reset (the signal that puts flip-flops into a known starting state) take effect immediately, but ends it only in step with the clock, so every flip-flop leaves reset on the same tick.

Expert. Asynchronous assertion, synchronous release. It keeps reset from ending too close to a clock edge, which would be a recovery or removal violation. Each asynchronous clock domain gets its own, and test logic needs a way to control its output during scan testing.

Resource sharing

Using one piece of hardware, such as an adder, for two operations that never happen in the same clock cycle, with a selector (multiplexer) in front to choose its inputs.

All three levels

Beginner. Using one piece of hardware for two jobs that never happen at the same time.

Novice. Using one piece of hardware, such as an adder, for two operations that never happen in the same clock cycle, with a selector (multiplexer) in front to choose its inputs.

Expert. Merging operators that are mutually exclusive, never active in the same cycle. It saves area but puts a multiplexer’s delay in front of the shared operator, so timing-driven tools weigh it against slack. Yosys proves the exclusivity with a SAT solver.

Respin

A new tapeout of a corrected design after the first chips fail or miss a target. A full respin needs a whole new mask set; a metal-only respin replaces only the masks for the wiring layers.

All three levels

Beginner. Making the chip again with fixes after the first version comes back with a problem.

Novice. A new tapeout of a corrected design after the first chips fail or miss a target. A full respin needs a whole new mask set; a metal-only respin replaces only the masks for the wiring layers.

Expert. Costs new masks plus another trip through the fab, about a quarter of a year at advanced nodes. Metal-only respins rewire spare cells and reuse the transistor-layer masks, so teams scatter spare cells over the design before the first tapeout.

Restoring (regenerative) logic

Each gate has a gain greater than one in its switching region and less than one near 0 and 1, so a slightly degraded input produces an output closer to the ideal 0 or 1. A chain of such gates restores full logic levels.

All three levels

Beginner. Logic that cleans up a slightly wrong signal, so small errors shrink as they pass through gates instead of growing.

Novice. Each gate has a gain greater than one in its switching region and less than one near 0 and 1, so a slightly degraded input produces an output closer to the ideal 0 or 1. A chain of such gates restores full logic levels.

Expert. Follows from ∣slope∣>1|\text{slope}| > 1 in the transition region of the VTC and <1< 1 in the valid regions. It is why static CMOS tolerates noise stage after stage and why the cross-coupled inverters of a latch or SRAM cell snap to a stable state.

retention flip-flop

A flip-flop (a one-bit storage cell) with a tiny backup cell powered by the always-on supply. It copies its value into the backup before the block powers off and copies it back on wake-up, so the block resumes where it stopped.

All three levels

Beginner. A memory bit with a tiny backup, so it keeps its value while the rest of its zone sleeps.

Novice. A flip-flop (a one-bit storage cell) with a tiny backup cell powered by the always-on supply. It copies its value into the backup before the block powers off and copies it back on wake-up, so the block resumes where it stopped.

Expert. A shadow latch on an always-on supply with save and restore controls. It trades area and leakage for fast wake-up without software state restore; the save/restore sequence and the always-on supply routing to every retention cell must be verified.

Reticle

The mask as it is loaded into the lithography machine. Its layout places one or more copies of the chip with test patterns, alignment marks and the scribe lines along which the wafer is later cut.

All three levels

Beginner. The stencil plate used in the printing machine. It usually holds a few copies of the chip side by side.

Novice. The mask as it is loaded into the lithography machine. Its layout places one or more copies of the chip with test patterns, alignment marks and the scribe lines along which the wafer is later cut.

Expert. The reticle layout is built during mask data preparation. Its maximum size is set by the lithography tool, which caps the size of a single die.

reticle limit

The largest area a lithography machine (the tool that prints a chip’s patterns onto the wafer) can expose in one shot, roughly 800 to 860 mm². A chip made as one piece of silicon must fit inside it.

All three levels

Beginner. The largest area a chipmaking machine can print in one shot, about the size of a postage stamp. One chip can’t be bigger than this.

Novice. The largest area a lithography machine (the tool that prints a chip’s patterns onto the wafer) can expose in one shot, roughly 800 to 860 mm². A chip made as one piece of silicon must fit inside it.

Expert. Set by the scanner’s exposure field, 26 mm × 33 mm on current tools. Dies near the limit yield poorly and few fit on a wafer; going beyond it takes stitching several exposures, splitting into chiplets, or wafer-scale integration.

Retimer

A chip that receives a fast PCIe signal, recovers the data, and transmits a fresh copy. It is used when the wire is too long or lossy for the signal to arrive cleanly.

All three levels

Beginner. A small chip partway along a long connection. It cleans up the signal and sends it on, so it can travel farther.

Novice. A chip that receives a fast PCIe signal, recovers the data, and transmits a fresh copy. It is used when the wire is too long or lossy for the signal to arrive cleanly.

Expert. A protocol-aware, software-transparent physical-layer device that splits a channel into two segments, each with its own loss budget. Unlike a redriver (an analog amplifier), it resets jitter and noise, at the cost of a few nanoseconds of latency, power and board area.

Retiming

Moving flip-flops forward or backward through the logic so each stage between them has a similar amount of work, without changing what the circuit outputs on each clock tick.

All three levels

Beginner. Moving the chip’s small memories a little earlier or later along a path, so the work between them is shared out more evenly.

Novice. Moving flip-flops forward or backward through the logic so each stage between them has a similar amount of work, without changing what the circuit outputs on each clock tick.

Expert. A transformation formalized by Leiserson and Saxe that moves registers across logic to minimize the clock period or the register count while keeping cycle-by-cycle input/output behavior. Flip-flops no longer match the RTL one-to-one, so equivalence checking needs guidance or slower sequential methods.

Return address stack

A small hardware stack that predicts where a function will return to: a call pushes the address after it, and a return pops it. It is much more accurate for returns than a branch target buffer.

All three levels

Beginner. A stack of notes that remembers where to come back to after each helper part of a program finishes.

Novice. A small hardware stack that predicts where a function will return to: a call pushes the address after it, and a return pops it. It is much more accurate for returns than a branch target buffer.

Expert. Typically tens of entries; fails on overflow or mismatched call/return. Its speculative pushes and pops must be repaired after a mispredict.

Ridge point (machine balance)

Peak compute divided by memory bandwidth, in FLOPs per byte. It is the least arithmetic intensity a workload needs to reach the chip’s peak speed.

All three levels

Beginner. The bend in the roofline chart. Jobs to its right can use all the chip’s math; jobs to its left are held back by memory.

Novice. Peak compute divided by memory bandwidth, in FLOPs per byte. It is the least arithmetic intensity a workload needs to reach the chip’s peak speed.

Expert. Peak FLOP/s ÷ peak byte/s, defined per memory level. It has drifted right as compute grew faster than bandwidth, so more kernels fall on the slope. Williams et al. read it as a measure of how hard a machine is to program to peak.

Rip-up and reroute

A repeat-until-done method: route every net, sometimes letting them overlap, then remove the nets involved in conflicts and route them again with the crowded spots made more expensive. Repeat until nothing conflicts.

All three levels

Beginner. Pulling out a wire that is in the way and drawing it again somewhere else, so another wire can fit.

Novice. A repeat-until-done method: route every net, sometimes letting them overlap, then remove the nets involved in conflicts and route them again with the crowded spots made more expensive. Repeat until nothing conflicts.

Expert. It converges only if costs remember history; otherwise nets trade places forever. The rip-up order, the region rerouted (whole net or a local window) and the cost schedule decide convergence. Global routers use it as negotiated congestion; detailed routers as search and repair.

RISC and CISC

RISC (reduced instruction set computer) has simple instructions of fixed length, and only loads and stores touch memory. CISC (complex instruction set computer) has many instructions of varying length, and some combine a memory access with arithmetic. Arm and RISC-V are RISC; x86 is CISC.

All three levels

Beginner. Two styles of command list. RISC uses only small, simple commands. CISC also has big commands that do several jobs at once.

Novice. RISC (reduced instruction set computer) has simple instructions of fixed length, and only loads and stores touch memory. CISC (complex instruction set computer) has many instructions of varying length, and some combine a memory access with arithmetic. Arm and RISC-V are RISC; x86 is CISC.

Expert. The 1980s argument: simple, regular instructions make pipelining and compilers easier at the cost of more instructions. Today x86 cores translate instructions into RISC-like micro-ops, and measurements find little intrinsic performance or energy difference between the styles.

RISC-V

An open, royalty-free RISC instruction set begun at UC Berkeley in 2010 and now managed by the non-profit RISC-V International. A small base set of integer instructions plus optional standard extensions.

All three levels

Beginner. A free, open list of commands for chips. Anyone can build a chip that follows it, without paying for the right.

Novice. An open, royalty-free RISC instruction set begun at UC Berkeley in 2010 and now managed by the non-profit RISC-V International. A small base set of integer instructions plus optional standard extensions.

Expert. A modular load-store ISA: RV32I/RV64I base (32 integer registers, x0 hardwired to zero, fixed 32-bit encodings in six formats) plus extensions such as M, A, F, D, C and V, with no exposed delay slots.

RoCE (RDMA over Converged Ethernet)

RDMA carried over Ethernet. Version 2 (RoCEv2) wraps the InfiniBand transport in standard IP and UDP headers, so it can be routed through normal data-center switches.

All three levels

Beginner. A way to run RDMA, the fast copy-straight-into-memory kind of networking, over Ethernet. Ethernet is the network most datacenters use.

Novice. RDMA carried over Ethernet. Version 2 (RoCEv2) wraps the InfiniBand transport in standard IP and UDP headers, so it can be routed through normal data-center switches.

Expert. RoCEv2 puts the IB transport header and payload inside Ethernet/IP/UDP (destination port 4791). The source UDP port per queue pair gives ECMP its entropy. Classic RoCE NICs use go-back-N recovery, so deployments add PFC to make the Ethernet fabric lossless.

Roofline model

A chart of the best speed a chip can reach on a task: attainable performance is the smaller of peak compute and memory bandwidth×arithmetic intensity\text{memory bandwidth} \times \text{arithmetic intensity}. On log-log axes the limit looks like a roof: a sloped line (limited by memory) meeting a flat one (limited by compute).

All three levels

Beginner. A simple chart that shows whether a job is limited by how fast the chip can calculate or by how fast it can fetch data.

Novice. A chart of the best speed a chip can reach on a task: attainable performance is the smaller of peak compute and memory bandwidth×arithmetic intensity\text{memory bandwidth} \times \text{arithmetic intensity}. On log-log axes the limit looks like a roof: a sloped line (limited by memory) meeting a flat one (limited by compute).

Expert. The ridge point, peak compute÷bandwidth\text{peak compute} \div \text{bandwidth}, is the lowest arithmetic intensity at which a task can reach peak. Lower “ceilings” show the limit when an optimization is missing. For chips with several memory levels, draw one roofline per level (DRAM, on-chip SRAM).

Root complex

The logic in the CPU that connects the PCIe links to the processor and its memory. Each link starts at one of its root ports.

All three levels

Beginner. The part of the processor where all its PCIe wires start.

Novice. The logic in the CPU that connects the PCIe links to the processor and its memory. Each link starts at one of its root ports.

Expert. The top of the PCIe hierarchy: root ports originate links, translate device DMA into memory accesses, and enforce ordering and isolation. Peer-to-peer traffic that must cross between root ports or sockets goes through it, often at reduced performance.

Round to nearest, ties to even

The default rounding mode for floating point. A value is rounded to the nearest representable number; exact ties go to the one whose last bit is 0. Breaking ties this way avoids a systematic upward drift.

All three levels

Beginner. The rule for rounding a number that doesn’t fit: pick the closest value the format can store, and when it’s exactly halfway, pick the one ending in an even digit.

Novice. The default rounding mode for floating point. A value is rounded to the nearest representable number; exact ties go to the one whose last bit is 0. Breaking ties this way avoids a systematic upward drift.

Expert. IEEE 754’s roundTiesToEven, required by the OCP FP8 and MX specifications for conversions. Unbiased on average, but deterministic: repeated tiny updates can round away to nothing, which is why low-precision training sometimes uses stochastic rounding instead.

Route guide

The output of global routing for one net: a set of rectangles, each a run of GCells on one layer, that the detailed router should keep that net’s wires inside.

All three levels

Beginner. The rough plan for one connection that the planning pass hands to the exact pass.

Novice. The output of global routing for one net: a set of rectangles, each a run of GCells on one layer, that the detailed router should keep that net’s wires inside.

Expert. A per-net list of GCell-aligned rectangles with a layer each. TritonRoute reads them in the ISPD-2018/2019 contest format. Detailed routers may stray outside the guides at a cost, so comparing guide coverage with final wiring shows how well the global model matched reality.

Router (on-chip network)

A small switch at each node of an on-chip network. It takes packets arriving from its neighbors or its own core and forwards each one toward its destination, one hop at a time.

All three levels

Beginner. The little intersection at each core that decides which way a piece of data goes next.

Novice. A small switch at each node of an on-chip network. It takes packets arriving from its neighbors or its own core and forwards each one toward its destination, one hop at a time.

Expert. Typically a few input buffers, a routing-computation step, an allocator and a crossbar, pipelined over one to a few cycles. Its buffer depth and flow-control scheme set how it behaves under congestion.

Routing fabric (programmable interconnect)

All the wire segments and programmable switches that connect logic blocks, memory, DSP blocks and pins. It takes most of an FPGA’s area and most of its delay.

All three levels

Beginner. The wires and switches that join an FPGA’s blocks. Most of the chip is wires and switches.

Novice. All the wire segments and programmable switches that connect logic blocks, memory, DSP blocks and pins. It takes most of an FPGA’s area and most of its delay.

Expert. Island-style channels of W tracks of mixed segment lengths, connection boxes (Fc) and switch boxes (Fs). Most configuration bits and 60–70% of power sit in it.

Routing overflow

How many more wires want to cross a GCell border than it has free tracks for. A report gives it per layer, as a total and as the worst single border. Zero overflow everywhere means the global plan fits.

All three levels

Beginner. How many more wires want to pass through a spot than there are lanes for.

Novice. How many more wires want to cross a GCell border than it has free tracks for. A report gives it per layer, as a total and as the worst single border. Zero overflow everywhere means the global plan fits.

Expert. Total overflow sums the excess over every border; maximum overflow shows the worst one. Small, scattered overflow often clears in detailed routing. Overflow clustered near large blocks or dense cells usually becomes rule violations and needs a placement or floorplan fix.

Routing track and pitch

One of the evenly spaced lines a wire’s center may run along on a layer. The spacing between tracks, the pitch, is about one minimum wire width plus one minimum gap. The number of tracks that cross a region is how many wires can pass through it.

All three levels

Beginner. An invisible lane on a wiring layer. Wires run along these lanes, which are spaced evenly so wires never get too close.

Novice. One of the evenly spaced lines a wire’s center may run along on a layer. The spacing between tracks, the pitch, is about one minimum wire width plus one minimum gap. The number of tracks that cross a region is how many wires can pass through it.

Expert. Each layer’s track pattern (starting offset, pitch and, on multi-patterned layers, mask color) fixes where wires may go; DEF records it as TRACKS statements built from LEF PITCH. Standard-cell height is quoted in tracks. Pins that sit off the track grid, and layers whose pitches don’t line up, make pin access and via placement harder.

Row and site

The core of the chip is tiled with horizontal rows, each exactly one cell tall. Each row is divided into equal slots called sites, the smallest step a cell can be placed at. A cell must start on a site boundary, and its width is a whole number of sites. Every other row is flipped so neighboring rows can share the power or ground wire along their common edge.

All three levels

Beginner. The neat lanes on a chip where the small parts must sit, split into equal slots like spaces in a parking lot.

Novice. The core of the chip is tiled with horizontal rows, each exactly one cell tall. Each row is divided into equal slots called sites, the smallest step a cell can be placed at. A cell must start on a site boundary, and its width is a whole number of sites. Every other row is flipped so neighboring rows can share the power or ground wire along their common edge.

Expert. The site comes from a LEF SITE statement (size, symmetry); DEF ROW statements lay rows out with an origin, an orientation and a site count and step. Multi-height cells span several rows and must match the power/ground rail pattern, which limits which rows and orientations are legal for them.

Row-stationary (RS)

The dataflow introduced with the Eyeriss chip: each PE computes a one-dimensional convolution of a filter row with an input row, keeping the filter row, part of the input row and its partial sum in a small local memory.

All three levels

Beginner. A way of sharing work across a grid where each cell handles one row of a small image filter, so every kind of number gets reused close by.

Novice. The dataflow introduced with the Eyeriss chip: each PE computes a one-dimensional convolution of a filter row with an input row, keeping the filter row, part of the input row and its partial sum in a small local memory.

Expert. Optimizes reuse of weights, activations and partial sums together in each PE’s register file, rather than one data type; Eyeriss reports 1.4–2.5× lower energy than WS, OS and no-local-reuse dataflows in AlexNet’s conv layers.

RowHammer

A DRAM disturbance error: activating the same row very many times between refreshes makes charge leak faster from cells in neighboring rows, flipping bits that were never accessed.

All three levels

Beginner. A flaw where reading one row of DRAM over and over can flip bits in the rows next to it.

Novice. A DRAM disturbance error: activating the same row very many times between refreshes makes charge leak faster from cells in neighboring rows, flipping bits that were never accessed.

Expert. Caused by repeated word-line toggling coupling into adjacent cells. It became a security problem because software can trigger it. Mitigations: more frequent refresh, target-row refresh of neighbors, and access counting in the controller or the DRAM.

RRAM (resistive RAM)

Resistive random-access memory: a thin oxide between two electrodes whose resistance can be switched, and often tuned to many levels, by voltage pulses. Non-volatile and small, so it can hold weights densely.

All three levels

Beginner. A memory cell that stores a number as how easily it lets electricity through. It keeps its value with the power off.

Novice. Resistive random-access memory: a thin oxide between two electrodes whose resistance can be switched, and often tuned to many levels, by voltage pulses. Non-volatile and small, so it can hold weights densely.

Expert. Conductive-filament devices; analog programming needs iterative write-verify. Device-to-device and cycle-to-cycle variation, conductance relaxation after programming, limited endurance and higher forming/programming voltages are the practical issues.

RTL-to-GDS flow

The sequence of software steps from the design’s code (RTL) to the final layout file (GDSII): turn the code into logic gates, plan the chip’s area, lay out the power wiring, place the gates, build the clock wiring, route the remaining wires, and run the final checks.

All three levels

Beginner. The full chain of programs that turns a chip’s written plan into the final drawing sent to the factory.

Novice. The sequence of software steps from the design’s code (RTL) to the final layout file (GDSII): turn the code into logic gates, plan the chip’s area, lay out the power wiring, place the gates, build the clock wiring, route the remaining wires, and run the final checks.

Expert. In practice a set of scripts (Makefiles, Tcl, Python) that runs the tools in order, passes the design between them in a database or in files, and checks quality numbers such as timing, area and wiring congestion between steps before letting the run continue. OpenROAD-flow-scripts and LibreLane are open examples; companies build theirs starting from vendor and foundry reference flows.

RUDY

Rectangular Uniform wire DensitY: a quick congestion estimate. Each connection’s expected wire is spread evenly over the rectangle around its pins, and the amounts from all connections are added up tile by tile to make a congestion map.

All three levels

Beginner. A quick way to guess where wires will pile up, by spreading each wire evenly over the box around the parts it connects.

Novice. Rectangular Uniform wire DensitY: a quick congestion estimate. Each connection’s expected wire is spread evenly over the rectangle around its pins, and the amounts from all connections are added up tile by tile to make a congestion map.

Expert. Per net, a density of (w+h)/(wh)(w + h)/(w h) over its bounding box, summed over nets. Router-independent and cheap, so placers call it inside the loop (OpenROAD gpl uses it for cell inflation). It ignores layer assignment and detours, so confirm with a global route.

Rule deck (runset)

A file supplied by the factory that encodes its manufacturing rules (minimum widths, gaps and so on) as a program a checking tool can run over a layout.

All three levels

Beginner. The factory’s rule book, written so a computer can check a chip drawing against it.

Novice. A file supplied by the factory that encodes its manufacturing rules (minimum widths, gaps and so on) as a program a checking tool can run over a layout.

Expert. A foundry-qualified program for one physical verification tool: layer definitions, derived layers built with Boolean operations, and hundreds to thousands of geometric checks, plus device and connectivity definitions for LVS. Decks are versioned, and the checklist records which version was run.

Rush current (in-rush current)

The surge of current at wake-up, when a powered-off block’s internal capacitance charges back up. If every switch turns on at once, the surge can pull down the supply for neighboring blocks that are still working.

All three levels

Beginner. The sudden rush of electricity when a switched-off part of the chip is turned back on.

Novice. The surge of current at wake-up, when a powered-off block’s internal capacitance charges back up. If every switch turns on at once, the surge can pull down the supply for neighboring blocks that are still working.

Expert. Limited by turning switches on in stages, for example through a daisy chain or in groups that wake one after another. The trade-off is peak current against wake-up time and the number of switch transistors.

Safety island

A small, separately protected part of a large chip, often lockstep processors plus watchdogs, that monitors the big compute blocks and takes the system to a safe state if they fail.

All three levels

Beginner. A small, extra-reliable helper inside a big chip. It watches the rest and steps in if something goes wrong.

Novice. A small, separately protected part of a large chip, often lockstep processors plus watchdogs, that monitors the big compute blocks and takes the system to a safe state if they fail.

Expert. Provides watchdogs, fault collection, diagnostics and safe-state control, independent of the monitored domain (separate clock and power where possible).

Safety mechanism

A piece of hardware or software that detects a fault, or keeps it from causing harm, and moves the system to a safe state. Examples are error-correcting memory, two processors comparing answers, and a watchdog timer.

All three levels

Beginner. A built-in checker that notices when something has gone wrong and reacts.

Novice. A piece of hardware or software that detects a fault, or keeps it from causing harm, and moves the system to a safe state. Examples are error-correcting memory, two processors comparing answers, and a watchdog timer.

Expert. Credited in the FMEDA with a diagnostic coverage per failure mode. It must itself be tested (by built-in self-test or periodic checks), or its own failure becomes a latent fault.

SAT solver

A program that, given a logic formula, either finds values for its variables that make the formula true or proves that no such values exist. Many formal tools work by turning a question about the design into such a formula.

All three levels

Beginner. A program that solves giant logic puzzles. It finds a way to make every condition true, or proves there is no way.

Novice. A program that, given a logic formula, either finds values for its variables that make the formula true or proves that no such values exist. Many formal tools work by turning a question about the design into such a formula.

Expert. Modern solvers learn a new clause from every dead end (conflict-driven clause learning) and can be called again and again with small changes (incremental solving). They power BMC, k-induction, IC3, equivalence checking and stimulus generation. SMT solvers extend SAT with arithmetic on bit-vectors and other theories.

Saturation region

The region where VGSV_{\mathrm{GS}} is above threshold and VDSV_{\mathrm{DS}} is large (at least VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}} in the simple model). The channel pinches off near the drain, and the current stops growing with VDSV_{\mathrm{DS}}, so the transistor acts like a current source set by the gate.

All three levels

Beginner. An on state where the flow has leveled off. More push across the transistor barely helps; only the gate can raise it.

Novice. The region where VGSV_{\mathrm{GS}} is above threshold and VDSV_{\mathrm{DS}} is large (at least VGS−VtV_{\mathrm{GS}} - V_{\mathrm{t}} in the simple model). The channel pinches off near the drain, and the current stops growing with VDSV_{\mathrm{DS}}, so the transistor acts like a current source set by the gate.

Expert. VDS≥VDSATV_{\mathrm{DS}} \ge V_{\mathrm{DSAT}}. Not to be confused with velocity saturation, which is a carrier-speed limit and changes where VDSATV_{\mathrm{DSAT}} is. In real devices IDI_{\mathrm{D}} still rises slowly with VDSV_{\mathrm{DS}} through channel-length modulation and DIBL.

Scale factor

A number stored once per tensor, row or block that the quantized values are multiplied by to get back the real values. It lets a narrow format cover whatever range the data actually has.

All three levels

Beginner. A shared multiplier stored alongside a group of small numbers, telling you how big they really are.

Novice. A number stored once per tensor, row or block that the quantized values are multiplied by to get back the real values. It lets a narrow format cover whatever range the data actually has.

Expert. Real-valued (often FP32) per-tensor or per-channel scales for INT8 and FP8; power-of-two E8M0 scales per 32-element block in MX; E4M3 scales per 16 elements plus an FP32 tensor scale in NVFP4. In a dot product the scales factor out and are applied once per block or per output.

Scale-out network

The datacenter network (InfiniBand or Ethernet, through layers of switches) that joins many servers or scale-up domains. It reaches far more chips but gives each one less bandwidth and more latency than the scale-up fabric.

All three levels

Beginner. The network that joins the servers across a whole building. It is slower than the links inside a team, but it reaches thousands of machines.

Novice. The datacenter network (InfiniBand or Ethernet, through layers of switches) that joins many servers or scale-up domains. It reaches far more chips but gives each one less bandwidth and more latency than the scale-up fabric.

Expert. NIC-based, message-oriented (RDMA) networks built as multi-tier Clos or similar topologies. Per-accelerator bandwidth is typically several times lower than scale-up, and latency includes NIC and multi-hop switch traversal, so only parallelism that tolerates that, such as data and pipeline parallelism, is placed across it.

Scale-up domain

The set of accelerators joined by a dedicated high-bandwidth, low-latency fabric, usually inside one server or one rack. Within it, chips can read and write each other’s memory and run group operations quickly. Also called an NVLink domain, a pod or a node, depending on the vendor.

All three levels

Beginner. A team of AI chips joined by such fast links that they can work almost like one giant chip.

Novice. The set of accelerators joined by a dedicated high-bandwidth, low-latency fabric, usually inside one server or one rack. Within it, chips can read and write each other’s memory and run group operations quickly. Also called an NVLink domain, a pod or a node, depending on the vendor.

Expert. The largest set of accelerators that share a memory-semantic fabric with roughly an order of magnitude more bandwidth per chip than the scale-out NIC. Its size (8, 16, 64, 72 chips in current systems) bounds tensor and expert parallelism, which is why designers stretch it to rack scale over copper.

Scan chain

Scan flip-flops wired output-to-input into one long shift register, from a scan-in pin to a scan-out pin. Each clock tick moves every bit one place along. Chips use many chains side by side to save time.

All three levels

Beginner. A long line of the chip’s memory cells linked together in test mode. Values go in one end and come out the other, like a bucket brigade.

Novice. Scan flip-flops wired output-to-input into one long shift register, from a scan-in pin to a scan-out pin. Each clock tick moves every bit one place along. Chips use many chains side by side to save time.

Expert. Shift time per pattern is set by the longest chain, so chains are balanced in length. Chains are usually grouped by clock domain and clock edge, with lockup latches where they must cross. Chain order is reworked after placement to shorten wires.

Scan diagnosis

Software that takes a failing chip’s wrong outputs, traces them back through the circuit and simulates likely faults, to list where the defect probably is and what kind it is.

All three levels

Beginner. Using a failed chip’s test results to work out where on the chip the flaw probably is.

Novice. Software that takes a failing chip’s wrong outputs, traces them back through the circuit and simulates likely faults, to list where the defect probably is and what kind it is.

Expert. Judged by resolution (how many candidates) and accuracy (whether the real defect is among them). Volume diagnosis correlates results across many failing chips to find systematic yield problems and choose chips for physical failure analysis.

Scan enable

A single control signal (SE) wired to every scan flip-flop. When it is 1, the flip-flops form chains and shift; when it is 0, they store what the logic computed, as in normal use.

All three levels

Beginner. The switch that flips the chip’s memory cells between normal work and passing test values along the chain.

Novice. A single control signal (SE) wired to every scan flip-flop. When it is 1, the flip-flops form chains and shift; when it is 0, they store what the logic computed, as in normal use.

Expert. A very high-fanout net, buffered much like a slow clock. Launch-on-capture lets it change slowly; launch-on-shift needs it to switch within one functional clock period across the whole chip, which makes it a timing-critical net.

Scan flip-flop

A flip-flop with a two-way switch (a multiplexer) in front of its input. In normal mode it stores the value the logic computed; in test mode it stores the value from the previous flip-flop in the chain instead.

All three levels

Beginner. A tiny memory cell with an extra side door, so in test mode it can pass values along a chain.

Novice. A flip-flop with a two-way switch (a multiplexer) in front of its input. In normal mode it stores the value the logic computed; in test mode it stores the value from the previous flip-flop in the chain instead.

Expert. The mux-D style is the usual choice in standard-cell flows (two-clock LSSD styles also exist). The mux adds delay on the functional data path and the chain connection adds a load on the output; libraries offer scan versions of each flop so scan replacement is a cell swap.

Scan-chain reordering

For testing, a chip’s flip-flops (one-bit memories) are linked into long chains so a tester can load and read every bit. After placement, the order of flip-flops in each chain is changed so each link joins nearby flip-flops. The chip’s normal behavior does not depend on that order.

All three levels

Beginner. Rearranging the chain of parts used for testing, so each link connects to a close neighbor instead of one far away.

Novice. For testing, a chip’s flip-flops (one-bit memories) are linked into long chains so a tester can load and read every bit. After placement, the order of flip-flops in each chain is changed so each link joins nearby flip-flops. The chip’s normal behavior does not depend on that order.

Expert. Driven by the DEF SCANCHAINS section: FLOATING lists may be reordered, ORDERED lists may not. Classically a traveling-salesman problem over flip-flop locations, constrained by chain partitions and clock domains. ATPG patterns must be regenerated for the new order.

Scoreboard

The part of a testbench that collects what the design actually produced and what the reference model predicted, pairs them up, and reports any mismatch.

All three levels

Beginner. The part of a test that keeps score. It compares what the chip did with what it should have done.

Novice. The part of a testbench that collects what the design actually produced and what the reference model predicted, pairs them up, and reports any mismatch.

Expert. Has to cope with outputs that arrive reordered, dropped or duplicated, and must check at the end of a test that nothing is still waiting. A scoreboard that compared nothing passes silently, so tests should also check that it saw traffic.

Scratchpad memory

Fast on-chip memory that software manages directly: a program or a DMA engine (a block that copies data without the processor’s help) explicitly moves data in and out. Unlike a cache it has no tags and never guesses, so its timing is predictable. Common in accelerators.

All three levels

Beginner. A fast on-chip memory that software fills and empties on purpose, instead of letting hardware guess what to keep.

Novice. Fast on-chip memory that software manages directly: a program or a DMA engine (a block that copies data without the processor’s help) explicitly moves data in and out. Unlike a cache it has no tags and never guesses, so its timing is predictable. Common in accelerators.

Expert. It saves the area and energy of tags and gives fixed latency, but the compiler or driver must plan every transfer. Sizing it is a joint hardware and software decision: too small and traffic to DRAM rises; too large and it dominates the die.

Scribe line

A thin strip left between neighboring dies so they can be sawn apart without damage. It also holds small test structures the fab uses to check the process.

All three levels

Beginner. The narrow empty lane between chips on a wafer, where the saw cuts them apart.

Novice. A thin strip left between neighboring dies so they can be sawn apart without damage. It also holds small test structures the fab uses to check the process.

Expert. Normally keep-out for product wiring, since it is destroyed at dicing and carries fab test structures. Wafer-scale designs negotiate special wiring across it so neighboring reticle fields join into one circuit.

Scrubbing

Regularly reading through memory (or an FPGA’s configuration), correcting any flipped bits, and writing the clean data back, so errors don’t accumulate.

All three levels

Beginner. Reading memory again and again and fixing any mistakes, before they pile up.

Novice. Regularly reading through memory (or an FPGA’s configuration), correcting any flipped bits, and writing the clean data back, so errors don’t accumulate.

Expert. The scrub interval is set against the upset rate, so the chance that two upsets land in one code word between scrubs stays acceptably low.

SDC (timing constraints)

Synopsys Design Constraints: the file of timing goals, written as short commands. create_clock says how fast the clock ticks; set_input_delay and set_output_delay say how much of each tick is used up outside the block.

All three levels

Beginner. Instructions that tell the tool how fast the chip must run.

Novice. Synopsys Design Constraints: the file of timing goals, written as short commands. create_clock says how fast the clock ticks; set_input_delay and set_output_delay say how much of each tick is used up outside the block.

Expert. A Tcl-based format that synthesis, timing analysis and layout tools all read. It defines clocks, timing budgets at the block’s inputs and outputs, exceptions (paths to ignore or to give extra cycles), clock margin and design-rule limits. A wrong constraint silently changes what every tool optimizes and checks, so SDC gets its own review.

secure element

A small, tamper-resistant processor with encryption hardware and sensors. It stores secret keys and carries out sensitive operations so the keys never leave the chip.

All three levels

Beginner. A tiny chip, like the one on a bank card, built to keep secret codes safe even if someone steals it.

Novice. A small, tamper-resistant processor with encryption hardware and sensors. It stores secret keys and carries out sensitive operations so the keys never leave the chip.

Expert. Designed against an attacker who holds the chip: probing, fault injection and side channels. Certified against a protection profile at high attack potential.

secure enclave

A walled-off part of a chip, with its own small processor, memory protection and encryption hardware, that keeps secrets such as keys and fingerprint data safe even if the main processor’s software is compromised.

All three levels

Beginner. A locked-off part of the chip that guards secrets like fingerprints and payment keys. Even the phone’s main processor can’t read them.

Novice. A walled-off part of a chip, with its own small processor, memory protection and encryption hardware, that keeps secrets such as keys and fingerprint data safe even if the main processor’s software is compromised.

Expert. A separate trust domain with a dedicated core, encrypted and authenticated memory, a true random number generator and a key hierarchy. Its physical design must also resist side-channel and fault-injection attacks.

Self-aligned gate

The gate is patterned first and then blocks the source/drain implant beneath it, so the source and drain end exactly at the gate edges without a separate alignment.

All three levels

Beginner. Using a switch’s gate, its on/off control, as the stencil for its two ends, so they always line up perfectly.

Novice. The gate is patterned first and then blocks the source/drain implant beneath it, so the source and drain end exactly at the gate edges without a separate alignment.

Expert. Removes the gate-overlap margin that misalignment once forced, cutting parasitic overlap capacitance. Sidewall spacers then offset the heavy implant from the channel; extension implants done before the spacers set the overlap that remains.

Self-heating

The temperature rise of a transistor from its own current. Fins and nanosheets are thin and surrounded by insulators that conduct heat poorly, so they warm up more than flat transistors did, which slows them and ages them faster.

All three levels

Beginner. A transistor warming itself up as it works, because the heat can’t escape fast enough from such a tiny space.

Novice. The temperature rise of a transistor from its own current. Fins and nanosheets are thin and surrounded by insulators that conduct heat poorly, so they warm up more than flat transistors did, which slows them and ages them faster.

Expert. Local channel heating set by power density and the thermal resistance to ambient. Confined fins and oxide-wrapped sheets raise that resistance; heat leaves mainly through the substrate, contacts and wiring. It shifts delay and accelerates BTI, HCI and electromigration.

Semiconductor

A solid with a full valence band and an empty conduction band separated by a modest band gap (1.12 eV in silicon). Heat, light and added impurities can put carriers into the bands, so its conductivity can be controlled over many orders of magnitude.

All three levels

Beginner. A material, like silicon, that carries electricity better than glass but much worse than metal. We can change how much it carries on purpose.

Novice. A solid with a full valence band and an empty conduction band separated by a modest band gap (1.12 eV in silicon). Heat, light and added impurities can put carriers into the bands, so its conductivity can be controlled over many orders of magnitude.

Expert. Distinguished from an insulator by degree: EgE_{\mathrm{g}} small enough that doping sets carrier densities at room temperature. Silicon wins on its native oxide, cost and crystal quality, not on mobility or band gap.

Sense amplifier

A circuit at the end of each bit line (or bit-line pair) that detects a small voltage difference and amplifies it to a full logic level. It lets the memory finish a read without waiting for the cell to swing the long bit line all the way.

All three levels

Beginner. A sensitive detector that notices a tiny change on a wire and turns it into a clear 1 or 0.

Novice. A circuit at the end of each bit line (or bit-line pair) that detects a small voltage difference and amplifies it to a full logic level. It lets the memory finish a read without waiting for the cell to swing the long bit line all the way.

Expert. Usually a clocked cross-coupled latch: fired once the bit-line split exceeds its offset plus noise, it regenerates to full rail. Its input offset (from transistor mismatch) sets the minimum swing, so it ties read speed to variation. In DRAM it also writes the value back into the cell.

Sequence pair

A way to describe how blocks sit relative to one another using two ordered lists of their names. Shuffling the lists gives a new arrangement, which lets a program search through many layouts.

All three levels

Beginner. A way to write down a layout as two lists of block names, so a computer can shuffle the lists to try new layouts.

Novice. A way to describe how blocks sit relative to one another using two ordered lists of their names. Shuffling the lists gives a new arrangement, which lets a program search through many layouts.

Expert. If aa comes before bb in both lists, aa is left of bb; before in the first and after in the second, aa is above bb. Decoding builds horizontal and vertical constraint graphs and takes longest paths. The (n!)2(n!)^2 space always contains an optimal packing, and annealing swaps names within or across the lists.

Sequence parallelism

Splitting work along the sequence (the list of tokens) instead of along the weights. In the Megatron sense it splits the layer parts that tensor parallelism leaves whole, such as normalization, so their memory is shared too.

All three levels

Beginner. Splitting a long piece of text into chunks handled by different chips.

Novice. Splitting work along the sequence (the list of tokens) instead of along the weights. In the Megatron sense it splits the layer parts that tensor parallelism leaves whole, such as normalization, so their memory is shared too.

Expert. Megatron sequence parallelism partitions LayerNorm and dropout along ss inside the TP group, replacing each all-reduce with a reduce-scatter plus all-gather at equal bandwidth and dividing all activations by tt.

SerDes

Serializer/deserializer: a transmitter turns wide, slow parallel data into a fast serial stream on one wire pair, and a receiver turns it back. Used for chip-to-chip links.

All three levels

Beginner. The special circuits that send data very fast between separate chips over a few wires.

Novice. Serializer/deserializer: a transmitter turns wide, slow parallel data into a fast serial stream on one wire pair, and a receiver turns it back. Used for chip-to-chip links.

Expert. Long-reach links with equalization and clock recovery, costing several pJ/bit and significant edge area; short on-wafer or on-interposer parallel links avoid most of that cost.

Set-associative cache

A cache divided into sets of a few lines (ways). An address picks one set; the line can go in any way of that set. Direct-mapped is 1 way per set; fully associative is one set holding every line.

All three levels

Beginner. A cache where each block can go in a few spots, not just one. That way, two blocks that want the same place can both stay.

Novice. A cache divided into sets of a few lines (ways). An address picks one set; the line can go in any way of that set. Direct-mapped is 1 way per set; fully associative is one set holding every line.

Expert. An N-way cache reads N tags in parallel and compares them, so more ways cost energy and hit time. Sets = capacity ÷ (line × ways); index bits = log2(sets). 8–16 ways is common in L1 and L2.

Setup time

The minimum time the data at a flip-flop’s input must be steady before the clock tick that captures it. Data that arrives later than that may be stored wrongly. The setup check: launch arrival + data delay must be no later than period + capture arrival − setup time.

All three levels

Beginner. How early the data must be ready before the clock tick so a memory cell grabs it correctly.

Novice. The minimum time the data at a flip-flop’s input must be steady before the clock tick that captures it. Data that arrives later than that may be stored wrongly. The setup check: launch arrival + data delay must be no later than period + capture arrival − setup time.

Expert. Stored in the Liberty library as a table indexed by the data and clock edge rates. The setup check pairs the slowest (late) data with the earliest possible capture clock, subtracts clock uncertainty, and adds back the common-path credit (CPPR).

Shallow trench isolation (STI)

Trenches etched into the silicon around each transistor’s active area and filled with silicon dioxide, then polished flat. They stop current leaking between neighboring transistors.

All three levels

Beginner. Little trenches filled with glass that wall off each switch from its neighbors.

Novice. Trenches etched into the silicon around each transistor’s active area and filled with silicon dioxide, then polished flat. They stop current leaking between neighboring transistors.

Expert. Pad oxide and nitride hard mask, trench etch, liner oxide, oxide fill, CMP stopping on the nitride, nitride strip. Replaced LOCOS around 250 nm because it doesn’t encroach on the active area.

Shared design database

A single copy of the whole design held in the computer’s memory: the gates and how they connect, where each one sits, the wires, and the timing goals. Every step inside the tool reads and changes it directly, instead of writing files and reading them back.

All three levels

Beginner. One shared copy of the chip design in memory. Every step reads and changes it, so nothing has to be saved and reloaded in between.

Novice. A single copy of the whole design held in the computer’s memory: the gates and how they connect, where each one sits, the wires, and the timing goals. Every step inside the tool reads and changes it directly, instead of writing files and reading them back.

Expert. Supports incremental optimization. When one engine resizes or moves a gate, the others are told through callbacks (functions registered to run on each change), so the timer recomputes only what changed and thousands of trial changes per second can be evaluated. OpenDB (in OpenROAD) and OpenAccess are examples; commercial suites each have their own.

Shared memory (LDS)

Fast on-chip memory inside each GPU core that the programmer controls directly. Threads in the same block use it to share data and to reuse it without going back to main memory. AMD calls it the local data share.

All three levels

Beginner. A small, fast notepad inside each GPU core that a team of workers can share.

Novice. Fast on-chip memory inside each GPU core that the programmer controls directly. Threads in the same block use it to share data and to reuse it without going back to main memory. AMD calls it the local data share.

Expert. A banked, software-managed scratchpad (32 four-byte banks on NVIDIA) carved from the same SRAM as L1 on recent NVIDIA parts. Bank conflicts serialize accesses. It is the staging area for GEMM and attention tiles.

Shift and capture

Shift: with scan enable on, clock the chains once per flip-flop in the chain to load a pattern (and push out the previous result). Capture: with scan enable off, give one normal clock tick (two for timing tests) so every flip-flop stores what the logic computed.

All three levels

Beginner. The two steps of a scan test: slide a pattern into the chip’s memory cells, then let the chip take one step and catch the result.

Novice. Shift: with scan enable on, clock the chains once per flip-flop in the chain to load a pattern (and push out the previous result). Capture: with scan enable off, give one normal clock tick (two for timing tests) so every flip-flop stores what the logic computed.

Expert. Shift is slow and long (one cycle per bit of the longest chain); capture is a few cycles and may run at full speed. Unloading pattern kk overlaps loading pattern k+1k+1. Shift power and capture power are separate limits.

Shmoo plot

A grid of pass/fail results for one test as two conditions are swept, usually supply voltage against clock speed, at a fixed temperature. The boundary between pass and fail shows the chip’s real operating limits.

All three levels

Beginner. A grid that shows where a chip works and where it fails as two settings change, such as speed and voltage (how hard the electricity pushes).

Novice. A grid of pass/fail results for one test as two conditions are swept, usually supply voltage against clock speed, at a fixed temperature. The boundary between pass and fail shows the chip’s real operating limits.

Expert. The shape is diagnostic. A boundary that slopes with voltage is the normal speed limit; fails at high voltage and low speed, holes inside the pass region, or ragged edges point to hold-time, noise or marginal analog problems. Repeated across many parts and temperatures to build the operating envelope.

Shoreline (beachfront)

The length of die edge devoted to an interface. Side-by-side links can only use bumps in a strip behind the facing edges, so bandwidth is quoted per millimeter of shoreline.

All three levels

Beginner. The edge of a chip that faces its neighbor and holds the connections between them. Like a coast, there’s only so much of it.

Novice. The length of die edge devoted to an interface. Side-by-side links can only use bumps in a strip behind the facing edges, so bandwidth is quoted per millimeter of shoreline.

Expert. Grows only with the perimeter while logic grows with area, which is why edge-limited I/O tightens as dies grow. UCIe specifies bump maps with a fixed shoreline per module, so finer pitch buys a shallower PHY, not a wider one.

Short-channel effects

The set of ways a transistor misbehaves once its gate is too short for the gate to dominate the channel: the drain voltage lowers the threshold (DIBL), the turn-off gets less steep (worse subthreshold swing), and the threshold falls as the gate gets shorter (roll-off).

All three levels

Beginner. Problems a transistor gets when it’s so short that the far end, the drain, starts to fight the gate. It leaks and doesn’t switch off cleanly.

Novice. The set of ways a transistor misbehaves once its gate is too short for the gate to dominate the channel: the drain voltage lowers the threshold (DIBL), the turn-off gets less steep (worse subthreshold swing), and the threshold falls as the gate gets shorter (roll-off).

Expert. 2D electrostatic effects that appear when LL is only a few natural lengths λ\lambda: VTV_{\mathrm{T}} roll-off, DIBL, degraded SS and, in the limit, punch-through. They grow roughly as exp⁡(−L/2λ)\exp(-L/2\lambda), so the fix is to shrink λ\lambda (thinner channel, thinner EOT, more gates) rather than only LL.

Short-circuit current

During a transition, the input passes through a range where both the NMOS and PMOS conduct, so current flows straight from the supply to ground for a moment.

All three levels

Beginner. A short burst of wasted electricity when a gate flips, while both its switches are partly on at once.

Novice. During a transition, the input passes through a range where both the NMOS and PMOS conduct, so current flows straight from the supply to ground for a moment.

Expert. Crowbar current during input transitions, when Vtn<Vin<VDD−∣Vtp∣V_{\mathrm{tn}} < V_{\mathrm{in}} < V_{\mathrm{DD}} - |V_{\mathrm{tp}}|. Usually under 10% of dynamic power when input and output edge rates are comparable; it grows with slow input slews and disappears when VDD<Vtn+∣Vtp∣V_{\mathrm{DD}} < V_{\mathrm{tn}} + |V_{\mathrm{tp}}|.

Si/SiGe superlattice

Alternating crystal layers of silicon and silicon-germanium, each a few nanometers thick, grown on the wafer one atomic layer after another. The silicon layers become the nanosheets; the silicon-germanium is later etched away.

All three levels

Beginner. A stack of very thin layers that take turns: silicon, then silicon mixed with germanium, then silicon again. Nanosheet transistors start from it.

Novice. Alternating crystal layers of silicon and silicon-germanium, each a few nanometers thick, grown on the wafer one atomic layer after another. The silicon layers become the nanosheets; the silicon-germanium is later etched away.

Expert. Epitaxial Si/Si1−xGex\mathrm{Si/Si_{1-x}Ge_x} stack (one published flow used about 30% Ge and layers about 9 nm thick). Sheet thickness and spacing are set by layer growth rather than by etching, as a fin’s width is.

side-channel attack

An attack that learns a secret by measuring a physical effect of the computation, such as how much power it draws, how long it takes or the radio noise it gives off, instead of breaking the encryption itself.

All three levels

Beginner. Stealing a secret by watching how a chip behaves, like how much power it uses, instead of cracking the code.

Novice. An attack that learns a secret by measuring a physical effect of the computation, such as how much power it draws, how long it takes or the radio noise it gives off, instead of breaking the encryption itself.

Expert. Includes simple and differential power analysis (SPA, DPA), electromagnetic analysis and timing attacks. Countermeasures include masking (randomizing intermediate values), hiding (constant-power logic, added noise) and protocol-level key updates, and they are validated by measurement on silicon.

Sidewall spacer

A conformal insulating film deposited over the gates and etched straight down, which leaves material only on the gate’s vertical sides. It sets how far the heavy source/drain doping sits from the channel.

All three levels

Beginner. A thin strip of insulator left on each side of a gate, used to keep the next implant a small distance away from it.

Novice. A conformal insulating film deposited over the gates and etched straight down, which leaves material only on the gate’s vertical sides. It sets how far the heavy source/drain doping sits from the channel.

Expert. Formed by conformal deposition and anisotropic etch-back, so its width is set by film thickness, not lithography. The same trick at larger scale is spacer patterning (SADP/SAQP) for fins, gates and tight metal pitches.

Signature aliasing

The chance that a faulty chip’s squeezed-down result happens to equal the good result, hiding the fault. For a kk-bit MISR it is about 1 in 2k2^k.

All three levels

Beginner. When a broken chip’s fingerprint happens to match a good chip’s, so the flaw slips by.

Novice. The chance that a faulty chip’s squeezed-down result happens to equal the good result, hiding the fault. For a kk-bit MISR it is about 1 in 2k2^k.

Expert. Negligible for practical sizes (32 bits or more), but never zero. Simpler compactors such as parity or counting 1s alias far more often.

Signoff

The checks run on the finished layout before it goes to the factory: that signals arrive on time, that every shape follows the factory’s rules, that the drawing matches the intended circuit, and that the power wiring is strong enough. Each check has an owner who approves it.

All three levels

Beginner. The last set of checks a chip design must pass before it goes to the factory.

Novice. The checks run on the finished layout before it goes to the factory: that signals arrive on time, that every shape follows the factory’s rules, that the drawing matches the intended circuit, and that the power wiring is strong enough. Each check has an owner who approves it.

Expert. The final verification of the routed, filled layout: timing in every scenario on extracted parasitics, physical verification (DRC, LVS, ERC, antenna, density), power integrity (IR drop and electromigration) and equivalence checking. It is run with trusted (“golden”) tools and the foundry’s own rule files, and the deliverable is a checklist in which every remaining exception has a recorded, reviewed reason.

Silicide (salicide)

A metal (titanium, cobalt or nickel) is deposited and heated. It reacts only where it touches bare silicon, forming a low-resistance silicide; the leftover metal is then etched away. No mask is needed, hence “self-aligned silicide”.

All three levels

Beginner. A thin layer of metal mixed with silicon on a switch’s ends and gate. It lets electricity get in and out more easily.

Novice. A metal (titanium, cobalt or nickel) is deposited and heated. It reacts only where it touches bare silicon, forming a low-resistance silicide; the leftover metal is then etched away. No mask is needed, hence “self-aligned silicide”.

Expert. Cuts sheet and contact resistance of source, drain and gate. Older flows used titanium silicide, annealed in nitrogen so the same step also formed titanium nitride.

Silicon bridge

A small piece of silicon with fine wiring layers, embedded in the package substrate (or in a molded layer) under the facing edges of two dies, so only that region needs fine-pitch bumps.

All three levels

Beginner. A small strip of silicon with very fine wires. It is tucked under the edges of two chips to join them.

Novice. A small piece of silicon with fine wiring layers, embedded in the package substrate (or in a molded layer) under the facing edges of two dies, so only that region needs fine-pitch bumps.

Expert. Gives interposer-class wiring density between neighbors without a full interposer under every die: the rest of each die keeps coarse bumps, and power need not pass through silicon. Each link can get its own bridge; examples include Intel’s EMIB.

silicon interposer

A slice of silicon with very fine wiring but no transistors. Several chips and memory stacks sit side by side on top of it, and it connects them to each other and, through vertical holes filled with metal, to the package below. This arrangement is called 2.5D packaging.

All three levels

Beginner. A thin slice of silicon that works like a tiny circuit board. It wires several chips together side by side.

Novice. A slice of silicon with very fine wiring but no transistors. Several chips and memory stacks sit side by side on top of it, and it connects them to each other and, through vertical holes filled with metal, to the package below. This arrangement is called 2.5D packaging.

Expert. Passive silicon carrying micrometer-scale routing between dies, at far higher density than an organic package substrate, with through-silicon vias down to the substrate. It can itself exceed one reticle, and its yield, cost and warpage become part of the product’s risk.

Silicon photonics

Building optical waveguides, modulators and detectors on silicon wafers with chip-making tools. Silicon does not emit light well, so the laser usually comes from a separate chip.

All three levels

Beginner. Making the parts that steer and switch light on silicon chips, using the same factories that make computer chips.

Novice. Building optical waveguides, modulators and detectors on silicon wafers with chip-making tools. Silicon does not emit light well, so the laser usually comes from a separate chip.

Expert. Waveguides, Mach-Zehnder or micro-ring modulators and germanium photodiodes in a CMOS-like process. Ring modulators are small and efficient but must be thermally tuned; MZMs are larger but more tolerant. Light is supplied by III-V lasers, integrated or external.

SIMD (single instruction, multiple data)

A processor design where one instruction operates on a short vector of values at the same time, for example adding eight pairs of numbers in one step.

All three levels

Beginner. One command that does the same thing to a whole row of numbers at once.

Novice. A processor design where one instruction operates on a short vector of values at the same time, for example adding eight pairs of numbers in one step.

Expert. Vector width is exposed to software, which must pack data into vectors and handle divergence itself with masks. GPUs build SIMD datapaths but present them through the SIMT thread model.

SIMT (single instruction, multiple threads)

The GPU execution model. You write code for one thread; the hardware groups threads into warps and issues each instruction once for the whole group, each thread using its own data and registers.

All three levels

Beginner. A way of running a program where a group of workers all do the same step at the same moment, each on its own piece of data.

Novice. The GPU execution model. You write code for one thread; the hardware groups threads into warps and issues each instruction once for the whole group, each thread using its own data and registers.

Expert. Like SIMD in hardware, but the ISA describes one scalar thread and the hardware handles grouping and divergence with active masks. Correctness ignores warp width; performance doesn’t.

Simulated annealing

A search method that tries random changes, always keeps improvements, and sometimes keeps a change that makes things worse, with a probability that shrinks as a “temperature” setting is lowered. Early on it explores freely; later it only accepts improvements.

All three levels

Beginner. A way to search for a good layout by trying random swaps, and sometimes keeping a worse one early on to get out of a dead end.

Novice. A search method that tries random changes, always keeps improvements, and sometimes keeps a change that makes things worse, with a probability that shrinks as a “temperature” setting is lowered. Early on it explores freely; later it only accepts improvements.

Expert. Propose a move, accept it if cost falls, otherwise accept with probability exp⁡(−Δcost/T)\exp(-\Delta\text{cost}/T), and lower TT geometrically. The basis of TimberWolf, the classic annealing placer of the 1980s. Quality can be high, but it converges slowly and scales poorly to millions of cells.

Simulation (logic simulation)

Running a software model of the design on an ordinary computer. The simulator works out the value of every wire, clock cycle by clock cycle, while a test feeds in inputs. It is far slower than a real chip, but any signal can be inspected at any moment.

All three levels

Beginner. Running the chip design as a program on an ordinary computer, to see what the chip would do.

Novice. Running a software model of the design on an ordinary computer. The simulator works out the value of every wire, clock cycle by clock cycle, while a test feeds in inputs. It is far slower than a real chip, but any signal can be inspected at any moment.

Expert. The simulator interprets or compiles the RTL into a program that computes every signal each cycle. It can run any testbench code and show any signal, but it runs far slower than silicon, and slower still as the design grows, which is why very long software runs move to emulation.

Simultaneous multithreading (SMT)

A core design that keeps several hardware threads and issues instructions from more than one of them in the same clock cycle, filling execution slots that one thread would leave empty. Intel’s brand name is Hyper-Threading.

All three levels

Beginner. Letting one core work on two (or more) jobs in the same tick of the clock, so gaps left by one job are filled by the other.

Novice. A core design that keeps several hardware threads and issues instructions from more than one of them in the same clock cycle, filling execution slots that one thread would leave empty. Intel’s brand name is Hyper-Threading.

Expert. Fills both vertical waste (cycles with nothing to issue, e.g. on a miss) and horizontal waste (part-filled cycles). Raises throughput by tens of percent at the cost of per-thread speed, shared-cache contention and weaker isolation between threads.

Single-event effect (SEE)

Any effect caused by the electric charge one particle leaves behind as it passes through a chip. Some are harmless once fixed, like a flipped bit or a brief glitch; others, like latch-up, can destroy the chip.

All three levels

Beginner. What happens when one fast particle hits a chip: a flipped bit, a glitch, or worse.

Novice. Any effect caused by the electric charge one particle leaves behind as it passes through a chip. Some are harmless once fixed, like a flipped bit or a brief glitch; others, like latch-up, can destroy the chip.

Expert. Measured in a beam test as a cross-section (errors divided by particles per cm²) at several values of LET, the energy a particle deposits per unit of path length. Combining that curve with an orbit’s particle population predicts an event rate per device per day.

Single-event functional interrupt (SEFI)

A particle strike in a chip’s control circuits that makes it freeze or misbehave. A soft one clears with a reset or reload; a hard one needs the power switched off and on.

All three levels

Beginner. A particle hit that makes a chip freeze or act strangely until it is restarted.

Novice. A particle strike in a chip’s control circuits that makes it freeze or misbehave. A soft one clears with a reset or reload; a hard one needs the power switched off and on.

Expert. Often an upset in a state machine, a configuration register or test logic. Because it can raise supply current, it can look like latch-up on a tester, so telling them apart needs visibility into the chip’s internal state.

Single-event latch-up (SEL)

A particle strike switches on an unintended path through the layers of silicon in a chip, letting a large current flow from the power supply to ground. It stays on until the power is removed and can destroy the chip.

All three levels

Beginner. A particle hit opens a hidden short circuit inside the chip. Too much electricity flows, and it can burn the chip out.

Novice. A particle strike switches on an unintended path through the layers of silicon in a chip, letting a large current flow from the power supply to ground. It stays on until the power is removed and can destroy the chip.

Expert. The unintended path is a parasitic thyristor formed by neighboring n- and p-type regions. Specified as an LET threshold below which latch-up must not occur (for example above 125 MeV·cm²/mg). Countered by the process (silicon-on-insulator is inherently immune), guard rings and well taps in layout, and current limiting with power cycling in the system.

Single-event transient (SET)

A brief false voltage pulse in a logic circuit caused by a particle strike. It does harm only if a flip-flop happens to store it when the clock ticks.

All three levels

Beginner. A quick false blip on a wire, caused by a particle hit.

Novice. A brief false voltage pulse in a logic circuit caused by a particle strike. It does harm only if a flip-flop happens to store it when the clock ticks.

Expert. A glitch on a data path that arrives at the clock edge is captured by all three copies of a basic TMR flip-flop at once, so TMR alone doesn’t stop it. Transient filters, tripled logic and voters, and timing offsets between copies address this.

Single-event upset (SEU)

A stored bit, in a memory cell or a flip-flop (a tiny circuit that holds one bit), changes value because a particle struck it. Nothing is damaged, but the data is now wrong.

All three levels

Beginner. A particle flips a stored 1 to a 0, or a 0 to a 1. Nothing breaks, but the data is wrong.

Novice. A stored bit, in a memory cell or a flip-flop (a tiny circuit that holds one bit), changes value because a particle struck it. Nothing is damaged, but the data is now wrong.

Expert. Handled by error-correcting codes on memories, TMR or DICE flip-flops on registers, and scrubbing. A single strike can flip several neighboring bits (a multi-bit upset), which defeats simple codes and redundant copies placed close together.

Single-point fault metric (SPFM)

The share of a chip’s possible random hardware failures that either can’t cause harm on their own or are caught by a safety mechanism. Targets: at least 90% for ASIL B, 97% for C and 99% for D.

All three levels

Beginner. Out of all the faults a chip could have, the share that are caught or can’t cause harm on their own.

Novice. The share of a chip’s possible random hardware failures that either can’t cause harm on their own or are caught by a safety mechanism. Targets: at least 90% for ASIL B, 97% for C and 99% for D.

Expert. One minus the fraction of safety-relevant failure rate that is single-point or residual (escaping the safety mechanism). Computed from FMEDA failure rates and diagnostic coverage, with the coverage claims backed by fault-injection evidence.

Skew group

A set of clock sinks that CTS must deliver the clock to at about the same time. Sinks in different groups don’t need to be matched with each other.

All three levels

Beginner. A set of memory cells that must all get the clock at about the same time.

Novice. A set of clock sinks that CTS must deliver the clock to at about the same time. Sinks in different groups don’t need to be matched with each other.

Expert. Skew groups tell CTS which sinks actually exchange data: they separate unrelated clock domains, leave out test-only flip-flops, balance across generated clocks, and give memory blocks or I/O flip-flops their own latency targets.

Skewed gate

A gate whose PMOS and NMOS are sized unequally on purpose, which moves its switching threshold away from the middle and makes one edge faster at the expense of the other.

All three levels

Beginner. A gate deliberately built so one kind of output change (up or down) happens faster than the other.

Novice. A gate whose PMOS and NMOS are sized unequally on purpose, which moves its switching threshold away from the middle and makes one edge faster at the expense of the other.

Expert. Beta ratio βp/βn≠1\beta_{\mathrm{p}}/\beta_{\mathrm{n}} \ne 1 (relative to the balanced point). High-skew gates favor the rising output, low-skew the falling one; the cost is a smaller noise margin on one side and a slower opposite edge. Used on critical edges in pulsed and domino-style paths.

Skewed input (wavefront)

Delaying the data entering row rr of a systolic array by rr clock cycles. Because values move one PE per cycle, the delay makes matching operands arrive at each PE together, and the active PEs form a diagonal wave.

All three levels

Beginner. Feeding the numbers into the grid in a staircase, one tick later for each row, so the right numbers meet in the right cell.

Novice. Delaying the data entering row rr of a systolic array by rr clock cycles. Because values move one PE per cycle, the delay makes matching operands arrive at each PE together, and the active PEs form a diagonal wave.

Expert. Implemented with delay registers or staggered buffer reads at the array edge; the outputs leave equally skewed and are de-skewed on the way out. The skew is one source of the fill-and-drain cycles in the runtime model.

Skin effect

At high frequencies current flows only in a thin layer near a conductor’s surface. The layer gets thinner as frequency rises, so resistance, and loss, grow roughly with the square root of frequency.

All three levels

Beginner. At high speeds, electricity crowds into the thin outer skin of a wire, so the wire acts thinner and wastes more signal.

Novice. At high frequencies current flows only in a thin layer near a conductor’s surface. The layer gets thinner as frequency rises, so resistance, and loss, grow roughly with the square root of frequency.

Expert. Skin depth δ=ρ/(πfμ)\delta = \sqrt{\rho / (\pi f \mu)}; conductor loss ∝f\propto \sqrt{f}. Together with dielectric loss (∝ftan⁡δ\propto f \tan\delta) it sets copper channel loss; surface roughness adds more above ~10 GHz.

SKY130

SkyWater’s 130 nm process, published as an open-source process design kit (PDK): transistor models, design rules and standard-cell libraries anyone can read.

All three levels

Beginner. A free, open recipe for making chips from the company SkyWater. Anyone can download it and design a chip with it.

Novice. SkyWater’s 130 nm process, published as an open-source process design kit (PDK): transistor models, design rules and standard-cell libraries anyone can read.

Expert. A mature 180–130 nm hybrid technology with 1.8 V core devices, five metal levels plus a local-interconnect layer, and open SPICE models, which makes it a source of real device numbers.

Slack

The spare time on a timing path: when a signal is required at its destination minus when it actually arrives. Negative slack means it arrives too late and the path fails.

All three levels

Beginner. How much spare time a signal has before it would arrive too late. Positive is good; negative means it is late.

Novice. The spare time on a timing path: when a signal is required at its destination minus when it actually arrives. Negative slack means it arrives too late and the path fails.

Expert. Required time minus arrival time at a path’s endpoint, for a setup or hold check. The worst value in a group of paths is WNS (worst negative slack) and the sum of the negative values is TNS (total negative slack). Synthesis computes it with a perfect clock and guessed wires, so teams aim for some positive margin.

Slicing floorplan

A layout you can build by repeatedly cutting a rectangle in two with straight cuts that run all the way across it. It can be stored as a tree whose branches are cuts and whose leaves are blocks.

All three levels

Beginner. A layout you can make by cutting a rectangle in two again and again, with each cut going all the way across.

Novice. A layout you can build by repeatedly cutting a rectangle in two with straight cuts that run all the way across it. It can be stored as a tree whose branches are cuts and whose leaves are blocks.

Expert. Represented by slicing trees or Polish expressions and often searched with simulated annealing. Cheap to evaluate, since block shape functions combine up the tree and give the minimum area in polynomial time, but it cannot represent non-slicing packings, in which no straight cut crosses the whole layout without splitting a block.

Slotting

Cutting small rectangular holes into wide metal, often power wires, to bring a window under the maximum density rule and keep the wide metal from lifting off during polishing.

All three levels

Beginner. Cutting holes into very wide metal wires so they don’t polish badly in the factory.

Novice. Cutting small rectangular holes into wide metal, often power wires, to bring a window under the maximum density rule and keep the wide metal from lifting off during polishing.

Expert. It narrows the wire’s cross-section, so its resistance, voltage drop and current density must be rechecked. Slots must avoid the areas where vias land. Some processes generate slots during mask preparation; others want them drawn on a slot datatype.

SmartNIC

A network interface card that can be programmed to process packets itself, for example applying a cloud’s security and routing rules, so the server’s CPUs are freed for customers. It can be built from an FPGA, many small CPU cores, or a fixed-function chip.

All three levels

Beginner. A network card with its own brain. It handles network jobs so the main processor doesn’t have to.

Novice. A network interface card that can be programmed to process packets itself, for example applying a cloud’s security and routing rules, so the server’s CPUs are freed for customers. It can be built from an FPGA, many small CPU cores, or a fixed-function chip.

Expert. A NIC with programmable data-plane offload (flow tables, encapsulation, encryption, load balancing). FPGA versions sit as a bump in the wire between host and switch; SoC versions trade latency and jitter for software programmability; ASIC versions trade flexibility for efficiency.

Snooping (cache coherence)

A coherence scheme in which each miss is broadcast on a shared bus or network and every cache checks (snoops) its tags. Simple and fast for a few cores; the broadcasts don’t scale to many.

All three levels

Beginner. A way to keep copies in agreement where every cache listens to every request, like everyone hearing an announcement.

Novice. A coherence scheme in which each miss is broadcast on a shared bus or network and every cache checks (snoops) its tags. Simple and fast for a few cores; the broadcasts don’t scale to many.

Expert. Relies on an ordered broadcast medium as the serialization point. Traffic and tag-lookup energy grow with core count, so large chips use directories or snoop filters that track which caches might hold a line.

SoC

System-on-chip: a single chip holding most of a computer, including the main processors, graphics, memory controllers, accelerators, input/output and often radio and analog circuits.

All three levels

Beginner. A system-on-chip: one chip that holds almost everything a device needs, from processors to radios.

Novice. System-on-chip: a single chip holding most of a computer, including the main processors, graphics, memory controllers, accelerators, input/output and often radio and analog circuits.

Expert. Mostly an integration problem: licensed and reused IP, the on-chip interconnect, clock and power domains, and chip-level verification of interactions no single IP team owns.

SOI (silicon on insulator)

A wafer with a thin layer of crystalline silicon on top of a buried oxide (insulator). Transistors built in a thin enough top layer have no deep silicon underneath for current to leak through.

All three levels

Beginner. A chip base where the transistors sit on a very thin layer of silicon above a layer of glass. The glass blocks leak paths underneath.

Novice. A wafer with a thin layer of crystalline silicon on top of a buried oxide (insulator). Transistors built in a thin enough top layer have no deep silicon underneath for current to leak through.

Expert. Ultra-thin-body SOI and double-gate SOI were the first answers to short-channel effects: making tSit_{\mathrm{Si}} small shrinks λ\lambda. Fully depleted SOI (FD-SOI) is still used for low-power planar processes; FinFETs moved the same idea into a vertical fin on bulk wafers.

Source and drain

Two heavily doped regions on either side of the gate. Carriers enter the channel at the source and leave at the drain. In an NMOS the source is the lower-voltage end; in a PMOS it is the higher-voltage end.

All three levels

Beginner. The two ends of a transistor switch. When the switch is on, electricity flows from one to the other.

Novice. Two heavily doped regions on either side of the gate. Carriers enter the channel at the source and leave at the drain. In an NMOS the source is the lower-voltage end; in a PMOS it is the higher-voltage end.

Expert. Physically symmetric in a standard MOSFET: which end is the source is set by the voltages, not the layout. Their junctions to the body add diffusion capacitance comparable to the gate capacitance.

Spare cell

Unused logic cells (inverters, NAND gates, flip-flops) scattered across the chip with their inputs tied to fixed values. If a bug turns up late, a fix can be built by rewiring nearby spares, which changes only the metal wiring layers.

All three levels

Beginner. Extra unused parts scattered around the chip, so a late bug fix can be made by changing only the wires.

Novice. Unused logic cells (inverters, NAND gates, flip-flops) scattered across the chip with their inputs tied to fixed values. If a bug turns up late, a fix can be built by rewiring nearby spares, which changes only the metal wiring layers.

Expert. Spread uniformly during placement so any late functional fix (a metal-only ECO) has spare gates nearby. They cost area and leakage. Too few spares near a patch means long ECO wires, which can cause timing violations and congestion.

Spare core (redundancy)

Extra cores (or memory rows, or links) built into a design. After testing, broken ones are switched off and spares are wired in their place, so a chip with a few defects still works.

All three levels

Beginner. An extra copy of a small part of the chip, kept in reserve so it can take over if the original turns out to be broken.

Novice. Extra cores (or memory rows, or links) built into a design. After testing, broken ones are switched off and spares are wired in their place, so a chip with a few defects still works.

Expert. Repair by remapping: test finds failing units, fuses or configuration record the map, and the network routes around them. Spares cost area and power up front; their placement decides which defect patterns can be repaired.

Sparse format (metadata)

A storage layout that keeps nonzero values and an index (metadata) saying where each sits. Common ones are CSR and CSC (compressed sparse row/column). The index costs memory and lookup time.

All three levels

Beginner. A small way to store a table full of zeros. Keep only the nonzero numbers, plus a note of where each one goes.

Novice. A storage layout that keeps nonzero values and an index (metadata) saying where each sits. Common ones are CSR and CSC (compressed sparse row/column). The index costs memory and lookup time.

Expert. CSR/CSC store a column (or row) index per nonzero plus pointers per row; with 8-bit values and 16-bit indices the metadata is 200% of the value storage. 2:4 needs only a 2-bit position per kept value; block formats need one index per block.

Sparsity

The fraction of values in a weight matrix or activation tensor that are exactly zero. 75% sparsity means three out of four values are zero. Weight sparsity comes from pruning; activation sparsity appears naturally, for example after a ReLU sets negatives to zero.

All three levels

Beginner. How many of the numbers in an AI model are zero. Anything times zero is zero. So that work can be skipped.

Novice. The fraction of values in a weight matrix or activation tensor that are exactly zero. 75% sparsity means three out of four values are zero. Weight sparsity comes from pruning; activation sparsity appears naturally, for example after a ReLU sets negatives to zero.

Expert. Static (weights, known ahead of time) or dynamic (activations, known only at run time), and unstructured or structured. The speedup it buys depends on the pattern and the hardware, not on the zero count alone.

Spatial (dataflow) architecture

A design that lays a computation out across many compute units at once, each with its own small memory, and passes data directly between them instead of sending everything through one shared memory. Contrast with CPUs and GPUs, where a central control steers many arithmetic units that all fetch from a shared memory hierarchy.

All three levels

Beginner. A chip design where many small workers sit side by side, each doing one part of the job and handing results straight to its neighbors.

Novice. A design that lays a computation out across many compute units at once, each with its own small memory, and passes data directly between them instead of sending everything through one shared memory. Contrast with CPUs and GPUs, where a central control steers many arithmetic units that all fetch from a shared memory hierarchy.

Expert. Computation is distributed in space: operators, or slices of them, are pinned to processing elements and data moves over an exposed on-chip network, usually under a schedule fixed at compile time. Covers systolic arrays, CGRAs and many-core meshes with local SRAM.

Spec freeze

A milestone after which the requirements change only through a formal change request that weighs the effect on schedule, cost and the other requirements.

All three levels

Beginner. The point where the team agrees to stop changing the plan, so building can start in earnest.

Novice. A milestone after which the requirements change only through a formal change request that weighs the effect on schedule, cost and the other requirements.

Expert. In practice a series: requirements freeze, architecture freeze, then feature freeze in the design code. Each later change needs an impact analysis across design code, verification, physical layout, software and manufacturing test.

Specification (spec)

The set of documents that turns a product idea into numbered, measurable requirements: what the chip must do, how fast, how much power it may use, how big and costly it may be, what it must connect to, and which test, safety and security rules it must meet. Every later stage is checked against it.

All three levels

Beginner. The written list of everything a chip must do and the limits it must stay within, agreed before anyone designs it.

Novice. The set of documents that turns a product idea into numbered, measurable requirements: what the chip must do, how fast, how much power it may use, how big and costly it may be, what it must connect to, and which test, safety and security rules it must meet. Every later stage is checked against it.

Expert. A controlled, versioned baseline: requirements with IDs, numeric targets, the conditions each applies under and the way each will be verified, plus the architecture and block-level specs derived from them. After it is frozen, changes go through a formal review, because each one ripples into the design code, the verification plan, the physical layout and the schedule.

Spectre (and Meltdown)

A family of attacks published in 2018. They make a processor speculatively run instructions that touch secret data; the wrong-path work is undone, but it leaves traces in the cache that can be measured. Meltdown is a related attack on out-of-order execution past a permission check.

All three levels

Beginner. A kind of attack, found in 2018, that tricks a processor’s guessing into leaving clues about secrets in its fast memory.

Novice. A family of attacks published in 2018. They make a processor speculatively run instructions that touch secret data; the wrong-path work is undone, but it leaves traces in the cache that can be measured. Meltdown is a related attack on out-of-order execution past a permission check.

Expert. Transient-execution attacks: mistrained predictors (Spectre) or delayed fault handling (Meltdown) let transient instructions encode secrets in microarchitectural state, read back through a cache timing side channel. Mitigations span hardware, OS (KPTI) and software (fences, process isolation).

Speculative decoding

An inference trick in which a small draft model proposes several tokens and the large model verifies them in one pass, keeping the ones it agrees with. The output is the same as the large model’s.

All three levels

Beginner. A small, quick model guesses the next few words, and the big model checks all the guesses at once.

Novice. An inference trick in which a small draft model proposes several tokens and the large model verifies them in one pass, keeping the ones it agrees with. The output is the same as the large model’s.

Expert. It spends spare compute in a memory-bound regime: verifying k tokens costs about one weight read, so accepted tokens per weight read rise. Gains depend on the draft’s acceptance rate and shrink as batching already uses the compute.

Speculative execution

Executing instructions before it is certain they should run, usually past a predicted branch. If the prediction was right, the results are kept; if wrong, they are discarded and execution restarts on the correct path.

All three levels

Beginner. Doing work early based on a guess, and throwing it away if the guess was wrong.

Novice. Executing instructions before it is certain they should run, usually past a predicted branch. If the prediction was right, the results are kept; if wrong, they are discarded and execution restarts on the correct path.

Expert. Architectural state is protected by in-order commit, but microarchitectural state (caches, predictors, buffers) changed by transient instructions is not rolled back, which Spectre and Meltdown exploit.

SPEF

Standard Parasitic Exchange Format: a text file listing, for every wire in the design, its resistance and capacitance. The timing tool reads it to compute real wire delays.

All three levels

Beginner. The file that lists the drag on every wire, so the timing tool can use it.

Novice. Standard Parasitic Exchange Format: a text file listing, for every wire in the design, its resistance and capacitance. The timing tool reads it to compute real wire delays.

Expert. For each net: a resistor network between its pins, grounded capacitances, and coupling capacitors to named neighbor nets. One SPEF is written per RC corner. The timer must report how many nets were annotated, because a net missing from the SPEF silently falls back to an estimate.

SPICE

A family of circuit simulators that solve the equations of a transistor-level netlist using device models from the process kit. It is accurate but far too slow to run on a whole chip, so it is used on small circuits such as single cells.

All three levels

Beginner. A computer program that simulates a circuit’s transistors in detail to predict exactly how its voltages and currents change over time.

Novice. A family of circuit simulators that solve the equations of a transistor-level netlist using device models from the process kit. It is accurate but far too slow to run on a whole chip, so it is used on small circuits such as single cells.

Expert. Transient analysis with compact device models (BSIM and similar) on the extracted netlist, including parasitic resistance and capacitance. Characterizers drive it in batches, one simulation (or a sweep) per arc, slew, load and corner; ngspice and Xyce are open-source options.

Split manufacturing

Building a chip in two factories: an untrusted one makes the transistors and lower wiring layers, and a trusted one adds the upper wiring layers that complete the connections.

All three levels

Beginner. Making a chip in two factories, so the one that isn’t trusted never sees the whole design.

Novice. Building a chip in two factories: an untrusted one makes the transistors and lower wiring layers, and a trusted one adds the upper wiring layers that complete the connections.

Expert. The lower part is called the front end of line (FEOL) and the upper part the back end of line (BEOL). Security depends on how much of the hidden wiring can be inferred from the FEOL. Proximity attacks exploit the fact that placement tools put connected cells near each other; defenses include placement perturbation, wire lifting and layout obfuscation.

SQNR (signal-to-quantization-noise ratio)

The ratio of the signal’s power to the power of the rounding error, in decibels (dB). Every extra bit of precision adds about 6 dB, which means roughly halving the error.

All three levels

Beginner. A score for how close a rounded copy is to the original: higher is better.

Novice. The ratio of the signal’s power to the power of the rounding error, in decibels (dB). Every extra bit of precision adds about 6 dB, which means roughly halving the error.

Expert. 10log⁡10 ⁣(∑x2/∑(x^−x)2)10 \log_{10}\!\big(\sum x^2 / \sum (\hat{x} - x)^2\big). For a uniform quantizer the noise power is Δ2/12\Delta^2/12, so each added bit (Δ\Delta halved) adds 20log⁡102≈6.0220 \log_{10} 2 \approx 6.02 dB. For floats SQNR depends on mantissa bits, not magnitude, as long as values stay in range.

Square-law (Shockley) model

The classic first-order MOSFET model: zero current below threshold, ID=β(VGS−Vt−VDS/2)VDSI_{\mathrm{D}} = \beta(V_{\mathrm{GS}} - V_{\mathrm{t}} - V_{\mathrm{DS}}/2)V_{\mathrm{DS}} in the linear region and ID=(β/2)(VGS−Vt)2I_{\mathrm{D}} = (\beta/2)(V_{\mathrm{GS}} - V_{\mathrm{t}})^2 in saturation, where β\beta depends on the device size and the process.

All three levels

Beginner. The simplest textbook rule for a transistor: push the gate twice as far past its tipping point and four times as much current flows.

Novice. The classic first-order MOSFET model: zero current below threshold, ID=β(VGS−Vt−VDS/2)VDSI_{\mathrm{D}} = \beta(V_{\mathrm{GS}} - V_{\mathrm{t}} - V_{\mathrm{DS}}/2)V_{\mathrm{DS}} in the linear region and ID=(β/2)(VGS−Vt)2I_{\mathrm{D}} = (\beta/2)(V_{\mathrm{GS}} - V_{\mathrm{t}})^2 in saturation, where β\beta depends on the device size and the process.

Expert. Derived from the gradual-channel approximation with constant mobility. β=μCoxW/L\beta = \mu C_{\mathrm{ox}} W/L. It is SPICE Level 1. It works for long channels but overestimates drive current and its growth with VDDV_{\mathrm{DD}} in short ones, because it ignores velocity saturation and mobility degradation.

SRAM (static RAM)

Static random-access memory. Each bit is a cell of (usually) six transistors: two inverters in a loop that holds the value, plus two access transistors. Fast and easy to build alongside logic, but large per bit. Used for caches, register files and on-chip buffers.

All three levels

Beginner. Fast memory built only from switches. Each bit sits in a tiny loop of switches that hold each other in place, as long as the power is on.

Novice. Static random-access memory. Each bit is a cell of (usually) six transistors: two inverters in a loop that holds the value, plus two access transistors. Fast and easy to build alongside logic, but large per bit. Used for caches, register files and on-chip buffers.

Expert. Bistable-latch memory, almost always the 6T cell in dense arrays (8T or 10T where reads must be decoupled or VDDV_{\mathrm{DD}} is low). Density is set by the bitcell and the periphery; stability is set by read SNM, write margin and VTV_{\mathrm{T}} mismatch, which together set the minimum operating voltage (VminV_{\mathrm{min}}).

Standard cell

A small, pre-designed circuit, such as a logic gate or a flip-flop, drawn once and reused thousands of times. All cells in a library have the same height so they line up in rows, and most come in several sizes: bigger ones are faster but take more area and power.

All three levels

Beginner. A small, ready-made building block that a chip is assembled from, such as a tiny yes/no circuit or a one-bit memory.

Novice. A small, pre-designed circuit, such as a logic gate or a flip-flop, drawn once and reused thousands of times. All cells in a library have the same height so they line up in rows, and most come in several sizes: bigger ones are faster but take more area and power.

Expert. A ready-to-place cell with a layout outline (LEF) for placement and routing and Liberty models for timing and power. Each function comes in several drive strengths (X1, X2, X4) and often several threshold-voltage flavors that trade speed for leakage. Synthesis picks function, size and flavor for every instance; a dont_use list keeps unwanted cells out.

Standard-cell power rail (followpin)

Thin supply and ground wires on the lowest wiring layer, running along the top and bottom edge of every row of logic cells. Cells are designed so their power connections touch these rails when they are placed side by side in a row.

All three levels

Beginner. A thin power wire that runs along each row of small parts and touches every one, like a local street.

Novice. Thin supply and ground wires on the lowest wiring layer, running along the top and bottom edge of every row of logic cells. Cells are designed so their power connections touch these rails when they are placed side by side in a row.

Expert. Called followpin rails because they follow the cell rows. They are the most resistive part of the grid, so stacks of vias from the straps above land on them at regular intervals. The rails take room inside every cell, which limits how short cells can be made; that is one motivation for buried and backside power.

Static noise margin (SNM)

The largest steady voltage disturbance an SRAM cell can tolerate without losing its value. It is read off the butterfly curve as the side of the largest square that fits inside it. Hold SNM is measured with the cell idle; read SNM with the cell being read, which is worse.

All three levels

Beginner. How much disturbance a memory cell can take before it accidentally flips its stored bit.

Novice. The largest steady voltage disturbance an SRAM cell can tolerate without losing its value. It is read off the butterfly curve as the side of the largest square that fits inside it. Hold SNM is measured with the cell idle; read SNM with the cell being read, which is worse.

Expert. Defined by Seevinck, List and Lohstroh (1987) as the side of the largest square nested between the two VTCs, which equals the largest DC series noise source the loop can absorb. Bounded by VDD/2V_{\mathrm{DD}}/2. Read SNM is the limiting case in 6T cells and shrinks with lower VDDV_{\mathrm{DD}}, lower cell ratio and VTV_{\mathrm{T}} mismatch.

Static timing analysis (STA)

A method that adds up the delays of every gate and wire on every path between storage cells and checks that each signal arrives in time for the clock tick. It uses arithmetic on known delays, not test inputs, which is why it is called static.

All three levels

Beginner. A way to check that every signal in the chip arrives on time. The chip never has to run.

Novice. A method that adds up the delays of every gate and wire on every path between storage cells and checks that each signal arrives in time for the clock tick. It uses arithmetic on known delays, not test inputs, which is why it is called static.

Expert. Computes, for every pin, the latest and earliest time a signal can arrive and the time it is required, then checks setup, hold and other rules on every path. It runs once per scenario, applies variation margins and removes double-counted margin on shared clock wiring (CPPR), adds crosstalk effects, and can re-time the worst paths individually (path-based analysis) to remove pessimism.

Statistical process control (SPC)

Plotting measurements such as film thickness or line width on control charts with a center line and limits, typically 3 standard deviations away. A point outside the limits signals that something has changed.

All three levels

Beginner. Keeping a running chart of each measurement and stopping to investigate when one looks unusual.

Novice. Plotting measurements such as film thickness or line width on control charts with a center line and limits, typically 3 standard deviations away. A point outside the limits signals that something has changed.

Expert. Fab data is nested (sites within wafers within lots), so limits must reflect the right variance component, or charts either cry wolf on known lot-to-lot shifts or miss within-wafer drift.

Statistical STA (SSTA)

Timing analysis that treats each delay as a spread of possible values with probabilities, rather than a single worst-case number, and reports how likely each path is to fail.

All three levels

Beginner. Timing checks that treat each delay as a range of likely values, not one number.

Novice. Timing analysis that treats each delay as a spread of possible values with probabilities, rather than a single worst-case number, and reports how likely each path is to fail.

Expert. Propagates delay distributions through the timing graph. Block-based SSTA writes every delay as a mean plus sensitivities to named sources of variation and approximates the max of two arrivals back into that form by matching moments. It yields sensitivities and criticality probabilities. POCV is the simpler statistical model most signoff flows run.

Steiner tree (RSMT)

A branching set of wires that connects all pins of a net and may add extra junctions, called Steiner points, where no pin exists. The rectilinear Steiner minimum tree (RSMT) is the shortest such tree built only from horizontal and vertical segments.

All three levels

Beginner. The shortest way to connect several points with wires when you may add new junction points along the way.

Novice. A branching set of wires that connects all pins of a net and may add extra junctions, called Steiner points, where no pin exists. The rectilinear Steiner minimum tree (RSMT) is the shortest such tree built only from horizontal and vertical segments.

Expert. Finding an RSMT is NP-hard, but Hanan’s theorem limits Steiner points to the grid formed by the pins’ xx and yy lines, and lookup-table methods such as FLUTE are exact and fast for nets with fewer than 10 pins. Global routers use them to break nets into two-pin connections.

Stochastic rounding

A rounding mode where a value between two representable numbers rounds up with probability equal to how close it is to the upper one. On average the rounding error is zero, so many small updates are not all lost.

All three levels

Beginner. Rounding up or down at random, more often toward the closer value, so tiny changes add up correctly on average.

Novice. A rounding mode where a value between two representable numbers rounds up with probability equal to how close it is to the upper one. On average the rounding error is zero, so many small updates are not all lost.

Expert. Round xx to ⌊x⌋\lfloor x \rfloor or ⌈x⌉\lceil x \rceil with P(up)=(x−⌊x⌋)/ulpP(\text{up}) = (x - \lfloor x \rfloor)/\mathrm{ulp}; unbiased (E[round(x)]=xE[\mathrm{round}(x)] = x). Needs a random source per conversion. Used for low-precision weight updates and gradient quantization (for example in FP4 training) where round-to-nearest would systematically drop small increments.

Streaming multiprocessor (SM) / compute unit (CU)

The GPU’s basic core. NVIDIA calls it a streaming multiprocessor, AMD a compute unit. It holds a register file, a small cache and shared memory, warp schedulers, and math and matrix units. A big GPU has around a hundred or more.

All three levels

Beginner. One of the many cores on a GPU. Each one has its own math units and its own small, fast memory.

Novice. The GPU’s basic core. NVIDIA calls it a streaming multiprocessor, AMD a compute unit. It holds a register file, a small cache and shared memory, warp schedulers, and math and matrix units. A big GPU has around a hundred or more.

Expert. A multithreaded core partitioned into (typically four) scheduler quadrants or SIMDs. All threads of a thread block run on one SM, and the SM’s registers, shared memory and warp slots set how many blocks fit.

Structured ASIC

A chip built on a prefabricated base of logic and memory shared by many customers. Each design customizes only one or a few top metal and via layers, so it needs only those masks: cheaper and faster than a full custom chip, more efficient than an FPGA.

All three levels

Beginner. A half-made chip. The bottom layers are the same for everyone. Only the top wires are made for your design.

Novice. A chip built on a prefabricated base of logic and memory shared by many customers. Each design customizes only one or a few top metal and via layers, so it needs only those masks: cheaper and faster than a full custom chip, more efficient than an FPGA.

Expert. Metal- or via-programmable base array. NRE falls with the number of custom layers, while delay and power improve with a few more (Ahmed et al.). Sits between FPGA and cell-based ASIC on unit cost, power and NRE; adoption has stayed narrow.

Structured sparsity

A sparsity pattern with rules about where zeros may go. Examples: 2:4 (at most two nonzeros in every group of four), blocks of zeros, or whole rows or channels removed. Unstructured sparsity puts zeros anywhere.

All three levels

Beginner. Zeros set out in a fixed pattern, like “two out of every four.” Then the chip knows where the gaps are.

Novice. A sparsity pattern with rules about where zeros may go. Examples: 2:4 (at most two nonzeros in every group of four), blocks of zeros, or whole rows or channels removed. Unstructured sparsity puts zeros anywhere.

Expert. Trades some accuracy at a given sparsity for regular memory access, cheap metadata and balanced work. N:M patterns (2:4 on recent GPUs) give a fixed 2× in hardware; block patterns let dense matrix engines skip whole tiles; channel pruning just makes a smaller dense layer.

Stuck-at fault

The classic fault model: one wire in the circuit is assumed to be permanently 0 (stuck-at-0) or permanently 1 (stuck-at-1), whatever drives it. Every gate input and output has both possibilities.

All three levels

Beginner. A pretend flaw where one wire acts as if it were stuck at 0 (always off) or at 1 (always on).

Novice. The classic fault model: one wire in the circuit is assumed to be permanently 0 (stuck-at-0) or permanently 1 (stuck-at-1), whatever drives it. Every gate input and output has both possibilities.

Expert. Single (one fault at a time), permanent, at gate pins. Simple and a good proxy for many shorts and opens, but blind to timing defects and to many defects inside cells. Lists are shrunk by equivalence collapsing before test generation.

Subnormal number

A floating-point value smaller than the smallest normal number. Its exponent field is all zeros and the hidden leading 1 becomes a 0, so it has fewer significant bits. Subnormals fill the gap between the smallest normal value and zero.

All three levels

Beginner. An extra-small number a format can still store by giving up some of its detail, instead of jumping straight to zero.

Novice. A floating-point value smaller than the smallest normal number. Its exponent field is all zeros and the hidden leading 1 becomes a 0, so it has fewer significant bits. Subnormals fill the gap between the smallest normal value and zero.

Expert. Encodings with E=0E = 0, valued (−1)S×21−bias×0.M(-1)^S \times 2^{1 - \text{bias}} \times 0.M. They give gradual underflow (x−y=0x - y = 0 only if x=yx = y) at the cost of precision; hardware sometimes flushes them to zero for speed. In FP16 they reach down to 2−242^{-24}.

Subthreshold conduction

Below the threshold voltage the channel isn’t fully formed, but some electrons still have enough energy to cross from source to drain. The current falls by a fixed factor for every step down in gate voltage instead of dropping to zero.

All three levels

Beginner. The small trickle that still flows when a transistor is “off,” like a tap that never quite stops dripping.

Novice. Below the threshold voltage the channel isn’t fully formed, but some electrons still have enough energy to cross from source to drain. The current falls by a fixed factor for every step down in gate voltage instead of dropping to zero.

Expert. Diffusion current over a gate-controlled barrier: ID∝exp⁡[(VGS−Vt)/(n kT/q)] (1−e−VDS/(kT/q))I_{\mathrm{D}} \propto \exp[(V_{\mathrm{GS}} - V_{\mathrm{t}})/(n\,kT/q)]\,(1 - e^{-V_{\mathrm{DS}}/(kT/q)}). It is the dominant leakage in modern logic and the reason VtV_{\mathrm{t}} cannot fall freely.

Subthreshold swing (slope)

The gate-voltage change needed to cut the off-state current by a factor of ten. Smaller is better: the switch turns off more completely for the same voltage.

All three levels

Beginner. How sharply a transistor shuts off as its gate voltage drops below the tipping point. Sharper means less leaking.

Novice. The gate-voltage change needed to cut the off-state current by a factor of ten. Smaller is better: the switch turns off more completely for the same voltage.

Expert. S=n (kT/q)ln⁡10S = n\,(kT/q)\ln 10, with a thermal limit of about 60 mV/decade at room temperature; real devices sit above it, and about 100 mV/decade is a common textbook planning value. It sets how many decades of leakage you buy per volt of VTV_{\mathrm{T}}.

Sum of products (SOP)

Writing a logic rule as several AND terms joined by OR, such as (a AND b) OR (NOT a AND c). Minimizing it means using as few and as short terms as possible.

All three levels

Beginner. Writing a logic rule as a list of “this AND that” cases joined by OR.

Novice. Writing a logic rule as several AND terms joined by OR, such as (a AND b) OR (NOT a AND c). Minimizing it means using as few and as short terms as possible.

Expert. The two-level form targeted by Quine–McCluskey and Espresso. It maps directly onto programmable logic arrays but blows up for functions like parity, which is why standard-cell synthesis works on multi-level networks and uses SOP minimization only locally.

Superscalar

A processor that fetches, issues and finishes several instructions per clock cycle, using several execution units side by side. A 4-wide core can start up to four instructions each cycle.

All three levels

Beginner. A processor that can start more than one step of a program in the same tick of its clock.

Novice. A processor that fetches, issues and finishes several instructions per clock cycle, using several execution units side by side. A 4-wide core can start up to four instructions each cycle.

Expert. Width applies separately to fetch, decode, rename, issue and commit; sustained IPC is far below peak width because of dependences, mispredictions and misses. Wider machines need quadratically more bypass and wakeup logic.

Surrogate model

A cheap model fitted to the results of slow experiments, used to guess the result of experiments not yet run.

All three levels

Beginner. A quick stand-in that guesses what a slow test would say, so you only run the slow test on the most promising options.

Novice. A cheap model fitted to the results of slow experiments, used to guess the result of experiments not yet run.

Expert. Gaussian processes, tree-based models and neural networks are common. A useful one reports how unsure it is along with its guess, which lets an optimizer balance trying unexplored settings against refining good ones.

Swamping

In floating-point addition the smaller number is shifted to line up with the larger one; if the difference in size exceeds the mantissa width, the small number is shifted out entirely and lost. Long sums in a narrow accumulator suffer from this.

All three levels

Beginner. When a small number added to a big total disappears, because the total can’t store enough digits to notice it.

Novice. In floating-point addition the smaller number is shifted to line up with the larger one; if the difference in size exceeds the mantissa width, the small number is shifted out entirely and lost. Long sums in a narrow accumulator suffer from this.

Expert. Loss of addends whose magnitude is below the accumulator’s ulp: roughly when the sum exceeds the addend by more than 2m+12^{m+1}. Worse for long reductions and non-zero-mean data. Mitigated by wider accumulators, chunked (pairwise) summation and stochastic rounding.

Switch

A chip with many ports that forwards traffic from any input port to any output port. Connecting every accelerator to switches gives every pair the same bandwidth and distance.

All three levels

Beginner. A hub chip that can connect any chip plugged into it to any other, so they don’t each need wires to everyone.

Novice. A chip with many ports that forwards traffic from any input port to any output port. Connecting every accelerator to switches gives every pair the same bandwidth and distance.

Expert. In scale-up, usually a single-stage, high-radix crossbar per plane, with each accelerator spreading its links across several parallel switch planes. Adds a hop of latency, power and cost, and caps one-level domains at the switch radix; can also do in-network reduction and multicast.

Switch box (switch block)

The set of programmable switches where a horizontal and a vertical routing channel meet. Each one lets a wire segment connect to segments on the other sides.

All three levels

Beginner. A spot where wires cross in an FPGA. Tiny switches there decide which wires join.

Novice. The set of programmable switches where a horizontal and a vertical routing channel meet. Each one lets a wire segment connect to segments on the other sides.

Expert. Characterized by its flexibility Fs (often 3) and pattern (disjoint, Wilton and others). With multi-length directional wires, switch and connection boxes merge into per-wire driver muxes.

Switch radix

The number of ports on a switch. A higher radix lets each switch reach more servers or more other switches, so the network needs fewer layers.

All three levels

Beginner. How many cables one switch can plug in at once.

Novice. The number of ports on a switch. A higher radix lets each switch reach more servers or more other switches, so the network needs fewer layers.

Expert. Port count kk at a given lane speed. Total capacity is fixed by the SerDes, so one chip can be run as fewer fast ports or more slow ones (for example 64 × 800G or 128 × 400G). Reach grows as k2k^2 for two tiers and k3k^3 for three.

Switching threshold (VM)

The input voltage at which the output equals the input on the transfer curve. Below it the gate treats the input as mostly a 0; above it, mostly a 1. Ideally it sits at half the supply voltage.

All three levels

Beginner. The input level where a gate is right on the edge between reading a 0 and a 1.

Novice. The input voltage at which the output equals the input on the transfer curve. Below it the gate treats the input as mostly a 0; above it, mostly a 1. Ideally it sits at half the supply voltage.

Expert. VMV_{\mathrm{M}}, where Vin=VoutV_{\mathrm{in}} = V_{\mathrm{out}} and both devices are saturated. Set by the PMOS/NMOS strength ratio and the thresholds; it moves only slowly (with the square root of the strength ratio in the square-law model) as you resize the devices.

Synchronizer (two-flop)

Two flip-flops in a row, both on the receiving clock. The first may become undecided; the second reads it one clock tick later, by which time it has almost certainly settled.

All three levels

Beginner. Two memory cells in a row that give an unsure signal extra time to settle before the rest of the chip uses it.

Novice. Two flip-flops in a row, both on the receiving clock. The first may become undecided; the second reads it one clock tick later, by which time it has almost certainly settled.

Expert. Safe only for single-bit signals that stay steady long enough to be sampled. The settling time is one receiving-clock period minus the flops’ own delays and the wire between them, so the two flops are placed side by side. Fast clocks, low voltage or temperature extremes may call for three stages.

Synthesizable subset

The parts of a hardware language that a tool can turn into real circuits. Time delays, printed messages and file reading are left out: they are useful when testing a design but mean nothing in hardware.

All three levels

Beginner. The part of a chip-design language that a tool can turn into real circuits.

Novice. The parts of a hardware language that a tool can turn into real circuits. Time delays, printed messages and file reading are left out: they are useful when testing a design but mean nothing in hardware.

Expert. Codified for Verilog in IEEE 1364.1 and assumed by tools such as Yosys. Code outside it can simulate correctly and then be ignored or rejected by synthesis, so the gates behave differently from the simulation: a simulation/synthesis mismatch.

systolic array

A grid of identical multiply-and-add cells. On every tick of the chip’s clock each cell passes its numbers to its neighbor, so a value read once from memory is used by a whole row or column of cells. It computes matrix products with few memory reads.

All three levels

Beginner. A grid of tiny calculators that pass numbers to their neighbors, step by step. Each number gets used many times instead of being fetched again.

Novice. A grid of identical multiply-and-add cells. On every tick of the chip’s clock each cell passes its numbers to its neighbor, so a value read once from memory is used by a whole row or column of cells. It computes matrix products with few memory reads.

Expert. A regular mesh of processing elements in which operands move one neighbor per cycle. Which operand stays put (weights, outputs or inputs) defines its dataflow. It trades flexibility for operand reuse and short wires; utilization drops on small or oddly shaped matrices, and the cycles to fill and drain the array matter for small batches.

Tail latency

A high percentile of response time, such as the 99th percentile: 99% of requests finish faster than this. Services set limits on it, not on the average.

All three levels

Beginner. How long the slowest few answers take. An app feels slow if even one answer in a hundred is slow.

Novice. A high percentile of response time, such as the 99th percentile: 99% of requests finish faster than this. Services set limits on it, not on the average.

Expert. Batching and deep queues raise throughput but stretch the tail, so a latency bound caps usable batch size. MLPerf Server reports throughput only at a load where the tail bound holds, estimated with an early-stopping statistical test.

Tap cell (well tap)

A small cell with no logic that connects the silicon under the transistors (the wells and the substrate) to the power and ground wires. Taps are laid down in a regular pattern before any logic, so no transistor is too far from one; this prevents latch-up, a fault in which a stray path through the silicon shorts power to ground.

All three levels

Beginner. A tiny helper part placed at regular spacing that connects the chip’s silicon base to power, so its switches behave reliably.

Novice. A small cell with no logic that connects the silicon under the transistors (the wells and the substrate) to the power and ground wires. Taps are laid down in a regular pattern before any logic, so no transistor is too far from one; this prevents latch-up, a fault in which a stray path through the silicon shorts power to ground.

Expert. LEF CLASS CORE WELLTAP. The foundry sets a design rule on the distance from any active area to the nearest tap, so taps go in a checkerboard at a fixed pitch (OpenROAD tapcell -distance) and are FIXED before placement. They cut rows into segments that legalization must treat separately.

Tapeout

The release of the finished, fully checked layout file to the foundry, which uses it to make the masks for manufacturing. The name comes from the reels of magnetic tape that once carried the file.

All three levels

Beginner. The moment the finished chip design file is sent to the factory to be made.

Novice. The release of the finished, fully checked layout file to the foundry, which uses it to make the masks for manufacturing. The name comes from the reels of magnetic tape that once carried the file.

Expert. A freeze as much as a technical step: one final layout file with a recorded checksum, the exact versions of the rule files, libraries and IP it was checked with, and a signed checklist. Any change after it means new masks.

TARA

Threat Analysis and Risk Assessment, the method in the car-cybersecurity standard ISO/SAE 21434: list what is worth protecting, how an attacker could reach it, and how bad the result would be.

All three levels

Beginner. A careful list of what a hacker might attack and how bad it would be.

Novice. Threat Analysis and Risk Assessment, the method in the car-cybersecurity standard ISO/SAE 21434: list what is worth protecting, how an attacker could reach it, and how bad the result would be.

Expert. Feeds chip requirements such as secure boot, key storage, debug lock-down and cryptographic acceleration, and must be maintained across the vehicle’s life.

Target impedance

The most the supply may sag per amp of sudden extra demand: the allowed ripple voltage divided by the largest current step, Z=ΔV/ΔIZ = \Delta V / \Delta I. Impedance is resistance generalized to currents that change over time.

All three levels

Beginner. A rule for how steady the power supply must stay when the chip suddenly needs more electricity.

Novice. The most the supply may sag per amp of sudden extra demand: the allowed ripple voltage divided by the largest current step, Z=ΔV/ΔIZ = \Delta V / \Delta I. Impedance is resistance generalized to currents that change over time.

Expert. Applied at every frequency: the regulator covers slow changes, board and package capacitors the middle range, and on-chip capacitance the fastest. The resonance between on-chip capacitance and package inductance makes an impedance peak that must also stay under the target.

Tcl

Tool Command Language, a small scripting language that programs can build in. Most chip design tools, open and commercial, take their commands in Tcl, so an engineer can save a sequence of commands as a script and rerun it exactly. Timing goals (SDC files) are written in Tcl syntax too.

All three levels

Beginner. A simple language for typing commands. Most chip design programs understand it, so engineers can save their steps instead of clicking.

Novice. Tool Command Language, a small scripting language that programs can build in. Most chip design tools, open and commercial, take their commands in Tcl, so an engineer can save a sequence of commands as a script and rerun it exactly. Timing goals (SDC files) are written in Tcl syntax too.

Expert. Written by John Ousterhout, who also wrote the Magic layout editor. Its small interpreter is easy to embed in a C or C++ program, which made it the usual command language of EDA tools. Newer tools add Python on the same internal data, but most flow scripts and every SDC file are still Tcl.

TCO (total cost of ownership)

Purchase price (spread over the years it is used) plus running costs such as electricity, cooling, space, networking and staff. Divide by the work done to get cost per unit of work.

All three levels

Beginner. Everything a computer costs over its life: buying it, plus the electricity and upkeep to run it.

Novice. Purchase price (spread over the years it is used) plus running costs such as electricity, cooling, space, networking and staff. Divide by the work done to get cost per unit of work.

Expert. Usually modeled per hour: capex ÷ (depreciation life in hours) + IT power × PUE × electricity price + other opex, then divided by delivered work per hour at the real duty cycle. Real prices are negotiated and rarely public.

TDP (thermal design power)

The steady power that the chip’s case and cooling are designed to carry away as heat while keeping the transistors below their maximum safe temperature. It is the chip’s power allowance for long-running work.

All three levels

Beginner. The amount of power, and so heat, a chip is designed to give off steadily without overheating.

Novice. The steady power that the chip’s case and cooling are designed to carry away as heat while keeping the transistors below their maximum safe temperature. It is the chip’s power allowance for long-running work.

Expert. A thermal number, not a peak electrical one. Short bursts can exceed it, because the chip and its cooling take time to heat up; the largest instantaneous current is a separate requirement for the power-delivery design. It follows from the package’s thermal resistance (degrees of temperature rise per watt), the ambient temperature and the maximum junction temperature.

Technology mapping

Rebuilding the optimized logic out of cells from the target library, keeping the function the same while making the result small, fast or both.

All three levels

Beginner. Choosing real parts from the catalog to build each piece of the simplified circuit.

Novice. Rebuilding the optimized logic out of cells from the target library, keeping the function the same while making the result small, fast or both.

Expert. A covering problem: tile the optimized logic graph (today usually an AIG) with library cells. Modern mappers list the small sub-networks (cuts) ending at each node, look up which cells compute the same function, pick the fastest cover by dynamic programming, then swap in smaller cells away from the critical path.

Tensor

A multi-dimensional array of numbers: a vector is 1-D, a matrix 2-D. Weights, activations and gradients in a neural network are all tensors, and a number format is chosen per tensor.

All three levels

Beginner. A big block of numbers arranged in rows and columns (or more dimensions), the way a neural network stores its data.

Novice. A multi-dimensional array of numbers: a vector is 1-D, a matrix 2-D. Weights, activations and gradients in a neural network are all tensors, and a number format is chosen per tensor.

Expert. An nn-dimensional array with a shape and element type. Quantization granularity is described along tensor axes: per-tensor, per-channel (one axis), per-block (kk contiguous elements, usually along the reduction dimension of the matrix multiplication).

Tensor core (matrix core)

A hardware unit that performs a small matrix multiply-and-add (for example 4×4 or larger tiles) as a single instruction. NVIDIA calls them tensor cores, AMD matrix cores. They now provide most of a GPU’s AI math.

All three levels

Beginner. A special unit in a GPU core that multiplies whole small blocks of numbers in one go, instead of one pair at a time.

Novice. A hardware unit that performs a small matrix multiply-and-add (for example 4×4 or larger tiles) as a single instruction. NVIDIA calls them tensor cores, AMD matrix cores. They now provide most of a GPU’s AI math.

Expert. A small systolic-style MMA datapath fed from registers (and, from Hopper on, directly from shared memory). One instruction does tens to hundreds of FMAs, amortizing fetch, decode and register-read energy. Supports low-precision inputs with wider accumulation.

Tensor parallelism

Dividing the weight matrices of each model layer among several accelerators. Each computes part of every layer, then they combine partial results, typically with all-reduce operations, several times per layer.

All three levels

Beginner. Splitting each step of an AI model’s math across several chips. They must swap part-answers many times per step.

Novice. Dividing the weight matrices of each model layer among several accelerators. Each computes part of every layer, then they combine partial results, typically with all-reduce operations, several times per layer.

Expert. Intra-layer model parallelism (Megatron-style column/row splits). Communication is on the critical path, per layer, with message sizes proportional to batch × sequence × hidden size, so it is normally confined to the scale-up domain.

Test (fault) coverage

The percentage of possible manufacturing flaws, modeled as simple faults such as a wire stuck at 0 or stuck at 1, that the factory’s test patterns would detect. Higher coverage means fewer bad chips slip through.

All three levels

Beginner. The share of possible manufacturing flaws that the factory tests can catch.

Novice. The percentage of possible manufacturing flaws, modeled as simple faults such as a wire stuck at 0 or stuck at 1, that the factory’s test patterns would detect. Higher coverage means fewer bad chips slip through.

Expert. Set at the spec stage from the outgoing-quality target, because the fraction of bad chips that escape depends on both yield and coverage. Targets for stuck-at faults (a node stuck at 0 or 1) and transition faults (a node too slow to switch) drive the scan test circuitry, extra test points, built-in memory self-test and tester time, so they belong in the requirements.

Test compression

On-chip hardware that expands a few streams of test data from the tester into many short internal scan chains, and squeezes the many chain outputs back into a few. It works because only a small share of the bits in each pattern actually matter.

All three levels

Beginner. Squeezing test patterns so the test machine sends far less data and the test runs faster.

Novice. On-chip hardware that expands a few streams of test data from the tester into many short internal scan chains, and squeezes the many chain outputs back into a few. It works because only a small share of the bits in each pattern actually matter.

Expert. A linear decompressor (an LFSR-like circuit fed continuously from the tester) is driven with data found by solving linear equations so that it reproduces the bits that matter. The output compactor, usually XOR-based, needs masking or X-tolerance because unknown values spoil the compacted result.

Test escape

A defective chip that passes manufacturing test and ships. Escapes come back as customer returns. The opposite mistake, failing a chip that was actually good, is called yield loss.

All three levels

Beginner. A broken chip that passes the factory test by mistake and gets shipped to a customer.

Novice. A defective chip that passes manufacturing test and ships. Escapes come back as customer returns. The opposite mistake, failing a chip that was actually good, is called yield loss.

Expert. Caused by defects no fault model describes, gaps in coverage, timing defects a slow test can’t see, and failures that appear only at certain voltages or temperatures. Escapes set the shipped defect level; overkill (failing good parts) costs yield.

Test pattern

One set of input values applied to the chip during test, plus the output values a fault-free chip would produce. A chip passes if every pattern’s output matches.

All three levels

Beginner. A list of 0s and 1s fed into a chip during testing, along with the answer a good chip should give back.

Novice. One set of input values applied to the chip during test, plus the output values a fault-free chip would produce. A chip passes if every pattern’s output matches.

Expert. In scan test, one load of the scan chains plus primary-input values, one or more capture clocks, and the expected unload, with unknown bits masked. Pattern count times chain length sets most of the test time.

Testbench

The test program wrapped around a design inside a simulator. It creates inputs, applies them to the design, watches the outputs and checks them against the right answers. It is never built into the chip; it exists only for testing.

All three levels

Beginner. A test setup built around a chip design inside a computer. It feeds the design inputs and checks what comes out.

Novice. The test program wrapped around a design inside a simulator. It creates inputs, applies them to the design, watches the outputs and checks them against the right answers. It is never built into the chip; it exists only for testing.

Expert. Usually layered. Sequences create transactions (whole operations, such as one write to memory), drivers turn them into signal changes, monitors turn observed signals back into transactions, and scoreboards and coverage collectors check and count. Keeping the checks independent of how inputs were chosen lets one environment serve every test and be reused from block to chip level.

Thermal-expansion mismatch

Silicon and circuit boards expand at different rates as temperature changes. On a small chip the difference is tiny; across a whole wafer it adds up to a large strain on the connections.

All three levels

Beginner. Different materials grow by different amounts when they warm up, which can bend or crack things glued together.

Novice. Silicon and circuit boards expand at different rates as temperature changes. On a small chip the difference is tiny; across a whole wafer it adds up to a large strain on the connections.

Expert. Shear strain grows with distance from the neutral point × Δ(CTE)×ΔT\Delta(\mathrm{CTE}) \times \Delta T, so it scales with part size. Large parts need compliant connectors, matched materials or careful assembly.

Thread block (workgroup)

A group of up to 1,024 threads that always runs on one GPU core, can share data through that core’s shared memory, and can synchronize at barriers. AMD calls it a workgroup.

All three levels

Beginner. A bigger group of GPU workers that share a small notepad and can wait for each other.

Novice. A group of up to 1,024 threads that always runs on one GPU core, can share data through that core’s shared memory, and can synchronize at barriers. AMD calls it a workgroup.

Expert. The unit of placement on an SM. Blocks are independent, which is what lets the same kernel scale across GPUs with different SM counts. Hopper adds clusters of blocks that can address each other’s shared memory.

Threat model

A document that names what must be protected (secret keys, software, user data), who might attack (remotely, with the chip in hand, or from inside the supply chain) and where they could get in. Security requirements are written from it.

All three levels

Beginner. A list of what an attacker might want from the chip, who they might be, and how they could try to get it.

Novice. A document that names what must be protected (secret keys, software, user data), who might attack (remotely, with the chip in hand, or from inside the supply chain) and where they could get in. Security requirements are written from it.

Expert. Covers logical attacks; physical attacks such as side-channel analysis (reading secrets from power use or timing) and fault injection (glitching voltage or clock to make the chip misbehave); and test and debug ports as a way in. It must exist before architecture, because countermeasures change the block diagram and the test design.

Threshold voltage (Vt)

The control voltage at which a transistor switch starts to conduct strongly. A lower threshold makes switching faster but lets more current leak when the switch is off; chip libraries offer several threshold flavors to choose between.

All three levels

Beginner. How hard the electricity must push before one of the chip’s tiny switches turns on.

Novice. The control voltage at which a transistor switch starts to conduct strongly. A lower threshold makes switching faster but lets more current leak when the switch is off; chip libraries offer several threshold flavors to choose between.

Expert. Sets the floor for the supply voltage: speed collapses as the supply approaches it. Below threshold, current falls by roughly a factor of ten for every 60–100 mV (the subthreshold slope), so every 100 mV taken off the threshold costs roughly ten times more off-state leakage.

Threshold-voltage mismatch

Random variation in the switching threshold of transistors drawn the same size, mostly from the random number and placement of dopant atoms. Smaller transistors vary more, which hurts circuits like SRAM cells that rely on two halves matching.

All three levels

Beginner. Tiny random differences between switches that were meant to be identical. At this size, a few atoms out of place make a difference.

Novice. Random variation in the switching threshold of transistors drawn the same size, mostly from the random number and placement of dopant atoms. Smaller transistors vary more, which hurts circuits like SRAM cells that rely on two halves matching.

Expert. Local (within-die) variation whose standard deviation scales roughly as 1/WL1/\sqrt{WL} (Pelgrom). It is why SRAM bitcells, built from the smallest devices on the die and replicated millions of times, must be designed for 5–6σ5\text{–}6\sigma tails and why their VminV_{\mathrm{min}} stays high.

Through-silicon via (TSV)

A metal-filled hole that runs straight through a thinned silicon chip, so chips stacked on top of each other, or a chip and the slab of silicon under it, can be wired together vertically.

All three levels

Beginner. A tiny metal-filled hole that goes straight through a silicon chip, so chips stacked on top of each other can connect.

Novice. A metal-filled hole that runs straight through a thinned silicon chip, so chips stacked on top of each other, or a chip and the slab of silicon under it, can be wired together vertically.

Expert. Large compared with logic gates and surrounded by keep-out zones where stress would change nearby transistors, so their number and position are early layout decisions. Each TSV process adds steps that can fail, which raises the value of testing each die before it is stacked.

TIA (transimpedance amplifier)

A transimpedance amplifier converts the small current a photodiode produces when light hits it into a voltage the rest of the receiver can work with.

All three levels

Beginner. A tiny amplifier that boosts the weak electrical signal from a light sensor so other chips can read it.

Novice. A transimpedance amplifier converts the small current a photodiode produces when light hits it into a voltage the rest of the receiver can work with.

Expert. Sets receiver noise and bandwidth. For linear-drive optics it must be linear (no limiting) so the host SerDes can equalize the whole optical path; together with the laser driver it forms the module’s analog front end.

tick-to-trade

The time from receiving a market-data message (a “tick”) to sending an order in response, measured on the network cable.

All three levels

Beginner. How long a trading machine takes to answer a price change with an order.

Novice. The time from receiving a market-data message (a “tick”) to sending an order in response, measured on the network cable.

Expert. Benchmarks fix the measurement points precisely. STAC-T0 measures from the last bit of inbound data needed for the decision to the first bit of the outbound order, and covers network I/O only, with essentially no trading logic.

Tiling (folding)

Splitting a matrix multiplication that is bigger than the array into blocks (tiles) that fit, running each, and combining the results. Edge tiles may only partly fill the array.

All three levels

Beginner. Cutting a big job into pieces the size of the grid and running them one after another.

Novice. Splitting a matrix multiplication that is bigger than the array into blocks (tiles) that fit, running each, and combining the results. Edge tiles may only partly fill the array.

Expert. A GEMM mapped onto an R×CR \times C array needs ⌈SR/R⌉×⌈SC/C⌉\lceil S_R/R \rceil \times \lceil S_C/C \rceil folds. Tile order sets how often each operand is re-read from the buffer; tile size relative to the matrix sets the fragmentation loss.

Time to train

The wall-clock time from first touching the training data until the model reaches a set quality target on held-out data. It rewards both raw speed and choices that help the model learn in fewer steps.

All three levels

Beginner. How long a machine takes to teach an AI model until it is good enough at its task. MLPerf Training scores machines this way.

Novice. The wall-clock time from first touching the training data until the model reaches a set quality target on held-out data. It rewards both raw speed and choices that help the model learn in fewer steps.

Expert. Chosen over throughput because tricks that raise samples/s (larger batches, lower precision) can raise the number of steps needed. Scored over several runs with the fastest and slowest dropped; convergence faster than the reference’s statistical bounds is rejected.

Timing arc

A pair of pins in a cell, such as input A to output Y, with tables for how long the output takes to respond and how sharp its edge is. A flip-flop also has arcs for its timing checks, such as setup and hold between D and CLK.

All three levels

Beginner. One cause-and-effect path through a building block: when this input changes, that output changes after a certain time.

Novice. A pair of pins in a cell, such as input A to output Y, with tables for how long the output takes to respond and how sharp its edge is. A flip-flop also has arcs for its timing checks, such as setup and hold between D and CLK.

Expert. A Liberty timing() group with related_pin, timing_sense (positive_unate, negative_unate, non_unate) and timing_type (combinational, rising_edge, setup_rising, hold_rising…). STA builds its graph from these arcs plus nets, so a missing or wrong arc is a silent timing hole.

Timing budget

When a signal must pass through two blocks within one tick of the clock, the share of that time each block may use. Each block’s team gets its share as a target, so it can finish its block without waiting for the other.

All three levels

Beginner. How much of its tiny time limit a signal may spend in each piece of a chip, so each team knows its target.

Novice. When a signal must pass through two blocks within one tick of the clock, the share of that time each block may use. Each block’s team gets its share as a target, so it can finish its block without waiting for the other.

Expert. Written into each block’s SDC constraints as input and output delays on its pins, derived from the top-level floorplan’s estimates of wire and buffer delay. Too generous a share on one side starves the other, so budgets are revisited as the blocks mature.

Timing closure

The loop of checking timing, changing the design, its constraints or the tool settings, and rebuilding, until every path meets the clock (slack of zero or more). On an FPGA it has to be done without changing the chip itself.

All three levels

Beginner. The work of tweaking a design until every signal gets where it is going before the next clock tick.

Novice. The loop of checking timing, changing the design, its constraints or the tool settings, and rebuilding, until every path meets the clock (slack of zero or more). On an FPGA it has to be done without changing the chip itself.

Expert. Iterating logic and implementation choices (pipelining, retiming, floorplanning, tool directives, physical optimization) until worst negative slack is zero or positive. On an FPGA the delay model is fixed by the device and its speed grade.

Timing path

The route a signal takes from one flip-flop (or an input) through a chain of gates to the next flip-flop (or an output). It must finish within one clock period, minus a little margin.

All three levels

Beginner. One route a signal takes from one small memory in the chip to the next.

Novice. The route a signal takes from one flip-flop (or an input) through a chain of gates to the next flip-flop (or an output). It must finish within one clock period, minus a little margin.

Expert. A start point (a flip-flop clock pin or an input port), the cells and nets it passes through, and an end point (a flip-flop data pin or an output port). Timing tools compute arrival and required times for every such path and report the worst per end point.

Timing-driven placement

Placement that checks, during the run, which signal paths come closest to missing the clock deadline, and tells the placer to keep the wires on those paths extra short.

All three levels

Beginner. Putting the parts on the slowest signal paths extra close together, so their signals arrive in time.

Novice. Placement that checks, during the run, which signal paths come closest to missing the clock deadline, and tells the placer to keep the wires on those paths extra short.

Expert. Usually net reweighting: at a few checkpoints the placer runs timing analysis on estimated wire parasitics and multiplies the wirelength weight of low-slack nets. Path-based and Lagrangian formulations also exist. Over-weighting a few nets can create density hot spots and lengthen everything else.

TLB (translation lookaside buffer)

A small cache of recent virtual-to-physical page translations. A hit translates in about a cycle; a miss triggers a page-table walk, one memory access per level.

All three levels

Beginner. A tiny, fast list of recent address translations, so the chip doesn’t have to look them up in memory every time.

Novice. A small cache of recent virtual-to-physical page translations. A hit translates in about a cycle; a miss triggers a page-table walk, one memory access per level.

Expert. Typically 32–128 highly associative entries at L1 plus a larger shared L2 TLB (Skylake: 64 and 1,536 for 4 KB pages). Reach = entries × page size; huge pages extend it. On x86 a hardware walker refills it; MIPS and Alpha used software.

Token

The unit of text a language model processes: a word, part of a word or a punctuation mark, mapped to a number. How text is split into tokens depends on the model’s tokenizer.

All three levels

Beginner. A small chunk of text, often a word or part of a word. Language models read and write text one token at a time.

Novice. The unit of text a language model processes: a word, part of a word or a punctuation mark, mapped to a number. How text is split into tokens depends on the model’s tokenizer.

Expert. Throughput and cost are quoted per token. Prefill processes all prompt tokens in parallel; decode emits one token per sequence per forward pass.

Topology

The shape of a network: which chips (or switches) are connected by links. Rings, meshes, tori and switched stars are common patterns.

All three levels

Beginner. The pattern of who is wired to whom.

Novice. The shape of a network: which chips (or switches) are connected by links. Rings, meshes, tori and switched stars are common patterns.

Expert. A graph whose properties (degree, diameter, bisection bandwidth, path diversity) bound latency and throughput for each traffic pattern, and whose regularity decides which collective algorithms map onto it without contention.

TOPS/W

Tera-operations per second per watt, the same as trillions of operations per joule. One multiply-accumulate usually counts as two operations. 10 TOPS/W means 100 femtojoules per operation.

All three levels

Beginner. How many trillion math steps a chip does for each joule of energy. A joule is about what it takes to lift an apple one meter. Higher is better.

Novice. Tera-operations per second per watt, the same as trillions of operations per joule. One multiply-accumulate usually counts as two operations. 10 TOPS/W means 100 femtojoules per operation.

Expert. Meaningless without the precision, the sparsity assumed, what is included (array only, macro, chip, system) and whether it is peak or sustained on a real workload. Some papers normalize to 1-bit operations, which inflates numbers by the product of the bit widths.

Torus

A mesh (grid) of chips in two or three dimensions, with extra wraparound links joining opposite edges. Each chip connects only to its 4 (2D) or 6 (3D) nearest neighbors.

All three levels

Beginner. A grid of chips where each one talks to its neighbors, and the edges wrap around so the last chip in a row links back to the first.

Novice. A mesh (grid) of chips in two or three dimensions, with extra wraparound links joining opposite edges. Each chip connects only to its 4 (2D) or 6 (3D) nearest neighbors.

Expert. A kk-ary nn-cube: constant degree 2n2n, diameter about nk/2nk/2, bisection bandwidth proportional to kn−1k^{n-1}. Cheap to cable with short links, well suited to dimension-ordered collectives, but non-neighbor traffic takes multiple hops and shares links.

Total ionizing dose (TID)

The total amount of radiation a chip absorbs over its life, measured in krad (thousands of rad, a unit of absorbed energy). It slowly changes how the transistors behave until the chip drifts out of spec or stops working.

All three levels

Beginner. All the radiation a chip soaks up over its life. It slowly wears the chip out.

Novice. The total amount of radiation a chip absorbs over its life, measured in krad (thousands of rad, a unit of absorbed energy). It slowly changes how the transistors behave until the chip drifts out of spec or stops working.

Expert. Charge trapped in the transistors’ gate oxide shifts their threshold voltage and raises leakage current. A spec states it as a minimum tolerance, for example 300 krad(Si), where “(Si)” means energy absorbed per kilogram of silicon. Enclosed-layout transistors, whose gate surrounds one terminal, raise the tolerance further.

Trace buffer

A small memory inside the chip that keeps recording a chosen set of internal signals and freezes when a trigger condition occurs, so engineers can read out what happened just before a failure through the debug port.

All three levels

Beginner. A small memory inside the chip that records recent signals, like a flight recorder, so engineers can see what happened before a failure.

Novice. A small memory inside the chip that keeps recording a chosen set of internal signals and freezes when a trigger condition occurs, so engineers can read out what happened just before a failure through the debug port.

Expert. Its size sets how many signals and how many clock cycles can be seen, so teams choose signals carefully, compress the data and share buffers between blocks. It works best for logic bugs that can be made to happen again.

Track assignment

The step between global and detailed routing that puts the long straight pieces of each planned route onto specific tracks, so the detailed router starts from a sensible arrangement.

All three levels

Beginner. Giving each planned wire its own lane within a stretch of the chip before the final drawing.

Novice. The step between global and detailed routing that puts the long straight pieces of each planned route onto specific tracks, so the detailed router starts from a sensible arrangement.

Expert. Usually greedy or graph-based, one strip of GCells (a panel) at a time. It is the first point where each wire’s neighbors are decided, so crosstalk-aware track assignment can keep sensitive nets apart or put shield wires beside them.

Track height

A standard cell’s height expressed as the number of horizontal routing tracks (wire lanes on the lowest metal layer) that fit in it. SKY130’s high-density library is 9 tracks tall; ASAP7’s main library is 7.5.

All three levels

Beginner. How tall a building block is, counted in the number of wire lanes that fit across it.

Novice. A standard cell’s height expressed as the number of horizontal routing tracks (wire lanes on the lowest metal layer) that fit in it. SKY130’s high-density library is 9 tracks tall; ASAP7’s main library is 7.5.

Expert. Sets how many device widths, internal wires and pin access points fit in a cell. Fewer tracks give smaller, lower-capacitance cells but weaker drive and harder pin access; libraries often come in several heights for density, speed or leakage.

Training

Running examples forward through the model, measuring the error, then running backward to compute how to change every weight, and updating them. It is done once per model, on large clusters, for weeks.

All three levels

Beginner. Teaching a neural network by showing it examples and nudging its numbers after each mistake.

Novice. Running examples forward through the model, measuring the error, then running backward to compute how to change every weight, and updating them. It is done once per model, on large clusters, for weeks.

Expert. Forward plus backward is about 3× the forward FLOPs (6N6N per token). It needs memory for weights, gradients, optimizer states and saved activations; large token batches keep the GEMMs compute-bound.

Training step

One iteration of training: a forward pass over a batch of examples, a backward pass that computes gradients, and an optimizer update of every weight. Large models take hundreds of thousands of steps.

All three levels

Beginner. One round of learning: the model tries a batch of examples, checks how wrong it was, and nudges all its numbers a little.

Novice. One iteration of training: a forward pass over a batch of examples, a backward pass that computes gradients, and an optimizer update of every weight. Large models take hundreds of thousands of steps.

Expert. Forward, backward and optimizer phases over a global batch, possibly split into microbatches whose gradients are accumulated. Synchronous training ends every step with all replicas holding identical weights.

Transaction-level model (TLM)

A software model of a chip, usually in SystemC, where blocks exchange whole operations (a read, a write, a packet) instead of individual wire changes on every clock tick. It runs fast enough to boot real software before the chip exists.

All three levels

Beginner. A fast, simplified model of a chip that counts messages between blocks instead of individual wires.

Novice. A software model of a chip, usually in SystemC, where blocks exchange whole operations (a read, a write, a packet) instead of individual wire changes on every clock tick. It runs fast enough to boot real software before the chip exists.

Expert. SystemC’s TLM-2.0 interfaces support two styles: loosely timed, for speed, and approximately timed, for more accurate performance numbers. Accuracy is traded for speed, so performance conclusions from a TLM need checking against the real design or silicon.

Transconductance (gm)

The slope of the drain current versus gate voltage: gm=ΔID/ΔVGSg_{\mathrm{m}} = \Delta I_{\mathrm{D}}/\Delta V_{\mathrm{GS}}, measured in siemens. It tells you how much extra current one more millivolt on the gate buys.

All three levels

Beginner. How much the current changes when the gate voltage changes a little. Higher means the transistor reacts more strongly.

Novice. The slope of the drain current versus gate voltage: gm=ΔID/ΔVGSg_{\mathrm{m}} = \Delta I_{\mathrm{D}}/\Delta V_{\mathrm{GS}}, measured in siemens. It tells you how much extra current one more millivolt on the gate buys.

Expert. gm=∂ID/∂VGSg_{\mathrm{m}} = \partial I_{\mathrm{D}}/\partial V_{\mathrm{GS}}. In the square law gm=βVGT=2ID/VGTg_{\mathrm{m}} = \beta V_{\mathrm{GT}} = 2I_{\mathrm{D}}/V_{\mathrm{GT}}; when fully velocity-saturated it tends to WCoxvsatW C_{\mathrm{ox}} v_{\mathrm{sat}}, independent of VGSV_{\mathrm{GS}}. The ratio gm/IDg_{\mathrm{m}}/I_{\mathrm{D}} peaks in subthreshold at 1/(n kT/q)1/(n\,kT/q).

Transformer

A neural network built from repeated blocks, each with an attention layer (which mixes information between positions in the input) and a feed-forward layer (two matrix multiplications applied to each position).

All three levels

Beginner. The design behind today’s chatbots. It reads a whole passage at once and works out which words matter to which.

Novice. A neural network built from repeated blocks, each with an attention layer (which mixes information between positions in the input) and a feed-forward layer (two matrix multiplications applied to each position).

Expert. Stacks of multi-head attention and position-wise MLP blocks with residual connections and normalization. Almost all of its FLOPs are GEMMs: the Q/K/V/output projections, the MLP, and attention’s QKTQK^T and PVPV products.

Transformer (distribution)

An electrical machine that changes AC voltage with almost no loss, for example from tens of thousands of volts to 480 volts. Datacenter transformers are typically more than 99% efficient.

All three levels

Beginner. A big device that lowers the strong push of electricity from power lines to a level the building can use.

Novice. An electrical machine that changes AC voltage with almost no loss, for example from tens of thousands of volts to 480 volts. Datacenter transformers are typically more than 99% efficient.

Expert. Losses split into no-load (core) loss, paid even at zero load, and load (copper) loss that grows with current squared, so efficiency peaks at partial load. U.S. minimum efficiencies in 10 CFR 431.196 are set at 50% load for liquid-immersed and medium-voltage dry-type units and at 35% for low-voltage dry-type units. Lead times for large units have become a constraint on new sites.

Transistor

A semiconductor device in which a small signal on one terminal controls the current between two others. Digital chips use it as an on/off switch; analog circuits use it as an amplifier.

All three levels

Beginner. A tiny electric switch with no moving parts. Electricity on its control connection turns it on or off.

Novice. A semiconductor device in which a small signal on one terminal controls the current between two others. Digital chips use it as an on/off switch; analog circuits use it as an amplifier.

Expert. In digital CMOS, almost always a MOSFET used as a voltage-controlled switch with finite on-resistance, a gate capacitance to charge, and nonzero off-current. Bipolar transistors survive in analog, RF and some I/O circuits.

Transistor stack

Transistors connected in series, source to drain, as in the pull-down network of a NAND. The current has to pass through every one, so a stack conducts less well than a single transistor of the same size.

All three levels

Beginner. Several switches in a row that must all be on for electricity to get through.

Novice. Transistors connected in series, source to drain, as in the pull-down network of a NAND. The current has to pass through every one, so a stack conducts less well than a single transistor of the same size.

Expert. Series devices: resistance adds, internal nodes add diffusion capacitance, and upper devices suffer body effect because their sources sit above the rail. Stacks are upsized by their depth for equal drive, and they cut off-state leakage roughly tenfold for a two-high stack.

Transition fault

A fault model for timing defects: a wire that is too slow to change from 0 to 1 (slow-to-rise) or from 1 to 0 (slow-to-fall). A test first sets the wire to its old value, then flips it and checks the result at the chip’s full clock speed.

All three levels

Beginner. A pretend flaw where a signal still gets to the right value, but too slowly.

Novice. A fault model for timing defects: a wire that is too slow to change from 0 to 1 (slow-to-rise) or from 1 to 0 (slow-to-fall). A test first sets the wire to its old value, then flips it and checks the result at the chip’s full clock speed.

Expert. Models a large delay at one spot. Each test is a stuck-at-style test preceded by an initializing pattern, so it needs two patterns and an at-speed capture. It misses small extra delays on short paths, which is what timing-aware and path delay tests are for.

Transition time (slew)

The time a signal takes to swing between two set fractions of the supply voltage, for example from 20% to 80%. The input transition is one of the two numbers used to look up a cell’s delay; the output transition becomes the next cell’s input transition.

All three levels

Beginner. How quickly a signal flips from 0 to 1, or from 1 to 0. A sharp edge flips fast; a slow edge flips gradually.

Novice. The time a signal takes to swing between two set fractions of the supply voltage, for example from 20% to 80%. The input transition is one of the two numbers used to look up a cell’s delay; the output transition becomes the next cell’s input transition.

Expert. Measured between the library’s slew thresholds (slew_lower/upper_threshold_pct) and scaled by slew_derate_from_library. In SKY130 hd these are 20% and 80% with a derate of 1, so a 10 ps index is a 16.7 ps full-swing ramp. STA propagates slew along every path; max_transition caps it.

Trial placement (prototype)

A quick, rough placement of the whole design into the floorplan, followed by an estimate of wiring crowding and a rough speed check, to see whether the plan can work. It is thrown away afterwards.

All three levels

Beginner. A quick, rough practice run of placing the parts, to check a floorplan before doing the real thing.

Novice. A quick, rough placement of the whole design into the floorplan, followed by an estimate of wiring crowding and a rough speed check, to see whether the plan can work. It is thrown away afterwards.

Expert. Often run on clusters of cells rather than individual cells, which shrinks the problem. Its congestion map and worst timing paths point back at macro positions, channel widths and pin locations. Its numbers are trends, not final results.

Triple modular redundancy (TMR)

Building three copies of a storage element or circuit and passing their outputs through a voting circuit that outputs the majority answer, so one wrong copy is outvoted.

All three levels

Beginner. Keeping three copies and going with the majority.

Novice. Building three copies of a storage element or circuit and passing their outputs through a voting circuit that outputs the majority answer, so one wrong copy is outvoted.

Expert. Flip-flop-level TMR masks a single upset but not a glitch that all three copies capture together; full TMR also triples the voters and logic. In FPGAs, tripling everything except clocks, resets and enables (distributed TMR) reduced errors more than tripling only flip-flops in a NASA beam test.

Trusted Foundry

A U.S. Defense Department program, managed by the Defense Microelectronics Activity (DMEA), that gives government users access to approved companies for designing, making, packaging and testing sensitive chips.

All three levels

Beginner. A government program that lets the military get chips made by factories that have been checked out.

Novice. A U.S. Defense Department program, managed by the Defense Microelectronics Activity (DMEA), that gives government users access to approved companies for designing, making, packaging and testing sensitive chips.

Expert. Run by DMEA’s Trusted Access Program Office. It guarantees access to accredited flows (shared multi-project wafer runs, dedicated prototype runs and production) for the low volumes government programs need, and protects integrity and confidentiality through chain-of-custody controls.

Trusted supplier (DMEA-accredited)

A company approved by DMEA to handle one step of making a sensitive chip: design, brokering, making the photomasks, manufacturing the silicon, packaging or testing.

All three levels

Beginner. A company the military has inspected and approved to handle one step in making secret chips.

Novice. A company approved by DMEA to handle one step of making a sensitive chip: design, brokering, making the photomasks, manufacturing the silicon, packaging or testing.

Expert. Accreditation covers the people and processes for a specific service (design, aggregation, brokering, mask making, foundry, post-processing, packaging/assembly, test). DoD Instruction 5200.44 requires custom or tailored ICs for a military end use in applicable systems to be procured from accredited suppliers.

Truth table

A table with one row per input combination (2n2^n rows for nn inputs) and the output value in the last column. It fully defines what a logic gate does.

All three levels

Beginner. A list of every possible combination of inputs and what the output is for each one.

Novice. A table with one row per input combination (2n2^n rows for nn inputs) and the output value in the last column. It fully defines what a logic gate does.

Expert. The complete specification of a combinational function. For a CMOS gate, each row also tells you which network conducts, which is useful when reasoning about leakage states and worst-case delay.

TUE (total-power usage effectiveness)

PUE multiplied by ITUE: total facility energy divided by the energy used by the computing parts. Because fans and power losses count no matter whether they sit in the building or the server, it can fairly compare designs that move them across that line.

All three levels

Beginner. A score for the whole site: how much electricity the building uses for each unit that reaches the chips.

Novice. PUE multiplied by ITUE: total facility energy divided by the energy used by the computing parts. Because fans and power losses count no matter whether they sit in the building or the server, it can fairly compare designs that move them across that line.

Expert. TUE=ITUE×PUE=total energy÷compute energy\mathrm{TUE} = \mathrm{ITUE} \times \mathrm{PUE} = \text{total energy} \div \text{compute energy}. It removes PUE’s incentive to shift cooling and conversion into the IT box, but inherits PUE’s blind spots (no measure of useful work) and is sensitive to temperature set points.

UCIe

Universal Chiplet Interconnect Express, an open standard for the links between chiplets in one package. It defines the electrical signals, the bump layout and the protocols, so chiplets from different designers can work together.

All three levels

Beginner. A shared set of rules that lets small chips from different makers talk to each other inside one package.

Novice. Universal Chiplet Interconnect Express, an open standard for the links between chiplets in one package. It defines the electrical signals, the bump layout and the protocols, so chiplets from different designers can work together.

Expert. Defines a standard-package PHY (100–130 µm bump pitch, up to 25 mm reach) and an advanced-package PHY (25–55 µm, up to 2 mm) with spare lanes for repair, plus mappings for PCIe and CXL traffic.

Ultra Ethernet

The Ultra Ethernet Consortium’s open specification (version 1.0 in June 2025) for a new transport and related link and switch features for AI and HPC networks over standard Ethernet and IP.

All three levels

Beginner. New rules for Ethernet, written by many companies together, to make it work better for AI.

Novice. The Ultra Ethernet Consortium’s open specification (version 1.0 in June 2025) for a new transport and related link and switch features for AI and HPC networks over standard Ethernet and IP.

Expert. Defines UET (semantic, packet-delivery, congestion-management and security sublayers, programmed via libfabric) with packet spraying, unordered reliable delivery, NSCC/RCCC congestion control, optional packet trimming, link-layer retry and credit-based flow control.

Uncore

The shared parts of a processor chip outside the cores: caches (small, fast on-chip memories), the on-chip connections, memory controllers and input/output. Adding cores does not add more of it.

All three levels

Beginner. The parts of a processor chip that are not the processor cores themselves, such as memory and connections to the outside.

Novice. The shared parts of a processor chip outside the cores: caches (small, fast on-chip memories), the on-chip connections, memory controllers and input/output. Adding cores does not add more of it.

Expert. Often treated as a fixed cost in early power and area models. The interface circuits (PHYs) are sized by the standard they implement, so they shrink little with each new manufacturing generation and can dominate a small chip.

Unknown value (X)

A value that simulation can’t predict, marked X. It comes from things like memories that were never written or paths too slow to settle. The tester must ignore X bits, or it would fail good chips.

All three levels

Beginner. A spot in the chip whose value during a test can’t be predicted, so the tester has to ignore it.

Novice. A value that simulation can’t predict, marked X. It comes from things like memories that were never written or paths too slow to settle. The tester must ignore X bits, or it would fail good chips.

Expert. Sources include uninitialized memories, non-scan flops, analog and black-box boundaries, and false or multicycle paths at speed. X’s cost observability, and in a MISR or XOR compactor one X can hide many good bits, so they are blocked at the source or masked.

Untestable (redundant) fault

A fault that no input pattern can both trigger and make visible at an output. Redundant logic, wires tied to a constant and blocked paths cause them.

All three levels

Beginner. A pretend flaw that no test can ever catch, because it doesn’t change what the chip does.

Novice. A fault that no input pattern can both trigger and make visible at an output. Redundant logic, wires tied to a constant and blocked paths cause them.

Expert. Reports separate faults proven undetectable (tied, blocked, redundant) from ATPG-untestable ones (not testable under the current test constraints) and aborted ones (the search gave up). Only the proven ones are fair to exclude from coverage.

UPF

The Unified Power Format (IEEE standard 1801): a file, kept separate from the design’s code, that lists which parts of the chip have their own power supply, which can be switched off, and what protects their neighbors while they are off.

All three levels

Beginner. A written plan that tells the design tools which parts of the chip can be switched off, and how.

Novice. The Unified Power Format (IEEE standard 1801): a file, kept separate from the design’s code, that lists which parts of the chip have their own power supply, which can be switched off, and what protects their neighbors while they are off.

Expert. Tcl-based power intent (domains, supply nets, switches, isolation, level shifting, retention) kept separate from RTL and refined through the flow. Simulation, synthesis, place-and-route and equivalence checking all read it, so mismatches between UPF and implementation are a classic bug source.

UPS (uninterruptible power supply)

Equipment between the grid and the IT load that rides through power disturbances using batteries or a flywheel. The common “double-conversion” type turns AC into DC and back to AC all the time, which costs a few percent of the power passing through it.

All three levels

Beginner. A big set of batteries that keeps the computers running if the power flickers or fails, until backup generators start.

Novice. Equipment between the grid and the IT load that rides through power disturbances using batteries or a flywheel. The common “double-conversion” type turns AC into DC and back to AC all the time, which costs a few percent of the power passing through it.

Expert. Double-conversion UPS efficiency rose from 85–90% in the 1990s to 95% or more; “eco” modes that bypass the inverter reach about 99% at some cost in conditioning. Efficiency falls at low load factor, which redundant (N+1, 2N) designs create. Some AI racks move short-term backup into battery backup units in the rack.

Useful skew

Delaying the clock to one flip-flop on purpose so the slow calculation feeding it gets more time. The time comes from somewhere: the next calculation, which starts at that flip-flop, gets less, and that flip-flop’s hold check gets tighter.

All three levels

Beginner. Making the clock tick reach one memory cell a bit late on purpose, to give a slow calculation more time.

Novice. Delaying the clock to one flip-flop on purpose so the slow calculation feeding it gets more time. The time comes from somewhere: the next calculation, which starts at that flip-flop, gets less, and that flip-flop’s hold check gets tighter.

Expert. Fishburn formalized it in 1990 as a linear program over every flip-flop’s clock arrival time. Production tools apply it step by step after CTS as concurrent clock and data (CCD) optimization, limited by hold checks and by loops of paths, around which the skews cancel.

Utilization (of compute)

Delivered throughput divided by peak throughput, as a percentage. It depends on the workload, the software and the system as much as on the chip.

All three levels

Beginner. How much of its top speed a chip actually uses on a real job. A chip at 30% utilization is busy only about a third of the time it could be.

Novice. Delivered throughput divided by peak throughput, as a percentage. It depends on the workload, the software and the system as much as on the chip.

Expert. Report what was counted: useful model FLOPs (MFU) or all FLOPs the hardware executed (HFU), over what interval, and against which peak (dense or sparse, which precision). Not the same as placement density, which is also called utilization.

Utilization (placement density)

The share of the space for logic cells that the cells actually cover. Overall (core) utilization is an average; target density is the most any small area may be filled. Leaving a good part free gives room for wires and for the cells that timing repair adds later.

All three levels

Beginner. How full the chip’s floor is. 70% means parts cover 70% of the space and 30% is left open.

Novice. The share of the space for logic cells that the cells actually cover. Overall (core) utilization is an average; target density is the most any small area may be filled. Leaving a good part free gives room for wires and for the cells that timing repair adds later.

Expert. Global utilization is an average; local bin density and pin density are what cause trouble. A block can be 60% utilized overall and still have bins at the cap next to macros. Placers enforce a per-bin target density (OpenROAD gpl defaults to 0.7), and routability modes inflate cells where congestion is predicted.

UVM (Universal Verification Methodology)

A standard library and set of conventions for building testbenches out of reusable parts, written in the SystemVerilog language and standardized as IEEE 1800.2.

All three levels

Beginner. A popular set of rules and building blocks for chip tests, so parts can be reused from one chip to the next.

Novice. A standard library and set of conventions for building testbenches out of reusable parts, written in the SystemVerilog language and standardized as IEEE 1800.2.

Expert. Supplies standard component types (agents containing drivers, sequencers and monitors), fixed run phases, a factory that lets a test swap a component for a variant without editing the rest, a configuration database and a register model. Reuse is the payoff; boilerplate and a steep learning curve are the cost.

Valence band

The band of allowed energies below the band gap, filled by the electrons that form the covalent bonds. A full band carries no current; an empty spot in it (a hole) can. Its top edge is written EvE_{\mathrm{v}}.

All three levels

Beginner. The lower energy level in silicon. It is full of electrons that hold the atoms together, so none of them can move.

Novice. The band of allowed energies below the band gap, filled by the electrons that form the covalent bonds. A full band carries no current; an empty spot in it (a hole) can. Its top edge is written EvE_{\mathrm{v}}.

Expert. Hole density p=Nv e−(EF−Ev)/kTp = N_{\mathrm{v}}\, e^{-(E_{\mathrm{F}} - E_{\mathrm{v}})/kT}, with Nv≈1.04×1019 cm−3N_{\mathrm{v}} \approx 1.04 \times 10^{19}\,\mathrm{cm^{-3}} in Si. Holes have a larger effective mass and about a third of the electron mobility in silicon.

VDD and VSS

VDD is the positive supply voltage, below 1 V in many modern chips; VSS is ground, 0 V. Every logic cell connects to both, and the current it uses flows in from VDD and out to VSS.

All three levels

Beginner. The two power connections every part of a chip needs, like the plus and minus ends of a battery.

Novice. VDD is the positive supply voltage, below 1 V in many modern chips; VSS is ground, 0 V. Every logic cell connects to both, and the current it uses flows in from VDD and out to VSS.

Expert. Both are networks with resistance. A cell sees its local VDD minus its local VSS, so a sag on the supply side and a rise on the ground side (ground bounce) come out of the same voltage budget, and both nets need analysis.

Vector (SIMD) extension

A set of instructions added to a processor’s instruction set that work on short vectors of numbers held in wide registers: SSE, AVX and AVX-512 on x86, NEON and SVE on Arm, and the V extension on RISC-V.

All three levels

Beginner. Extra instructions that let a core do the same math on a whole row of numbers in one step.

Novice. A set of instructions added to a processor’s instruction set that work on short vectors of numbers held in wide registers: SSE, AVX and AVX-512 on x86, NEON and SVE on Arm, and the V extension on RISC-V.

Expert. Fixed-width families (SSE 128, AVX 256, AVX-512 512 bits; NEON 128) bake the register width into the code; scalable ones (SVE, RVV) leave it to the implementation and expose it at run time. Masks or predicates handle conditionals and loop tails.

Vector-length agnostic

A style of vector instruction set (Arm SVE, RISC-V V) where the program doesn’t fix the vector width. It asks the hardware how many elements fit and loops accordingly, so one compiled program runs on narrow and wide hardware alike.

All three levels

Beginner. A way of writing vector programs so the same program works on chips whose row-of-numbers units are different widths.

Novice. A style of vector instruction set (Arm SVE, RISC-V V) where the program doesn’t fix the vector width. It asks the hardware how many elements fit and loops accordingly, so one compiled program runs on narrow and wide hardware alike.

Expert. VL is a run-time value: SVE derives loop control from predicates (whilelt) and increments by the element count (incd); RVV’s vsetvli returns vl from the requested length. No remainder loop, but data layouts and register spills can’t assume a compile-time size.

Vectorless analysis

A way of estimating the worst supply dip without simulating the chip running real work: instead of a recording of every signal, it assumes how often each part of the chip switches.

All three levels

Beginner. A way to find the worst sag without running real programs on the chip, by guessing how busy each part could get.

Novice. A way of estimating the worst supply dip without simulating the chip running real work: instead of a recording of every signal, it assumes how often each part of the chip switches.

Expert. Dynamic IR analysis whose currents come from assumed switching activity or from limits on each block’s current, instead of from a simulation waveform file (VCD). Fast and available early, but only as good as its assumptions, and bounding methods tend to be pessimistic. Paired with waveform-based runs on the stress patterns that matter.

Velocity saturation

At high electric fields, carriers scatter off the silicon lattice and their speed stops rising, leveling off near 107 cm/s10^7\,\mathrm{cm/s} for electrons. In a short channel the field from VDSV_{\mathrm{DS}} is high enough for this to cap the current before pinch-off does.

All three levels

Beginner. A speed limit for electrons in silicon. Past a certain push they bump into the crystal so often that they can’t go any faster.

Novice. At high electric fields, carriers scatter off the silicon lattice and their speed stops rising, leveling off near 107 cm/s10^7\,\mathrm{cm/s} for electrons. In a short channel the field from VDSV_{\mathrm{DS}} is high enough for this to cap the current before pinch-off does.

Expert. Carrier velocity rolls off from μE\mu E toward vsatv_{\mathrm{sat}} once E=VDS/LE = V_{\mathrm{DS}}/L approaches the critical field. The fully saturated limit is ID=WCox(VGS−Vt)vsatI_{\mathrm{D}} = W C_{\mathrm{ox}} (V_{\mathrm{GS}} - V_{\mathrm{t}}) v_{\mathrm{sat}}: linear in overdrive and independent of LL. Real devices are partly velocity-saturated, which the α\alpha-power law fits.

Verification plan (testplan)

A document that lists what will be checked, how (simulating the design in software, mathematical proof, running it on special hardware, analysis, or testing real chips) and when each check counts as done. Each entry points back to a requirement or feature.

All three levels

Beginner. The list of checks the team will run to prove the chip does what the spec says.

Novice. A document that lists what will be checked, how (simulating the design in software, mathematical proof, running it on special hardware, analysis, or testing real chips) and when each check counts as done. Each entry points back to a requirement or feature.

Expert. Usually a machine-readable list of testpoints, each mapped to features or requirement IDs and to the tests and coverage measurements that close it. The results of the automated test runs then roll up into a per-requirement status report.

Via

A small vertical metal plug through the insulator that joins a wire on one layer to the layer directly above or below. It has a cut (the plug itself) and a metal landing pad on each side. Each via adds a little electrical resistance and takes room that nearby wires could have used.

All three levels

Beginner. A tiny vertical plug that connects a wire on one layer to a wire on the layer above or below, like an elevator between floors.

Novice. A small vertical metal plug through the insulator that joins a wire on one layer to the layer directly above or below. It has a cut (the plug itself) and a metal landing pad on each side. Each via adds a little electrical resistance and takes room that nearby wires could have used.

Expert. Defined in LEF and placed by name in DEF. Rules cover the spacing between cuts, how far the landing pad must extend past the cut (often different in each direction) and a minimum number of cuts on wide wires. Via resistance adds delay, which motivates double vias, via pillars and routers that count vias in their cost.

Via array (power via stack)

A cluster of vias, the tiny vertical connectors between wiring layers, placed where two power wires on different layers cross. Stacks of them carry current down from the thick upper wires to the thin rails.

All three levels

Beginner. A cluster of tiny upright links that joins power wires on different layers.

Novice. A cluster of vias, the tiny vertical connectors between wiring layers, placed where two power wires on different layers cross. Stacks of them carry current down from the thick upper wires to the thin rails.

Expert. Sized for resistance and electromigration, since vias are a common failure point. Large arrays on the layers in between can grow into metal patches that block signal routes, so tools limit their rows and columns or keep the pass-through layers at minimum width.

Virtual memory

Each program uses virtual addresses that the hardware translates, page by page (often 4 KiB), into physical addresses using page tables the operating system keeps in memory. It isolates programs and lets memory be moved or shared.

All three levels

Beginner. A trick that gives every program its own made-up set of addresses. The chip quietly turns them into real memory addresses.

Novice. Each program uses virtual addresses that the hardware translates, page by page (often 4 KiB), into physical addresses using page tables the operating system keeps in memory. It isolates programs and lets memory be moved or shared.

Expert. Translation by multi-level radix page tables (e.g., four 9-bit levels plus a 12-bit offset in x86-64 and RISC-V Sv48), cached in TLBs and page-walk caches. Larger pages (2 MiB, 1 GiB) cut walk depth and raise TLB reach.

Volatile memory

Memory that needs power to keep its data. SRAM and DRAM are volatile; flash is non-volatile.

All three levels

Beginner. Memory that forgets everything when the power goes off.

Novice. Memory that needs power to keep its data. SRAM and DRAM are volatile; flash is non-volatile.

Expert. Volatility is a property of the storage mechanism: a latch (SRAM) or a leaking capacitor (DRAM) versus charge isolated by thick insulators (flash). Non-volatility costs write speed, write energy and endurance.

Voltage area (power-domain area)

A fenced region of the chip for one power domain: a group of logic that runs at its own supply voltage or can be switched off on its own. The domain’s cells must sit inside the fence, and its power wiring connects to that domain’s supply.

All three levels

Beginner. A fenced-off part of a chip that can run on less power, or be switched off, apart from the rest.

Novice. A fenced region of the chip for one power domain: a group of logic that runs at its own supply voltage or can be switched off on its own. The domain’s cells must sit inside the fence, and its power wiring connects to that domain’s supply.

Expert. Implemented as an exclusive region (OpenROAD set_domain_area; a FENCE region in DEF). Its boundary is where level shifters and isolation cells go, and a domain that switches off also needs room for its power-switch cells.

Voltage transfer curve (VTC)

A plot of a gate’s steady-state output voltage against its input voltage. For an inverter it is high for low inputs, falls steeply in the middle and is low for high inputs.

All three levels

Beginner. A graph showing what a gate outputs for every input, from fully low to fully high.

Novice. A plot of a gate’s steady-state output voltage against its input voltage. For an inverter it is high for low inputs, falls steeply in the middle and is low for high inputs.

Expert. The DC transfer characteristic Vout(Vin)V_{\mathrm{out}}(V_{\mathrm{in}}), found where pull-up and pull-down currents are equal. Its steep middle (high gain) is what gives a gate noise margin and lets logic restore degraded levels. Five regions follow from which device is off, linear or saturated.

VPR (Versatile Place and Route)

The academic FPGA packing, placement and routing tool from the University of Toronto, part of the Verilog-to-Routing (VTR) project. It can target any FPGA described in a file, which makes it the standard tool for FPGA research.

All three levels

Beginner. A free research program that places and wires designs on pretend FPGAs, so scientists can test ideas for new chips.

Novice. The academic FPGA packing, placement and routing tool from the University of Toronto, part of the Verilog-to-Routing (VTR) project. It can target any FPGA described in a file, which makes it the standard tool for FPGA research.

Expert. VTR’s pack, place, route and timing engine (since 1997): a routing-resource graph built from an architecture description, annealing placement with an adaptive schedule, PathFinder-based timing-driven routing. F4PGA also uses it for real devices.

VRM (voltage regulator module)

A DC-to-DC converter that steps a supply such as 12 V down to the chip’s core voltage, under 1 V, and holds it steady while the chip’s current jumps around. Big chips need hundreds of amps from it.

All three levels

Beginner. A circuit right next to a chip. It turns the board’s voltage, its push of electricity, down to the weak push the chip runs on.

Novice. A DC-to-DC converter that steps a supply such as 12 V down to the chip’s core voltage, under 1 V, and holds it steady while the chip’s current jumps around. Big chips need hundreds of amps from it.

Expert. Usually a multiphase buck converter with a digital controller, placed as close as possible to the load (point of load). Efficiency, transient response (droop under load steps), phase count and the board area it takes are the main trade-offs.

Wafer

A slice of single-crystal silicon, today usually 300 mm across and under a millimeter thick, cut from a large crystal ingot and polished. Each copy of a chip on it is a die.

All three levels

Beginner. A thin, mirror-polished disc of very pure silicon. Hundreds of chips are built side by side on one wafer and cut apart at the end.

Novice. A slice of single-crystal silicon, today usually 300 mm across and under a millimeter thick, cut from a large crystal ingot and polished. Each copy of a chip on it is a die.

Expert. Czochralski-grown, (100)-oriented for CMOS, lightly doped, often with an epitaxial top layer. Film thickness and flatness are worst near the edge, so an edge exclusion ring is written off and edge dies yield worse.

Wafer map

Test results (pass, fail, or which grade) plotted at each chip’s position on the wafer. Patterns such as a ring of failures at the edge, a cluster, or a line along a scratch point to a cause in manufacturing.

All three levels

Beginner. A picture of a wafer with each chip colored by whether it passed the test.

Novice. Test results (pass, fail, or which grade) plotted at each chip’s position on the wafer. Patterns such as a ring of failures at the edge, a cluster, or a line along a scratch point to a cause in manufacturing.

Expert. Stacking maps from many wafers separates random loss from loss tied to a position. Grouping neighboring dies into pairs, triples and so on (the windowing technique) turns one product’s maps into an estimate of defect density and of the loss that doesn’t depend on die size.

Wafer sort (probe test)

The first production test. A probe card with fine needles touches each chip’s pads while the chips are still on the wafer, before it is cut up. Failing chips are marked and never packaged.

All three levels

Beginner. Testing each chip with tiny needles while it is still on the round slice of silicon it was made on, before the slice is cut up.

Novice. The first production test. A probe card with fine needles touches each chip’s pads while the chips are still on the wafer, before it is cut up. Failing chips are marked and never packaged.

Expert. Limited by probe contact, power delivery and speed, so it may run a reduced or slower test set. Its main job is to keep bad dies out of packages; for multi-die packages it is the screen that defines a known-good die.

Wafer-scale integration

Making one working part from all, or nearly all, of a wafer. Because no wafer is flawless, the design has to include spare pieces and route around the broken ones.

All three levels

Beginner. Building one single, giant chip out of a whole silicon wafer instead of cutting the wafer into many small chips.

Novice. Making one working part from all, or nearly all, of a wafer. Because no wafer is flawless, the design has to include spare pieces and route around the broken ones.

Expert. A monolithic part spanning many reticle fields, joined by wiring across the scribe lines, with redundancy at core and link level so every wafer ships. A variant bonds known-good chiplets onto an interconnect wafer instead.

Waiver

A written, approved explanation for why a reported violation can safely stay, for example a rule flagged inside a block the factory has already proven. Waivers are reviewed because an unexamined one can hide a real bug.

All three levels

Beginner. Written permission to leave a reported problem in place, because it is known to be harmless.

Novice. A written, approved explanation for why a reported violation can safely stay, for example a rule flagged inside a block the factory has already proven. Waivers are reviewed because an unexamined one can hide a real bug.

Expert. A recorded exception to a signoff check with its reason, scope, owner and approver, matched by location or pattern so it can’t silently cover new violations. Carrying waivers across deck versions or projects without review is a known source of escapes.

Warp (wavefront)

The group of threads a GPU schedules as one unit: 32 threads on NVIDIA GPUs, 64 on AMD’s data-center GPUs (where it’s called a wavefront). One instruction is issued for the whole group.

All three levels

Beginner. A team of 32 (or 64) workers on a GPU who always do the same step together.

Novice. The group of threads a GPU schedules as one unit: 32 threads on NVIDIA GPUs, 64 on AMD’s data-center GPUs (where it’s called a wavefront). One instruction is issued for the whole group.

Expert. The unit of scheduling and of SIMD execution. Its context (PC, active mask, registers) stays on chip for its lifetime, so switching warps costs nothing. Threads are packed into warps in thread-ID order.

Warp scheduler

Hardware in each GPU core that, every clock cycle, picks one warp whose next instruction is ready and sends that instruction to the math or memory units.

All three levels

Beginner. The part of a GPU core that decides, every tick of the clock, which team of workers gets to go next.

Novice. Hardware in each GPU core that, every clock cycle, picks one warp whose next instruction is ready and sends that instruction to the math or memory units.

Expert. One per SM quadrant on recent NVIDIA parts, one issue per cycle. Policies such as loose round-robin or greedy-then-oldest trade fairness, cache locality and how well warps stagger their long-latency stalls.

Weight streaming

An execution style in which weights live in external memory and are streamed onto the chip one layer at a time, so the model can be larger than the chip’s own memory.

All three levels

Beginner. Keeping a model’s learned numbers in a separate memory box and sending them into the chip a piece at a time.

Novice. An execution style in which weights live in external memory and are streamed onto the chip one layer at a time, so the model can be larger than the chip’s own memory.

Expert. Decouples model size from on-chip capacity: on-chip SRAM holds activations and the layer in flight, while external memory and a reduction network handle weight storage, gradients and updates. It trades on-chip weight bandwidth for external bandwidth.

Weight-stationary (WS)

A dataflow in which each processing element holds one weight for a long time while activations stream past it and partial sums move between PEs. Each weight is read from memory once and reused for every input row.

All three levels

Beginner. A way of running a grid of calculators where each cell keeps one weight and the inputs flow past it.

Novice. A dataflow in which each processing element holds one weight for a long time while activations stream past it and partial sums move between PEs. Each weight is read from memory once and reused for every input row.

Expert. Maximizes weight reuse in the PE; partial sums travel through the array to accumulators outside it. Loading a new weight tile costs cycles without compute unless it is double-buffered. Used in TPU matrix units.

Weights (parameters)

The learned values inside a neural network’s layers. A model’s size is usually given as its parameter count, such as 7 billion; at 2 bytes each that is 14 GB to store.

All three levels

Beginner. The numbers a neural network learned during training. They are its knowledge, and there can be billions of them.

Novice. The learned values inside a neural network’s layers. A model’s size is usually given as its parameter count, such as 7 billion; at 2 bytes each that is 14 GB to store.

Expert. Fixed during inference and updated every step during training. Whether they can be reused across many inputs per fetch (batching) decides much of a layer’s arithmetic intensity.

Well

A doped region a few micrometers deep: NMOS transistors sit in a p-type well, PMOS in an n-type well. Both kinds on one wafer is what makes CMOS possible.

All three levels

Beginner. A zone of the silicon given a particular mix of added atoms, so one kind of switch can be built in it.

Novice. A doped region a few micrometers deep: NMOS transistors sit in a p-type well, PMOS in an n-type well. Both kinds on one wafer is what makes CMOS possible.

Expert. Formed by masked implants, today high-energy retrograde implants whose concentration peaks below the surface, which lowers well resistance and avoids a long high-temperature drive-in.

Wire bond

The traditional way to connect a chip to its package: fine wires of gold, copper, silver or aluminum, each welded from a pad on the chip to a pad on the package. It is the cheapest and most flexible option and is used for most packages.

All three levels

Beginner. Connecting a chip to its package with very thin wires, each stitched from a pad on the chip to a pad on the package.

Novice. The traditional way to connect a chip to its package: fine wires of gold, copper, silver or aluminum, each welded from a pad on the chip to a pad on the package. It is the cheapest and most flexible option and is used for most packages.

Expert. Pads must sit around the die edge, which caps the connection count, and each wire’s inductance limits signal speed and how much supply current it can deliver cleanly. Flip-chip takes over when pad count, speed or current exceed what the edge can carry.

Wire-load model

A table that guesses a wire’s length and electrical load from how many pins it connects to. Tools used it to estimate wire delay before any layout existed.

All three levels

Beginner. A rough guess at how long the wires will be, made before anything is actually laid out.

Novice. A table that guesses a wire’s length and electrical load from how many pins it connects to. Tools used it to estimate wire delay before any layout existed.

Expert. A statistical lookup, stored in older Liberty files, from fanout (the number of pins a net drives) to estimated capacitance and resistance. It predicts the average net well but individual nets badly, so the worst paths are mispredicted once wires dominate delay. Physical-aware synthesis replaces it with distances from a quick placement.

Word line

A horizontal wire running across one row of a memory array. Raising it turns on the access transistors of every cell in that row, connecting them to their bit lines.

All three levels

Beginner. The wire that ‘opens the door’ to every memory cell in one row at once.

Novice. A horizontal wire running across one row of a memory array. Raising it turns on the access transistors of every cell in that row, connecting them to their bit lines.

Expert. Driven by the row decoder through a large buffer; its RC delay and the number of cells it loads set part of the access time. Its voltage can be lowered on reads (underdrive) to protect stability or raised on writes (boost) to help writability.

Work function

The energy needed to take an electron from a material’s Fermi level out into free space (the vacuum). The difference between the gate’s and the silicon’s work functions shifts the threshold voltage.

All three levels

Beginner. How tightly a material holds on to its electrons. Different gate materials hold them differently, which changes when a transistor turns on.

Novice. The energy needed to take an electron from a material’s Fermi level out into free space (the vacuum). The difference between the gate’s and the silicon’s work functions shifts the threshold voltage.

Expert. ψg−ψs\psi_{\mathrm{g}} - \psi_{\mathrm{s}} is the main part of VfbV_{\mathrm{fb}}. Poly gates were doped n+\mathrm{n^+} or p+\mathrm{p^+} to set it; metal gates use different metals (or work-function layers) for NMOS and PMOS, and in FinFETs a few work-function options provide the multiple VTV_{\mathrm{T}} flavors.

Write margin

How much room is left when writing an SRAM cell: how far below the switching point the bit line manages to pull the stored 1. If the cell’s pull-up is too strong, the node never gets low enough and the write fails.

All three levels

Beginner. How easily a memory cell takes a new value when you write to it.

Novice. How much room is left when writing an SRAM cell: how far below the switching point the bit line manages to pull the stored 1. If the cell’s pull-up is too strong, the node never gets low enough and the write fails.

Expert. Measured several ways: the write noise margin (smallest square between the write-mode VTCs), the highest bit-line voltage that still flips the cell, or the word-line voltage margin. It degrades at fast-PMOS/slow-NMOS corners and low VDDV_{\mathrm{DD}}; write assists such as a boosted word line restore it.

Write-back (versus write-through)

Write-back caches update only the cache on a store and mark the line dirty; the line is written to the next level when evicted. Write-through caches send every store on immediately.

All three levels

Beginner. A cache that keeps changes to itself and only saves them to main memory when the block is thrown out.

Novice. Write-back caches update only the cache on a store and mark the line dirty; the line is written to the next level when evicted. Write-through caches send every store on immediately.

Expert. Write-back cuts traffic for repeated stores and pairs naturally with invalidation coherence (the M state). Write-through simplifies coherence and error recovery but needs a write buffer and much more bandwidth. Most CPU caches are write-back with write-allocate.

WUE (water usage effectiveness)

Annual on-site water use in liters divided by IT energy in kilowatt-hours. Evaporative cooling saves electricity but uses water, so low PUE and low WUE can pull in opposite directions.

All three levels

Beginner. A score for how much water a datacenter uses for each unit of electricity its computers use.

Novice. Annual on-site water use in liters divided by IT energy in kilowatt-hours. Evaporative cooling saves electricity but uses water, so low PUE and low WUE can pull in opposite directions.

Expert. Site WUE counts water used at the facility; source WUE adds water used to generate the electricity. The 2024 LBNL report estimated a U.S. average around 0.36 L/kWh (site).

X-propagation

Simulators use a value X for “unknown,” for example a storage bit that was never reset. X-propagation is how that unknown spreads to other signals. Standard simulation sometimes quietly turns an unknown into a definite choice, which can hide a real bug.

All three levels

Beginner. How an “unknown” value inside a simulation spreads to other parts of the design, sometimes hiding a bug.

Novice. Simulators use a value X for “unknown,” for example a storage bit that was never reset. X-propagation is how that unknown spreads to other signals. Standard simulation sometimes quietly turns an unknown into a definite choice, which can hide a real bug.

Expert. RTL semantics can be X-optimistic: an if-statement treats an X condition as false and takes one branch, where real hardware might go either way. X-propagation modes make RTL simulation pessimistic so these show up; two-state simulators such as Verilator replace X with constants or random values.

Yield

The fraction of manufactured chips that come out free of defects that stop them working. Bigger chips have lower yield, because each one is more likely to contain a defect; yield rises as a factory process matures.

All three levels

Beginner. The share of chips from the factory that work. If 90 out of 100 work, the yield is 90%.

Novice. The fraction of manufactured chips that come out free of defects that stop them working. Bigger chips have lower yield, because each one is more likely to contain a defect; yield rises as a factory process matures.

Expert. Sets cost per good die and, together with fault coverage, the defect level of shipped parts. New processes and designs start low and climb through yield learning, which draws on inline inspection, test structures and volume diagnosis of failing chips.

Yield learning

The work that raises the share of good chips on a new process or design: collect test and inspection data, work out where failing chips are broken, look for causes they share, and change the process or the design.

All three levels

Beginner. Finding and fixing the reasons chips fail in the factory, so the share of good chips climbs.

Novice. The work that raises the share of good chips on a new process or design: collect test and inspection data, work out where failing chips are broken, look for causes they share, and change the process or the design.

Expert. Driven by wafer maps, inline inspection, test structures, scan diagnosis of many failing dies, and physical failure analysis. Its speed sets how quickly a product reaches volume and what each good die costs.

Yosys

An open-source synthesis tool. It reads Verilog, cleans up the logic and maps it onto a chip’s building blocks, using the ABC program for the final optimization and mapping. It has scripts for many FPGA families and is also used in open chip (ASIC) flows.

All three levels

Beginner. A free program that reads hardware code and turns it into a list of parts. Anyone can use it or change it.

Novice. An open-source synthesis tool. It reads Verilog, cleans up the logic and maps it onto a chip’s building blocks, using the ABC program for the final optimization and mapping. It has scripts for many FPGA families and is also used in open chip (ASIC) flows.

Expert. Open RTL synthesis framework (YosysHQ), begun by Claire Wolf. Per-family synth_* scripts infer RAM, DSP and carry primitives and call ABC (abc, abc9) for LUT mapping; output is JSON for nextpnr or BLIF/EDIF for other tools.

ZeRO / FSDP (sharded data parallelism)

A way to save memory in data parallelism: instead of every replica storing the full optimizer state, gradients and weights, each stores only its share and fetches the rest when needed.

All three levels

Beginner. A way for chips holding copies of a model to split their notes among themselves, instead of each chip keeping all of them.

Novice. A way to save memory in data parallelism: instead of every replica storing the full optimizer state, gradients and weights, each stores only its share and fetches the rest when needed.

Expert. ZeRO stage 1 shards optimizer state, stage 2 adds gradients, stage 3 (PyTorch FSDP) adds weights. Stages 1–2 move the same 2Ψ2\Psi per step as plain DP; stage 3 moves 3Ψ3\Psi, because weights are all-gathered before use.