Transistors · Chapter 4 of 8 · The device

Speed and power

Every time a switch flips, it uses a little energy and takes a little time. Giving the chip a stronger push of electricity makes it faster, but it runs hotter and drains the battery sooner.

Gate delay comes from charging capacitance through a transistor’s on-resistance. Switching power grows with capacitance, the square of the voltage and the clock frequency, and leakage flows even when nothing switches. Lower voltage saves power but slows every gate.

RC and alpha-power delay models, CV²f switching power, subthreshold and gate leakage, multi-threshold cell libraries, and why supply voltage stopped falling in step with transistor size: the end of Dennard scaling.

Every time a transistor switch flips, it takes a tiny bit of time and uses a tiny bit of energy. A chip flips billions of switches billions of times a second. So those tiny bits add up to the two things people care about most: how fast the chip is, and how much power it uses. That power turns into heat, and in a phone, into a dead battery.

That drip is called . It wastes energy even when the chip is doing nothing.

The main knob designers can turn is the supply voltage. Voltage is the push behind the electricity, like the water pressure in the taps. Turn it up and the chip runs faster but gets hot. Turn it down and it saves power but slows down. This chapter is about that balancing act.

Chip designers spend much of their time trading two quantities against each other: speed, measured as how long a logic gate takes to respond (its ) and therefore how fast the clock can tick; and power, measured in watts and, for a battery or a datacenter bill, as energy per operation.

Both come from the same physical picture, which this chapter builds step by step:

  • Delay is the time a transistor takes to charge or discharge the at a gate’s output, like filling a small tank through a resistor.
  • Switching power is the energy spent on that charging, paid every time a signal flips. It grows with capacitance, with the square of the supply voltage, and with how often things flip.
  • Leakage power is the current that flows through transistors that are supposed to be off. It grows exponentially as the transistor’s turn-on voltage (its , VTV_{\mathrm{T}}) is lowered.

The supply voltage (VDDV_{\mathrm{DD}}) is the strongest knob. Lowering it saves power quickly but slows every gate, and the slowdown gets steep as VDDV_{\mathrm{DD}} approaches VTV_{\mathrm{T}}. Lowering VTV_{\mathrm{T}} buys back speed at the cost of leakage. For decades, shrinking transistors let engineers lower both voltages and make each switch faster and cheaper in energy. That stopped during the 2000s, and the reasons why shape almost every chip designed since.

This chapter builds on the switch model of The switch and the current-voltage behavior in The I-V curve.

You know a MOSFET is a voltage-controlled switch. To predict how fast and how hungry a circuit built from it will be, you need three more pieces: how much current the switch carries when on, how much it leaks when off, and how much capacitance it must charge. This chapter turns those into four working models:

  1. Delay: an RC charging model, refined with the so that delay depends on VDDV_{\mathrm{DD}} and VTV_{\mathrm{T}} the way real short-channel devices do.
  2. Switching power: P=αCVDD2fP = \alpha C V_{\mathrm{DD}}^2 f, and where the energy actually goes.
  3. Leakage: subthreshold conduction, exponential in −VT/(n kT/q)-V_{\mathrm{T}}/(n\,kT/q), plus gate tunneling.
  4. The VDD/VTV_{\mathrm{DD}}/V_{\mathrm{T}} trade: why the two voltages have to be chosen together, how libraries let a design use different points on the curve in different places, and where energy per operation bottoms out.

Then the history that explains the present: Dennard’s 1974 scaling rules, which kept power density constant as transistors shrank, and why they broke when threshold voltage hit a floor set by thermodynamics rather than by lithography.

V_DD1inputPMOSNMOSload C_L00→1 flips0From supply0.00 fJ
Output
Leakage

The output is at 0 V: the NMOS holds the load capacitance empty. Set the output to 1.

A CMOS inverter driving its load capacitance. Every 0→1 at the output costs time and C·V² of energy; the transistor that is off leaks all the time. Values from the chapter’s worked example.Share freely with credit: ‘Figure from chipfieldguide.com’

When a gate changes its answer from 0 to 1, its output wire has to fill up with electric charge. Only then do the next gates notice. Think of the wire as a small bucket that has to fill.

A bigger bucket takes longer. A long wire, or a gate that feeds many other gates, is a bigger bucket. A stronger tap is quicker. More voltage turns a switch on harder, so it fills the bucket faster.

One flip takes a few trillionths of a second. But a chip has a clock: a steady beat that tells every part when to take its next step. On each beat, a signal may pass through dozens of gates in a row, so every gate’s delay matters.

Every gate output is connected to something: a wire, and the inputs of the next gates. Each of those has capacitance, the ability to store charge, measured in femtofarads (fF, 10−1510^{-15} farad). Together they make up the CLC_{\mathrm{L}}. To switch the output from 0 V up to the supply voltage VDDV_{\mathrm{DD}}, a transistor in the gate must push a charge Q=CLVDDQ = C_{\mathrm{L}} V_{\mathrm{DD}} onto it.

A transistor that is switched on behaves roughly like a resistor, with an RR. Charging a capacitor through a resistor is the classic RC circuit: the voltage rises quickly at first and then levels off, and after a time 0.69 RC0.69\,RC it has reached half of VDDV_{\mathrm{DD}}. That half-way time is the usual definition of .

A worked example. Suppose a transistor’s on-resistance is 5 kΩ and it drives 2 fF. Then RC=5,000 Ω×2×10−15 FRC = 5{,}000\,\Omega \times 2 \times 10^{-15}\,\mathrm{F} = 10 picoseconds, and the delay is about 0.69×10≈7 ps0.69 \times 10 \approx 7\,\mathrm{ps}. Double the load, by driving twice as many gates, and the delay doubles. Make the transistor twice as wide and its resistance halves, so the delay halves; but the wider transistor is itself a bigger load for whatever drives it.

Where the voltages come in. The on-resistance isn’t fixed. A transistor conducts because its gate voltage is above its VTV_{\mathrm{T}}, and it carries more current the further above it is. That excess, VDD−VTV_{\mathrm{DD}} - V_{\mathrm{T}}, is called the overdrive. Lower the supply and you lower the overdrive, so the current falls and the delay rises. At VDD=1.0 VV_{\mathrm{DD}} = 1.0\,\mathrm{V} with VT=0.3 VV_{\mathrm{T}} = 0.3\,\mathrm{V} there is 0.7 V of overdrive; at VDD=0.5 VV_{\mathrm{DD}} = 0.5\,\mathrm{V} there is only 0.2 V, and the same gate is about two and a half times slower. (The Expert level shows the formula behind that number.)

From gates to clock speed. Data moves between , storage elements that capture a value on each clock tick. Between two flip-flops it passes through a chain of gates. The slowest such chain, the , must finish within one clock period. With 30 gates of 15 ps each, the path takes 450 ps, so the clock period can be no shorter than that: a maximum frequency of 1/450 ps≈2.2 GHz1/450\,\mathrm{ps} \approx 2.2\,\mathrm{GHz}. (Real designs also reserve time for the flip-flops themselves and for clock uncertainty.)

The RC model

Treat the switching transistor as an effective resistance ReffR_{\mathrm{eff}} charging the load CLC_{\mathrm{L}}. The step response of an RC node crosses 50% at t=ln⁡2⋅RC≈0.69 RCt = \ln 2 \cdot RC \approx 0.69\,RC, so to first order tpd≈0.69 ReffCLt_{\mathrm{pd}} \approx 0.69\,R_{\mathrm{eff}} C_{\mathrm{L}}. For a chain or tree of RC segments the Elmore delay, the sum over each resistor of RiR_{i} times all the capacitance downstream of it, generalizes this. ReffR_{\mathrm{eff}} scales as 1/W1/W (transistor width), and CL=Cwire+∑Cin+CselfC_{\mathrm{L}} = C_{\mathrm{wire}} + \sum C_{\mathrm{in}} + C_{\mathrm{self}}, where each driven gate’s input capacitance also scales with its width. That coupling is why sizing is an optimization rather than “make everything wide.” A common process-independent unit for the result is the : an inverter driving four copies of itself.

Putting the voltages in

ReffR_{\mathrm{eff}} hides the voltage dependence. A more useful first-order form treats the transistor as a current source discharging the load: the time to move the output by VDD/2V_{\mathrm{DD}}/2 is

tpd≈CL VDD2 Iont_{\mathrm{pd}} \approx \frac{C_{\mathrm{L}}\, V_{\mathrm{DD}}}{2\, I_{\mathrm{on}}}

and everything interesting is in IonI_{\mathrm{on}}. The long-channel square law gives Ion∝(VDD−VT)2I_{\mathrm{on}} \propto (V_{\mathrm{DD}} - V_{\mathrm{T}})^2. Short-channel devices don’t follow it: once the lateral field is high enough, carrier velocity saturates and current grows more slowly with overdrive (see The I-V curve). Sakurai and Newton’s captures this with one fitted exponent:

Ion∝(VDD−VT)α,1≤α≤2I_{\mathrm{on}} \propto (V_{\mathrm{DD}} - V_{\mathrm{T}})^{\alpha}, \qquad 1 \le \alpha \le 2

α=2\alpha = 2 recovers the square law, and α\alpha falls toward 1 as velocity saturation gets more severe: Sakurai and Newton note that measured short-channel IDI_{\mathrm{D}}–VGSV_{\mathrm{GS}} curves are close to linear. Their closed-form inverter delay is a term proportional to the input transition time plus CLVDD/(2ID0)C_{\mathrm{L}} V_{\mathrm{DD}}/(2 I_{\mathrm{D0}}), where ID0I_{\mathrm{D0}} is the drain current at VGS=VDS=VDDV_{\mathrm{GS}} = V_{\mathrm{DS}} = V_{\mathrm{DD}}. With ID0∝(VDD−VT)αI_{\mathrm{D0}} \propto (V_{\mathrm{DD}} - V_{\mathrm{T}})^{\alpha}, delay goes as VDD/(VDD−VT)αV_{\mathrm{DD}}/(V_{\mathrm{DD}} - V_{\mathrm{T}})^{\alpha}. Fitted values for modern processes land in between; Harris quotes α≈1.3\alpha \approx 1.3 for a 65 nm process.

Step by step, with α=1.3\alpha = 1.3 and VT=0.3 VV_{\mathrm{T}} = 0.3\,\mathrm{V} (delay in arbitrary units, VDD/(VDD−VT)1.3V_{\mathrm{DD}}/(V_{\mathrm{DD}} - V_{\mathrm{T}})^{1.3}):

  • VDD=1.0 VV_{\mathrm{DD}} = 1.0\,\mathrm{V}: 1.0/0.71.3=1.0/0.629=1.591.0/0.7^{1.3} = 1.0/0.629 = 1.59.
  • VDD=0.7 VV_{\mathrm{DD}} = 0.7\,\mathrm{V}: 0.7/0.41.3=0.7/0.304=2.300.7/0.4^{1.3} = 0.7/0.304 = 2.30. That is 45% slower, for a 51% cut in switching energy (0.72=0.490.7^2 = 0.49).
  • VDD=0.5 VV_{\mathrm{DD}} = 0.5\,\mathrm{V}: 2.5× slower than at 1.0 V. At 0.4 V: 5× slower. The curve turns sharply upward as the overdrive vanishes.
  • Lowering VTV_{\mathrm{T}} from 0.3 V to 0.2 V cuts delay by 25% at VDD=0.7 VV_{\mathrm{DD}} = 0.7\,\mathrm{V} but only 16% at 1.0 V. Threshold matters most when the supply is low.

Below VTV_{\mathrm{T}} the alpha-power law no longer applies. The transistor conducts only by diffusion, current is exponential in VGSV_{\mathrm{GS}}, and so is delay. That regime is where the minimum-energy story in “Under the hood” plays out.

Real libraries don’t evaluate formulas. They characterize every cell in circuit simulation over a grid of input slews and output loads and store the delays in tables, one set per process, voltage and temperature corner (see From devices to a cell library). The models here are for intuition and early estimates.

FFFFcritical path: 30 gates between flip-flopsOne gate4 fF14 psPath delay0 ps400 ps800 ps1200 psf_max ≈ 2.4 GHz
Fan-out (gates driven)

Each gate: 0.69 × 5 kΩ × 4 fF ≈ 14 ps. 30 gates: 414 ps, so the clock can tick at most 2.4 GHz.

A critical path between two flip-flops. Delay per gate grows with its load (here, how many gates it drives); the path’s total sets the fastest clock. Values from the chapter’s worked example; wires and flip-flop overheads left out.Share freely with credit: ‘Figure from chipfieldguide.com’

Back to the buckets. Every time a wire goes from 0 to 1, a bucket fills from the power supply. When it goes back to 0, that water goes down the drain. Bigger buckets waste more each time.

Voltage matters even more, and this is the surprise. Double the voltage and each flip costs four times the energy. Cut the voltage by 30% and each flip costs about half.

Power is the energy used each second. So it also depends on how often wires flip. A faster clock means more flips each second. Luckily, most wires don’t flip on most beats.

To charge a load capacitance CC to voltage VV, the supply delivers charge Q=CVQ = CV at voltage VV, so it supplies energy CV2CV^2. Half of that ends up stored on the capacitor, and the other half turns into heat in the transistor that did the charging. When the output later falls back to 0, the stored half is dumped as heat in the transistor that discharges it. So each full up-and-down cycle costs CV2CV^2, no matter how fast or slow it happens.

Across a whole chip that gives the most important equation in low-power design:

Pdynamic=αCVDD2fP_{\mathrm{dynamic}} = \alpha C V_{\mathrm{DD}}^2 f
  • CC is the total capacitance that could switch: every gate input and every wire.
  • VDDV_{\mathrm{DD}} is the supply voltage, squared.
  • ff is the clock frequency.
  • α\alpha is the : the fraction of that capacitance that actually charges up in a typical clock cycle. A clock wire switches every cycle; ordinary logic far less often.

A worked example. Take a block with C=1 nFC = 1\,\mathrm{nF} of total switchable capacitance (very roughly a million gates’ worth), α=0.1\alpha = 0.1, VDD=0.8 VV_{\mathrm{DD}} = 0.8\,\mathrm{V}, f=1 GHzf = 1\,\mathrm{GHz}:

P=0.1×1×10−9 F×0.82 V2×1×109 Hz=0.064 W=64 mW\begin{aligned} P &= 0.1 \times 1 \times 10^{-9}\,\mathrm{F} \times 0.8^2\,\mathrm{V^2} \\ &\quad \times 1 \times 10^{9}\,\mathrm{Hz} \\ &= 0.064\,\mathrm{W} = 64\,\mathrm{mW} \end{aligned}

Drop the supply to 0.6 V at the same clock and power falls to 36 mW, a 44% cut from a 25% voltage reduction. That is the default point in the simulation below; set the supply slider to 0.6 V to see it. (The “Low power” preset also halves the clock, which halves the power again.)

Each factor has its own tools. Designers cut α\alpha with (stopping the clock to idle blocks), cut CC with smaller cells and shorter wires, and cut VV and ff together with , dynamic voltage and frequency scaling, which lowers both when there is little work to do.

A small extra cost, , flows during each transition while both halves of a gate are briefly on at once. In a well-designed circuit, where input and output signals change at similar speeds, it stays under about a tenth of switching power.

Where the energy goes

During a 0→1 transition the supply delivers ∫VDD i dt=VDDQ=CLVDD2\int V_{\mathrm{DD}}\, i\, dt = V_{\mathrm{DD}} Q = C_{\mathrm{L}} V_{\mathrm{DD}}^2. The capacitor ends with 12CLVDD2\tfrac{1}{2} C_{\mathrm{L}} V_{\mathrm{DD}}^2; the other half is dissipated in the pull-up network’s resistance, independent of that resistance’s value. On 1→0 the stored 12CLVDD2\tfrac{1}{2} C_{\mathrm{L}} V_{\mathrm{DD}}^2 is dissipated in the pull-down, and the supply delivers nothing. So the energy per complete cycle is CLVDD2C_{\mathrm{L}} V_{\mathrm{DD}}^2, and transistor sizing or switching speed doesn’t change it, only CC and VV do.

Summed over all nodes, with α\alpha defined as the probability a node makes a 0→1 transition in a cycle, Psw=αCVDD2fP_{\mathrm{sw}} = \alpha C V_{\mathrm{DD}}^2 f. Conventions differ: some texts count both edges and write 12αCV2f\tfrac{1}{2} \alpha C V^2 f, so check which α\alpha a tool or paper reports. For a clock net α=1\alpha = 1 under this definition. Completely random data gives α=0.25\alpha = 0.25 (a node is 0 in one cycle and 1 in the next with probability 12×12\tfrac{1}{2} \times \tfrac{1}{2}), and typical logic sits near 0.1, because data passing through ANDs and ORs is biased toward one value. Glitches, spurious transitions before a node settles, push real activity back up.

The cubic lever

Switching power is quadratic in VDDV_{\mathrm{DD}} at fixed ff. But fmaxf_{\mathrm{max}} itself depends on VDDV_{\mathrm{DD}}: from the alpha-power delay, fmax∝(VDD−VT)α/VDDf_{\mathrm{max}} \propto (V_{\mathrm{DD}} - V_{\mathrm{T}})^{\alpha}/V_{\mathrm{DD}}, which is roughly linear in VDDV_{\mathrm{DD}} over a typical DVFS range. Scaling VV and ff together therefore gives P∝V3P \propto V^3 approximately, while energy per operation goes as V2V^2. That is the case for running wide and slow: Chandrakasan, Sheng and Brodersen showed in 1992 that you can lower the supply and recover the lost throughput with parallel or pipelined hardware, trading area for a large drop in power at the same throughput. Looking back on the paper, Chandrakasan describes this architectural voltage scaling as improving the power-delay product by an order of magnitude without losing performance.

Example. Two copies of a datapath at half the clock deliver the same throughput as one at full clock. If halving ff lets VDDV_{\mathrm{DD}} fall from 1.0 V to 0.7 V, each copy burns 0.49×0.50.49 \times 0.5 of the original, so the pair burns 0.49: half the power, for twice the area plus the cost of splitting and merging the data.

Short-circuit power

While the input slews through the window VT,n<Vin<VDD−∣VT,p∣V_{\mathrm{T,n}} < V_{\mathrm{in}} < V_{\mathrm{DD}} - |V_{\mathrm{T,p}}|, both networks conduct and a crowbar current flows. Its energy scales with input transition time, so keeping input and output slews comparable keeps it small: under 10% of dynamic power by Harris’s estimate. If VDD<VT,n+∣VT,p∣V_{\mathrm{DD}} < V_{\mathrm{T,n}} + |V_{\mathrm{T,p}}| the window closes altogether, so short-circuit current vanishes at very low supply voltages.

Vcharge Q →0.60.6 VheatstoredOne 0→1, 2 fF wireSupply gives C·V²0.72 fJStored ½C·V²0.36 fJHeat in pull-up0.36 fJBlock, α = 0.1, 1 GHz: 36 mW(1.00× the 0.6 V cost)
Transition

V_DD = 0.60 V: the supply gives C·V² = 0.72 fJ to fill a 2 fF wire, 1.00× the 0.6 V cost. Half is stored, half is heat in the pull-up.

Charging a 2 fF wire: the supply delivers charge Q = CV at voltage V, the whole box (C·V²). The wire stores the lower triangle; the upper one is heat. Block power assumes the chapter’s 1 nF, α = 0.1, 1 GHz.Share freely with credit: ‘Figure from chipfieldguide.com’

A switch that is “off” isn’t perfectly off. A tiny bit of electricity still trickles through, like a dripping tap. One drip is nothing. But billions of drips, all day long, add up to . That is one reason a phone sitting in your pocket still loses charge.

Two things make the drip worse. First, a switch built to turn on more easily is faster, but it doesn’t close as tightly. Making it a little easier to turn on can make the drip ten times bigger. Second, heat: hot switches leak more, and leaking makes them hotter.

Leakage has several sources, but the biggest in most logic is subthreshold current: current that flows from drain to source even when the gate voltage is below the . Below threshold the transistor doesn’t switch off abruptly. Its current falls exponentially as the gate voltage drops, by a fixed factor of ten for every so many millivolts.

That number of millivolts per factor of ten is the . Physics sets a best case of about 60 mV per decade at room temperature for a conventional transistor; real devices are somewhat worse: planar transistors of the 2000s were typically around 90–100 mV per decade, and FinFETs get close to 60–70. The consequence is stark:

  • A transistor that is off has a gate voltage of 0, which is VTV_{\mathrm{T}} below the point where it turns on. Its leakage is roughly its current at threshold divided by 10 for every slope’s-worth of millivolts in VTV_{\mathrm{T}}.
  • With a slope of 90 mV per decade, lowering VTV_{\mathrm{T}} from 0.35 V to 0.26 V makes every off transistor leak ten times more. Lowering it to 0.17 V makes it a hundred times more.

Leakage power is simply Pleak=VDD IleakP_{\mathrm{leak}} = V_{\mathrm{DD}}\, I_{\mathrm{leak}}, the supply voltage times the total off-current of every transistor on the chip. Unlike switching power, it doesn’t depend on the clock or on activity. It does rise with temperature, because the threshold voltage drops as a chip heats up (and the subthreshold slope, proportional to absolute temperature, gets worse).

The second source is . The silicon dioxide insulator under the gate was thinned generation after generation, to about 1.2 nm, to keep the gate in control of the channel. At that thickness electrons tunnel straight through it. The industry’s fix, first used in 45 nm microprocessors in 2007, was a different insulator with a higher dielectric constant (“high-k”) and a metal gate. A high-k layer can be physically thicker while giving the gate the same control, and it cut gate leakage by more than ten times.

Subthreshold conduction

Below threshold the channel is weakly inverted and current flows by diffusion, so it follows the Boltzmann distribution of carriers over the source barrier:

Isub=I0exp⁡ ⁣(VGS−VTn kT/q)×(1−exp⁡ ⁣(−VDSkT/q))\begin{aligned} I_{\mathrm{sub}} &= I_0 \exp\!\left(\frac{V_{\mathrm{GS}} - V_{\mathrm{T}}}{n\,kT/q}\right) \\ &\quad \times \left(1 - \exp\!\left(-\frac{V_{\mathrm{DS}}}{kT/q}\right)\right) \end{aligned}

where kT/q≈25.9 mVkT/q \approx 25.9\,\mathrm{mV} at 300 K and n≥1n \ge 1 is the ideality factor, set by how the gate capacitance divides with the depletion capacitance underneath it. For VDSV_{\mathrm{DS}} more than a few kT/qkT/q the last factor is 1. Converting to base 10, current changes one decade per millivolts: 59.5 mV/decade for n=1n = 1 at 300 K, and 89 mV/decade for n=1.5n = 1.5 (the value used in the simulation).

Setting VGS=0V_{\mathrm{GS}} = 0 gives the off-current Ioff≈I0⋅10−VT/SI_{\mathrm{off}} \approx I_0 \cdot 10^{-V_{\mathrm{T}}/S}. Three things make it worse in practice:

  • . In short channels the drain field lowers the source barrier, so VTV_{\mathrm{T}} falls with VDSV_{\mathrm{DS}} (VT,eff=VT−ηVDSV_{\mathrm{T,eff}} = V_{\mathrm{T}} - \eta V_{\mathrm{DS}}, with η\eta around 0.1 in Harris’s 65 nm example). Leakage then rises exponentially with VDDV_{\mathrm{DD}} as well, which is one reason lowering VDDV_{\mathrm{DD}} cuts leakage power by more than the linear VDDV_{\mathrm{DD}} factor suggests.
  • Temperature. Heat lowers VTV_{\mathrm{T}}, and kT/qkT/q, and with it SS, grows in proportion to absolute temperature (SS at 85 °C is 71 mV/decade even with n=1n = 1). Off-current therefore rises with temperature while on-current falls, which is why leakage is estimated at the hot corner.
  • Stacking helps. Two off transistors in series leak about ten times less than one, because the node between them floats up and gives the upper device a negative VGSV_{\mathrm{GS}}. Designers exploit this during sleep by choosing input values that turn off stacked devices.

Gate and junction leakage

is carriers tunneling through the gate dielectric. It is exponentially sensitive to oxide thickness: negligible above about 2 nm, critically important at 65 nm and below, where the oxide is around 1 nm and gate leakage approaches subthreshold leakage. SiO2\mathrm{SiO_2} gate insulators stopped thinning at about 1.2 nm. High-k dielectrics with metal gates, first used in 45 nm microprocessors in 2007, restored a physically thicker barrier at the same capacitance and cut gate leakage by more than ten times. Junction leakage comes from the reverse-biased source and drain diodes, including band-to-band tunneling and gate-induced drain leakage at high doping levels.

The on/off ratio

The useful figure of merit is Ion/IoffI_{\mathrm{on}}/I_{\mathrm{off}}. With VDD=0.8 VV_{\mathrm{DD}} = 0.8\,\mathrm{V}, VT=0.35 VV_{\mathrm{T}} = 0.35\,\mathrm{V} and S=89 mV/decadeS = 89\,\mathrm{mV/decade}, the simulation’s device has Ion/Ioff≈1.4×105I_{\mathrm{on}}/I_{\mathrm{off}} \approx 1.4 \times 10^5. Every 89 mV taken off VTV_{\mathrm{T}} multiplies IoffI_{\mathrm{off}} by ten while raising IonI_{\mathrm{on}} only modestly through the larger overdrive. Raising VDDV_{\mathrm{DD}} instead raises IonI_{\mathrm{on}} but costs switching power as VDD2V_{\mathrm{DD}}^2. That is the trade the next section is about.

S0 VG0 VDV_DDp-type body (0 V)source n+drain n+gatesubthresholdgatejunction
Threshold V_T
Temperature
Gate insulator

Transistor off (V_GS = 0, drain at V_DD): normal V_T; gate leakage through 1.2 nm SiO₂. Tap a path.

An NMOS that is off, in cross-section (not to scale), with its three leakage paths. Arrow widths are relative and illustrative; the 10× step for a 90 mV lower V_T and the >10× cut from high-k come from the chapter’s sources.Share freely with credit: ‘Figure from chipfieldguide.com’

Put the pieces together and you get two knobs that pull in opposite ways.

The first is the supply voltage. Higher makes switches faster but burns much more power. Lower saves a lot of power but slows everything down.

The second is built into each switch: its , the push it needs before it turns on. A lower threshold makes switches faster but leakier. A higher one leaks less but is slower.

Engineers pick the pair that fits the job. A phone chip resting in your pocket wants low leakage above all. A big server chip working nonstop wants speed. Many chips also change their voltage as they run: high when there is work, low when there isn’t.

Delay and the two kinds of power pull VDDV_{\mathrm{DD}} and VTV_{\mathrm{T}} in different directions:

ChangeDelaySwitching powerLeakage power
Raise VDDV_{\mathrm{DD}}ShorterUp with VDD2V_{\mathrm{DD}}^2Up (at least in proportion to VDDV_{\mathrm{DD}})
Lower VDDV_{\mathrm{DD}}Longer, steeply near VTV_{\mathrm{T}}Down with VDD2V_{\mathrm{DD}}^2Down
Lower VTV_{\mathrm{T}}ShorterAbout the sameUp about tenfold per 60–100 mV
Raise VTV_{\mathrm{T}}LongerAbout the sameDown about tenfold per 60–100 mV

So the two voltages have to be chosen together. A useful rule of thumb is that speed depends mostly on the overdrive VDD−VTV_{\mathrm{DD}} - V_{\mathrm{T}}, while switching power depends on VDDV_{\mathrm{DD}} and leakage on VTV_{\mathrm{T}}. To keep the same speed at a lower supply you must lower VTV_{\mathrm{T}} as well, by a good fraction of the supply drop (about half in the figure below), and pay for it in leakage: roughly ten times more for every 100 mV.

Multi-threshold libraries

Foundries make transistors with several threshold voltages on the same chip, for example by adding an extra implant step that adjusts the doping under the gate. A cell library then offers each logic gate in several , commonly called high-VTV_{\mathrm{T}} (HVT), regular (RVT), low (LVT) and super-low (SLVT). The flavors of one gate do the same job and can be swapped for each other. The open SkyWater SKY130 process offers low-VTV_{\mathrm{T}} versions of its 1.8 V NMOS and PMOS transistors and a high-VTV_{\mathrm{T}} PMOS, and its cell libraries mix them: the high-speed library uses low-VTV_{\mathrm{T}} transistors and “has the highest speed and the highest leakage,” while a high-density, low-leakage library uses the high-VTV_{\mathrm{T}} PMOS and leaks 5–10 times less than its regular counterpart. The ASAP7 predictive 7 nm kit models four flavors, because, in its authors’ words, multiple threshold voltages have “become essential to meet both performance and standby power constraints.”

The rule is to use low-VTV_{\mathrm{T}} cells only where speed is critical. Implementation tools start from slower, low-leakage cells and swap in faster ones only where a path would otherwise miss its clock. The open-source OpenROAD flow, for instance, performs this “VTV_{\mathrm{T}} swap” as part of timing repair and can list which cells are threshold equivalents.

Changing voltage at run time

Because switching power falls with VDD2V_{\mathrm{DD}}^2 and the clock can fall with VDDV_{\mathrm{DD}}, many chips use : a power manager picks from a set of voltage and frequency pairs, called operating points, according to the workload. Blocks that are idle for long can be disconnected from the supply entirely with , which stops their leakage too.

Choosing VDDV_{\mathrm{DD}} and VTV_{\mathrm{T}} together

Fix a target delay. From the alpha-power model, delay ∝VDD/(VDD−VT)α\propto V_{\mathrm{DD}}/(V_{\mathrm{DD}} - V_{\mathrm{T}})^{\alpha}, so a family of (VDD,VT)(V_{\mathrm{DD}}, V_{\mathrm{T}}) pairs meets it, roughly a line of constant overdrive with a gentle tilt. Along that line, switching energy falls as VDD2V_{\mathrm{DD}}^2 while leakage rises as 10−VT/S10^{-V_{\mathrm{T}}/S}. Total power has a minimum where the marginal saving in αCVDD2f\alpha C V_{\mathrm{DD}}^2 f equals the marginal cost in VDDIoffV_{\mathrm{DD}} I_{\mathrm{off}}.

Horowitz frames this as matching marginal costs. At the optimum, the energy spent per unit of delay saved must be the same whether it is bought with VDDV_{\mathrm{DD}} or with VTV_{\mathrm{T}}, and both equal the slope of the energy-performance Pareto curve. So neither voltage is set directly by scaling, and leakage was allowed to rise because that lowered total power. The National Research Council’s 2011 report puts the resulting balance at static leakage power of roughly 30% of dynamic power.

The optimum depends on activity. A busy datapath (high α\alpha) can afford more leakage per unit of switching and wants a lower VTV_{\mathrm{T}}; mostly idle logic and memory (low α\alpha) leak for a long time per useful transition and want a higher one. There is no single best VTV_{\mathrm{T}}, which is why a process offers several.

Multi-VTV_{\mathrm{T}} assignment

In SKY130 the low-VTV_{\mathrm{T}} NMOS has the same cross-section as the standard device “except for the VTV_{\mathrm{T}} adjust implants.” The kit exposes nfet_01v8_lvt, pfet_01v8_lvt and pfet_01v8_hvt alongside the standard nfet_01v8 and pfet_01v8, and its seven cell libraries are largely defined by which pair they use: hs pairs the two low-VTV_{\mathrm{T}} devices, hd the two standard ones, and hdll and lp the standard NMOS with the high-VTV_{\mathrm{T}} PMOS. ASAP7 provides SLVT, LVT, RVT and SRAM flavors, in decreasing order of drive strength, with off-current dropping by about an order of magnitude per step.

Tools use the flavors in two directions. During timing repair they swap cells on failing paths to a faster flavor (OpenROAD’s repair_timing does this by default unless told to skip it). During power recovery they resize cells on paths with positive slack to save power, which with a multi-VTV_{\mathrm{T}} library includes moving them to slower, lower-leakage flavors. The result is a design that is mostly high-VTV_{\mathrm{T}} by cell count, with low-VTV_{\mathrm{T}} concentrated on the critical paths; Harris’s worked example of a billion-transistor chip assumes high-VTV_{\mathrm{T}} in all memories and 95% of logic gates. Because leakage is exponential in VTV_{\mathrm{T}} and rises with temperature, leakage is checked at the hot, fast corner and timing at the slow corner (see ).

DVFS and power gating

moves along the delay-VDDV_{\mathrm{DD}} curve at run time. Each operating point needs its own timing signoff, a regulator that can slew between voltages, and a clock that can change frequency, so designs support a handful of points rather than a continuum. Lowering VDDV_{\mathrm{DD}} cuts leakage power more than linearly because of DIBL, but leakage energy per operation rises once each operation takes much longer: that tension produces the minimum energy point described in “Under the hood.” For long idle periods removes the supply with a header or footer switch, eliminating leakage at the cost of state loss (or retention flip-flops) and wake-up time, and the switch’s own voltage drop slows the gated block while it runs.

The two voltagesV_DDV_T0.800.35overdrivePower at equal speedminimumswitchingleakage0.60.81.0V_DD (V) →
Activity α

V_DD = 0.80 V needs V_T = 0.35 V for the same 15 ps gates. Switching 64.0 mW, leakage 6.4 mW, total 70.4 mW (lowest: 68.6 mW at 0.76 V).

Pairs of V_DD and V_T that keep the same gate delay, with the power of the chapter’s 1 nF block at 1 GHz. The best pair depends on activity. Same toy model as the simulation; illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

In 1974 a team at IBM led by Robert Dennard found a recipe. Shrink every part of a transistor by the same amount, and lower its voltage by the same amount. The smaller transistor is faster and uses less power. And a chip packed with them gives off the same heat for its size as before. This became known as .

For about thirty years the rule held. Each new generation of chips held more transistors and ran faster. (Real chips also pushed their clocks harder, so they did get hotter.)

Then, in the early 2000s, the voltage stopped going down. To keep switches fast at a lower voltage, you must make them easier to turn on, and that makes them leak far more. So the voltage got stuck. Transistors kept shrinking, and packing in more of them meant more heat in the same space.

That is why chip clocks stopped getting faster around 2005. Instead, chips got more cores: several smaller processors working side by side. It is also why a chip today can’t run every part at full speed at once.

Dennard and his colleagues at IBM showed in 1974 how to shrink a MOSFET by a factor κ\kappa (kappa) and keep it working the same way: divide every dimension (gate length, width, oxide thickness) by κ\kappa, divide the voltages by κ\kappa, and multiply the doping by κ\kappa. Their table of results:

QuantityScales by
Dimensions, voltage, current1/κ1/\kappa
Capacitance1/κ1/\kappa
Delay per circuit (VC/IVC/I)1/κ1/\kappa
Power per circuit (VIVI)1/κ21/\kappa^2
Power density (VI/areaVI/\text{area})1, unchanged

In the paper’s words, “even if many more circuits are placed on a given integrated circuit chip, the cooling problem is essentially unchanged.” With κ≈1.4\kappa \approx 1.4 (2\sqrt{2}) per generation, each new process doubled the transistors per area and cut gate delay by about 30%, all at the same power per square millimeter. You can check it with P=αCV2fP = \alpha C V^2 f: CC and VV each shrink by 1/κ1/\kappa and ff grows by κ\kappa, so each gate’s power falls by 1/κ21/\kappa^2, which exactly cancels the κ2\kappa^2 more gates per area.

Why it stopped. The recipe needs the threshold voltage to shrink along with the supply, so that the overdrive VDD−VTV_{\mathrm{DD}} - V_{\mathrm{T}} keeps up. But leakage grows about tenfold for every 100 mV that VTV_{\mathrm{T}} drops, and that rate is set by temperature (through kT/qkT/q), not by size, so shrinking doesn’t improve it. By roughly the 130 nm generation, at the start of the 2000s, leakage had grown large enough that VTV_{\mathrm{T}} stopped falling, and supply voltage scaling slowed with it to keep VDD−VTV_{\mathrm{DD}} - V_{\mathrm{T}} roughly constant.

With voltage nearly fixed, P=αCV2fP = \alpha C V^2 f per gate no longer falls fast enough to offset the density gain, so power per area rises with each generation unless frequency or activity comes down. Processor makers hit this “power wall” around 2005: clock frequencies stopped rising, and the extra transistors went into more cores instead. Extrapolating, Esmaeilzadeh and colleagues estimated in 2011 that a growing fraction of a chip would have to be kept off at any moment, which they called . Specialized hardware, the subject of the AI chapters in the Architectures guide, is one answer: chips that do one job with far less energy per operation. What the extra cores bought, and what they cost, is the subject of Multicore and vector units.

Dennard’s constant-field rules

Dennard, Gaensslen, Yu, Rideout, Bassous and LeBlanc scaled toxt_{\mathrm{ox}}, LL and WW by 1/κ1/\kappa, the substrate doping by κ\kappa and all voltages by 1/κ1/\kappa, which keeps the electric-field pattern inside the device unchanged. Their Table I gives current per device 1/κ1/\kappa, capacitance 1/κ1/\kappa, delay VC/I∝1/κVC/I \propto 1/\kappa, power per circuit VI∝1/κ2VI \propto 1/\kappa^2, and power density constant. Harris’s version of the same table adds switching energy per gate, CV2CV^2, falling as 1/κ31/\kappa^3.

The paper flagged the exception itself: “one must accept the fact that the subthreshold behavior does not scale as desired.” The subthreshold slope is n (kT/q)ln⁡10n\,(kT/q) \ln 10, and kT/qkT/q is a property of the operating temperature, not of the device size. For two decades that was harmless, because VTV_{\mathrm{T}} was large compared with the slope. Historically VTV_{\mathrm{T}} sat around 800 mV, and each generation that lowered it paid about a decade of off-current per 100 mV removed.

What the data show

  • Supply voltage. CPU DB, a database of decades of microprocessors, shows the knee: processors stayed at 5 V until about the 0.6 µm generation, then scaled roughly as the square root of feature size down to 0.13 µm, and scaling slowed from 0.13 µm on.
  • Power density. From the 80386 to 2005, feature size scaled about 16×, supply voltage about 4× and frequency about 200×, and processor power density rose more than 32×.
  • Frequency. In the processor trend data the fastest clock rises from about 200 MHz in 1993 to 3.65 GHz in 2004 and never passes 4.7 GHz afterwards.

The ASAP7 authors summarize the cause from a device-kit perspective: Dennard scaling “has given way due to difficulty in reducing VDDV_{\mathrm{DD}} due to Ion/IoffI_{\mathrm{on}}/I_{\mathrm{off}} considerations, i.e., vanishing room for threshold voltage scaling while maintaining gate overdrive, and variability.”

Scaling with fixed voltage

Scale dimensions by 1/κ1/\kappa with VDDV_{\mathrm{DD}} and VTV_{\mathrm{T}} fixed. CC per gate still falls as 1/κ1/\kappa and density rises as κ2\kappa^2, so switching power per area rises as κ\kappa at constant frequency, and as κ2\kappa^2 if you also take the frequency gain. Every later technique in this guide and the next is partly a response: FinFETs and gate-all-around devices improve gate control and therefore nn and the subthreshold slope (ASAP7’s modeled FinFETs sit at about 63 mV/decade, close to the 60 mV limit; see Shrinking), while multicore, DVFS, power gating and specialization spend transistors to save energy.

Esmaeilzadeh et al. combined device projections with multicore performance models and estimated that 21% of a fixed-size chip would have to be powered off at 22 nm, rising to more than 50% at 8 nm. The exact percentages depend on their projections; the lasting point is that design became limited by power rather than by area. How architects spend a fixed power budget on cores, clock and vector width is worked through in Multicore and vector units.

Fixed die area1× transistorsPower per area, all onbudget1×01234generation →
Scaling
1 / 5

Generation 0. Step forward to scale by κ ≈ 1.4 each time, voltage included.

Scaling by κ ≈ 1.4 per generation doubles the transistors per area. Dennard’s rules keep power density constant; with V_DD stuck (and the clock held) it rises by κ each generation. Idealized scaling rules, not measured data.Share freely with credit: ‘Figure from chipfieldguide.com’

This is a pretend piece of a chip. It has a chain of 30 switches that must all finish before each clock beat. Move the sliders for supply voltage, turn-on voltage and clock speed. Watch the time per switch and the power bar.

Try the “Low power” button, then “Turbo,” then “Low VT” (VT is short for turn-on voltage). See which part of the power bar grows. Then push the clock up until it fails, and find a voltage that makes it work again.

The simulation models a block with 1 nF of switchable capacitance, an activity factor of 0.1 and a critical path of 30 gates. Change VDDV_{\mathrm{DD}}, VTV_{\mathrm{T}} and the clock frequency. The chart shows gate delay against supply voltage on a logarithmic scale, with the clock’s limit as a red dashed line: points above it miss timing. Things to try:

  • From the default (0.8 V, 1 GHz), lower VDDV_{\mathrm{DD}} to 0.6 V. Switching power falls by 44%.
  • Lower VTV_{\mathrm{T}} from 0.35 V to 0.26 V and watch leakage grow about tenfold.
  • Tick “Clock at the maximum frequency” and slide VDDV_{\mathrm{DD}} down toward VTV_{\mathrm{T}}. The maximum clock collapses while the power keeps falling.

The model: IonI_{\mathrm{on}} follows the alpha-power law (α=1.3\alpha = 1.3) above threshold and an exponential with n=1.5n = 1.5 (S≈89 mV/decadeS \approx 89\,\mathrm{mV/decade}) below it, joined smoothly. Gate delay is KVDD/IonK V_{\mathrm{DD}}/I_{\mathrm{on}}, calibrated to 15 ps at 0.8 V and VT=0.35 VV_{\mathrm{T}} = 0.35\,\mathrm{V}, and fmax=1/(30 tpd)f_{\mathrm{max}} = 1/(30\, t_{\mathrm{pd}}). Leakage is the same current model at VGS=0V_{\mathrm{GS}} = 0, scaled with CC. DIBL and temperature are left out. You also get α\alpha and CC, and a second chart: energy per cycle when clocked at fmaxf_{\mathrm{max}} for each VDDV_{\mathrm{DD}}, decomposed into αCV2\alpha C V^2 and leakage energy. Find the minimum, then see how it moves when you change α\alpha (try 0.02) or VTV_{\mathrm{T}}, and try the “Near threshold” preset.

Loading simulation…
Supply voltage, older free chip kit
1.8 volts
Supply voltage, modern free chip kit
0.7 volts
Energy saved by running very slowly
up to ~10×

What these numbers mean:

  • 1.8 and 0.7 volts. For comparison, an AA battery gives 1.5 volts. Two free chip-making kits show how far voltage fell as transistors shrank. The older one runs at 1.8 volts. A model of a modern process runs at 0.7. Each flip at 0.7 volts costs less than a sixth of the energy.
  • About 10× either way. Running a chip just above the point where its switches turn on can use up to ten times less energy for each job. But the chip also runs about ten times slower.
  • Ten times the leak. In the modern kit, each step to a faster kind of switch leaks about ten times more.
Best subthreshold slope at room temperature
60 mV/decade
Leakage increase per 100 mV lower VT (typical)
~10×
ASAP7 nominal VDD
0.7 V

Sources: the slope limit and typical values from university lecture notes; the tenfold rule from the National Research Council’s 2011 report; the supply from the ASAP7 paper.

Four threshold flavors in an open 7 nm kit

The ASAP7 predictive kit models a FinFET process with four threshold-voltage choices. Its NMOS numbers per fin (typical corner, 25 °C) show the trade directly:

FlavorThreshold (V)On-current (µA)Off-current (nA)
SRAM (highest VTV_{\mathrm{T}})0.2528.60.001
RVT (regular)0.1737.90.019
LVT (low)0.1045.20.242
SLVT (super-low)0.0450.82.444

Going from RVT to SLVT buys 34% more drive current for about 130 times more leakage. That asymmetry is why low-VTV_{\mathrm{T}} cells are spent only where timing needs them.

Two libraries in SKY130

SKY130’s high-density library (hd) and its high-density, low-leakage sibling (hdll) share the same cell height and pin grid. The low-leakage version uses the high-VTV_{\mathrm{T}} PMOS. Its documented leakage is 0.08 nA per thousand gates against 0.86 nA for hd (typical leakage corner, 1.8 V, 25 °C), roughly ten times lower.

Threshold flavor and delay in SKY130

The SKY130 device documentation lists a 143-stage, fanout-of-one inverter chain for each device pairing at the typical corner. Changing only the threshold flavor moves the stage delay by roughly −10% to +20%:

NMOS / PMOSFO1 delay, TT
nfet_01v8_lvt / pfet_01v828.6 ps
nfet_01v8 / pfet_01v831.8 ps (24.7 ps FF, 44.1 ps SS)
nfet_01v8 / pfet_01v8_hvt38 ps

The process corners matter as much as the flavor: the standard pair spans 24.7 ps at the fast corner to 44.1 ps at the slow one.

A textbook leakage budget

Harris works through an example: a one-billion-transistor chip at 1.0 V with subthreshold leakage of 100 nA/µm for normal-VTV_{\mathrm{T}} and 10 nA/µm for high-VTV_{\mathrm{T}} devices, high-VTV_{\mathrm{T}} in all memories and 95% of logic gates, and gate leakage of 5 nA/µm. The result is 584 mA of subthreshold and 275 mA of gate leakage: 859 mW of static power before a single gate switches. The same slides give typical 65 nm values of IoffI_{\mathrm{off}} = 100, 10 and 1 nA/µm at VTV_{\mathrm{T}} = 0.3, 0.4 and 0.5 V, with S=100 mV/decadeS = 100\,\mathrm{mV/decade} and a DIBL coefficient η=0.1\eta = 0.1.

Near threshold, measured

  • Scaling from a nominal 1.1 V to 400–500 mV gives up to about 10× better energy efficiency, for roughly 10× lower performance. In an industrial 45 nm process the FO4 delay at 400 mV is 10× the delay at 1.1 V.
  • The minimum energy point is typically 250–350 mV, and it is shallow: going from near-threshold to subthreshold saves only about 2× more energy while delay rises 50–100×.
  • Calhoun, Wang and Chandrakasan measured the minimum at 250 mV for two filter circuits in 0.18 µm, with a ring oscillator slowing from 330 kHz at 300 mV to 29 kHz at 200 mV: an order of magnitude per 100 mV below threshold.

Every choice in this chapter gives something up:

  • Faster costs power. Raising the voltage makes switches quicker, but power climbs much faster than speed. That is part of why a phone gets warm when you play a game.
  • Saving power costs speed. Lower voltage saves energy, but very low voltage makes the chip slow.
  • Fast switches leak. Using fast, leaky switches everywhere would waste power all the time, even when the chip is idle.
  • More parts can save power. Two slow copies of a circuit can do the work of one fast copy with less power. But they take up more room, and room on a chip costs money.

Designers can’t plan for an average day. A chip must still work when it is hot, when the voltage dips, and when the factory made some switches a bit slow. So they check every mix of bad luck they can think of.

What you give up for what you get

ChoiceYou getYou give up
Higher VDDV_{\mathrm{DD}}Faster gates, a higher clockSwitching power rising with V2V^2, and more heat to remove
Lower VDDV_{\mathrm{DD}}Much less energy per operationSpeed, and robustness: near VTV_{\mathrm{T}}, small variations swing delay widely
Lower VTV_{\mathrm{T}} (LVT cells)Speed, especially at low VDDV_{\mathrm{DD}}Leakage, roughly tenfold per 60–100 mV
Parallel, slower hardwareSame throughput at lower VDDV_{\mathrm{DD}} and powerArea, and the cost of splitting the work
Power gating idle blocksNo leakage while offWake-up time, lost state, extra switch transistors

Race to idle, or slow and steady?

Suppose a task must be done within a deadline. You can run fast at high voltage and then switch the block off (“race to idle”), or run just fast enough at a lower voltage. If switching energy dominates, slow and steady wins, because energy per operation falls with V2V^2. If leakage dominates and the block can be power-gated when done, racing can win, because it shortens the time spent leaking. The right answer depends on the leakage share, and the leakage share changes with temperature and workload.

How it goes wrong

  • Leakage checked at the wrong temperature. Off-current rises as the chip heats up, so a leakage estimate made at room temperature can be badly low.
  • Too much low-VTV_{\mathrm{T}}. Fixing timing by swapping cells to LVT is easy, and it quietly raises standby power, so it pays to watch what share of cells are low-VTV_{\mathrm{T}}.
  • Supply droop. The voltage a gate actually sees is lower than the regulator’s, because of resistance in the power wiring (see Power planning). Lower voltage means slower gates, so timing must be checked at the drooped voltage. SKY130’s high-speed library, for example, ships timing data for 10% and 20% supply drop.
  • Ignoring heat feedback. Leakage heats the chip, and heat raises leakage. If cooling is marginal, that loop can push a chip past its thermal limit.

Energy, delay and the knee of the curve

Plot energy per operation against delay for every (VDD,VT,size)(V_{\mathrm{DD}}, V_{\mathrm{T}}, \text{size}) choice and the best designs form a convex Pareto front. At the fast end, a little more speed costs a lot of energy (high VDDV_{\mathrm{DD}}, low VTV_{\mathrm{T}}, upsized gates). At the slow end, below the minimum energy point, you lose both. Practical designs sit near the knee, and DVFS lets them slide along it. Metrics such as energy-delay product (E⋅tE \cdot t) pick a point near the knee; pure energy per operation picks the minimum energy point.

Variation near threshold

Above threshold, delay depends on VTV_{\mathrm{T}} through (VDD−VT)−α(V_{\mathrm{DD}} - V_{\mathrm{T}})^{-\alpha}, a mild dependence. Near and below threshold the dependence of drive current on VTV_{\mathrm{T}}, VDDV_{\mathrm{DD}} and temperature approaches exponential. Dreslinski et al. report that performance variation from global process variation alone grows from about 1.3× at nominal voltage to about 5× at 400 mV, and that temperature and supply ripple can each add another 2×, for a total uncertainty near 20× against about 1.5× at nominal. Local mismatch from random dopant fluctuation and line-edge roughness also threatens circuits that hold state by positive feedback, SRAM above all. Near-threshold designs therefore need variation-tolerant clocking and architecture and special SRAM, which eats into the energy gain.

Failure modes in signoff

  • Corner selection. Setup is checked at slow process and low voltage, leakage at fast process and high temperature (see ). In SKY130 the same standard inverter is 24.7 ps per stage at the fast corner and 44.1 ps at the slow one, a spread larger than any threshold flavor change.
  • Activity assumptions. Dynamic power is only as good as α\alpha. A default toggle rate can be far from what a real workload does, in either direction, so signoff power estimates use switching activity recorded from simulations of representative workloads.
  • Dynamic IR drop. Simultaneous switching draws current spikes that sag the local supply. A gate whose VDD−VTV_{\mathrm{DD}} - V_{\mathrm{T}} shrinks by a few tens of millivolts at 0.6 V loses far more speed than it would at 1.0 V, so low-voltage designs are more sensitive to supply noise.
  • Mode transitions. DVFS changes must lower frequency before lowering voltage, and raise voltage before raising frequency; getting the order wrong gives a brief window where the clock is faster than the gates.
deadlinepower-gatedRace to idlef, 1.0 VE 0.65Slow and steadyf/2, 0.7 VE 0.45 ✓switchingleakagetime →
Racer when done

Race to idle: 0.65 units (switching 0.50, leakage 0.15). Slow and steady: 0.45 (switching 0.25, leakage 0.21). Steady wins.

The same work done by a deadline: full clock at 1.0 V then idle, or half clock at 0.7 V throughout. Area = energy (blue switching, red leakage). Illustrative units; leakage taken as proportional to V_DD.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

Fitting the alpha-power law

The alpha-power model replaces the square law’s two discrepancies with short-channel measurements: the saturation-region current no longer grows as the square of the overdrive, and the drain saturation voltage is shifted from the square-law value. Both come from velocity saturation. Its parameters, including VTV_{\mathrm{T}} and α\alpha, are fitted to measured I-V curves. With it, Sakurai and Newton derived closed-form expressions for inverter delay that include the input transition time. The simple CV/ICV/I form used in this chapter is the step-input limit.

Two cautions when you use it. First, it is an above-threshold model: it predicts zero current at VGS=VTV_{\mathrm{GS}} = V_{\mathrm{T}}, so anything near threshold needs a subthreshold term. The simulation joins the two with a soft-plus function, I∝[sln⁡(1+e(V−VT)/s)]αI \propto \big[s \ln\big(1 + e^{(V - V_{\mathrm{T}})/s}\big)\big]^{\alpha} with s=αn kT/qs = \alpha n\,kT/q, which tends to (V−VT)α(V - V_{\mathrm{T}})^{\alpha} well above threshold and to an exponential with the right slope, one decade per n (kT/q)ln⁡10n\,(kT/q) \ln 10, well below it. Second, α\alpha itself drifts with VDDV_{\mathrm{DD}}, because velocity saturation weakens at low fields. Fits over a narrow DVFS range are good; extrapolation to near threshold is not.

The minimum energy point, derived

Clock a circuit at its maximum frequency for each VDDV_{\mathrm{DD}}. One operation (one clock cycle) costs

E(V)=αCV2+VIleak tcycle(V)tcycle=LDP tpd(V)\begin{gathered} E(V) = \alpha C V^2 + V I_{\mathrm{leak}}\, t_{\mathrm{cycle}}(V) \\ t_{\mathrm{cycle}} = L_{\mathrm{DP}}\, t_{\mathrm{pd}}(V) \end{gathered}

where LDPL_{\mathrm{DP}} is the logic depth. This is the same decomposition Calhoun, Wang and Chandrakasan use for subthreshold circuits: switching energy plus VDDIOFFTDV_{\mathrm{DD}} I_{\mathrm{OFF}} T_{\mathrm{D}}.

With tpd=kCgV/Ion(V)t_{\mathrm{pd}} = k C_{\mathrm{g}} V/I_{\mathrm{on}}(V) and the leakage of the whole block proportional to its size, the ratio of leakage to switching energy is

EleakEdyn∝LDPα⋅IoffIon(V)\frac{E_{\mathrm{leak}}}{E_{\mathrm{dyn}}} \propto \frac{L_{\mathrm{DP}}}{\alpha} \cdot \frac{I_{\mathrm{off}}}{I_{\mathrm{on}}(V)}

Each term in that ratio tells you something:

  • LDP/αL_{\mathrm{DP}}/\alpha: deep logic and low activity mean many idle gates leaking for a long cycle per useful transition. Leakage matters more, and the minimum moves to a higher voltage.
  • Ioff/Ion(V)I_{\mathrm{off}}/I_{\mathrm{on}}(V): above threshold this changes slowly with VV. Below threshold Ion(V)/Ioff≈eV/(n kT/q)I_{\mathrm{on}}(V)/I_{\mathrm{off}} \approx e^{V/(n\,kT/q)}, so the leakage share grows exponentially as VV falls while αCV2\alpha C V^2 shrinks only quadratically. Their sum has a minimum.

In the subthreshold limit, write x=V/(n kT/q)x = V/(n\,kT/q) and E=αCV2(1+Ae−x)E = \alpha C V^2 (1 + A e^{-x}), with AA collecting the constants above. Setting dE/dV=0dE/dV = 0 gives 2(1+Ae−x)=Axe−x2(1 + A e^{-x}) = A x e^{-x}, or ex=A(x−2)/2e^x = A(x - 2)/2. The optimum depends on AA only logarithmically, which fits the observation that minimum energy points land in a narrow band, typically 250–350 mV, across very different circuits. Its position relative to VTV_{\mathrm{T}} still shifts with activity and leakage.

The simulation’s numbers (VT=0.35 VV_{\mathrm{T}} = 0.35\,\mathrm{V}, 30-gate path) show the activity dependence directly:

α\alphaMinimum atEnergy saved vs 0.8 VSlowdown vs 0.8 V
0.020.50 V1.6×2.6×
0.10.40 V2.7×6.1×
0.50.32 V4.3×19×

These are properties of the toy model, not of any process. They do reproduce the qualitative findings of the near-threshold literature: savings of several times in energy per operation, paid for with roughly an order of magnitude in speed, and an optimum that sits near threshold.

Assigning thresholds: a sensitivity-driven swap

Choosing a VTV_{\mathrm{T}} flavor for each of millions of cells is a discrete optimization: minimize total leakage subject to every timing path meeting its constraint. Exact solutions are impractical at that scale, so implementation tools use incremental, sensitivity-driven heuristics that work alongside sizing and buffering. The general shape is:

  1. Start from a low-leakage assignment (mostly HVT or RVT) and run static timing analysis.
  2. For cells on violating paths, estimate each candidate swap’s benefit as delay gained per unit of leakage added, Δt/ΔPleak\Delta t/\Delta P_{\mathrm{leak}}, and apply the best ones.
  3. Re-time incrementally, since a faster cell also changes the slew and load seen by its neighbors.
  4. Run the reverse pass for power recovery: on paths with positive slack, swap cells to slower, lower-leakage flavors while the slack stays above a margin.
FFRVTA1 ×132 psRVTA2 ×1.857 psRVTA3 ×0.825 psRVTA4 ×1.548 psRVTA5 ×1.238 psFFA: 200.3 psslack -10.3 psFFRVTB1 ×0.825 psRVTB2 ×132 psRVTB3 ×0.929 psRVTB4 ×0.722 psFFB: 108.1 psslack +81.9 psLeakage90all LVT: 900HVTRVTLVT
1 / 5

1. Start low-leakage (here: every cell RVT) and run static timing. Path A misses the 190 ps clock. Leakage 90 units (all-LVT would be 900).

Sensitivity-driven V_T assignment on two paths with a 190 ps clock and a 10 ps recovery margin. Stage delays: SKY130 FO1 inverter-chain values (HVT-PMOS pair 38 ps, standard 31.8 ps, LVT-NMOS pair 28.6 ps) × load; leakage per cell 1 / 10 / 100, illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

OpenROAD exposes both directions: VTV_{\mathrm{T}} swapping during repair_timing (on by default) and a power-recovery option that works on a chosen percentage of paths. Because the swap preserves footprint and pins, it can run late, even after routing, which makes it one of the cheapest last-mile timing fixes available.

Novice · 0 of 4 correct
  1. Q1Using P=αCV2fP = \alpha C V^2 f, what happens to switching power if supply voltage drops from 1.0 V to 0.7 V and the clock is halved?

  2. Q2A process has a subthreshold slope of 100 mV per decade. Lowering the threshold voltage by 100 mV changes leakage how?

  3. Q3What mainly sets how long a CMOS gate takes to switch?

  4. Q4How does a design typically use a multi-threshold (multi-Vt) cell library?

Sources

Show Hide 20 sources
  1. Design of Ion-Implanted MOSFET’s with Very Small Physical DimensionsRobert H. Dennard, Fritz H. Gaensslen, Hwa-Nien Yu, V. Leo Rideout, Ernest Bassous, Andre R. LeBlanc · IEEE Journal of Solid-State Circuits SC-9(5), reprinted as Appendix D of “The Future of Computing Performance” (National Academies Press, 2011) · 1974Table I: dimensions and voltage scaled by 1/κ, doping by κ, give delay 1/κ, power per circuit 1/κ² and constant power density; the paper also warns that subthreshold behavior does not scale.
  2. Lecture 5: DC & Transient Response (CMOS VLSI Design, 4th ed.)David Harris · Harvey Mudd CollegeRC delay model with an effective transistor resistance (tpd ≈ RC); Elmore delay of RC ladders.
  3. A Simple MOSFET Model for Circuit Analysis and Its Application to CMOS Gate Delay Analysis and Series-Connected MOSFET StructureTakayasu Sakurai, A. Richard Newton · UC Berkeley EECS technical report UCB/ERL M90/19 · 1990Short-channel model with saturation current ∝ (VGS − VT)^n, reducing to the Shockley square law for n = 2; two discrepancies of the square law, both from velocity saturation; closed-form CMOS inverter delay with an input-transition term and τ = C0·VDD/ID0.
  4. Lecture 4: Nonideal Transistor Theory (CMOS VLSI Design, 4th ed.)David Harris · Harvey Mudd CollegeAlpha-power law with 1 < α < 2 (≈ 1.3 at 65 nm); DIBL; subthreshold slope ≈ 100 mV/decade with n of 1.3–1.7; gate leakage; ION falls and IOFF rises with temperature.
  5. Lecture 7: Power (CMOS VLSI Design, 4th ed.)David Harris · Harvey Mudd CollegeCV² energy per transition, P = αCV²f, activity of random data (0.25) and typical logic (≈ 0.1), short-circuit power under 10%, static power terms, stack effect, multiple Vt, gate leakage and high-k, power gating.
  6. A JSSC Classic Paper: Low-power CMOS Digital DesignAnantha P. Chandrakasan · IEEE Solid-State Circuits Society newsletter, on the author’s MIT group site · 2003Retrospective on the 1992 JSSC paper by Chandrakasan, Sheng and Brodersen: architectural voltage scaling uses hardware concurrency to make up for lower throughput at lower voltage, improving the power-delay product by an order of magnitude without performance loss, trading silicon area for lower power.
  7. EEC 216 Lecture #8: LeakageRajeevan Amirtharajah · University of California, Davis · 2008S = n·(kT/q)·ln 10: 60 mV/decade for n = 1, about 90 mV/decade for a typical n = 1.5; high-Vt devices off the critical path, low-Vt on it.
  8. The High-k SolutionMark T. Bohr, Robert S. Chau, Tahir Ghani, Kaizad Mistry · IEEE Spectrum · 2007SiO₂ gate insulators thinned to about 1.2 nm and leaked by tunneling; high-k plus metal gate transistors cut gate leakage by more than 10× and reached 45 nm microprocessors in 2007.
  9. Power Is Now Limiting Growth in Computing Performance (chapter 3 of The Future of Computing Performance)National Research Council · National Academies Press · 2011Leakage rises about 10× per 100 mV of Vth reduction; Vth stopped scaling and supply scaling slowed from roughly the 130 nm node; the optimum balances leakage at about 30% of dynamic power.
  10. Device Details (SkyWater SKY130 PDK documentation)SkyWater Technology · SKY130 open PDK documentation1.8 V NMOS and PMOS devices with low-VT and (PMOS) high-VT variants made by VT-adjust implants, with threshold voltages and inverter delays for each.
  11. SkyWater Foundry Provided Standard Cell LibrariesSkyWater Technology · SKY130 open PDK documentationSeven libraries (hs, ms, ls, lp, hd, hdll, hvl), the transistor flavors each uses, and hdll’s 5–10× lower leakage than hd.
  12. ASAP7: A 7-nm finFET predictive process design kitLawrence T. Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, Greg Yeric · Microelectronics Journal 53 (open access), copy in the ASAP7 PDK repository · 2016Nominal VDD 0.7 V; four threshold flavors (SLVT, LVT, RVT, SRAM) with per-fin drive and off-current; Dennard scaling ended for lack of room to scale Vth.
  13. Gate Resizer (rsz)The OpenROAD Project · OpenROAD documentationrepair_timing performs threshold-voltage swaps by default and offers power recovery; report_equiv_cells -vt lists HVT/RVT/LVT/SLVT equivalents.
  14. Scaling, Power and the Future of CMOS (talk slides)Mark Horowitz · 2006 Workshop on On- and Off-Chip Interconnection Networks for Multicore Systems (OCIN), workshop site · 2006Voltage scaling has stopped because kT/q does not scale and lowering Vth costs leakage power; at the optimum the energy/delay sensitivities of Vdd and Vth are equal to the slope of the Pareto curve, so neither is set directly by scaling; the optimal leakage share depends on activity factor.
  15. CPU DB: Recording Microprocessor HistoryAndrew Danowitz, Kyle Kelley, James Mao, John P. Stevenson, Mark Horowitz · ACM Queue · 2012Supply stayed at 5 V until about 0.6 µm, scaled roughly as the square root of feature size until about 0.13 µm and then slowed; power density rose more than 32× from the 80386 to 2005; by about 2005 processors hit the power wall.
  16. Dark Silicon and the End of Multicore ScalingHadi Esmaeilzadeh, Emily Blem, Renée St. Amant, Karthikeyan Sankaralingam, Doug Burger · ISCA 2011, author copy hosted by the University of Wisconsin–Madison · 2011With Dennard scaling failed and supply voltage scaling slowed, 21% of a fixed-size chip must be off at 22 nm and more than 50% at 8 nm.
  17. Lecture 15: Scaling & Economics (CMOS VLSI Design, 4th ed.)David Harris · Harvey Mudd CollegeDennard scaling table including switching energy 1/S³; tox and VDD scaling slowed around 65 nm because of gate tunneling and leakage.
  18. Near-Threshold Computing: Reclaiming Moore’s Law Through Energy Efficient Integrated CircuitsRonald G. Dreslinski, Michael Wieckowski, David Blaauw, Dennis Sylvester, Trevor Mudge · Proceedings of the IEEE 98(2), author copy on Trevor Mudge’s University of Michigan site · 2010Near-threshold operation gives about 10× energy savings for about 10× lower performance; the energy minimum (typically 250–350 mV) arises because leakage energy grows as delay rises; about 5× more performance variation.
  19. Device Sizing for Minimum Energy Operation in Subthreshold Circuits (CICC 2004 slides)Benton Calhoun, Alice Wang, Anantha Chandrakasan · MIT Energy-Efficient Circuits and Systems Group · 2004E_total = C·VDD² + VDD·I_OFF·T_D; delay rises exponentially below VT; measured minimum energy at 250 mV in 0.18 µm.