Transistors · Chapter 2 of 8 · The device

CMOS logic

Join two switches so that when one is on, the other is off. Now the output is always the opposite of the input: on becomes off, and off becomes on. Every logic gate on a chip is built from pairs like this.

An inverter stacks a PMOS pull-up above an NMOS pull-down. NAND and NOR gates extend the idea with switches in series and in parallel. Because one side is always off, a CMOS gate draws almost no current while it holds a value.

Complementary pull-up and pull-down networks, the voltage transfer curve and noise margins, why CMOS gates are naturally inverting, sizing PMOS against NMOS for balanced edges, and why stacking limits a gate’s inputs.

In the previous chapter a transistor was a tiny switch with no moving parts. Electricity turns it on or off. Now we use those switches to make decisions. Almost every chip made today does it with a trick called .

The trick is to use two kinds of switch that act in opposite ways. One kind turns on when its control wire is high, meaning it has electricity on it. The other kind turns on when its control wire is low.

Stack one of each. The top switch connects to the power wire. The bottom one connects to ground, the wire that electricity drains away to. The output comes from the middle. This circuit is an . It flips whatever you give it, as the picture below shows.

Every other is the same trick with more switches. This chapter shows how.

The switch introduced the MOSFET as a voltage-controlled switch. This chapter turns switches into , the circuits that compute 0s and 1s. Nearly every digital chip uses one circuit style for this: , complementary metal-oxide-semiconductor.

“Complementary” means each gate uses both kinds of transistor. An transistor conducts when its gate is high; a conducts when its gate is low. In a CMOS gate, NMOS transistors form a between the output and ground, and PMOS transistors form a between the supply and the output. For every input combination exactly one of the two networks conducts, so the output is always driven firmly to 0 or to 1.

That one rule has a big payoff. Because the two networks are never on together in a steady state, no current flows from the supply to ground while a gate holds its value. C. T. Sah and Frank Wanlass of Fairchild presented it in 1963 as logic that drew close to zero power on standby, and Wanlass patented it. Chips of the 1970s mostly used NMOS-only logic, which was cheaper to make but drew power even when idle; from the 1980s on, CMOS took over for its low idle power.

The rest of the chapter covers:

  • how the , NAND and NOR gates are wired, and why every single CMOS gate inverts;
  • the and , which explain how gates tolerate noise;
  • sizing PMOS against NMOS, and why gates rarely have more than four inputs.

Given a MOSFET as a switch with a threshold, an on-resistance and a gate capacitance (the switch), static builds every gate from two complementary networks: an NMOS (PDN) to GND and a PMOS (PUN) to VDDV_{\mathrm{DD}}, with the PUN the series/parallel dual of the PDN. Duality guarantees that for every input vector exactly one network conducts: no contention (both on, the “crowbar” state) and no floating output (both off). The properties designers rely on fall out of that:

  • Rail-to-rail levels. The conducting network ties the output all the way to VDDV_{\mathrm{DD}} or GND, so VOH=VDDV_{\mathrm{OH}} = V_{\mathrm{DD}} and VOL=0V_{\mathrm{OL}} = 0 with no static load.
  • Ideally zero static current. No DC path exists in either logic state. What is left is leakage, which is far from negligible at nanometer nodes.
  • High gain around the switching point, which makes logic levels regenerate stage after stage.
  • Inverting stages only, and 2n2n transistors for an nn-input gate.

Historically the alternative was ratioed NMOS logic. A resistor pull-up forces a trade between noise margin and speed, and a current-source pull-up that eases the trade still burns current whenever the output is low. CMOS, presented by Sah and Wanlass in 1963, won once processes caught up in performance, because it offered the best way to manage power density as integration grew. This chapter covers the DC side: topology, the transfer curve, noise margins, P/N sizing and fan-in. Delay and energy get their own chapter in Speed and power.

V_DDPMOSon when lowON1outNMOSon when highOFFGND0inIdle currentV_DD → GNDnone
Circuit style
Input

Input low: PMOS on, NMOS off. The output is tied to the supply (1), and with the NMOS off there is no path to ground.

The CMOS inverter against the NMOS-only inverter that came before it. Both invert; only CMOS has no supply-to-ground current while it holds a value.Share freely with credit: ‘Figure from chipfieldguide.com’

The inverter

An inverter is two switches in a stack: one on top joined to power, one on the bottom joined to ground. Whatever the input, one switch is on and the other is off. So the output is always the opposite of the input.

Two inputs: NAND

Now give the gate two inputs, A and B. Put two bottom switches in a row, one after the other. Electricity can only drain to ground if both are on. So the output drops to 0 only when A and B are both 1.

Put the two top switches side by side. Then either one alone can connect the output to power. This gate is a gate. Its output is 0 only when both inputs are 1.

The mirror image: NOR

Now swap the pattern. Put the bottom switches side by side and the top ones in a row. Either input can now pull the output down to 0. This is a gate. Its output is 1 only when both inputs are 0.

Why the answer always comes out flipped

Each kind of switch is good at only one job. The bottom kind is good at passing a 0, and the top kind is good at passing a 1. So the bottom switches always sit next to ground, and the top ones next to power.

A high input turns on a bottom switch, and that can only pull the output down. So every simple gate flips its answer. A plain AND gate gives a 1 only when both inputs are 1. To make one, a chip puts an inverter after a NAND.

The inverter

The CMOS is one PMOS from VDDV_{\mathrm{DD}} (the supply voltage) to the output and one NMOS from the output to GND (0 V), with both gates connected to the input.

  • Input 0 V: the NMOS sees no gate voltage and is off. The PMOS sees its full gate drive and is on. The output is connected to VDDV_{\mathrm{DD}}, a logic 1.
  • Input VDDV_{\mathrm{DD}}: the NMOS is on and the PMOS is off. The output is connected to GND, a logic 0.

In both states one transistor is off, so no current flows from VDDV_{\mathrm{DD}} to GND, and the output sits exactly at one of the supply rails. Engineers call this “rail-to-rail” output.

Series and parallel

Larger gates combine transistors in two ways.

  • In series (one after another): the path conducts only if all of them are on. A series chain of NMOS is an AND of their inputs.
  • In parallel (side by side): the path conducts if any of them is on. Parallel NMOS are an OR of their inputs.

A PMOS conducts when its input is 0, so the same shapes in the pull-up network act on inverted inputs.

NAND and NOR

A two-input puts two NMOS in series to ground and two PMOS in parallel to VDDV_{\mathrm{DD}}. The output is pulled low only when A and B are both 1; if either is 0, one of the parallel PMOS pulls it high. Its :

ABNMOS (series)PMOS (parallel)Y = NOT(A AND B)
00both off: openboth on: conducts1
01one off: openone on: conducts1
10one off: openone on: conducts1
11both on: conductsboth off: open0

A is the mirror image: NMOS in parallel, PMOS in series. Any input at 1 pulls the output low; only when both inputs are 0 do both series PMOS conduct and pull it high.

The duality rule

Notice that in every row of the table exactly one network conducts. That is no accident. The pull-up network is built as the dual of the pull-down network: every series connection becomes parallel and every parallel connection becomes series. If both networks could conduct at once, current would flow straight from VDDV_{\mathrm{DD}} to GND and the output would settle somewhere in between; if neither conducted, the output would float, connected to nothing. Duality rules out both.

Complex gates

Nothing limits a network to pure series or pure parallel. An AND-OR-INVERT gate, AOI21, computes NOT((A1 AND A2) OR B): A1 and A2 in series, that pair in parallel with B, and the dual (A1 and A2 in parallel, then in series with B) on top. The open SKY130 library’s a21oi cell is exactly this, six transistors in one stage. Such can implement any inverting function a series-parallel network can express.

Why single CMOS gates always invert

An NMOS can pull a node all the way down to 0 V, but it can only pull a node up to about VDDV_{\mathrm{DD}} minus its , a “degraded 1,” and it gets there slowly. A PMOS is the reverse: strong 1, degraded 0. So NMOS belong next to GND and PMOS next to VDDV_{\mathrm{DD}}. With that placement, a rising input can only turn on NMOS (pulling the output down) and turn off PMOS. Raising any input can never raise the output, so every one-stage CMOS gate is inverting. Non-inverting gates take two stages: the SKY130 and2 cell is a small NAND2 followed by an inverter, six transistors in total.

Networks and duality

A static CMOS gate computes Y=f(x)‾Y = \overline{f(x)}, where ff is the conduction function of the PDN: series composition is AND, parallel is OR, over uncomplemented inputs. The PUN, built from PMOS that conduct on 0, must conduct exactly when ff is false. By De Morgan, f(x)‾\overline{f(x)} evaluated on PMOS (which see x‾\overline{x}) is ff with AND and OR swapped, so the PUN is the series/parallel dual of the PDN. Consequences worth keeping in mind:

  • Two failure states are excluded by construction. Both networks on is contention (a “crowbar” DC path whose output level is set by a resistive divider); both off is a floating, high-Z output. Any non-dual network, for example from a schematic error or a missing transistor, creates one of these for some input vector.
  • Transistor count is 2n2n for an nn-input single-stage gate, one NMOS and one PMOS per input.
  • Series depth is split between the networks. NAND puts its depth in the NMOS stack, NOR in the PMOS stack, and a complex gate trades them off (AOI21: PDN depth 2, PUN depth 2). Depth sets sizing, as the sizing section shows.

Why a single stage must invert

An NMOS passing a high level turns itself off when its source reaches VDD−VTnV_{\mathrm{DD}} - V_{\mathrm{Tn}} (and later than that, given body effect), so it delivers a degraded 1 and approaches it slowly; a PMOS cannot pull below ∣VTp∣|V_{\mathrm{Tp}}|. Restricting NMOS to the PDN and PMOS to the PUN keeps every transistor in its strong direction, with its source on a rail or on a stack node near one. Under that restriction a 0→10 \to 1 input edge can only turn NMOS on and PMOS off, so each stage is a monotone-decreasing function of its inputs. Non-inverting functions cost a second stage. The SKY130 and2_1 is a 0.42 µm NAND2 driving a 0.65/1.0 µm inverter. That cost is why a technology mapper, choosing cells to cover a design’s logic, prefers inverting NAND, NOR and AOI/OAI cells where it can (see Logic synthesis).

Pass-transistor and transmission-gate styles break the rule on purpose. A transmission gate pairs an NMOS and a PMOS in parallel so the switch passes both levels fully. These styles are used for multiplexers and latches, but because their outputs are not driven from the rails, a chain of them does not restore levels the way a static CMOS gate does.

Complex gates in a real library

implement any series-parallel inverting function in one stage. Open libraries carry many of them. SKY130’s high-density library includes a21oi, a22oi, a211oi, a2111oi, o21ai, o22ai and a long list of other AOI/OAI variants alongside nand2–4 and nor2–4. Its a21oi_1 netlist shows the duality directly: NMOS A1–A2 in series and that pair in parallel with B1; PMOS A1∥A2\mathrm{A1} \parallel \mathrm{A2} in series with B1.

sky130_fd_sc_hd__a21oi_1.spice (bulk terminals as published)spice
.subckt sky130_fd_sc_hd__a21oi_1 A1 A2 B1 VGND VNB VPB VPWR Y
X0 a_199_47# A2 VGND VNB sky130_fd_pr__nfet_01v8 w=650000u l=150000u
X1 a_113_297# A1 VPWR VPB sky130_fd_pr__pfet_01v8_hvt w=1e+06u l=150000u
X2 Y B1 a_113_297# VPB sky130_fd_pr__pfet_01v8_hvt w=1e+06u l=150000u
X3 VGND B1 Y VNB sky130_fd_pr__nfet_01v8 w=650000u l=150000u
X4 VPWR A2 a_113_297# VPB sky130_fd_pr__pfet_01v8_hvt w=1e+06u l=150000u
X5 Y A1 a_199_47# VNB sky130_fd_pr__nfet_01v8 w=650000u l=150000u
.ends
  1. 1L1Pins: three logic inputs, output Y, plus ground, supply and the two body (bulk) connections.
  2. 2L2Bottom of the NMOS series pair: A2 from GND to the internal node a_199_47#. Device order is drain, gate, source, body.
  3. 3L3A1 and A2 PMOS (lines 3 and 6) both connect VPWR to node a_113_297#: they are in parallel.
  4. 4L4B1 PMOS connects that node to Y: the parallel pair is in series with B1, the dual of the pull-down.
  5. 5L5B1 NMOS alone from Y to GND: in parallel with the A1–A2 stack.
  6. 6L7Top of the NMOS series pair: A1 from the internal node to Y.
SeriesV_DD1A0Bconducts if A AND BParallelV_DD1A0Bconducts if A OR B
Transistor type
A
B

Series open, parallel conducts. With NMOS, series is AND and parallel is OR of the inputs.

Series and parallel switches. With NMOS, series is AND and parallel is OR. PMOS act on inverted inputs, so a gate’s pull-up uses the opposite shape: the dual.Share freely with credit: ‘Figure from chipfieldguide.com’
V_DD = 1.8 V1gateONoutput (load)V_out0.00.91.8wanted 1.8 Vstuck ≈ 1.16 V
Transistor
Pass a

NMOS passing a 1: weak. It turns itself off once the output is one threshold below V_DD, about 1.16 V, and gets there slowly.

A switch that is fully on, passing a 1 or a 0 to an output. Thresholds from SKY130’s 1.8 V transistors; the body effect, ignored here, makes the weak cases worse.Share freely with credit: ‘Figure from chipfieldguide.com’

Voltage is the push that moves electricity along a wire. A 1 is a strong push and a 0 is almost none. But real signals are never a perfect 0 or 1. Wires pick up small nudges from their neighbors, and the power supply wobbles. So what does a gate do with an input that is only “sort of high”?

Engineers draw a graph to find out. The input goes along the bottom and the output goes up the side. For an inverter, the line stays high, drops steeply in the middle, then stays low.

The flat parts are the key. An input a little off from 0 still gives an output very close to a perfect 1. So each gate cleans up the signal before passing it on, and errors shrink instead of growing. The is how big a nudge a signal can take before this stops working.

Reading the curve

The (VTC) plots a gate’s output voltage against its input voltage after everything has settled. Each point is where the current the PMOS can supply equals the current the NMOS can sink. Sweeping the input of an inverter from 0 to VDDV_{\mathrm{DD}}:

  1. Input below the NMOS threshold: the NMOS is off and the output sits at VDDV_{\mathrm{DD}}.
  2. Input a little higher: the NMOS starts to conduct, the output sags slightly below VDDV_{\mathrm{DD}}.
  3. Around the middle: both transistors conduct strongly and the output falls steeply. Here the gate has high gain: a small input change makes a large output change.
  4. Input higher still: the output approaches 0 V.
  5. Input above VDDV_{\mathrm{DD}} minus the PMOS threshold: the PMOS is off and the output sits at 0 V.

The input voltage where the output equals the input is the , VMV_{\mathrm{M}}. In a well-balanced inverter it is near VDD/2V_{\mathrm{DD}}/2.

Logic levels

Four voltages define what counts as a 0 or a 1.

  • VOLV_{\mathrm{OL}}: the highest voltage a gate outputs when it means 0.
  • VOHV_{\mathrm{OH}}: the lowest voltage a gate outputs when it means 1.
  • VILV_{\mathrm{IL}}: the highest input voltage a gate still reads as 0.
  • VIHV_{\mathrm{IH}}: the lowest input voltage a gate still reads as 1.

Inputs between VILV_{\mathrm{IL}} and VIHV_{\mathrm{IH}} are a no-man’s-land. The usual choice puts VILV_{\mathrm{IL}} and VIHV_{\mathrm{IH}} where the curve’s slope is exactly −1-1. Outside those points the gate shrinks disturbances; between them it amplifies them.

Noise margins

A is how much a valid output can be disturbed before the next gate might misread it:

  • NML=VIL−VOL\mathrm{NM}_{\mathrm{L}} = V_{\mathrm{IL}} - V_{\mathrm{OL}} for a 0.
  • NMH=VOH−VIH\mathrm{NM}_{\mathrm{H}} = V_{\mathrm{OH}} - V_{\mathrm{IH}} for a 1.

A small example. Suppose VDD=1.8 VV_{\mathrm{DD}} = 1.8\,\mathrm{V} and an inverter’s curve gives VOL=0.10 VV_{\mathrm{OL}} = 0.10\,\mathrm{V}, VIL=0.78 VV_{\mathrm{IL}} = 0.78\,\mathrm{V}, VIH=1.02 VV_{\mathrm{IH}} = 1.02\,\mathrm{V} and VOH=1.70 VV_{\mathrm{OH}} = 1.70\,\mathrm{V} (the values the simulator below produces for its balanced inverter). Then NML=0.78−0.10=0.68 V\mathrm{NM}_{\mathrm{L}} = 0.78 - 0.10 = 0.68\,\mathrm{V} and NMH=1.70−1.02=0.68 V\mathrm{NM}_{\mathrm{H}} = 1.70 - 1.02 = 0.68\,\mathrm{V}. A 0 can be pushed up by 0.68 V, or a 1 pulled down by 0.68 V, and the next gate still reads it correctly. An ideal inverter, with an infinitely steep middle, would reach VDD/2=0.9 VV_{\mathrm{DD}}/2 = 0.9\,\mathrm{V} on both sides.

Restoring logic

Because the gain is less than one near 0 and 1 and greater than one in the middle, a CMOS gate pulls a slightly wrong input toward a clean level. Send a degraded signal through a few inverters and it comes out at full 0 or VDDV_{\mathrm{DD}}. This property is what lets billions of gates be chained without noise piling up. It is also how two inverters connected in a loop hold a bit, the heart of the SRAM cell in Memory cells.

Regions of the inverter VTC

The DC transfer curve is the locus Idsn(Vin,Vout)=∣Idsp∣(Vin,Vout)I_{\mathrm{dsn}}(V_{\mathrm{in}}, V_{\mathrm{out}}) = |I_{\mathrm{dsp}}|(V_{\mathrm{in}}, V_{\mathrm{out}}), most easily found by load-line analysis: overlay the NMOS and PMOS output curves for each VinV_{\mathrm{in}} and read off the crossing. Classifying each device as off, linear or saturated gives five regions:

RegionVinV_{\mathrm{in}}NMOSPMOSVoutV_{\mathrm{out}}
A<VTn< V_{\mathrm{Tn}}cutofflinearVDDV_{\mathrm{DD}}
Bjust above VTnV_{\mathrm{Tn}}saturatedlinearslightly below VDDV_{\mathrm{DD}}
C≈VM\approx V_{\mathrm{M}}saturatedsaturatedsteep drop (slope set by output resistance)
Dabove VMV_{\mathrm{M}}linearsaturatedslightly above 0
E>VDD−∣VTp∣> V_{\mathrm{DD}} - |V_{\mathrm{Tp}}|linearcutoff0

In region C both devices are current sources, so the slope is limited only by their output resistance: Av=−(gmn+gmp)(ron∥rop)A_{\mathrm{v}} = -(g_{\mathrm{mn}} + g_{\mathrm{mp}})(r_{\mathrm{on}} \parallel r_{\mathrm{op}}), which can be large. In an ideal square-law device with no channel-length modulation the slope there is infinite; real short-channel devices have far lower output resistance, which makes region C less steep and eats into the noise margins (see The I-V curve for why).

Defining the levels

VILV_{\mathrm{IL}} and VIHV_{\mathrm{IH}} are conventionally taken at the two unity-gain points (slope −1-1), with VOH=Vout(VIL)V_{\mathrm{OH}} = V_{\mathrm{out}}(V_{\mathrm{IL}}) and VOL=Vout(VIH)V_{\mathrm{OL}} = V_{\mathrm{out}}(V_{\mathrm{IH}}). Choosing the levels there maximizes the noise margins. The justification is regeneration: where ∣slope∣<1|\text{slope}| < 1, a perturbation at the input shrinks at the output; where ∣slope∣>1|\text{slope}| > 1 it grows. A chain of stages therefore drives any input outside [VIL,VIH][V_{\mathrm{IL}}, V_{\mathrm{IH}}] to the rail. The simulator’s balanced inverter (square law, VTn=∣VTp∣=0.5 VV_{\mathrm{Tn}} = |V_{\mathrm{Tp}}| = 0.5\,\mathrm{V}, λ=0.1 V−1\lambda = 0.1\,\mathrm{V}^{-1}, VDD=1.8 VV_{\mathrm{DD}} = 1.8\,\mathrm{V}) gives VIL≈0.78 VV_{\mathrm{IL}} \approx 0.78\,\mathrm{V}, VIH≈1.02 VV_{\mathrm{IH}} \approx 1.02\,\mathrm{V} and NML=NMH≈0.68 V\mathrm{NM}_{\mathrm{L}} = \mathrm{NM}_{\mathrm{H}} \approx 0.68\,\mathrm{V}, about 38% of VDDV_{\mathrm{DD}}.

What the margins have to absorb

Static noise margins budget DC-like disturbances: IR drop and ground bounce that shift a driver’s rails relative to a receiver’s, and coupling from switching neighbors. Coupling from an aggressor wire onto a quiet victim is . The Cornell notes show that as feature sizes shrank, crosstalk glitches stronger than 80% could occur. A glitch that large exceeds any static margin. Whether it causes a failure then depends on how long it lasts and whether a flip-flop captures it, which a DC transfer curve cannot tell you.

Margins also shrink with VDDV_{\mathrm{DD}}. With VTnV_{\mathrm{Tn}} and ∣VTp∣|V_{\mathrm{Tp}}| roughly fixed, lowering the supply shrinks the region where both devices conduct and lowers the gain, so near-threshold logic trades margin for energy. The trade is quantified in Speed and power.

1.8 V0 V?0.60in1.79node 10.00node 21.80node 30.00outred: noise ΔV
First gate’s output
Noise on

Nudge 0.50 V, within the 0.68 V noise margin: every gate reads correctly, and each output is back near a rail.

Four inverters in a chain, from the chapter’s simulator model (illustrative). Bars show each gate’s input; the band is V_IL to V_IH. A nudge inside the 0.68 V noise margin always comes out right.Share freely with credit: ‘Figure from chipfieldguide.com’

Making both sides equally strong

The two kinds of switch are not equally strong. The top kind lets electricity through less easily than the bottom kind. If both were the same size, the output would drop fast but rise slowly. So designers make the top switch wider, like widening a doorway so a crowd moves as fast one way as the other.

Switches in a row are slow

Pushing electricity through several switches in a row is harder than through one. Think of water in a hose with several kinks. A gate with many inputs needs a long row of switches, so it is slow. Chips use gates with at most about four inputs, and build bigger jobs from several small gates.

Why PMOS is made wider

Current in an NMOS is carried by electrons; in a PMOS, by “holes,” missing electrons that behave like positive charges. Holes have lower : they drift more slowly for the same electric field. For the same size, a PMOS therefore conducts roughly half as much current as an NMOS, or even less.

A transistor’s current grows with its width (WW). To give the pull-up the same strength as the pull-down, the textbook inverter makes the PMOS about twice as wide as the NMOS: Wp≈2WnW_{\mathrm{p}} \approx 2W_{\mathrm{n}}. That puts the switching threshold at VDD/2V_{\mathrm{DD}}/2 and makes rising and falling edges take about the same time.

Sizing is a dial, not a fixed rule. Making the PMOS stronger moves VMV_{\mathrm{M}} up; making the NMOS stronger moves it down. A gate sized unevenly on purpose is a : one output edge gets faster and the other slower. VMV_{\mathrm{M}} moves surprisingly little: changing the width ratio tenfold shifts it by only a few hundred millivolts.

Stacks and fan-in

Transistors in series form a . An on transistor behaves roughly like a resistor, so two in series have about twice the resistance of one. To keep a NAND2’s pull-down as strong as an inverter’s, each series NMOS has to be twice as wide; a NAND3 needs them three times as wide. A NOR has the same problem in its PMOS, which were already twice as wide, so a NOR2 ends up with PMOS four times the width of a unit NMOS.

, the number of inputs, sets how tall the stacks get. Every extra input adds another series transistor, more width to make up for it and more capacitance for the inputs to drive. Delay grows faster than the input count, so real libraries stop around four inputs. The open SKY130 high-density library has NAND and NOR gates with two, three and four inputs, and nothing wider. A wider function becomes a small tree of gates, and choosing between many simple stages and fewer high-fan-in stages is a speed decision in its own right.

The same reasoning explains a common preference for NAND over NOR. A NAND’s stack is made of NMOS, the stronger device; a NOR’s stack is made of PMOS, the weaker one. For the same speed, a NOR needs more area and loads its inputs more heavily.

The RC view of sizing

Model an on transistor as a resistor inversely proportional to width. With a unit NMOS of resistance RR, a unit-width PMOS has about 2R2R, from the mobility ratio. The ratio μ\mu is 2–3 in classic planar processes. Sizing for equal drive then means: every path through the PDN totals RR and every path through the PUN totals RR. A device in a kk-high stack is therefore sized kk (NMOS) or kμk\mu (PMOS) units wide:

Gate (μ=2\mu = 2)NMOS widthsPMOS widthsInput cap per input
INV123
NAND22, 2 (series)2, 2 (parallel)4
NOR21, 1 (parallel)4, 4 (series)5
NAND33, 3, 32, 2, 25

Input capacitance relative to the inverter’s 3 is the : 4/3 for NAND2 and 5/3 for NOR2, and in general (n+2)/3(n+2)/3 for an nn-input NAND and (2n+1)/3(2n+1)/3 for NOR. Effort is the price of the function; the stack is where it comes from.

Balanced is not fastest

Balancing rise and fall is not the same as minimizing delay. For an inverter driving an identical inverter with PMOS width PP (NMOS = 1), the fall delay is proportional to P+1P + 1 and the rise delay to (P+1)μ/P(P + 1)\mu/P. Their average is minimized at P=μP = \sqrt{\mu}, not P=μP = \mu. The wider PMOS speeds its own edge but loads the previous stage with more capacitance. That is one reason production cells often use a P/N ratio below the mobility ratio. The switching threshold barely cares: VMV_{\mathrm{M}} depends on kp/kn\sqrt{k_{\mathrm{p}}/k_{\mathrm{n}}}, so even a 10× change in Wp/WnW_{\mathrm{p}}/W_{\mathrm{n}} moves it by well under VDD/2V_{\mathrm{DD}}/2. A deliberately unbalanced moves VMV_{\mathrm{M}} toward one rail to speed one edge, and pays with a smaller noise margin on the other side.

What stacks cost

  • Resistance and internal capacitance. Upsizing a stack restores its resistance but adds diffusion capacitance at every internal node, and the Elmore delay of the stack charges those nodes through the devices below them. For a NAND3 driving hh identical gates, summing the RC terms of Harris’s NAND3 example gives a falling delay of (12+5h)RC(12 + 5h)RC (the slide prints 11+5h11 + 5h) versus (9+5h)RC(9 + 5h)RC rising: the parasitic term is larger on the stacked side.
  • . Every device above the bottom of a stack has its source above the body while switching, which raises its threshold and lowers its drive.
  • Input capacitance. Logical effort grows linearly with nn, which slows the gates that drive this one.

Together these make delay grow faster than linearly in , so libraries cap NAND and NOR at about four inputs (SKY130 HD stops at nand4/nor4; the ASAP7 library goes to a weak NAND5 and NOR5). Wider functions become trees, and the choice between “many simple stages” and “fewer high-fan-in stages” is a delay-versus-area decision made per path. The cell-level view of these decisions, drive strengths and characterized delays, is in From devices to a cell library.

NAND2YPMOSeach ×2NMOSeach ×2not in kit123456fan-in →5101520delay (×τ) ↑NORNAND
Gate
Fan-in

NAND2: 2 NMOS in series, each 2× wide; PMOS in parallel at 2×. Each input drives 4 units of width, a logical effort of 4/3. Delay driving four copies ≈ 7.3, against 5 for an inverter.

Textbook sizing with PMOS twice as weak as NMOS (μ = 2): box widths are transistor widths. The chart is Harris’s logical-effort delay for driving four copies, in units of an ideal inverter’s delay.Share freely with credit: ‘Figure from chipfieldguide.com’

Pick a gate along the top. For most gates, click the inputs and watch which switches turn on (bright green). The output connects either to power (amber) or to ground (blue).

For the inverter (INV), drag the input from low to high and watch the dot follow the curve. The red hump shows electricity wasted straight from power to ground. It only appears while the input is in the middle.

The simulator draws each gate’s pull-up and pull-down networks. Toggle inputs and check the truth table: in every row exactly one network conducts, and the VDDV_{\mathrm{DD}}-to-GND current stays at zero. On the inverter, sweep VinV_{\mathrm{in}} and watch the operating point move along the transfer curve. The dashed lines mark VILV_{\mathrm{IL}}, VIHV_{\mathrm{IH}}, VOLV_{\mathrm{OL}} and VOHV_{\mathrm{OH}}; the panel reports both noise margins. Try parking the input near 0.9 V to see both transistors on at once and the short-circuit current peak.

Gates are generated from a series/parallel pull-down tree; the pull-up is its computed dual, and widths (×) come from series depth with μ=2\mu = 2, so the readouts give each input’s logical effort and the gate’s parasitic delay. Check NAND2 (4/3, 2), NOR2 (5/3, 2) and AOI21 (2 and 5/3, 7/3). On the inverter, the Wp/WnW_{\mathrm{p}}/W_{\mathrm{n}} slider skews the square-law VTC (VTn=∣VTp∣=0.5 VV_{\mathrm{Tn}} = |V_{\mathrm{Tp}}| = 0.5\,\mathrm{V}, λ=0.1 V−1\lambda = 0.1\,\mathrm{V}^{-1}, VDD=1.8 VV_{\mathrm{DD}} = 1.8\,\mathrm{V}). Sweep it from 0.25 to 8, a 32× strength range, and note how little VMV_{\mathrm{M}} moves while one noise margin grows at the other’s expense.

Loading simulation…
Switches in an inverter
2
Switches in a 2-input NAND
4
Most inputs on a gate in one free chip kit
4
Year CMOS was first shown
1963

What these numbers mean:

  • Each input needs two switches, one on top and one on the bottom. An inverter has one input and two switches. A two-input NAND has four.
  • A free chip-making kit called SKY130, used by schools and hobbyists, offers NAND and NOR gates with up to four inputs and no more. Bigger gates would be too slow.
  • CMOS was first shown in 1963 as logic that used almost no power while waiting. More than sixty years later, nearly all chips still work this way.
SKY130 inverter: NMOS / PMOS width
0.65 / 1.0 µm
SKY130 NMOS vs PMOS current, same size
≈ 2.6×
Logical effort: INV / NAND2 / NOR2
1 / 4/3 / 5/3
Largest NAND/NOR in SKY130 HD
4 inputs

Real transistor strengths. The open SKY130 process documents its 1.8 V transistors. At the same size (7 µm wide, 0.15 µm long), the typical NMOS conducts 3.51 mA when fully on, the standard PMOS 1.35 mA, and the high-threshold PMOS 1.00 mA. So the NMOS is about 2.6 times stronger than the standard PMOS, which is why PMOS are drawn wider.

Real cells. SKY130’s smallest inverter, inv_1, uses a 0.65 µm NMOS and a 1.0 µm high-threshold PMOS, both 0.15 µm long. Its two-input NAND keeps the same widths: two 1.0 µm PMOS in parallel and two 0.65 µm NMOS in series. The NOR2 has the mirror arrangement with the same widths. These cells trade some speed for small size; the library also offers larger versions of each.

Effort. In the logical-effort model, an inverter has effort 1, a NAND2 4/3 and a NOR2 5/3. A gate’s delay in units of a basic inverter delay is effort×fanout+a fixed “parasitic” part\text{effort} \times \text{fanout} + \text{a fixed “parasitic” part}; an inverter driving four copies of itself (the delay, a common yardstick) takes 5 units, about 15 ps in a 65 nm process.

GateTransistorsSeries NMOSSeries PMOSLogical effort per input
Inverter2111
NAND24214/3
NOR24125/3
NAND48412
NOR48143

Logical effort values from Harris’s lecture notes.

SKY130 Idsat at 7/0.15 µm: NMOS / PMOS / PMOS-hvt
3.51 / 1.35 / 1.00 mA
SKY130 inv_1 Wp/Wn
1.0 / 0.65 = 1.54
ASAP7 INVx1 fins (n / p)
3 / 3
65 nm example: static vs dynamic power
0.86 W vs 6.1 W

Device strengths in an open kit

SKY130’s device tables (typical corner, W/L=7/0.15 μmW/L = 7/0.15\,\mu\mathrm{m}) give the NMOS VT≈0.645 VV_{\mathrm{T}} \approx 0.645\,\mathrm{V} and Idsat=3.512 mAI_{\mathrm{dsat}} = 3.512\,\mathrm{mA}; the standard PMOS VT≈−0.781 VV_{\mathrm{T}} \approx -0.781\,\mathrm{V} and 1.347 mA; the high-VTV_{\mathrm{T}} PMOS VT≈−0.888 VV_{\mathrm{T}} \approx -0.888\,\mathrm{V} and 1.003 mA. The effective N/P drive ratio is 2.6 with the standard PMOS and 3.5 with the HVT one. The same page lists a fanout-of-1 inverter stage delay of 31.8 ps (nominal) for the standard NMOS/PMOS pair and 38 ps with the HVT PMOS: the slower pull-up is the price of lower PMOS leakage.

How the cells are actually sized

The library does not follow the textbook. inv_1 pairs a 0.65 µm NMOS with a 1.0 µm HVT PMOS, a width ratio of 1.54 against a drive ratio of about 3.5. By the RC model its pull-up is roughly half as strong as its pull-down, so VMV_{\mathrm{M}} sits below VDD/2V_{\mathrm{DD}}/2 and rising edges are slower than falling ones. Since 3.5≈1.9\sqrt{3.5} \approx 1.9, a ratio of 1.54 is in the range the least-average-delay argument favors over full balancing. Stacks are not upsized either: nand2_1 uses two series 0.65 µm NMOS, nor2_1 two series 1.0 µm PMOS, and nand4_1 four series 0.65 µm NMOS. At drive strength 1 the library minimizes area and input capacitance and lets characterization capture the weaker stacked edge. The same function also comes in larger drive strengths (nand2_2, nand2_4, nand2_8) for paths that need them.

The FinFET ASAP7 library makes different choices, in whole fins. Its INVx1 uses 3 fins for both NMOS and PMOS (P/N = 1). NAND2x1 doubles the series NMOS to 6 fins against 3-fin PMOS, and NOR2x1 doubles the series PMOS to 6 fins against 3-fin NMOS: textbook stack compensation. NAND5xp2 stacks five 3-fin NMOS under 2-fin PMOS.

asap7sc7p5t_28_R.cdl (excerpt)spice
.SUBCKT INVx1_ASAP7_75t_R A VDD VSS Y
MM0 Y A VSS VSS nmos_rvt w=81.0n l=20n nfin=3
MM1 Y A VDD VDD pmos_rvt w=81.0n l=20n nfin=3
.ENDS

.SUBCKT NAND2x1_ASAP7_75t_R A B VDD VSS Y
MM3 net16 A VSS VSS nmos_rvt w=162.00n l=20n nfin=6
MM2 Y B net16 VSS nmos_rvt w=162.00n l=20n nfin=6
MM1 Y B VDD VDD pmos_rvt w=81.0n l=20n nfin=3
MM0 Y A VDD VDD pmos_rvt w=81.0n l=20n nfin=3
.ENDS
  1. 1L2Width is quantized: nfin = 3 fins. The w value is the effective width the netlist reports.
  2. 2L3Same fin count for the PMOS: a 1:1 ratio, unlike SKY130’s planar inverter.
  3. 3L7Series NMOS doubled to 6 fins each, so the two-high stack matches one 3-fin device.
  4. 4L9Parallel PMOS stay at 3 fins: either one alone matches the inverter’s pull-up.

Static power is not zero

Harris’s worked example of a 1-billion-transistor chip in a 1.0 V, 65 nm process estimates 6.1 W of switching power at 1 GHz and 859 mW of static power: 584 mA of subthreshold leakage (assuming high-VTV_{\mathrm{T}} devices in all memory and 95% of logic) plus 275 mA of gate leakage. “Near-zero static current” is a statement about the circuit topology; the devices still leak, and the next chapter covers how much.

  • Two switches per input. CMOS needs twice as many switches as older designs that used one kind. In return, gates waste almost no power while sitting still. With billions of gates on a chip, that matters far more.
  • Only flipped answers. One gate can only give a flipped answer. A plain AND costs an extra inverter, which takes a little extra time.
  • Bigger gates are slower. More inputs mean longer rows of switches. Several small gates often beat one big one.
  • Power when switching. Gates only save power while they sit still. Every flip uses a little energy. That is the topic of Speed and power.
  • Tiny leaks. Modern switches never turn fully off, so a tiny bit of electricity always trickles through. Across billions of gates, it adds up.

What you give up for what you get

  • Area vs idle power. A CMOS gate needs one NMOS and one PMOS per input. The older NMOS-only style used a single pull-up device per gate, but whenever its output was low, current flowed continuously through that pull-up to ground. CMOS spends transistors to remove that current.
  • Inverting logic. Every single-stage gate inverts, so AND and OR take two stages. Designers and synthesis tools work around this by building logic from NAND, NOR and complex gates directly.
  • Fan-in vs depth. One four-input gate or a tree of two-input gates? The big gate has fewer stages but tall stacks; the tree has more stages but each one is quick.
  • Balanced vs skewed sizing. Making one edge faster with uneven sizing shrinks the noise margin on the other side.

Ways designs go wrong

  • A floating input. An input left unconnected can drift to a middle voltage. As the simulator shows, an inverter with its input near VDD/2V_{\mathrm{DD}}/2 has both transistors on and draws current continuously. Unused inputs are tied to VDDV_{\mathrm{DD}} or GND (through “tie” cells in a standard-cell flow).
  • Contention. If two gates drive the same wire to opposite values, or a network is wired wrong so both pull-up and pull-down conduct, current flows straight from VDDV_{\mathrm{DD}} to GND and the output sits at an undefined level.
  • A floating output. If neither network conducts, the output keeps whatever charge it had and slowly drifts. Correct dual networks never do this; mistakes and some special circuit styles can.
  • Slow input edges. The longer an input spends in the middle range, the longer both transistors conduct. With reasonably sharp edges this is under about a tenth of switching power; with sluggish edges it grows.
  • Too much noise. Crosstalk or supply droop larger than the noise margin flips a bit. On a chip, the error shows up as a wrong value captured somewhere downstream.

Static CMOS against the alternatives

  • Ratioed logic (NMOS with a resistive or always-on PMOS load) uses n+1n + 1 transistors and presents less input capacitance, but VOLV_{\mathrm{OL}} depends on the pull-down/load ratio and a DC current flows whenever the output is low. With a resistor load, noise margin and speed are in direct tension.
  • Pass-transistor and transmission-gate logic save devices in multiplexers and XORs but pass levels without restoring them, so chains need buffering. An NMOS-only pass device also degrades the 1 by VTnV_{\mathrm{Tn}}.
  • Static complementary CMOS pays 2n2n devices, the PMOS input load and inverting-only stages for rail-to-rail restoring outputs, no static current and robustness to sizing. That robustness is why it is the default for standard-cell logic.

Sizing and topology trade-offs

  • P/N ratio. Balanced (P=μP = \mu) gives VM=VDD/2V_{\mathrm{M}} = V_{\mathrm{DD}}/2 and equal edges; P=μP = \sqrt{\mu} gives less average delay and input capacitance; skewing favors one edge. Libraries choose per cell and per drive strength, as the SKY130 numbers show.
  • NAND vs NOR. With μ>1\mu > 1, a NOR’s PMOS stack makes it costlier than a NAND of the same fan-in (g=5/3g = 5/3 vs 4/34/3 at n=2n = 2, 3 vs 2 at n=4n = 4). Logic is restructured toward NANDs and AOIs where possible. The size of the penalty depends on μ\mu, which differs between processes; ASAP7’s 1:1 inverter shows a kit whose designers chose equal PMOS and NMOS fin counts.
  • Input ordering in stacks. The parasitic delay of a stacked gate depends on which input switches last. For a NAND2 falling edge, Harris’s estimate is 2τ2\tau if the input nearest the output (the “inner” input) arrives last, and 2.33τ2.33\tau if the input nearest the rail does, because then the internal node still has to be discharged. The rule: connect the latest-arriving signal to the inner input.
  • Leakage depends on the input vector. Series off devices leak about 10× less than a single off device (the stack effect), so a gate’s leakage depends on which inputs are 0. Sleep modes can park inputs in their lowest-leakage state.

Failure modes

  • Floating inputs put an inverter in region C, both devices on, with continuous current and an output that amplifies any noise on the input. Tie cells and input pull resistors exist for this.
  • Bus contention (two tristate drivers enabled at once) creates the crowbar state between gates rather than within one: a DC path, heating and an undefined level.
  • Margin erosion. Supply droop at the driver, ground bounce at the receiver and crosstalk stack up against the same NM\mathrm{NM}. At low VDDV_{\mathrm{DD}} the margins themselves shrink.
  • Short-circuit power stays under about 10% of dynamic power only while input and output slews are comparable. A weak driver on a long wire breaks that assumption at the receiver, which is one reason flows enforce max-transition limits.
V_DDPMOSONNMOSOFF0in1outV_DDI_DDV_DD → GND≈ 0
Case

Healthy, input 0: PMOS on, NMOS off. The output is a solid 1 and no current flows from V_DD to GND.

An inverter in its two healthy states and three faults: a floating input, contention (both networks on) and a floating output (both off). The meter is qualitative.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

Solving for the switching threshold

At VMV_{\mathrm{M}} both devices are saturated and carry the same current. In the long-channel square law, with kn=μnCoxWn/Lnk_{\mathrm{n}} = \mu_{\mathrm{n}} C_{\mathrm{ox}} W_{\mathrm{n}}/L_{\mathrm{n}} and kpk_{\mathrm{p}} likewise:

12kn(VM−VTn)2=12kp(VDD−VM−∣VTp∣)2\tfrac{1}{2} k_{\mathrm{n}} (V_{\mathrm{M}} - V_{\mathrm{Tn}})^2 = \tfrac{1}{2} k_{\mathrm{p}} (V_{\mathrm{DD}} - V_{\mathrm{M}} - |V_{\mathrm{Tp}}|)^2

Taking square roots and solving:

VM=VTn+r (VDD−∣VTp∣)1+rV_{\mathrm{M}} = \frac{V_{\mathrm{Tn}} + r\,(V_{\mathrm{DD}} - |V_{\mathrm{Tp}}|)}{1 + r}
r=kp/knr = \sqrt{k_{\mathrm{p}}/k_{\mathrm{n}}}
  • With r=1r = 1 and VTn=∣VTp∣V_{\mathrm{Tn}} = |V_{\mathrm{Tp}}|, VM=VDD/2V_{\mathrm{M}} = V_{\mathrm{DD}}/2. Since μn≈2μp\mu_{\mathrm{n}} \approx 2\mu_{\mathrm{p}}, that needs Wp≈2WnW_{\mathrm{p}} \approx 2W_{\mathrm{n}}.
  • As kn≫kpk_{\mathrm{n}} \gg k_{\mathrm{p}}, VM→VTnV_{\mathrm{M}} \to V_{\mathrm{Tn}}; as kp≫knk_{\mathrm{p}} \gg k_{\mathrm{n}}, VM→VDD−∣VTp∣V_{\mathrm{M}} \to V_{\mathrm{DD}} - |V_{\mathrm{Tp}}|.
  • The square root compresses sizing changes. With the simulator’s values (VDD=1.8 VV_{\mathrm{DD}} = 1.8\,\mathrm{V}, VT=0.5 VV_{\mathrm{T}} = 0.5\,\mathrm{V}), r=0.5r = 0.5 gives VM=0.77 VV_{\mathrm{M}} = 0.77\,\mathrm{V} and r=2r = 2 gives 1.03 V: a 16× change in kp/knk_{\mathrm{p}}/k_{\mathrm{n}} moves VMV_{\mathrm{M}} by 0.27 V. That insensitivity is why VMV_{\mathrm{M}} is robust to process variation in the P/N ratio.

Velocity-saturated devices change the exponent: current is closer to linear than quadratic in overdrive. Redo the algebra with I∝kVovI \propto k V_{\mathrm{ov}} and rr becomes kp/knk_{\mathrm{p}}/k_{\mathrm{n}} instead of its square root, so the same imbalance moves VMV_{\mathrm{M}} further from VDD/2V_{\mathrm{DD}}/2. The shape of the argument is the same. The model behind that is in The I-V curve.

Gain and two ways to estimate noise margins

Linearizing around VMV_{\mathrm{M}} gives the small-signal gain Av=−(gmn+gmp)(ron∥rop)A_{\mathrm{v}} = -(g_{\mathrm{mn}} + g_{\mathrm{mp}})(r_{\mathrm{on}} \parallel r_{\mathrm{op}}). A quick piecewise-linear estimate replaces the curve with three straight segments: flat at VDDV_{\mathrm{DD}}, a line of slope AvA_{\mathrm{v}} through (VM,VM)(V_{\mathrm{M}}, V_{\mathrm{M}}), flat at 0. Then:

VIL=VM−VDD−VM∣Av∣V_{\mathrm{IL}} = V_{\mathrm{M}} - \frac{V_{\mathrm{DD}} - V_{\mathrm{M}}}{|A_{\mathrm{v}}|}
VIH=VM(1+1∣Av∣)V_{\mathrm{IH}} = V_{\mathrm{M}} \left(1 + \frac{1}{|A_{\mathrm{v}}|}\right)
NML=VIL,NMH=VDD−VIH\mathrm{NM}_{\mathrm{L}} = V_{\mathrm{IL}}, \qquad \mathrm{NM}_{\mathrm{H}} = V_{\mathrm{DD}} - V_{\mathrm{IH}}

For the simulator’s balanced inverter, gm=kVov≈40 μA/Vg_{\mathrm{m}} = k V_{\mathrm{ov}} \approx 40\,\mu\mathrm{A/V} per device at Vov=0.4 VV_{\mathrm{ov}} = 0.4\,\mathrm{V} and I≈8.7 μAI \approx 8.7\,\mu\mathrm{A}, so ro≈1/(λI)≈1.1 MΩr_{\mathrm{o}} \approx 1/(\lambda I) \approx 1.1\,\mathrm{M}\Omega and ∣Av∣≈80 μA/V×0.57 MΩ≈46|A_{\mathrm{v}}| \approx 80\,\mu\mathrm{A/V} \times 0.57\,\mathrm{M}\Omega \approx 46 (the sim’s numerical derivative, which includes the (1+λVds)(1 + \lambda V_{\mathrm{ds}}) factor in the current, reports about 54). The piecewise estimate then gives NML≈NMH≈0.88 V\mathrm{NM}_{\mathrm{L}} \approx \mathrm{NM}_{\mathrm{H}} \approx 0.88\,\mathrm{V}. The unity-gain definition gives 0.68 V for the same curve. The piecewise version assumes the steep segment runs straight to the rails, ignoring the rounded knees of regions B and D, so it is optimistic. Use it for intuition about gain and the unity-gain numbers for budgets.

Sizing stacks and logical effort

Logical effort expresses a gate’s delay in units of τ\tau (the delay of an ideal inverter driving an identical one) as d=gh+pd = gh + p, where h=Cout/Cinh = C_{\mathrm{out}}/C_{\mathrm{in}} is the electrical effort. The procedure behind the gg values, which the simulator follows:

  1. Size every device so the worst-case path through each network matches a unit inverter (NMOS 1, PMOS μ\mu). A device on a path of kk series devices gets width kk (or kμk\mu). For nested series-parallel networks, kk is the longest series path through that device.
  2. g=(sum of widths that input drives)/(1+μ)g = (\text{sum of widths that input drives})/(1 + \mu) for each input. For an nn-input NAND with μ=2\mu = 2: (n+2)/3(n + 2)/3. For NOR: (2n+1)/3(2n + 1)/3.
  3. p≈(diffusion width on the output node)/(1+μ)p \approx (\text{diffusion width on the output node})/(1 + \mu), counting the devices that touch the output. That gives p=np = n for an nn-input NAND or NOR, and p=1p = 1 for the inverter. For AOI21 it gives gA=2g_{\mathrm{A}} = 2, gB=5/3g_{\mathrm{B}} = 5/3 and p=7/3p = 7/3.

The FO4 delay, d=1⋅4+1=5d = 1 \cdot 4 + 1 = 5, is the common yardstick; Harris quotes about 15 ps in a 65 nm process. Logical effort counts only output diffusion in pp, so it understates the cost of tall stacks. Elmore delay adds the internal nodes. In the NAND3 example (NMOS width 3, PMOS 2, driving hh identical NAND3s), summing Ri×CdownstreamR_i \times C_{\mathrm{downstream}} along the stack gives a falling delay of (12+5h)RC(12 + 5h)RC against (9+5h)RC(9 + 5h)RC rising. For an nn-high stack whose internal nodes each carry capacitance proportional to nn, the internal-node terms sum to something proportional to n(n−1)/2n(n - 1)/2. That quadratic growth is the formal version of “keep fan-in at four or below.”

The best P/N ratio

For an inverter with NMOS width 1 and PMOS width PP driving an identical inverter, the unit-RC delays are tpdf∝P+1t_{\mathrm{pdf}} \propto P + 1 and tpdr∝(P+1)μ/Pt_{\mathrm{pdr}} \propto (P + 1)\mu/P. Minimize the average:

tpd=(P+1)(1+μ/P)2t_{\mathrm{pd}} = \frac{(P + 1)(1 + \mu/P)}{2}
dtpddP=1−μ/P22=0  ⇒  P=μ\frac{dt_{\mathrm{pd}}}{dP} = \frac{1 - \mu/P^2}{2} = 0 \;\Rightarrow\; P = \sqrt{\mu}

With μ=2\mu = 2, P=1.41P = 1.41 gives tpd≈2.91t_{\mathrm{pd}} \approx 2.91 against 3.00 for P=2P = 2: about 3% faster on average and 20% less input capacitance, at the cost of unequal edges. Harris tabulates the same optimization for NAND2 and NOR2. Libraries land between the two answers, and real data (such as SKY130’s 1.54 ratio with a 3.5× drive gap) reflects the area, leakage and layout limits the RC model ignores.

24681234PMOS width P (NMOS = 1) →delay (RC units) ↑√μμfall ∝ P + 1rise ∝ (P + 1)μ/Paverage

P = 2.00, μ = 2.0: fall 3.00, rise 3.00, average 3.00 RC units, 2.9% above the minimum 2.91 at P = √μ = 1.41. Input capacitance 3.00 against 3.0 balanced.

Unit-RC delays of an inverter (NMOS 1, PMOS P) driving an identical one (Harris): fall ∝ P + 1, rise ∝ (P + 1)μ/P. The average bottoms out at √μ; balancing edges (P = μ) costs a little delay and more input capacitance. Preset: SKY130 inv_1 (1.54, with μ ≈ 3.5 against the HVT PMOS).Share freely with credit: ‘Figure from chipfieldguide.com’

Leakage and the stack effect

Subthreshold current falls by about a decade for every 100 mV of threshold: Harris’s 65 nm example gives Ioff≈100 nA/μmI_{\mathrm{off}} \approx 100\,\mathrm{nA}/\mu\mathrm{m} at VT=0.3 VV_{\mathrm{T}} = 0.3\,\mathrm{V}, 10 nA/µm at 0.4 V and 1 nA/µm at 0.5 V. In a two-high stack of off NMOS, the node between them settles at a small Vx>0V_x > 0. The upper device then sees Vgs=−VxV_{\mathrm{gs}} = -V_x, a body-effect VTV_{\mathrm{T}} increase and a smaller VdsV_{\mathrm{ds}} (less drain-induced barrier lowering), and its current drops about 10×; three-high stacks cut it further. Leakage-aware flows exploit this through high-VTV_{\mathrm{T}} cells, stacking and input-vector control in sleep. All of them trade against delay, the subject of Speed and power.

Novice · 0 of 4 correct
  1. Q1In a CMOS NOR2 gate, how are the transistors arranged?

  2. Q2An inverter’s transfer curve gives VOL=0.1 VV_{\mathrm{OL}} = 0.1\,\mathrm{V} and VIL=0.8 VV_{\mathrm{IL}} = 0.8\,\mathrm{V}. What is its low noise margin?

  3. Q3Why is the PMOS in a textbook CMOS inverter often drawn about twice as wide as the NMOS?

  4. Q4Where does a CMOS inverter draw significant current from VDDV_{\mathrm{DD}} to GND?

Sources

Show Hide 19 sources
  1. 1963: Complementary MOS Circuit Configuration is InventedComputer History Museum, The Silicon EngineSah and Wanlass at Fairchild presented complementary MOS logic that drew close to zero standby power; Wanlass’s patent was filed in 1963; CMOS later won because it managed power density as transistor counts grew.
  2. Lecture 1: Circuits & Layout (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd CollegenMOS-only processes drew power while idle; complementary pull-up and pull-down networks; crowbar and float states; series/parallel conduction rules; the conduction-complement rule; compound (AOI) gates.
  3. Lecture 12: Digital Circuits (II), MOS Inverter Circuits (6.012 Microelectronic Devices and Circuits)MIT OpenCourseWare · 2009Resistor and current-source pull-ups and their idle current; the CMOS inverter’s rail-to-rail levels and zero idle current; the switching-threshold equation; Wp ≈ 2Wn for a symmetric inverter; small-signal gain and noise-margin estimates.
  4. Lecture 5: DC & Transient Response (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd CollegeDegraded levels through pass transistors; the inverter DC transfer curve and its five operating regions; beta ratio and skewed gates; noise margins at the unity-gain points; the RC model (unit pMOS = 2R); Elmore delay of a 3-input NAND.
  5. ECE 4740 Lecture 4: The CMOS inverterCornell University, ECE Open Courseware · 2018Switching threshold and its weak sensitivity to Wp/Wn; noise sources such as crosstalk and supply noise; logic-level definitions at slope −1; the regenerative property; the ideal inverter’s VDD/2 noise margins.
  6. Lecture 6: Logical Effort (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd Colleged = gh + p; logical effort (n+2)/3 for an n-input NAND and (2n+1)/3 for NOR; parasitic delay n; FO4 delay of 5 units, about 15 ps in 65 nm.
  7. Lecture 9: Combinational Circuit Design (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd Collegeμ = 2–3 for an inverter; the P/N ratio for least average delay is √μ; choosing between many simple stages and fewer high-fan-in stages.
  8. Lecture 4: Nonideal Transistor Theory (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd CollegeBody effect: raising a transistor’s source voltage above its body raises its threshold voltage.
  9. Lecture 7: Power (CMOS VLSI Design, 4th ed. slides)David Harris · Harvey Mudd CollegeShort-circuit current under 10% of dynamic power; static power components; a 1-billion-transistor 65 nm example (6.1 W dynamic, 859 mW static); subthreshold leakage vs threshold; the ~10× stack effect.
  10. Device Details: 1.8V NMOS, PMOS and high-VT PMOS FETsSkyWater PDK Authors · SkyWater SKY130 PDK documentationTypical-corner threshold voltages and saturation currents at W/L = 7/0.15 µm, and nominal fanout-of-1 inverter delays, for the 1.8 V devices.
  11. sky130_fd_sc_hd__inv_1.spiceSkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)The inverter: one 0.65 µm NMOS and one 1.0 µm high-Vt PMOS, both 0.15 µm long.
  12. sky130_fd_sc_hd__nand2_1.spiceSkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)Two parallel 1.0 µm PMOS and two series 0.65 µm NMOS.
  13. sky130_fd_sc_hd nand2 cell directorySkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)NAND2 in drive strengths 1, 2, 4 and 8 (nand2_1 … nand2_8).
  14. sky130_fd_sc_hd__nor2_1.spiceSkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)Two series 1.0 µm PMOS and two parallel 0.65 µm NMOS.
  15. sky130_fd_sc_hd__nand4_1.spiceSkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)Four series 0.65 µm NMOS and four parallel 1.0 µm PMOS.
  16. sky130_fd_sc_hd__and2_1.spiceSkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)A small NAND2 (0.42 µm devices) followed by an inverter output stage: six transistors.
  17. sky130_fd_sc_hd__a21oi_1.spiceSkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)AND-OR-invert: NMOS A1–A2 in series, in parallel with B1; PMOS A1 ∥ A2 in series with B1.
  18. sky130_fd_sc_hd cell directorySkyWater PDK Authors · skywater-pdk-libs-sky130_fd_sc_hd (GitHub)The library’s cell list: NAND and NOR up to four inputs, plus many AOI/OAI complex gates.
  19. asap7sc7p5t_28_R.cdl (ASAP7 standard-cell netlists, regular Vt)Lawrence T. Clark and Vinay Vashishtha (Arizona State University) · asap7sc7p5t_28 (GitHub)FinFET cells sized in whole fins: INVx1 uses 3 fins for both NMOS and PMOS; NAND2x1 uses 6-fin NMOS and 3-fin PMOS; NOR2x1 the reverse; NAND5xp2 stacks five NMOS.