In the previous chapter a transistor was a tiny switch with no moving parts. Electricity turns it on or off. Now we use those switches to make decisions. Almost every chip made today does it with a trick called .
The trick is to use two kinds of switch that act in opposite ways. One kind turns on when its control wire is high, meaning it has electricity on it. The other kind turns on when its control wire is low.
Stack one of each. The top switch connects to the power wire. The bottom one connects to ground, the wire that electricity drains away to. The output comes from the middle. This circuit is an . It flips whatever you give it, as the picture below shows.
Every other is the same trick with more switches. This chapter shows how.
The switch introduced the MOSFET as a voltage-controlled switch. This chapter turns switches into , the circuits that compute 0s and 1s. Nearly every digital chip uses one circuit style for this: , complementary metal-oxide-semiconductor.
“Complementary” means each gate uses both kinds of transistor. An transistor conducts when its gate is high; a conducts when its gate is low. In a CMOS gate, NMOS transistors form a between the output and ground, and PMOS transistors form a between the supply and the output.2 For every input combination exactly one of the two networks conducts, so the output is always driven firmly to 0 or to 1.
That one rule has a big payoff. Because the two networks are never on together in a steady state, no current flows from the supply to ground while a gate holds its value. C. T. Sah and Frank Wanlass of Fairchild presented it in 1963 as logic that drew close to zero power on standby, and Wanlass patented it.1 Chips of the 1970s mostly used NMOS-only logic, which was cheaper to make but drew power even when idle; from the 1980s on, CMOS took over for its low idle power.2
The rest of the chapter covers:
- how the , NAND and NOR gates are wired, and why every single CMOS gate inverts;
- the and , which explain how gates tolerate noise;
- sizing PMOS against NMOS, and why gates rarely have more than four inputs.
Given a MOSFET as a switch with a threshold, an on-resistance and a gate capacitance (the switch), static builds every gate from two complementary networks: an NMOS (PDN) to GND and a PMOS (PUN) to , with the PUN the series/parallel dual of the PDN.2 Duality guarantees that for every input vector exactly one network conducts: no contention (both on, the “crowbar” state) and no floating output (both off). The properties designers rely on fall out of that:
- Rail-to-rail levels. The conducting network ties the output all the way to or GND, so and with no static load.3
- Ideally zero static current. No DC path exists in either logic state. What is left is leakage, which is far from negligible at nanometer nodes.9
- High gain around the switching point, which makes logic levels regenerate stage after stage.5
- Inverting stages only, and transistors for an -input gate.
Historically the alternative was ratioed NMOS logic. A resistor pull-up forces a trade between noise margin and speed, and a current-source pull-up that eases the trade still burns current whenever the output is low.3 CMOS, presented by Sah and Wanlass in 1963, won once processes caught up in performance, because it offered the best way to manage power density as integration grew.1 This chapter covers the DC side: topology, the transfer curve, noise margins, P/N sizing and fan-in. Delay and energy get their own chapter in Speed and power.
Input low: PMOS on, NMOS off. The output is tied to the supply (1), and with the NMOS off there is no path to ground.
The inverter
An inverter is two switches in a stack: one on top joined to power, one on the bottom joined to ground. Whatever the input, one switch is on and the other is off. So the output is always the opposite of the input.
Two inputs: NAND
Now give the gate two inputs, A and B. Put two bottom switches in a row, one after the other. Electricity can only drain to ground if both are on. So the output drops to 0 only when A and B are both 1.
Put the two top switches side by side. Then either one alone can connect the output to power. This gate is a gate. Its output is 0 only when both inputs are 1.
The mirror image: NOR
Now swap the pattern. Put the bottom switches side by side and the top ones in a row. Either input can now pull the output down to 0. This is a gate. Its output is 1 only when both inputs are 0.
Why the answer always comes out flipped
Each kind of switch is good at only one job. The bottom kind is good at passing a 0, and the top kind is good at passing a 1.4 So the bottom switches always sit next to ground, and the top ones next to power.
A high input turns on a bottom switch, and that can only pull the output down. So every simple gate flips its answer. A plain AND gate gives a 1 only when both inputs are 1. To make one, a chip puts an inverter after a NAND.16
The inverter
The CMOS is one PMOS from (the supply voltage) to the output and one NMOS from the output to GND (0 V), with both gates connected to the input.3
- Input 0 V: the NMOS sees no gate voltage and is off. The PMOS sees its full gate drive and is on. The output is connected to , a logic 1.
- Input : the NMOS is on and the PMOS is off. The output is connected to GND, a logic 0.
In both states one transistor is off, so no current flows from to GND, and the output sits exactly at one of the supply rails. Engineers call this “rail-to-rail” output.3
Series and parallel
Larger gates combine transistors in two ways.2
- In series (one after another): the path conducts only if all of them are on. A series chain of NMOS is an AND of their inputs.
- In parallel (side by side): the path conducts if any of them is on. Parallel NMOS are an OR of their inputs.
A PMOS conducts when its input is 0, so the same shapes in the pull-up network act on inverted inputs.
NAND and NOR
A two-input puts two NMOS in series to ground and two PMOS in parallel to . The output is pulled low only when A and B are both 1; if either is 0, one of the parallel PMOS pulls it high.2 Its :
| A | B | NMOS (series) | PMOS (parallel) | Y = NOT(A AND B) |
|---|---|---|---|---|
| 0 | 0 | both off: open | both on: conducts | 1 |
| 0 | 1 | one off: open | one on: conducts | 1 |
| 1 | 0 | one off: open | one on: conducts | 1 |
| 1 | 1 | both on: conducts | both off: open | 0 |
A is the mirror image: NMOS in parallel, PMOS in series. Any input at 1 pulls the output low; only when both inputs are 0 do both series PMOS conduct and pull it high.
The duality rule
Notice that in every row of the table exactly one network conducts. That is no accident. The pull-up network is built as the dual of the pull-down network: every series connection becomes parallel and every parallel connection becomes series.2 If both networks could conduct at once, current would flow straight from to GND and the output would settle somewhere in between; if neither conducted, the output would float, connected to nothing.2 Duality rules out both.
Complex gates
Nothing limits a network to pure series or pure parallel. An AND-OR-INVERT gate, AOI21, computes NOT((A1 AND A2) OR B): A1 and A2 in series, that pair in parallel with B, and the dual (A1 and A2 in parallel, then in series with B) on top. The open SKY130 library’s a21oi cell is exactly this, six transistors in one stage.17 Such can implement any inverting function a series-parallel network can express.2
Why single CMOS gates always invert
An NMOS can pull a node all the way down to 0 V, but it can only pull a node up to about minus its , a “degraded 1,” and it gets there slowly. A PMOS is the reverse: strong 1, degraded 0.4 So NMOS belong next to GND and PMOS next to . With that placement, a rising input can only turn on NMOS (pulling the output down) and turn off PMOS. Raising any input can never raise the output, so every one-stage CMOS gate is inverting. Non-inverting gates take two stages: the SKY130 and2 cell is a small NAND2 followed by an inverter, six transistors in total.16
Networks and duality
A static CMOS gate computes , where is the conduction function of the PDN: series composition is AND, parallel is OR, over uncomplemented inputs. The PUN, built from PMOS that conduct on 0, must conduct exactly when is false. By De Morgan, evaluated on PMOS (which see ) is with AND and OR swapped, so the PUN is the series/parallel dual of the PDN.2 Consequences worth keeping in mind:
- Two failure states are excluded by construction. Both networks on is contention (a “crowbar” DC path whose output level is set by a resistive divider); both off is a floating, high-Z output.2 Any non-dual network, for example from a schematic error or a missing transistor, creates one of these for some input vector.
- Transistor count is for an -input single-stage gate, one NMOS and one PMOS per input.
- Series depth is split between the networks. NAND puts its depth in the NMOS stack, NOR in the PMOS stack, and a complex gate trades them off (AOI21: PDN depth 2, PUN depth 2). Depth sets sizing, as the sizing section shows.
Why a single stage must invert
An NMOS passing a high level turns itself off when its source reaches (and later than that, given body effect), so it delivers a degraded 1 and approaches it slowly; a PMOS cannot pull below .4 Restricting NMOS to the PDN and PMOS to the PUN keeps every transistor in its strong direction, with its source on a rail or on a stack node near one. Under that restriction a input edge can only turn NMOS on and PMOS off, so each stage is a monotone-decreasing function of its inputs. Non-inverting functions cost a second stage. The SKY130 and2_1 is a 0.42 µm NAND2 driving a 0.65/1.0 µm inverter.16 That cost is why a technology mapper, choosing cells to cover a design’s logic, prefers inverting NAND, NOR and AOI/OAI cells where it can (see Logic synthesis).
Pass-transistor and transmission-gate styles break the rule on purpose. A transmission gate pairs an NMOS and a PMOS in parallel so the switch passes both levels fully.4 These styles are used for multiplexers and latches, but because their outputs are not driven from the rails, a chain of them does not restore levels the way a static CMOS gate does.
Complex gates in a real library
implement any series-parallel inverting function in one stage.2 Open libraries carry many of them. SKY130’s high-density library includes a21oi, a22oi, a211oi, a2111oi, o21ai, o22ai and a long list of other AOI/OAI variants alongside nand2–4 and nor2–4.18 Its a21oi_1 netlist shows the duality directly: NMOS A1–A2 in series and that pair in parallel with B1; PMOS in series with B1.17
.subckt sky130_fd_sc_hd__a21oi_1 A1 A2 B1 VGND VNB VPB VPWR Y
X0 a_199_47# A2 VGND VNB sky130_fd_pr__nfet_01v8 w=650000u l=150000u
X1 a_113_297# A1 VPWR VPB sky130_fd_pr__pfet_01v8_hvt w=1e+06u l=150000u
X2 Y B1 a_113_297# VPB sky130_fd_pr__pfet_01v8_hvt w=1e+06u l=150000u
X3 VGND B1 Y VNB sky130_fd_pr__nfet_01v8 w=650000u l=150000u
X4 VPWR A2 a_113_297# VPB sky130_fd_pr__pfet_01v8_hvt w=1e+06u l=150000u
X5 Y A1 a_199_47# VNB sky130_fd_pr__nfet_01v8 w=650000u l=150000u
.ends- 1L1Pins: three logic inputs, output Y, plus ground, supply and the two body (bulk) connections.
- 2L2Bottom of the NMOS series pair: A2 from GND to the internal node a_199_47#. Device order is drain, gate, source, body.
- 3L3A1 and A2 PMOS (lines 3 and 6) both connect VPWR to node a_113_297#: they are in parallel.
- 4L4B1 PMOS connects that node to Y: the parallel pair is in series with B1, the dual of the pull-down.
- 5L5B1 NMOS alone from Y to GND: in parallel with the A1–A2 stack.
- 6L7Top of the NMOS series pair: A1 from the internal node to Y.
Series open, parallel conducts. With NMOS, series is AND and parallel is OR of the inputs.
NMOS passing a 1: weak. It turns itself off once the output is one threshold below V_DD, about 1.16 V, and gets there slowly.
Voltage is the push that moves electricity along a wire. A 1 is a strong push and a 0 is almost none. But real signals are never a perfect 0 or 1. Wires pick up small nudges from their neighbors, and the power supply wobbles.5 So what does a gate do with an input that is only “sort of high”?
Engineers draw a graph to find out. The input goes along the bottom and the output goes up the side. For an inverter, the line stays high, drops steeply in the middle, then stays low.
The flat parts are the key. An input a little off from 0 still gives an output very close to a perfect 1. So each gate cleans up the signal before passing it on, and errors shrink instead of growing.5 The is how big a nudge a signal can take before this stops working.
Reading the curve
The (VTC) plots a gate’s output voltage against its input voltage after everything has settled. Each point is where the current the PMOS can supply equals the current the NMOS can sink.4 Sweeping the input of an inverter from 0 to :
- Input below the NMOS threshold: the NMOS is off and the output sits at .
- Input a little higher: the NMOS starts to conduct, the output sags slightly below .
- Around the middle: both transistors conduct strongly and the output falls steeply. Here the gate has high gain: a small input change makes a large output change.3
- Input higher still: the output approaches 0 V.
- Input above minus the PMOS threshold: the PMOS is off and the output sits at 0 V.
The input voltage where the output equals the input is the , . In a well-balanced inverter it is near .3
Logic levels
Four voltages define what counts as a 0 or a 1.4
- : the highest voltage a gate outputs when it means 0.
- : the lowest voltage a gate outputs when it means 1.
- : the highest input voltage a gate still reads as 0.
- : the lowest input voltage a gate still reads as 1.
Inputs between and are a no-man’s-land. The usual choice puts and where the curve’s slope is exactly . Outside those points the gate shrinks disturbances; between them it amplifies them.45
Noise margins
A is how much a valid output can be disturbed before the next gate might misread it:
- for a 0.
- for a 1.5
A small example. Suppose and an inverter’s curve gives , , and (the values the simulator below produces for its balanced inverter). Then and . A 0 can be pushed up by 0.68 V, or a 1 pulled down by 0.68 V, and the next gate still reads it correctly. An ideal inverter, with an infinitely steep middle, would reach on both sides.5
Restoring logic
Because the gain is less than one near 0 and 1 and greater than one in the middle, a CMOS gate pulls a slightly wrong input toward a clean level. Send a degraded signal through a few inverters and it comes out at full 0 or .5 This property is what lets billions of gates be chained without noise piling up. It is also how two inverters connected in a loop hold a bit, the heart of the SRAM cell in Memory cells.
Regions of the inverter VTC
The DC transfer curve is the locus , most easily found by load-line analysis: overlay the NMOS and PMOS output curves for each and read off the crossing.4 Classifying each device as off, linear or saturated gives five regions:43
| Region | NMOS | PMOS | ||
|---|---|---|---|---|
| A | cutoff | linear | ||
| B | just above | saturated | linear | slightly below |
| C | saturated | saturated | steep drop (slope set by output resistance) | |
| D | above | linear | saturated | slightly above 0 |
| E | linear | cutoff | 0 |
In region C both devices are current sources, so the slope is limited only by their output resistance: , which can be large.3 In an ideal square-law device with no channel-length modulation the slope there is infinite; real short-channel devices have far lower output resistance, which makes region C less steep and eats into the noise margins (see The I-V curve for why).
Defining the levels
and are conventionally taken at the two unity-gain points (slope ), with and . Choosing the levels there maximizes the noise margins.4 The justification is regeneration: where , a perturbation at the input shrinks at the output; where it grows. A chain of stages therefore drives any input outside to the rail.5 The simulator’s balanced inverter (square law, , , ) gives , and , about 38% of .
What the margins have to absorb
Static noise margins budget DC-like disturbances: IR drop and ground bounce that shift a driver’s rails relative to a receiver’s, and coupling from switching neighbors.5 Coupling from an aggressor wire onto a quiet victim is . The Cornell notes show that as feature sizes shrank, crosstalk glitches stronger than 80% could occur.5 A glitch that large exceeds any static margin. Whether it causes a failure then depends on how long it lasts and whether a flip-flop captures it, which a DC transfer curve cannot tell you.
Margins also shrink with . With and roughly fixed, lowering the supply shrinks the region where both devices conduct and lowers the gain, so near-threshold logic trades margin for energy. The trade is quantified in Speed and power.
Nudge 0.50 V, within the 0.68 V noise margin: every gate reads correctly, and each output is back near a rail.
Making both sides equally strong
The two kinds of switch are not equally strong. The top kind lets electricity through less easily than the bottom kind.3 If both were the same size, the output would drop fast but rise slowly. So designers make the top switch wider, like widening a doorway so a crowd moves as fast one way as the other.
Switches in a row are slow
Pushing electricity through several switches in a row is harder than through one. Think of water in a hose with several kinks. A gate with many inputs needs a long row of switches, so it is slow. Chips use gates with at most about four inputs, and build bigger jobs from several small gates.18
Why PMOS is made wider
Current in an NMOS is carried by electrons; in a PMOS, by “holes,” missing electrons that behave like positive charges. Holes have lower : they drift more slowly for the same electric field. For the same size, a PMOS therefore conducts roughly half as much current as an NMOS, or even less.37
A transistor’s current grows with its width (). To give the pull-up the same strength as the pull-down, the textbook inverter makes the PMOS about twice as wide as the NMOS: . That puts the switching threshold at and makes rising and falling edges take about the same time.34
Sizing is a dial, not a fixed rule. Making the PMOS stronger moves up; making the NMOS stronger moves it down. A gate sized unevenly on purpose is a : one output edge gets faster and the other slower.4 moves surprisingly little: changing the width ratio tenfold shifts it by only a few hundred millivolts.5
Stacks and fan-in
Transistors in series form a . An on transistor behaves roughly like a resistor, so two in series have about twice the resistance of one.4 To keep a NAND2’s pull-down as strong as an inverter’s, each series NMOS has to be twice as wide; a NAND3 needs them three times as wide. A NOR has the same problem in its PMOS, which were already twice as wide, so a NOR2 ends up with PMOS four times the width of a unit NMOS.
, the number of inputs, sets how tall the stacks get. Every extra input adds another series transistor, more width to make up for it and more capacitance for the inputs to drive. Delay grows faster than the input count, so real libraries stop around four inputs. The open SKY130 high-density library has NAND and NOR gates with two, three and four inputs, and nothing wider.18 A wider function becomes a small tree of gates, and choosing between many simple stages and fewer high-fan-in stages is a speed decision in its own right.7
The same reasoning explains a common preference for NAND over NOR. A NAND’s stack is made of NMOS, the stronger device; a NOR’s stack is made of PMOS, the weaker one. For the same speed, a NOR needs more area and loads its inputs more heavily.6
The RC view of sizing
Model an on transistor as a resistor inversely proportional to width. With a unit NMOS of resistance , a unit-width PMOS has about , from the mobility ratio.4 The ratio is 2–3 in classic planar processes.7 Sizing for equal drive then means: every path through the PDN totals and every path through the PUN totals . A device in a -high stack is therefore sized (NMOS) or (PMOS) units wide:
| Gate () | NMOS widths | PMOS widths | Input cap per input |
|---|---|---|---|
| INV | 1 | 2 | 3 |
| NAND2 | 2, 2 (series) | 2, 2 (parallel) | 4 |
| NOR2 | 1, 1 (parallel) | 4, 4 (series) | 5 |
| NAND3 | 3, 3, 3 | 2, 2, 2 | 5 |
Input capacitance relative to the inverter’s 3 is the : 4/3 for NAND2 and 5/3 for NOR2, and in general for an -input NAND and for NOR.6 Effort is the price of the function; the stack is where it comes from.
Balanced is not fastest
Balancing rise and fall is not the same as minimizing delay. For an inverter driving an identical inverter with PMOS width (NMOS = 1), the fall delay is proportional to and the rise delay to . Their average is minimized at , not .7 The wider PMOS speeds its own edge but loads the previous stage with more capacitance. That is one reason production cells often use a P/N ratio below the mobility ratio. The switching threshold barely cares: depends on , so even a 10× change in moves it by well under .5 A deliberately unbalanced moves toward one rail to speed one edge, and pays with a smaller noise margin on the other side.4
What stacks cost
- Resistance and internal capacitance. Upsizing a stack restores its resistance but adds diffusion capacitance at every internal node, and the Elmore delay of the stack charges those nodes through the devices below them. For a NAND3 driving identical gates, summing the RC terms of Harris’s NAND3 example gives a falling delay of (the slide prints ) versus rising: the parasitic term is larger on the stacked side.4
- . Every device above the bottom of a stack has its source above the body while switching, which raises its threshold and lowers its drive.8
- Input capacitance. Logical effort grows linearly with , which slows the gates that drive this one.6
Together these make delay grow faster than linearly in , so libraries cap NAND and NOR at about four inputs (SKY130 HD stops at nand4/nor4; the ASAP7 library goes to a weak NAND5 and NOR5).1819 Wider functions become trees, and the choice between “many simple stages” and “fewer high-fan-in stages” is a delay-versus-area decision made per path.7 The cell-level view of these decisions, drive strengths and characterized delays, is in From devices to a cell library.
NAND2: 2 NMOS in series, each 2× wide; PMOS in parallel at 2×. Each input drives 4 units of width, a logical effort of 4/3. Delay driving four copies ≈ 7.3, against 5 for an inverter.
Pick a gate along the top. For most gates, click the inputs and watch which switches turn on (bright green). The output connects either to power (amber) or to ground (blue).
For the inverter (INV), drag the input from low to high and watch the dot follow the curve. The red hump shows electricity wasted straight from power to ground. It only appears while the input is in the middle.
The simulator draws each gate’s pull-up and pull-down networks. Toggle inputs and check the truth table: in every row exactly one network conducts, and the -to-GND current stays at zero. On the inverter, sweep and watch the operating point move along the transfer curve. The dashed lines mark , , and ; the panel reports both noise margins. Try parking the input near 0.9 V to see both transistors on at once and the short-circuit current peak.
Gates are generated from a series/parallel pull-down tree; the pull-up is its computed dual, and widths (×) come from series depth with , so the readouts give each input’s logical effort and the gate’s parasitic delay. Check NAND2 (4/3, 2), NOR2 (5/3, 2) and AOI21 (2 and 5/3, 7/3). On the inverter, the slider skews the square-law VTC (, , ). Sweep it from 0.25 to 8, a 32× strength range, and note how little moves while one noise margin grows at the other’s expense.
- Switches in an inverter
- 2
- Switches in a 2-input NAND
- 4
- Most inputs on a gate in one free chip kit
- 4
- Year CMOS was first shown
- 1963
What these numbers mean:
- Each input needs two switches, one on top and one on the bottom. An inverter has one input and two switches. A two-input NAND has four.
- A free chip-making kit called SKY130, used by schools and hobbyists, offers NAND and NOR gates with up to four inputs and no more.18 Bigger gates would be too slow.
- CMOS was first shown in 1963 as logic that used almost no power while waiting.1 More than sixty years later, nearly all chips still work this way.
- SKY130 inverter: NMOS / PMOS width
- 0.65 / 1.0 µm
- SKY130 NMOS vs PMOS current, same size
- ≈ 2.6×
- Logical effort: INV / NAND2 / NOR2
- 1 / 4/3 / 5/3
- Largest NAND/NOR in SKY130 HD
- 4 inputs
Real transistor strengths. The open SKY130 process documents its 1.8 V transistors. At the same size (7 µm wide, 0.15 µm long), the typical NMOS conducts 3.51 mA when fully on, the standard PMOS 1.35 mA, and the high-threshold PMOS 1.00 mA.10 So the NMOS is about 2.6 times stronger than the standard PMOS, which is why PMOS are drawn wider.
Real cells. SKY130’s smallest inverter, inv_1, uses a 0.65 µm NMOS and a 1.0 µm high-threshold PMOS, both 0.15 µm long.11 Its two-input NAND keeps the same widths: two 1.0 µm PMOS in parallel and two 0.65 µm NMOS in series.12 The NOR2 has the mirror arrangement with the same widths.14 These cells trade some speed for small size; the library also offers larger versions of each.
Effort. In the logical-effort model, an inverter has effort 1, a NAND2 4/3 and a NOR2 5/3. A gate’s delay in units of a basic inverter delay is ; an inverter driving four copies of itself (the delay, a common yardstick) takes 5 units, about 15 ps in a 65 nm process.6
| Gate | Transistors | Series NMOS | Series PMOS | Logical effort per input |
|---|---|---|---|---|
| Inverter | 2 | 1 | 1 | 1 |
| NAND2 | 4 | 2 | 1 | 4/3 |
| NOR2 | 4 | 1 | 2 | 5/3 |
| NAND4 | 8 | 4 | 1 | 2 |
| NOR4 | 8 | 1 | 4 | 3 |
Logical effort values from Harris’s lecture notes.6
- SKY130 Idsat at 7/0.15 µm: NMOS / PMOS / PMOS-hvt
- 3.51 / 1.35 / 1.00 mA
- SKY130 inv_1 Wp/Wn
- 1.0 / 0.65 = 1.54
- ASAP7 INVx1 fins (n / p)
- 3 / 3
- 65 nm example: static vs dynamic power
- 0.86 W vs 6.1 W
Device strengths in an open kit
SKY130’s device tables (typical corner, ) give the NMOS and ; the standard PMOS and 1.347 mA; the high- PMOS and 1.003 mA.10 The effective N/P drive ratio is 2.6 with the standard PMOS and 3.5 with the HVT one. The same page lists a fanout-of-1 inverter stage delay of 31.8 ps (nominal) for the standard NMOS/PMOS pair and 38 ps with the HVT PMOS: the slower pull-up is the price of lower PMOS leakage.10
How the cells are actually sized
The library does not follow the textbook. inv_1 pairs a 0.65 µm NMOS with a 1.0 µm HVT PMOS, a width ratio of 1.54 against a drive ratio of about 3.5.1110 By the RC model its pull-up is roughly half as strong as its pull-down, so sits below and rising edges are slower than falling ones. Since , a ratio of 1.54 is in the range the least-average-delay argument favors over full balancing.7 Stacks are not upsized either: nand2_1 uses two series 0.65 µm NMOS, nor2_1 two series 1.0 µm PMOS, and nand4_1 four series 0.65 µm NMOS.121415 At drive strength 1 the library minimizes area and input capacitance and lets characterization capture the weaker stacked edge. The same function also comes in larger drive strengths (nand2_2, nand2_4, nand2_8) for paths that need them.13
The FinFET ASAP7 library makes different choices, in whole fins. Its INVx1 uses 3 fins for both NMOS and PMOS (P/N = 1). NAND2x1 doubles the series NMOS to 6 fins against 3-fin PMOS, and NOR2x1 doubles the series PMOS to 6 fins against 3-fin NMOS: textbook stack compensation. NAND5xp2 stacks five 3-fin NMOS under 2-fin PMOS.19
.SUBCKT INVx1_ASAP7_75t_R A VDD VSS Y
MM0 Y A VSS VSS nmos_rvt w=81.0n l=20n nfin=3
MM1 Y A VDD VDD pmos_rvt w=81.0n l=20n nfin=3
.ENDS
.SUBCKT NAND2x1_ASAP7_75t_R A B VDD VSS Y
MM3 net16 A VSS VSS nmos_rvt w=162.00n l=20n nfin=6
MM2 Y B net16 VSS nmos_rvt w=162.00n l=20n nfin=6
MM1 Y B VDD VDD pmos_rvt w=81.0n l=20n nfin=3
MM0 Y A VDD VDD pmos_rvt w=81.0n l=20n nfin=3
.ENDS- 1L2Width is quantized: nfin = 3 fins. The w value is the effective width the netlist reports.
- 2L3Same fin count for the PMOS: a 1:1 ratio, unlike SKY130’s planar inverter.
- 3L7Series NMOS doubled to 6 fins each, so the two-high stack matches one 3-fin device.
- 4L9Parallel PMOS stay at 3 fins: either one alone matches the inverter’s pull-up.
Static power is not zero
Harris’s worked example of a 1-billion-transistor chip in a 1.0 V, 65 nm process estimates 6.1 W of switching power at 1 GHz and 859 mW of static power: 584 mA of subthreshold leakage (assuming high- devices in all memory and 95% of logic) plus 275 mA of gate leakage.9 “Near-zero static current” is a statement about the circuit topology; the devices still leak, and the next chapter covers how much.
- Two switches per input. CMOS needs twice as many switches as older designs that used one kind. In return, gates waste almost no power while sitting still. With billions of gates on a chip, that matters far more.1
- Only flipped answers. One gate can only give a flipped answer. A plain AND costs an extra inverter, which takes a little extra time.16
- Bigger gates are slower. More inputs mean longer rows of switches. Several small gates often beat one big one.
- Power when switching. Gates only save power while they sit still. Every flip uses a little energy. That is the topic of Speed and power.
- Tiny leaks. Modern switches never turn fully off, so a tiny bit of electricity always trickles through. Across billions of gates, it adds up.9
What you give up for what you get
- Area vs idle power. A CMOS gate needs one NMOS and one PMOS per input. The older NMOS-only style used a single pull-up device per gate, but whenever its output was low, current flowed continuously through that pull-up to ground.3 CMOS spends transistors to remove that current.
- Inverting logic. Every single-stage gate inverts, so AND and OR take two stages. Designers and synthesis tools work around this by building logic from NAND, NOR and complex gates directly.
- Fan-in vs depth. One four-input gate or a tree of two-input gates? The big gate has fewer stages but tall stacks; the tree has more stages but each one is quick.7
- Balanced vs skewed sizing. Making one edge faster with uneven sizing shrinks the noise margin on the other side.4
Ways designs go wrong
- A floating input. An input left unconnected can drift to a middle voltage. As the simulator shows, an inverter with its input near has both transistors on and draws current continuously. Unused inputs are tied to or GND (through “tie” cells in a standard-cell flow).
- Contention. If two gates drive the same wire to opposite values, or a network is wired wrong so both pull-up and pull-down conduct, current flows straight from to GND and the output sits at an undefined level.2
- A floating output. If neither network conducts, the output keeps whatever charge it had and slowly drifts. Correct dual networks never do this; mistakes and some special circuit styles can.2
- Slow input edges. The longer an input spends in the middle range, the longer both transistors conduct. With reasonably sharp edges this is under about a tenth of switching power; with sluggish edges it grows.9
- Too much noise. Crosstalk or supply droop larger than the noise margin flips a bit. On a chip, the error shows up as a wrong value captured somewhere downstream.5
Static CMOS against the alternatives
- Ratioed logic (NMOS with a resistive or always-on PMOS load) uses transistors and presents less input capacitance, but depends on the pull-down/load ratio and a DC current flows whenever the output is low. With a resistor load, noise margin and speed are in direct tension.39
- Pass-transistor and transmission-gate logic save devices in multiplexers and XORs but pass levels without restoring them, so chains need buffering. An NMOS-only pass device also degrades the 1 by .4
- Static complementary CMOS pays devices, the PMOS input load and inverting-only stages for rail-to-rail restoring outputs, no static current and robustness to sizing. That robustness is why it is the default for standard-cell logic.
Sizing and topology trade-offs
- P/N ratio. Balanced () gives and equal edges; gives less average delay and input capacitance; skewing favors one edge. Libraries choose per cell and per drive strength, as the SKY130 numbers show.7
- NAND vs NOR. With , a NOR’s PMOS stack makes it costlier than a NAND of the same fan-in ( vs at , 3 vs 2 at ). Logic is restructured toward NANDs and AOIs where possible.6 The size of the penalty depends on , which differs between processes; ASAP7’s 1:1 inverter shows a kit whose designers chose equal PMOS and NMOS fin counts.19
- Input ordering in stacks. The parasitic delay of a stacked gate depends on which input switches last. For a NAND2 falling edge, Harris’s estimate is if the input nearest the output (the “inner” input) arrives last, and if the input nearest the rail does, because then the internal node still has to be discharged. The rule: connect the latest-arriving signal to the inner input.7
- Leakage depends on the input vector. Series off devices leak about 10× less than a single off device (the stack effect), so a gate’s leakage depends on which inputs are 0. Sleep modes can park inputs in their lowest-leakage state.9
Failure modes
- Floating inputs put an inverter in region C, both devices on, with continuous current and an output that amplifies any noise on the input. Tie cells and input pull resistors exist for this.
- Bus contention (two tristate drivers enabled at once) creates the crowbar state between gates rather than within one: a DC path, heating and an undefined level.2
- Margin erosion. Supply droop at the driver, ground bounce at the receiver and crosstalk stack up against the same . At low the margins themselves shrink.5
- Short-circuit power stays under about 10% of dynamic power only while input and output slews are comparable.9 A weak driver on a long wire breaks that assumption at the receiver, which is one reason flows enforce max-transition limits.
Healthy, input 0: PMOS on, NMOS off. The output is a solid 1 and no current flows from V_DD to GND.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
Solving for the switching threshold
At both devices are saturated and carry the same current. In the long-channel square law, with and likewise:3
Taking square roots and solving:
- With and , . Since , that needs .3
- As , ; as , .3
- The square root compresses sizing changes. With the simulator’s values (, ), gives and gives 1.03 V: a 16× change in moves by 0.27 V. That insensitivity is why is robust to process variation in the P/N ratio.5
Velocity-saturated devices change the exponent: current is closer to linear than quadratic in overdrive. Redo the algebra with and becomes instead of its square root, so the same imbalance moves further from . The shape of the argument is the same. The model behind that is in The I-V curve.
Gain and two ways to estimate noise margins
Linearizing around gives the small-signal gain .3 A quick piecewise-linear estimate replaces the curve with three straight segments: flat at , a line of slope through , flat at 0. Then:3
For the simulator’s balanced inverter, per device at and , so and (the sim’s numerical derivative, which includes the factor in the current, reports about 54). The piecewise estimate then gives . The unity-gain definition gives 0.68 V for the same curve. The piecewise version assumes the steep segment runs straight to the rails, ignoring the rounded knees of regions B and D, so it is optimistic. Use it for intuition about gain and the unity-gain numbers for budgets.
Sizing stacks and logical effort
Logical effort expresses a gate’s delay in units of (the delay of an ideal inverter driving an identical one) as , where is the electrical effort.6 The procedure behind the values, which the simulator follows:
- Size every device so the worst-case path through each network matches a unit inverter (NMOS 1, PMOS ). A device on a path of series devices gets width (or ). For nested series-parallel networks, is the longest series path through that device.
- for each input. For an -input NAND with : . For NOR: .6
- , counting the devices that touch the output. That gives for an -input NAND or NOR, and for the inverter.6 For AOI21 it gives , and .
The FO4 delay, , is the common yardstick; Harris quotes about 15 ps in a 65 nm process.6 Logical effort counts only output diffusion in , so it understates the cost of tall stacks. Elmore delay adds the internal nodes. In the NAND3 example (NMOS width 3, PMOS 2, driving identical NAND3s), summing along the stack gives a falling delay of against rising.4 For an -high stack whose internal nodes each carry capacitance proportional to , the internal-node terms sum to something proportional to . That quadratic growth is the formal version of “keep fan-in at four or below.”
The best P/N ratio
For an inverter with NMOS width 1 and PMOS width driving an identical inverter, the unit-RC delays are and . Minimize the average:7
With , gives against 3.00 for : about 3% faster on average and 20% less input capacitance, at the cost of unequal edges. Harris tabulates the same optimization for NAND2 and NOR2.7 Libraries land between the two answers, and real data (such as SKY130’s 1.54 ratio with a 3.5× drive gap) reflects the area, leakage and layout limits the RC model ignores.11
P = 2.00, μ = 2.0: fall 3.00, rise 3.00, average 3.00 RC units, 2.9% above the minimum 2.91 at P = √μ = 1.41. Input capacitance 3.00 against 3.0 balanced.
Leakage and the stack effect
Subthreshold current falls by about a decade for every 100 mV of threshold: Harris’s 65 nm example gives at , 10 nA/µm at 0.4 V and 1 nA/µm at 0.5 V.9 In a two-high stack of off NMOS, the node between them settles at a small . The upper device then sees , a body-effect increase and a smaller (less drain-induced barrier lowering), and its current drops about 10×; three-high stacks cut it further.9 Leakage-aware flows exploit this through high- cells, stacking and input-vector control in sleep. All of them trade against delay, the subject of Speed and power.
Q1In a CMOS NOR2 gate, how are the transistors arranged?
Q2An inverter’s transfer curve gives and . What is its low noise margin?
Q3Why is the PMOS in a textbook CMOS inverter often drawn about twice as wide as the NMOS?
Q4Where does a CMOS inverter draw significant current from to GND?
Sources
Show Hide 19 sources
- 1963: Complementary MOS Circuit Configuration is InventedSah and Wanlass at Fairchild presented complementary MOS logic that drew close to zero standby power; Wanlass’s patent was filed in 1963; CMOS later won because it managed power density as transistor counts grew.
- Lecture 1: Circuits & Layout (CMOS VLSI Design, 4th ed. slides)nMOS-only processes drew power while idle; complementary pull-up and pull-down networks; crowbar and float states; series/parallel conduction rules; the conduction-complement rule; compound (AOI) gates.
- Lecture 12: Digital Circuits (II), MOS Inverter Circuits (6.012 Microelectronic Devices and Circuits)Resistor and current-source pull-ups and their idle current; the CMOS inverter’s rail-to-rail levels and zero idle current; the switching-threshold equation; Wp ≈ 2Wn for a symmetric inverter; small-signal gain and noise-margin estimates.
- Lecture 5: DC & Transient Response (CMOS VLSI Design, 4th ed. slides)Degraded levels through pass transistors; the inverter DC transfer curve and its five operating regions; beta ratio and skewed gates; noise margins at the unity-gain points; the RC model (unit pMOS = 2R); Elmore delay of a 3-input NAND.
- ECE 4740 Lecture 4: The CMOS inverterSwitching threshold and its weak sensitivity to Wp/Wn; noise sources such as crosstalk and supply noise; logic-level definitions at slope −1; the regenerative property; the ideal inverter’s VDD/2 noise margins.
- Lecture 6: Logical Effort (CMOS VLSI Design, 4th ed. slides)d = gh + p; logical effort (n+2)/3 for an n-input NAND and (2n+1)/3 for NOR; parasitic delay n; FO4 delay of 5 units, about 15 ps in 65 nm.
- Lecture 9: Combinational Circuit Design (CMOS VLSI Design, 4th ed. slides)μ = 2–3 for an inverter; the P/N ratio for least average delay is √μ; choosing between many simple stages and fewer high-fan-in stages.
- Lecture 4: Nonideal Transistor Theory (CMOS VLSI Design, 4th ed. slides)Body effect: raising a transistor’s source voltage above its body raises its threshold voltage.
- Lecture 7: Power (CMOS VLSI Design, 4th ed. slides)Short-circuit current under 10% of dynamic power; static power components; a 1-billion-transistor 65 nm example (6.1 W dynamic, 859 mW static); subthreshold leakage vs threshold; the ~10× stack effect.
- Device Details: 1.8V NMOS, PMOS and high-VT PMOS FETsTypical-corner threshold voltages and saturation currents at W/L = 7/0.15 µm, and nominal fanout-of-1 inverter delays, for the 1.8 V devices.
- sky130_fd_sc_hd__inv_1.spiceThe inverter: one 0.65 µm NMOS and one 1.0 µm high-Vt PMOS, both 0.15 µm long.
- sky130_fd_sc_hd__nand2_1.spiceTwo parallel 1.0 µm PMOS and two series 0.65 µm NMOS.
- sky130_fd_sc_hd nand2 cell directoryNAND2 in drive strengths 1, 2, 4 and 8 (nand2_1 … nand2_8).
- sky130_fd_sc_hd__nor2_1.spiceTwo series 1.0 µm PMOS and two parallel 0.65 µm NMOS.
- sky130_fd_sc_hd__nand4_1.spiceFour series 0.65 µm NMOS and four parallel 1.0 µm PMOS.
- sky130_fd_sc_hd__and2_1.spiceA small NAND2 (0.42 µm devices) followed by an inverter output stage: six transistors.
- sky130_fd_sc_hd__a21oi_1.spiceAND-OR-invert: NMOS A1–A2 in series, in parallel with B1; PMOS A1 ∥ A2 in series with B1.
- sky130_fd_sc_hd cell directoryThe library’s cell list: NAND and NOR up to four inputs, plus many AOI/OAI complex gates.
- asap7sc7p5t_28_R.cdl (ASAP7 standard-cell netlists, regular Vt)FinFET cells sized in whole fins: INVx1 uses 3 fins for both NMOS and PMOS; NAND2x1 uses 6-fin NMOS and 3-fin PMOS; NOR2x1 the reverse; NAND5xp2 stacks five NMOS.