The first chapter treated a transistor as a simple switch: on or off. Real ones are more interesting. A can be off, a little bit on, or fully on.
How much electricity flows through it depends on two voltages at once. A voltage is an electric push, like the one a battery gives. The gate voltage is the hand on the switch: it decides how far the switch opens. The voltage across the switch drives electricity through once it’s open.
Engineers draw this on a chart called an I-V curve. The I stands for current, which is how much electricity flows. The V stands for voltage. The chart shows three moods: off, rising and leveled off.
A has three terminals that matter here: the gate, the source and the drain. The previous chapters used it as an on/off switch controlled by the gate. This chapter looks at the current it actually carries, the , as a function of two voltages: the gate-to-source voltage and the drain-to-source voltage .
Plotting that function gives the I-V curves, and they divide into three regions:
- : is below the and the device is nearly off.
- : the device is on and is small, so current rises with like a resistor.
- : the device is on and is large, so current levels off and is set mostly by the gate.
A simple formula, the , describes these regions well for large transistors.1 Modern transistors are so short that several extra effects bend the curves away from it. This chapter builds the simple model first, then adds those effects one at a time, and ends with how to read a real I-V plot.
The I-V characteristic is the transistor’s contract with the circuit: . Every number a digital designer cares about comes from it. The sets gate delay, the sets standby leakage, and the shape of the curve in between sets how a gate’s output slews and how much current flows while it switches. Analog designers read and off the same curves.
The long-channel gives the three regions and a useful first-order picture.1 Short-channel devices depart from it in predictable ways:
- flattens the growth of drive current with overdrive.
- and tilt the saturation curves.
- replaces “off” with an exponential tail.2
This chapter works through each effect with its standard hand model, ties the models to open SKY130 data, and ends with how SPICE compact models and threshold-extraction methods turn measured curves into the numbers in a .
V_GS = 0.0 V is below V_T: cutoff. Only a tiny leak flows, whatever V_DS is.
Under the gate is a thin strip of silicon called the channel. When the gate voltage is low, the channel is empty and the switch is off. Raise the gate voltage past its tipping point, the . Now the gate pulls electrons into the channel, which makes a path, and current can flow.
Next, the voltage across the switch takes over. When it’s small, more push means more flow. Double the push and the current roughly doubles.
When it’s large, the far end of the path gets squeezed thin, because the voltage there partly cancels the gate’s pull. The current stops growing and levels off. That’s the waterfall from the tap picture. Engineers call this . Now only the gate can raise the current.
Take an NMOS transistor with its source and body at 0 V. The gate and the silicon underneath form a capacitor with a thin insulating oxide between them. A positive gate voltage pulls electrons to the surface. Once passes , there are enough of them to form a conducting layer (the channel) linking the two heavily doped regions at the source and the drain.1
How much current flows depends on two things: how much charge is in the channel, and how fast it moves.1
- Charge comes from the gate capacitor. It is proportional to how far the gate voltage is past threshold, the .
- Speed comes from the sideways electric field that drags the charge from source to drain, where is the channel length. In a long channel, speed is that field times the , a property of the material.
The three regions
| Region | Condition (NMOS) | What the channel looks like | Behaves like |
|---|---|---|---|
| No channel | An open switch (almost) | ||
| , | Channel runs from source to drain, thinner at the drain end | A resistor set by the gate | |
| , | Channel pinched off near the drain | A current source set by the gate |
The boundary between linear and saturation is called . The voltage between gate and drain is . When it falls to , the gate can no longer hold a channel at the drain end, and the channel vanishes there. Electrons still cross the short gap to the drain, carried by the strong field, but raising further doesn’t bring more of them. The current saturates.1
PMOS works the same way with every polarity flipped. The source is the higher-voltage terminal, and the device turns on when is more negative than its (negative) threshold. Holes carry its current, and holes are about two to three times less mobile than electrons. So a PMOS has to be wider than an NMOS to carry the same current.1 The CMOS logic chapter uses that ratio to size gates.
By convention the source is the lower-voltage diffusion of an NMOS, so . Take the body at the source potential. The operating region follows from comparing and with .1
- (): no inversion layer anywhere. The long-channel model sets . In reality the device is in weak inversion and conducts a diffusion current that is exponential in (see “Below threshold”).
- (, ): the channel is inverted from source to drain. The local channel voltage rises from 0 to , so the inversion charge per area, , tapers toward the drain.
- (, i.e. ): the inversion charge reaches zero at the drain end (). The extra drops across a short depleted region next to the drain, and to first order the current no longer depends on .
A digital gate visits all three regions in every transition. When an inverter’s input rises, the NMOS leaves cutoff and enters saturation, because its drain is still at . It then slides into the linear region as the output falls below , and the PMOS goes the other way.3 That is why both (a saturation-region number) and the linear-region resistance matter for delay. The Speed and power chapter turns them into delay models.
For PMOS, flip every sign. Hole mobility is about 2–3× lower than electron mobility (120 vs. 350 cm²/V·s in the 0.6 µm example process below), so PMOS devices are drawn wider for matched drive.1
Linear: the channel runs end to end, thinner at the drain. I_D = 246 µA.
For big, old-style transistors there’s a neat rule. Once the current has leveled off, it grows with the square of how far the gate is past its tipping point. Push the gate twice as far past it, and you get four times the current.
This is the , the first formula students learn. It gets the shape of the curves right. But tiny modern transistors don’t follow it closely, as the next section shows.
Multiply the channel charge by its speed and you get the current. Working that out along the channel gives the (Shockley) model:1
- Cutoff:
- Linear:
- Saturation:
Here . The current grows with the mobility , with (the gate capacitance per unit area, larger for thinner oxide), and with the width-to-length ratio . A wider transistor is like several narrow ones side by side. A shorter one gives the charge less distance to travel.
The two formulas meet at , the saturation voltage . Plug that into the linear formula and you get the saturation formula, so the curve is smooth at the knee.
The gradual-channel approximation treats the channel as a sheet with charge per area and constant mobility. Integrating from source to drain gives the Shockley equations:14
cutoff VGS < Vt ID = 0
linear VDS < VDSAT ID = β (VGS − Vt − VDS/2) VDS
saturation VDS ≥ VDSAT ID = (β/2) (VGS − Vt)²
β = μ Cox W / L Cox = εox / tox VDSAT = VGS − Vt- 1L2Inverted-parabola in VDS. Its peak, at VDS = VGS − Vt, is where pinch-off happens.
- 2L3Holding the peak value beyond pinch-off gives a flat line: an ideal current source.
- 3L5β (sometimes written K or k′W/L) is the only process-and-geometry parameter besides Vt.
Three properties of the model are worth keeping in mind, because the short-channel effects attack each one:
- Square-law drive. , with . So the grows with bias.
- . The knee moves out linearly with overdrive.
- Flat saturation. , so the is infinite.
In MIT’s 6.012 formulation a body-effect factor also appears. It divides and the saturation current, and it is usually close to 1.4 Harris’s 0.6 µm example (, , ) gives . A device then carries about 2.2 mA at , which matches the plotted curves for that process.1 For long channels this is a good model. The next section shows how far it drifts for a 0.15 µm one.
Overdrive 2.00 V → I_D = (β/2) × 2.00² = 480 µA in saturation.
Modern transistors are amazingly short. The channel can be just a few hundred atoms long, or less. That changes the rules.
The big change is a speed limit. Electrons in silicon can only go so fast. Push them harder and they just bump into the crystal more often. In a short channel they hit that limit, so extra voltage adds less current than the square law promises. Engineers call this .
Short channels bring two smaller changes too. After the current levels off, it still creeps up a little. And the far end of the channel is so close that its voltage helps open the switch, so it leaks more when off.
Velocity saturation
The square law assumes carriers speed up in proportion to the field pushing them: . That stops being true at high fields. The carriers scatter off the vibrating silicon lattice so often that their speed levels off at a saturation velocity , about for electrons and for holes.2
The field along the channel is roughly . A 0.15 µm channel with 1.8 V across it sees , which is 120,000 V/cm. Even at the 350 cm²/V·s mobility of the older example above, would come to , four times the speed limit. So in short channels the carriers hit before the channel pinches off. The current then levels off at a lower than the square law predicts, and it grows more slowly with gate voltage. Once carriers move at a fixed speed, current is just charge times that speed, and charge is proportional to the overdrive. So the current grows roughly in proportion to instead of as its square.2 Course notes place the effect as visible from about 250 nm channel lengths downward.5
Real devices sit in between. A popular fit, the , replaces the exponent 2 with a fitted between 1 and 2. A typical value for a 65 nm process is about 1.3.2
Channel-length modulation
In saturation, the depleted region around the drain widens as rises, and it eats into the channel. The effective length gets shorter, so the current rises a little. Models capture this with a factor , where is fitted to measurements.2 On the plot, the saturation lines tilt upward instead of running flat. The shorter the channel, the bigger the tilt, because the same depletion width is a larger fraction of it.
Drain-induced barrier lowering
In a long channel only the gate controls the energy barrier that keeps electrons in the source. In a short one the drain is close enough to pull that barrier down too. The threshold voltage then drops as rises.2 This (DIBL) adds to the tilt in saturation, and it raises leakage, as the next section shows. The Shrinking chapter covers how FinFETs and gate-all-around transistors fight it.
Mobility degradation and velocity saturation
Two high-field effects bend the long-channel curves.2
- Mobility degradation. The vertical field (about ) presses carriers against the oxide interface, where they scatter more. The effective falls as rises.
- Velocity saturation. The lateral field (about ) can’t accelerate carriers past for electrons and for holes.
Compare the two limits:
- Long channel: .
- Fully velocity-saturated: .2
The second is linear in overdrive and has no in it. Once carriers run at , a shorter channel no longer buys more saturation current through the term, only through higher and lower . Real devices are partly saturated, and fit them with an exponent between 1 and 2. Harris gives the -power saturation current and (line 1 below); the linear-region and terms (lines 2 and 3) are Sakurai and Newton’s:26
IDsat = Pc (β/2) (VGS − Vt)^α VDSAT = Pv (VGS − Vt)^(α/2)
VDS < VDSAT: ID = IDsat · (2 − VDS/VDSAT) · (VDS/VDSAT)
VDS ≥ VDSAT: ID = IDsat · (1 + λ VDS)- 1L1α = 2, Pc = Pv = 1 recovers the square law. α → 1 is the fully velocity-saturated limit, with VDSAT ∝ √VGT.
- 2L2Sakurai and Newton’s linear-region form: it meets IDsat with zero slope at VDSAT, so the curve is smooth.
- 3L3The (1 + λVDS) factor is their finite drain conductance in saturation.
Sakurai and Newton’s 1990 report states the limits directly. Their model reduces to Shockley for , and it can also express short-channel theory’s prediction that and .6 Fitted values sit in between: about 1.3 for a 65 nm process.2 Two consequences follow for digital design:
- Drive current grows as about rather than . Raising the supply buys less speed than the square law suggests.
- is well below . A switching transistor reaches its saturation current at a lower , and the linear region is compressed.
Channel-length modulation
The reverse-biased drain junction’s depletion width grows with , so shrinks and keeps rising in saturation.2 The usual empirical forms are , or in MIT’s 6.012 notation, with by analogy to the bipolar Early voltage.4 is not the layout . It is a fitted coefficient, and it grows as shrinks.2 The follows directly: .
DIBL and other threshold shifts
isn’t a constant.2
- Body effect. rises with source-to-body reverse bias.
- . falls with , modeled as .
- Short-channel effect. depends on . Some processes show a reverse short-channel effect, where is higher at short .
In saturation, DIBL raises with on top of channel-length modulation, so the measured of a short device is lower than alone predicts. Its biggest effect is below threshold, where a shift becomes an exponential change in current.
L = 0.60 µm: field ≈ 30 kV/cm. µE would be 1.1×10⁷ cm/s (1.1× v_sat); the electrons actually move at about 0.5×10⁷ cm/s.
A perfect switch would let nothing through when off. A transistor can’t quite do that. Below its tipping point, a small trickle still gets through. It’s called current, which means “below the threshold.”
The trickle follows a neat pattern. Lower the gate voltage by about a tenth of a volt, and the trickle shrinks to a tenth. It never reaches zero.
That matters because most of a chip’s billions of transistors are off at any moment. Their trickles add up to real wasted power, called . It flows even when the chip is doing nothing.
The square law says below . Measure a real transistor and you find a current that falls off smoothly instead: .2 Below threshold the gate hasn’t formed a full channel, but it still controls the height of an energy barrier between the source and the channel. A few electrons are energetic enough to get over it, and their number grows exponentially as the gate lowers the barrier.7
On an ordinary plot this current is invisible. On a log-scale plot of against it becomes a straight line. Its steepness is the : the gate-voltage drop that cuts the current tenfold (one decade).
- The physics sets a floor of about 60 mV per decade at room temperature. Real devices are worse by a factor , so . One MIT example device has , which gives 75 mV/decade.7
- Typical values are and .2
The slope ties speed to leakage. Suppose and you lower by 100 mV to make the transistor faster. The current at , the , goes up about ten times. The Speed and power chapter follows this into chip-level . Two more things make it worse:
- Heat. Higher temperature lowers , so off-current rises as the chip warms up.2
- DIBL. In short channels, a high drain voltage lowers too, so the device leaks more with across it than at low .
In weak inversion the channel potential at the surface tracks through a capacitive divider (oxide capacitance against depletion capacitance), and drain current is diffusion over the source barrier. The 6.012 derivation gives:7
ID ≈ I0 · exp[(VGS − Vt) / (n·kT/q)] · [1 − exp(−VDS / (kT/q))]
n = 1 + Cdep/Cox (≈ 1.3–1.7 in practice)
S = n · (kT/q) · ln 10 ≈ 60·n mV/decade at 300 K
Vt → Vt − η·VDS (DIBL) Vt → Vt + γ-term (body effect)- 1L1Exponential in VGS. The drain factor saturates within a few kT/q (~26 mV each), so ID barely depends on VDS except through DIBL.
- 2L3n is the same capacitive-divider factor that appears as α in the 6.012 above-threshold model.
- 3L4kT/q = 25.85 mV at 300 K, and ln 10 ≈ 2.303, which gives the 59.5 ≈ 60 mV/decade floor.
MIT’s notes plot this for a 3 nm oxide and device: , so 75 mV/decade. They also note that the current at threshold is small next to strong-inversion current even 0.1 V above .7 Harris quotes and at room temperature, and writes the full expression with DIBL () and body-effect terms.2
For design, the useful form is
adjusted for DIBL at . Each of threshold reduction costs one decade of leakage, and each of DIBL shift between and costs another. Temperature attacks twice:
- grows, so grows in proportion to absolute temperature.
- falls. So rises steeply with temperature while falls, because mobility drops.2
Subthreshold current isn’t the only leakage. Harris lists gate tunneling (critical once reached about 10.5 Å at 65 nm) and junction leakage (reverse diode current, band-to-band tunneling and gate-induced drain leakage). It names subthreshold conduction as the dominant source in contemporary transistors.2
Note what a hand model like the -power law does here: nothing. Sakurai and Newton say outright that their model is a rough approximation near and below threshold, which they judged unimportant for delay.6 Leakage analysis needs the full all-region compact model.
V_GS = 0.00 V, 400 mV below V_T: I_D is 4.0 decades below its value at threshold and full V_DS.
To test a new transistor, engineers set the two voltages, measure the current, and repeat thousands of times. The result is a family of curves. Each one climbs from the left and flattens out to the right. Each curve is one gate setting: the higher the gate voltage, the higher the curve.
Engineers look first at how high the top curve gets. That’s the most current the switch can pass, and more means a faster chip. They also like the curves flat on the right, because flat means the current holds steady.
A second chart shows how well the switch turns off. Its scale stretches tiny numbers apart, so you can see the trickle. The steeper the line, the cleaner the switch.
Device engineers measure or simulate two standard plots:
- Output characteristics: against , one curve per step. Read off:
- the , the top curve at ;
- where each curve bends from linear into saturation;
- how much the saturated part tilts upward.
- Transfer characteristics: against at a fixed , often on a log scale. Read off:
- the threshold voltage, and the at , 2;
- the subthreshold slope;
- with two curves, one at low and one at high , the DIBL shift between them.
Currents are usually quoted per micrometer of channel width (µA/µm), because current scales with width and the number per micrometer lets you compare devices of any size.
“Threshold voltage” is less exact than it sounds. The current has no sharp corner, so labs define by a convention, for example the where the current reaches a set value, or where a straight line through the steepest part of the curve hits zero. Measured on the same 0.18 µm transistors, different conventions disagree by tens of millivolts.8 When you compare thresholds, compare the method too.
You don’t need a lab to make these plots. A circuit simulator such as SPICE sweeps the voltages through a model of the transistor. The deck below uses the square-law model with the 0.6 µm numbers from earlier, and it reproduces the 2.2 mA top curve.1
* NMOS output characteristics, square-law (Level 1) model
.model nch nmos (level=1 vto=0.7 kp=120u lambda=0)
M1 d g 0 0 nch W=1.2u L=0.6u
Vd d 0 0
Vg g 0 0
.dc Vd 0 5 0.05 Vg 0 5 1
.control
run
plot -i(Vd)
.endc
.end- 1L2Level 1 is the Shockley square law. vto is Vt, kp is μCox (120 µA/V²) and lambda is channel-length modulation (off here).
- 2L3W/L = 2, so β = 240 µA/V². Drain, gate, source, body.
- 3L6Sweep VDS from 0 to 5 V for each VGS from 0 to 5 V in 1 V steps: one curve per gate voltage.
- 4L9SPICE reports a voltage source’s current as flowing into its + terminal, so the drain current is −i(Vd).
A device characterization run produces a standard set of sweeps. Each one isolates a parameter:
| Sweep | Plot | What you extract |
|---|---|---|
| at low (tens of mV) | Linear –, plus | , low-field mobility, series resistance (from how the curve bends at high ) |
| at low and high | Log – | (mV/dec), , , (mV/V) |
| at several | – family | , versus (velocity saturation), saturation slope (, ) |
| Same sweeps at several and temperatures | Families of the above | roll-off or reverse roll-off, temperature coefficients, corner spreads |
Things to check when you read one:
- Normalization. µA/µm at what , temperature and ? An without its conditions is not comparable. Corners shift the curves noticeably. Harris’s example uses ±10% supply and 0–125 °C.2
- Which . Linear-extrapolation is skewed by mobility degradation and series resistance. Constant-current depends on the chosen current. Second-derivative and methods use other definitions again. On 0.18 µm NMOS at , the methods gave 481 mV (linear extrapolation), 490 (second derivative), 520 () and 501 (constant current).8
- Spacing of the output family. With equal steps, curve spacing that grows with signals square-law behavior. Spacing that stays roughly constant signals velocity saturation ( independent of ).
- Saturation slope. Compare it at low and high . Channel-length modulation gives a slope proportional to . DIBL adds a component that is largest near threshold, where a shift matters most.
In simulation the same sweeps run against the PDK’s . The ngspice deck below uses Level 1 with Harris’s 0.6 µm parameters, which reproduces his 2.2 mA top curve.1 Swapping the .model line for a PDK’s BSIM model, and the M line for the kit’s device, turns it into a real characterization run.
* NMOS output characteristics, square-law (Level 1) model
.model nch nmos (level=1 vto=0.7 kp=120u lambda=0)
M1 d g 0 0 nch W=1.2u L=0.6u
Vd d 0 0
Vg g 0 0
.dc Vd 0 5 0.05 Vg 0 5 1
.control
run
plot -i(Vd)
.endc
.end- 1L2Level 1 = Shockley. kp is μCox. Setting lambda > 0 adds (1 + λVDS); Level 1 has no velocity saturation or subthreshold current at all.
- 2L6Nested sweep: VDS inner, VGS outer, giving the output family. Swap the order (.dc Vg … Vd …) for transfer curves.
- 3L9Plot log(-i(Vd)) on a transfer sweep to see the subthreshold region. With Level 1 it is simply missing.
Output characteristics: I_D against V_DS, one curve per V_GS. Tap a feature.
Move the two sliders and watch the dot. It marks your transistor’s current on the chart. The panel says whether the switch is off, on with more push giving more flow, or on with the flow leveled off.
Then flip from “Textbook” to “Real, tiny.” Watch the curves get squashed and tilted, and spread out more evenly.
The thick line is the – curve for your . The gray family shows from 0.4 to 1.0 V, and the dashed line marks the edge of saturation (). You can drag on the plot to move . Switch between the long-channel square law and the short-channel model, and turn on Compare to overlay the other one. Notice three things in the short-channel model: the knee moves left, the top curves bunch together, and the saturated lines tilt up.
Two models of one generic NMOS (, , currents per µm):
- the square law;
- an -power law () with and DIBL of 80 mV/V, matched to the square law at .
Both use a smooth subthreshold tail with . Switch to the log – view to read . The dotted curve at shows the DIBL shift. The readouts give , (infinite in long-channel saturation), and .
These numbers come from SKY130, a free chipmaking recipe that anyone can read online. They describe a typical small transistor made with it.
- Turn-on voltage of a small SKY130 transistor (typical)
- 0.645 V
- Most current it passes, fully on
- 3.5 mA
- Same-size transistor of the opposite type
- 1.3 mA
- Top speed of electrons in silicon
- 100 km/s
What these numbers mean:
- The switch starts to turn on at a bit under two-thirds of a volt. An AA battery gives one and a half volts.9
- Fully on, it passes 3.5 thousandths of an amp (mA). An amp is the unit for current. That’s tiny, but billions of transistors together run the whole chip.9
- The other kind of transistor, at the same size, passes less than half as much. So designers make those wider.9
- Electrons in silicon top out at about 100 kilometers per second. That’s fast, but in a tiny transistor it’s the limit that matters.2
SKY130 is an open 130 nm process kit. Its documentation lists model numbers next to the foundry’s test specifications for each transistor type. For the 1.8 V devices, the SPICE models are valid for and up to 1.95 V.9 The figures below are typical-corner (TT) model values.
- NMOS Vt, W/L = 7/0.15 µm
- 0.645 V
- NMOS saturation current, same device
- 3.51 mA
- PMOS saturation current, same size
- 1.35 mA
- Electron saturation velocity
- ~10⁷ cm/s
| SKY130 1.8 V device, typical corner | Saturation current | Per micrometer of width | |
|---|---|---|---|
| NMOS, standard | 0.645 V | 3.512 mA | ≈ 0.50 mA/µm |
| NMOS, low- | 0.611 V | 4.010 mA | ≈ 0.57 mA/µm |
| PMOS, standard | −0.781 V | 1.347 mA | ≈ 0.19 mA/µm |
| PMOS, high- | −0.888 V | 1.003 mA | ≈ 0.14 mA/µm |
All four devices are 7 µm wide and 0.15 µm long. Source: SKY130 device documentation; per-µm values divide by the 7 µm width.9 What the table shows:
- The NMOS carries about 2.6 times the current of the same-size PMOS, in line with electrons’ higher mobility.
- The low- NMOS gains 14% more current for a threshold just 34 mV lower.
For scale: at 1.8 V across 0.15 µm, the field along the channel is about 120,000 V/cm. That is far into the range where carrier speed approaches its limit of about .2
The SKY130 device tables give model values against e-test specifications, across TT/FF/SS/FS/SF corners, for the 1.8 V FETs. The models are valid for .9
- SKY130 NMOS 7/0.15 Idsat (TT / FF / SS)
- 3.51 / 3.95 / 3.08 mA
- SKY130 NMOS Vt, 7/8 vs 7/0.15 (TT)
- 0.538 vs 0.645 V
- α-power exponent, 65 nm (Harris)
- ≈ 1.3
- Typical subthreshold slope at room temperature
- ≈ 100 mV/dec
| SKY130 1.8 V parameter (TT) | (micrometers) | Standard | Low- / high- |
|---|---|---|---|
| NMOS , long and wide | 7/8 | 0.538 V | 0.434 V (LVT) |
| NMOS , short and wide | 7/0.15 | 0.645 V | 0.611 V (LVT) |
| NMOS | 7/0.15 | 3.512 mA | 4.010 mA (LVT) |
| PMOS | 7/0.15 | −0.781 V | −0.888 V (HVT) |
| PMOS | 7/0.15 | 1.347 mA | 1.003 mA (HVT) |
Values from the SKY130 device tables.9 Readings worth making:
- Corner spread. FF to SS spans 3.95 to 3.08 mA, about ±12% around TT. Timing signoff has to cover that whole spread.
- versus . The short device has a higher threshold than the long one (0.645 vs. 0.538 V). That matches what Harris calls the reverse short-channel effect.2
- Threshold flavors. The LVT device trades 34 mV of for 14% more . At a typical 100 mV/decade, 34 mV is worth roughly 2× in subthreshold leakage. The tables report leakage too, in log amps.
- NMOS/PMOS ratio. at equal size, consistent with the 2–3× mobility ratio.1
For physical scale, at and the average lateral field is 120 kV/cm. with any plausible channel mobility then exceeds the electron of , so these devices sit well inside the partly velocity-saturated regime the -power law describes.2
Every way of getting more current out of a transistor costs something.
A lower tipping point makes the switch stronger when on, but it leaks more when off. So engineers use a few “eager” transistors where speed matters most, and “lazy” ones everywhere else.
A bigger supply voltage pushes more current. But because of the speed limit, tiny transistors give back less than you’d hope, while the power cost keeps climbing.
Heat works against you both ways. A hot transistor is weaker when on and leakier when off.2
And the neat textbook rule can fool you. For tiny transistors, its guesses can be way off. So designers use detailed computer models, built by measuring real transistors.
What you trade
- Threshold voltage: on-current against off-current. Lowering raises the overdrive, so goes up. Every subthreshold slope’s worth of reduction (60–100 mV) also multiplies by ten. Processes therefore offer several threshold flavors. SKY130’s low- NMOS, for example, gets 14% more current than the standard one.9
- Supply voltage: speed against power. More means more overdrive and more current. With velocity saturation, current grows more slowly than the square law predicts, while switching power keeps growing with .2
- Channel length: drive against steadiness. Shorter channels give more current per micrometer. They also have stronger and : the saturation curves tilt more, and the device leaks more at full drain voltage. Analog circuits that need a steady current often use longer-than-minimum transistors.
- Temperature. Heat lowers mobility, so on-current falls. It also lowers , so off-current rises.2
Ways it goes wrong
- Trusting the square law for short devices. It overestimates drive at high gate voltage and misses leakage entirely. Sakurai and Newton showed it can badly misjudge relative gate delays, too.6
- Mixing up two “saturations”. The saturation region (current flat with ) and velocity saturation (carriers at their speed limit) are different things. Harris warns students not to confuse them.2
- Comparing numbers measured differently. A threshold voltage or on-current means little without its bias, temperature and extraction method.8
The trade space
- selection. , while . Near , a 100 mV cut buys roughly 20–35% drive () for 10× leakage at . Multi- libraries let synthesis spend leakage only on critical paths. SKY130 ships LVT and HVT variants alongside the standard devices.9
- scaling. Because grows only as about rather than , raising the supply shortens gate delay () less than the square law suggests, while switching energy still grows as . The speed bought per unit of power falls faster than a square-law estimate predicts.2
- Channel length. Below the length where velocity saturation dominates, stops scaling as . Meanwhile , DIBL and sensitivity to all get worse. Intrinsic gain falls, which is why analog blocks use longer devices than logic. Gate control over the channel is the lever that the Shrinking chapter’s FinFET and GAA structures pull.
- Temperature. falls with (mobility), but rises (lower , larger ). Corner selection reflects this: Harris’s table puts the cycle-time corner at SS/low /high and the leakage corner at fast devices, high and high temperature.2
Failure modes
- Hand models outside their range. Sakurai and Newton’s motivating example: with velocity saturation, an 8-input NAND is only about 4–5× slower than an inverter (fast input, large load), while the Shockley model always predicts more than 8×. The series stack shares and , which relieves velocity saturation in each device.6 Use the square law for intuition, never for sizing.
- Extrapolating fitted parameters. -power and similar fits hold only near the , and corner they were extracted at. Sakurai and Newton note that a parameter set is valid only for a narrow range of channel length.6
- Ignoring DIBL in leakage budgets. must be read at , not from a low- transfer curve. The gap is , often several times.
- Threshold apples and oranges. from linear extrapolation, from constant current at high , and a SPICE model’s VTH0 parameter are different quantities. On one 0.18 µm process, linear extrapolation and the method differed by roughly 20–40 mV.8
Reference point: V_DD = 1.0 V. Move the slider to compare speed and energy against it.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
From hand models to compact models
Circuit simulators need , terminal charges and their derivatives at every bias, in every region, evaluated millions of times per transient. The ladder of models trades simplicity for accuracy.
- SPICE Level 1 (Shockley). The three-region square law with an optional . Sakurai and Newton showed it is “far from accurate in the short-channel region” and insufficient even for first-order relative delay.6
- Power-law and unified hand models. Sakurai and Newton’s nth-power model uses a handful of appearance-oriented parameters. It evaluated in about a third of the time of SPICE Level 3, and its parameters come from a few points on measured curves.6 The unified model taught in digital IC courses takes and , so one expression covers linear, saturation and velocity saturation.5
- Physics-based . BSIM4 extends BSIM3 into the sub-100 nm regime. It has been used from 0.13 µm down to 22/20 nm, and its development is coordinated through the Compact Model Coalition.10 BSIM-CMG covers FinFETs and other multi-gate devices with a surface-potential formulation; its CMC standard releases date back to 2012.11 Models such as BSIM3, BSIM4 and PSP run to more than ten thousand lines of C.5
One equation across all regions
Piecewise models have kinks at region boundaries. The unified model’s min() has a discontinuous derivative at , and Level 1 has no current at all below . Newton–Raphson in a circuit simulator needs smooth currents and conductances, and leakage analysis needs the subthreshold tail. So production models use single expressions that blend weak and strong inversion. Charge-based all-region models, such as the ACM model used by Schneider and colleagues, are one approach.8
The simulation on this page uses the same idea in miniature. It replaces the overdrive with a softplus that becomes exponential below threshold:
φt = kT/q = 25.85 mV n = 1.4 Vt = 0.35 V − η·VDS
w = α · n · φt
VGT* = w · ln(1 + exp((VGS − Vt)/w)) (softplus overdrive)
Isat = K · VGT*^α
Vsat = sqrt( (Pv · VGT*^(α/2))² + (2φt)² )
x = min(VDS / Vsat, 1)
ID = Isat · (2 − x) · x · (1 + λ·VDS)
long: α = 2, K = 2442 µA/V²/µm, Pv = 1, λ = 0, η = 0
short: α = 1.3, K = 1051, Pv = 0.46, λ = 0.15/V, η = 0.08 V/V- 1L2Scaling the softplus width by α makes VGT*^α ∝ exp((VGS − Vt)/(nφt)) below threshold, so S = n·φt·ln10 ≈ 83 mV/dec for both models.
- 2L3Above threshold VGT* → VGS − Vt, recovering the square law (α = 2) or the α-power law.
- 3L5The 2φt floor keeps VDSAT finite in subthreshold, standing in for the (1 − e^(−VDS/φt)) drain factor.
- 4L7Sakurai–Newton linear-region form: smooth at VDSAT, with CLM applied everywhere so ID stays continuous.
- 5L10K and Pv are chosen so both models agree at VGS − Vt = 0.3 V; all values are illustrative, not from a PDK.
Extracting parameters from measured curves
Fitting a model means choosing parameters so the model’s curves pass through measured ones. Sakurai and Newton made this nearly closed-form.6
- Pick a handful of points on the measured I-V curves.
- Divide out the channel-length-modulation factor from the saturation points, then solve one equation for the threshold by bisection (convergence within about 10 iterations).
- Get and from the ratio of two saturation currents.
- Get and from two linear-region points.
The whole fit runs in under a second, with no multi-parameter optimizer. The parameters are valid only near the channel length they came from, so a design uses two or three sets.6 Full BSIM extraction applies the same idea at much larger scale, fitting many more parameters to sweeps across geometry, bias and temperature.
Defining and extracting the threshold
Because is smooth, is defined by a procedure. Most methods use a transfer curve at low (below in Schneider et al.’s analysis).8
- ELR (extrapolation in the linear region). Fit a tangent at the point of maximum and take its intercept. Simple, but biased by mobility degradation, series resistance and the nonlinear charge–voltage relation.
- Transconductance change / second derivative. The peak of . Physically meaningful, but the second derivative of measured data is very noisy.
- Constant current. The at which reaches a chosen value. Simple and accurate, provided the effective and are known.
- . The where falls to half its subthreshold peak.
ELR: Tangent at the point of maximum g_m (V_GS = 0.57 V), extrapolated to I_D = 0: V_T = 435 mV. Mobility degradation pulls the max-g_m point, and with it the answer.
On 0.18 µm devices, these methods spread over about 25–45 mV at each channel length. The and second-derivative results differ systematically by about , which is at least 30 mV.8 A PDK’s quoted is therefore always tied to a method and a bias. SKY130’s table, for instance, lists separate e-test parameters for each geometry.9
The next chapter, Speed and power, turns , and the -power law into delay and power. The cell-library chapter shows how these device curves become the timing tables synthesis and place-and-route read.
Q1An NMOS transistor has , and . Which region is it in, in the square-law model?
Q2In the square-law model, doubling the gate overdrive () in saturation changes the drain current by about how much?
Q3On a log-scale plot of drain current against gate voltage, the curve below threshold is a straight line. What does that tell you?
Q4In a measured – plot, the saturation curves slope gently upward instead of being flat. What is the usual reason?
Sources
Show Hide 11 sources
- Lecture 3: CMOS Transistor Theory (CMOS VLSI Design, 4th ed. slides)Cutoff, linear and saturation; channel charge and carrier velocity; the Shockley I-V equations with β = μCoxW/L and Vdsat = Vgs − Vt; a 0.6 µm example (tox 100 Å, μ 350 cm²/V·s, Vt 0.7 V); pMOS mobility 2–3× lower.
- Lecture 4: Nonideal Transistor Theory (CMOS VLSI Design, 4th ed. slides)Ion and Ioff definitions; mobility degradation; velocity saturation (vsat ≈ 10⁷ cm/s for electrons, 8×10⁶ for holes); the α-power law (α ≈ 1.3 at 65 nm); channel-length modulation; DIBL; subthreshold leakage (n ≈ 1.3–1.7, S ≈ 100 mV/decade); temperature effects; process corners.
- Lecture 5: DC and Transient Response (CMOS VLSI Design, 4th ed. slides)Which region the nMOS and pMOS of an inverter are in as Vin and Vout move: cutoff for Vin < Vtn, saturated for Vout > Vin − Vtn, linear below that.
- Lecture 11 – MOSFETs II; Large Signal Models (6.012 Microelectronic Devices and Circuits)Gradual-channel model in cutoff, linear and saturation; channel-length modulation as L(vDS), iD ≈ K(vGS − VT)²[1 + λ(vDS − VDSat)]/2, with λ = 1/VA.
- Prerequisite 1: Devices — MOS transistors, models, scaling (ET4293 Digital IC Design slides)Operating regimes including velocity saturation, visible at 250 nm and below; the unified model for hand analysis with V_min = min(V_GT, V_DS, V_DSAT); SPICE models such as BSIM3, BSIM4 and PSP run to more than 10k lines of C.
- A Simple MOSFET Model for Circuit Analysis and Its Application to CMOS Gate Delay Analysis and Series-Connected MOSFET Structure (Memorandum UCB/ERL M90/19)The nth-power-law model (Idsat ∝ (VGS − VTH)ⁿ, VDSAT ∝ (VGS − VTH)ᵐ) that reduces to Shockley at n = 2; why SPICE Level 1 fails for short channels; parameter extraction from a few points on measured curves; rough below threshold by design.
- Lecture 12 – Sub-threshold MOSFET Operation (6.012 Microelectronic Devices and Circuits)Derivation of the subthreshold current as diffusion over a gate-controlled barrier: exponential in (vGS − VT)/(nkT/q), a (1 − e^(−qvDS/kT)) drain factor, and a log-scale slope of 60 × n mV/decade (75 mV/decade for n = 1.25).
- About the Concept of Threshold in MOS TransistorsThreshold-voltage definitions and extraction methods (linear extrapolation, second derivative, gm/ID, constant current) measured at low VDS on 0.18 µm devices; the methods disagree by tens of millivolts.
- Device DetailsModel-versus-e-test tables for the 1.8 V NMOS and PMOS (standard, low-Vt and high-Vt): threshold voltages and saturation currents at W/L = 7/8 and 7/0.15 µm, across TT/FF/SS corners; SPICE models valid for |VDS|, |VGS| up to 1.95 V.
- BSIM4BSIM4 extends BSIM3 into the sub-100 nm regime as a physics-based MOSFET SPICE model, used from 0.13 µm to 22/20 nm, with improvements charted by the Compact Model Coalition.
- BSIM-CMGA surface-potential-based compact model for common multi-gate FETs (FinFETs), with CMC standard releases since 2012.