Transistors · Chapter 5 of 8 · On the chip

Shrinking

For decades, transistors shrank every couple of years. At tiny sizes they started to leak, so their shape changed: from flat, to a thin fin, to a path the switch’s control wraps all the way around.

As channels got shorter the gate lost control, and transistors leaked even when off. FinFETs wrap the gate around three sides of a fin. Gate-all-around transistors wrap it around stacked sheets. Newer ideas stack NMOS on PMOS and move the power wiring under the transistors.

Short-channel effects (DIBL, degraded subthreshold slope), the planar → FinFET → gate-all-around nanosheet transition, stacked CFETs, backside power delivery, and why node names stopped describing a physical dimension. FinFET-era numbers come from the open ASAP7 predictive kit.

For about fifty years, chipmakers made transistors smaller every couple of years. Smaller switches are cheaper, faster and use less energy. So every shrink made computers better.

A transistor has a control part called the gate. The gate decides whether electricity can flow along a narrow path called the channel. But make the switch too short, and the far end of the channel starts to interfere. Then the switch leaks, even when it should be off.

The fix was to change the transistor’s shape, so the gate wraps around more of the channel.

A transistor (more precisely a MOSFET) is a switch: a voltage on its gate lets current flow from its source to its drain, or stops it. The distance from source to drain under the gate is the . Shrinking it, along with everything around it, made chips cheaper and faster for decades.

Shrinking has a limit set by electrostatics. The gate controls current by holding up an energy barrier between source and drain. When the gate is only a few times longer than the reach of the drain’s electric field, the drain lowers that barrier too, and current leaks through a switch that should be off. These are measured mainly by the (how sharply the switch turns off) and (how much the drain voltage weakens it).

The industry’s answer was to give the gate more of the channel to hold: flat (planar) transistors gave way to in the early 2010s and are now giving way to stacked . Next come , which stack the two transistor types on top of each other, and , which moves power wiring under the transistors. Along the way, process names like “7 nm” stopped describing any real dimension.

This chapter is about the electrostatic limit of MOSFET scaling and the architectural moves that pushed it back. The central quantity is the λ\lambda, the distance over which source and drain potentials decay into the channel. When the gate length LL is many λ\lambda long, the gate alone sets the barrier; when LL approaches a few λ\lambda, the drain shares control, which shows up as , a degraded , threshold roll-off and finally punch-through.

λ\lambda shrinks with a thinner channel, a thinner and more gated sides. Planar bulk devices ran out of the first two; and then the attacked channel thickness and gate count, and take both to their practical limit. Beyond that, density comes from stacking (), from the wiring () and from , not from a shorter gate. FinFET-era numbers here come from ASAP7, an open, predictive 7 nm from Arizona State University and Arm Research.

sourcedraingate1. Planar, long gategate on 1 sidegatesource / drain (n+)silicon body
1 / 5

A planar MOSFET, off. The gate controls the channel from above only. With a long gate it still holds the source-to-drain barrier up.

From planar to CFET. Each architecture gives the gate more of the channel to hold, so the transistor can be shorter before the drain takes over. Sketches, not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

In 1965, Gordon Moore noticed a pattern. The number of parts that could be cheaply packed onto one chip was doubling about every year. That pattern became known as “Moore’s law.”

In 1974, engineers at IBM wrote a recipe for shrinking. Make every part of a transistor smaller. Lower the voltage, the push that drives electricity, by the same amount. Then each transistor gets faster and uses less power, and a chip packed with them stays just as cool. This recipe is called .

So shrinking was a triple win: more switches, faster switches and no extra heat.

But the recipe had a catch, and the 1974 paper said so. The voltage can only go so low before the switch can’t turn fully off. Around 2005, simple shrinking slowed down. Since then, progress has come from new materials and new transistor shapes.

Moore’s 1965 paper was about cost: the number of components at which a chip is cheapest per component was doubling roughly every year. Shrinking every feature is the main way to get there. For decades each generation made features about 30% smaller in both directions, which halves the area of a transistor and doubles how many fit in the same space.

Robert Dennard and colleagues at IBM showed in 1974 why shrinking also made each transistor better. Divide every length, and the voltage, by the same factor κ\kappa (say 1.4, a 0.7× shrink), and raise the doping by κ\kappa:

  • Each transistor’s capacitance falls by κ\kappa, so it charges faster: delay falls by κ\kappa.
  • Voltage and current both fall by κ\kappa, so power per circuit falls by κ2\kappa^2.
  • Area falls by κ2\kappa^2 too, so power per square millimeter stays the same. Chips could double their transistor count without needing better cooling.

Dennard flagged the exception in the same paper: the steepness with which a transistor turns off does not scale. At room temperature it takes about 60 millivolts of gate voltage to cut the off-current tenfold, however small the device. So the (the gate voltage at which the switch turns on) can’t keep falling without the off-state rising exponentially, and the supply voltage has to stay well above the threshold for speed. Once supply voltage stalled, ended. The Speed and power chapter tells that part of the story.

Shrinking continued anyway, because density still lowers cost. But the gate length itself stopped shrinking as fast as the spacing between transistors, and performance came from other tricks: strained silicon, new gate materials, and new transistor shapes.

Dennard’s constant-field rules (Table I of the 1974 paper) scale dimensions, VV and II by 1/κ1/\kappa and doping by κ\kappa: delay per circuit 1/κ1/\kappa, power per circuit 1/κ21/\kappa^2, power density 1. Two caveats in the same paper defined the next forty years. First, the subthreshold slope is (kT/q)ln⁡10⋅m(kT/q) \ln 10 \cdot m and doesn’t scale with dimensions, so VTV_{\mathrm{T}} and VDDV_{\mathrm{DD}} can’t follow 1/κ1/\kappa without an exponential IoffI_{\mathrm{off}} penalty. Second, the source and drain depletion regions must not merge under the gate, or threshold control is lost and the device punches through; the paper’s shallow junctions and surface implant were there to prevent that.

The first caveat ended constant-field scaling around 2005. The ASAP7 authors summarize the post-Dennard regime: VDDV_{\mathrm{DD}} stalled because of Ion/IoffI_{\mathrm{on}}/I_{\mathrm{off}} constraints and variability, gains came from strain and , FinFETs relieved short-channel effects and allowed a little more VDDV_{\mathrm{DD}} scaling, and ASAP7 assumes only about 100 mV or less of VDDV_{\mathrm{DD}} reduction from 14 nm to 7 nm, with drive current up about 15% per node. After about 2005, progress was “equivalent scaling” (HKMG, FinFETs, strain) rather than purely geometric.

The second caveat is the subject of the rest of this chapter. Gate length is the one dimension electrostatics won’t let you shrink freely, which is why it stopped tracking pitch (it shrank faster than pitch from the mid-1990s, then slower) and why node names drifted away from it.

fixed area of silicon4 transistors1× (no change)lessmoreTransistors per area1.0×Speed (1 / delay)1.0×Power per transistor1.0×Power per area1.0×
Supply voltage

Move the slider: each step shrinks every length and the voltage to 0.7×.

Dennard’s 1974 rule: shrink every length and the voltage by the same factor, and power per area stays constant. The ‘voltage stuck’ case is illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Think of the channel as a path with a hill in the middle. Electrons start at one end, called the source. They want to reach the other end, called the drain. The gate sets how high the hill is. A high hill means the switch is off. Flat ground means it’s on.

The drain sits lower, at the bottom of a slope. In a long channel, that slope is far from the hill and doesn’t matter. In a very short channel, the slope eats into the side of the hill and pulls it down. Now electrons trickle over even when the switch is “off.”

So the switch leaks and wastes power all the time. It also needs a bigger push from the gate to shut off. These problems are called . They’re why you can’t just keep shrinking the same flat transistor.

Below its threshold voltage a transistor is “off,” but not perfectly: a small current flows by electrons spilling over the source-to-drain barrier, and it rises exponentially as the gate lowers the barrier. Two numbers describe how good that off state is.

  • (SS) is the gate voltage needed to change the off-current by a factor of ten, in mV/decade. Physics sets a floor of about 60 mV/decade at room temperature for any ordinary transistor; older flat transistors ran about 20 to 50% above it because part of the gate’s voltage is wasted on the silicon underneath. A 60 mV/decade switch goes from 1 nA to 1 µA (three decades) in 180 mV of gate voltage; a 100 mV/decade switch needs 300 mV.
  • (drain-induced barrier lowering) is how much the drain voltage lowers the threshold, in mV per volt. With DIBL of 100 mV/V and a 0.7 V supply, the threshold is 70 mV lower whenever the drain is high. At 60 mV/decade, that’s more than ten times the leakage.

A third symptom is threshold roll-off: the shorter the gate, the lower the threshold, so small variations in gate length become large variations in leakage.

What sets “too short”? Every transistor design has a , the distance over which the source’s and drain’s fields spread into the channel. It depends on how thin the region the gate controls is, how thin the gate insulator is, and how many sides the gate covers. A gate several natural lengths long behaves well; at around three natural lengths a flat bulk transistor is at its limit. So the way to keep shrinking is to shrink the natural length: make the channel thinner and wrap more gate around it.

In the off state the drain current is diffusion over the source-channel barrier, I∝exp⁡(−Eb/kT)I \propto \exp(-E_{\mathrm{b}}/kT). A long-channel gate moves EbE_{\mathrm{b}} by ΔVG/m\Delta V_{\mathrm{G}}/m, where m=1+Cdep/Coxm = 1 + C_{\mathrm{dep}}/C_{\mathrm{ox}} is the body factor, so

SS=dVGd(log⁡10ID)=m kTqln⁡10≈60 mV/dec×mat 300 K\begin{aligned} \mathrm{SS} &= \frac{dV_{\mathrm{G}}}{d(\log_{10} I_{\mathrm{D}})} = m\, \frac{kT}{q} \ln 10 \\ &\approx 60\,\mathrm{mV/dec} \times m \quad \text{at } 300\,\mathrm{K} \end{aligned}

with mm about 1.2–1.5 in conventional bulk devices. A fully depleted thin body with no depletion charge to speak of pushes mm toward 1, which is the first gift of SOI and multigate devices.

Short channels add a second loss. The channel potential obeys a 2D Poisson equation; reduced to one dimension along the channel, the deviation from the long-channel potential decays from each junction as exp⁡(−y/λ)\exp(-y/\lambda). The barrier peak sits where the two tails meet, so both its height and the gate’s leverage on it fall as the tails overlap:

  • Roll-off and DIBL. The peak is lowered by roughly 2Eb(Eb+qVDS) e−L/2λ2\sqrt{E_{\mathrm{b}}(E_{\mathrm{b}} + qV_{\mathrm{DS}})}\, e^{-L/2\lambda}. Its VDSV_{\mathrm{DS}} dependence is DIBL; its LL dependence is VTV_{\mathrm{T}} roll-off. DIBL falls exponentially with L/λL/\lambda.
  • Degraded SS. The source and drain pin part of the barrier, so ∂Epeak/∂Eb,gate\partial E_{\mathrm{peak}}/\partial E_{\mathrm{b,gate}} drops below 1 and SS rises above m×60 mV/decm \times 60\,\mathrm{mV/dec}.
  • Punch-through. When the tails overlap completely, no barrier is left and the gate cannot turn the device off.

For a single gate over a channel (or depletion region) of thickness tcht_{\mathrm{ch}} and oxide toxt_{\mathrm{ox}}, the classic approximation is

λ≈εSiεox tch tox\lambda \approx \sqrt{\frac{\varepsilon_{\mathrm{Si}}}{\varepsilon_{\mathrm{ox}}}\, t_{\mathrm{ch}}\, t_{\mathrm{ox}}}

and the minimum usable channel length is a multiple of it, about 3λ3\lambda for planar bulk. A second gate on the far side of a thin body shortens λ\lambda further, which shows up in the UC Berkeley design rules of thumb: tSit_{\mathrm{Si}} below Lg/4L_{\mathrm{g}}/4 for single-gate ultra-thin-body SOI and below 2Lg/32L_{\mathrm{g}}/3 for double-gate devices. Every architecture in the next section is a way to make tcht_{\mathrm{ch}} smaller or put more gate around it.

10 pA1 nA100 nA1 µA0.00.10.20.30.40.5gate voltage (V) →drain current per µm of width0.05 V drainSS 63 mV/decI_off 5.8 pA/µm
Gate length
Drain voltage

Long gate, drain at 0.7 V: swing 63 mV/decade, DIBL 20 mV/V, so the threshold drops 13 mV at full drain voltage. Off-current at 0 V gate: 5.8 pA/µm.

Drain current against gate voltage on a log scale, for an illustrative long and short gate. The slope shows the subthreshold swing; the sideways shift with drain voltage is DIBL.Share freely with credit: ‘Figure from chipfieldguide.com’

Flat. For most of history, transistors were flat. The channel was a thin layer at the top of the silicon, and the gate sat on top of it. The gate could only press from above, so leaking current could sneak underneath.

Fin. Researchers at UC Berkeley stood the channel up as a thin fin. Then they draped the gate over it, like a saddle on a horse. Now the gate grips from both sides and the top. This design is called a , and it went into factories in 2011.

Gate-all-around. The next step slices the channel into a stack of flat ribbons, often three, and fills the gaps with gate. Each ribbon is wrapped on every side. This is called a transistor, and the newest chips are moving to it.

Each step gives the gate a better grip. A better grip means the transistor can be made shorter before the drain takes over.

The optional chapter Fins and nanosheets looks closer at each shape and the names companies give them.

Planar: the gate on one side

In a planar transistor the gate sits on a thin insulator on top of the silicon. Current flows in a very thin layer at the surface, but the silicon below is part of the same crystal. As the source and drain got closer, charge could leak between them under the surface, where the gate’s field is weak. Chipmakers fought this with heavier doping, which slows electrons and makes transistors vary more from one to the next.

FinFET: the gate on three sides

A builds the channel as a fin of silicon a few nanometers wide and a few tens of nanometers tall, and wraps the gate over its two sides and top. In the ASAP7 kit, for example, fins are 6.5 nm thick and 32 nm tall. Because the fin is thin, no part of it is far from the gate. The first FinFETs were built at UC Berkeley in 1998 with gates down to 17 nm long, and it took about ten years to bring them into volume production. Intel switched to FinFETs for its 22 nm process in 2011; its fins were 8 nm wide.

Fins come with two limits:

  • Width comes in whole fins. All fins in a process have the same height, so a transistor’s strength is set by how many fins it uses: one, two, three. Designers can no longer size a transistor to an arbitrary width.
  • The bottom is still open. The fin is joined to the wafer at its base, where the gate doesn’t reach, so some leakage can still flow underneath. Processes add a heavily doped “punch-through stopper” there.

Gate-all-around nanosheets: the gate on every side

A transistor replaces the fin with a stack of horizontal silicon sheets, each surrounded by the gate on all sides. It’s built by growing alternating layers of silicon and silicon-germanium, then etching away the silicon-germanium with a chemical that leaves silicon alone. The silicon sheets are left as bridges between source and drain, and the gate materials are deposited into the gaps one atomic layer at a time.

Because the sheets lie flat, their width is drawn in the layout, so designers get back a continuous choice: wide sheets for drive current, narrow ones for low power. IBM showed stacks of three sheets from 8 to 50 nm wide.

The optional deep dive Fins and nanosheets covers these devices in more detail: effective width, supply voltage, how they are made, and each company’s name for them.

Planar bulk and the doping trap

In planar bulk the “channel thickness” in λ\lambda is effectively the depletion depth, which can only be cut by raising the channel doping. That raises mm (worse SS), lowers mobility and, because a small gate covers few dopant atoms, adds random dopant fluctuation to VTV_{\mathrm{T}}. Ultra-thin-body fixed the geometry directly (the body is physically thin, so it can be undoped), and multigate devices put a second gate on the far side of that body.

FinFET design

The vertical thin-body idea goes back to Hitachi’s DELTA transistor, published in 1990; Berkeley’s 1998 FinFET made it a self-aligned double gate that a planar-style process could build.

A FinFET is a double gate (two sidewalls) or tri-gate (plus the top) device on a body of width WfinW_{\mathrm{fin}}. The fin width sets DIBL; Berkeley’s data put the requirement at Lg/WfinL_{\mathrm{g}}/W_{\mathrm{fin}} above about 1.5. For DIBL of 100 mV/V, published design curves need WSi≈Lg/2W_{\mathrm{Si}} \approx L_{\mathrm{g}}/2 for a FinFET versus HSi≈Lg/5H_{\mathrm{Si}} \approx L_{\mathrm{g}}/5 for single-gate UTB SOI; tri-gates relax both somewhat. Other consequences:

  • Quantized width. Effective width per fin is about 2Hfin+Wfin2H_{\mathrm{fin}} + W_{\mathrm{fin}}, so drive comes in fin increments and fin height becomes a process-wide trade between layout efficiency and design flexibility. ASAP7’s 7.5-track cell allows at most three fins per nMOS and pMOS device.
  • Sub-fin leakage. On bulk wafers the fin’s base, below the gated region, needs a punch-through stopper (super-steep retrograde well).
  • VTV_{\mathrm{T}} by work function. With undoped fins, threshold flavors come from gate work-function engineering rather than channel implants. ASAP7’s four flavors (SLVT, LVT, RVT, SRAM) share one drawn gate length and differ mainly in work function; the SRAM device also drops the lightly doped drain implant.
  • Variability. Performance becomes sensitive to fin width, and work-function variation replaces dopant fluctuation as the main VTV_{\mathrm{T}} variation source in undoped channels.
  • Parasitics. Gate-to-source/drain fringing capacitance and series resistance in the narrow fin grow in importance as LgL_{\mathrm{g}} shrinks.

Nanosheets

Stacked nanosheets go further on both λ\lambda knobs: the sheet thickness is set by growing the Si/SiGe layers rather than by etching a fin, and the gate covers the bottom face, closing the sub-fin path that a FinFET leaves connected to the body. The width becomes a layout parameter again, but the total stack height, the sheet count, the inner spacers between gate and source/drain and the n-p spacing in the cell now bound what you can draw. Variants trim the n-p spacing: the puts a dielectric wall between the n and p sheet stacks, and imec expects it to carry the nanosheet roadmap to around its A10 node.

Fins and nanosheets goes further on the devices: width per footprint, capacitance, self-heating, the inner-spacer and release modules, and how brand names map to these types.

STIgateFinFETgate: 3 sidesgated width141 nm= 2 × 70.5 nm10 nmcurrent flowsinto the page
Architecture
Fins per transistor

FinFET: gate on three sides. Width comes in whole fins, about 70.5 nm of channel each (2 × 32 + 6.5). Tap a part.

Planar, FinFET and nanosheet cross-sections at a common scale, with ASAP7 fin dimensions. Tap a part; change the width to compare whole fins with drawn sheets. Sheet thickness and spacing are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Once the gate wraps all the way around, there’s no more grip to gain. So the newest ideas save space in other ways.

Stacking. Every logic circuit uses two kinds of transistor, normally placed side by side. A stacks one right on top of the other, like bunk beds. That could nearly double how many fit in the same space. The three biggest chipmakers showed working ones in 2023. They may reach products in about seven to ten years.

Power from below. A chip has a dozen or more layers of wiring above its transistors. Some wires carry power, and the rest carry signals, the 1s and 0s. Until now, they all shared that space. With , the power wires move to the back of the chip. Signals get the whole top to themselves. Power arrives by a shorter, wider path that wastes less.

CFET: stacking n on p

A CMOS logic gate always pairs an nMOS transistor (on when the gate is high) with a pMOS transistor (on when it’s low); the CMOS logic chapter explains why. In a standard cell they sit in two rows side by side, separated by a gap. A (complementary FET) builds the nMOS and pMOS nanosheet stacks on top of each other in one structure. That removes the n-p gap from the cell’s height for the first time.

At the 2023 IEEE International Electron Devices Meeting (IEDM), Intel, Samsung and TSMC all reported CFETs. Intel built an inverter with gates spaced 60 nm apart and contacted the bottom transistor from below the silicon; Samsung and TSMC reached 48 nm spacing. Experts expected CFETs in products in seven to ten years. imec’s roadmap places its monolithic CFET at its A7 node.

Backside power delivery

Power reaches every transistor through the same stack of wiring layers that carries signals. In advanced processes, power wiring uses at least 20% of the routing resources, and the thin lower wires are resistive, so the supply sags on its way down (IR drop). builds the power network on the other side of the wafer:

  1. Build the transistors and the signal wiring on the front as usual.
  2. Bond the front side to a carrier wafer and flip it over.
  3. Grind and polish the original wafer down from the back until only a thin layer remains below the transistors.
  4. Connect power to the transistor level through tiny, deep, metal-filled holes called nano through-silicon vias (Intel drills them from the front before bonding and exposes them by thinning; imec etches them from the back after thinning), then build thick power wires on the back.

In Intel’s PowerVia test chip, this gave more than a 6% higher clock frequency, 30% less power lost in delivery, and areas of logic packed up to 95% full, because signal wires were no longer competing with power. imec measured seven times lower IR drop in its test structures. The costs: extra process steps, a changed path for heat to escape, and new methods to debug chips. The Power planning stage shows how designers lay out a power grid.

CFET

A stacks the n and p nanosheet stacks vertically. Two integration routes compete: monolithic, where top and bottom channels come from one epitaxial superlattice and are processed together, and sequential, where the top layer is bonded onto a finished bottom tier. Either way the process has to isolate the top device from the bottom one, give each its own gate work function, and reach the bottom source/drain without going through the top device. imec’s answer is to contact the bottom pFET directly from the backside; its IEDM 2025 work combines a frontside-contacted (100) nFET on top with a backside-contacted (110) pFET below, choosing crystal orientations that suit each carrier type. The density case is cell height: with n-p spacing gone, imec targets cells of about four metal tracks. The 2023 IEDM demonstrations (Intel at 60 nm CPP with three sheets per device and a backside contact; Samsung at 48 and 45 nm CPP; TSMC at 48 nm with a high-germanium SiGe isolation layer) show the pieces working, but not yet at product yield.

Backside power

Frontside power delivery spends cell height on wide VDDV_{\mathrm{DD}}/VSSV_{\mathrm{SS}} rails and at least a fifth of the routing on the power grid, while the resistance of thin lower-metal vias and wires grows each node. imec’s flow uses buried power rails (metal lines sunk into the front-end, below the devices) connected by nano-TSVs to a backside network after carrier bonding and thinning to a SiGe etch stop; the test vehicle cut IR drop 7×. Intel’s PowerVia drills its nano-TSVs from the front and exposes them by thinning; its test core reported more than 6% frequency gain, 30% lower delivery loss and local cell utilization up to 95%, with new thermal design and debug methods required. For CFET, backside access stops being optional: the bottom device’s source/drain is reached from below.

For physical design this changes the cell template (no frontside rails), the PDN (no power straps competing with signals in the BEOL), IR-drop and electromigration signoff, and debug, which needed new methods in Intel’s test chips. See the Power planning and Routing stages.

wiringdevicessiliconsupply gridM1M2M3M4pMOSnMOScell height7.5 trackstracks for signals≤ 80%signalpower
n and p devices
Power delivery

Side by side: nMOS and pMOS rows with a gap between (a 7.5-track FinFET cell, as in ASAP7). Front-side power: rails and straps take at least 20% of the routing.

A standard cell in cross-section: transistors at the bottom, wiring above. Try stacking n on p (CFET) and moving power to the back side. Schematic; track counts from ASAP7 and imec.Share freely with credit: ‘Figure from chipfieldguide.com’

Chipmakers label each generation with a size, like “7 nanometer” or “2 nanometer.” A nanometer is a millionth of a millimeter. Long ago, that number really was the length of the transistor’s gate. Since the mid-1990s, the names and the real sizes have drifted apart.

Today a name is a brand for a generation, a bit like a phone’s model number. In one “7 nm” process, the gates are about 20 nm long and about 50 nm apart. The name isn’t the gate length or the spacing of anything. (The fins happen to be about that thin, but that’s not where the name comes from.) A smaller name still means a newer, denser process. But you can’t compare two companies by their names.

A name once matched the transistor’s gate length. In the mid-1990s, chipmakers began shrinking gate length faster than everything else to gain speed, so the two uncoupled: the “130 nm” node had 70 nm gates. Later the reverse happened. Short-channel effects made gate length hard to shrink, while density kept improving, so the name kept falling faster than the gate: Intel’s “22 nm” FinFETs of 2011 had 26 nm gates.

What actually sets how many transistors fit in a square millimeter is the spacing of things, not the gate length:

  • the (CPP), the center-to-center spacing of gates, which sets the width of a logic cell;
  • the of the tightest wires, times the number of wiring “tracks” a cell is tall, which sets its height;
  • and, from CFET onward, how many layers of devices are stacked.

Paolo Gargini, chairman of the IEEE International Roadmap for Devices and Systems (IRDS), proposed naming processes by exactly these three numbers, and the IRDS roadmap uses them. A “5 nm” process with a 48 nm gate pitch, a 36 nm metal pitch and one layer of devices would be G48M36T1. Another proposal, the LMC metric, states the density of logic, of main memory and of the connections between them directly.

Node names were once both the gate length and the metal half-pitch, which for a long time were about the same number, and each generation cut features by about 30% to double density. From the mid-1990s gate length was scaled ahead of pitch for performance (the 130 nm node shipped 70 nm gates), and in the FinFET era the opposite: electrostatics held LgL_{\mathrm{g}} near 20–26 nm while the label kept falling. Intel’s 22 nm FinFETs (2011) had 26 nm gates, a 40 nm half-pitch and 8 nm fins. ASAP7’s “7 nm” has a 21 nm physical gate on a 54 nm CPP and 36 nm M1–M3 pitch; the authors deliberately relaxed CPP, noting that a larger pitch eases pin access and can give better logic density at iso-power and performance.

To first order, cell area = CPP × (cell width in gate pitches) × Mx\mathrm{M}_x pitch × (track height), divided by device tiers. That’s why IRDS’s G-M-T notation (G48M36T1 for “5 nm”) captures density where the node name can’t, and why much of each generation’s density gain now comes from moves that reduce track height or gate pitches per cell (fewer fins per device and fewer tracks, tighter diffusion breaks, buried or backside power rails, stacked devices) rather than from LgL_{\mathrm{g}}. The alternative “LMC” metric reports the density of logic, of main memory and of the interconnect linking them.

“7 nm” (ASAP7, predictive)0 nm50 nm100 nm150 nmnode name7 nmgate length21 nmgate pitch (CPP)54 nmmetal pitch (M1–M3)36 nmfin width6.5 nm
Node

In ASAP7’s “7 nm”, the gate is 21 nm and gates are 54 nm apart. Density is set by the pitches, not the name.

Node names against measured or published dimensions, drawn to one scale (1 nm = 3 px). 7 nm values are from the predictive ASAP7 kit.Share freely with credit: ‘Figure from chipfieldguide.com’

Pick a transistor shape, then drag the slider to shrink its gate. The left picture shows how much of the channel the gate wraps. The right picture shows the hill electrons must climb. Watch the drain pull the hill down as you shrink. Which shape can you shrink the furthest before it starts to leak?

Choose planar, FinFET or gate-all-around, then shorten the gate. The plot shows the energy barrier from source to drain with the transistor off, at a low drain voltage (dashed) and at the full 0.7 V supply (solid). The readouts show subthreshold swing, DIBL and off-current. Find the shortest gate each shape can manage while staying under 90 mV/decade and 100 mV/V.

A quasi-2D natural-length model: λ\lambda from tcht_{\mathrm{ch}}, EOT and an effective gate count NN per architecture; the off-state conduction band Ec(y)E_{\mathrm{c}}(y) from the 1D screened Poisson solution; SS from the gate’s leverage on the barrier peak, DIBL from the peak’s drain dependence, IoffI_{\mathrm{off}} from the peak height. The FinFET case is calibrated so that at L=21 nmL = 21\,\mathrm{nm} it lands near ASAP7’s predictive RVT nFET. The full model is shown below the plot.

Loading simulation…

These numbers come from a free practice kit that researchers built to show what a “7 nm” process looks like. No factory uses it, but it’s realistic, and anyone can download it. A nanometer (nm) is a millionth of a millimeter. A volt (V) measures the electric push.

Gate length in a “7 nm” chip
21 nm
Space from one gate to the next
54 nm
Fin thickness
6.5 nm
Power supply
0.7 V

What they mean:

  • The gate is three times longer than the “7” in the name. The space between gates is nearly eight times the “7.”
  • The fin is only a couple of dozen atoms thick. That thinness is what lets the gate hold the channel firmly.
  • The power supply is less than half the push of one AA battery, which gives 1.5 volts.
  • The kit offers four versions of each transistor, from fast and leaky to slow and tight. Each step down leaks about ten times less.

All of these are predictive numbers from the ASAP7 kit, an open 7 nm FinFET process design kit from Arizona State University and Arm Research, made for teaching and research and not for manufacturing.

ASAP7 dimension (predictive)Value
Fin height × thickness32 nm × 6.5 nm
Fin pitch27 nm
Gate length (physical / drawn)21 nm / 20 nm
Contacted poly pitch54 nm
M1–M3 pitch (EUV)36 nm
Standard-cell height7.5 tracks, up to 3 fins per transistor
Nominal supply0.7 V

ASAP7 is predictive: its assumptions were derived from published 14–10 nm data and trends, it isn’t tied to any foundry, and its authors note it will be inaccurate in some details. Its electrostatic targets were explicit: SS approaching the 60 mV/dec limit and DIBL around 30 mV/V, “following the best to date FinFET published result,” both eased by the relaxed 54 nm CPP allowing a longer gate.

ASAP7 nFET, TT, 25 °C, per finSRAMRVTLVTSLVT
IdsatI_{\mathrm{dsat}} (µA)28.637.945.250.8
IoffI_{\mathrm{off}} (nA)0.0010.0190.2422.444
VT,sat/VT,linV_{\mathrm{T,sat}} / V_{\mathrm{T,lin}} (V)0.25 / 0.270.17 / 0.190.10 / 0.120.04 / 0.06
SS (mV/dec)62.463.062.963.3
DIBL (mV/V)19.221.322.322.6
pFET SS / DIBL64.3 / 24.164.5 / 30.464.4 / 31.164.9 / 31.8

Things to notice:

  • The subthreshold swing of every device is within about 4–8% of the 60 mV/decade physical limit.
  • DIBL is about 20–32 mV/V: at the full 0.7 V supply, the drain lowers the threshold by only 13–22 mV.
  • Choosing a lower-threshold device buys more current at the cost of about ten times more leakage per step: from 1 pA per fin for the SRAM device to 2.4 nA for the fastest one.

Fins 32 nm tall and 6.5 nm thick on a 27 nm pitch; LgL_{\mathrm{g}} 21 nm (20 nm drawn); CPP 54 nm; M1–M3 36 nm pitch, single-patterned EUV; M4–M5 48 nm and M6–M7 64 nm with self-aligned double patterning; VDDV_{\mathrm{DD}} 0.7 V.

Points worth noticing in the table:

  • IoffI_{\mathrm{off}} rises 10–19× per VTV_{\mathrm{T}} flavor step while VT,satV_{\mathrm{T,sat}} drops 60–80 mV, which is what an SS near 63 mV/dec predicts. IeffI_{\mathrm{eff}}, the inverter-relevant drive, is roughly half of IdsatI_{\mathrm{dsat}}.
  • The VT,lin−VT,satV_{\mathrm{T,lin}} - V_{\mathrm{T,sat}} difference (0.02 V for the RVT nFET) is consistent with DIBL×(0.7−0.05) V≈14 mV\mathrm{DIBL} \times (0.7 - 0.05)\,\mathrm{V} \approx 14\,\mathrm{mV}, given the table’s 10 mV rounding.
  • PMOS drive is about 90% of NMOS, an assumption the authors base on strain trends from 32 nm planar to 14 nm FinFETs.
  • The gear ratio between the 27 nm fin pitch and the 36 nm M2 pitch gives whole numbers of fins at 6, 7.5, 9 and 12 tracks; ASAP7 chose 7.5 tracks (270 nm) with up to three fins per device.

For cell layouts built on these rules, see From devices to a cell library; for how the 27 nm fins and 36 nm EUV wires are printed, see Making them.

Every new transistor shape fixes one problem and brings new ones.

  • Harder to make. Fins and stacked ribbons need many more factory steps than flat transistors. Some layers are only a few atoms thick. More steps mean more cost and more ways to fail.
  • Less freedom. A fin transistor can only be one, two or three fins wide. So designers can’t fine-tune its strength.
  • Heat and testing. Moving power wires under the chip changes how heat escapes. It also changes how engineers look inside a chip to find faults.
  • Shrinking costs more. Printing ever-finer patterns needs extra steps and hugely expensive machines. Cost is now the main thing slowing shrinking down. So not every chip needs the newest generation.
  • Gate control vs capacitance and resistance. Wrapping the gate around a thin channel improves control but adds gate area next to the source and drain contacts (extra capacitance to charge every switch) and squeezes current through narrow, resistive fins or sheets.
  • Discrete sizing. FinFET widths come in whole fins, which complicates circuits that need exact ratios, such as memory cells. Nanosheets restore continuous widths within limits.
  • Threshold vs leakage. Even a near-ideal 60 mV/decade switch gains ten times more leakage for every 60 mV of lower threshold, so processes offer several threshold “flavors” and designers mix them.
  • Variation. A FinFET’s performance is very sensitive to the width of its fin, which is only a few nanometers, so tiny manufacturing differences show up as differences between transistors.
  • Cost. Once ordinary lithography reached its limits and needed several exposures per layer, cost became the main obstacle to scaling; extreme-ultraviolet lithography, new architectures and backside processing all add steps. Not every design needs the newest node.
  • Electrostatics vs parasitics. Multigate devices win SS and DIBL but pay in gate-to-source/drain fringing capacitance and in series resistance through narrow fins, and series resistance matters more as LgL_{\mathrm{g}} shrinks.
  • Relaxed pitch vs density. ASAP7 deliberately chose a less aggressive CPP for low power and short-channel control, citing DTCO results where a larger transistor pitch improved pin access and gave better logic density at iso-power and performance.
  • Width quantization vs SRAM stability. FinFET SRAM cells come in fin-count ratios (1-2-3, 1-1-2, 1-1-1); the densest 1-1-1 cell needs read and write assist circuits. See Memory cells.
  • Variability. Undoped channels remove random dopant fluctuation but expose fin-width and work-function variation; below about 10 nm, source/drain dopant gradients trade performance against variability.
  • Backside: wins and new failure modes. Backside power frees routing and cuts IR drop, but adds wafer bonding and thinning, changes the thermal path, and required new debug methods.
  • CFET integration risk. Top/bottom isolation, separate n and p gate stacks and bottom-device contacts all have to work in one stack at product yield; in 2023, estimates put production 7–10 years out.
  • The model’s limit. Scale-length scaling has a floor: for atomically thin channels, published simulations find λ\lambda set by the physical oxide thickness, not by the channel.
DriveI_dsat (µA)38050100150LeakageI_off (log)1 pA10 pA100 pA1 nA10 nAone flavor faster ≈ 10× leakage19 pA
Threshold flavor
Fins

RVT, 1 fin: 37.9 µA of drive, 19 pA of leakage. Each faster flavor costs roughly 10× the leakage; width only comes in whole fins.

ASAP7 nFET per-fin values (typical corner, 25 °C), times the fin count. The ticks under the drive bar are every strength the four flavors and one to three fins allow.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

The natural-length model, step by step

Take the off state of an n-channel device with source at y=0y = 0 and drain at y=Ly = L. Write the channel potential as the long-channel value set by the gate plus a correction that satisfies the source and drain boundary conditions. Gauss’s law on a thin slice of channel, with the vertical field approximated as linear through the body and oxide, turns the 2D Poisson equation into

d2ψdy2−ψ−ψlongλ2=0\frac{d^2\psi}{dy^2} - \frac{\psi - \psi_{\mathrm{long}}}{\lambda^2} = 0

whose solution decays from each junction as sinh⁡(y/λ)\sinh(y/\lambda). This is the screening-length picture behind the exp⁡(−L/λ)\exp(-L/\lambda)-type dependence of DIBL. In energy terms, with EbE_{\mathrm{b}} the long-channel barrier and the drain at −qVD-qV_{\mathrm{D}}:

Ec(y)=Eb−Ebsinh⁡((L−y)/λ)sinh⁡(L/λ)−(Eb+qVD)sinh⁡(y/λ)sinh⁡(L/λ)\begin{aligned} E_{\mathrm{c}}(y) &= E_{\mathrm{b}} - E_{\mathrm{b}} \frac{\sinh\big((L - y)/\lambda\big)}{\sinh(L/\lambda)} \\ &\quad - (E_{\mathrm{b}} + qV_{\mathrm{D}}) \frac{\sinh(y/\lambda)}{\sinh(L/\lambda)} \end{aligned}
0.40.0-0.4E_c (eV)0λ2λ4λ6λ8λ10λposition along the channel, y →drain, −qV_DsourceE_b (long channel)ΔE 192 meVsource taildrain tailE_c(y), both
Drain voltage V_D

L/λ = 4.0, V_D = 0.7 V: peak at y = 0.38L, lowered by 192 meV (approximation 198 meV). DIBL ≈ 98 mV/V.

The two sinh tails from the source and drain, subtracted from the long-channel barrier E_b = 0.46 eV (the sim’s FinFET case, gate at 0 V). Distance in units of λ. Shorten L/λ and watch the peak sink; the approximation holds for L ≫ λ.Share freely with credit: ‘Figure from chipfieldguide.com’

The barrier peak EpeakE_{\mathrm{peak}} sets the off-current, Ioff∝exp⁡(−Epeak/kT)I_{\mathrm{off}} \propto \exp(-E_{\mathrm{peak}}/kT). Three figures of merit follow directly:

  1. SS =60 mV/dec⋅m/g= 60\,\mathrm{mV/dec} \cdot m/g, with g=∂Epeak/∂Ebg = \partial E_{\mathrm{peak}}/\partial E_{\mathrm{b}}. Long channel: g=1g = 1. Short channel: the junctions pin part of the barrier and g<1g < 1.
  2. DIBL =[Epeak(VD,lin)−Epeak(VDD)]/(VDD−VD,lin)= \big[E_{\mathrm{peak}}(V_{\mathrm{D,lin}}) - E_{\mathrm{peak}}(V_{\mathrm{DD}})\big] / (V_{\mathrm{DD}} - V_{\mathrm{D,lin}}).
  3. IoffI_{\mathrm{off}} =I0⋅10−(Epeak−Eth)/60 mV= I_0 \cdot 10^{-(E_{\mathrm{peak}} - E_{\mathrm{th}})/60\,\mathrm{mV}}, where EthE_{\mathrm{th}} is the barrier at threshold and I0I_0 the constant-current threshold definition.

For L≫λL \gg \lambda the peak is near the middle and the lowering is approximately

ΔE≈2Eb(Eb+qVD) exp⁡(−L/2λ)\Delta E \approx 2\sqrt{E_{\mathrm{b}}(E_{\mathrm{b}} + qV_{\mathrm{D}})}\, \exp(-L/2\lambda)

so roll-off, DIBL and SS degradation all decay as exp⁡(−L/2λ)\exp(-L/2\lambda). Doubling L/λL/\lambda squares the suppression factor. In this model, for a given body factor, the shortest usable LL is roughly a fixed multiple of λ\lambda (about 4λ4\lambda for FinFET and GAA here; planar needs about 5λ5\lambda because its m=1.25m = 1.25 already costs SS), so an architecture that cuts λ\lambda by 30% allows a 30% shorter gate, about one traditional generation.

λ\lambda for each architecture

For a single gate, λ≈εSi tch tox/εox\lambda \approx \sqrt{\varepsilon_{\mathrm{Si}}\, t_{\mathrm{ch}}\, t_{\mathrm{ox}} / \varepsilon_{\mathrm{ox}}} with tcht_{\mathrm{ch}} the body or depletion thickness. The sim extends it to multigate devices in the simplest way, by dividing the term under the root by an effective number of gated sides NN. That is a modeling choice for illustration, not a derived result. Its inputs (εSi/εox≈3\varepsilon_{\mathrm{Si}}/\varepsilon_{\mathrm{ox}} \approx 3):

Architecturetcht_{\mathrm{ch}}EOTNNmmλ\lambdaShortest usable LL in the sim
Planar bulk10 nm (depletion depth)1.0 nm11.255.5 nmabout 29 nm
FinFET6.5 nm (fin width, ASAP7)0.9 nm21.03.0 nmabout 12 nm
GAA nanosheet5 nm (sheet)0.9 nm31.02.1 nmabout 9 nm

“Usable” here means SS≤90 mV/dec\mathrm{SS} \le 90\,\mathrm{mV/dec} and DIBL≤100 mV/V\mathrm{DIBL} \le 100\,\mathrm{mV/V}. The 100 mV/V DIBL figure is the target used in published multigate design curves. The model is optimistic about absolute lengths (real processes keep margin for variation, parasitic resistance and leakage targets), but the ordering and the ratio are the point. At L=21 nmL = 21\,\mathrm{nm} the FinFET case gives SS≈64 mV/dec\mathrm{SS} \approx 64\,\mathrm{mV/dec} and DIBL≈22 mV/V\mathrm{DIBL} \approx 22\,\mathrm{mV/V}, close to ASAP7’s predictive RVT nFET (63.0 mV/dec, 21.3 mV/V).

Assumptions and what the model leaves out

  • One-dimensional reduction: vertical fields linear through body and oxide; fringing fields through the spacers ignored. Quantitatively wrong for LL below about 3λ3\lambda, which the sim flags.
  • Boltzmann statistics and drift-diffusion: no source-to-drain tunneling (significant below about 10 nm), no quantum confinement shift of VTV_{\mathrm{T}} in very thin bodies.
  • The bulk body factor mm is a constant, not a function of doping and bias; NN is an effective number, not derived from a 3D solution.
  • Off-state only: no velocity saturation, series resistance or mobility, so it says nothing about IonI_{\mathrm{on}}. The I-V curve chapter covers the on state.
  • At atomic-scale channel thickness, λ\lambda saturates at a value set by the physical oxide thickness.

How production models handle it

Compact models used in real kits fold the same physics into fitted parameters. ASAP7’s transistors use BSIM-CMG, a compact model for multigate devices, with parameters derived from public data and historical trends. Compact models fit short-channel behavior with parameters extracted from measurements or device simulation, so they track SS and DIBL across gate lengths far more closely than a hand formula like this one.

Novice · 0 of 4 correct
  1. Q1Under classic Dennard scaling, shrinking every dimension and the voltage by the same factor kept which quantity constant?

  2. Q2A transistor has a subthreshold swing of 90 mV/decade. Roughly how much does its off-current change if its threshold voltage is lowered by 180 mV?

  3. Q3What does a DIBL of 100 mV/V mean?

  4. Q4Why is a FinFET’s width “quantized” while a nanosheet’s is not?

Sources

Show Hide 14 sources
  1. Cramming More Components onto Integrated CircuitsGordon E. Moore · Electronics 38(8), reprinted with Intel’s permission as Appendix C of “The Future of Computing Performance” (National Academies Press, 2011) · 1965Cost per component falls as more components fit on one chip, until yield losses raise it; the complexity at minimum component cost had increased by roughly a factor of two per year.
  2. Design of Ion-Implanted MOSFET’s with Very Small Physical DimensionsRobert H. Dennard, Fritz H. Gaensslen, Hwa-Nien Yu, V. Leo Rideout, Ernest Bassous, Andre R. LeBlanc · IEEE Journal of Solid-State Circuits SC-9(5), reprinted with permission of IEEE and the author as Appendix D of “The Future of Computing Performance” (National Academies Press, 2011) · 1974Table I: scaling dimensions and voltage by 1/κ cuts delay by κ and power per circuit by κ², keeping power density constant. The subthreshold slope does not scale, and merging source and drain depletion regions cost threshold control.
  3. Correlated Electron Materials and Field Effect Transistors for Logic: A ReviewYou Zhou, Shriram Ramanathan · arXiv:1212.2684 (Harvard University) · 2012Section II: subthreshold swing = (1 + C_d/C_ox)·ln(10)·kT/q; body factor 1.2–1.5 in conventional MOSFETs; about 60 mV/decade minimum at room temperature for any MOSFET.
  4. FinFET: History, Fundamentals and Future (2012 Symposium on VLSI Technology short course)Tsu-Jae King Liu · University of California, Berkeley · 2012Gate coupling and DIBL; thin-body rules (UTB t_Si below L_g/4, double gate below 2L_g/3); Hitachi’s DELTA vertical thin-body MOSFET (published 1990); FinFET L_g/W_fin above 1.5 to suppress DIBL; first FinFETs in 1998 down to 17 nm; quantized width; punch-through stopper in bulk FinFETs; ~10 years to volume production; gate length scaling slower than pitch.
  5. Scaling Theory of Two-Dimensional Field Effect TransistorsSaurabh V. Suryavanshi, Chris D. English, H.-S. Philip Wong, Eric Pop · arXiv:2105.10791 (Stanford University) · 2021Minimum channel length is a multiple of the scale length Λ (about 3Λ for planar bulk); short-channel effects (DIBL, higher SS, V_T roll-off); thinner channels give shorter Λ, which drove SOI, FinFETs, surround gates and nanosheets; Λ ≈ √(ε_ch·t_ch·t_ox/ε_ox); DIBL falls exponentially with L/Λ.
  6. ASAP7: A 7-nm finFET predictive process design kitLawrence T. Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, Greg Yeric · Microelectronics Journal 53 (open access, CC BY-NC-ND), copy in the ASAP7 PDK repository · 2016Predictive, not tied to a foundry. Fins 32 nm tall, 6.5 nm thick, 27 nm pitch; CPP 54 nm; L_g 21 nm; M1–M3 pitch 36 nm (EUV); 7.5-track cells with up to 3 fins per device; VDD 0.7 V; four V_t flavors; Tables 3–4 per-fin I_dsat, I_off, SS and DIBL; limited V_DD scaling from 14 nm to 7 nm.
  7. ASAP7 predictive PDK and standard-cell librariesArizona State University and Arm Research · The OpenROAD Project (GitHub)A 7-nm FinFET predictive process design kit and cell libraries for research and teaching, BSD-3-Clause licensed, not for manufacturing.
  8. The Nanosheet Transistor Is the Next (and Maybe Last) Step in Moore’s LawPeide D. Ye, Thomas Ernst, Mukesh V. Khare · IEEE Spectrum · 2019Why planar devices leaked; the FinFET gate on three sides; fin height can’t vary freely and the fin bottom stays connected to the body; nanosheets wrapped on all sides with adjustable width; built from a Si/SiGe superlattice with the SiGe etched away.
  9. Intel, Samsung, and TSMC Demo 3D-Stacked TransistorsSamuel K. Moore · IEEE Spectrum · 2023IEDM 2023 CFET demonstrations: nFET and pFET stacked in one structure, close to double the density; Intel inverter at 60 nm CPP with a backside contact; Samsung and TSMC at 48 nm; commercial use expected in 7–10 years.
  10. Imec puts complementary FET (CFET) on the logic technology roadmapimec (Julien Ryckaert, Naoto Horiguchi) · imec · 2022In a CFET the n- and pMOS devices are stacked on top of each other; monolithic and sequential integration compared; aimed at 4-track cells beyond the 1 nm node.
  11. Performance boosters to scale monolithic CFET across multiple logic technology nodesSheng Yang, Anne Vandooren, Geert Hellings, Naoto Horiguchi · imec · 2026Stacking removes the n-p separation from cell height; the forksheet extends nanosheets to the A10 node and monolithic CFET is expected at A7; bottom pFET contacted directly from the backside.
  12. Backside power deliveryNaoto Horiguchi, Eric Beyne · imec · 2022Power wiring takes at least 20% of routing resources; buried power rails and nano-TSVs; bonding to a carrier wafer and thinning; 7× lower IR drop in a test vehicle.
  13. Intel Is All-In on Backside Power DeliverySamuel K. Moore · IEEE Spectrum · 2023PowerVia test chip: more than 6% higher frequency, 30% less power loss, cell areas up to 95% filled; nano-TSVs, carrier wafer and thinning; new thermal and debug work; planned with RibbonFET, Intel’s nanosheet (gate-all-around) transistor.
  14. A Better Way to Measure Progress in SemiconductorsSamuel K. Moore · IEEE Spectrum · 2020Node names and gate length uncoupled in the mid-1990s (the 130-nm node had 70-nm gates); Intel’s 22-nm FinFETs in 2011 had 26-nm gates, a 40-nm half-pitch and 8-nm fins; IRDS’s G-M-T metric (5 nm ≈ G48M36T1) and the LMC metric (densities of logic, main memory and the connections between them).