Every chip that computes also has to remember. It keeps the numbers it is working on, the steps it is following and the photo you just took. That memory is built from the same tiny switches as the rest of the chip. They are grouped into , and each cell holds one bit: a 1 or a 0.
There are three main kinds of memory cell, and each one uses a different trick:
- SRAM keeps a bit in a loop of switches that hold each other in place. It is fast, but each cell is big.
- DRAM keeps a bit as a tiny bucket of electric charge that slowly leaks. It fits far more bits in the same space, but it is slower.
- Flash locks electrons behind a wall, so the bit stays even with the power off. It fits the most bits of all, but saving to it is slow.
A memory is a grid of . Each row shares a that selects it, and each column shares a that carries data in and out. Around the grid sit decoders that turn an address into one active word line, and column circuits that write the bit lines or sense them.1
Three cell designs cover almost all memory in use today:
| Cell | How it stores a bit | Parts per bit | Keeps data without power? | Typical home |
|---|---|---|---|---|
| Two inverters in a loop that hold each other’s value | 6 transistors | No | Caches, registers and buffers on the processor chip | |
| Charge on a capacitor, reached through one transistor | 1 transistor + 1 capacitor | No, and it leaks even with power on | Main memory, on separate chips or stacked next to the processor | |
| Electrons trapped in an insulated layer inside a transistor | 1 transistor, often holding 2–4 bits | Yes | Storage: SSDs, phones, memory cards |
The order of that table is also the order of speed (fastest first) and the reverse order of density. Memory closer to the processor is smaller and faster; memory farther away is larger, slower and cheaper per bit.2 SRAM and DRAM are : they lose their data when power is removed.
This chapter builds on two earlier ones: the transistor as a voltage-controlled switch, and the CMOS inverter and its transfer curve. An SRAM cell is, at heart, two inverters wired back to back.
Memory cells are where device physics meets system architecture most directly. The cell’s storage mechanism fixes its area per bit, its access time, its energy per access and whether it needs power or refresh to keep its data, and those four numbers set the shape of the above it.2
- 6T is a bistable latch built from logic transistors. It lives on the logic die, reads non-destructively and is fast, but six devices per bit make it the least dense. Its design problem is ratioed: the same access transistors must be weak enough not to disturb the cell on a read and strong enough to overpower it on a write.1
- 1T1C stores charge on a dedicated capacitor. An SRAM cell takes roughly 20 times its silicon area per bit: about 120–140 against about 6 , where is the process’s smallest feature.9 But a DRAM read is destructive, its signal is a small fraction of the supply after , and it must be refreshed every 64 ms.11
- NAND stores charge in a floating gate or charge trap isolated by oxide, shifting the cell’s threshold voltage. It keeps data unpowered and stores several bits per cell, at the cost of slow, high-voltage programming, block erase and limited endurance.14
The emphasis here is the circuit-level detail that decides whether a cell works at all: read stability, writability, the and the for SRAM; the sensing budget and refresh for DRAM; and the charge-storage and wear mechanisms of flash. Numbers come from the open SKY130 process where possible.
Power on: each cell holds a bit. Switch the power off to see which ones are volatile. Tap a cell for details.
Start with an , the simplest circuit on a chip. It flips whatever goes in: put in a 1, get out a 0. Now connect two inverters in a circle, so each one’s output feeds the other’s input.
Say the left side holds 1. Then the right side must hold 0. That 0 feeds back into the left side and keeps it at 1. The loop locks itself in place. It works just as well the other way round. So the loop can remember one bit, as long as the power stays on.8
To use the loop, add one more switch on each side. Think of these as doors. Two outside wires, called , run past the cell. A control wire called the opens both doors at once. Four switches in the loop plus two doors makes six. That’s why this is called a six-transistor cell.
- Writing: set one bit line to 1 and the other to 0, then open the doors. The bit lines push harder than the loop can, so the loop flips to the new value.
- Reading: fill both bit lines with charge, so both read 1, then open the doors. The side holding 0 drains its bit line a little. A sensitive detector notices which wire dropped.2
Reading is the delicate part. The cell must not flip while you peek at it.
An is one PMOS transistor connecting the output to the supply () and one NMOS transistor connecting it to ground; a high input turns the NMOS on and pulls the output low, a low input turns the PMOS on and pulls it high. Wire two inverters in a ring and you have a latch with two stable states, (Q = 1, QB = 0) and (Q = 0, QB = 1), plus an unstable balance point in the middle that any small push tips one way or the other.8
The 6T cell names its six transistors by job:
- Pull-downs (N1, N2): the inverters’ NMOS transistors, which hold a node at 0.
- Pull-ups (P1, P2): the inverters’ PMOS transistors, which hold a node at 1.
- Access transistors (A1, A2): NMOS switches between each storage node and its bit line (BL for Q, BLB for QB), turned on by the word line.
Holding
With the word line low, both access transistors are off and the cell is isolated. In each inverter one transistor is on and the other off, so apart from leakage no current flows: an idle SRAM cell draws almost no power, but it does need the supply.
Reading
- both bit lines to .
- Raise the word line. Say Q = 0: current flows from BL through A1 and N1 to ground, so BL starts to fall. BLB, connected to the node holding 1, stays high.
- A fires as soon as the two lines differ by a small amount, and turns that difference into a full 0 or 1.
The sense amplifier matters because a bit line is long and connected to every cell in its column, so it has a large capacitance, and one small cell discharges it slowly. Waiting for a full swing would be slow; sensing a small swing is fast.1 SRAM reads are non-destructive: the cell keeps its value afterward.
Writing
- Write drivers force the bit lines to opposite values: BL = 0 and BLB = 1 to write Q = 0.
- Raise the word line. A1 drags Q down even though P1 is trying to hold it up. Once Q falls past the switching point of the other inverter, QB rises, which turns P1 off and N1 on, and the loop finishes the flip by itself.1
- Lower the word line. The cell holds the new value.
Notice the tension: on a read the access transistor must be too weak to flip the cell, and on a write it must be strong enough to flip it. The next section is about how designers satisfy both.
The 6T cell is two cross-coupled CMOS inverters (pull-ups P1, P2; pull-downs N1, N2) plus two NMOS access transistors A1, A2 gated by the and connecting the storage nodes Q, QB to a differential pair. It is the cell in most commercial chips: cell area dominates array area, and six logic transistors give a compact cell that is static, differential and readable without destroying its data.1
Read. Both bit lines are precharged and equalized to , then the word line rises. On the side storing 0, the access transistor (saturated, ) and the pull-down (linear, ) form a ratioed divider, so Q rises from 0 to a read bump while the bit line discharges through the series pair. Bit-line delay goes as with a large (every cell’s drain junction on the column) and a small (two minimum-size transistors in series), so arrays fire a clocked, latch-type as soon as the swing is large enough to resolve, and twist bit-line pairs so coupling noise lands on both lines equally.1
Write. Write drivers pull one bit line to 0 and hold the other at . The access transistor on the 0 side (linear, ) fights the pull-up holding that node at 1. If it wins by enough to pull the node below the opposite inverter’s trip point, positive feedback completes the flip; the other side helps a little, since its access transistor pulls the 0 node up from the bit line.4
Variants. Every access transistor added to the storage nodes adds another read disturbance, so multiported register files and low-voltage SRAMs isolate reads from the storage nodes with a separate read stack: an 8T cell (one read-only port) or a 10T cell. The price is area: the SKY130 dual-port 8T bitcell is 6.162 µm² against 1.896 µm² for the single-port 6T cell.115 A 10T sub-threshold design from MIT, whose separate read buffer removes the read disturbance, read correctly down to 320 mV.6
Input 1, output 0. The NMOS is on and pulls the output to ground.
An SRAM cell is a tug-of-war. The cell’s own switches pull one way, and the doors to the bit lines pull the other. Designers choose how strong each one is.
- For safe reads, the switch holding the 0 must be stronger than the door. If not, the bit line pushes in and lifts the 0. The whole loop flips, and the read wipes out the bit it was reading.
- For working writes, the door must be stronger than the switch holding the 1. If not, the cell refuses the new value.
Engineers measure how hard a cell is to knock over. They call it the cell’s : the biggest bump it can take without flipping.6
Here’s the catch. No two switches come out of the factory exactly alike. In a memory with millions of cells, a few will be lopsided. The memory only works if even those cells hold their bits. So designers plan for the unluckiest cell, not the average one.
Two sizing ratios, both measured against the access transistor, decide whether a 6T cell reads and writes correctly:
- = pull-down width ÷ access width. During a read the node holding 0 is pulled up through the access transistor and down through the pull-down. The stronger the pull-down, the smaller the bump. Commercial cells typically use 1.5 to 2.5.4
- = pull-up width ÷ access width. During a write the access transistor must overpower the pull-up. Cells usually make both minimum size, a ratio of about 1, and keep it below a limit that guidelines put between about 1.8 and 3.410
Bigger ratios cost area, and cell area is most of the memory’s area. Dropping the cell ratio from 2 to 1 saves about 25% of the cell but costs about 25% of its read margin.4
The butterfly curve
To measure stability, imagine cutting the loop. Plot the of one inverter (QB as a function of Q), then plot the other inverter’s curve flipped across the diagonal (Q as a function of QB) on the same axes. The two curves cross at three points: the two stored states in the corners and the unstable balance point in the middle. Between the crossings they enclose two “wings”, which is why this plot is called a .8
Fit the largest square inside each wing. The side of the smaller square, in volts, is the (SNM): the largest steady disturbance the cell can absorb without flipping.6 Even with perfect inverters it can’t exceed .10
- Hold SNM is measured with the word line off. The curves are steep and the wings are wide.
- Read SNM is measured with the word line on and both bit lines at , as during a read. The access transistors pull the 0 side up, the curves sag and the wings shrink. Read SNM is the worst case, and it grows with the cell ratio.510
Writes and mismatch
A write works the other way: it should leave the butterfly with only one crossing, so the cell has no choice but the new value. If the pull-up is too strong, a second stable point survives and the write fails.6 The measures how far the cell is from that failure.
Finally, means the two halves of a real cell never match. Its differ at random, so one wing is smaller than the other. Smaller transistors vary more, and SRAM cells use the smallest transistors on the chip.4 Across millions of cells the spread of SNM, not its average, decides whether the memory works.6
Read stability
During a read, the node storing 0 settles at , set by the access/pull-down divider. The read is non-destructive if stays below the trip point of the inverter it drives, approximately that inverter’s switching threshold. Raising the lowers ; commercial cells sit at . A minimum-size cell (, ) is about 25% smaller than a cell and has about 25% less read SNM; its read SNM can be restored with a 10% word-line underdrive during reads, which weakens the access transistor.4
Writability
With , the access transistor must pull the 1 node below the opposite trip point against the pull-up. PR is usually 1 and must stay below a limit that guidelines put between about 1.8 and 3.410 Write ability is quantified as a : the write noise margin (the side of the smallest square spanning the two write-mode VTCs), the highest bit-line voltage that still writes, or the word-line voltage margin, which is easier to measure in an array.410 On the butterfly, a successful write is a plot with a single crossing (a “negative SNM”); a failed write still has two lobes.6
SNM and the butterfly
Seevinck, List and Lohstroh defined the static noise margin in 1987 as the side of the largest square nested between the two VTCs, equivalent to the largest DC noise voltage that can be inserted in series with both inverter inputs (in the adverse direction) before the loop loses bistability.6 Read SNM is measured by holding the word line and both bit lines at , breaking the loop and sweeping each half-cell.5 The measures are not interchangeable:
- Hold SNM (WL = 0) is bounded by , even for ideal inverters.10
- Read SNM (WL = 1, BL = BLB = ) is the worst case in a 6T cell and rises with CR.610
- Write margin is not read off the hold or read butterfly; it needs the asymmetric write configuration.105
SNM is a DC metric. It ignores how long a disturbance lasts: a short enough noise pulse can exceed the SNM without flipping the cell, and dynamic noise-margin criteria, which depend on pulse width and the cell’s own response time, capture that.7
Variation sets Vmin
Random dopant fluctuation gives each transistor a threshold offset whose spread grows as channel area shrinks.4 between N1 and N2 (or between the access devices) skews the butterfly: one lobe grows, the other shrinks, and the cell’s SNM is the smaller lobe. An array needs every cell to have positive margin, so the SNM distribution is evaluated far into its tail, out to .6 That is why SRAM often sets a chip’s minimum operating voltage: in 2006, 65 nm 6T memories usually ran at 0.9 V or more (the lowest reported was 0.7 V), while a 10T cell that removes the read disturbance read at 320 mV and wrote at 380 mV at room temperature.6
Word line low: the access transistors are off. Raise it to see the read and the write fights.
Worst of 1 million cells ≈ 180 − 4.75×25 = 61 mV of read SNM: still positive, so every read is safe.
A cell is simpler than an SRAM cell. It has one switch and one tiny bucket for electric charge, called a capacitor. A full bucket means 1, and an empty one means 0.11
Two things make DRAM tricky:
- Reading empties the cell. So after every read, the memory fills the cell back up.
- The buckets leak. So the memory reads and rewrites every cell about 15 times a second. This is called .11
In return, one switch and one bucket take far less room than six switches. That’s why your computer’s main memory, many gigabytes of it, is DRAM.
A cell is one access transistor and one capacitor. The word line turns the transistor on, connecting the capacitor to the bit line. A charged capacitor (near ) is a 1; a discharged one (0 V) is a 0. Cells sharing a word line form a row, and roughly 512 cells share each bit line.1112
Reading by charge sharing
- every bit line to .
- Raise one word line. Each capacitor in the row shares its charge with its bit line, nudging the bit line slightly above (stored 1) or slightly below (stored 0). This is .
- Turn on the , one per bit line. Each detects which way its bit line moved and drives it all the way to or 0.
- Because the cell is still connected, driving the bit line also recharges the cell to its full value. The row of sense amplifiers now holds the whole row’s data, and reads and writes go through it.
- Lower the word line and precharge again before opening another row.
The nudge is small because the bit line is much bigger than the cell. The size of the change is
where is the cell’s capacitance and the bit line’s. The ratio is typically only 1–10%.10 Step 2 also wipes out the cell’s own charge, so a DRAM read is destructive; step 4 is what puts the data back.12
Refresh
The capacitor leaks, so a stored 1 slowly decays toward 0. A row is refreshed simply by opening it, since the sense amplifiers restore every cell in the row. Standard DRAM must refresh every row at least every 64 milliseconds; the memory controller issues a refresh command about every 7.8 microseconds, and each one makes the chip briefly unavailable.11
Activation. All bit lines in a bank are precharged to . ACTIVATE raises one word line and the selected row’s cells share charge with their bit lines, moving each by . With a charge transfer ratio of 1–10%, the signal is only a few percent of the cell’s voltage swing.10 The sense amplifiers (cross-coupled latches on each bit line, collectively the row buffer) then regenerate the bit line to the rail.1711 The bit line reaches a usable level ( or ) after tRCD, when READ or WRITE may be issued, and the cell is fully restored after tRAS; PRECHARGE then returns the bit lines to in tRP.12
Bit-line capacitance is the latency. A long bit line lowers and slows both amplification and precharge, because the sense amplifier must move at a fixed drive. But each bit line needs its own sense amplifier, so short bit lines multiply sense-amplifier area. In transistor-level simulations based on a 55 nm DDR3 process, cutting a bit line from 512 to 32 cells reduced tRCD from 15 ns to 8.2 ns and tRC from 52.5 ns to 23.1 ns, but grew the die 3.76×. Commodity DRAM chooses long bit lines and cheap bits.12
Refresh. Charge leaks only off the storage node (a stored 0 can’t gain charge), so retention errors are 1→0. The 64 ms interval is set by the leakiest cells; most cells retain data far longer. The controller issues auto-refresh every (3.9 µs above 85 °C), and each command blocks the rank for tRFC, on the order of 300 ns, closing all open rows.11 That works out to refresh commands per interval and roughly of the rank’s time at normal temperature. Because tRFC grows roughly linearly with chip density, refresh was projected to cost nearly 50% of throughput at 64 Gb densities in the extended temperature range, without changes such as retention-aware refresh.11
Disturbance. Dense cells couple. Repeatedly activating one row between refreshes accelerates leakage in neighbors (): a 2014 study induced errors in 110 of 129 modules from three manufacturers, with as few as 139K activations, and found up to one cell in 1.7K susceptible.13
Cell stores 1 (1.2 V). The bit line is precharged to V_DD/2 = 0.6 V.
memory keeps its data with the power off. Each cell is one switch with an extra part inside: a tiny island, sealed off by insulation on every side. Electrons can be parked there.14
Parked electrons make the switch harder to turn on. So to read the cell, the memory gives it a gentle test push. If the switch turns on, the island is empty. If it stays off, electrons are parked there. The insulation is so good that they stay for years.
Parking electrons takes a big electric push that forces them through the insulation. Each push does a little damage. After a thousand or more rewrites, the cell wears out.14
Modern flash even squeezes two, three or four bits into one cell. It parks different amounts of charge, like a glass that can be empty, a quarter full, half full or full.
A flash cell is a transistor with a charge-storage layer between its gate (the control gate) and the channel, insulated above and below. Electrons stored there raise the transistor’s , and the threshold encodes the data.14 Two storage layers are in use (both called here):
- A floating gate: a conducting layer completely surrounded by oxide. Used in older, planar flash.
- A charge trap: an insulating layer that holds electrons in traps. Used by most 3D NAND, where cells are stacked in dozens of layers.14
Operations:
- Program: a high voltage on the control gate pulls electrons from the channel through the thin tunnel oxide into the storage layer, in short pulses with a check after each one until the threshold reaches its target.
- Erase: a high voltage on the substrate pulls the electrons back out. Because all cells in a block share the substrate, a whole block is erased at once, and a cell can’t be rewritten until its block is erased.
- Read: apply a reference voltage to the control gate and see whether the cell conducts.14
A stores 2 bits (MLC, 4 threshold levels) or 3 bits (TLC, 8 levels) by dividing the same voltage range more finely. That multiplies capacity but makes each level narrower and easier to misread.14 NAND flash also chains cells in series along a bit line, which gives high capacity and low cost per bit but rules out reading single cells at random the way NOR flash can.10
Mechanism. Program and erase move charge through the tunnel oxide by Fowler–Nordheim tunneling, , which is exponential in the oxide field. Programming uses incremental step-pulse programming (ISPP): high-voltage pulses on the selected word line, each followed by a verify, until the cell crosses its target threshold. ISPP can only add charge, so data changes require a block erase (the substrate is raised to with the control gates grounded).14
Levels and errors. SLC uses two threshold states, MLC four, TLC eight, each assigned a window within the same overall range, so each extra bit halves the window per state. Programmed thresholds drift: stress-induced leakage through the tunnel oxide slowly moves charge off the gate after programming (retention errors), reads disturb unselected cells on the same string, and neighbors couple during programming.14
Endurance. Each program/erase (P/E) cycle damages the oxide. 50–59 nm MLC endured about 10,000 P/E cycles per block; 15–19 nm MLC and TLC about 3,000 and 1,000. 3D NAND reversed the trend for a while: with 48–64 stacked layers it could use larger features (about 50–54 nm), and its charge-trap cells are less prone to oxide breakdown, raising endurance by over an order of magnitude, though charge-trap cells leak faster soon after programming, including sideways along the vertical channel.14
Read voltage below the cell’s threshold: no current. Stored electrons raised V_T above this reference.
Here is one SRAM cell. Press Write 0, Write 1 or Read and watch each step. Green boxes are switches that are on. The circles show what each side holds. Orange dashes show where electricity flows.
Now make the cell weaker. Slide its grip on the 0 down and the manufacturing unevenness up, then read a 0. At some point, reading flips the bit. Or turn up the strength of the side holding 1, and writes start to fail.
Write and read the 6T cell and watch which transistors conduct (green) and where current flows (orange). The butterfly plot shows the same cell’s two transfer curves with the largest square in each wing; the status bar gives hold SNM, read SNM and write margin. Things to try:
- Switch to Read mode and compare it with Hold: the wings shrink.
- Lower the and raise the threshold mismatch until read SNM hits zero, then press Read with Q = 0. The cell flips while being read.
- Raise the to about 2 and try writing: the access transistor can no longer overpower the pull-up.
The model is a long-channel square-law 6T cell (, , PMOS at half the NMOS strength per width); curves are solved by bisection and the SNM found with the 45° rotation method described in Under the hood. Mismatch adds to N1 and to N2, so the Q = 0 lobe is the weak one. With at 0, check that read SNM at is roughly a quarter below ; find the PR at which the write margin closes; and see how much a cell tolerates before a read upsets it. The lower panel models DRAM sensing: choose and the time since refresh, and see when a decayed 1 falls below an illustrative ±30 mV sense threshold.
- SKY130 6T bitcell
- 1.896 µm²
- SKY130 8T dual-port bitcell
- 6.162 µm²
- DRAM refresh interval
- 64 ms
- Energy: DRAM vs register file
- ≈200×
Some real numbers, and what they mean:
- About 1.9 square micrometers. That’s one SRAM cell on an older chip-making process that anyone can use. A micrometer is a thousandth of a millimeter.15 About 50 cells in a row would be as wide as a thin hair.
- About 15 times a second. That’s how often every row of a DRAM chip gets refreshed, even when nobody is using it.11
- 800 steps. A fast processor waiting for main memory could do about 800 steps of work in the meantime, and the wait is only a ten-millionth of a second. That’s why chips keep fast SRAM close by.2
- 200 times. In an AI chip, fetching a number from DRAM can use about 200 times the energy of fetching it from a tiny memory right next to where the math happens.16
- About 1,000 rewrites. That’s how many times one common kind of flash from the mid-2010s could be rewritten before it wore out.14
| Quantity | Value | What it shows |
|---|---|---|
| SKY130 single-port 6T bitcell | 1.896 µm²15 | The foundry cell uses special dense “core” rules |
| SKY130 dual-port 8T bitcell | 6.162 µm²15 | A second port more than triples the cell |
| Typical cell ratio in commercial 6T cells | 1.5–2.54 | Pull-downs are made wider than access transistors for read stability |
| DRAM charge-transfer ratio | 1–10%10 | Only a small fraction of the cell’s voltage reaches the sense amplifier |
| Cells per DRAM bit line | ≈51212 | Long bit lines keep sense-amplifier area down |
| DRAM refresh interval / command spacing | 64 ms / 7.8 µs11 | Each row refreshed about 15 times a second |
| DRAM latency seen by a 2 GHz four-issue processor | ≈100 ns = 800 instructions2 | Why caches exist |
| Energy per access relative to a register file, in an accelerator | buffer 6×, DRAM 200×16 | Why AI chips spend so much area on SRAM |
| Flash endurance at 15–19 nm (MLC / TLC) | ≈3,000 / ≈1,000 cycles14 | More bits per cell, less endurance |
Array efficiency in an open process. OpenRAM’s second SKY130 test chip (MPW2) lists its macros’ outlines.15 Multiplying bits by the 1.896 µm² bitcell shows how much of each macro is actually storage:
| Macro (1 read/write port) | Outline | Macro area | Bitcell area | Storage share |
|---|---|---|---|---|
| 1 KB, 32 × 256 | 478.4 × 223.4 µm | 0.107 mm² | 8,192 × 1.896 µm² = 0.016 mm² | ≈15% |
| 8 KB, 64 × 1024 | 830.7 × 541.7 µm | 0.450 mm² | 65,536 × 1.896 µm² = 0.124 mm² | ≈28% |
Decoders, sense amplifiers, write drivers, control and power rings are a fixed overhead, so small macros are mostly periphery and density per bit improves with macro size. Measurements of the project’s first test chip, a 1 KB dual-port (8T) macro, show how much voltage headroom an SRAM array has in hold: written at 1.8 V, the first retention errors appeared at 440 mV cold and 410 mV hot.15
Margins and timing. Read SNM falls about 25% when CR goes from 2 to 1.4 In DRAM, the activate-to-read delay tRCD is 15 ns and the row cycle tRC 52.5 ns for a 512-cell bit line, against 8.2 ns and 23.1 ns for 32 cells at 3.76× the die area; over an 11-year span DRAM cost per bit fell 16× while tRCD and tRC improved by only about 30% and 26%.12
Energy. In a spatial accelerator, normalized energy per access is about 1× for a 0.5–1 kB register file, 2× for a neighboring processing element, 6× for a 100–500 kB global SRAM buffer and 200× for DRAM.16
No memory cell is fast, small, cheap and permanent all at once. So chips mix them. A little fast SRAM sits right next to the parts that compute. The bulk of the data lives farther away, in DRAM or flash.
A few things that go wrong:
- Squeezing SRAM too small. Smaller cells are weaker and more uneven. Some start flipping when read, or refuse new values.
- Leaky DRAM. A hot chip leaks faster, so hot DRAM gets refreshed twice as often.11
- Hammering. Reading one row of DRAM over and over can flip bits in the rows next door. This flaw is called .13
- Worn-out flash. Every rewrite wears flash a little. So phones and drives spread their saving evenly across all the cells.
For AI chips, the gap between small, fast SRAM and big, slow DRAM is one of the biggest problems of all. The Architectures guide picks up that story in The memory wall.
Density against speed: the memory hierarchy
SRAM is built from ordinary logic transistors, so it sits on the same die as the processor, is fast to reach and needs no refresh, but six transistors per bit make it the least dense. DRAM’s one-transistor cell is far denser but needs its own process for the capacitor, a destructive read, sense amplifiers and refresh. Flash is denser still and keeps data without power, but writes are slow and wear the cells out.210
Those differences produce the : small of SRAM next to the processor, sized so they can still answer within a cycle or two,3 backed by gigabytes of DRAM, backed by flash storage. The same split shows up in AI accelerators, where the energy of moving data dominates: a DRAM access costs about 200 times a register-file access, so designers fill the die with SRAM buffers and arrange computation to reuse each fetched value many times.16 The Architectures guide develops this in The memory wall, and SRAM-heavy designs in Wafer-scale and SRAM-heavy designs. How a CPU organizes its SRAM into levels of cache is the subject of Caches and coherence.
Inside each cell type
- SRAM: stability against area. Wider pull-downs improve read stability; smaller pull-ups help writes; both cost or save area. Minimum-size cells save about 25% area but lose about 25% of read margin.4
- DRAM: density against latency. More cells per bit line means fewer sense amplifiers and a smaller chip, but a weaker signal and slower access. Commodity DRAM picks density.12
- Flash: capacity against reliability. More bits per cell multiplies capacity but narrows the voltage gap between levels, which raises error rates and cuts endurance.14
How memories fail
- SRAM read upsets and write failures, concentrated in the few cells where random variation made one side much weaker. They get worse as voltage drops.6
- DRAM retention failures, where the leakiest cells lose a 1 before refresh, and disturbance errors such as .1113
- Flash wear-out and retention loss, managed by error-correcting codes and by spreading writes evenly across blocks.14
On-chip SRAMs arrive as pre-built blocks that the physical design flow places as (see Floorplanning), and they are tested with built-in memory self-test (see Design for test).
Where each technology sits
| 6T SRAM | 1T1C DRAM | NAND flash | |
|---|---|---|---|
| Process | Logic process, foundry bitcell with core rules | Dedicated DRAM process (capacitor, low-leakage access device) | Dedicated NAND process, 3D stacks of word-line layers |
| Read | Differential, non-destructive, small-swing sensing | Single-ended, destructive, charge sharing then restore | Threshold compare against reference voltages, page at a time |
| Retention | While powered | 64 ms worst-case, refreshed | Years, degrading with wear |
| Limiting margin | Read SNM / write margin at the cell | against sense-amp offset after leakage | Threshold window per state after drift |
The hierarchy follows directly: an L1 is limited by the largest SRAM that keeps hit time at 1–2 cycles,3 DRAM supplies capacity at a latency (tens of nanoseconds per activate) that barely improved over the 11 years one study examined, because manufacturers spend scaling on cost per bit,12 and flash supplies persistence. For accelerators, the energy ratio matters as much as the latency: with DRAM accesses about 200× a register-file access, dataflow is chosen to maximize on-chip reuse, and on-chip SRAM capacity becomes a first-order architectural parameter.16 The Architectures chapter The memory wall quantifies this for inference, including HBM, the stacked DRAM placed in the package (see Packaging and chiplets).
Failure modes and what designers do about them
- SRAM . Read SNM collapses first as drops and mismatch grows; write failures appear at fast-PMOS/slow-NMOS corners. Remedies: upsized or 8T/10T cells for low-voltage or multiport arrays, read/write assists, a separate boosted word-line supply, and row/column redundancy that replaces failing cells after test.64
- DRAM refresh overhead and retention tails. Retention is log-normally distributed with a tail of leaky cells; refreshing everything at the leakiest cell’s rate wastes energy and bandwidth that grow with density, which motivates retention-aware refresh.11
- DRAM disturbance. RowHammer showed that inter-cell coupling at small pitch is a correctness and security issue, not only a reliability one.13
- Flash retention, disturb and wear, handled in the controller with ECC, read-retry and wear leveling.14
Choose a source to fetch an operand from. Bars show energy per access relative to the register file.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
1. The read bump from the square law
Use the long-channel model from The I-V curve: in saturation and in the linear region, with proportional to . During a read with Q = 0 and QB = :
- The access transistor has its gate at , drain on the bit line at and source at . , so it is saturated: .
- The pull-down has and a small , so it is linear: .
Setting the two equal with gives a quadratic in :
Its smaller root has a tidy closed form:
With and : gives 0.205 V, gives 0.158 V and gives 0.128 V. The bump shrinks with CR, but slowly: doubling CR from 1 to 2 removes only about a third of it. The read is safe while stays below the trip point of the inverter driven by Q, which is itself raised slightly by its own access transistor pulling QB up. This is the calculation behind the rule “pull-down much stronger than access”.14
2. The write condition
Writing Q from 1 to 0: BL = 0, QB = 0, so the pull-up P1 has its full . The access transistor (source on BL at 0 V, ) is linear; the pull-up is saturated while is below :
which gives
as long as the result is below (beyond that the pull-up leaves saturation and must be solved numerically). With and , , well below the opposite inverter’s trip point (about 0.5 V here, raised slightly because its own access transistor is pulling QB up from BLB), so the write succeeds. Raising PR raises ; once it crosses the trip point the write fails. In this model that happens near , consistent with the guideline of keeping PR below about 1.8.10 The write noise margin generalizes this to the full VTCs.4
3. Computing SNM: the 45° rotation
Finding the largest square between two curves by brute force is awkward. Because both VTCs are monotonically decreasing, the largest square in a lobe has two opposite corners on the two curves and its diagonal along the direction . Rotate coordinates by 45°:
In the rotated frame each curve is a single-valued function , and the gap at a given is the diagonal of a square that fits between the curves at that position. So
taken separately for and , and the cell’s SNM is the smaller of the two.7 If, in one lobe, the gap never has the expected sign, the curves cross only once on that side: the cell is monostable in that state and the margin is zero (or negative, if you keep the sign). The simulation above uses exactly this procedure on 90-point curves.
Assumptions to keep in mind: the method is static (DC), it treats the two halves as independent once the loop is cut, and it measures the margin to a noise source applied in the most damaging direction on both nodes at once. Dynamic criteria, which include the noise pulse width and the cell’s internal time constants, are needed when pulses are short.7
u = -0.200 V: Δv = 388 mV, square side Δv/√2 = 274 mV. This lobe peaks at 391 mV (u = -0.350); cell hold SNM = min of the lobes = 391 mV.
4. Designing for the tail
For an array of cells with independent margins, the probability that every cell works is , where is the probability one cell fails. Keeping a 1 MB array () failure-free with, say, 99% probability needs , about for a Gaussian margin. Because local mismatch has a standard deviation that grows as channel area shrinks,4 and analytic models are used to estimate the SNM distribution’s tail (plain Monte Carlo would need billions of samples to see a event),6 is set by the worst cells, and redundancy (spare rows and columns mapped in after test) relaxes the requirement on .
5. DRAM sensing budget
Charge conservation at the moment the access transistor turns on:
so .10 With the signals for a stored 1 and 0 are symmetric only when the 1 is fresh. A stored 1 leaks toward 0 between refreshes, so its signal shrinks first; a 0 cannot gain charge.11 The sense amplifier needs larger than its input offset (set by mismatch between its own transistors) plus coupling noise from neighboring bit lines. That gives the design equation the simulation plots: for a given ratio and minimum signal , a 1 is still read correctly while
and the refresh interval must be shorter than the time the leakiest cell takes to fall to that level. Longer bit lines (larger ) raise the required cell voltage and shorten the allowed interval, which is the density-versus-retention trade in one line.12
6. Refresh arithmetic
Refresh overhead . With and , about 3.8% of the rank’s time is spent refreshing; at extended temperature () it doubles. Since tRFC scales roughly with the number of rows refreshed per command, and therefore with density, the overhead grows with every generation. RAIDR’s observation is that only a tiny fraction of cells need 64 ms: in a 32 GB system only about 30 cells could not tolerate a 128 ms interval and about 1,000 could not tolerate 256 ms, so binning rows by retention time lets most rows be refreshed far less often.11
7. Flash programming
Fowler–Nordheim current density rises exponentially with the oxide field, which is why programming uses pulses of high voltage and why a slightly weakened oxide leaks measurably over months. ISPP raises the program voltage by a fixed step per pulse and verifies after each one, so the final threshold distribution width is set by the step size: smaller steps give tighter distributions (needed for TLC and QLC) at the cost of more pulses and slower programming.14
Q1During an SRAM read, both bit lines start at the supply voltage and the word line turns on. What could go wrong, and which design choice prevents it?
Q2What does the static noise margin read from a butterfly curve tell you?
Q3A DRAM cell is connected to a bit line with ten times its capacitance. Roughly how much of the cell’s voltage difference shows up on the bit line?
Q4Why does storing 3 bits per flash cell (TLC) reduce how many times the cell can be rewritten, compared with 1 bit (SLC)?
Sources
Show Hide 17 sources
- Lecture 19: SRAM (CMOS VLSI Design, 4th ed. slides)6T cell read and write sequences, read stability (pull-down much stronger than access) and writability (access much stronger than pull-up), sense amplifiers triggered on a small bit-line swing, cell sizes in λ, multiported cells.
- Instruction Set Architecture, MIT 6.5900 Lecture L02 (memory technology and the 6T SRAM cell)Read and write steps of a 6T cell; register, SRAM and DRAM ordered by size and latency; a 2 GHz four-issue core could run 800 instructions during one 100 ns DRAM access; flash is slower but denser than DRAM.
- Caches (continued), MIT 6.5900 Lecture L03The L1 cache is sized as the biggest cache that doesn’t push hit time past 1–2 cycles.
- A 65-nm Reliable 6T CMOS SRAM Cell with Minimum Size TransistorsCell ratio CR = Wpd/Wacc, commonly 1.5–2.5; pull-up ratio usually below 3 and often 1; CR = 1 saves 25% area but cuts read SNM by 25%; V_READ vs V_TRIP; write noise margin; word-line underdrive as a read assist; Vt variation grows as channel area shrinks.
- SRAM Static Characterization (course handout)Read SNM: hold WL and both bit lines at VDD, break the loop, plot an inverter’s VTC and its inverse; SNM is the side of the largest square in the butterfly. Write noise margin from the asymmetric write butterfly.
- A 256kb Sub-threshold SRAM in 65nm CMOS (ISSCC 2006 slides)SNM as the side of the largest embedded square (after Seevinck, List and Lohstroh, 1987); read SNM is the worst case; variation spreads SNM and sets yield; a successful write leaves one stable point; a 10T cell with a separate read buffer reads to 320 mV.
- Static and Dynamic Stability Criteria of 6T SRAM Bit CellComputing SNM by rotating the butterfly 45° and taking the largest gap between the rotated curves divided by √2; static vs dynamic noise margin.
- Lecture 32: MOS Memory (EE 331 Microelectronic Circuit Design)Cross-coupled inverters are bistable with an unstable equilibrium in between; the butterfly diagram; SRAM keeps data through a read, DRAM must be refreshed.
- The Memory Hierarchy (CSE 240A lecture slides)Memory cell sizes in F² (F = smallest feature): a 65 nm 6T SRAM cell is 0.52 µm², 123–140 F² (after ITRS 2008); a DRAM cell is about 6 F², about 20× denser than SRAM.
- Lecture 16, ECE 122A VLSI Principles (SRAM noise margins, DRAM and flash cells)SNM is limited to VDD/2; read SNM rises with cell ratio; pull-up ratio below 1.8 for writes; 1T DRAM charge redistribution ΔV = (V_BIT − V_PRE)·Cs/(Cs + C_BL), charge transfer ratio 1–10%, destructive read; floating-gate flash, NAND vs NOR.
- RAIDR: Retention-Aware Intelligent DRAM Refresh1T1C cell, precharge to VDD/2, sensing and restore; 64 ms refresh interval, tREFI 7.8 µs (3.9 µs above 85 °C), tRFC around 300 ns; retention set by the leakiest cells; charge only leaks off a capacitor.
- Tiered-Latency DRAM: A Low Latency and Low Cost DRAM ArchitectureAbout 512 cells share a bit line; charge sharing, sensing and restore; bit-line capacitance is the main source of DRAM latency; short bit lines cut tRCD from 15 to 8.2 ns but cost 3.76× die area; cost per bit fell 16× while latency barely moved.
- Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance ErrorsRepeatedly activating one row corrupts nearby rows: errors in 110 of 129 modules, as few as 139K accesses, up to one in 1.7K cells susceptible.
- Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State DrivesFloating-gate and charge-trap cells; data stored as threshold voltage; SLC/MLC/TLC states; Fowler–Nordheim tunneling and incremental step-pulse programming; block erase; ~3,000 (MLC) and ~1,000 (TLC) P/E cycles at 15–19 nm; 3D NAND with 48–64 layers.
- SRAM Design with OpenRAM in SkyWater 130nmSKY130 foundry bitcells: single-port 6T 1.896 µm², dual-port 8T 6.162 µm²; MPW2 macro dimensions; on the first (OR1) dual-port test chip, retention errors first appeared at 410–440 mV after writing at 1.8 V.
- Efficient Processing of Deep Neural Networks: A Tutorial and SurveyNormalized energy per access in an accelerator: register file 1×, neighbor PE 2×, global buffer 6×, DRAM 200×; DRAM costs two orders of magnitude more energy per access than a small on-chip memory.
- Dynamic random-access memoryAll cells in the open row are sensed and their sense-amplifier outputs latched; rows are refreshed every 64 ms or less.