Picture a chef who can chop a thousand carrots a second. That speed is wasted if the carrots arrive slowly. If the pantry is down a long hallway and one helper carries one basket at a time, the chef mostly stands around.
AI chips have the same problem. The parts that do math are incredibly fast. But the numbers they work on are kept in memory, and memory can only hand them over so fast. Engineers call this gap the .
It hurts most when a chatbot writes. It writes its answer one (a word or part of a word) at a time. For every token, the chip reads the whole AI model out of memory: billions of numbers. The math is quick. The fetching is slow.
Every computation needs two things: arithmetic, measured in (operations per second), and data delivered from memory, measured as (bytes per second). Over the last few decades chips have gained arithmetic far faster than memory has gained bandwidth. William Wulf and Sally McKee named the result the in 1995: no matter how fast the arithmetic, a program that waits on memory runs at the speed of memory.1
Large language models make the problem sharp. When a model generates text, each new (a word or word fragment) needs one pass through the whole network, and that pass reads every weight from memory. A 70-billion-parameter model stored with 2 bytes per weight is 140 GB. A chip that streams 4 TB/s needs 35 ms just to read it, so a single user can never get more than about 28 tokens per second, however fast the chip’s math units are. A job like this, limited by data movement rather than arithmetic, is called .
The chapter covers:
- the data on compute outgrowing memory, and what moving data costs in energy;
- , the stacked DRAM placed beside AI chips;
- on-chip , and why it has become expensive;
- why text generation is memory-bound, and the memory a conversation’s history takes up;
- the main ways around the wall, and what each costs.
It builds on the roofline model from What the workload needs.
The says attainable throughput is the lesser of peak compute and bandwidth × .3 The memory wall is the observation that the first term has grown much faster than the second for decades, so the keeps moving right and more real kernels fall on the bandwidth slope.1 The wall has three faces: bandwidth (bytes per second to the compute), capacity (bytes that fit close enough to be fast), and latency, which has improved least of all.
LLM inference shows all three. The phase does about 2 FLOPs per parameter per token while reading every parameter once per step, an intensity of 1–2 FLOPs per byte at batch 1 against ridge points of several hundred on current accelerators. Batching raises the weight-read intensity linearly, but each sequence also carries a that is read every step, is not shared across the batch, and competes with the weights for capacity.17
This chapter quantifies the trend, then walks the hierarchy from HBM to on-chip SRAM, builds the decode model (memory time versus compute time, the batch “knee”, the capacity bound), and evaluates the mitigations: batching and its scheduling, quantization of weights and KV cache, grouped-query attention, speculative decoding and splitting the phases. The full model, with its assumptions, is in “Under the hood”.
Fetch 35.0 ms, math 0.14 ms: about 28 tokens/s for one user, math units busy 0.4% of the time.
Over the past twenty years, the math speed of AI computers grew about 60,000 times. The speed of their memory grew only about 100 times.1 The chef keeps getting faster, but the hallway barely gets wider.
Distance also wastes energy. Fetching a number from memory outside the chip uses about 200 times more energy than fetching it from right beside the math.2 So designers try to keep numbers close and use each one many times.
That gives a simple rule: a job runs well when it does lots of math with each number it fetches. Reading your question is a job like that, because each number gets used for every word of the question at once. Writing the answer is not. Each number is used for one token, then fetched again for the next.
The trend
Gholami and colleagues at UC Berkeley compared AI hardware over twenty years. Peak compute grew about 3.0× every two years (60,000× in total), DRAM bandwidth about 1.6× every two years (100×), and the bandwidth of the links between chips about 1.4× every two years (30×).1 Models grew faster still: between 2018 and 2022 the parameter count of the largest Transformer models rose about 410× every two years, while the memory on a single accelerator rose about 2×.1 So the models no longer fit on one chip, and the chip that holds part of a model can’t read it fast enough.
Moving data costs energy as well as time
Every chip has a : tiny, fast storage beside each math unit, larger on-chip buffers, and large off-chip . In one widely used accelerator analysis, reading a value from a small register file costs 1 unit of energy, from a neighboring processing element 2, from a shared on-chip buffer 6, and from DRAM 200.2 A costs about the same as the register-file read, so a design that fetches each weight from DRAM for a single use spends almost all its energy on movement.
The rule of thumb: operations per byte
is the number of operations a job does per byte it moves. A chip has its own number to compare it with, the : peak compute divided by bandwidth. A hypothetical chip with 1,000 TFLOPS ( operations per second) and 4 TB/s has a ridge point of 250 operations per byte.
- A job with intensity above 250 can keep the math units busy. It is compute-bound.
- A job with intensity below 250 leaves them idle part of the time. It is , and runs at bandwidth × intensity. At intensity 2, this chip delivers 8 TFLOPS, under 1% of its peak.
The draws exactly this: a slanted line for the bandwidth limit meeting a flat line for the compute limit.3 As compute outgrows bandwidth, the corner where they meet moves to the right, and jobs that used to be compute-bound become memory-bound on newer chips.
Rates, compounded
Fitted over about twenty years of server-class AI hardware, peak FLOPS grew 3.0× per two years, DRAM bandwidth 1.6× and interconnect bandwidth 1.4×, which over the period is 60,000×, 100× and 30×.1 The ratio of the first two, about 1.9× per two years, is the rate at which the ridge point drifts right. Two consequences:
- A fixed algorithm slides down the roofline over generations. A kernel at 100 FLOP/B that was compute-bound on a chip with a 60 FLOP/B ridge becomes memory-bound when the ridge passes 100, without any change to the code.
- Scaling out doesn’t escape it. Spreading a model over more chips trades on-chip traffic for chip-to-chip traffic, which has grown slowest of the three.1 See Why the network looks this way in Systems.
Capacity scales no better. Between 2018 and 2022 the parameter count of state-of-the-art Transformers grew about 410× per two years against about 2× for single-accelerator memory, so every frontier model is now partitioned.1 Latency, the third face, has been even harder to improve than bandwidth.1 Accelerators therefore hide it behind many outstanding requests (many warps on a GPU, long DMA streams on a systolic design) rather than wait on it.
Energy per access by level
Sze, Chen, Yang and Emer normalize the energy of fetching an operand to the energy of the ALU operation itself (1×): about 1× from the register file, 2× from a neighboring PE over the array network, 6× from a 100–500 kB global buffer and 200× from DRAM.2 Two orders of magnitude between a small on-chip SRAM and DRAM is the reason dataflow design exists: a weight fetched once from DRAM and reused 200 times from local storage costs about the same energy per use as the arithmetic. The same logic gives one roofline per level. A kernel tiled for on-chip SRAM is judged against the SRAM bandwidth ridge; the DRAM ridge only binds for the traffic that actually leaves the chip.
Intensity of the transformer phases
Counting a multiply-accumulate as 2 FLOPs, a dense layer applied to tokens does FLOPs per weight. If the weight is read once in bytes, intensity from weights alone is .
- processes every prompt token together, so is the prompt length times the batch: thousands of FLOPs per byte, above any current ridge.
- processes one token per sequence per step, so , the batch size. At with 16-bit weights the intensity is 1 FLOP/B.
Profiling makes the same point: GPT-2 and BERT-Base have nearly the same configuration and total FLOPs, yet the autoregressive model is much slower, because its matrix–vector steps have far lower intensity.1 The chips in “By the numbers” below have ridge points from about 140 to about 330 FLOP/B.
Year 0: compute 1.0×, DRAM bandwidth 1.0×, so the ridge point is 0.4 FLOP/B. Decode at 1 FLOP/B is compute-bound.
Fetch from DRAM: 200× a math operation. Reused 1×, energy per use is 201×, 100% of it data movement.
A laptop’s memory sits a few centimeters from its processor, joined by a fairly small number of wires. AI chips use a different setup called , short for high-bandwidth memory.
In HBM, several memory chips are stacked like pancakes. The stack sits right next to the processor, joined to it by more than a thousand tiny wires.5
HBM has limits. Only a few stacks fit around a chip, and they are hard and expensive to build. A big AI model can need more memory than one chip has.
DRAM and its interface
A cell is one transistor and one tiny capacitor; the bit is the charge on the capacitor. That makes it dense and cheap per bit, but it is built in its own manufacturing process, separate from the logic chip. The Memory cells chapter of the Transistors guide shows how the cell works.
Bandwidth is interface width times the data rate on each wire. Ordinary graphics memory (GDDR) uses a 32-bit interface per chip and runs each wire very fast.5 goes wide instead:
- DRAM dies are stacked and joined by , vertical conductors through each thinned die, and microbumps.
- The stack presents a 1,024-bit interface (2,048 bits in the newest generation, HBM4).59
- Thousands of wires can’t run through a circuit board, so the stack and the processor sit side by side on a , a slab of silicon with fine wiring.5
One HBM3 stack runs up to 6.4 Gb/s per wire across 1,024 wires: 6.4 × 1,024 ÷ 8 ≈ 819 GB/s. An HBM4 stack runs 8 Gb/s across 2,048 wires, for up to 2 TB/s.79 A chip gets its total bandwidth by placing several stacks around it. NVIDIA’s H100, for example, used five HBM3 stacks for over 3 TB/s at launch.10
What it costs
- Capacity comes in stack-sized steps. HBM4 allows up to 64 GB per stack (16 dies of 32 Gb).9 Total capacity is stacks × stack size, and both are limited.
- Room around the chip is limited. Every stack needs its own interface circuits on the edge of the processor die and its own patch of interposer, so the number of stacks is set by the die’s perimeter and the interposer’s size.
- The package is complex. Stacking, interposers and bonding are covered in Packaging and chiplets in the Systems guide.
Bandwidth arithmetic
. Over the generations, both terms moved:
- First-generation HBM, as shipped in 2015, ran 1 Gb/s per pin, so four stacks on a 4,096-bit interface gave 512 GB/s, 128 GB/s per stack.5 HBM2 reached 2.4 Gb/s for 307 GB/s; HBM3 runs 6.4 Gb/s for 819 GB/s; HBM3E parts run 9.8 Gb/s for more than 1.2 TB/s.678
- JEDEC’s JESD270-4 (HBM4, April 2025) doubles the width to 2,048 bits and the independent channels from 16 to 32 (each split into 2 pseudo-channels), at up to 8 Gb/s, for up to 2 TB/s per stack. It defines 4-, 8-, 12- and 16-high stacks of 24 Gb or 32 Gb dies, up to 64 GB.9
The host side mirrors this. H100 SXM5 drives five HBM3 stacks from ten 512-bit memory controllers, a 5,120-bit interface, for over 3 TB/s at launch; that works out to roughly 4.7 Gb/s per pin, below the HBM3 maximum.107 MI300X uses eight HBM3 stacks for 5.3 TB/s.11
Constraints that come with it
- Shoreline. Each stack’s PHY occupies die edge on the processor, and the stack must sit within a few millimeters of it. Stack count is therefore bounded by the processor’s perimeter.
- Interposer area. Logic die plus stacks must fit on the interposer, whose size, yield and cost are packaging problems in their own right (see Packaging and chiplets).
- Capacity and bandwidth are coupled. More stacks give both; taller stacks give capacity only. Doubling capacity at constant bandwidth halves how often the chip can read all of its memory, which matters for decode, where every step reads most of it.
- Peak is not sustained. Refresh, bank conflicts and read/write turnarounds take a share; streaming kernels typically reach a large fraction of peak but not all of it, so the decode model in the simulation is an upper bound.
HBM3: 1,024 bits × 6.4 Gb/s ÷ 8 = 819 GB/s per stack; × 4 = 3.3 TB/s. Tap a part of the stack to read about it.
Inside the chip itself is a second kind of memory, called . It sits right beside the math parts, so it is very fast and uses little energy.
The catch is size. SRAM takes much more room per number than the memory in an HBM stack. A typical big AI chip holds a few hundred megabytes of it. A large model needs hundreds of times more than that.12
So chips use SRAM like a workbench, not a warehouse. They bring in a slice of the model, use it as many times as they can, then swap in the next slice. Some designs instead spread a whole model over many chips full of SRAM. The Wafer-scale and SRAM-heavy designs chapter covers them.
Why SRAM is fast and big
An cell is usually six transistors forming two cross-coupled inverters plus two access switches. It is built in the same process as the logic, so it can sit beside the math units with short, wide connections. A DRAM cell is one transistor and a capacitor, built in a separate process. Trading density for speed this way makes SRAM the material of registers, and . The Memory cells chapter shows the 6T cell and why it needs margins that limit how small it can get.
A concrete comparison on one GPU, the A100:
- HBM: 40–80 GB at 1.5–2.0 TB/s.
- SRAM: 192 KB in each of 108 streaming multiprocessors (about 20 MB in all) at an estimated 19 TB/s.4
That is about ten times the bandwidth with a two-thousandth of the capacity or less. FlashAttention, a widely used attention algorithm, gets its speed by splitting the work into tiles that fit in that SRAM, so the large intermediate matrices never travel to HBM.4
Big last-level memories
Accelerators also carry one large shared on-chip memory: a 50 MB L2 cache on H100, a 256 MiB last-level cache on MI300X, and a 128 MiB software-managed memory (CMEM) on TPU v4.101213 Google reports that one of its production models, an RNN with small weights run at small batch, benefits significantly from CMEM bandwidth compared with HBM: exactly the case where HBM bandwidth would otherwise bind. CMEM made that model 2× faster, against 1.2× across Google’s production workloads overall.13 For a model of tens of gigabytes, a few hundred megabytes is a buffer, not a home.
What SRAM costs in area
At a 2025 conference, TSMC and Intel each reported SRAM blocks on their newest (2 nm-class) processes with a density of 38.1 megabits per square millimeter.16 At that best-case density, one gigabyte (8,000 megabits) takes about 210 mm² of silicon. A 70-billion-parameter model at 8 bits per weight would need about 15,000 mm², roughly 25 dies the size of a TPU v4 (under 600 mm²), before any logic.13
The scaling stall
At the 3 nm generation the high-density SRAM bitcell stopped shrinking. TSMC’s N3 cell is 0.0199 µm², and the N3E variant, which is believed to relax some pitches for performance and yield, discloses 0.021 µm², larger than N3’s.14 Trade press summarized the 2022 disclosures as SRAM scaling stalling at 3 nm, with direct effects on the die size and cost of high-performance chips such as GPUs and processors; scaling resumed to some degree at 2 nm, TSMC’s and Intel’s first nanosheet node.1516 The bitcell is built from the smallest devices in the process, and its read and write margins depend on how well they match (the butterfly curve and static noise margin in Memory cells), one reason it is harder to shrink than logic.
Area budget arithmetic
Take the densest SRAM blocks TSMC and Intel reported at ISSCC 2025, 38.1 Mb/mm² on 2 nm-class processes.16 Real arrays lose density to sense amplifiers, decoders, redundancy and the wiring that connects banks to compute, so treat the following as a lower bound on area:
- → about 210 mm².
- A 256 MiB last-level memory (MI300X’s size) → about 56 mm², already a meaningful slice of an advanced-node die budget.
- A 70B model at 8 bits → about 15,000 mm²: dozens of reticle-class dies.
Hence the two strategies. Most accelerators keep weights in HBM and size SRAM for reuse: enough to hold tiles of weights, activations and KV blocks so that each HBM byte is used many times. SRAM-heavy and wafer-scale designs instead keep weights entirely on chip and pay for capacity with chip count, which turns the memory wall into an interconnect problem (see Wafer-scale and SRAM-heavy designs).
Two rooflines
With an SRAM level, each kernel has two intensities: FLOPs per HBM byte and FLOPs per SRAM byte, each judged against its own ridge. FlashAttention’s contribution is to make attention IO-aware: it tiles the computation so that the score matrix lives only in SRAM, cutting HBM reads and writes.4 That moves attention right on the HBM roofline. It does not change what decode must read from HBM each step, which is set by the weights and the KV cache.
1 GB of SRAM at 38.1 Mb/mm² needs about 210 mm²: 35% of a 600 mm² die, before any logic.
A chatbot answers in two stages. First it reads your whole question in one trip through the model. Then it writes the answer, and every token needs another whole trip. The second stage is the slow one, and memory is the reason, not math.19
There is more to fetch, too. The model keeps notes on everything said so far, so it doesn’t have to reread the whole chat for each token. These notes are called the . They are read every step, and they grow as the chat gets longer. A long chat with a big model can need over a gigabyte of notes.18
Two phases
Language-model inference has two phases with very different needs:
- runs the whole prompt through the model in one pass. All prompt tokens share each weight read, so it is compute-intensive.
- then generates one token per step. Each step reads all the weights for a single token per conversation, so it is memory-intensive.1918
The arithmetic of one decode step
Each weight takes part in one multiply and one add per token: 2 operations. If a weight takes 2 bytes, the step does 2 operations for every 2 bytes read, an intensity of 1 operation per byte. The ridge point of a current accelerator is a few hundred. So at one conversation at a time, the math units are busy well under 1% of the time.
decodes conversations in the same step. The weights are read once and used times, so the intensity rises to about operations per byte (at 16-bit). Total throughput rises almost -fold while each user’s speed barely drops, until the math units catch up with memory.
The KV cache
Attention lets each new token look back at every earlier one. To do that cheaply, the model stores two vectors per earlier token in every layer, a key and a value: the . Its size per token is:
For Meta’s Llama 3 70B (80 layers, 8 KV heads of 128 numbers each) at 2 bytes per number, that is bytes, or 320 KiB per token.20 A 4,096-token conversation then needs about 1.3 GB. An older 13-billion-parameter model without the shared-head trick needs 800 KB per token, up to 1.6 GB for one 2,048-token request.18
This hurts in two ways:
- Capacity. Every conversation in the batch needs its own KV cache in memory next to the weights. Serving a 13B model on a 40 GB GPU, about 65% of memory holds the weights and about 30% the KV cache, so only a few dozen requests fit.18 Capacity, not compute, often sets the batch size.
- Bandwidth. Each conversation’s KV cache is read every step and can’t be shared, so it adds bytes in proportion to batch size and context length. At small batches and short contexts, reading the weights dominates; at large batches and long contexts, reading the KV cache does.17
Model designers respond by sharing keys and values across attention heads. Llama 3 uses with 8 KV heads, specifically to speed up inference and shrink the KV cache.20
Memory time and compute time
Pope et al. give the standard first-order model. Weights and KV cache are the two large tensors, and each is transferred from HBM to the compute once per forward pass; that transfer time is the “memory time”. An -parameter decoder needs matmul FLOPs per token; at peak that is the “compute time”.17 For one decode step at batch and context :
- Bytes: , with and the KV bytes per token.
- FLOPs: ; the second term is attention ( and the weighted sum over ), small next to until reaches thousands.
- Step time ; per-user rate ; aggregate .
At small and the term dominates; at large and (Pope et al. cite 2,048+ tokens at batch 512+) the KV term does.17 For a 500B+ model with full multi-head attention at batch 512 and context 2,048, the KV cache totals 3 TB, three times the parameters, and must be read once per generated token while the compute is essentially idle.17
Why batching stops helping
The weight term’s intensity is , so it rises with batch. The attention term’s is not: per layer and sequence it does FLOPs and reads bytes ( KV heads of size ). With query heads, , so
independent of both and . For Llama 3 70B () at 16 bits that is 8 FLOP/B.20 As grows, the step’s blended intensity levels off at one sequence’s FLOPs per KV byte, and as grows that ratio falls toward this small number, so at long context decode stays memory-bound at any batch size. The simulation below shows this as a curve that never reaches its compute roof.
Capacity sets the batch, and fragmentation wastes it
The largest batch that fits is . Static per-request allocation wastes much of that: in the systems vLLM measured, only 20.4–38.2% of KV-cache memory held actual token states, the rest lost to reservation for maximum length and fragmentation.18 Paging the KV cache in fixed blocks (PagedAttention) nearly eliminates the waste and gave 2–4× throughput at equal latency.18 More sequences per batch is the whole game in the memory-bound regime.
A latency floor
At batch 1, . PaLM 540B with int8 weights on 64 TPU v4 chips gives . The measured low-batch latency was 29 ms per token.1713 The gap is the part of the problem a single-chip model leaves out: chip-to-chip communication (which shrinks slowly or not at all as chips are added), KV traffic, and sustained bandwidth below peak.17
A six-token prompt arrives. Step through prefill and decode.
This is a pretend AI chip running a chatbot. The solid line shows how many tokens per second it writes for everyone together. The dashed line shows what each user gets. Start at 1 user, then slide “Users at once” up and watch the total climb. Then set “Size of each number” to 4-bit, and try a very long conversation.
A generic, illustrative accelerator decodes a dense language model. Each step reads all the weights plus each sequence’s KV cache, and time per step is the larger of memory time and compute time. Things to try:
- At batch 1, compare the memory and compute bars. Then raise the batch: aggregate tokens/s grows almost linearly until the green “knee” where compute takes over.
- Switch 16-bit → 8-bit → 4-bit at batch 1 and watch per-user speed double each time.
- Press “Chat” (16-bit, 2K context): memory fills at 77 conversations; at this context length compute never becomes the limit. Press “Long” (32K context): the KV cache dominates and the curve flattens without ever reaching the compute limit. “Huge” doesn’t fit at all.
The model is the one in “Why decode is memory-bound”: , , step = max of the two times, ideal overlap and 100% of peak. Peak compute doubles per halving of the weight format. Layer counts follow open dense models; edit layers, KV heads, head dimension and KV-cache format directly. Things to check:
- The knee moves with the ridge point and with ; note that changing the format moves it only when the KV format is decoupled from the weight format.
- Set KV heads to 64 (multi-head) on the default model and watch and the knee.
- At 32K context, intensity levels off at one sequence’s FLOPs per KV byte, far below the ridge however large the batch. It tends toward only as context grows further; quantizing the KV cache to 2 bits raises it.
- Growth in peak AI-hardware compute, 20 years
- ≈ 60,000×
- Growth in DRAM bandwidth, same period
- ≈ 100×
- HBM4 bandwidth per stack (JEDEC maximum)
- 2 TB/s
- Llama 3 70B KV cache per token, 16-bit
- 320 KiB
Sources: Gholami et al. for the growth rates, the JEDEC HBM4 announcement for the stack bandwidth, and the Llama 3 model shape for the KV cache.1920
What these numbers mean:
- 60,000× versus 100×. Math got about 600 times faster compared with memory. The chef sped up far more than the hallway widened.
- 2 TB/s per stack means 2 trillion bytes a second (a byte is about one letter of text). That could read a two-hour HD movie hundreds of times a second. It is still the slow part.
- 320 KiB per token is the size of the notes for each token. A 4,000-token chat needs more than a gigabyte of notes, for just one user.
- Ridge in the table is how many math steps a chip must do with each byte it fetches to stay busy. Writing a chatbot answer does only one or two. Newer chips need more, so the wall is getting taller.
Reading the table:
- Ridge points have risen from about 140 to over 300 FLOPs per byte across these generations. Decode at batch 1 runs at 1–2 FLOPs per byte, so on every chip in the table a single conversation uses well under 1% of peak compute.
- The largest on-chip memories are 32–256 MiB, hundreds of times smaller than the HBM beside them, and HBM itself (32–192 GB) is smaller than many current models at 16 bits.
- Bandwidth per chip grew about 6× from TPU v3 to MI300X, mostly by adding stacks and moving to HBM3.
HBM generations, per stack:
| Generation | Per-pin rate | Interface | Bandwidth per stack |
|---|---|---|---|
| HBM | 1.0 Gb/s | 1,024 bits | 128 GB/s |
| HBM2 | 2.4 Gb/s | 1,024 bits | 307 GB/s |
| HBM3 | 6.4 Gb/s | 1,024 bits | 819 GB/s |
| HBM3E | 9.8 Gb/s | 1,024 bits | ≈ 1.25 TB/s |
| HBM4 | 8 Gb/s | 2,048 bits | 2 TB/s |
56789 The first row is the 1 Gb/s rate first-generation HBM shipped at in 2015; the others are the standards’ maxima, except HBM3E, which is one vendor’s part.
KV-cache bytes per token at 16 bits for the Llama 3 family, from Table 3 of the paper (2 × layers × 8 KV heads × 128 × 2 bytes): 8B, 128 KiB; 70B, 320 KiB; 405B, 504 KiB.20 At a 128K-token context the 70B model’s cache is 40 GiB per sequence, about a third of its 16-bit weights.
Every way around the memory wall costs something:
- Serve more users at once. Total output goes way up. But each user waits a little longer, and everyone’s notes must fit in memory.
- Use smaller numbers. Store each number in 4 bits instead of 16, and there is a quarter as much to fetch. One method like this ran more than three times faster.25 Shrink them too far, though, and the answers get worse.
- Keep shorter notes. Models can be built to keep much smaller notes for each token.2026
- Guess ahead. A small, quick model guesses the next few tokens. The big model checks all the guesses in one trip.24
- Buy more memory or more chips. That works, but it costs money and power.
Batching: throughput versus latency
Batching raises total tokens per second almost linearly while decode is memory-bound, but each user’s per-token time grows a little with every added sequence, and once compute-bound it grows in direct proportion. Serving systems pick a batch that meets a per-user latency target. Two refinements matter in practice:
- . Requests finish at different times. Orca schedules one iteration at a time, so finished requests return immediately and new ones join without waiting for the whole batch.23
- Efficient KV memory. vLLM stores the KV cache in fixed-size pages, like an operating system’s virtual memory, so little space is wasted and more requests fit, for 2–4× throughput.18
Fewer bits per weight means fewer bytes per step. Weight-only 4-bit quantization (AWQ) with its runtime ran more than 3× faster than a standard 16-bit implementation on desktop and mobile GPUs.25 The KV cache can be quantized too: KIVI stores it in 2 bits, cutting peak memory about 2.6× and allowing up to 4× larger batches, for 2.35–3.47× throughput.26 The cost is accuracy, which depends on the method and the model and has to be measured. The Number formats chapter covers the formats themselves.
Smaller KV caches by design
Multi-query attention shares one key/value head across all query heads, which greatly reduces the memory traffic of decoding with only minor quality loss.21 is the middle ground, a few shared KV heads, with quality close to full multi-head attention.22 Both are choices made when the model is trained, so hardware teams can only plan for them.
In the memory-bound regime the math units are mostly idle, so checking several proposed tokens costs about the same as generating one. A small draft model proposes; the large model verifies in one pass and the output is identical to the large model’s own. On T5-XXL this gave 2–3× speedups.24
Splitting the phases
Prefill wants compute and decode wants bandwidth. Splitwise runs them on separate machines, and notes that the decode phase doesn’t need the compute of the latest GPUs.19
More hardware
More or newer HBM stacks raise bandwidth and capacity, at the cost of a bigger, more expensive and harder-to-cool package. Spreading the model over several chips multiplies bandwidth, but adds chip-to-chip communication, the slowest-growing of the three trends.1
What each lever actually moves
| Lever | Changes | Helps most when | Costs |
|---|---|---|---|
| Larger batch | Weight intensity | Short context, weights dominate | Per-token latency; KV capacity |
| Weight quantization | Small batch, latency-bound serving | Accuracy; dequantization work | |
| KV quantization, GQA/MQA | Long context, large batch | Accuracy; retraining (GQA/MQA) | |
| Paged KV, continuous batching | Effective | Mixed request lengths | Allocator and kernel complexity |
| Speculative decoding | Accepted tokens per weight read | Small batch, idle compute | Draft model compute and memory |
| More HBM / more chips | BW and C | Any memory-bound case | Cost, power, packaging, communication |
Failure modes and fine print
- Format and the ridge. If a chip doubles peak FLOPS at each halving of precision and the KV cache follows the weights, the knee batch does not move: bytes and FLOP time shrink together. Quantization buys per-user speed and capacity, not a lower knee.
- Weight-only quantization is a memory optimization. If weights are dequantized to 16 bits for the math, compute time stays at the 16-bit rate, which only matters once batching has used up the memory-bound headroom.25
- Speculation competes with batching for the same idle compute. At large batch the compute is no longer free, and the draft model’s own weights and KV cache take bandwidth and capacity.24
- MQA and GQA change the model. MQA brings some quality loss; converting an existing multi-head checkpoint to GQA takes about 5% of its original pre-training compute.2122
- Partitioning. Spreading a model across chips divides per chip but adds communication that grows in relative importance as chip count rises; for lowest latency the model is partitioned as far as it profitably can be, and small batches then cost more per token.17 See Scale-up fabrics.
- Computing in memory. The radical way around the wall is to stop moving weights at all, covered in Sparsity, in-memory and analog compute.
Step 37.7 ms (memory-bound): 27 tokens/s per user, 212 in total. Memory used 151 GB of 192 GB.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
The decode model, step by step
Notation: parameters, bytes per weight, layers, model width , query heads and KV heads of size (so ), bytes per KV element, batch , context , bandwidth , peak (FLOP/s at the chosen format), capacity . This is the model in the simulation and, with chip count set to 1, the cost model of Pope et al.17
- Weight bytes: .
- KV bytes per token: ; per sequence, .
- Bytes per step: .
- FLOPs per token: (dense matmuls plus the two attention products).
- Step time: . Per-user rate , aggregate .
- Capacity: , so .
The knee
Set memory time equal to compute time:
Solving for :
- With and : , the ridge point times bytes per weight over two. For a ridge of 250 FLOP/B and 16-bit weights, ; at 8-bit with doubled (ridge 500), still 250.
- The denominator is the bytes a sequence “earns” per step () minus the bytes it costs (). When , there is no knee: decode is memory-bound at every batch size. With and , the condition is . For a 70B model with 16-bit GQA cache () on a ridge-250 chip, that happens above about tokens.
- The usable regime is . Above , aggregate throughput saturates at tokens/s and per-user throughput falls as ; above , the batch doesn’t fit.
Worked example: the simulation’s defaults
70B parameters, 8-bit weights and KV, 80 layers, 8 KV heads × 128, , , (8-bit), :
- ; ; per sequence.
- : memory time 17.5 ms, compute time 0.07 ms, so about 57 tokens/s. : 70.7 GB per step, 17.7 ms, about 57 tokens/s per user and 450 in total, math units busy 3% of the step.
- ; ; the compute roof is about 14,100 tokens/s.
What the model leaves out
- Sustained versus peak. Real kernels reach a fraction of peak bandwidth and FLOPS, and memory and compute overlap imperfectly, so measured throughput is below both roofs and the knee is rounded, not sharp.
- Communication. With the model on chips, , and per chip divide by , but collective operations add time that shrinks slowly with . The PaLM 540B example above (7 ms bound, 29 ms measured) shows how much that and other overheads can add.17
- Mixed phases. Real servers often batch prefill work together with decode steps, which raises the step’s intensity; or they separate the phases onto different machines.19
- Ragged contexts. Sequences in a batch have different lengths; is a sum over sequences, not . Allocator waste lowers the effective .18
- Activations and logits. Small compared with at modest batch, ignored here; the vocabulary projection is counted inside .
- Mixture-of-experts. Only the routed experts’ weights do work per token, but across a large batch most experts can get used, so bytes per step can approach the full parameter count while FLOPs per token follow the active count. The model above would need separate for bytes and for FLOPs.
Roofline view
Dividing through by bytes gives the roofline form:3
rises with and saturates at , the per-sequence ceiling set by the KV cache. The knee exists only if exceeds the ridge , which is the same condition as above. Every mitigation in the trade-offs table either raises (batching, smaller , smaller , speculation) or raises the slope (more or faster HBM, more chips, SRAM residency).
Batch 1: 2.0 FLOP/B, 8 TFLOPS (0.40% of peak). The KV cache caps intensity at 1,685 FLOP/B; the knee is at batch 352.
Q1A 70-billion-parameter model stored in 16-bit format runs at batch size 1 on a chip with 4 TB/s of memory bandwidth. Roughly how many tokens per second can one user get at best?
Q2What is the KV cache?
Q3Going from 16-bit to 8-bit weights roughly doubles single-user decode speed on a memory-bound chip. Why?
Q4Why is HBM placed right next to the processor on a silicon interposer instead of on the circuit board like ordinary memory chips?
Sources
Show Hide 26 sources
- AI and Memory WallPeak server FLOPS grew 3.0×/2 yrs (60,000× in 20 years) against 1.6×/2 yrs for DRAM bandwidth (100×) and 1.4×/2 yrs for interconnect (30×); Wulf and McKee coined “memory wall” in 1995; model size grew 410×/2 yrs against 2×/2 yrs for accelerator memory; decoder inference is bandwidth-bound at low batch.
- Efficient Processing of Deep Neural Networks: A Tutorial and SurveyFig. 22: normalized energy per access is 1× for a register file, 2× from a neighboring processing element, 6× for a global buffer and 200× for DRAM; DRAM costs two orders of magnitude more energy per access than a small on-chip memory.
- Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore ArchitecturesAttainable performance = min(peak floating-point performance, peak memory bandwidth × operational intensity).
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessA100: 40–80 GB of HBM at 1.5–2.0 TB/s and 192 KB of SRAM in each of 108 SMs at about 19 TB/s; operations are increasingly bottlenecked by HBM accesses; tiling cuts HBM–SRAM traffic for attention.
- AMD’s Next Generation GPU and High Bandwidth Memory Architecture: FURYFirst-generation HBM: DRAM dies stacked on a base die, joined by TSVs and µbumps, on an interposer beside the GPU; per package a 1,024-bit bus at 1 Gb/s for over 100 GB/s per stack, against a 32-bit GDDR5 chip at 7 Gb/s; Fury’s four stacks give a 4,096-bit interface and 512 GB/s.
- JEDEC Updates Groundbreaking High Bandwidth Memory (HBM) StandardJESD235B (December 2018): 2.4 Gb/s per pin, up to 307 GB/s across a 1,024-bit interface divided into 8 channels; up to 24 GB per stack. The page sits behind a bot check but opens normally in a browser.
- JEDEC Publishes HBM3 Update to High Bandwidth Memory (HBM) StandardJESD238 HBM3 (January 2022): up to 6.4 Gb/s per pin, 819 GB/s per device; 16 channels, each with 2 pseudo-channels. The page sits behind a bot check but opens normally in a browser.
- Samsung Electronics Holds Memory Tech Day 2023: Unveiling New Innovations To Lead the Hyperscale AI EraHBM3E “Shinebolt”: 9.8 Gb/s per pin, more than 1.2 TB/s per stack.
- JEDEC® and Industry Leaders Collaborate to Release JESD270-4 HBM4 Standard: Advancing Bandwidth, Efficiency, and Capacity for AI and HPCHBM4 (April 2025): up to 8 Gb/s across a 2,048-bit interface, up to 2 TB/s; independent channels doubled from 16 (HBM3) to 32, each with 2 pseudo-channels; 4- to 16-high stacks of 24 Gb or 32 Gb dies, up to 64 GB (32 Gb, 16-high). The page sits behind a bot check but opens normally in a browser.
- NVIDIA Hopper Architecture In-DepthH100 SXM5 (preliminary launch specs): 80 GB HBM3 in 5 stacks on ten 512-bit controllers, over 3 TB/s, 50 MB L2, 1,000 dense BF16 TFLOPS; A100: 40 GB, 1,555 GB/s, 40 MB L2, 312 dense BF16 TFLOPS.
- AMD Instinct MI300 Series microarchitectureMI300 Series: up to 8 compute dies and 8 HBM3 stacks; MI300X peak memory bandwidth 5.3 TB/s; 1,307.4 TFLOPS matrix BF16/FP16 and 2,614.9 FP8.
- AMD GPU specificationsMI300X: 192 GiB VRAM, 256 MiB L3 cache, 32 MiB L2 (4 per compute die).
- TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsTable: TPU v4 275 peak TFLOPS (bf16), 7 nm, under 600 mm², 128 MiB CMEM + 32 MiB VMEM, 32 GiB HBM2 at 1,200 GB/s; TPU v3 123 TFLOPS, 32 GiB at 900 GB/s; CMEM gives 1.2× overall on production workloads but 2× for RNN1, whose small weights and small batch benefit from CMEM bandwidth.
- IEDM 2022 – TSMC 3nmFallback source (the IEDM papers are paywalled): TSMC N3 high-density SRAM cell 0.0199 µm²; the N3E paper discloses 0.021 µm², larger than N3’s; N3E is believed to relax some pitches for performance and yield.
- Intel, TSMC Tout SRAM Breakthroughs At 2nmFallback source (no open primary covers the stall): SRAM scaling stalled at the 3 nm node in 2022, which affected the die sizes and costs of high-performance chips such as GPUs and processors; it came back to some degree at 2 nm.
- Intel, Synopsys, TSMC All Unveil Record Memory DensitiesISSCC 2025: Intel 18A and TSMC N2, the two companies’ first nanosheet processes (Samsung moved a generation earlier), reach 38.1 Mb/mm² in their densest SRAM blocks; SRAM has been particularly hard to shrink.
- Efficiently Scaling Transformer InferenceWeights and KV cache are read from HBM once per step (“memory time”); 2N FLOPs per token (“compute time”); weights dominate at small batch and short context, KV cache at large batch and long context; a 500B+ model’s KV cache reaches 3 TB at batch 512 and context 2,048; multiquery attention shrinks it by the number of heads; 29 ms/token for PaLM 540B on 64 TPU v4 chips with int8 weights.
- Efficient Memory Management for Large Language Model Serving with PagedAttentionGeneration is memory-bound; for a 13B model on a 40 GB A100, about 65% of memory holds weights and about 30% the KV cache; OPT-13B needs 800 KB of KV cache per token and up to 1.6 GB per 2,048-token request; earlier systems used only 20.4–38.2% of KV memory for real tokens; vLLM gives 2–4× throughput.
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingInference has a compute-intensive prompt phase and a memory-intensive token-generation phase; generation does not need the compute of the latest GPUs.
- The Llama 3 Herd of ModelsTable 3: 8B/70B/405B have 32/80/126 layers, model dimension 4,096/8,192/16,384, 32/64/128 attention heads and 8 key/value heads; GQA with 8 KV heads reduces the KV cache.
- Fast Transformer Decoding: One Write-Head is All You NeedMulti-query attention shares one key/value head across all query heads, cutting memory bandwidth in incremental decoding, with minor quality loss.
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head CheckpointsGrouped-query attention uses more than one but fewer KV heads than query heads; uptraining a multi-head checkpoint takes 5% of the original pre-training compute.
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsIteration-level scheduling: the batch is re-formed every iteration, so finished requests leave and new ones join without waiting for the whole batch.
- Fast Inference from Transformers via Speculative DecodingA small model drafts several tokens and the large model checks them in parallel, with identical outputs; 2–3× faster on T5-XXL.
- AWQ: Activation-aware Weight Quantization for LLM Compression and AccelerationLow-bit weight-only quantization (4-bit); its TinyChat runtime is more than 3× faster than the FP16 Hugging Face implementation on desktop and mobile GPUs.
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache2-bit KV cache (keys per channel, values per token): about 2.6× less peak memory, up to 4× larger batch, 2.35–3.47× throughput.