is the day a chip design leaves the design team and goes to the factory. Before it, the team draws and checks. After it, the factory builds. A few months later, the first chips come back.
The name comes from the past. The finished files once went to the factory on reels of tape.1 Today they travel over a network, but chip teams still call that day “tapeout.”
What gets sent is one file. It holds a drawing of every layer of the chip, stacked like the floors of a building. Tiny switches called sit at the bottom, with many layers of metal wires on top.2 The last weeks before tapeout are about making that file complete and correct.
By this stage the chip exists as a layout: an exact drawing, layer by layer, of every transistor and every wire, with shapes measured in billionths of a meter. The layout has also passed , the final round of checks that it is fast enough, draws acceptable power and obeys the factory’s rules (covered in the Signoff chapter).
is the release of that layout to the , the company that will manufacture the chip.1 What the foundry receives is one file, in a format called or its newer, more compact successor, . GDSII is usually the final output of the whole design process and the file the foundry builds from.2 Getting that file right takes a short stage of its own:
- Assemble. Until now, the small prebuilt building blocks of the chip were handled as outlines. Their full drawings are now put in.
- Fill. Add , small unconnected squares of metal, so every patch of every layer has a similar amount of metal. The factory needs this to polish each layer flat.
- Check again. Rerun the rule checks and the circuit comparison on the complete file, because only now does it contain every shape.
- Sign and send. Close a checklist, freeze the file and transfer it.
The foundry then does work of its own before any wafer moves. It adds structures such as a seal ring around the chip, arranges copies of the chip with test patterns on the mask plate, and reshapes the drawing so it prints correctly. A design can still bounce back at this point if it fails the foundry’s checks. And the first chips are rarely the last: most designs go through more than one manufacturing round, called a “spin,” before production.1 How the factory then builds the chip layer by layer is the subject of the Transistors guide’s Making them chapter.
You know a design ends as a layout. Tapeout turns that layout into the one file the foundry will build from, and freezes it. That file is the only thing the factory sees, which has three blunt consequences:
- Anything missing from the file does not exist on the chip.
- A block still represented by its outline (an abstract) instead of its real drawing becomes an empty hole on the masks.
- A shape tagged with the wrong layer number lands on the wrong mask, or on none.
So the work is half physical checking and half bookkeeping: one top-level cell, consistent units, unique cell names, the foundry’s layer numbering, fill kept separate from real wiring, and a record of exactly which rule files, libraries and third-party blocks the clean results came from. This chapter walks through each.
Data volume is the second theme. Fill on every metal layer, and later the foundry’s corrections to the shapes, multiply the number of shapes. GDSII files have grown to many tens of gigabytes, and GDSII has only two ways to avoid repeating itself: referencing a cell, and placing a cell in a regular array.5 That pressure is why OASIS was created.54
The third theme is cost. A full mask set at 7 nm was estimated at $8–10 million in 2016,17 and a wafer spends about 12 weeks in the fab on average, 14–20 weeks at the most advanced processes: three to five months.2324 A mistake that escapes tapeout costs both, which is why the last checks are so careful.
Design: The layout is still being finished and checked. Any shape can still change. Mistake found now: edit the layout and rerun the checks. Days of work.
- Mask set cost estimate, 28 nm (2016)
- ~$2M
- Mask set cost estimate, 7 nm (2016)
- $8–10M
- Fab time, industry average
- ~12 weeks
- Fab time at the most advanced processes
- 14–20 weeks
Sources: Baas, citing AMD’s CTO in 2016,17 Wikipedia on semiconductor fabrication,23 and the Semiconductor Industry Association (2021).24
The layout still has some parts drawn as empty boxes. So their full drawings come in too, with the factory’s rules. Out comes one finished file.
| Direction | What | Typical format |
|---|---|---|
| In | Routed top level from place and route | DEF or GDSII, with cells as abstracts |
| In | Full layouts of standard cells | GDSII/OASIS from the library vendor |
| In | Hard IP: memories, input/output pads, clock generators, analog blocks | GDSII/OASIS (plus the LEF abstract used during place and route) |
| In | Layer map and fill rules | Text tables; fill rule file |
| In | Rule files for DRC, LVS, density and antenna checks | Checking-tool scripts (“decks” or “runsets”) |
| In | Reference circuit for LVS | Verilog netlist plus each cell’s transistor netlist (SPICE/CDL) |
| Out | Merged, filled final layout | GDSII (chip_top.gds) or OASIS (chip_top.oas) |
| Out | Final DRC, LVS and density results, plus approved exceptions (waivers) | Reports, error-marker databases |
| Out | Signed tapeout checklist and release record | Checklist, checksums, tool and rule-file versions |
A few words in that table need unpacking.
- Place and route is the earlier stage that positioned every building block and drew the wires between them. It saves its result as DEF, a text file listing each block’s position and every wire, or streams it straight out as GDSII.
- Standard cells are the chip’s small prebuilt building blocks: logic gates and one-bit storage elements, each a few transistors, supplied as a library. Hard IP blocks are bigger prebuilt pieces, such as memories, bought or reused with their layout already finished.
- The netlist is the list of every cell in the design and the wires connecting them. LVS compares the layout against it.
- Antenna rules limit long stretches of wire connected to a transistor’s gate, which can damage the gate while the chip is being made. Small protective diodes are added to fix violations.21
The key word in the table is . During place and route, a standard cell or a memory block is only its outline: its size, its connection points (pins) and the areas wires must avoid, stored in a format called . The transistors and internal wiring live in separate GDSII files from the library and IP vendors. Tapeout is where the two views come together. The open-source layout viewer KLayout can do it in one step: it reads the DEF and swaps each LEF outline for the real layout of the same name from a list of GDSII or OASIS files, using a map file to number the layers.8
Most merge failures come from mismatched conventions rather than wrong shapes. Three to check before anything else:
- Units. GDSII stores every coordinate as a whole number of “database units,” and a
UNITSrecord says how long one unit is.3 A block written with a different unit must be scaled when it is read in, or it lands at the wrong size. - Layer numbers. Every block must use the foundry’s layer map, or its shapes end up on the wrong masks.
- Cell names. Names must be unique across all sources. If two vendors both ship a cell called
via_1x2, either one silently replaces the other or the file is invalid, depending on the tool.7
The open-source mpw_precheck, written for SKY130 shuttle projects built on the Caravel harness, tests for related problems, such as shapes off the manufacturing grid and changes outside the designer’s area.22
Six kinds of input, three outputs. Tap a file to see what it is and why it matters.
The last stretch has four steps.
- Fill in the boxes. While the chip was laid out, some parts were drawn as empty boxes, like furniture drawn as rectangles on a floor plan. Now their full drawings go in. The file finally holds every switch and every wire.
- Even out the metal. The factory builds the chip in layers on a , a thin slice of silicon about the size of a dinner plate. After each layer, a machine polishes the wafer flat.13 But spots with lots of metal and spots with little wear down at different speeds. So software adds tiny squares of unused metal to empty spots. This is called . Spots with too much metal get holes cut in them instead.12
- Check again. The finished file goes through all the checks one more time. Only now does it hold every shape that will be built.
- Sign and send. Each person in charge of a check signs off. The file is locked, and it goes to the factory.
What’s inside a GDSII file
GDSII is a binary file format (written for programs to read, not people) developed by the company Calma and released in 1978.2 A file is a library of named cells. Each cell holds shapes (mostly polygons, plus wire paths and text labels) and references to other cells, either one copy at a time or as a regular grid of copies.3 This nesting, called hierarchy, keeps files manageable: a standard cell is drawn once and placed a million times by reference, and one cell at the top contains everything else.
Coordinates are whole numbers counted in a tiny “database unit,” and the file states how long that unit is.3 One nanometer is a common choice, and it is the default in the open-source gdstk library.7
Every shape carries a , each from 0 to 255 in the original specification.3 The numbers mean nothing by themselves; the system reading the file decides what they stand for.6 For a given manufacturing process, the foundry’s fixes them. In an illustrative map, wires on the first metal layer (M1) are 15/0 (layer 15, datatype 0) and dummy fill on M1 is 15/1. Keeping fill on its own datatype lets every later tool tell real wiring from dummy metal.
Put together, the file is a stack of flat drawings, one per layer. A layout viewer draws each layer in its own color; lift each one to its real height in the chip and the file turns back into a picture of the chip itself.
(SEMI standard P39, version 1.0 in 2004) carries the same information more compactly.4 It stores small numbers in fewer bytes, describes a rectangle by its width, height and corner where GDSII lists every corner, lets a shape leave out values such as the datatype when they repeat the previous shape’s, and stores a set of identical shapes as one shape plus a list of positions. Each cell can also be compressed separately.4 Fill, with millions of identical squares, shrinks especially well.5
Merging the final file
The place-and-route tool knows standard cells and IP blocks only as abstracts. To produce the final file, a merge step replaces each abstract with the full layout of the same name from the vendor’s GDSII file. In OpenROAD-flow-scripts, an open-source flow, KLayout does this merge (and the final rule checks on public processes).10 Small scripts can do the same: the gdstk library’s replace swaps cells by name and updates every reference to them, and remap renumbers layers and datatypes.7 After the merge, two checks: exactly one top cell should remain, and no cell should still be empty, because an empty abstract becomes a hole in the chip.
Metal fill and density rules
A chip is built layer upon layer. After each layer, grinds the wafer flat with a rotating pad and a chemical slurry, so the next layer has a level surface to build on.13 How much material the polish removes depends on how densely the layer below is patterned, so foundries set density rules on many layers. A 1998 paper gives a typical example: every 10 × 10 µm square of a metal layer must contain between 35 µm² and 70 µm² of metal, that is, 35% to 70% coverage. Uneven density leads to dishing (wide metal scooped out in the middle) and other irregularities, and very wide wires such as power lines must be slotted so they don’t lift off during polishing.12
A checker steps a across the whole chip and measures the coverage in each position. KLayout’s rule-checking language does this with with_density, which reports every window whose density falls in a given range, for a given window size and step.9 A fill tool then adds unconnected metal shapes in the sparse windows, keeping a safe distance from real wires. OpenROAD’s density_fill, for example, reads a rule file that gives, for each layer, the layer and datatype to put fill on, the fill shape sizes to try, and the spacing to other fill and to real wires.11 Windows that are too dense need the opposite fix: cuts holes into wide metal.12
Fill is not free. Metal near a wire adds capacitance, extra charge the wire must move every time its signal switches, which can slow the signal and couple noise into it.14 So fill goes in before the final timing check, and designers can mark areas, such as sensitive analog circuits, where no fill is allowed.9
Final checks on the merged file
Two checks from signoff run once more. (design rule checking) tests every shape against the foundry’s manufacturing rules, such as minimum widths and spacings. (layout versus schematic) extracts the circuit from the drawn shapes, transistor by transistor and wire by wire, and confirms it matches the netlist. At tapeout they run on the merged, filled file with the foundry’s current rule files. This is the first run that sees the inside of every cell and block, every fill shape, and every place where one block’s edge meets its neighbor.
The open-source mpw_precheck, built for shared manufacturing runs on the SKY130 process, automates a similar list: DRC with two different tools (Magic and KLayout); an XOR, a shape-by-shape comparison proving nothing changed outside the designer’s area; a check that every corner sits on the manufacturing grid; a metal density check; and LVS.22
The checklist and the sign-offs
Tapeout ends with a checklist that ties every result to one exact file. A typical list:
- DRC, LVS, antenna and density checks clean on the final file.
- Timing signed off with fill included: every signal path still arrives before its clock deadline.
- Power delivery signed off: the supply voltage doesn’t sag too far across the chip (IR drop), and no wire carries enough current to wear out over the chip’s life (electromigration).
- passed: a formal proof that the final netlist still computes the same logic as the design that was verified.
- Every waiver (an approved exception to a rule) signed by its owner.
- Tool, rule-file and library versions recorded, plus the file’s checksum: a short fingerprint computed from its contents, which changes if even one byte changes.
The owners of the layout checks, the timing and the project each sign their part, and the file goes to the foundry over a secure transfer.
What’s inside a GDSII file
A GDSII file is a library of named cells (the format calls them structures). Each cell holds elements: polygons (boundaries), wire paths, text labels, and references to other cells, either a single placement (SREF) or a regular grid of placements (AREF). Coordinates are 4-byte integers in database units, and every shape carries a layer and a datatype from 0 to 255.3 The numbers mean nothing until the foundry’s layer map assigns them.6 The exact byte layout is in “Under the hood.” What matters here is that the file is a stack of flat drawings, one per layer and datatype, held together by cell references. Draw each layer at its real height in the chip and the file becomes a 3D picture of the die.
Database hygiene
Most tapeout-week emergencies are bookkeeping, not geometry. Five things to check, and why each goes wrong:
- Units. A block streamed at a different database unit must be scaled on import, or it lands at the wrong size. Libraries such as gdstk track both a user unit and a precision for exactly this reason.7
- Grid. Every corner must sit on the manufacturing grid. Rotating or scaling a block, or drawing at angles other than 90°, can push corners off it.22
- Names. A hierarchical merge needs unique cell names, so IP is often delivered or renamed with a prefix.
- Labels. LVS finds the chip’s pins by reading text labels in the file, so a missing or misplaced top-level label shows up as an LVS mismatch at that pin.
- Layers. List every layer/datatype pair in the merged file and compare it with the layer map. An unexpected pair from an IP delivery is either junk or a mask that will be missing.7
How fill is decided
The simplest approach is rule-based: find every empty area big enough and fill it with a fixed pattern. It is fast, but the gap between the densest and sparsest windows usually stays fairly large, and the result isn’t tied to any model of how polishing actually behaves.14 Model-driven fill works in two steps instead: first compute how much fill each small tile of the layout needs, then place real shapes to deliver it.13 “Under the hood” walks through those algorithms.
Whatever the method, fill has to be visible to extraction, the step that computes each wire’s resistance and capacitance from the layout. A 1998 paper warned that without an accurate estimate of fill and slotting, extraction, delay calculation and timing and noise analysis would all be inaccurate.12 The effect is smaller than it sounds for tightly packed wires, whose capacitance is dominated by their same-layer neighbors, and small fill shapes add less total capacitance than large ones.14 Two practical consequences follow. First, fill must run before final extraction and timing, and any later fix that removes wiring needs local refill and re-extraction. Second, fill is a data problem: a flat fill layer on every metal can dominate the file, so fill is placed as arrays of a few fill cells, which GDSII stores as AREFs and OASIS as repetitions.5
What the foundry checks
After submission the foundry runs its own intake checks and its own chip finishing, and a design can fail there for manufacturing reasons before a single wafer starts.1 Shared-run services publish theirs: mpw_precheck also ran on its shuttle provider’s platform, apart from LVS, which was enabled only for local runs.22 The design team’s defense is to run the foundry’s exact rule-file versions, keep the waiver list short and documented, and make the file self-describing: one top cell, a recorded layer list, and checksums for every file in the release.
The routed layout from place and route. Every standard cell and the memory block is still an abstract: an outline with pins.
The factory turns each layer of the drawing into a : a glass plate that works like a stencil. Light shines through it to print that layer’s pattern onto the wafer. A chip needs a whole set of them, often dozens, used one after another.16 The shapes are so small that light blurs them. So the factory bends the stencil shapes on purpose to make them print right.18
Masks cost a lot. One estimate put a full set at about $2 million for older chips and up to $10 million for the most advanced ones.17 That is why a mistake found in the first chips hurts so much. Fixing it means new masks and another trip through the factory.
Small teams, schools and hobbyists split the cost instead. On a , many designs share the same masks and wafers, like friends carpooling. Each pays only for its own space.20 A project called Tiny Tapeout goes further. Hundreds of small designs share one chip,20 and each gets a tile smaller than a grain of salt.21
Then comes the wait. Tiny Tapeout tells people to expect six to nine months.21 When the first chips arrive, engineers switch them on and test them. The After tapeout page tells that story. The Transistors guide’s Making them chapter shows how the factory builds a chip. Tapeout day itself is worth a party: years of work just left the building.1
Mask data preparation
Chips are printed with light, through : quartz plates that each carry the pattern of one layer. A chip needs a set of them, at least one per patterned layer, used one after another.16 Turning the layout into those masks starts with , which the foundry does in three parts:15
- Chip finishing adds structures such as a seal ring around the chip’s edge and fill.
- Reticle layout arranges one or more copies of the chip on the , the mask plate loaded into the printing machine, together with test patterns and alignment marks.
- Layout-to-mask preparation corrects the shapes so they print well, then breaks every polygon into the rectangles and trapezoids the mask-writing machine can draw, a step called fracturing.
The main correction is . The features are so small that the light printing them blurs: lines come out narrower or wider than drawn, corners round off and line ends shrink back. OPC pre-distorts the shapes to compensate, for example moving edges, widening line ends into “hammerheads,” and adding thin assist bars beside isolated lines, too narrow to print themselves, that help those lines print like their crowded neighbors.18
Making the masks
Each mask is written into a light-sensitive coating with an electron beam or a laser, developed, and etched into a chrome film. It is then measured for feature size and placement, inspected for defects, repaired, cleaned and covered with a pellicle, a thin film that keeps dust off the pattern.17
Cost, and why respins hurt
Mask cost climbs steeply as features shrink: roughly $2 million at 28 nm, $4 million at 14/16 nm and $8–10 million at 7 nm in a 2016 estimate. Finer features and multiple patterning, which needs several masks for one layer, drive the increase.17 Masks are a large part of a chip’s (non-recurring engineering), the one-time cost paid before the first chip is sold. A pays for masks again and adds another full trip through the fab: about 12 weeks on average, and 14–20 weeks at the most advanced processes.2324
Multi-project wafers
A , or shuttle, shares the cost of masks and wafers among many designs. MOSIS, set up by the US defense research agency DARPA, began in 1981; current programs include CMC Microsystems in Canada, EUROPRACTICE in Europe, Muse Semiconductor and Tiny Tapeout.20 Tiny Tapeout connects up to 512 designs on a shared chip through a multiplexer, a switch that selects one design at a time.20 Each participant gets a tile of about 160 × 100 µm on the open SKY130 process, and an automated online workflow runs the open-source OpenLane flow that turns each design into a GDSII file.21 The catch: each provider sets its own schedule, design rules, die sizes and recommended tools, and you get a limited number of chips.20
First silicon and metal-only respins
When the wafers come back, the team brings up the first chips in the lab (the After tapeout page covers this). If a bug turns up, the cheapest fix changes only the metal wiring layers, because a full rebuild needs new masks for every layer.25 That is only possible if the design already contains unused gates to rewire. , scattered across the chip when the cells were placed, are exactly that: a (engineering change order, a late fix) rewires them into the logic without moving any transistor.26
Tapeout is also a milestone worth marking. Teams usually celebrate the day the file leaves, because it closes months or years of work and starts the wait for silicon.1
What the foundry does to the shapes
After intake, mask data preparation adds a seal ring and other finishing structures, lays out the reticle with test patterns and alignment marks, applies resolution enhancement, and fractures every polygon into rectangles and trapezoids for the mask writer.15 The resolution enhancement has evolved in three steps:
- Rule-based OPC looks up a correction in a table indexed by feature width and spacing. It is fast and gives simple masks, but it only corrects local effects and can’t find the best mask overall.19
- Model-based OPC simulates how each shape will print and moves its edges until the simulated result matches the target. For the critical layers of a full chip that takes large compute farms.1819
- treats the mask as an unknown to solve for, often pixel by pixel. It is more precise, but its complex curved shapes make masks harder and costlier to write.19
Data volume grows at every step. The 2002 industry roadmap projected that the fractured data for a single critical layer would reach hundreds of gigabytes at the transition from 130 nm to 90 nm.5
Respin economics
A full respin pays for a new mask set, $8–10 million at 7 nm in 2016,17 plus another three to five months of fab time.2324 A metal-only respin replaces only a few masks, typically the metal and via layers, and reuses the expensive transistor layers.2526 Teams buy that option before the first tapeout by scattering spare cells over the design during placement. It only works if there are enough spares near the logic that needs fixing: when they are too few or too far away, the patch ends up with long wires, timing violations and routing congestion.26
Choosing a shuttle
A shuttle’s turnaround and cost depend on the process, and each provider sets its own design rules, die sizes, tools and schedule.20 Tiny Tapeout, at the small end, sells tiles of about 160 × 100 µm on SKY130 and quotes six to nine months to manufacture before boards are assembled and tested.21 Shuttles suit test chips that characterize a new block, risky analog circuits before a full product tapeout, research and teaching. Parts come back in small numbers, so plan the package and test board early.
Mask = drawn shapes. Blurring pulls the line ends back, rounds the outer corner and fills the inner one: about 29% of the target area prints wrong.
A dedicated mask set: one design carries the whole $2M.
Each square is a patch of the chip’s first metal layer. Press “Run checks” to measure how much metal each patch holds. Some patches have too little, and one has too much. Both would polish unevenly.
Press “Insert fill” to add metal fill. It only fixes the empty patches. The crowded patch has a wide power wire in it, so press “Slot wide metal” to cut holes in that wire. Check again. When every patch passes, the checklist is done and the file is ready to send. “Reset” starts over.
The chip is split into 8 × 6 density windows on the first metal layer (M1), and the rule is that every window must be between 20% and 80% metal (illustrative numbers). “Run checks” flags the windows below 20% and the one above 80%. “Insert fill” adds dummy squares to the sparse windows, keeping clear of real wires, until each reaches about 35%. Notice that fill fixes only the sparse windows; the dense window stays flagged until “Slot wide metal” cuts slots in its wide wire and brings it under 80%. Rerun the checks after each change: fill and slots alter the layout and the wire capacitance, so the DRC and timing items untick until you run the checks again on the new layout. The tapeout checklist (DRC clean, LVS clean, antenna clean, timing signed off, density clean) is complete when every item passes, and the export panel shows the file name chip_top.gds.
The Expert view adds the per-window numbers and the layer map: M1 wires on 15/0, fill on 15/1 and slots on 15/2. Before pressing anything, think about three questions. Why fill to about 35% rather than just above 20%? (Margin, in case the foundry’s checker measures slightly differently from yours, and a smaller spread in density across the chip.) Why do fill and slots go on their own datatypes? (So extraction, LVS and mask data preparation can treat them differently from real wires.) And what does slotting the wide wire do to its resistance, voltage drop and current density? Slotting shrinks the wire’s cross-section, so the power analysis has to be rerun.12
Picture the last week. On a big screen is the tapeout checklist: one row per check, each with an owner. An engineer opens the final file and zooms from the whole chip down to single switches. They make sure every part is really there, not an empty box.
When the last row turns green, everyone signs. Someone saves a fingerprint of the file, a short code that changes if even one tiny detail changes. That way nobody sends the wrong version by accident. Then the file is uploaded to the factory.
First, an illustrative layer map for a generic process. Real numbers come from the foundry; these match the simulation above. “Drawing” means real, intended shapes; V1 is the layer of vias, the vertical plugs that connect M1 to M2.
| Layer | Purpose | GDS layer/datatype |
|---|---|---|
| M1 | Drawing (signal and power wires) | 15/0 |
| M1 | Fill (unconnected dummy metal) | 15/1 |
| M1 | Slot (holes cut into wide metal) | 15/2 |
| M1 | Pin label (text, read by LVS) | 15/10 (texttype) |
| V1 | Drawing | 16/0 |
| M2 | Drawing | 17/0 |
| M2 | Fill | 17/1 |
| Fill keep-out | Marker for the fill tool, never made into a mask | 200/0 |
Next, a short merge script in Python using the open-source gdstk library. It reads the routed design, in which every cell is still an empty abstract, swaps in the real cell layouts, fixes one supplier’s numbering, runs two sanity checks and writes the final file in both formats.7
# merge_gds.py (illustrative, gdstk; generic layer numbers)
import gdstk
lib = gdstk.read_gds("results/chip_top_routed.gds")
stdcells = gdstk.read_gds("lib/stdcells.gds")
sram = gdstk.read_gds("ip/sram_64x32.gds")
needed = {c.name for c in lib.cells}
lib.replace(*[c for c in stdcells.cells if c.name in needed])
lib.replace(*sram.cells)
empty = [c.name for c in lib.cells if not (c.polygons or c.paths or c.references)]
assert not empty, f"cells still empty after merge: {empty}"
lib.remap({(15, 7): (15, 1)})
top = lib.top_level()
assert len(top) == 1, f"expected one top cell, got {[c.name for c in top]}"
print(sorted(lib.layers_and_datatypes()))
lib.write_gds("out/chip_top.gds")
lib.write_oas("out/chip_top.oas")- 1L4The routed design exported from place and route. Standard cells and the memory are empty cells that carry the right names.
- 2L5The full transistor-level layouts of the standard cells, from the library supplier.
- 3L6A memory block (SRAM) delivered as GDSII by its supplier.
- 4L8Collect the names of the cells the design actually uses. Bringing in the whole library would leave hundreds of unused cells floating in the file.
- 5L9replace swaps each cell for the one with the same name and updates every reference to it. If two suppliers used the same name, the wrong cell would be swapped in without any warning.
- 6L12Any cell that is still empty is an abstract that never got its real layout: it would be a hole in the chip.
- 7L15This memory supplier put M1 fill on datatype 7. The foundry’s layer map wants 15/1, so renumber it.
- 8L17Cells that nothing references are ‘top’ cells. A stray test cell or unused library cell shows up here; the foundry expects exactly one.
- 9L19List every layer/datatype pair in the file. Compare it with the layer map: an unexpected pair is junk or a missing mask.
- 10L21Write GDSII for tools that need it, and OASIS for a much smaller file to transfer.
An illustrative final tapeout checklist, kept as the release record.
TAPEOUT CHECKLIST chip_top rev B0 (illustrative)
File: chip_top.oas SHA-256 9f3c...e21a top cell: chip_top
Units: 1 um user, 1 nm database Layer map: generic_10M v3.2
PHYSICAL VERIFICATION (merged, filled file) owner status
DRC, foundry deck v1.7 0 violations PV PASS
LVS vs chip_top.cdl CORRECT PV PASS
Antenna 0 violations PV PASS
Density, all layers, 50 um windows 20%-80%, clean PV PASS
Waivers: 3 (ESD clamp spacing, IP vendor-approved) PV SIGNED
TIMING AND POWER (extracted with fill)
STA, all signoff scenarios WNS +0.004 ns STA PASS
IR drop static/dynamic within budget PI PASS
EM, 10-year target clean PI PASS
LOGIC
Equivalence: final netlist vs synthesis EQUIVALENT LEC PASS
RELEASE
IP versions: sram_64x32 r4, io_ring r2, pll r1 Lead CHECKED
Tool and deck versions recorded Lead CHECKED
Signed: PV lead, timing lead, PI lead, project lead Lead DONE- 1L2The SHA-256 checksum, a fingerprint of the file’s contents, ties every result below to one exact file. Any later edit, even to a text label, changes the checksum and means rerunning the checks.
- 2L3Database unit and layer map version. A wrong unit scales the whole chip; a wrong map moves shapes to the wrong mask.
- 3L5Owners: PV is physical verification (the layout checks), STA is static timing analysis, PI is power integrity, LEC is logic equivalence checking, Lead is the project lead.
- 4L6DRC run on the merged file with the foundry’s current rule file (“deck”). A deck update can turn a clean layout dirty.
- 5L7LVS compares the layout against a CDL netlist, the final gate-level netlist expanded down to each cell’s transistors.
- 6L9Density after fill and slotting, for every layer that has a density rule.
- 7L10Each waiver, an approved exception to a rule, has a reason and an owner. Here the exceptions sit inside a protection circuit (ESD clamp) and are approved by the block’s supplier.
- 8L12WNS, worst negative slack, is the margin of the tightest signal path. Positive means every path meets its clock deadline, here with 4 ps to spare, checked with fill in place because fill adds capacitance.
- 9L13IR drop: how far the supply voltage sags across the chip, both on average (static) and during bursts of switching (dynamic).
- 10L14EM, electromigration: no wire carries enough current to wear out within the target lifetime.
- 11L16A formal proof that the final netlist still computes the same logic as the synthesized design.
- 12L18Exact versions of every bought-in block matter: an older memory view with a known bug is a classic escape.
A KLayout rule script that adds M1 fill only where density is low, keeping clear of wires and of marked keep-out areas. It is written in KLayout’s Ruby-based rule language; units are µm.9
# fill_m1.drc (illustrative KLayout DRC; generic layers, units in um)
source($input)
target($output)
m1 = input(15, 0)
keepout = input(200, 0)
core = extent.sized(-5.0)
low = m1.with_density(0.0 .. 0.35, tile_size(50.um), tile_step(25.um))
space = (low & core) - m1.sized(0.4)
pattern = fill_pattern("FILL_M1").shape(15, 1, box(0.0, 0.0, 0.5, 0.5))
space.fill(pattern, hstep(1.0), vstep(1.0), auto_origin, fill_exclude(keepout))- 1L3fill writes cells into a target layout; it does not work with a report as output.
- 2L6Keep-out areas drawn by designers over analog blocks and timing-critical wires.
- 3L7The whole layout’s outline, shrunk by 5 µm, so fill stays clear of the chip’s edge, its seal ring and pads.
- 4L9Windows of 50 µm, stepped by 25 µm so they overlap, whose M1 density is under 35%. Fill goes only where it is needed.
- 5L11The area that may be filled: low-density windows inside the core, at least 0.4 µm from any M1 wire.
- 6L13The fill shape: one 0.5 µm square on M1’s fill datatype, 15/1, in a cell named FILL_M1.
- 7L14A 1 µm pitch covers at most 25% of the empty area, so fill alone can’t push a window over the maximum. auto_origin lets the tool pick the best starting point for each empty island.
After fill, rerun density, DRC and extraction. The new shapes can violate spacing rules to wires the script didn’t anticipate, and they change the capacitance the timing check saw. Placing fill as many copies of one small cell, rather than millions of separate squares, keeps the file small, because the output becomes cell references instead of flat polygons.
Each shape’s layer/datatype pair, from the illustrative map. Tap a row to highlight its shapes.
- A missing piece. A part still drawn as an empty box becomes an empty hole in the chip. Engineers look inside every part before sending.
- The wrong version. Someone sends an older file than the one that passed the checks. The file’s fingerprint on the checklist catches it.
- Fill in the wrong place. Metal fill too close to an important wire can slow down its signals. Teams mark spots where fill is not allowed.
- A real bug. The checks prove the drawing matches the plan. They can’t prove the plan itself is right. So some chips need a second try.
- Outlines left in the merge. A cell or block that is still an empty abstract becomes a hole in the chip. An empty-cell check catches it, and so does LVS, which finds no transistors there.
- Numbering mismatches. A block delivered with its own datatype conventions puts fill or pin labels on the wrong datatype. Listing every layer/datatype pair in the merged file and comparing it with the layer map catches it.7
- Two cells with one name. Two suppliers each ship a cell with the same name, and one silently replaces the other. Checking names before the merge catches it.
- Fill after timing. Fill added after the final timing check adds capacitance nobody analyzed.14 Recompute the wires’ capacitance with fill in place before signing off timing.
- Trying to fix high density with fill. A window that is too dense needs slotting or rerouting; more fill only makes it worse.12
- Changes outside your area. On a shared run, an XOR against the provider’s frame catches any accidental edit outside your block.22
- Rule-file drift. Signing off with one version of the foundry’s rule file while the foundry checks with a newer one. Record the version in the checklist and rerun on every update.
- Off-grid and angled shapes. Rotated blocks, scaled imports and 45° geometry produce corners off the manufacturing grid, which the foundry rejects.22
- Late fixes without refill. A last-minute wiring change removes metal and leaves a window under minimum density, or existing fill now overlaps a new wire. Refill locally, then rerun DRC and extraction.
- Slotting a power wire without re-analysis. Slots shrink the wire’s cross-section, which raises its voltage drop and its current density.12
- Data volume surprises. Flat fill on every layer turns a manageable file into one that takes hours to transfer and check. Place fill hierarchically and ship OASIS.5
- No respin insurance. Too few spare cells, or none near the logic that needs to change, can make a metal-only fix impractical, and a full rebuild needs new masks for every layer.2625
The SRAM is still an abstract after the merge: its outline and pins, but no transistors. On the masks it would be a hole.
This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.
GDSII stream records
A GDSII file is one long sequence of records. Each record starts with a 4-byte header: a 2-byte length (counting the header itself), a 1-byte record type (what the record is, such as LAYER or XY) and a 1-byte data type (how to read what follows: nothing, 2-byte integers, 4-byte integers, 8-byte reals or text). The records must come in a fixed order: HEADER, BGNLIB, LIBNAME, UNITS, any number of structures (cells), then ENDLIB. A structure is BGNSTR, STRNAME, its elements, and ENDSTR. Each element (BOUNDARY for a polygon, PATH, SREF, AREF, TEXT…) is a run of records closed by ENDEL. A boundary has at most 200 points, with the first repeated as the last; layers and datatypes run from 0 to 255; and an AREF stores a column and row count plus three points: the array’s origin, the origin displaced by the column count times the column spacing, and the origin displaced by the row count times the row spacing.3 Here is a complete file containing one 1 × 0.5 µm rectangle on layer 15, as written by gdstk:
0006 0002 0258 HEADER version 600
001C 0102 007E 000A 0002 0000 ... BGNLIB modified/accessed dates
0008 0206 4C49 4200 LIBNAME "LIB"
0014 0305 3E41 8937 4BC6 A7F0 UNITS 0.001 user units per db unit
3944 B82F A09B 5A54 1e-9 m per db unit
001C 0502 007E 000A 0002 0000 ... BGNSTR dates
0008 0606 544F 5000 STRNAME "TOP"
0004 0800 BOUNDARY
0006 0D02 000F LAYER 15
0006 0E02 0000 DATATYPE 0
002C 1003 0000 0000 0000 0000 XY (0,0)
0000 03E8 0000 0000 (1000,0)
0000 03E8 0000 01F4 (1000,500)
0000 0000 0000 01F4 (0,500)
0000 0000 0000 0000 (0,0) closes the polygon
0004 1100 ENDEL
0004 0700 ENDSTR
0004 0400 ENDLIB- 1L1Length 6 bytes, record type 0x00 (HEADER), data type 0x02 (2-byte integer), value 0x0258 = 600.
- 2L4Two 8-byte reals. First byte 0x3E: exponent 62 − 64 = −2. Mantissa 0x4189374BC6A7F0 / 2⁵⁶ = 0.256. Value 0.256 × 16⁻² = 0.001: a database unit is a thousandth of a user unit (µm).
- 3L5Same format: exponent 0x39 − 64 = −7, so 0.268… × 16⁻⁷ = 1e-9. The database unit is 1 nm.
- 4L8Element records carry no data, only the 4-byte header.
- 5L110x2C = 44 bytes: 4 of header plus five points of two 4-byte integers each. Five points for a rectangle is exactly the redundancy OASIS removes.
The 8-byte real is not the IEEE 754 format modern computers use. It has a sign bit, a 7-bit exponent stored with 64 added to it (“excess-64”) that raises 16, not 2, to its power, and a 56-bit mantissa normalized to at least and below 1.3 A writer that converts from IEEE doubles has to renormalize in hex digits, and a careless converter produces a UNITS value that is off in its last digits, which some tools then round differently.
OASIS compaction
OASIS attacks each source of GDSII redundancy in turn:
- Variable-length integers. An OASIS integer takes as many bytes as its value needs, so a small offset doesn’t pay the fixed four bytes GDSII spends on every coordinate.54
- Repetitions. Any shape or cell placement can carry a repetition: a regular grid, a row with uneven spacing, or an arbitrary list of displacements. Chen et al. describe eight repetition types and note that GDSII’s
AREFcan express only the first three, the regular grids. Writing fill with the full set gave files on average about half the size of those limited to GDSII-style arrays.5 - Compact shapes and compressed cells. Rectangles and trapezoids have short dedicated records instead of a full corner list, and each cell can be stored compressed; gdstk, for example, exposes both as options when writing OASIS.74
Fill synthesis: windows, tiles and linear programs
A density rule applies to every possible window position, but no checker can test infinitely many. Industry practice uses a fixed -dissection instead.13 Step by step:
- Take the rule’s window size (say 50 µm) and a resolution (say 2).
- Cut the layout into square tiles of size (25 µm).
- Check every window made of neighboring tiles. Windows now start every , so they overlap; a larger comes closer to checking every position.
r = 1: tiles of 50 µm, 8 windows checked. All pass; the sparsest is 32%, yet a sparse patch sits on the grid corner.
Fill synthesis then has two steps: decide how much fill goes in each tile, then place real fill shapes there.13 The classic way to do the first step is a linear program, a standard optimization in which every quantity is a variable and every rule is a linear inequality:
- Each tile gets a variable , the fill area to add, between zero and the tile’s slack (the area where fill can legally go).
- No window may end up denser than the upper bound .
- An extra variable must be no larger than any window’s density, so is a lower bound on all of them.
- Maximize . Raising the floor as far as the ceiling allows minimizes the spread between the densest and sparsest windows.
This is the min-variation formulation.13 A min-fill variant instead adds as little fill as possible while keeping densities within bounds, which matters because fill adds coupling capacitance.14
The linear program has a variable for every tile, of them on a die units across, and solving it takes time roughly cubic in the number of variables. Its answer is also only optimal for the chosen , not for a finer dissection.13 The Monte-Carlo alternative trades exactness for speed:
- Give every tile a priority.
- Pick a tile at random, with probability proportional to its priority.
- Add one unit of fill there, and update the priorities it affects.
- When a window reaches , lock it: its tiles drop out of the draw. Repeat until no tile can take more.
The priority decides the behavior. Slack priority favors tiles with the most empty area. Minimal priority, proportional to minus the density of the sparsest window containing the tile, favors tiles in the emptiest windows and behaves like a randomized version of a greedy solver for the linear program. A quadrisection tree, which splits the layout into quadrants recursively and stores the sum of priorities in each, makes each random pick and update take time for tiles. The result came close to the linear program’s optimum while running much faster.13 Later work iterated Monte-Carlo and greedy passes and added hierarchical fill, trading solution quality against runtime and output size.14
Hierarchy is the unresolved tension. Wiring layers have little natural hierarchy, and the same IP block sits in different surroundings at each place it is used, so fill has traditionally been computed flat and emitted as one flat layer.5 Hierarchical fill speeds up verification and shrinks the data, but forcing every copy of a cell to carry identical fill conflicts with hitting the density target in each copy’s neighborhood.14 A related problem spans layers: polishing an upper layer sees the bumps left by the layers below, so multiple-layer formulations track a cumulative density through the stack rather than treating each layer alone.14 Compression-aware fill closes the loop with the file format: given the fill amount per tile, choose positions that need the fewest OASIS repetition records to describe.5
CMP models
Fill exists because polishing results follow density. In the oxide polishing model Chen et al. adopt (from Stine), raised areas are worn down at the blanket polish rate (the rate on an unpatterned wafer) divided by the local pattern density, until the step between high and low areas is gone; after that the whole surface recedes evenly. The final insulator thickness at each point therefore depends on the starting pattern density there.14 That density is not simply the metal fraction in a window. The polishing pad bends over a characteristic distance, so the effective density is a weighted sum of the surrounding tile densities, with an elliptical weighting function of distance from the window’s center whose constants are fitted to measurements.14 Because both density definitions are weighted sums of tile densities, the same fill algorithms work for either. Density matters at the high end too: very wide wires such as power buses are slotted to keep them from lifting off during polishing, which is why maximum-density rules exist alongside minimum ones.12
Model-based OPC and inverse lithography
Edge-based OPC splits each polygon edge into segments and, iteration by iteration, moves each segment along its normal (perpendicular to the edge) to cancel imaging errors.19 Each iteration needs a forward model of the printing:
- Optics. The image formed by the lens is computed by convolving the mask with a set of optical kernels and adding up the results with weights: the sum-of-coherent-systems (SOCS) decomposition of the imaging system.
- Resist. A threshold model turns that light intensity into a printed shape: wherever the intensity exceeds a threshold, the resist prints.
- Error. The edge placement error, the distance between the predicted edge and the target at sampled evaluation points, is the quantity driven toward zero.19
Assist features, serifs and hammerheads come out of the same machinery or from rules, and the full-chip job runs on large compute farms.18
Inverse lithography turns the mask into a grid of pixels and adjusts every pixel by gradient descent through the lithography model. The freedom gives better precision, but the curved results have to be broken into rectangles a mask writer can draw, which raises cost and can bring back errors.19 Current research tries to keep inverse lithography’s gradient-driven quality while moving only edge segments, so masks stay manufacturable; DiffOPC reports lower edge placement error than prior methods at about half the manufacturing cost.19 For the design team, the practical point is that the shapes in chip_top.gds are a target. What goes on the mask is the output of an optimization the foundry runs on that target, and many rules in the design rule deck exist to keep that optimization solvable.
Q1In a GDSII file, what is the datatype of a shape used for?
Q2Why must the final rule check (DRC) and circuit comparison (LVS) run on the merged layout file, not on the place-and-route result?
Q3A density window on M1 measures 85% against a 20–80% rule. What fixes it?
Q4Which features make OASIS files smaller than GDSII for the same layout?
Q5What makes a metal-only respin cheaper and faster than a full respin?
Sources
Show Hide 26 sources
- Tape-outDefinition; origin in paper and magnetic tape reels carrying the files used to make photomasks; foundry checks, chip finishing and RET after submission; tapeout as a cause for celebration; respins (“spins”) and why they happen.
- GDSIIBinary hierarchical layout format from Calma (1978); shapes grouped by layer number and datatype; usually the final output handed to the foundry; originally written on magnetic tape.
- GDSII formatRecord layout (2-byte length, record type, data type), record codes (HEADER, BGNLIB, UNITS, BGNSTR, BOUNDARY, SREF, AREF, XY, ENDLIB…), excess-64 base-16 8-byte reals, layer and datatype 0–255, 200-point limit, stream BNF.
- Open Artwork System Interchange StandardSEMI P39, version 1.0 in March 2004; GDSII files of tens of gigabytes; variable-length integers, per-cell compression, rectangles stored as width and height, repetition records, values omitted when they repeat the previous record.
- Evaluation of the New OASIS Format for Layout Fill CompressionGDSII files of many tens of gigabytes; fill added by physical verification tools as a flat layer and merged at mask data prep; GDSII’s only compression operators are SREF and AREF; OASIS unsigned integers of variable length; AREF covers three of eight OASIS repetition types; full OASIS about 2× smaller than restricted; fractured data of hundreds of GB per layer.
- Getting StartedLayer and datatype have no predefined meaning; the system using the file decides; reading and writing GDSII and OASIS; references and arrays.
- gdstk.LibraryLibrary constructor (unit 1e-6, precision 1e-9 by default); add does not check for duplicate cell names; replace swaps cells by name and updates references; remap of (layer, type) pairs; layers_and_datatypes; top_level; write_gds; write_oas with per-cell compression and an optional CRC32 or checksum.
- LEFDEFReaderConfiguration class referencemacro_layout_files substitutes real layouts for LEF macros when reading DEF; macro_resolution_mode and FOREIGN; map_file maps LEF/DEF layers to GDS layer/datatype.
- DRC Reference: Layer Objectwith_density with tile_size and tile_step; fill with fill_pattern, hstep, vstep, origin, auto_origin, multi_origin, fill_exclude; fill output to a target layout.
- OpenROAD Flow Scripts User GuideYosys for synthesis, OpenROAD from floorplan to detailed routing, KLayout for GDS merge, DRC and LVS on public PDKs.
- Metal fill (density_fill) documentationInserts floating metal fill to meet density rules while obeying DRC; JSON rules per layer with GDS layer, datatype, fill sizes and spacing to fill and non-fill.
- Filling and Slotting: Analysis and AlgorithmsDensity rules for CMP and etch; example rule of 35–70 µm² of metal per 10 × 10 µm window; wide power buses slotted to avoid problems in CMP; fill and slotting change extraction results.
- Monte-Carlo Algorithms for Layout Density ControlCMP with a rotating pad and slurry; fixed r-dissection with tiles and overlapping windows; min-variation LP; Monte-Carlo tile selection by priority, quadrisection data structure, near-LP quality with better scaling.
- Area Fill Synthesis for Uniform Layout DensityILD thickness variation controlled through window density; Stine’s CMP model; effective density with an elliptical weighting for pad deformation; Min-Var and Min-Fill LPs; iterated Monte-Carlo and greedy; hierarchical fill; fill changes capacitance.
- Mask data preparationTranslating layout polygons into mask-writer instructions: chip finishing (seal rings, fill), reticle layout with test patterns and alignment marks, OPC/ILT, fracturing into rectangles and trapezoids.
- PhotomaskMask set (one mask per layer, used in turn); written by electron-beam or laser writers; CD-SEM measurement; pellicles against particles.
- EEC 116 lecture handout: PhotomasksMask set costs of about $2 million at 28 nm, $4 million at 14/16 nm and $8–10 million at 7 nm (Papermaster, 2016); multiple patterning adds masks; mask-making steps from write to pellicle.
- Optical proximity correctionRule-based vs. model-based OPC; serifs and hammerheads; sub-resolution assist features; large compute farms for full-chip OPC.
- Differentiable Edge-based OPCEdge-based OPC moves edge segments to reduce edge placement error; SOCS lithography model with a threshold resist; ILT optimizes mask pixels by gradient descent but yields masks that are hard and costly to write.
- Multi-project wafer serviceSharing mask and wafer cost among designs; MOSIS, set up by DARPA, from 1981; providers including CMC Microsystems, EUROPRACTICE, Muse Semiconductor and Tiny Tapeout (up to 512 designs behind a multiplexer); each provider sets its own design rules, die sizes, tools and timing; turnaround and cost depend on the technology.
- Tiny Tapeout FAQSKY130 PDK; tiles of about 160 × 100 µm (TT04–TT10); a GitHub Action runs the OpenLane flow that turns each design into GDS; OpenLane adds antenna diodes to protect transistor gates during manufacture; 6 to 9 months to manufacture.
- mpw_precheckOpen-source pre-tapeout checks for SKY130 shuttle projects built on the Caravel harness, run on the submission platform: XOR outside the user area, Magic and KLayout DRC, off-grid shapes, maximum metal density, LVS (local runs only), documentation and license checks.
- Semiconductor device fabricationFabrication at 14/10/7 nm takes up to 15 weeks, 11–13 weeks on average.
- Chipmakers Are Ramping Up Production to Address Semiconductor Shortage. Here’s Why that Takes TimeWafer cycle time of about 12 weeks on average and 14–20 weeks for advanced processes; up to 26 weeks from order to finished chip.
- Engineering change orderAfter masks are made, a change confined to a few (typically metal) layers costs much less than a rebuild, which needs new masks for all layers; designers sprinkle unused gates to make this possible.
- Resource-Aware Functional ECO Patch GenerationMetal-only ECO after placement is frozen; spare cells spread over the design during placement and rewired; too few or too distant spares cause timing violations and routing congestion.