Systems · Chapter 5 of 7 · Between the boxes

Optics

Over long distances, copper wires lose too much signal. So networks send the data as flashes of light through thin glass threads.

Pluggable optical modules turn electrical signals into light at the edge of a switch or server. Linear-drive optics remove some electronics to save power, and co-packaged optics move the optical parts right next to the switch chip.

Reach and energy per bit for copper and optics, SerDes and DSP power, pluggable, linear-drive (LPO) and co-packaged optics (CPO), and the reliability and serviceability trade-offs among them.

An AI cluster is thousands of chips that talk to each other all the time. Each conversation travels over a cable. For a short hop, the cable can be plain copper wire. But today’s cables carry so much data, so fast, that copper only works for about a meter. Beyond that, the signal fades away.

For anything longer, the data travels as light. A small converter turns the electrical signal into flashes of laser light. A glass thread called a fiber, thinner than a hair, carries the light hundreds of meters. At the far end, another converter turns it back.

Those converters add up. A big AI datacenter has hundreds of thousands of them. Together they use about as much electricity as a small town. This chapter is about where copper stops working and how engineers make the light converters use less power.

Every link in a datacenter network starts and ends at a , the circuit on a switch or network chip that sends bits one after another down a pair of wires at tens of billions of bits per second. Each of those serial channels is a ; a switch port bundles several lanes, for example 8 lanes of 100 Gb/s for an 800 Gb/s port. A 51.2 Tb/s switch chip has 512 lanes of 100 Gb/s.

Over a meter or two, those lanes can run straight through a copper cable. Copper loses more signal the longer the cable and the higher the frequency, so each time lane speeds double, copper’s reach shrinks: about 5 m at 25 Gb/s per lane, 3 m at 50, 2 m at 100, and about 1 m at 200. Past that, the signal has to become light in an optical fiber, which loses very little over hundreds of meters or kilometers.

Converting to light costs power, and with hundreds of thousands of links in an AI cluster it adds up to a real share of the facility. One switch vendor estimated transceiver power at about 2.3 MW for 100,000 servers in a traditional cloud datacenter and about 40 MW in an AI datacenter of the same server count. The figure of merit is : watts divided by bits per second, usually given in picojoules per bit (pJ/bit).

The chapter covers four ways to build a link, from cheapest and shortest to the newest:

  • Copper cables, passive or with signal-restoring chips in the plugs.
  • : hot-swappable boxes at the front of a switch, each with its own signal-processing chip.
  • : the same modules with that chip removed.
  • : the optics moved into the switch chip’s own package.

Each step saves power and gives up something in return, usually how easy the link is to repair or how freely parts from different vendors can be mixed.

At the edge of every switch ASIC, NIC and accelerator sits a bank of lanes. A 51.2 Tb/s switch has 512 lanes at 100 Gb/s; its 102.4 Tb/s successor has 512 at 200 Gb/s. How those lanes reach the other end of a link, over passive copper, retimed copper, a DSP-based pluggable module, a linear-drive pluggable, or optics co-packaged with the ASIC, sets the link’s reach, its energy per bit, and what happens when a part fails.

Two pressures drive the design space. First, electrical reach shrinks with every lane-rate doubling: OIF’s long-reach electrical projects went from 5 m of copper cable at 25 Gb/s to 3 m at 56, 2 m at 112 and 1 m at 224 Gb/s. Second, I/O power grows faster than the rest of the system. OIF’s 224G framework notes that SerDes power, on the host and in pluggable modules, is rising faster than any other part of system power. In OIF’s 12.8 Tb/s example the pluggable modules already drew nearly as much as the switch chip, 384 W against 432 W. At 51.2 Tb/s, 64 DSP-based 800G modules at 15–16 W each add about 1,000 W of optics.

This chapter works through the options in order of integration. It covers copper loss and SerDes equalization, what the DSP in a pluggable does and costs, what linear drive removes and what it couples together, and how CPO’s optical engines and external lasers change both the power budget and the . Energy figures are quoted per port end, in pJ/bit, with the host SerDes shown separately wherever the source allows.

0.5 mchip Achip Bcopper cable (twinax)sentreceivedreadable ✓
Distance
Medium

0.5 m of copper at 200 Gb/s per lane: inside the ~1 m reach, so the switch chip’s own SerDes drives the cable with no extra parts.

One lane at 200 Gb/s between two chips, over copper or over fiber with an optical converter at each end. Copper loss uses the sim’s model; distances not to scale.Share freely with credit: ‘Figure from chipfieldguide.com’

Why a wire fades

Data travels down a copper cable as electricity switching on and off very fast. The cable soaks up a bit of the signal along every inch. So the longer the cable, the weaker the signal at the end. And the faster the signal switches, the more the cable soaks up.

The chip at the far end can clean up a weak signal, up to a point. Past that point, it can’t tell a 1 from a 0. So each time engineers make wires faster, copper cables have to get shorter. Today that limit is about one meter. That is enough to link machines in the same cabinet, and not much more.

Stretching copper a little

A plain copper cable has no chips inside. It is the cheapest link and uses no extra power. To reach a few meters more, some cables hide a small chip in each plug. The chip reads the faded signal and sends out a fresh copy. That is an . It reaches farther, but its chips need power.

How a SerDes sends bits

A SerDes lane sends a stream of symbols. Up to 25 Gb/s per lane, each symbol was one of two voltage levels (, one bit per symbol). From 50 Gb/s per lane upward, Ethernet uses : four levels, two bits per symbol. A 100 Gb/s lane sends about 53 billion symbols per second; a 200 Gb/s lane, 106 billion.

Loss grows with frequency and length

The highest frequency a lane really has to carry is about half its symbol rate, the : 26.6 GHz for a 100 Gb/s lane, 53.1 GHz for 200 Gb/s. How much of the signal a channel loses at that frequency is its , in decibels: 10 dB means a tenth of the power arrives, 20 dB a hundredth, 30 dB a thousandth.

Two effects make copper lose more at higher frequencies:

  • The . Fast-changing current crowds into the outer skin of the wire, so the wire behaves as if it were thinner. Its resistance and loss rise with the square root of frequency.
  • Dielectric loss. The insulation between the wires absorbs a little energy on every cycle, and that loss rises in direct proportion to frequency.

Both losses accumulate along the cable, so loss in dB grows in proportion to length. Double the lane rate and every meter loses more; keep the same loss budget and you get fewer meters.

What the standards allow

The receiving SerDes undoes much of the damage with : it boosts the frequencies the cable weakened and subtracts the smeared-out echoes of earlier symbols. Standards set how much loss a SerDes must handle. OIF’s long-reach electrical interfaces, the class used for backplanes and copper cables, have been specified at 25 dB at 12.5 GHz (25 Gb/s lanes), 30 dB at 14 GHz (56 Gb/s) and 28 dB at 28 GHz (112 Gb/s), with copper cable reach falling from 5 m to 3 m to 2 m, and a target of 1 m at 224 Gb/s.

Ethernet follows the same curve. The 100 Gb/s-per-lane standard, IEEE 802.3ck, set out to support copper cables of at least 2 m; the 200 Gb/s-per-lane project, IEEE P802.3dj, targets at least 1.0 m. Intel and TE demonstrated a 200 Gb/s-class lane over a 1 m passive cable with about 36 dB of end-to-end loss at 53 GHz.

Not all of that budget goes to the cable. The package, the circuit-board traces from the chip to the front panel, vias and connectors all take their share, and at 200 Gb/s per lane the host board and package loss budget becomes a challenge in its own right.

Passive and active copper

A is just twin-axial copper pairs with a plug at each end; the switch chip’s own SerDes drives it. An puts a in each plug: a chip that fully receives the signal, recovers the bits, and transmits them again. A variant without retiming, called an active copper cable, uses a linear redriver that only equalizes. Active copper uses less power than optics and reaches farther, in less bulky cables, than passive copper. At 200 Gb/s per lane, host board and package loss make passive cable hard enough that IEEE contributors proposed retimed cables as the easier path to the 1 m objective. In PCIe systems, for example, retimed active cables reach up to 7 m, enough to connect GPUs across racks.

The cost is power. At 800 Gb/s, one Cisco engineering comparison puts a passive cable at under 1 W, an active electrical cable at about 8 W and an active optical cable at about 16 W.

The channel

A copper link is SerDes → package → host PCB → connector → cable → connector → host PCB → package → SerDes. Each segment’s in dB is additive and, for transmission lines, linear in length, with a per-unit-length attenuation that has a conductor term growing as f\sqrt{f} once the is smaller than the conductor and a dielectric term growing as ftan⁡δf \tan\delta. Losses are specified at the fN=baud/2f_{\mathrm{N}} = \mathrm{baud}/2: 26.56 GHz for 100G lanes at 53.125 GBd, 53.125 GHz for 200G lanes at 106.25 GBd.

Budgets by generation

Lane rateSignalingLong-reach channel lossCopper cable reach
25 Gb/sNRZ25 dB at 12.5 GHz5 m
56 Gb/sPAM430 dB at 14 GHz3 m
112 Gb/sPAM428 dB at 28 GHz2 m
224 Gb/sPAM4 (expected; not final in the 2022 framework)Not yet specified1 m

OIF CEI long-reach projects, from the CEI-224G framework. The matching IEEE objectives are at least 2 m of twinax for 802.3ck’s 100 Gb/s lanes, alongside 28 dB at 26.56 GHz for backplanes, and at least 1.0 m of twinax for P802.3dj’s 200 Gb/s lanes, alongside 40 dB die-to-die at 53.125 GHz for backplanes.

The cable gets only part of the end-to-end budget. A baseline proposal for 802.3ck’s cable PHY divided a 29 dB channel at 26.56 GHz as 2 × (7 dB host + 1.6 dB connector) + 11.8 dB for 2 m of bulk cable and its wire terminations, against 12.62 dB for the 3 m bulk cable of the 50 Gb/s-lane generation, whose budget was quoted at 13.28 GHz. Roughly the same number of dB buys a cable two thirds as long when the Nyquist frequency doubles. At 200 Gb/s per lane the host side gets worse too: IEEE contributors flagged the host PCB and package loss budget as a challenge for 200 Gb/s-lane cables.

What the SerDes pays for reach

a 30–40 dB channel takes several stages: a transmit FFE (feed-forward equalizer, a short digital filter that pre-distorts each symbol), a receive CTLE (continuous-time linear equalizer, an analog boost of the high frequencies), and usually an ADC followed by a DSP running a longer FFE and a DFE (decision-feedback equalizer, which subtracts the echoes of symbols already decided), plus clock recovery. A survey of published 100 Gb/s-per-lane SerDes chips found 4.5–6.5 pJ/bit for long-reach designs handling 41.5–50 dB of loss and about 1.6–1.7 pJ/bit for extra-short-reach (XSR) designs handling 7–11.5 dB; the same survey projected 4–5 pJ/bit for 200 Gb/s-lane SerDes. One open analysis models a 224G long-reach SerDes at about 5 pJ/bit and a 112G short-reach (XSR) one at about 1 pJ/bit. That 3–5× gap between LR and XSR is the energy prize behind shortening the electrical path, which is what LPO partly and CPO fully do.

Two other costs come with the loss. PAM4’s four levels sit at a third of the NRZ eye height, so links run at raw bit-error ratios around 10−410^{-4} and rely on to clean up. The Intel and TE 1 m, 36 dB demonstration ran at a pre-FEC BER of 3–4×10−43\text{--}4 \times 10^{-4}. And retiming anywhere in the path, in an AEC or in an optical module, adds a full receive-and-transmit SerDes per lane at each point.

Active copper

IEEE’s terminology distinguishes a full , with a clock-and-data-recovery () device in each plug, from a non-retimed active copper cable whose plugs hold a linear redriver that equalizes the receive direction. Retiming splits the channel in three: host → plug, cable, plug → host, each with its own budget, so the cable segment can be much longer and thinner. Active copper cables at 100 Gb/s per lane were already on the market in 2023, at lower power than optical modules. At 200 Gb/s per lane, host PCB and package loss budgets squeeze passive cable, which is why fully retimed cables were proposed as the easier route. A 2024 Cisco comparison at 800 Gb/s put an AEC at about 8 W (about 10 pJ/bit) and 3 m or more, against under 1 W for a 1 m DAC and 16 W for a 20 m active optical cable.

sent · PAM4, 2 bits per symbolreceived, after equalizingTXRX0 m1 m2 m3 m4 m5 m6 mreach ≈ 2.0 mloss at Nyquist (26.6 GHz)9.0 dB ✓equalizer budget ≈ 12 dB50+ dB
Lane rate

100G PAM4, Nyquist 26.6 GHz: 9.0 dB of cable loss over 1.5 m against about 12 dB the receiver can equalize. The link works. Reach ≈ 2.0 m.

A lane through twinax copper, using the sim’s loss model: loss per meter grows with frequency, total loss with length. The equalizer undoes the loss but boosts noise with it (shaded band); past its budget the levels blur together. Waveforms are illustrative.Share freely with credit: ‘Figure from chipfieldguide.com’

Most light links today use a . It is a box about the size of a pack of gum that slides into the front of a switch. A fiber clips into its other end. Inside are a laser, a light sensor and a chip that cleans up the signal.

Because the module plugs in, a broken one can be swapped in a minute. The rest of the switch keeps running.

The catch is power. Older, slower modules use about 12 watts each, as much as a bright LED bulb. Today’s fastest use 15–16 watts. A big switch holds up to 128 of them. More than half of that power can go to the signal-cleaning chip.

Form factors

A pluggable module follows a multi-source agreement (MSA): a specification that several companies publish together so that modules and switches from different vendors fit and work together. The two common families today are QSFP-DD and OSFP; both carry 8 electrical lanes.

What is inside

  • The , a digital signal processor. On the side facing the switch, it receives the electrical lanes and fully recovers the bits, like a . On the side facing the fiber, it shapes the signal for the optical transmitter and equalizes what the optical receiver delivers.
  • Transmitter. A driver amplifier and either a laser modulated directly or a steady laser followed by a modulator, which imprints the data on the light.
  • Receiver. A photodiode turns light into a tiny current, and a turns that current into a voltage for the DSP.

Because the DSP re-creates the bits, the link is split into independent pieces: switch → module (electrical), module → module (optical), module → switch (electrical). Each piece can be specified and tested on its own, which is why modules from many vendors interoperate.

Reach classes

IEEE Ethernet names optical links by reach. For 200 Gb/s-per-lane Ethernet, P802.3dj defines single-mode fiber links of at least 500 m and at least 2 km, plus 10 km and longer for some rates. By convention those are the DR (500 m) and FR (2 km) classes; LR is 10 km. For 100 Gb/s-per-lane Ethernet, multimode fiber links (SR) are specified for at least 50 m and 100 m.

Power

In OIF’s example of a 12.8 Tb/s switch, the 32 × 400G modules drew 12 W each, 384 W in total, against 432 W for the switch chip. Per bit, 12 W ÷ 400 Gb/s is 30 pJ/bit, on top of the SerDes in the switch chip that drives the module. Current 800 Gb/s modules are rated at up to about 17 W, about 21 pJ/bit, and a 2024 Cisco comparison puts a typical 800G pluggable at 15–16 W, about 20 pJ/bit. Roughly half of a module’s power goes to the DSP.

Architecture

A DSP-based (fully retimed) module presents a standard chip-to-module electrical interface to the host and a standard optical PMD to the fiber. The host side is a VSR-class channel: OIF’s CEI-112G-VSR covers about 16 dB at 29 GHz, against 28–30 dB for LR. Inside, the retimes and equalizes the signal in both directions and converts between digital and analog; in the receive direction it retimes and cleans up the optical signal before sending it to the switch. The analog front end is a modulator or laser driver and a . Light comes from directly modulated lasers, electro-absorption modulated lasers or silicon-photonic modulators fed by continuous-wave lasers.

The key property is that the link is retimed twice. Electrical and optical budgets are closed independently against reference receivers, which is what makes multi-vendor interoperability tractable, and what LPO gives up.

Form factors and power classes

QSFP-DD and OSFP are both 8-lane form factors with defined power classes and thermal designs. Power is the binding constraint on the faceplate: 32 or 64 cages, each dissipating up to about 30 W, must be cooled by front-to-back airflow. OIF’s 12.8 Tb/s example lists a typical 12 W per 400G port, 384 W for 32 ports beside a 432 W switch ASIC.

Where the module’s power goes

In NVIDIA’s description, a 1.6 Tb/s transceiver draws around 30 W, with the DSP consuming more than half, and the electrical path from switch ASIC to module is 14–16 inches of board trace or copper cable. One open analysis, from a photonics vendor, uses 16 pJ/bit for a DSP-based 112G-lane module on top of 5 pJ/bit for the host LR SerDes, 21 pJ/bit in total. Cisco’s 15–16 W per 800G pluggable comes to about 19–20 pJ/bit for the module alone. Treat the split inside the module as approximate: roughly half DSP, and the rest drivers, TIAs, lasers, their temperature control and the module’s microcontroller.

Failures

OIF puts current pluggable module failure rates at about 300 , failures per billion device-hours. With hundreds of thousands of modules in a cluster, that is a few failures a day, each a hot swap at the faceplate. Optical transceivers are among the most frequently replaced components in a datacenter. A study of 350,000 optical links across 15 datacenters of a major cloud provider traced packet corruption to faulty and decaying transceivers, damaged fiber and dirty connectors, and found that the first repair attempt fixed a link only about half the time.

module power, ≈ 15–16 W per 800GDSP 56%driver + TIA 25%lasers 19%front panelswitchSerDestracenext port’s moduleDSPdriverlaserTIAPDfiber

A DSP-based pluggable module in its cage at the front panel. Tap a part, send data, or pull the module.

A DSP-based pluggable module. About 15–16 W per 800G module (Cisco); the split is approximate (roughly half in the DSP; “lasers” includes the rest).Share freely with credit: ‘Figure from chipfieldguide.com’

If the signal-cleaning chip uses about half the power, why not leave it out? That is the idea behind , or LPO. The module keeps its laser and light sensor, but drops the cleaning chip. The big switch chip already has a strong cleaner of its own, so it does the job instead.

In one test, a linear module used under 9 watts, against 15 or 16 for a regular one. That saves over 40 percent. The modules still plug in, so they are just as easy to replace.

The catch: a linear module only works with switch chips strong enough to drive it. So mixing parts from different companies takes more testing.

What linear means

In a linear module, the transmitter’s driver and the receiver’s are linear amplifiers: they make the signal bigger without deciding what the bits are. Nothing in the module re-creates the bits. The SerDes in the switch chip at one end and in the chip at the far end must equalize everything in between: their own circuit-board traces, both modules’ analog parts and the fiber. OIF describes the same idea for co-packaged engines as a “linear amplified” interface, in which the host ASIC’s transmitter and receiver “must have greater capability” because they compensate for the entire link.

What it saves

  • Power. Cisco puts an 800G LPO module under 9 W against 15–16 W for a DSP module. That is under about 11 pJ/bit instead of about 20, for the modules alone.
  • Cost and latency. One fewer chip per module, and no extra retiming step on the way through.

What it costs

  • Interoperability. With the DSP gone, the electrical and optical parts of the link are no longer separate. The host, module and fiber now share one analog link budget, so LPO has needed its own specification of the host-to-module interface, for hosts with DSP-based SerDes.
  • Strong host SerDes. In LPO the host chip does the forward error correction, retiming and conversion between analog and digital for the whole link, so it needs long-reach-class SerDes with powerful equalization.
  • Reach. Without a DSP to clean up the optical signal, links run with more errors and shorter distances. The LPO MSA’s first specification covers 100 Gb/s lanes up to 500 m.

A middle path keeps a DSP in the transmit direction only and makes the receive path linear. It goes by several names: half-retimed, linear receive optics (LRO), or, in OIF’s 2025 agreement, retimed transmitter linear receiver (RTLR).

One budget instead of three

A DSP module closes three budgets: host TX → module RX (electrical, against a reference receiver), module TX → module RX (optical, against a reference receiver), module TX → host RX (electrical). In LPO the module’s driver and TIA are linear, so impairments accumulate from host TX through host channel, module driver, modulator, fiber, photodiode, TIA and far-end host channel to host RX, and only the far-end host receiver’s DSP removes them. OIF’s co-packaging framework states the consequence for its linear-amplified interface: the drive signal must be relatively linear, without amplitude compression, and the ASIC SerDes “must compensate for the entire link, from SerDes Tx to SerDes Rx.”

That has three engineering consequences. Host SerDes must be LR-class or better, so LPO saves module power but not host power. The module’s analog front end must be linear over the full PAM4 swing, with controlled bandwidth and noise, which raises its own power somewhat. And the link’s performance is a property of a specific host-module-host combination, which is why the LPO MSA defines test points spanning host, module and fiber, plus an optional host-to-host link test, and why link diagnostics move into the host ASIC. The MSA’s 100G-DR-LPO specification defines 53.125 GBd PAM4 optical interfaces up to 500 m and host-module electrical interfaces for hosts with DSP-based SerDes and RS(544,514) FEC, building on OIF’s CEI-112G-LINEAR-PAM4. Since an LPO module does not retime, even a basic loss-of-lock indication means nothing in the module, and receive-side signal-quality monitoring (SNR, BER, FEC statistics) moves to the host ASIC.

Numbers

For a 51.2 Tb/s switch with 64 ports of 800G, Cisco’s per-port figures (under 9 W for LPO, 15–16 W for DSP modules) mean roughly 400 W or more saved, all of it in the optics. An open modeling paper uses 8 pJ/bit for a DR8 LPO module plus 5 pJ/bit of host LR SerDes, 13 pJ/bit in total against 21 for DSP pluggables, and expects LPO hosts to need DSP-based SerDes of 4.5–6 pJ/bit and co-design with the host platform. The half-retimed RTLR variant keeps the module’s transmit DSP, so the optical transmitter can still be tested against IEEE specifications, and makes only the receive path linear, using the host’s DSP SerDes.

electricalopticalelectricalhosthostDSPlaserPDDSPtracemodule Afibermodule Btracesignal damage since the last retiming (illustrative)long-reach host can equalize✓module power: 15–16 W (800G)
Module type
Host SerDes

The DSPs retime twice, so the link is three short pieces, each tested on its own. The far host only equalizes its own trace.

One direction of an optical link. The line is the signal damage built up since the last retiming point, in illustrative units; a DSP resets it. Module power from Cisco’s 800G comparison.Share freely with credit: ‘Figure from chipfieldguide.com’

In a regular switch, the signal from the switch chip travels across the circuit board to a module at the front. That is a hand-width or more. The trip costs power, because the chip has to “shout” to be heard at the other end.

moves the light parts right next to the switch chip, a finger’s width away. Now the chip only has to “whisper.” Fibers then run from there to the front of the switch.

The lasers usually stay in their own plug-in boxes at the front. Lasers dislike heat, and they break more often than other parts. So it makes sense to keep them cool and easy to replace.

One chip maker says this cuts the power for the light parts by about 70 percent. The price is repair. The light parts can’t be unplugged. If one fails, it can take many connections down with it.

What gets co-packaged

OIF defines co-packaging as attaching optical or electrical communication devices to the same first-level substrate as the host ASIC. With the optical engine that close, the electrical channel is short enough for higher-speed, lower-power I/O drivers. Each pairs a photonic chip (modulators and photodetectors) with an electronic chip (drivers and TIAs). The photonic chips are usually built with , using chip-making tools to pattern optical components on silicon wafers.

External lasers

Silicon does not make good lasers, so the light comes from separate laser chips. Most CPO designs put them in an : a front-panel pluggable that sends steady, unmodulated light over fiber to the engines. OIF’s ELSFP agreement defines one such module, with a blind-mate optical connector so failed lasers can be replaced safely. Lasers tolerate much lower junction temperatures than silicon chips, so keeping them away from the hot switch package also lets them be cooled properly. The cost is extra optical loss through fibers and connectors, made up by running the lasers at higher power.

What it saves

Broadcom’s 51.2 Tb/s CPO switch, Bailly, integrates eight 6.4 Tb/s optical engines with the switch chip, serves 128 × 400G FR4 ports, supports field-replaceable remote laser modules, and is quoted at more than 70% lower optical interconnect power than standard pluggables. One system builder reports more than 30% lower power for the whole switch. Per 800G port, a Cisco comparison gives 15–16 W for DSP pluggables, under 9 W for LPO and under 6 W for CPO: about 20, under 11 and under 7.5 pJ/bit.

NVIDIA says its CPO switches cut the number of lasers, using 2 continuous-wave lasers per 1.6 Tb/s of bandwidth instead of 8 modulated lasers in pluggable modules, and claims about 3.5× better power efficiency.

What it costs

OIF’s framework is direct about it: co-packaged engines are “by nature less field serviceable than front panel pluggable designs.” A failed engine can’t be swapped at the faceplate, and one engine serves many ports, so CPO needs much more reliable parts. Lasers stay pluggable for exactly that reason.

Interfaces

OIF’s framework lists four host-to-engine interfaces: re-timed (the engine has a CDR/DSP, with the host using the CEI-112G-XSR extra-short-reach interface), linear amplified (no CDR/DSP in the engine; host SerDes compensates for the whole link), half-retimed, and direct drive, where the host drives the modulator directly and only the functions needed for a linear optical channel remain in the engine. OIF’s 3.2 Tb/s CPO module agreement takes the XSR route: 32 × 112G XSR lanes, optics as 8 × 400G-DR4 or FR4, with an internal laser or an external one fed over polarization-maintaining fiber. Energy-efficiency estimates in the framework, covering the engine-side electrical interface, CDR, photonics and laser but not the switch-side SerDes, are ≤ 15 pJ/bit for datacenter switching and HPC/AI, and 5–10 pJ/bit for AI training and disaggregation; a 3.2 Tb/s engine at 10 pJ/bit dissipates 32 W in roughly 20 × 20 mm (4 cm²), about 8 W/cm².

Engines and modulators

Engines differ mainly in modulator type and bandwidth density. Mach-Zehnder modulators are large, with optical bandwidth over 10 nm, and dominate the commercial market; micro-ring modulators are small and resonant, with optical bandwidth below 0.1 nm, which lets each ring add or drop its own wavelength for wavelength multiplexing. Rings drift with temperature and fabrication variation, so they need closed-loop thermal tuning. Ayar Labs reported a 1 nm resonance shift for a 16 °C change, and thermal transients of about 100 °C/s for an engine 3 mm from a 500 W ASIC, both of which its tuning loops must track.

Laser count and placement

Moving from externally modulated lasers in each pluggable to shared continuous-wave sources cuts laser count: NVIDIA cites 8 lasers per 1.6 Tb/s for a pluggable and 2 for its CPO engine, with micro-ring modulators and an external laser source module. Each of its external laser source modules contains eight lasers and powers 32 transmit lanes; Spectrum-X Ethernet engines carry 3.2 Tb/s over 16 transmit and 16 receive lanes, Quantum-X engines 1.6 Tb/s over eight each way. Broadcom’s 51.2 Tb/s design pairs its eight engines with field-replaceable remote laser modules.

Accounting for the host SerDes

CPO figures usually exclude the switch’s own SerDes, and that matters. An open analysis, by a photonics vendor, charged 5 pJ/bit of LR SerDes to every option and got 21 pJ/bit for DSP pluggables, 13 for LPO and 12 for 2.5D CPO with its laser, concluding that CPO with large beachfront expansion, and so long-reach host SerDes, “does not deliver substantially different energy efficiency than linear pluggable optics.” The larger win needs the short electrical path to be exploited: XSR SerDes around 1.6 pJ/bit instead of LR at 4.5–6.5, or direct drive. Intel’s Hot Chips 2024 slides make the same point, putting serial direct-drive CPO under 10 pJ/bit with 3–5 pJ/bit of that in the SerDes, and a die-to-die attached optical chiplet under 5 pJ/bit.

switch, top viewfront panelASICmodulemodulemodulemodulemodulemoduleoptics power per 800G port15–16 W
Optics placement

Front-panel pluggables: each lane crosses the board to its module, and each module has its own DSP. Tap a part.

A switch from above, schematic. Optics at the front panel or co-packaged with the switch chip. Power per 800G port from a Cisco comparison.Share freely with credit: ‘Figure from chipfieldguide.com’

Pick how far apart two machines are and how fast each wire runs. The chart shows how far each kind of link can reach. The dashed line is your cable. The cards below show which links work and how much power each uses. Make the wires faster and watch copper’s reach shrink.

Set the link length (log scale, 30 cm to 2 km) and the lane rate. Copper reach comes from a simple loss model: loss per meter rises with frequency, and a link works while total cable loss stays under what the SerDes can equalize. Each card shows energy per bit for one port end, including the switch chip’s SerDes, and the total for a 51.2 Tb/s switch (1 pJ/bit × 51.2 Tb/s = 51.2 W). The lowest-energy option that reaches is outlined. Numbers are illustrative, drawn from the published figures in this chapter.

Twinax loss is α(f)=1.1f+0.012f dB/m\alpha(f) = 1.1\sqrt{f} + 0.012 f\,\mathrm{dB/m} at fNf_{\mathrm{N}}; passive reach is a per-rate cable allowance divided by α\alpha, calibrated to the 5/3/2/1 m reach classes, and a retimer at each end adds 24 dB of allowance. Energy bars split each port into host SerDes, DSP or retimer, driver/TIA or optical engine, and laser, using rounded values from the measured and modeled figures above; each card also shows the typical range. Try 200G at 1 m and at 3 m, and compare LPO with CPO, whose host SerDes is set at 3 pJ/bit, between XSR and LR.

Loading simulation…
Copper reach at 200 Gb/s per lane
≈ 1 m
DSP pluggable optics, per 800G port
15–16 W
Co-packaged optics, per 800G port
< 6 W
Pluggable module failure rate
~300 FIT

What these numbers mean:

  • About 1 meter. At the newest speeds, a plain copper cable can only link machines sitting close together. Anything much farther needs light.
  • 15–16 watts vs under 6 watts. For each connection, a regular plug-in module uses as much power as a bright desk lamp. Moving the light parts next to the switch chip cuts that by about two thirds.
  • About 300 failures per billion hours. That sounds tiny. But a big AI cluster has hundreds of thousands of modules, so some break every day. That’s fine when they unplug. It’s a bigger problem when they’re built in.

Energy per bit, side by side

OptionOptics onlyWith host SerDesSource
Passive copper0~5 pJ/bit
DSP pluggable (800G, Cisco comparison)≈ 19–20 pJ/bit—
LPO (800G, Cisco comparison)< 11 pJ/bit—
CPO engines + lasers (800G, Cisco comparison)< 7.5 pJ/bit—
DSP pluggable / LPO / CPO (modeled)16 / 8 / 7 pJ/bit21 / 13 / 12 pJ/bit

The Cisco rows divide its per-port power at 800G by 0.8 Tb/s; they count the optics, not the switch ASIC’s SerDes, which is the same in all three. The modeled row charges 5 pJ/bit of long-reach host SerDes to each option. OIF’s 2026 energy-efficient interfaces framework sets targets of about 10 pJ/bit for front-end networks to 2 km, 5–10 pJ/bit for back-end networks to 300 m, and under 5 pJ/bit for compute interconnects to 20 m.

Every step closer to the switch chip saves power and gives up something:

  • Copper is cheapest and almost never breaks. But it is short, thick and heavy.
  • Plug-in modules reach far and swap out in a minute. But they use the most power.
  • Linear modules save a good chunk of power and still swap out easily. But they only work with strong switch chips.
  • Co-packaged optics uses the least power. But when a part fails, you may lose many connections at once.

Repair and failure domain

The biggest trade-off is what happens when something breaks. A pluggable module that fails takes down one link, and a technician swaps it at the front panel. A co-packaged engine that fails can disable many ports and a larger part of the network; Meta has asked whether a failed CPO port could cost a whole switch node. That makes the , how much breaks together, much larger.

Early field data is encouraging. Meta reported about 15 million CPO device-hours, with roughly 5× better mean time between failures (MTBF) than comparable pluggables and about 65% lower optics power.

Link failures matter more for AI training than for ordinary web traffic, because one job spans thousands of links. In Alibaba’s training network, 0.057% of the links between servers and their first switch failed each month and about 0.051% of those switches hit critical errors; together, that meant a single large training job crashed once or twice a month.

Mixing vendors

Pluggable modules with DSPs interoperate because each piece of the link is specified separately. LPO joins the pieces into one analog link, so the LPO MSA had to define its own host, module and fiber test points to keep parts from different vendors interoperable. CPO goes further: the optics are chosen when the switch is built, and a buyer can no longer pick modules from several suppliers for the same switch.

Heat

Moving optics into the package moves their heat there too. A 3.2 Tb/s engine at 10 pJ/bit dissipates 32 W in a few square centimeters, and lasers in particular cannot run as hot as silicon. That is one reason lasers stay outside, and why NVIDIA’s CPO switches are liquid cooled.

Copper is not dead

Where it reaches, copper still wins on cost, power and reliability, so designers try to keep the busiest links short enough for copper and use optics only where distance forces it. The scale-up chapter shows how that boundary shapes accelerator domains, and the scale-out chapter shows the networks the optics build.

Reliability arithmetic

Failure rates in FIT (failures per 10910^{9} device-hours) add across the parts a system needs. OIF’s reliability targets for a 51.2 Tb/s CPO switch with sixteen 3.2 Tb/s engines are 10 FIT per engine (lasers excluded), 50 FIT per 3.2 Tb/s laser source, 100 FIT for the rest of the system, and a total annualized failure rate under 1%, against roughly 300 FIT for today’s pluggable modules. Sixteen engines at a pluggable-like 300 FIT would be 4,800 FIT, an engine failure about every 24 years per switch; across 10,000 switches, more than one a day, each a board replacement. The framework adds that redundancy, and modes that run with fewer lanes, may be needed to reach system targets.

Failure domain and repair time

Availability is failure rate × impact × time to repair. A pluggable failure takes one port down for a minutes-scale swap. A CPO engine serves dozens of lanes: in Broadcom’s Bailly each of eight engines serves 16 of the 128 400G ports, so losing one takes out 16 ports at once. AI training is less tolerant of link loss than traditional datacenter networks, which rely on redundant links; OIF notes that AI applications are typically less tolerant of a link failure because of their higher aggregate bandwidth and connectivity requirements. The design answers are external, field-replaceable lasers, which the ELSFP agreement exists to provide; burn-in and known-good-engine screening; laser sparing and power boosting; and graceful lane degradation.

Field data

Meta’s reported results so far cover about 15 million CPO device-hours and two million pluggable-module hours, with no unserviceable CPO failures, about 5× better MTBF for CPO and about 65% lower optics power. In about one million 400G-equivalent port device-hours of Meta’s high-temperature lab testing, no link flaps were observed. This is early data, from one operator and one product generation.

Other costs

  • Supply chain and test. CPO turns optics from a separately purchased, separately qualified commodity into part of the switch build. A package with many engines multiplies integration yield risk, so engines must be screened for defects before attach.
  • Fiber management. Besides its data fibers to the faceplate, an engine with an external laser needs polarization-maintaining fibers from the laser module, all routed inside the box.
  • Power density. About 8 W/cm² for a 32 W, 20 × 20 mm engine, far above what optical components usually face.
  • Host SerDes. LPO requires LR-class host SerDes, and CPO’s energy win is largest only when the host interface shrinks to XSR or direct drive.
128 ports × 400G = 51.2 Tb/s0 / 128 down1 moduleper portgreen: port up
Optics

Front-panel pluggables: one module per port, each its own failure domain. Break one to see.

A 51.2 Tb/s switch with 128 ports of 400G. With co-packaged optics, one engine serves 16 ports, as in the chapter’s example. Squares are ports.Share freely with credit: ‘Figure from chipfieldguide.com’

This part goes deeper, into the math, models and algorithms behind the chapter. It’s written for the Expert level.

1. A copper reach model

Model twinax loss per meter at the Nyquist frequency as a skin-effect term plus a dielectric term:

α(f)=ksf+kd f    [dB/m]f=fN=baud/2\begin{gathered} \alpha(f) = k_{\mathrm{s}}\sqrt{f} + k_{\mathrm{d}}\, f \;\; [\mathrm{dB/m}] \\ f = f_{\mathrm{N}} = \mathrm{baud}/2 \end{gathered}

The conductor term follows from skin depth shrinking as 1/f1/\sqrt{f}; the dielectric term from loss tangent absorption, proportional to ff. Loss in dB is linear in length because attenuation is exponential in distance. If the SerDes can equalize a total budget BB and the host side (packages, traces, vias, connectors) takes H(f)H(f), the cable gets A=B−H(f)A = B - H(f) and

reach=Aα(fN)\mathrm{reach} = \frac{A}{\alpha(f_{\mathrm{N}})}

Worked example from the 802.3ck cable baseline proposal: B=29 dBB = 29\,\mathrm{dB} at 26.56 GHz, H=2×(7+1.6)=17.2 dBH = 2 \times (7 + 1.6) = 17.2\,\mathrm{dB}, so A=11.8 dBA = 11.8\,\mathrm{dB} for a 2 m cable, about 5.9 dB/m. If loss were pure skin effect, doubling to 53.1 GHz would raise α\alpha by 2\sqrt{2} to about 8.3 dB/m; with the dielectric term it is a little more. Holding AA constant would give about 1.4 m. In practice HH grows too at 200 Gb/s, which is why the P802.3dj objective is 1.0 m. The sim uses exactly this model with per-rate allowances fitted to the standards’ reach classes.

2. Why reach falls faster than the loss curve

PAM4 puts three eyes in the swing that NRZ uses for one, so each eye is a third the height: a 20log⁡103≈9.5 dB20 \log_{10} 3 \approx 9.5\,\mathrm{dB} SNR penalty at equal swing and noise. Equalization recovers loss by boosting high frequencies, but it boosts noise and crosstalk with them, so the usable budget BB is bounded by the receiver’s noise floor, not by how many equalizer taps can be built. That is why budgets have stayed in the 25–40 dB range across generations while the frequencies they apply at have quadrupled. Links therefore run at pre-FEC error ratios around 10−410^{-4} and lean on FEC.

3. Energy accounting

Energy per bit is power over throughput, and the unit conversion is convenient:

1 pJ/bit×1 Tb/s=10−12 J×1012 s−1=1 W\begin{aligned} 1\,\mathrm{pJ/bit} \times 1\,\mathrm{Tb/s} &= 10^{-12}\,\mathrm{J} \times 10^{12}\,\mathrm{s^{-1}} \\ &= 1\,\mathrm{W} \end{aligned}

So a 51.2 Tb/s switch’s ports cost 51.2 W per pJ/bit. Charging each option per port end (one side of the link, which is what a switch pays), with illustrative values from the sources above:

OptionHost SerDesDSP / retimerDriver, TIA, engineLaserTotal51.2 Tb/s
Passive copper5———5256 W
Active copper511——16819 W
DSP pluggable5943211,075 W
LPO5—5313666 W
CPO3—4.52.510512 W

Where the numbers come from: host LR SerDes 4.5–6.5 pJ/bit and XSR about 1.6; AEC about 8 W per 800G end, or 10 pJ/bit, which the table rounds up to 11; DSP module 16–20 pJ/bit with the DSP over half; LPO module 8 to under 11 pJ/bit; CPO engines and lasers about 7 pJ/bit together, under 7.5 in Cisco’s comparison, split here as 4.5 and 2.5. The CPO host SerDes value of 3 is a midpoint between XSR and LR; Intel quotes 3–5 pJ/bit for SerDes driving CPO directly. The splits within the optics (driver versus laser, DSP versus the rest) are approximate.

4. Laser power budget for an external source

In an external-laser design the laser’s light passes a connector at the laser module, fiber, a connector at the engine, fiber-to-chip coupling, splitters and the modulator before it carries any data, and each loss is paid in laser power. In dB terms:

Plaser [dBm]=Pneeded at modulator+Lconnectors+Lcoupling+10log⁡10Nsplit+margin\begin{aligned} P_{\mathrm{laser}}\,[\mathrm{dBm}] = {} & P_{\mathrm{needed\ at\ modulator}} \\ & + L_{\mathrm{connectors}} + L_{\mathrm{coupling}} \\ & + 10 \log_{10} N_{\mathrm{split}} + \mathrm{margin} \end{aligned}

Electrical power is then PlaserP_{\mathrm{laser}} divided by the laser’s wall-plug efficiency, plus its cooling. OIF notes that external sources improve reliability and thermal environment “at the cost of higher insertion loss,” compensated by increasing laser output power, and that integrated lasers have fewer losses to overcome. Sharing one high-power laser across several lanes (the 1:N splitting cases in OIF’s framework) cuts laser count, at the price of a larger loss per lane and a larger blast radius if that laser fails.

5. Reliability: from FIT to outages

For constant failure rates, λsystem=∑iNiλi\lambda_{\mathrm{system}} = \sum_i N_i \lambda_i and the annualized failure rate is AFR≈λ×8760 h\mathrm{AFR} \approx \lambda \times 8760\,\mathrm{h}. OIF’s targets for a 51.2 Tb/s CPO switch:

16×10⏟engines+16×50⏟lasers+100⏟rest=1060 FIT≈0.93 %/year\begin{aligned} & \underbrace{16 \times 10}_{\text{engines}} + \underbrace{16 \times 50}_{\text{lasers}} + \underbrace{100}_{\text{rest}} \\ & = 1060\,\mathrm{FIT} \approx 0.93\,\%/\mathrm{year} \end{aligned}

With lasers field-replaceable, only the engine and rest-of-system terms (260 FIT, about 0.23%/year) force a board-level repair. For a fleet of SS switches, expected board-level events per year ≈S×AFR\approx S \times \mathrm{AFR}; for 10,000 switches at 0.23%, about 23 a year, against several pluggable swaps a day at ~300 FIT per module across the same fleet’s millions of ports. Whether CPO wins on availability depends on whether its per-port failure rate drops by more than its failure domain grows, which is what the early field data is starting to measure.

FIT per switch (per 10⁹ h)1,060 FIT ≈ 0.93%/yr07501,500forces a board repair: 260 FIT ≈ 0.23%/yrengines 16×10 = 160rest 100lasers 16×50 = 800board repairs/yr23in 10,000 switcheslaser swaps/yr70front panel, minutespluggable swaps/yr1,68264 × 800G at 300 FIT
Fleet

λ = 16×10 + 16×50 + 100 = 1,060 FIT ≈ 0.93%/yr. Board-level: 260 FIT ≈ 0.23%/yr → about 23 board repairs a year across 10,000 switches, against 1,682 pluggable swaps.

Failure arithmetic for a 51.2 Tb/s CPO switch, starting from OIF’s targets (10 FIT per engine, 50 FIT per laser source, 100 FIT for the rest). Constant failure rates; AFR ≈ λ × 8760 h. Try a pluggable-like 300 FIT per engine, or lasers that can’t be swapped.Share freely with credit: ‘Figure from chipfieldguide.com’
Novice · 0 of 4 correct
  1. Q1Ethernet’s copper cable reach was at least 2 m for 100 Gb/s lanes and dropped to at least 1 m for 200 Gb/s lanes. Why?

  2. Q2An 800 Gb/s pluggable module draws 16 W. What is its energy per bit?

  3. Q3In a linear-drive (LPO) link, which part has to equalize the optical path that the module’s DSP used to handle?

  4. Q4Why do most co-packaged optics designs keep the lasers in separate front-panel modules?

Sources

Show Hide 50 sources
  1. Co-Packaged Silicon Photonics Switches for Gigawatt AI FactoriesGilad Shainer · NVIDIA, Hot Chips 2025 · 2025Transceiver power of 2.3 MW (cloud) vs 40 MW (AI) per 100K servers; 8 lasers per 1.6 Tb/s for pluggables vs 2 for CPO; micro-ring modulators; external laser source; 3.5× power-efficiency claim; Spectrum-X and Quantum-X Photonics switch configurations, liquid cooled.
  2. Broadcom Ships Tomahawk 5, Industry’s Highest Bandwidth Switch Chip to Accelerate AI/ML WorkloadsBroadcom Inc. · Broadcom (news release) · 2022Announced August 16, 2022: 51.2 Tb/s of Ethernet switching in one monolithic 5 nm die; 64 ports of 800GbE; 512 × 100G PAM4 SerDes.
  3. Next Generation CEI-224G Framework (OIF-FD-CEI-224G-01.0)OIF · 2022Table 1: CEI-LR loss budgets and copper cable reach of 5, 3, 2 and 1 m at 25, 56, 112 and 224 Gb/s; copper increasingly bandwidth limited; SerDes power rising faster than other system power.
  4. Broadcom Now Shipping World’s First 102.4 Tbps Switch in Production VolumeBroadcom Inc. · Broadcom (news release) · 2026Tomahawk 6 doubles Tomahawk 5’s throughput to 102.4 Tb/s, with an option for 512 × 200G or 1,024 × 100G SerDes on a single chip.
  5. Broadcom Delivers Industry’s First 51.2-Tbps Co-Packaged Optics Ethernet Switch Platform for Scalable AI SystemsBroadcom Inc. · Broadcom (news release) · 2024March 14, 2024: Bailly integrates eight 6.4 Tb/s silicon-photonics optical engines with Tomahawk 5; 128 × 400G FR4 ports in an air-cooled 4RU system; remote laser modules for field replaceability; more than 70% lower optical interconnect power than pluggables.
  6. Broadcom Announces Third-Generation Co-Packaged Optics (CPO) Technology with 200G/lane CapabilityBroadcom Inc. · Broadcom (news release) · 2025May 15, 2025: Micas Networks’ TH5-Bailly switch system delivers more than 30% system-level power savings against traditional pluggable modules; Delta’s TH5-Bailly comes air- or liquid-cooled.
  7. ECEN720 High-Speed Links, Lecture 2: Channel Components, Wires & Transmission LinesSam Palermo · Texas A&M University · 2025Resistive (skin-effect) loss grows as √f, dielectric loss in proportion to f; both attenuate the signal exponentially with distance.
  8. Food for thought on active copper cablesAdee Ran · IEEE P802.3dj electrical ad hoc · 2023Active copper cables at 100G/lane exist: lower power than optical modules, longer reach and less bulky than passive copper; 200G/lane host PCB and package budgets challenge passive cable, so retimed active cable is easier to specify.
  9. 112 Gbps Electrical Interfaces – An OIF Update on CEI-112G (OFC 2020)OIF · 2020CEI-112G reach classes XSR, VSR, MR, LR with their loss at 28 GHz; VSR at 16 dB; IEEE baud rate of 53.1 GBd.
  10. Baseline proposals for electrical interfaces at 200 Gb/s per laneAdee Ran et al. · IEEE P802.3dj Task Force · 2024Signaling rate of 106.25 GBd for the 200 Gb/s-per-lane CR and KR PHYs.
  11. IEEE P802.3ck ObjectivesIEEE 802.3 · 2018Single-lane 100 Gb/s PHYs over twinax copper cable of at least 2 m, and over backplanes with ≤ 28 dB loss at 26.56 GHz.
  12. Adopted IEEE P802.3dj ObjectivesIEEE P802.3dj Task Force · 2024200 Gb/s-per-lane PHYs: twinax of at least 1.0 m, backplanes ≤ 40 dB die-to-die at 53.125 GHz, SMF of at least 500 m and 2 km; 10, 20 and 40 km for 800G.
  13. Demonstration of a 224Gbps-PAM4-LR SERDES in Supporting a 1 Meter Passive DAC Long Reach ChannelMike Li and Nathan Tracy · IEEE P802.3dj Task Force · 2023End-to-end channel with a 1 m DAC at about 36 dB loss at 53 GHz; pre-FEC BER of 3–4 × 10⁻⁴.
  14. Copper Cable TerminologyKent Lusted and Adee Ran · IEEE P802.3df Task Force · 2022Defines the fully retimed active electrical cable (CDR in each plug) and the non-retimed active copper cable (linear redriver).
  15. Astera Labs Unveils PCI Express® Cable Solutions that Accelerate Cloud and AI Infrastructure Deployments at ScaleAstera Labs · Astera Labs (news release) · 2024Retimed PCIe/CXL active electrical cables with seven meters of reach over copper, to connect PCIe 5.0 GPUs across racks.
  16. 400G, 800G, and Terabit Pluggable Optics: What You Need to Know (BRKOPT-2699)Mark Nowell and Errol Roberts · Cisco Live 2024 · 2024Slide 23, power and reach at 800G: backplane and DAC under 1 W (0.5 m, 1 m), linear active copper 2.5 W (2–3 m), AEC 8 W (3 m+), AOC 16 W (20 m), pluggable optics 15–16 W, LPO under 9 W, CPO under 6 W.
  17. Baseline Proposal: Cable assembly, Host, MTF, and Channel Insertion LossChris DiMinico · IEEE P802.3ck Task Force · 2019Proposed 29 dB channel at 26.56 GHz = 2 × (7 + 1.6) + 11.8 dB for 2 m of bulk cable and wire attachment (20 dB for the whole cable assembly); 12.62 dB for the 3 m bulk cable of 802.3cd.
  18. Power Considerations for 200G/L AUITobey P.-R. Li and Mau-Lin Wu · IEEE P802.3df Task Force · 2022Survey of published 100G/lane SerDes: 4.5–6.5 pJ/bit for long reach (41.5–50 dB), about 1.6–1.7 pJ/bit for XSR; 4–5 pJ/bit forecast at 200G/lane.
  19. Accelerating Frontier MoE Training with 3D Integrated OpticsBernadskiy et al. (Lightmatter) · arXiv (IEEE HotI 2025) · 2025Energy model: 5 pJ/bit LR host SerDes plus 16 (DSP module), 8 (LPO) or 7 (2.5D CPO with laser) pJ/bit; XSR about 1 pJ/bit; LPO hosts need 4.5–6 pJ/bit DSP SerDes; CPO with LR host SerDes not much better than LPO.
  20. Co-Packaging Framework Document (OIF-Co-Packaging-FD-01.0)OIF · 2022Definition of co-packaging; re-timed, linear amplified, half-retimed and direct-drive interfaces; 5–15 pJ/bit engine targets; external lasers; 12 W per 400G port in a 12.8T switch; ~300 FIT pluggables vs 10/50 FIT CPO engine/laser targets; less field serviceable.
  21. A New Era in Data Center Networking with NVIDIA Silicon Photonics-based Network SwitchingBrad Smith · NVIDIA Technical Blog · 2025A 1.6 Tb/s transceiver uses around 30 W, more than half in the DSP; 14–16 inch electrical path to pluggables vs under half an inch with CPO.
  22. QSFP-DD/QSFP-DD800/QSFP-DD1600 Hardware Specification, Rev 7.1QSFP-DD MSA · 2024Eight electrical lanes (8 × 100G for QSFP-DD800, 8 × 200G for QSFP-DD1600); module power classes.
  23. OSFP Module Specification, Rev 5.2OSFP MSA · 2025Eight transmit and eight receive differential pairs; power classes; maximum module power above 30 W (40 W for OSFP1600); hot-plug sequencing.
  24. 802.3 SMF PMD Nomenclature Proposal with Application to 400GBASE-LR4Chris Cole · IEEE 802.3cu Task Force · 2019Default single-mode reach names: DR 500 m, FR 2 km, LR 10 km, ER 40 km, ZR 80 km.
  25. IEEE P802.3db Adopted ObjectivesIEEE 802.3 · 2020100/200/400 Gb/s over multimode fiber of at least 50 m and 100 m.
  26. Cisco OSFP 800G Transceiver Modules Data SheetCisco · 2026Maximum power of 800G OSFP modules: DR8 16–17 W, 2×FR4 14.5 W, VR8 16 W.
  27. Linear Pluggable Optics Save Energy In Data CentersBryon Moyer · Semiconductor Engineering · 2025The module DSP retimes, equalizes and handles FEC, and cleans up the received signal; its power is roughly 50% of the module’s.
  28. 3D optoelectronics and co-packaged optics: when solving the wrong problems stalls deploymentYasha Yi and Danny Wilkerson · arXiv · 2026Optical transceivers are among the most frequently replaced datacenter components; a failed optical interface in an integrated node can compromise a whole package.
  29. Understanding and Mitigating Packet Corruption in Data Center NetworksDanyang Zhuo, Monia Ghobadi, Ratul Mahajan et al. · ACM SIGCOMM 2017 (author copy) · 2017350K switch-to-switch optical links in 15 datacenters of a major cloud provider; corruption caused by faulty and decaying transceivers, damaged fiber, dirty connectors; first repair attempt succeeds about half the time.
  30. Link Diagnostics in LPO Applications (white paper)LPO MSA · 2024LPO removes the module DSP for lower power, cost and latency, keeping a linear interface; the host ASIC’s DSP is far more capable than standards’ reference receivers; LPO modules do not retime.
  31. 100G-DR-LPO Specification, Revision 1.0LPO MSA · 202553.125 GBd PAM4 optical interfaces up to 500 m; host-module interfaces for hosts with DSP SerDes and RS(544,514) FEC; host performs FEC, retiming and A/D and D/A conversion; builds on CEI-112G-LINEAR-PAM4.
  32. Linear Pluggable Optics – An OverviewSujit Ramachandra and Farnood Rezaie · IEEE Electronics Packaging Society · 2026LPO uses high-linearity TIAs and no DSP/CDR, keeps hot-swappable pluggables; drawbacks of shorter distances from higher BER and more complex host SerDes; LRO keeps the transmit DSP.
  33. Implementation Agreement for 112 Gb/s Retimed Transmitter Linear Receiver (OIF-EEI-112G-RTLR-01.0)OIF · 2025Retimed optical transmitter with a linear optical receiver; no DSP in the module’s receive path, using the host ASIC’s DSP SerDes.
  34. External Laser Small Form Factor Pluggable (ELSFP) Implementation Agreement (OIF-ELSFP-01.0)OIF · 2023Front-panel pluggable continuous-wave laser source with a blind-mate optical connector for field replacement; lasers tolerate lower junction temperatures than silicon.
  35. Roadmapping the Next Generation of Silicon PhotonicsSudip Shekhar, Wim Bogaerts, Lukas Chrostowski, John Bowers, Michael Hochberg, Richard Soref and Bhavin Shastri · arXiv (Nature Communications) · 2023Mach-Zehnder vs ring modulators: size, optical bandwidth (>10 nm vs <0.1 nm); market dominated by traveling-wave MZMs; rings multiplex by resonant add/drop.
  36. Implementation Agreement for a 3.2Tb/s Co-Packaged (CPO) Module (OIF-Co-Packaging-3.2T-Module-01.0)OIF · 202332 × 112G XSR host lanes; 8 × 400G DR4 or FR4 optics; internal laser or external laser over polarization-maintaining fibers.
  37. Co-packaged optics (CPO): status, challenges, and solutionsMin Tan et al. · Frontiers of Optoelectronics (open access, PMC) · 2023Micro-ring modulators are sensitive to fabrication, temperature and laser variation and need closed-loop thermal tuning.
  38. A UCIe Optical I/O Retimer Chiplet for AI Scale-upVladimir Stojanovic · Ayar Labs, Hot Chips 2025 · 2025Up to 8.192 Tb/s bidirectional; 16-wavelength micro-ring links; external multi-wavelength light source; 1 nm = 16 °C; ~100 °C/s transients 3 mm from a 500 W ASIC.
  39. How Industry Collaboration Fosters NVIDIA Co-Packaged OpticsAshkan Seyedi · NVIDIA Technical Blog · 2025Spectrum-X 3.2 Tb/s engines with 16 transmit and 16 receive lanes; Quantum-X 1.6 Tb/s engines with eight each way; eight lasers per external laser source, powering 32 transmit lanes; 115.2 Tb/s over 144 × 800G ports.
  40. 4 Tb/s Optical Compute Interconnect Chiplet for XPU-to-XPU ConnectivitySaeed Fathololoumi · Intel, Hot Chips 2024 · 2024Serial direct-drive CPO under 10 pJ/bit with SerDes at 3–5 pJ/bit; D2D-attached chiplet under 5 pJ/bit; 64 lanes at 32 Gb/s over 8 wavelengths per fiber = 4 Tb/s.
  41. VCSEL-based CPO for Scale-Up in A.I. Datacenter – Status and PerspectivesM. Kohli and J. Teissier · arXiv · 2026Copper reach about 1 m at 200 Gb/s per lane, at about 5 pJ/bit in total.
  42. Energy Efficient Interfaces Framework (OIF-EEI-FD-01.0)OIF · 2026Targets: ~10 pJ/bit front-end to 2 km, 5–10 pJ/bit back-end to 300 m, under 5 pJ/bit compute interconnect to 20 m.
  43. The Third Time Will Be The Charm For Broadcom Switch Co-Packaged OpticsTimothy Prickett Morgan · The Next Platform · 2025About 5.5 W per 800G port for Bailly and, as Broadcom says, about 3.5 W for Davisson.
  44. Broadcom Announces Tomahawk® 6 – Davisson, the Industry’s First 102.4-Tbps Ethernet Switch with Co-Packaged OpticsBroadcom Inc. · Broadcom (news release) · 2025October 8, 2025: TH6-Davisson, 102.4 Tb/s with 16 × 6.4 Tb/s optical engines at 200 Gb/s per channel and field-replaceable ELSFP laser modules; 70% lower optical interconnect power than pluggables; shipping, with the device sampling to early-access customers.
  45. Liaison letter: IEEE P802.3dj status (draft D3.0)IEEE 802.3 Working Group · IEEE 802.3 · 2026P802.3dj completed Working Group ballot and progressed to IEEE Standards Association ballot.
  46. IEEE P802.3dj Standards Association ballot (reflector notice)John D’Ambrosia, IEEE P802.3dj Task Force chair · IEEE 802.3 B400G reflector archive · 2026February 17, 2026: the initial IEEE SA ballot opened on draft D3.0, closing March 26, 2026.
  47. IEEE 802.3 ballot and Task Force review announcementsIEEE 802.3 Working Group · IEEE 802.3 · 2026As of late September 2026: IEEE P802.3dj draft D3.3 in its 3rd Standards Association recirculation ballot, September 25 to October 10, 2026.
  48. PECC Summit: Meta’s Drew Alduino on AI Networking Reliability WallsJim Carroll · Converge Digest · 2025About 15 million CPO device-hours and 2 million pluggable-module hours; about 5× better MTBF and 65% lower power for CPO; a CPO failure can disable multiple ports.
  49. Broadcom Showcases Industry-Leading Quality and Reliability of Co-Packaged OpticsBroadcom Inc. · Broadcom (news release) · 2025October 1, 2025: one million cumulative 400G-equivalent port device hours of link-flap-free CPO operation in Meta’s high-temperature lab testing; CPO cut optics power 65% against pluggable modules; cites Meta’s ECOC 2025 paper.
  50. Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian et al. · ACM SIGCOMM 2024 (author copy) · 20240.057% of NIC-ToR links fail each month, causing 1–2 crashes a month for a single LLM training job.