Pipeline parallelism
Beginner
Splitting a model into stages, like an assembly line, with each chip (or part of a chip) doing a few steps and passing the work on.
Novice
Placing different layers of a model on different processors. Data flows from one stage to the next, and several inputs can be in flight at once, one per stage.
Expert
Each stage holds only its layers’ weights, so capacity scales with stage count; only activations cross stage boundaries. Costs: pipeline fill and drain, latency per hop, and load imbalance between stages.
Explained in Wafer-scale and SRAM-heavy designs (Architectures).
See also: Parameter (weight).