Weight streaming
Beginner
Keeping a model’s learned numbers in a separate memory box and sending them into the chip a piece at a time.
Novice
An execution style in which weights live in external memory and are streamed onto the chip one layer at a time, so the model can be larger than the chip’s own memory.
Expert
Decouples model size from on-chip capacity: on-chip SRAM holds activations and the layer in flight, while external memory and a reduction network handle weight storage, gradients and updates. It trades on-chip weight bandwidth for external bandwidth.
Explained in Wafer-scale and SRAM-heavy designs (Architectures).
See also: Parameter (weight).