Glossary

Weight streaming

Beginner

Keeping a model’s learned numbers in a separate memory box and sending them into the chip a piece at a time.

Novice

An execution style in which weights live in external memory and are streamed onto the chip one layer at a time, so the model can be larger than the chip’s own memory.

Expert

Decouples model size from on-chip capacity: on-chip SRAM holds activations and the layer in flight, while external memory and a reduction network handle weight storage, gradients and updates. It trades on-chip weight bandwidth for external bandwidth.