Pruning
Beginner
Setting the least useful numbers in a trained AI model to zero. Then the model practices a bit to make up for them.
Novice
Setting selected weights to zero, usually the smallest in magnitude, then fine-tuning the network so the remaining weights compensate. Done after or during training.
Expert
Magnitude pruning is the baseline; criteria using second-order information (e.g. SparseGPT) allow one-shot pruning of very large models. One-shot or iterative, global or per-layer, with or without retraining; the pattern (unstructured, N:M, block, channel) is a separate choice.
Explained in Sparsity, in-memory and analog compute (Architectures).
See also: Sparsity.