EfficientML.ai Lecture 3 - Pruning and Sparsity (Part I) (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
Pruning reduces neural-network memory and computation by removing weights, neurons, or channels while controlling the accuracy loss; fine-tuning and iterative pruning can recover accuracy even after removing up to 90% of weights in older models such as AlexNet. The lecture compares flexible but hardware-challenging unstructured sparsity with structured and 2:4 patterns that are easier to accelerate, then surveys pruning criteria including weight norms, channel scaling, second-order loss estimates, activation sparsity, and regression-based reconstruction.
Key takeaways
- Memory movement is the central efficiency motivation for pruning: the lecture's energy comparison puts memory access at more than two orders of magnitude above arithmetic operations.
- Pruning can remove about 80–90% of parameters in older networks when surviving weights are fine-tuned or pruned iteratively, but removing 95% can noticeably damage output quality.
- Unstructured sparsity offers more freedom to preserve accuracy, while structured pruning—especially channel or row removal—maps more readily to existing GPU kernels.
- NVIDIA 2:4 sparsity constrains every four weights to contain two zeros, trading arbitrary sparsity for hardware-friendly execution and reported gains of up to 2× peak performance.
- Pruning decisions can use weight magnitude, L1/L2 group norms, learned channel scales, Hessian-based loss estimates, or the fraction of zero ReLU activations; each criterion captures a different notion of importance.
- Regression-based pruning treats channel selection and weight adjustment as a reconstruction problem, alternating between optimizing the channel mask and the retained weights.
Chapters
- The lecture introduces pruning and sparsity as ways to reduce parameter counts, alongside quantization, neural architecture search, and knowledge distillation in the EfficientML.ai course.
- Memory movement consumes more than 100 times the energy of arithmetic operations in the comparison presented, motivating removal of redundant weights.
- Pruning is formulated as minimizing model loss while constraining the number of nonzero weights in the pruned model to a target budget.
- Biological synapses rise from about 2,500 per neuron at birth to roughly 15,000 in early childhood, then decline to around 7,000 in adulthood—a natural example of pruning.
- Neural-network pruning can remove individual connections or whole neurons; removing a neuron also removes its incoming and outgoing connections.
- For older models such as AlexNet, fine-tuning after pruning can recover accuracy after removing about 80% of parameters; iterative prune-and-retrain cycles can reach roughly 90%.
- Pruning limits vary by structure: older fully connected layers may tolerate about 90% removal, while convolutional layers and modern large language models often tolerate less.
- Iterative pruning reduced parameters in historical image models including AlexNet, VGG, GoogLeNet, ResNet-50, and SqueezeNet, with reported compression ratios varying by architecture.
- Image-captioning examples remained coherent after 90% pruning, but at 95% pruning a soccer caption became confused, illustrating that compression has an accuracy limit.
- Sparse attention can skip irrelevant attention positions even though the attention operation itself has no weights; transformer weights in Q, K, V, O, and feed-forward layers can also be pruned.
- NVIDIA hardware supports 2:4 sparsity—two zeros in each group of four weights—with up to 2× peak performance and about 1.5× measured speedup reported.
- Fine-grained unstructured pruning can remove any individual matrix entry, offering flexibility and high compression but producing irregular computations that are difficult to parallelize on GPUs.
- Structured pruning removes larger units such as rows, vectors, convolution kernels, or channels; row pruning turns a weight matrix into a smaller dense matrix that can use existing GPU kernels.
- Convolutional pruning choices span input channels, output channels, and kernel dimensions, with granularities ranging from single weights to complete channels.
- Coarser pruning is easier to accelerate but gives the pruning algorithm fewer choices, making accuracy preservation and high compression harder.
- In 2:4 sparsity, each four-weight group retains two nonzeros; the nonzero values and their indices can be stored separately to reduce representation size.
- Reported 2:4 results across models including ResNet, SSD, and BERT show sparse accuracy close to dense accuracy.
- Channel pruning reduces model dimensions and can run with unchanged dense GPU kernels, but typically offers less compression than fine-grained pruning.
- Layer-wise sparsity need not be uniform: assigning different pruning ratios per layer can improve the accuracy-latency tradeoff, and automated compression searches for better allocations.
- Magnitude pruning removes the smallest absolute-valued weights; in the example y = 10x₀ − 8x₁ + 0.1x₂, the 0.1 weight is the intuitive candidate.
- L1 and L2 norms rank weights or groups; for row pruning, summing absolute values selects the row with the smaller total magnitude.
- Channel scaling factors provide a learned importance signal: a channel with a scale near 0.1 can be removed before one with a scale near 1.17.
- Batch normalization's learned scale parameter can be reused to rank channels for structured pruning.
- Second-order pruning estimates loss change with a Taylor expansion; after training, the near-zero gradient and ignored cross-terms leave a score proportional to the Hessian diagonal times the squared weight.
- Hessian-based scoring can better reflect loss impact than simple magnitude, but computing it is costly for large models.
- ReLU activations create zeros that can indicate low channel utility; the example compares zero counts across two batches of 4×4 feature maps and prunes the channel with the highest zero fraction.
- Removing a neuron or channel eliminates its associated connections and reduces the corresponding weight-matrix or convolution dimensions.
- Regression-based pruning chooses channels and adjusts weights to minimize the difference between the original output Z = XW and the pruned reconstruction.
- The demonstration uses four input channels and a constraint that only three channels remain; a selection vector β marks a pruned channel with zero.
- Optimization alternates between fixing weights and selecting channels, then fixing channel selection and solving for weights, repeating until convergence.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.