EfficientML.ai Lecture 4 - Pruning and Sparsity (Part II) (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
The lecture explains how to choose layer-wise pruning rates, recover accuracy, and turn theoretical sparsity into measured speedups. It compares AMC’s reinforcement-learning search and NetAdapt’s iterative latency-constrained pruning, then examines hardware and software approaches including EIE, NVIDIA’s 2:4 sparse Tensor Cores, and sparse point-cloud convolution.
Key takeaways
- Uniform pruning ignores layer sensitivity: on the lecture’s ImageNet comparison, non-uniform pruning reached about 70.5% accuracy at 70 ms, versus about 120 ms for uniform pruning.
- AMC searches continuous per-layer pruning ratios with DDPG, using layer dimensions and a reward tied to accuracy and compute or measured latency; it reduced MobileNet from 569 to about 285 million MACs and from 119 ms to 64 ms on a Snapdragon phone.
- NetAdapt meets device-level latency or energy targets incrementally: it uses lookup-table estimates, tests candidate layer reductions with about 10,000 fine-tuning iterations, and chooses the highest-accuracy candidate at each step.
- Pruning is easier to recover from when applied gradually with a reduced learning rate; L1 or L2 regularization can further encourage weights or channel scales that are suitable for later removal.
- Theoretical sparsity does not guarantee speed: EIE must handle dynamic activation patterns and index overhead, while NVIDIA’s structured 2:4 pattern makes sparse execution more regular and approaches a 2× speedup on sufficiently large workloads.
- Sparse point-cloud convolution benefits from preserving spatial activation sparsity, but efficient execution depends on adaptive grouping or coordinate matching to balance parallelism against padding and irregular-workload overhead.
Chapters
0:00
Pruning Granularity: Accuracy Flexibility Versus Hardware Regularity
- Pruning reduces model parameters to lower deployment costs on edge devices and in cloud inference while aiming to preserve accuracy.
- Fine-grained pruning can remove arbitrary weights, but irregular sparsity is difficult to accelerate; channel pruning is more regular and uses standard kernels but offers less flexibility.
- NVIDIA A100 Sparse Tensor Cores support 2:4 weight sparsity, where two of every four elements are zero.
5:00
Why Uniform Layer Pruning Gives Poorer Accuracy–Latency Trade-offs
- Uniform shrinking applies the same pruning percentage to every layer, even though layers differ in their sensitivity to parameter removal.
- On an ImageNet accuracy-versus-latency comparison, automated non-uniform pruning reached about 70.5% accuracy at 70 ms, compared with about 120 ms for uniform pruning.
- At roughly 50 ms, non-uniform pruning exceeded 69% accuracy while uniform pruning fell below 67%.
8:00
Measure Layer Sensitivity Before Assigning Pruning Rates
- To estimate sensitivity, prune one layer at several rates and measure the accuracy of the whole network after each change.
- A layer whose pruning curve stays nearly flat even at 70–80% pruning is more robust than one whose removal quickly degrades accuracy.
- A chosen accuracy-loss threshold can be intersected with each layer’s curve to set its pruning ratio, but this requires repeated evaluations and ignores interactions between layers.
13:00
AMC Uses Reinforcement Learning to Search Per-Layer Sparsity
- Automated Model Compression (AMC) formulates layer-wise pruning as a reinforcement-learning problem with a continuous pruning-ratio action for each layer.
- The agent receives layer characteristics such as channel count, kernel size, dimensions, and FLOPs, then chooses pruning ratios under a model-size or compute constraint.
- AMC uses a DDPG agent for continuous actions and can reward lower error alongside lower log-FLOPs; a latency lookup table can target measured device latency instead.
21:00
AMC Results: Faster ResNet-50 and MobileNet Pruning
- Manual ResNet-50 sensitivity tuning took about a week and reached roughly 29% of the original model size without hurting accuracy; AMC found a model around 20% size within a few hours.
- AMC’s ResNet-50 layer-density pattern pruned 3×3 convolutions more aggressively than 1×1 convolutions, consistent with greater redundancy in the larger kernels.
- On a Samsung phone with a Snapdragon chip, AMC reduced MobileNet’s compute from 569 to about 285 million MACs and latency from 119 ms to 64 ms, roughly a 1.8× speedup.
- A uniformly scaled 0.75× MobileNet had a worse accuracy–latency trade-off than AMC’s non-uniformly pruned model; automated results should still be interpreted and checked.
30:00
NetAdapt Iteratively Meets Device Latency Constraints
- NetAdapt targets a global resource constraint such as latency or energy by reducing latency in incremental steps, ΔR, across candidate layers.
- A prebuilt lookup table estimates the latency impact of changing each layer’s filter count, avoiding repeated measurements on a mobile device during the search.
- For each candidate layer, NetAdapt prunes to the current ΔR target, briefly fine-tunes for about 10,000 iterations, and selects the candidate with the best accuracy.
- The iterative process produces models at different cost points, enabling a device to trade accuracy for latency or energy as battery conditions change.
36:00
Recover Pruned-Model Accuracy with Small Steps and Fine-Tuning
- Fine-tuning starts from a converged model, so its learning rate is typically reduced to about one-tenth or one-hundredth of the original training rate.
- Iterative pruning—such as pruning to 30%, fine-tuning, then pruning further to 50% and 70%—is generally easier to recover from than one aggressive pruning step.
- L1 or L2 regularization during fine-tuning can encourage small weights or channel-scaling factors that are easier to prune in the next iteration.
- The goal is to recover accuracy after pruning and, with gradual iterations, improve the accuracy–sparsity trade-off.
41:00
EIE Exploits Static Weight and Dynamic Activation Sparsity
- The Efficient Inference Engine (EIE) is a specialized accelerator designed to skip zero-valued weights and activations rather than perform redundant multiply-accumulate operations.
- A fully connected layer with 90% weight sparsity needs only about 10% of the dense weight computations, while sparse storage also records indices for nonzero values.
- Activation sparsity is dynamic: different inputs produce different zero patterns, unlike static weight sparsity, whose pattern can be reused across inputs.
- The lecture estimates that removing about 70% of activations can add roughly a 3× compute reduction, which can multiply with savings from weight sparsity.
46:00
EIE Processing Elements Skip Zeros but Add Indexing Overhead
- EIE divides matrix-vector work among processing elements that store nonzero weights and their locations rather than full dense rows.
- At runtime, zero activations are skipped; nonzero activations are broadcast to processing elements holding matching nonzero weights, and the partial results are accumulated.
- Sparse execution can reduce memory traffic and computation, but pointers, index decoding, and control flow create overhead for each multiplication and addition.
- Index encoding can be compressed with shorter distance fields and padding, but that introduces design trade-offs when adjacent nonzero weights are far apart.
53:00
Specialized Sparse Hardware Must Balance Efficiency and Generality
- Combining roughly 10× weight-sparsity savings with roughly 3× activation-sparsity savings can yield much larger theoretical compute reductions, though layers without activation sparsity gain less.
- EIE demonstrated that specialized hardware can accelerate fine-grained sparse operations, but irregular execution creates control-flow, indexing, and parallelization costs.
- A retrospective on EIE highlights the trade-off between specialization and generality: hardware that accommodates new model architectures can remain useful as networks evolve.
- The lecture’s broad efficiency principle is to avoid work on zeros, extending to token sparsity in transformers, temporal sparsity in video, and spatial sparsity in point clouds.
59:00
NVIDIA 2:4 Sparsity Compresses Weights for Sparse Tensor Cores
- In 2:4 sparsity, each group of four weights retains two nonzeros; the compressed representation stores those values plus metadata indicating their original positions.
- During matrix multiplication, the metadata selects the corresponding two activation values so the hardware computes only the retained products.
- NVIDIA Sparse Tensor Core benchmarks showed speedups from about 1.2× for small dimensions to near 1.9× for larger ones, approaching the theoretical 2× as overhead becomes less significant.
- The reported FP16 sparse ResNet-50 result was 76.1 versus 76.2 for dense FP16, and the sparse format also combined with INT8 quantization.
1:04:30
Sparse Point-Cloud Convolution Preserves Activation Patterns
- Ordinary convolution tends to make activations denser in deeper layers, while sparse convolution preserves the input’s spatial sparsity pattern in its output.
- Sparse point-cloud convolution computes only at occupied output locations, skipping kernel positions that cannot produce nonzero outputs.
- GPU implementations group work with different numbers of valid point interactions; without grouping, uneven workloads can leave parallel compute resources idle.
- Grouping trades padding overhead against regular execution, since larger, more uniform groups can improve parallelism but also introduce wasted operations.
1:09:00
Adaptive Grouping Balances Sparse-Workload Regularity
- Adaptive grouping clusters sparse-convolution work items with similar workloads, reducing imbalance while avoiding the large padding cost of dense convolution.
- Reducing the number of groups can improve regularity and speed until groups become too large and padded, invalid computations erode efficiency.
- The best grouping strategy balances throughput and wasted work rather than maximizing either sparsity or regularity alone.
1:12:00
Point Accelerator Uses Merge Operations to Match Sparse Coordinates
- A point-cloud accelerator can convert shifted input and output coordinates into sorted indices and use merge-style matching to identify valid convolution pairs.
- Equal indices reveal which input and output points interact for a shifted kernel position, allowing hardware to avoid scanning dense spatial grids.
- The lecture concludes with AMC and NetAdapt for layer-wise pruning, hardware and software support for different sparsity patterns, and quantization as the next topic.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.