Save this video — free

EfficientML.ai Lecture 4 - Pruning and Sparsity (Part II) (MIT 6.5940 Fall 2026)

MIT HAN Lab · 1:14:06 · Watch on YouTube

EfficientML.ai Lecture 4 - Pruning and Sparsity (Part II) (MIT 6.5940 Fall 2026) Watch on YouTube →

Overview

The lecture explains how to choose layer-wise pruning rates, recover accuracy, and turn theoretical sparsity into measured speedups. It compares AMC’s reinforcement-learning search and NetAdapt’s iterative latency-constrained pruning, then examines hardware and software approaches including EIE, NVIDIA’s 2:4 sparse Tensor Cores, and sparse point-cloud convolution.

Key takeaways

Chapters

0:00 Pruning Granularity: Accuracy Flexibility Versus Hardware Regularity
5:00 Why Uniform Layer Pruning Gives Poorer Accuracy–Latency Trade-offs
8:00 Measure Layer Sensitivity Before Assigning Pruning Rates
13:00 AMC Uses Reinforcement Learning to Search Per-Layer Sparsity
21:00 AMC Results: Faster ResNet-50 and MobileNet Pruning
30:00 NetAdapt Iteratively Meets Device Latency Constraints
36:00 Recover Pruned-Model Accuracy with Small Steps and Fine-Tuning
41:00 EIE Exploits Static Weight and Dynamic Activation Sparsity
46:00 EIE Processing Elements Skip Zeros but Add Indexing Overhead
53:00 Specialized Sparse Hardware Must Balance Efficiency and Generality
59:00 NVIDIA 2:4 Sparsity Compresses Weights for Sparse Tensor Cores
1:04:30 Sparse Point-Cloud Convolution Preserves Activation Patterns
1:09:00 Adaptive Grouping Balances Sparse-Workload Regularity
1:12:00 Point Accelerator Uses Merge Operations to Match Sparse Coordinates

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.