Save this video — free

EfficientML.ai Lecture 3 - Pruning and Sparsity (Part I) (MIT 6.5940 Fall 2026)

MIT HAN Lab · 41:58 · Watch on YouTube

EfficientML.ai Lecture 3 - Pruning and Sparsity (Part I) (MIT 6.5940 Fall 2026) Watch on YouTube →

Overview

Pruning reduces neural-network memory and computation by removing weights, neurons, or channels while controlling the accuracy loss; fine-tuning and iterative pruning can recover accuracy even after removing up to 90% of weights in older models such as AlexNet. The lecture compares flexible but hardware-challenging unstructured sparsity with structured and 2:4 patterns that are easier to accelerate, then surveys pruning criteria including weight norms, channel scaling, second-order loss estimates, activation sparsity, and regression-based reconstruction.

Key takeaways

Chapters

0:00 Efficient Inference Begins with Memory-Efficient Pruning
4:26 Fine-Tuning and Iterative Pruning Preserve Accuracy
11:35 Pruning Results, Caption Quality, and Hardware Adoption
18:10 Unstructured, Pattern-Based, and Channel-Level Sparsity
23:13 2:4 Sparsity and Layer-Wise Compression Budgets
27:02 Magnitude, Norm, and Scaling-Factor Pruning Criteria
32:20 Second-Order Scores and Activation-Based Channel Pruning
37:55 Regression-Based Pruning Minimizes Reconstruction Error

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.