Save this video — free

EfficientML.ai Lecture 6 - Quantization (Part II) (MIT 6.5940 Fall 2026)

MIT HAN Lab · 1:01:48 · Watch on YouTube

EfficientML.ai Lecture 6 - Quantization (Part II) (MIT 6.5940 Fall 2026) Watch on YouTube →

Overview

The lecture develops post-training quantization (PTQ) and quantization-aware training (QAT) as ways to reduce model precision while preserving accuracy, moving from per-tensor and per-channel scales to grouped formats such as NVIDIA Blackwell NVFP4 and microscaling (MX). It also explains calibration through clipping and rounding, the straight-through estimator for QAT, binary and ternary arithmetic using XNOR/popcount, and reinforcement-learning-based mixed-precision assignment for balancing accuracy, latency, size, and energy.

Key takeaways

Chapters

0:00 Review: K-Means, Pruning, and Integer Linear Quantization
5:25 PTQ Granularity: Per-Tensor Versus Per-Channel Scales
13:15 Group Quantization and Two-Level Scaling for Low-Bit Formats
17:10 NVFP4 and MX Formats: Counting Scale Bits
25:45 Activation Calibration and Selecting a Quantization Range
31:00 Clipping with KL Divergence and Rounding Beyond Nearest
37:45 QAT Setup: Fake Quantization and Full-Precision Weights
42:55 Straight-Through Estimator and QAT Accuracy Recovery
45:45 Binary Quantization: One-Bit Weights and Scaling
49:55 XNOR-Net: Replacing Binary Dot Products with XNOR and Popcount
54:45 Ternary Quantization and Trainable Ternary Weights
57:25 Hardware-Aware Mixed Precision and Lecture Takeaways

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.