EfficientML.ai Lecture 6 - Quantization (Part II) (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
The lecture develops post-training quantization (PTQ) and quantization-aware training (QAT) as ways to reduce model precision while preserving accuracy, moving from per-tensor and per-channel scales to grouped formats such as NVIDIA Blackwell NVFP4 and microscaling (MX). It also explains calibration through clipping and rounding, the straight-through estimator for QAT, binary and ternary arithmetic using XNOR/popcount, and reinforcement-learning-based mixed-precision assignment for balancing accuracy, latency, size, and energy.
Key takeaways
- Per-channel or grouped scales reduce error when layer channels have very different ranges, while coarser groups remain easier for hardware to execute in parallel.
- NVFP4 and MX-style hierarchical scaling amortize scale metadata across groups, making low-bit representations practical without paying for an independent high-precision scale per value.
- PTQ calibration must choose a middle-ground clipping threshold: KL divergence or reconstruction error can guide the choice between wasting levels on rare tails and saturating too many values.
- QAT preserves a full-precision master weight and uses the straight-through estimator to pass gradients through otherwise nondifferentiable quantization steps.
- Binary dot products can be implemented with XNOR and popcount; for vector length n, the result follows from twice the match count minus n.
- Mixed-precision quantization should account for layer sensitivity and hardware costs; hardware-aware reinforcement learning can search assignments against accuracy, latency, model-size, or energy objectives.
Chapters
- K-means weight quantization stores a cluster index for each weight and a shared floating-point centroid codebook; a 2-bit index selects among four centroids.
- Combining pruning with quantization can push model size to roughly 3–4% before accuracy falls sharply, compared with about 8% for pruning alone.
- Linear quantization maps integers to real values with a scale and zero point; symmetric weight quantization commonly sets the weight zero point to zero.
- PTQ selects scales and zero points after training, without retraining the model.
- A per-tensor scale is easy to implement but can fail when channel ranges differ substantially—by as much as 100× in some large language models.
- Per-channel quantization assigns each output channel its own range and scale, reducing quantization error compared with one scale for the whole tensor.
- Group quantization trades accuracy for hardware efficiency: smaller groups offer finer scaling, while larger groups are easier to parallelize.
- Per-vector scaling combines a global tensor scale with a local vector scale; with 4-bit values and a 4-bit scale per 16 values, the effective storage is 4.25 bits per value, excluding the global scale.
- NVIDIA Blackwell’s FP4 Tensor Cores are cited at 18 PFLOPS, versus 4.5 PFLOPS for FP16 and 9 PFLOPS for FP8 or INT8.
- NVFP4-style formats use hierarchical scales, including a low-cost local scale and a broader FP16 scale, to improve accuracy without storing a full-precision scale for every value.
- MX formats share exponent information across small groups; the lecture describes an exponent-only scale shared by every two values and a wider scale shared across 16.
- Effective bits per value include both the value bits and each shared scale amortized over its group—for example, the MX6 accounting totals about six bits per value.
- Weights can be calibrated offline, but activation ranges vary with images, tokens, and other inputs, so PTQ needs representative calibration data.
- Activation statistics can be tracked with an exponential moving average or estimated from a calibration batch.
- For assumed Laplacian activation distributions, the lecture gives optimal clipping ranges proportional to the distribution scale: 2.83× for 2-bit, 3.89× for 3-bit, and 5.03× for 4-bit quantization.
- Clipping avoids wasting quantization levels on sparse distribution tails, but clipping too aggressively saturates too many values; the best threshold balances both errors.
- KL-divergence calibration compares the original and quantized distributions to select a clipping threshold; the lecture also frames range selection as minimizing quantization error.
- Instead of always rounding to the nearest level, learned or stochastic rounding can choose between floor and ceiling values to reduce reconstruction error.
- Quantization-aware training inserts weight and activation quantization operators into the forward pass while retaining a full-precision master copy of each weight for small gradient updates.
- The quantized values are simulated in full-precision computation, so the approach is called fake or simulated quantization.
- A hard quantization step has zero derivative almost everywhere, preventing ordinary backpropagation from learning through the quantizer.
- The straight-through estimator (STE) treats the quantization operator as an identity function during backpropagation, passing its incoming gradient through unchanged.
- The lecture reports QAT recovering accuracy lost by PTQ on MobileNet-family examples, including MobileNetV1, MobileNetV2, and NASNet Mobile.
- Fake quantization preserves full-precision storage and gradients during training while constraining forward-pass values to the chosen quantized levels.
- Binary quantization maps each weight to −1 or +1, reducing weight storage from floating-point values to one bit per weight.
- For weights alone, binarization removes the multiplications but still requires accumulation; the lecture estimates about 2× less computation and 32× less weight memory.
- A learned or calculated scale, such as the mean absolute weight magnitude, improves reconstruction: the example error falls from 9.88 to 9.24 after scaling.
- When both weights and activations are binary, their dot product can be computed using XNOR followed by a population count.
- For a vector of length n, the binary dot product is recovered from the XNOR match count using a factor of two and an offset of −n.
- XNOR and popcount replace conventional multiply-and-accumulate operations with bitwise operations and counting, enabling parallel work across register bits.
- Ternary quantization adds a zero level: values above a positive threshold become +1, values below a negative threshold become −1, and values between become zero.
- The lecture gives a threshold near 0.7 times the expected weight magnitude as a useful rule of thumb and reports about 65% accuracy for a ternary example versus about 60% for a binary network.
- Trained Ternary Quantization (TTQ) learns separate scales for positive and negative weights; the cited result improves accuracy from 65.3% to 66.6%.
- Layers have different quantization sensitivities, so assigning one precision uniformly can waste capacity or harm accuracy; the lecture estimates 64 weight/activation choices per layer.
- Hardware-aware quantization uses an actor–critic reinforcement-learning setup with hardware simulation to evaluate accuracy and latency, and can optimize model-size, latency, or energy constraints.
- The lecture concludes with PTQ, QAT, NVFP4 and MX formats, binary/ternary methods, and mixed precision before the course moves on to neural architecture search.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.