EfficientML.ai Lecture 5 - Quantization (Part I) (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
Quantization reduces neural-network storage and data movement by mapping continuous weights and activations to lower-bit representations; moving from 32-bit to 4-bit values can cut memory traffic by 8×, important because data movement can consume up to 200× the energy of computation. The lecture explains integer and floating-point formats, K-means codebook quantization, and affine linear quantization, including how integer matrix multiplication and convolution can run with quantized weights and activations while largely preserving accuracy.
Key takeaways
- Replacing FP32 values with 4-bit values reduces storage and data movement by 8×; this matters because data movement can use up to 200× the energy of computation.
- BF16 preserves FP32's 8-bit exponent range while using only 7 fraction bits, making it a practical 16-bit format when range matters more than precision.
- K-means quantization represents each weight with a short centroid index, but its hardware must handle codebook lookups and still performs floating-point computation.
- Pruning before quantization preserves flexibility for fine-tuning; combining the methods reduced models to roughly 3–4% of their original size before accuracy loss in the cited experiments.
- Affine linear quantization uses a scale and integer zero point; with precomputed corrections and integer shifts, matrix multiplication and convolution can run in the integer domain.
- SqueezeNet's compact architecture combined with Deep Compression reached about 500× smaller size than AlexNet at the same reported accuracy, showing that architecture design and compression can compound.
Chapters
- Quantization reduces the bits used to represent each parameter or activation, complementing pruning's reduction in parameter count.
- Replacing a 32-bit float with a 4-bit integer uses 8× fewer bits and can reduce memory movement.
- Data movement can consume up to 200× more energy than computation, making lower-precision storage valuable.
- An n-bit unsigned integer represents values from 0 through 2ⁿ−1; an 8-bit example assigns bit weights from 2⁰ to 2⁷.
- Sign-magnitude representation uses a sign bit but has two encodings for zero, wasting one code.
- Two's complement avoids duplicate zero and represents the most negative n-bit value as −2ⁿ⁻¹.
- Fixed-point values encode fractional bits as negative powers of two and can be calculated by scaling an integer by a power of two.
- IEEE FP32 allocates 1 sign bit, 8 exponent bits, and 23 fraction bits; normal values use an implicit leading 1.
- The 8-bit exponent uses a bias of 127, so the represented exponent is the stored exponent minus 127.
- Subnormal values use an all-zero exponent and omit the implicit leading 1, allowing zero and values as small as 2⁻¹⁴⁹.
- An all-one exponent encodes infinity when the fraction is zero, and NaN when the fraction is nonzero.
- IEEE FP16 uses 5 exponent bits and 10 fraction bits, offering more fraction precision than BF16 but less dynamic range.
- Google Brain Float 16 (BF16) keeps FP32's 8-bit exponent and reduces the fraction to 7 bits, preserving FP32-like range in a 16-bit format.
- BF16 is useful for AI training and inference because dynamic range often matters more than fraction precision.
- FP8 formats make a similar trade-off: E4M3 provides 4 exponent and 3 mantissa bits, while E5M2 offers greater range for uses such as gradients.
- An FP16 exercise decodes sign 1, stored exponent 17, and fraction 0.75 as −(1 + 0.75) × 2² = −7.
- A BF16 exercise encodes 2.5 as a positive value with exponent 1 and fraction 0.25.
- NVIDIA's FP8 E4M3 and E5M2 formats allocate bits differently to balance precision and dynamic range.
- FP8 is used in AI workloads including attention and feed-forward network layers.
- Four-bit designs distribute three non-sign bits between exponent and mantissa; the choices include integer-like formats and E2M1.
- E2M1 places more representable values near zero and spaces larger magnitudes farther apart, fitting concentrated weight distributions.
- An exponent-only E3M0 format spans powers of two such as 1, 2, 4, 8, 16, 32, and 64, but offers coarse spacing.
- Some open models are released with 4-bit weights and activations, illustrating the push toward lower-bit representations.
- Quantization maps continuous values to a discrete set; K-means groups similar weights and represents each by a centroid.
- In the 4×4 example, four centroids—−1, 0, 1.5, and 2—are stored, while each weight stores a 2-bit codebook index.
- The toy matrix shrinks from 64 bytes of FP32 weights to 20 bytes: 4 bytes of indices plus a 16-byte FP32 codebook, or 3.2× smaller.
- For a large matrix, codebook overhead becomes negligible; n-bit indices approach a 32/n storage reduction versus FP32.
- K-means quantized training aggregates gradients for weights assigned to each centroid, then updates that centroid using the learning rate.
- In Deep Compression experiments, pruning and quantization together allowed model size to fall to roughly 3–4% before accuracy began to drop, outperforming either alone.
- AlexNet convolution layers could be quantized to about 4 bits and fully connected layers to about 2 bits without accuracy loss in the reported results.
- Retraining shifts centroids slightly to recover accuracy after quantization.
- Deep Compression combines pruning, K-means quantization, and Huffman coding; frequent indices receive shorter variable-length codes.
- The reported pipeline reaches about 9–13× size reduction after pruning and 27–31× after quantization, with Huffman coding applied afterward.
- Distillation followed by pruning and then quantization is a common ordering because pruning first removes redundant weights while preserving flexibility for fine-tuning.
- SqueezeNet's 1×1 and 3×3 convolution design combined with Deep Compression achieved about 500× smaller size than AlexNet at the same reported accuracy.
- Unlike K-means indices, linear quantization directly operates on integer values and maps them to real values using a scale and zero point.
- The zero point ensures real zero maps exactly to an integer code, avoiding quantization error for zero.
- For real range [Rmin, Rmax] and integer range [Qmin, Qmax], the scale is S = (Rmax − Rmin)/(Qmax − Qmin).
- For the example's 2-bit range [−2, 1] and real extrema −1.08 and 2.12, S is about 1.07 and the rounded zero point is −1.
- Linear quantization reconstructs each real tensor as its scale multiplied by the quantized integer minus its zero point.
- Expanding Y = WX produces integer matrix products plus zero-point correction terms that can often be precomputed.
- When weight distributions are approximately symmetric around zero, the weight zero point can be set to zero, simplifying the arithmetic.
- Scale factors can be implemented using integer multiplication and shifts, while accumulations commonly use 32-bit integers.
- For Y = WX + B, the bias can be quantized to the product scale, allowing its integer representation to join the accumulator.
- The same affine-quantization derivation applies to convolution because convolution is a linear operation.
- A typical integer data path convolves quantized inputs and weights, adds quantized bias, applies scaling, and adds the output zero point.
- The result is an integer-domain computation path for both matrix multiplication and convolution.
- Reported 8-bit integer results on ResNet-50 and Inception-v3 largely recovered floating-point accuracy while improving latency–accuracy trade-offs.
- K-means quantization reduces storage with integer indices and a floating-point codebook, but computation remains floating point after lookup.
- Linear quantization can reduce both storage and computation by using integer weights, activations, and arithmetic.
- The lecture concludes with binary and ternary quantization as the next topic and points students to the EfficientML.ai lab materials.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.