EfficientML.ai Lecture 1 - Introduction (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
Song Han introduces MIT 6.5940 as a hands-on course on efficient AI computing, arguing that model growth is outpacing GPU memory and making compression, optimized algorithms, and system design essential. Examples range from microcontroller training and faster image generation to four-bit LLM deployment with AWQ; five graded labs culminate in running a seven-billion-parameter model locally on a laptop.
Key takeaways
- Model sizes are growing much faster than GPU memory capacity, so pruning, sparsity, and quantization are central to reducing the gap between AI demand and available compute.
- Efficient AI must optimize more than arithmetic: data movement across memory and networks can be more expensive than computation, making communication-aware system design essential.
- On-device training can preserve privacy and enable local adaptation, but it is substantially more memory-intensive than inference; Han previews a reduction from over 600 MB to 141 KB.
- Four-bit quantization paired with AWQ and a tailored inference engine can make multi-billion-parameter language models run locally, including a seven-billion-parameter model on a Jetson Nano.
- The course connects algorithm design to practical performance measurement: its labs pair techniques such as quantization with deployment systems so students can verify real speedups and build a laptop chatbot.
Chapters
0:00
Song Han Introduces MIT 6.5940 and Its Hands-On Goal
- Song Han describes the course’s growth from 30 students in 2022 to 300 students this term.
- The practical curriculum emphasizes code and demonstrations over equations, with five graded labs and a goal of running a seven-billion-parameter model on a laptop.
- Han’s work includes Deep Compression and AWQ, an activation-aware quantization method reported to have more than 60 million downloads.
7:06
Why Model Growth Outpaces GPU Memory
- Han contrasts relatively steady GPU-memory growth—from about 16 GB on the P100 to 192 GB on the B200—with models that have grown from millions to hundreds of billions or trillions of parameters.
- The widening gap between compute supply and model demand motivates compression techniques such as pruning, sparsity, and quantization.
- The course frames the goal as reducing memory and compute while preserving model accuracy.
10:04
Vision Models: From ImageNet Accuracy to On-Device Learning
- ImageNet error rates fell from about 28% in 2010 to roughly 3% with ResNet, but higher accuracy generally required more multiply-accumulate operations.
- Hardware-aware architecture design and neural architecture search can move the accuracy–compute frontier; Han cites achieving comparable or better accuracy with up to 14 times less computation.
- Phone applications include local photo tagging, face recognition, and pose estimation; MCU-based systems can detect faces, masks, and people with only hundreds of kilobytes of SRAM.
- On-device learning offers privacy and personalization, but training is more demanding than inference; the course previews reducing training memory from over 600 MB to 141 KB.
18:15
EfficientViT-SAM Speeds Up Image Segmentation
- Segment Anything Model (SAM) uses a user-selected point to segment objects, with potential applications in augmented and virtual reality.
- Han presents EfficientViT-SAM as retaining accuracy while delivering up to 48 times the speed in a cited comparison.
- One demo increases segmentation throughput from about 11 to 182 images per second on the same GPU.
20:02
Making Diffusion-Based Image and Video Generation Cheaper
- Image and video generation are compute-intensive: Han cites Stable Diffusion training costing about $600,000 on 256 A100 GPUs.
- Attention computation grows quadratically with token count, motivating techniques such as linear attention, block-sparse attention, and token reduction.
- GAN Compression uses pruning to reduce computation by an order of magnitude; a CycleGAN horse-to-zebra demo improves from 12 to 40 frames per second.
- Anycost GAN shares weights across smaller and full models, enabling fast interactive edits before running the full model for a final result.
26:50
Distributed Inference and Efficient 3D Perception
- Accelerated image inpainting reduces a cited Stable Diffusion workload from 1,800 to about 500 GMACs and from 369 ms to under 100 ms.
- Multi-GPU inference requires managing communication: four GPUs can produce one coherent image in about four seconds rather than generating artifacts.
- Han emphasizes that moving data across memory and networks can cost more than computation, motivating gradient compression and quantization.
- For autonomous driving, Pass the Lighter Net raises LiDAR perception from 5 to 47 frames per second, while camera–LiDAR fusion enables bird’s-eye-view segmentation on mobile GPUs.
32:05
LLM Capability Comes with Model-Size and Token Costs
- Large language models support code generation, translation, and zero-shot or few-shot tasks, but accuracy improvements diminish as parameter counts grow.
- Chain-of-thought prompting can improve multi-step reasoning, as in the example of calculating that 23 apples minus 20 plus 6 equals 9.
- Generating intermediate reasoning requires many more tokens than returning a short answer, increasing inference cost.
- Training and serving large models also demand substantial GPU infrastructure, electricity, cooling, and physical space.
37:29
Sparse Attention and Four-Bit LLM Deployment
- Sparse attention reduces work by dropping less-relevant tokens; the sentiment example is shortened from 11 tokens to “treat,” “film,” and “perfect,” then to two salient tokens.
- Quantization reduces model storage by using fewer bits, but activation outliers make straightforward four-bit conversion difficult.
- SmoothQuant redistributes activation and weight ranges to make quantization easier; Han previews implementing it alongside a compact inference engine.
- AWQ-based deployment can run a seven-billion-parameter LLaMA model on a Jetson Nano at about 30 tokens per second; another comparison reports throughput rising from 50 to 166 tokens per second.
43:09
Vision-Language Models, Robotics, and Physical AI
- Vision-language models combine image or video tokens with text, using a vision transformer and projector to align visual inputs with language.
- AWQ can preserve useful visual reasoning after four-bit compression; examples include identifying the Mona Lisa and interpreting a meme.
- A robot ping-pong system runs at 15 Hz on an RTX 5090 laptop, while a world action model predicts future visual states to guide a robot making coffee.
- Han cites the high cost of landmark systems such as AlphaGo to motivate efficiency across algorithms, hardware, and data.
49:10
Hardware Trends Make Software and Parallelism Essential
- As single-thread performance and typical power plateau, multicore parallelism and algorithm–system co-design become increasingly important.
- NVIDIA GPU inference performance improved by more than 300 times over eight years in Han’s comparison, helped by tensor-core matrix operations, INT8 support, and structured sparsity.
- The course focuses on software techniques that better use existing hardware rather than designing new chips.
- The hardware survey spans cloud GPUs, Jetson edge systems, Qualcomm Snapdragon and Neural Processing Units, and microcontrollers with only hundreds of kilobytes of memory.
55:44
Course Modules, Schedule, and Prerequisites
- The course has three modules: efficient inference, efficient training, and application-specific optimization for LLMs, image and video generation, and point clouds.
- Inference topics include pruning, quantization, neural architecture search, and distillation; training topics include gradient compression, on-device training, and federated learning.
- Lectures meet Tuesdays and Thursdays, with Thursday office hours; Piazza is used for discussion and Canvas for submissions.
- The course expects machine-learning and computer-architecture preparation and includes algorithm and system code design.
58:43
Five Labs, Final Project, and Course Outcomes
- The five graded labs cover GPU basics, pruning, neural architecture search, quantization, and LLM deployment; an introductory PyTorch Lab 0 is ungraded.
- Labs account for 70% of the grade and a group project for 30%; Han recommends combining the quantization and deployment labs to build a laptop chatbot.
- The deployment labs support macOS, Linux, and Windows, require at least 8 GB of laptop memory and 5 GB of storage, and target local LLM execution.
- Han closes by emphasizing measured efficiency, practical trade-off analysis, current research, and hands-on model deployment.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.