Save this video — free

EfficientML.ai Lecture 8 - Neural Architecture Search (Part II) (MIT 6.5940 Fall 2026)

MIT HAN Lab · 1:09:16 · Watch on YouTube

EfficientML.ai Lecture 8 - Neural Architecture Search (Part II) (MIT 6.5940 Fall 2026) Watch on YouTube →

Overview

The lecture explains how Neural Architecture Search (NAS) can reduce search cost and find models tailored to real hardware, covering weight-sharing and accuracy estimation, hardware-aware latency measurement, Once-for-All subnetworks, zero-shot scoring, and neural–accelerator co-search. Applications include Jet-Nemotron, which combines attention types to improve throughput while retaining accuracy, plus efficient models for transformers, point clouds, GANs, pose estimation, and flexible large language models.

Key takeaways

Chapters

0:00 NAS Goals: Accuracy, Latency, Memory, and Energy
2:00 Estimating Candidate Accuracy Without Full Retraining
8:05 ProxylessNAS Uses Direct Hardware Feedback
17:53 Predicting Latency and Specializing Models by Device
23:00 Once-for-All: One Training Run, Many Hardware-Specific Models
28:00 Elastic Kernel, Depth, Width, and Resolution in Once-for-All
34:25 Jet-Nemotron Searches Hybrid Attention for LLM Efficiency
41:38 Zero-Shot NAS Scores Architectures Without Training
45:25 Co-Searching Neural Networks, Accelerators, and Compilers
54:00 Co-Search Results: Lower Energy-Delay Product and Higher Accuracy
56:53 Once-for-All Transformers and Point-Cloud Models
1:01:00 Anycost GANs and Lightweight On-Device Pose Estimation
1:06:10 Flextron Makes Pretrained LLMs Adaptable to Different Devices

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.