EfficientML.ai Lecture 8 - Neural Architecture Search (Part II) (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
The lecture explains how Neural Architecture Search (NAS) can reduce search cost and find models tailored to real hardware, covering weight-sharing and accuracy estimation, hardware-aware latency measurement, Once-for-All subnetworks, zero-shot scoring, and neural–accelerator co-search. Applications include Jet-Nemotron, which combines attention types to improve throughput while retaining accuracy, plus efficient models for transformers, point clouds, GANs, pose estimation, and flexible large language models.
Key takeaways
- FLOPs are not a reliable substitute for measured latency: GPU parallelism can make a wider model nearly as fast, while additional depth can increase kernel-launch overhead.
- Once-for-All amortizes architecture training across an estimated 10^19 subnetworks by sharing weights and varying kernel size, depth, width, and input resolution.
- Jet-Nemotron combines mostly linear attention with a small number of full or sliding-window attention layers to target linear-attention efficiency without giving up full-attention-level accuracy.
- Hardware-aware NAS should measure or predict performance on the deployment target because GPU, CPU, mobile, and Raspberry Pi platforms favor different model shapes.
- Joint neural–accelerator co-search can improve both efficiency and accuracy: the reported example cuts energy-delay product by about 4× and gains 2.7% accuracy when Once-for-All NAS is added.
- Shared supernetworks support practical deployment choices beyond model size, including Anycost GAN’s fast previews and high-quality finalization and Flextron’s device-specific LLM configurations.
Chapters
- NAS searches for architectures that improve the Pareto frontier across accuracy, latency, throughput, memory footprint, and energy.
- The lecture builds on pruning, quantization, and primitive-operation cost analysis to focus on efficient architecture design.
- Its central challenge is making both the search process and the resulting models efficient on target hardware.
- Training every candidate from scratch is expensive: searching roughly 12,000 architectures on CIFAR-10 was estimated at about 22,000 GPU-hours.
- Net2Net reuses weights when widening a network by splitting units and adjusting their weights, preserving the original function.
- To deepen a network, Net2Net inserts identity layers so the expanded model initially computes the same function.
- A hypernetwork can generate candidate weights from graph neural network embeddings of an architecture, but the approach is described as uncommon in current practice.
- Proxy-based NAS may search on CIFAR-10, small search spaces, short training runs, or FLOPs instead of the actual task and deployment hardware.
- ProxylessNAS trains an over-parameterized network with competing operations, such as 3×3 and 5×5 convolutions, and learns architecture probabilities alongside weights.
- The trained architecture parameters select paths that can be pruned, leaving a single active path at inference and reducing memory use.
- Searching directly on ImageNet and using measured latency as a reward avoids relying on FLOPs as a proxy for speed.
- Latency differs by hardware: increasing hidden width can have little effect on a GPU but raise latency nearly linearly on a Raspberry Pi with limited parallelism.
- Increasing network depth can add latency through extra kernel launches, even when FLOPs are similar.
- A latency lookup table can sum per-operator costs; a learned predictor can estimate latency from features such as kernel size, width, and input resolution.
- ProxylessNAS produced a model with about 1.88× lower latency than a hand-designed mobile model at roughly 74.7% accuracy; device-specific searches also favored wider, shallower GPU models and deeper CPU models.
- Searching separately for GPUs, CPUs, phones, and microcontrollers multiplies the cost of training and evaluation.
- Once-for-All trains a supernetwork once, then selects independently usable child networks for platforms such as H100 GPUs, desktop GPUs, mobile GPUs, or Cortex-M microcontrollers.
- The same model can adapt to new and old phones or reduce its compute demands in battery-saving mode without storing separately trained large, medium, and small models.
- The lecture states that a supernetwork can support approximately 10^19 subnetworks.
- Kernel sizes shrink progressively from 7×7 to 5×5 to 3×3 by reusing centered weights from the larger kernels.
- Elastic depth allows subnetworks to exit after different numbers of layers, while elastic width selects different channel counts.
- Channels are ranked by importance, for example using L2 norms, so narrower subnetworks retain the most important channels.
- The FPGA and cross-device results show Once-for-All models improving the accuracy–efficiency trade-off across phones, Intel CPUs, and FPGAs.
- Jet-Nemotron uses NAS on a pretrained language model to choose among full attention, linear attention, and sliding-window attention rather than training a new model from scratch.
- Full attention preserves capacity but grows KV-cache memory with context length; linear attention has fixed memory but can lose accuracy.
- The search uses only a small number of full or sliding-window attention layers, selects Gated DeltaNet among linear-attention choices, and adds input-conditioned dynamic convolution to restore capacity.
- The reported system reaches up to 2,800 tokens per second, with roughly 100–200 MB of KV cache versus about 1,800 MB for cited full-attention baselines, while maintaining strong benchmark accuracy.
- Zero-shot NAS estimates candidate quality from architecture behavior without the costly step of training each model.
- ZEN-NAS uses input perturbations to test whether model outputs respond meaningfully, and also considers batch-normalization variance.
- GradSign scores whether gradients from different samples share the same sign, treating denser sample-wise local minima as a favorable signal.
- These heuristics predate current large language models, so their effectiveness for modern LLM architectures remains uncertain.
- Neural architecture and accelerator choices are tightly coupled, so optimizing them separately can miss better energy and speed trade-offs.
- The search space includes accelerator buffers, processing-element count and connectivity; compiler loop order, tiling, and dataflow; and network depth, width, attention type, precision, and sparsity.
- Loop mappings contain non-numeric choices such as which tensor dimensions to parallelize and the order of nested loops.
- Importance-based encoding ranks dimensions numerically, then maps the top-ranked dimensions to spatial parallelism and the remaining dimensions to temporal execution.
- An evolutionary search jointly explores neural architectures, accelerator designs, and mappings using hardware performance estimates.
- Hardware architecture search reduces normalized energy-delay product by about 4× in the reported example.
- Adding Once-for-All NAS to the hardware search raises accuracy by 2.7% while retaining the efficiency improvement.
- The resulting design includes a 496-byte buffer, a 107 KB output buffer, and an 18×10 array, with a dataflow unlike the traditional human-designed mapping.
- A Once-for-All transformer shares weights across large GPU, medium CPU, and small IoT configurations; the lecture’s demo uses a transformer chip for intent and token classification.
- For point-cloud understanding in autonomous driving, elastic depth and channel counts are combined with evolutionary search over accuracy and speed.
- The point-cloud example improves throughput from about 3.4 frames per second for MinkowskiNet to 9.1 frames per second for the searched model.
- The same supernetwork-and-subnetwork strategy supports device-specific deployment without retraining an independent model for each target.
- Anycost GAN supports a fast, low-cost subnetwork for interactive image editing and a larger, slower subnetwork for one-time high-quality finalization.
- The demo adjusts attributes such as smiles, apparent age, hair color, glasses, and facial hair, with small subnetworks making previews more responsive.
- LightPose specializes pose-estimation subnetworks for edge hardware such as Qualcomm Snapdragon chips and reports about 4.9× lower latency with improved accuracy.
- The lecture also highlights on-device detection, gaze estimation, and segmentation as tasks where lightweight searched networks can run locally.
- Flextron derives models of different sizes from a pretrained LLM, targeting mobile devices, laptops, and GPUs without training every size from scratch.
- It ranks attention heads and feed-forward channels, then selects different proportions, such as 25%, 50%, or 75%, for subnetworks.
- A trained router chooses how much of the model to activate; the lecture describes producing 7B, 5B, 4B, and 3B variants.
- The lecture concludes with performance estimation, hardware-aware NAS, zero-shot NAS, co-search, and applications, then previews knowledge distillation.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.