EfficientML.ai Lecture 7 - Neural Architecture Search (Part I) (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
Song Han frames neural architecture search (NAS) as a way to find models that balance accuracy against latency, energy, memory, and storage, building on efficient primitives such as grouped and depthwise convolution and transformer attention. He explains how to define and narrow cell- and network-level search spaces, then compares grid and random search, reinforcement learning, differentiable architecture search, and evolutionary search; TinyML illustrates why the search space must fit real memory constraints.
Key takeaways
- A model's parameter count is not a reliable proxy for training-memory efficiency: MobileNetV2's 6× channel expansion can create large intermediate activations despite reducing weights and computation.
- Depthwise convolution saves compute by processing channels independently, but pointwise convolution or channel mixing is needed to restore cross-channel communication.
- Standard self-attention has O(N²D) compute in both its score and value-mixing stages, making long token sequences a central efficiency challenge.
- A NAS search space must be broad enough to contain useful architectures but constrained enough to search effectively; TinyNAS uses memory-aware FLOPs as a heuristic for choosing promising designs.
- Differentiable NAS can optimize latency as well as accuracy by combining operation-selection probabilities with premeasured hardware latencies.
- Grid search is systematic but quickly becomes expensive, while random search is a useful early debugging baseline and evolutionary search explores through mutation and crossover.
Chapters
- Efficient model design balances accuracy against latency, energy, memory, and storage rather than optimizing one metric in isolation.
- For language models, latency includes time to first token, while throughput measures tokens generated per second.
- Model size affects cloud decoding, edge-device storage, over-the-air updates, and app downloads; energy also constrains data centers, phones, robots, and cars.
- A fully connected layer maps CI input channels to CO outputs with CI × CO multiply-accumulates per example.
- A 2D convolution's compute scales with input and output channels, kernel height and width, and output height and width.
- Grouping channels into G independent groups reduces convolution compute by roughly G; the depthwise extreme uses one group per channel and relies on pointwise convolution to mix channels.
- A bottleneck block uses a 1×1 projection to shrink channels before an expensive 3×3 convolution, then a second 1×1 layer restores the output width.
- For the example, the middle layer operates at 512 channels instead of roughly 2,000, lowering the dominant spatial-convolution cost.
- The lecture estimates this arrangement at about 8.5× less computation than a direct 3×3 convolution between the wider input and output dimensions.
- MobileNetV2's inverted bottleneck expands channels by 6×, applies a cheap depthwise 3×3 convolution, then projects back to the original width.
- With 160 input channels, the expansion produces 960 channels; the depthwise layer uses only 960 × 9 kernel parameters.
- The design can cut parameters substantially, but expanded intermediate activations can raise memory use: the lecture cites 80% higher peak activation memory in one comparison and only about a 10% activation reduction in another.
- 1×1 group convolution lowers computation but prevents channels in different groups from exchanging information.
- Channel shuffle rearranges channel assignments between grouped layers so subsequent groups can combine information from earlier groups.
- The resulting building block combines grouped pointwise convolution, channel shuffle, depthwise 3×3 convolution, and another grouped pointwise layer.
- Multi-head self-attention projects inputs into query, key, and value tensors; attention scores come from QKᵀ, followed by scaling, optional causal masking, softmax, and multiplication by V.
- For N tokens and hidden dimension D, both the score calculation and attention-weighted value calculation cost O(N²D).
- Attention maps can expose token relationships—for example, strong video–game attention—and may support dynamic token pruning even though the attention matrix itself has no weights to prune.
- NAS automates proposing, evaluating, and refining architectures across choices such as layer count, kernel size, channels, connectivity, and resolution.
- Cell-based designs repeat normal cells that preserve resolution and insert reduction cells that lower resolution and may increase channel count.
- An RNN controller can construct a cell by selecting two inputs, choosing an operation for each, and combining their outputs with operations such as addition or concatenation.
- With two choices for each of two inputs, M operations per input, and N ways to combine outputs, a B-layer cell search has 4ᴮ × M²ᴮ × Nᴮ possibilities.
- Using M = 5, N = 2, and B = 5 yields roughly 10¹¹ candidate designs.
- Network-level choices add depth, image resolution, channel width, kernel size, and topology; deeper or higher-resolution models can improve capacity but cost more compute and memory.
- TinyML targets microcontrollers with about 320 KB of memory and 1 MB of storage, far below mobile devices with several gigabytes of memory.
- TinyNAS first narrows an overly broad design space, then specializes models to the target resource constraints.
- Under a fixed memory footprint, the lecture uses higher FLOPs as a heuristic for greater model capacity; the cited search-space comparison reports 78.7% accuracy for a designed space versus 74.7% for random search, while a ResNet-18-derived space reached 80.3% but ran out of memory.
- Grid search exhaustively tests combinations such as EfficientNet's depth, width, and resolution scaling, but its cost grows rapidly with each added dimension.
- Random search samples configurations rather than systematically enumerating every combination.
- Random search is useful as an implementation check: seeing no improvement after several trials can signal a bug, though improvement alone does not prove the search code is correct.
- An RNN controller treats architecture construction as sequential decisions, trains each proposed child network, and uses its accuracy reward to update the controller.
- Differentiable NAS represents candidate edges as weighted choices and combines their activations, allowing architecture probabilities to be optimized with gradients.
- Expected latency can be added to the loss as a probability-weighted sum of premeasured operation latencies, alongside cross-entropy and regularization.
- Evolutionary NAS evaluates a population with a fitness function that rewards both accuracy and efficiency, then generates candidates through mutation and crossover.
- Mutations can change stage depth or replace a 3×3 convolution with a 5×5 or 7×7 operation.
- Crossover selects layers from different parent architectures to form the next generation; the lecture closes by previewing hardware-aware NAS for the following session.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.