EfficientML.ai Lecture 2 - Basics of Neural Networks (MIT 6.5940 Fall 2026)
Watch on YouTube →
Overview
EfficientML.ai Lecture 2 connects neural-network building blocks—from fully connected and convolutional layers to Transformers—with the efficiency costs that shape their use. It compares AlexNet, VGG-16, ResNet-50, and MobileNetV2, then develops practical metrics for parameters, model and activation memory, MACs, FLOPs, latency, throughput, and data movement, emphasizing that moving data can consume far more energy than arithmetic.
Key takeaways
- Batching improves GPU efficiency by applying the same fully connected layer weights to multiple inputs in a matrix-matrix operation, but serving systems must balance that reuse against the time users wait for a batch to form.
- For stride-1 convolutions with kernel size K, each added layer expands the receptive field by K − 1; strided convolution expands it more quickly but reduces feature-map size.
- Latency and throughput are distinct: four engines each taking 100 ms per image can deliver 40 images per second while each image still has 100 ms latency.
- Reducing arithmetic alone may not save much energy when memory traffic remains high: the lecture cites about 640 pJ for a DRAM access versus about 0.1 pJ for a 32-bit integer addition.
- Parameter count is not enough to predict whether a model fits on a device: MobileNet’s channel expansion can increase peak activation memory even as depthwise convolution reduces computation and weights.
- Quantization directly reduces weight storage: a 14-billion-parameter model at 4 bits per parameter requires about 7 GB, compared with about 56 GB at 32 bits.
Chapters
- AI compute demand has grown faster than GPU memory supply, widening the gap between available hardware and model requirements.
- The lecture previews neural-network layers, classic architectures, efficiency metrics, and a PyTorch tutorial.
- Lab Zero introduces PyTorch, GPU basics, FLOP counting, and Google Colab; it is ungraded and prepares students for later labs.
- A neural network combines inputs using learned weights and applies activations; depth counts layers, width counts activations, and parameters count learned weights.
- A fully connected layer maps CI input features to CO outputs using a CO × CI weight matrix and CO biases.
- Batching turns matrix-vector operations into matrix-matrix operations, reusing weights across inputs to improve GPU efficiency.
- Dynamic batching trades higher throughput for possible waiting time while requests accumulate.
- Convolution filters combine local spatial regions across input channels; a 2D image is represented as a tensor with spatial dimensions plus channels such as RGB.
- With stride 1 and no padding, output height is input height minus kernel height plus one; a 5 × 5 input with a 3 × 3 kernel produces 3 × 3 outputs.
- Padding, including zero, reflection, or replication padding, can preserve feature-map dimensions; a 3 × 3 kernel with padding 1 preserves height and width.
- For stride-1 layers with kernel size K, L stacked convolutions have receptive field L(K − 1) + 1; strided convolution expands receptive field faster while shrinking output maps.
- Grouped convolution partitions channels so connections occur only within each group, reducing parameters and computation; depthwise convolution is the extreme case with one group per channel.
- Max pooling selects the largest value in each window, while average pooling computes the window mean; a 2 × 2 window is a common example.
- Batch normalization, layer normalization, instance normalization, and group normalization differ in which dimensions are used to calculate mean and standard deviation.
- Normalization typically applies learnable scale and shift parameters after centering and scaling.
- ReLU is inexpensive but can leave neurons inactive below zero and has an unbounded positive range that complicates quantization.
- ReLU6 caps outputs at 6; leaky ReLU preserves a negative-side gradient, while Swish has nonzero gradients but requires more expensive exponential and division operations.
- Hard-Swish approximates Swish with a piecewise function, offering a hardware-friendlier activation.
- Transformer attention uses query, key, and value projections: normalized QK scores are passed through softmax and used to combine values; a deeper Transformer treatment is reserved for the second half of the course.
- AlexNet starts from 224 × 224 RGB images, uses an 11 × 11 convolution with stride 4 to reduce activation size, and ends with fully connected layers for 1,000 classes.
- VGG-16 uses regular 3 × 3 convolutions and channel widths such as 64, 128, and 256; its regularity aids implementation but the model is bulky.
- ResNet-50 adds residual connections so blocks learn a delta; its bottleneck uses 1 × 1 convolutions to reduce and then restore channels around a 3 × 3 convolution.
- MobileNetV2 uses depthwise convolution with 1 × 1 expansion and projection layers; expanding channels by as much as six times supplies capacity while keeping depthwise computation lightweight.
- Efficiency has three practical goals: smaller models for storage and downloads, faster responses, and lower energy use on phones and in data centers.
- Latency measures the time to finish one request; throughput measures how many images or tokens are processed per second.
- One engine taking 50 ms per image has 50 ms latency and 20 images/s, while four parallel engines taking 100 ms each have 100 ms latency and 40 images/s.
- Higher throughput does not necessarily mean lower latency, and lower latency does not necessarily mean higher throughput.
- Latency can be reduced by pipelining computation with data movement—for example, loading the next image while the current image is being computed.
- Compute time depends on model operations divided by processor operations per second; weight-transfer time depends on model size divided by memory bandwidth.
- Activation-transfer time depends on input and output activation sizes and the processor’s memory bandwidth.
- The cited energy comparison puts a 32-bit integer add at about 0.1 pJ and a DRAM access at about 640 pJ, making data movement roughly two orders of magnitude more expensive than arithmetic.
- A fully connected layer with CI inputs and CO outputs has CI × CO weights; convolution weights additionally scale with kernel height and width and input/output channels.
- Grouped convolution divides the corresponding convolution parameter count by the number of groups; depthwise convolution reduces it to kernel area times channel count.
- A 14-billion-parameter model stored at 4 bits per parameter needs about 7 GB for weights, motivating quantization.
- AlexNet’s roughly 61 million parameters occupy about 244 MB at 32 bits per weight or 61 MB at 8 bits per weight.
- Total activation memory sums layer activations, while peak activation memory is governed by the largest simultaneously needed set, often input plus output activations.
- MobileNet can reduce parameter count substantially versus ResNet while increasing peak activation needs: channel expansion in depthwise blocks raises activation memory by about 1.8× in the cited comparison.
- For inference, peak activation memory can determine whether a model fits; early high-resolution layers can bottleneck deployment on devices with only 256 KB of memory.
- During CNN training, peak activation memory remains important even when model parameters shrink; AlexNet is cited as having about 932,000 total activations and roughly 440,000 peak activations.
- One MAC is one multiplication plus one accumulation; a matrix multiplication with dimensions M × K and K × N requires MKN MACs.
- Convolution MAC counts scale with output channels, output height and width, input channels, and kernel area; grouped convolution divides the count by group count.
- AlexNet is estimated at about 724 million MACs, equivalent to roughly 1.4 billion FLOPs because each MAC counts as two floating-point operations.
- FLOPS measures floating-point operations per second; for quantized integer workloads, operations per second is a more general measure.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, MIT HAN Lab.