CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 12
Watch on YouTube →
Overview
Lecture 12 completes the CNN backpropagation review, showing how gradients pass through convolution, pooling, and resampling before examining transform invariance, object localization, and parameter-efficient depthwise convolutions. It connects learned visual features and data augmentation to CNN history, highlighting LeNet-5 and AlexNet’s 2012 ImageNet result, which helped move top-five error from roughly 25% to 18% (and 15% with an ensemble).
Key takeaways
- Backpropagation through CNN operations can be derived by reversing the forward computation: max pooling routes each gradient to its selected maximum, mean pooling distributes it evenly, and resampling passes gradients only through entries that depend on the input.
- Strided convolution is easier to reason about as ordinary convolution plus downsampling, while fractional-stride convolution can be decomposed into upsampling plus convolution; backward implementations must preserve input sizes at boundaries.
- CNN weight sharing gives position invariance, but rotation and scale invariance require transformed filters or training examples; augmentation is typically more practical because enumerating transforms can multiply feature channels.
- Object localization needs spatial predictions and labeled coordinates: a bounding-box head can be trained alongside classification, and additional heads can predict orientation or pose joints.
- CNN receptive fields expand with depth, supporting a hierarchy from oriented lines to object parts and whole objects; this pattern resembles visual-processing findings associated with Hubel and Wiesel.
- AlexNet’s 2012 ImageNet result—about 18% top-five error versus roughly 25% previously, and 15% with a seven-model ensemble—demonstrated that large-scale CNNs could sharply improve visual recognition.
Chapters
- The lecture begins with informal conversation before the class starts around 6:00.
- Attendance and name tags are discussed in a class of roughly 350 students.
- CNNs detect patterns regardless of position by scanning inputs with shared filters; the same idea applies to 1D, 2D, or 3D data.
- Image-level labels can be weak: a flower image is labeled as containing a flower without specifying its location.
- Training uses a forward pass followed by backpropagation of the loss through the network.
- Activation derivatives are computed pointwise using the chain rule at each location in an activation map.
- Input gradients use transposed, spatially flipped filters convolved with appropriately padded output gradients.
- Filter gradients come from convolving input channels with the corresponding affine-output gradients, then accumulating shared-filter contributions.
- For max pooling, only the input that supplied each window’s maximum receives that output’s gradient.
- Overlapping max-pooling windows require gradients to be added when the same input is the maximum in multiple windows.
- For mean pooling, each output gradient is distributed evenly over every input in its window.
- Downsampling copies selected input entries, so backward propagation copies their gradients back and leaves dropped positions at zero.
- Upsampling inserts zeros that do not depend on the input; gradients at those inserted positions do not propagate, while copied entries pass gradients back.
- Convolution with an integer stride can be viewed as convolution followed by downsampling; fractional-stride convolution can be viewed as upsampling followed by convolution.
- A practical backpropagation strategy is to reverse the operations and loops used in the forward computation.
- For a multiply-and-add operation, gradients are propagated to both inputs using the product rule and accumulated when variables are reused.
- The method also handles strided convolution and pooling by reversing the corresponding loops, provided the forward code computes the intended operation.
- A CNN’s shared filters provide position-invariant pattern detection, but do not inherently provide invariance to rotation, reflection, or scale.
- A mathematically explicit approach transforms each filter and scans with every transformed copy, producing additional output channels.
- Gradients from transformed filters must be mapped back and combined to update the original filter.
- Enumerating rotations, reflections, and scales can make channel counts and computation grow rapidly across layers.
- Training-data augmentation instead supplies transformed examples, such as rotated or reflected images, without multiplying model channels.
- Augmentation only encourages invariance to the transformations represented in training; it does not guarantee exact general invariance.
- A spatial output map can indicate where a flower detector fires, preserving location clues that a single pooled classification score would discard.
- Object localization adds a prediction head for bounding-box coordinates and trains it with location labels and a suitable loss.
- Additional heads can predict orientation or body-joint coordinates, as in pose estimation systems that regress joint positions.
- ResNet-style skip connections add a convolution block’s output back to its input, helping gradients pass through very deep networks.
- The lecture introduces depthwise convolution as a way to reduce computation by convolving input channels individually rather than computing a separate full set of channel convolutions for every filter.
- In the described channel-combination scheme, output channels differ through weights used to combine the individually convolved channels, trading expressive power for fewer parameters and operations.
- For n input channels and k standard filters, the lecture counts n × k channel-wise convolutions before producing k output channels.
- The described depthwise approach performs n convolutions, then uses different channel weights to form outputs.
- A review question emphasizes that channels are processed separately rather than convolved together in one operation.
- A first-layer filter responds to patterns within its local input patch.
- Later-layer filters operate on earlier feature maps, so their effective receptive fields cover larger regions of the original image.
- Deeper representations combine increasingly complex functions of earlier features, making a single fixed visual pattern harder to assign to each filter.
- Early CNN filters often learn oriented-line detectors, echoing receptive-field patterns observed in Hubel and Wiesel’s studies of visual cortex.
- Examples show deeper layers responding to face parts such as eyes and noses, car parts such as wheels, and eventually whole object types.
- For a particular input, backpropagation can help identify which regions contributed to a deeper-layer response.
- Large CNNs can require optimization aids such as Adam, momentum, and batch normalization, and may outgrow the available labeled data.
- Common augmentations include rotations, translations, scaling, flips, shears, and small distortions that preserve the class label.
- CNNs are used beyond images, including speech, audio, and text; LeNet-5, associated with Yann LeCun, classified handwritten digits before the ImageNet era.
- ImageNet’s large-scale recognition task used millions of images across 1,000 classes; earlier top-five error was around 25%.
- AlexNet trained on about 1.2 million high-resolution images with 60 million parameters, five convolutional layers, and three fully connected layers.
- Its design included ReLU activations and training across two GPUs; the reported top-five error fell to about 18%, or 15% when seven networks voted together.
- AlexNet’s final-layer features grouped semantically similar images: querying with a flower, elephant, ship, or jack-o’-lantern retrieved related examples using Euclidean distance.
- The lecture traces reported ImageNet top-five error improvements from AlexNet to later systems, including VGG at 7.3%, GoogLeNet at 6.7%, and ResNet at about 3.5%.
- The combination of strong accuracy and semantically meaningful representations helped catalyze the modern deep-learning boom.
- The lecture concludes that CNNs remain widely used across vision and other domains, with many architectural variants still appearing in research.
- The next class will switch topics to time-series models.
- Students are directed to Piazza or office hours for follow-up questions.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Carnegie Mellon University Deep Learning.