CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 10
Watch on YouTube →
Overview
The lecture traces convolutional neural networks from Hubel and Wiesel’s studies of cat visual cortex through Kunihiko Fukushima’s unsupervised neocognitron to Yann LeCun’s supervised LeNet. It explains convolution, feature maps, pooling, padding, stride, downsampling, and upsampling, then connects these operations to CNN architecture choices such as channel counts, spatial resolution, and final classification with an MLP.
Key takeaways
- Hubel and Wiesel’s simple-cell and complex-cell hierarchy provides a biological analogy for CNN feature detection followed by local response aggregation.
- Fukushima’s neocognitron learned increasingly complex visual patterns without labels, while LeCun added a softmax classifier to train CNNs for supervised digit recognition on MNIST.
- A convolution filter spans every input channel and produces one output feature map; therefore, the number of filters determines the output channel count.
- With stride 1 and no padding, an m×m input and n×n filter produce an output of size m − n + 1; zero padding can preserve spatial dimensions but introduces artificial boundary values.
- Stride-based downsampling reduces spatial computation, while upsampling inserts zeros and generally needs a following convolution to recover a useful dense representation.
- Reducing spatial resolution shrinks the number of values per channel, so CNNs commonly increase channel counts across stages; this is an architectural heuristic, not a universal guarantee of information preservation.
Chapters
- The opening minutes include attendance and classroom conversation before the lecture begins.
- A network issue disables camera control, so remote students are told to follow the slides and listen to verbal descriptions.
- The lecture recaps CNNs as local pattern scanners that share parameters and reduce the number of learned weights.
- It shifts from this computational explanation to the biological history of visual processing and Gestalt perception.
- Figures made from disconnected shapes can appear as a cube, dog, or spiked sphere because perception fills in missing structure.
- The lecture links this idea to inattentional-blindness demonstrations, including the gorilla experiment and searching for red objects before being asked about yellow ones.
- In 1959, David Hubel and Torsten Wiesel studied receptive fields in cats’ striate cortex, corresponding broadly to human visual area V1.
- Their experiments recorded activity from individual neurons while controlled light patterns were moved across the cats’ visual fields.
- Some visual-cortex neurons responded most strongly to oriented slits, such as vertical lines, while surrounding regions could inhibit firing.
- Hubel and Wiesel distinguished simple cells, which detect oriented patterns, from complex cells, which combine responses and are less vulnerable to noise.
- Repeated simple- and complex-cell stages form a hierarchy that builds increasingly complex visual representations.
- Kunihiko Fukushima’s 1980 neocognitron computationalized the simple-cell/complex-cell hierarchy.
- Its repeated modules alternate learned S-cell planes with fixed C-cell planes; C cells combine nearby responses to improve robustness to position shifts.
- Replicated cells within a plane share response patterns while observing different locations, allowing features to be detected across the visual field.
- Fukushima used an unsupervised, winner-selecting learning rule that updates weights associated with the strongest response.
- With training on characters or digits, early units respond to line segments and later units to more complex shapes.
- The final stages can respond to individual digits even when inputs are noisy, without semantic labels.
- A final softmax classifier supplies external supervision, turning learned visual features into labeled predictions.
- Yann LeCun’s model replaces repeated identical S cells with a shared filter scanned across the input, producing a convolution.
- LeCun’s CNN approach achieved strong handwritten-digit recognition on MNIST, a dataset motivated by automating postal-code reading.
- CNN convolutional layers correspond to S-cell planes and produce feature maps, also called activation maps or channels.
- Each filter spans all input maps, computes a weighted sum plus bias at each spatial location, and applies an activation function.
- One filter produces one output map; adding filters increases the number of output maps, with each filter learning its own weights.
- For an input of width m and a filter of width n with stride 1, an unpadded convolution produces width m − n + 1.
- A 5×5 input convolved with a 3×3 filter yields a 3×3 output when no padding is used.
- Zero padding can preserve spatial size, but its artificial boundary values may affect edge responses; odd padding amounts cannot be split perfectly symmetrically.
- Max pooling takes the largest value in a local window—for example, a 2×2 window containing 3, 1, 4, and 6 outputs 6.
- Mean pooling instead averages the values in each window, and pooling is applied independently to each channel.
- Max pooling can retain the winning index as well as the value, which is useful for routing gradients during backpropagation.
- Downsampling by a factor of two keeps every other row and column, reducing each spatial dimension by roughly half.
- Rather than compute every convolution or pooling position and then discard outputs, a stride of two combines the scanning and downsampling steps.
- Pooling is commonly paired with downsampling to avoid redundant nearby outputs.
- Upsampling by a factor of two inserts rows and columns of zeros between existing values to enlarge a map.
- The inserted zeros alone add no information, so upsampling is typically followed by convolution to interpolate and smooth the result.
- The phrase “fractional stride” describes zero insertion followed by convolution; it does not mean a filter literally moves by half a pixel.
- The architecture combines learned convolutions, fixed or alternative pooling operations, and explicit resampling through downsampling or upsampling.
- Pooling functions need not be fixed: a small learned neural network could replace max pooling, though the lecture notes limited practical benefit.
- The choice to downsample trades redundant computation for reduced spatial detail.
- If a stride-2 pooling step reduces an n×n map to roughly n/2×n/2, each channel retains about one quarter as many spatial values.
- The lecture’s dimensionality-counting heuristic says at least four output channels are needed to match the input’s number of scalar values after that reduction.
- Repeating spatial reductions generally motivates increasing the number of feature maps, though a task may deliberately discard irrelevant information.
- A color image enters the network as three channels—red, green, and blue—and each first-layer filter spans all three.
- Common spatial filter sizes include 5×5, 3×3, and 1×1; a 1×1 filter mixes channel values at one location without examining neighboring pixels.
- A convolution with K filters produces K output channels, and its parameter count is the number of filters times filter area times input channels, plus biases.
- Pooling design includes the window size, stride, and pooling operation; max pooling stores the maximum’s location for the backward pass.
- Convolution and pooling progressively build feature maps; sufficiently deep layers can compress each channel to a single value that summarizes the whole image.
- The final maps are flattened and concatenated into a vector for an MLP that performs classification.
- Training uses labeled image–class pairs, a loss function, and backpropagation; shared convolutional parameters require accounting for repeated filter use.
- Architecture choices include the number of layers and filters, filter sizes, pooling windows and strides, and the final MLP’s depth and width.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Carnegie Mellon University Deep Learning.