Carnegie Mellon University 11-785 Introduction to Deep Learning
Professor Bhiksha Raj · Carnegie Mellon University · 10 lectures with notes
Students in this class: ask your lecturer for the class code, and these lectures will already be in your library when you sign up.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 13
State-space RNNs carry history in recurrent hidden states, enabling BPTT to learn from sequence-wide dependencies.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 12
CNNs combine shared filters and backpropagation, while AlexNet’s 2012 ImageNet breakthrough demonstrated their potential at scale.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 11
A worked derivation of CNN backpropagation for convolution filters, input maps, and max- and mean-pooling layers.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 10
CNNs translate hierarchical visual processing into learned convolutions, pooling, and a classifier.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 9
CNNs make pattern detection location-independent by scanning with shared filters and building complex features from local patterns.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 8
How cross-entropy, batch normalization, and regularization make neural-network training more stable and generalizable.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 7
Mini-batches balance update speed and gradient reliability; momentum, RMSProp, and Adam further stabilize noisy training.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 6
Different parameter curvatures make one global learning rate inefficient; adaptive step rules and momentum can improve convergence.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 5
Backpropagation computes neural-network parameter gradients by chaining local derivatives from the output back through a stored forward pass.
CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 14
LSTM gates create an input-controlled memory path that addresses ordinary RNNs’ short memory and vanishing-gradient problems.