Save this video — free

CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 6

Carnegie Mellon University Deep Learning · 1:28:09 · Watch on YouTube

CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 6 Watch on YouTube →

Overview

Lecture 6 moves from the chain-rule mechanics of backpropagation to the limits of gradient-based training: a differentiable proxy loss can favor stable solutions over perfect training-set separation, and neural-network loss landscapes can contain saddle points and local minima. Using scalar and multivariable quadratics, the class explains why one global learning rate struggles with different curvatures, then introduces learning-rate schedules, Rprop, momentum, and Nesterov’s accelerated method as ways to improve optimization.

Key takeaways

Chapters

0:00 Attendance and the Lecture’s Training-Algorithm Focus
5:12 Empirical Risk Minimization and Gradient Descent Recap
8:10 Backpropagation as Repeated Chain-Rule Multiplication
16:45 The Weight-Gradient Rule and Backpropagation’s Role
18:08 Proxy Losses, Perceptrons, and Outlier Sensitivity
22:48 Why a Neural Network May Prefer Stability to Perfect Separation
25:16 Saddle Points and Local Minima in Neural-Network Losses
30:51 Convexity and Three Types of Optimization Behavior
33:00 Learning-Rate Bounds for a One-Dimensional Quadratic
40:54 Anisotropic Curvature Makes One Step Size Difficult
49:32 Hessian Conditioning and Curvature Normalization
51:50 Learning-Rate Schedules for Escaping and Converging
59:54 Why Adaptive Per-Parameter Learning Rates Are Appealing
1:05:11 Rprop Uses Gradient Signs to Adapt Each Step
1:14:55 Rprop Caveats and Quickprop’s Curvature Approximation
1:16:14 Momentum Averages Gradients to Reduce Oscillation
1:24:36 Nesterov Acceleration and the Optimization Takeaways

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Carnegie Mellon University Deep Learning.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.