Save this video — free

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 17: RL Value-Based Methods

Stanford Online · 1:17:40 · Watch on YouTube

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 17: RL Value-Based Methods Watch on YouTube →

Overview

Stanford Online's Lecture 17 on RL Value-Based Methods introduces SARSA and Q-learning as key model-free reinforcement learning algorithms. The lecture details the transition from Monte Carlo to Temporal Difference (TD) learning, highlighting TD's advantages in variance reduction and online updates. It then explores the on-policy (SARSA) versus off-policy (Q-learning) distinction, demonstrating how Q-learning's ability to learn from a different behavior policy enables it to converge to the optimal policy, unlike SARSA which learns the optimal exploratory policy. Finally, the lecture addresses scaling these methods to high-dimensional spaces using value function approximation and introduces Deep Q-Networks (DQN) with experience replay and fixed Q-targets as crucial techniques for stabilizing deep reinforcement learning.

Key takeaways

Chapters

0:00 Introduction to Learning-Based Control and RL Categories
1:48 Prediction vs. Control in RL
3:21 Limitations of Dynamic Programming and Introduction to Model-Free Learning
4:55 Visualizing Rollouts and Value Estimation
7:08 Monte Carlo vs. Temporal Difference Learning
9:55 Extending Value Estimation to Q-Functions
10:31 Generalized Policy Iteration for Control
13:31 Introduction to SARSA and Q-Learning
15:06 Advantages of Temporal Difference Learning
18:47 Deriving SARSA from Monte Carlo Control
23:26 SARSA Algorithm Pseudocode and Epsilon-Greedy Policy
27:34 Windy Grid World Example for SARSA
30:31 Optimal Policy and SARSA Convergence in Windy Grid World
35:13 Monte Carlo Limitations in Non-Terminating Scenarios
39:17 On-Policy vs. Off-Policy Learning Definitions
43:45 Motivation and Formalism for Off-Policy Learning
48:29 Decoupling Behavior and Target Policies
49:10 Modifying SARSA for Off-Policy Learning: Towards Q-Learning
52:00 Q-Learning Algorithm Definition
57:24 Q-Learning Algorithm Pseudocode
1:03:33 Cliff Walking Example: SARSA vs. Q-Learning
1:15:44 Scaling to High-Dimensional State Spaces: Value Function Approximation

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.