Save this video — free

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 19: Model-Based RL

Stanford Online · 1:21:50 · Watch on YouTube

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 19: Model-Based RL Watch on YouTube →

Overview

Stanford Online's Lecture 19 on Model-Based RL covers policy optimization methods like TRPO and PPO, contrasting them with value-based methods. The lecture then delves into model-based RL, outlining a basic recipe involving learning a dynamics model and planning, and highlighting challenges like extrapolation errors. A key focus is uncertainty quantification, exploring how Bayesian approaches, Gaussian Processes, and ensembles can mitigate model overfitting and improve planning by accounting for model confidence.

Key takeaways

Chapters

0:00 Review of Model-Free RL: Policy Optimization
12:22 Limitations and Motivation for Model-Based RL
29:00 Recap of Model-Free RL Paradigms
40:06 The Core RL Skeleton and Trade-offs
43:43 Introduction to Model-Based RL Recipe
49:14 Challenges with Simple Model-Based Approaches
55:07 Heuristic Improvements: Receding Horizon and Refitting
57:20 The Crucial Role of Uncertainty Quantification
1:02:15 Defining and Modeling Uncertainty
1:08:45 Bayesian Approaches for Model Uncertainty

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.