Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 19: Model-Based RL
Watch on YouTube →
Overview
Stanford Online's Lecture 19 on Model-Based RL covers policy optimization methods like TRPO and PPO, contrasting them with value-based methods. The lecture then delves into model-based RL, outlining a basic recipe involving learning a dynamics model and planning, and highlighting challenges like extrapolation errors. A key focus is uncertainty quantification, exploring how Bayesian approaches, Gaussian Processes, and ensembles can mitigate model overfitting and improve planning by accounting for model confidence.
Key takeaways
- Policy optimization algorithms like TRPO and PPO aim to optimize policy gradients using surrogate objectives, with PPO offering a simpler clipped objective.
- Model-based RL faces challenges in learning complex dynamics and extrapolating, which can be addressed by uncertainty quantification.
- Uncertainty quantification, using Bayesian methods, Gaussian Processes, or ensembles, helps models avoid overfitting and guides planning by considering model confidence.
- Ensembles of neural networks can approximate posterior distributions by converging to different modes, capturing model uncertainty.
- Model-based RL, exemplified by PETS, often achieves superior sample efficiency compared to model-free methods like PPO and SAC.
- Autonomous systems employ hierarchical decision-making, combining high-level goal setting (DP), trajectory generation (open-loop), and low-level tracking (MPC, PID), with learning increasingly integrated across these layers.
Chapters
- Model-free RL policy optimization methods like REINFORCE and Actor-Critic optimize policy gradients.
- Policy optimization algorithms (TRPO, PPO) aim to optimize a surrogate objective function.
- TRPO uses a trust region constraint (KL divergence) to limit policy updates.
- PPO simplifies TRPO by using a clipped surrogate objective or penalty-based approach.
- TRPO's complexity with conjugate gradient methods and weak performance with deep networks led to PPO.
- PPO aims to achieve TRPO's objective without second-order optimization.
- PPO's popular clipped objective function constrains the policy ratio (r) between 1-epsilon and 1+epsilon.
- Value-based methods use Generalized Policy Iteration (policy evaluation + improvement).
- Policy optimization methods directly learn a parametric policy (pi_theta) via gradient descent.
- Policy gradients can suffer from high variance, addressed by baselines and critics.
- The general RL skeleton involves generating samples, estimating returns/values, and improving the policy.
- Trade-offs exist in sample efficiency, stability, action spaces (continuous/discrete), and horizon (episodic/infinite).
- On-policy methods (e.g., policy gradient) are less sample efficient than off-policy methods (e.g., Q-learning).
- Model-based RL requires learning a model of the environment's dynamics (state, action -> next state).
- A basic recipe: collect transitions, fit a dynamics model (e.g., regression), and plan using the model.
- Simple model fitting works for linear dynamics but struggles with complex, non-linear systems.
- Extrapolating from observed data to out-of-distribution states is difficult and can be misleading.
- Learned models can exploit errors in the positive direction if not carefully handled.
- Model Predictive Control (MPC) uses a receding horizon to replan periodically.
- Continuously refitting the model with new data from deployed policies can close the distribution gap.
- High-capacity models can overfit limited data, leading to exploitable errors during optimization.
- Uncertainty quantification aims to measure how confident the model is in its predictions.
- This helps prevent optimizers from exploiting model inaccuracies.
- Uncertainty can be expressed as a distribution over possible outcomes, not just a point estimate.
- Distinguishes between aleatoric uncertainty (inherent noise) and epistemic uncertainty (model uncertainty).
- Bayesian methods model uncertainty as a distribution over model parameters (theta) given data.
- This allows for calculating the predictive posterior distribution by averaging over parameter distributions.
- As more data is collected, the posterior distribution becomes narrower and more accurate.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.