Save this video — free

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 18: RL Policy Optimization

Stanford Online · 1:20:18 · Watch on YouTube

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 18: RL Policy Optimization Watch on YouTube →

Overview

This lecture introduces policy optimization as a model-free reinforcement learning method, contrasting it with value-based approaches. It details the derivation of policy gradients, leading to the REINFORCE algorithm, and discusses variance reduction techniques like baselines. The session culminates in an explanation of actor-critic methods, which combine policy gradient and value-based learning to improve sample efficiency and stability, exemplified by the AlphaGo system.

Key takeaways

Chapters

0:06 Introduction to Policy Optimization in RL
0:21 Reinforcement Learning Formalism: MDPs and Trajectories
0:47 Parametric Policy Representation
10:14 Gradient Ascent for Policy Optimization
12:06 Estimating the Objective Expectation via Sampling
15:18 Deriving the Policy Gradient
20:02 Log-Derivative Trick for Policy Gradient
23:23 Decomposing the Log-Probability of Trajectories
27:08 Actionable Policy Gradient Expression
30:26 The REINFORCE Algorithm
34:12 Intuition Behind Policy Gradients
45:53 High Variance in Policy Gradient Estimators
54:13 Variance Reduction: Causality and Baselines
1:08:26 Baselines and Average Return
1:12:03 On-Policy Nature of Policy Gradients
1:15:35 Pros and Cons of Policy Gradient Methods

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.