Save this video — free

Stanford CS229 Machine Learning | Spring 2026 | Lecture 18: GMM (EM), PCA

Stanford Online · 1:16:25 · Watch on YouTube

Stanford CS229 Machine Learning | Spring 2026 | Lecture 18: GMM (EM), PCA Watch on YouTube →

Overview

This lecture introduces reinforcement learning (RL) as sequential decision-making without direct supervision, relying instead on rewards. It defines the Markov Decision Process (MDP) framework, comprising states (S), actions (A), transition dynamics (P), rewards (R), and discount factor (gamma), using a 1D robot navigation example. The core goal is to find an optimal policy (pi*) that maximizes expected cumulative reward, often by solving the Bellman equation. The lecture then details the Policy Gradient (REINFORCE) algorithm, a method for optimizing parameterized stochastic policies (pi_theta) by estimating the gradient of the expected return using sampled trajectories.

Key takeaways

Chapters

0:00 Introduction to Reinforcement Learning and its Applications
1:59 Defining Sequential Decision Making
5:34 Exploration vs. Exploitation Trade-off
7:23 RL's Lack of Supervision and Reliance on Rewards
10:50 The Markov Decision Process (MDP) Framework
13:18 MDP Component: States (S)
18:35 MDP Component: Actions (A)
19:58 MDP Component: Transition Dynamics (P)
30:12 MDP Component: Reward Function (R)
41:49 Defining Return/Payoff for a Trajectory
48:59 The Complete MDP Tuple
53:48 The Markov Property and Policies
1:06:43 Value Functions: V(S) and V*(S)
1:12:23 The Bellman Equation

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.