Save this video — free

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 9: Stochastic Dyn. Program

Stanford Online · 1:17:01 · Watch on YouTube

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 9: Stochastic Dyn. Program Watch on YouTube →

Overview

This lecture introduces Markov Decision Processes (MDPs) as a framework for optimal control under uncertainty, extending dynamic programming to discrete-time stochastic systems. Key concepts include state transitions affected by disturbances (WK), control constraints, and optimizing expected costs. The lecture details the dynamic programming recursion for finite-horizon MDPs, illustrating with an inventory control example, and then discusses infinite-horizon MDPs, stationary policies, and the Bellman equation for value functions and Q-functions, setting the stage for reinforcement learning.

Key takeaways

Chapters

0:05 Introduction to Stochastic Dynamic Programming and MDPs
0:25 Modeling Disturbances and Probability Distributions
0:55 Cost Formulation in Stochastic Settings
2:15 Key Assumptions and Problem Structure
5:43 Principle of Optimality in Stochastic Settings
9:51 Dynamic Programming Algorithm for Finite Horizon MDPs
23:40 Inventory Control Example: Problem Formulation
26:54 Inventory Control Example: Cost and Disturbances
33:20 Inventory Control Example: Solving with Dynamic Programming
40:23 Inventory Control Example: Optimal Policy and Results
45:09 Stochastic Linear Quadratic Regulator (LQR)
48:29 Solving Stochastic LQR with Dynamic Programming
53:29 Stochastic LQR Solution and Implications
56:40 Infinite Horizon Markov Decision Processes (MDPs)
58:51 Infinite Horizon Bellman Equation and Stationary Policies
1:10:00 Q-Functions and Their Role
1:11:40 Q-Functions for Learning and Action Selection
1:13:20 Policy Evaluation for a Given Policy

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.