Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 10: Reachibility Analysis
Watch on YouTube →
Overview
Stanford Online's Lecture 10 on Optimal and Learning-Based Control introduces continuous-time optimal control, extending dynamic programming to the Hamilton-Jacobi-Bellman (HJB) equation. The lecture covers both standard optimal control and differential games with adversarial disturbances, deriving the HJB equation and its application to Linear Quadratic Regulator (LQR) problems and reachability analysis. Key concepts include value iteration, policy iteration, and the formulation of continuous-time problems using differential equations and integral costs.
Key takeaways
- The Hamilton-Jacobi-Bellman (HJB) equation is the continuous-time analog of the Bellman equation, describing the evolution of the optimal cost-to-go.
- Differential games extend optimal control to adversarial settings, where a disturbance player minimizes the objective, leading to the Hamilton-Jacobi-Isaacs (HJI) equation.
- The LQR problem in continuous time is solved by assuming a quadratic cost-to-go and solving a Riccati differential equation, resulting in linear feedback control.
- Reachability analysis, using differential game formulations, computes sets of states from which a target set is guaranteed to be reached or avoided, crucial for safety-critical planning.
- Non-anticipatory strategies for the adversarial player are key in differential games, granting them a slight advantage that influences the problem's formulation (e.g., minimax structure).
Chapters
- Recap of infinite horizon MDPs and the fixed-point Bellman equation for optimal value functions.
- Discussion of value function vs. Q-function in infinite horizon settings.
- Introduction to solving fixed-point equations via value iteration and policy iteration.
- Value iteration iteratively re-optimizes the Bellman equation starting from an initial guess (e.g., zero vector).
- Guaranteed to converge to the optimal value function V*.
- Intuition based on backward recursion from a very long finite horizon.
- Policy iteration operates in policy space, directly optimizing policies.
- Two-step procedure: policy evaluation (solving for V_pi) and policy improvement (finding a better policy).
- Policy improvement uses a one-step Bellman equation with the evaluated policy's value function as a surrogate.
- Policy iteration converges in a finite number of steps for finite state and control spaces.
- Each step strictly improves the policy's value if it's not already optimal.
- Requires knowledge of the MDP model, including transition probabilities.
- Shift from discrete-time to continuous-time reasoning for dynamics and costs.
- Continuous-time dynamics: x_dot = f(x, u).
- Continuous-time cost: integral of g(x, u, t) dt.
- Derivation of the Hamilton-Jacobi-Bellman (HJB) equation as the continuous-time counterpart to the Bellman equation.
- HJB equation is a partial differential equation (PDE) for the optimal value function J(x, t).
- Introduction of a disturbance term D(t) representing an active adversarial player.
- Player 1 (controller) maximizes J, Player 2 (disturbance) minimizes J.
- Formulation as a minimax problem: find U to maximize J against the worst-case D.
- Illustrates differential games with a pedestrian (Player 1) avoiding a car (Player 2).
- Car is faster but has curvature constraints; pedestrian has lower top speed.
- Focus on reachability analysis: determining initial states for survival.
- Player 2 (disturbance) has non-anticipatory strategies: observes Player 1's history up to current time.
- This gives Player 2 a slight advantage, influencing the minimax formulation.
- Contrast with full information where Player 2 sees Player 1's entire future control sequence.
- Objective: find U to maximize J under the worst-case D.
- The problem is formulated as finding U and D simultaneously.
- Player 2's advantage is captured by the outer minimization in the minimax structure.
- Applying the principle of optimality to the continuous-time adversarial problem.
- Using first-order Taylor approximations for short time intervals.
- The HJI equation emerges, embedding a max-min optimization problem.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.