Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 17: RL Value-Based Methods
Watch on YouTube →
Overview
Stanford Online's Lecture 17 on RL Value-Based Methods introduces SARSA and Q-learning as key model-free reinforcement learning algorithms. The lecture details the transition from Monte Carlo to Temporal Difference (TD) learning, highlighting TD's advantages in variance reduction and online updates. It then explores the on-policy (SARSA) versus off-policy (Q-learning) distinction, demonstrating how Q-learning's ability to learn from a different behavior policy enables it to converge to the optimal policy, unlike SARSA which learns the optimal exploratory policy. Finally, the lecture addresses scaling these methods to high-dimensional spaces using value function approximation and introduces Deep Q-Networks (DQN) with experience replay and fixed Q-targets as crucial techniques for stabilizing deep reinforcement learning.
Key takeaways
- Temporal Difference (TD) learning offers lower variance and online update capabilities compared to Monte Carlo methods, making it more suitable for continuous or non-terminating tasks.
- Q-learning, an off-policy algorithm, decouples the behavior policy (used for exploration) from the target policy (the optimal policy being learned), enabling it to learn the optimal greedy policy.
- SARSA, an on-policy algorithm, learns the optimal epsilon-greedy policy, which is often safer during exploration but may not be the absolute optimal policy.
- Value Function Approximation (VFA) using parametric models like neural networks is essential for scaling RL to high-dimensional state spaces, reducing memory and enabling generalization.
- Deep Q-Networks (DQN) leverage experience replay (to decorrelate samples) and fixed Q-targets (using a separate, delayed target network) to stabilize learning with deep neural networks.
Chapters
- Recap of imitation learning and introduction to reinforcement learning (RL) through trial and error.
- Focus on value-based methods within model-free RL for the remainder of the course.
- References include Sutton & Barto's textbook and David Silver's RL course.
- Prediction: estimating the value of a given policy.
- Control: learning the policy that optimizes performance.
- Previous discussion covered dynamic programming (policy iteration, value iteration) for MDPs.
- Dynamic programming requires perfect knowledge of system dynamics (transition function T).
- Reinforcement learning learns without a dynamics model through interaction.
- Key paradigms: Monte Carlo learning and Temporal Distance (TD) learning.
- White nodes represent states, black nodes represent actions in environment rollouts.
- Goal is to estimate the expected value over all possible leaf nodes in rollouts.
- Dynamic programming uses one-step calculations based on reward and discounted future value.
- Monte Carlo updates use the full return (GT) from a rollout until a terminal state.
- TD learning uses a one-step estimate (reward + discounted future value) as the target.
- TD learning bootstraps (updates towards another estimate), while Monte Carlo does not.
- Concepts for state-value functions (V) apply equally to state-action value functions (Q).
- Monte Carlo requires full episode completion.
- TD learning allows online updates and works with incomplete sequences.
- Alternates between policy evaluation (estimating value function for a policy) and policy improvement (deriving a better policy from the value function).
- Aims to converge to the optimal value function and policy.
- Monte Carlo Control is a V0 RL algorithm alternating policy evaluation (MC Q-function estimate) and improvement (epsilon-greedy).
- Focus on SARSA and Q-learning, value-based methods.
- Deep dive into on-policy vs. off-policy algorithms.
- Introduction to value function approximation for scaling.
- TD learning provides a lower variance estimate compared to Monte Carlo.
- TD samples only once per step, reducing stochasticity compared to full rollouts.
- TD's one-step updates enable online learning and handling of incomplete sequences.
- Replacing Monte Carlo Q-function evaluation with TD updates in Monte Carlo Control leads to SARSA.
- SARSA update requires a quintuple of events: (state, action, reward, next state, next action).
- The name SARSA comes from these five elements: S, A, R, S', A'.
- Initialize Q-function (e.g., to zeros).
- In state Xt, choose action Ut using an epsilon-greedy policy (ε random, 1-ε argmax over Q).
- Observe reward Rt and next state Xt+1, then choose next action Ut+1 using epsilon-greedy policy.
- Update Q(Xt, Ut) using the TD backup based on Rt, Xt+1, and Ut+1.
- Grid world with a 'wind' that pushes the agent upwards.
- Goal: reach the goal state in minimum time (negative reward per step).
- Hyperparameters: epsilon (exploration), alpha (step size), gamma (discount factor = 1).
- Optimal policy involves moving to the ceiling to avoid wind and then proceeding to the goal.
- SARSA converges to the optimal policy.
- Training dynamics show initial random performance improving over time as TD backups propagate positive rewards.
- Monte Carlo methods require terminating episodes.
- In scenarios like the windy grid world, if the agent gets stuck in a loop, Monte Carlo may not update values.
- TD learning's online updates allow it to handle such situations by accumulating negative rewards and learning to avoid them.
- On-policy: evaluating/improving the policy currently used for decision-making.
- Off-policy: evaluating/improving a target policy different from the behavior policy generating data.
- SARSA is on-policy because it improves the same epsilon-greedy policy used for interaction.
- Off-policy learning evaluates a target policy (π) using data from a behavior policy (μ).
- Useful when direct interaction with the target policy is impossible (e.g., robotics, learning from demonstrations).
- Allows reusing past data generated by older policies, avoiding waste.
- Off-policy methods decouple the exploratory behavior policy from the target exploiting policy.
- This enables learning the optimal policy even if the behavior policy is suboptimal or random.
- Key reason for the importance of off-policy learning.
- In SARSA, the next action Ut+1 is sampled from the behavior policy μ.
- For off-policy, use a different action Ut+1' sampled from the target policy π.
- This leads to updating the Q-function towards the value of the alternative action Ut+1'.
- Target policy π is the greedy policy with respect to Q (argmax over actions).
- Behavior policy μ is the epsilon-greedy policy.
- The TD target uses the reward plus the discounted maximum Q-value over possible next actions: R + γ * max_u Q(S', u).
- Initialize Q-function.
- In state Xt, choose action Ut using epsilon-greedy policy (behavior policy μ).
- Observe reward Rt and next state Xt+1.
- Update Q(Xt, Ut) using the Q-learning TD target: Rt + γ * max_u Q(Xt+1, u).
- Grid world with a 'cliff' region yielding a large negative reward (-100).
- Goal: reach the goal state with maximum reward.
- SARSA (on-policy) converges to the optimal epsilon-greedy policy, which is safer to avoid the cliff during exploration.
- Q-learning (off-policy) converges to the optimal greedy policy, but its training performance is lower due to occasional cliff falls during exploration.
- Tabular methods (lookup tables) are infeasible for large state spaces (e.g., 10^170 in Go).
- Value function approximation uses parametric functions (e.g., neural networks) to represent V(s) or Q(s, a).
- Benefits: reduced memory cost and generalization across states.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.