Stanford CS229 Machine Learning | Spring 2026 | Lecture 20: GMM (EM), PCA
Watch on YouTube →
Overview
This lecture delves into advanced reinforcement learning techniques, starting with a detailed explanation of the policy gradient algorithm and its intuitive interpretation. It then explores the Proximal Policy Optimization (PPO) algorithm, highlighting its importance sampling mechanism for reusing trajectories and its clipped surrogate objective function designed to prevent large policy updates. Finally, the lecture discusses the application of these RL concepts to large language models (LLMs) for training "chain-of-thought" reasoning capabilities, framing LLM generation as an MDP with rewards based on final answer correctness.
Key takeaways
- The policy gradient algorithm weights action probabilities by rewards, but its gradient estimation is complex.
- PPO uses importance sampling and a clipped surrogate objective to enable stable, multi-update learning from sampled trajectories.
- Chain-of-Thought (CoT) prompting, especially zero-shot CoT ('Let's think step by step'), significantly improves LLM reasoning.
- Reinforcement learning can train LLMs for CoT by rewarding correct final answers, treating generation as an MDP.
- A key challenge in RL for LLMs is obtaining an initial policy that produces non-zero rewards, often requiring careful initialization or pre-training.
Chapters
- Recap of the policy gradient algorithm and its loss function based on total return.
- The challenge of estimating the gradient when it's embedded within an expectation.
- Introduction to PPO (Proximal Policy Optimization) as an extension of policy gradient.
- PPO's popularity for training large language models and enabling long reasoning capabilities.
- The policy gradient equation maximizes log probability of actions weighted by reward.
- High rewards increase the weighting for a trajectory's log likelihood.
- Low or zero rewards result in zero weighting, meaning those trajectories don't contribute to maximization.
- Demonstration that the expectation of the gradient of the log probability of an action sampled from a policy is zero.
- This is because the gradient of a probability distribution integrated over itself is zero.
- This property is crucial for simplifying policy gradient calculations.
- Decomposing the total return into rewards from the current time step onwards (reward-to-go).
- The expectation of the log gradient multiplied by rewards from past time steps is zero due to the Markov property.
- This simplification leads to weighting the log likelihood by the 'reward-to-go'.
- The ability to add a constant baseline (dependent only on the current state) without changing the expected gradient.
- Baselines are introduced to reduce the variance of the gradient estimator.
- Subtracting a baseline (e.g., average reward-to-go) can center the rewards around zero, improving stability.
- PPO addresses the 'on-policy' nature of standard policy gradient, where samples are used only once.
- It uses importance sampling to reuse trajectories from older policies.
- The goal is to allow multiple gradient updates from the same set of samples.
- The challenge of estimating the gradient of a new policy using samples from an old policy.
- Importance sampling corrects for the distribution mismatch between the old and new policies.
- The ratio of probabilities (pi_theta / pi_theta_old) is used to re-weight the samples.
- PPO introduces a clipped surrogate objective function to limit policy updates.
- The objective clips the probability ratio (r_t) based on an advantage function (A_hat).
- This clipping prevents excessively large policy updates, enhancing stability.
- Case 1: Positive advantage (good action) and high probability ratio (r_t > 1 + epsilon_high) -> no update (clip to 0).
- Case 2: Positive advantage and moderate ratio -> standard gradient update.
- Case 3: Negative advantage (bad action) and low probability ratio (r_t < 1 - epsilon_low) -> no update (clip to 0).
- Case 4: Negative advantage and moderate ratio -> standard gradient update.
- CISPO (Constrained Policy Optimization) modifies PPO's clipping strategy.
- On the positive advantage side, CISPO clips the ratio to a constant (1 + epsilon_high) instead of zeroing it out.
- On the negative advantage side, CISPO retains a linear update for ratios slightly below (1 - epsilon_low), unlike PPO's zeroing.
- Introduction to Chain-of-Thought (CoT) prompting for LLMs.
- CoT involves generating intermediate 'thinking' tokens before the final answer.
- Demonstrates improved performance on complex reasoning tasks compared to direct answering.
- Few-shot CoT: providing examples of step-by-step reasoning in the prompt.
- Zero-shot CoT: adding simple phrases like 'Let's think step by step' to the prompt.
- Zero-shot CoT significantly boosts accuracy without requiring example data.
- Using RL to train LLMs for CoT reasoning, bypassing the need for labeled thinking processes.
- The LLM generation process is framed as an MDP.
- States are the input and previous generations; actions are the next tokens.
- Rewards are assigned based on the correctness of the final answer.
- Reward is typically binary (1 for correct answer, 0 otherwise), requiring ground truth.
- A common baseline strategy involves averaging rewards from multiple sampled trajectories for a given prompt.
- This average reward is subtracted from individual trajectory rewards to create an adjusted advantage.
- Importance of initializing with a policy that yields non-zero rewards (i.e., some correct answers).
- Models like Llama 4 initially lack CoT capabilities, leading to zero rewards.
- Models like Qwen and DeepSeek show initial CoT ability, allowing for bootstrapping via RL.
- Pre-training data can sometimes contain lengthy thinking processes that aid initialization.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.