Save this video — free

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL

Stanford Online · 1:12:39 · Watch on YouTube

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL Watch on YouTube →

Overview

This lecture explores train-time scaling techniques to enhance AI model reasoning capabilities, focusing on three papers: STaR, DeepSeekMath, and DAPO. These methods aim to improve model performance on complex reasoning tasks like the AIME benchmark by leveraging self-generated data and reinforcement learning, demonstrating that increased training compute can compensate for fewer model parameters. Key insights include the effectiveness of rationalization, the challenges of RL implementation, and the importance of verifiability in domains like mathematics and coding.

Key takeaways

Chapters

0:00 Introduction to Train Time Scaling and Key Papers
5:07 The Train Time Scaling Loop and Compute Trade-offs
10:13 Comparing Train Time vs. Test Time Compute Scaling
12:00 The Role of Reasoning in AI Model Capabilities
27:04 STaR: Boosting Reasoning with Rationales
35:13 STaR's Iterative Process and Assumptions
41:53 STaR Algorithm Details and Data Sets
47:10 STaR Results and Limitations
1:00:53 V-STaR, Quiet-STaR, and STaR's Bounds
1:07:33 DeepSeekMath: Scaling RL for Mathematical Reasoning
1:12:30 GRPO: A Memory-Efficient RL Algorithm

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.