Save this video — free

Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents

Stanford Online · 1:12:27 · Watch on YouTube

Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents Watch on YouTube →

Overview

This lecture explores self-improving AI agents, detailing AlphaCode and AlphaCode 2's approaches to competitive programming problem-solving through massive sampling, filtering, and clustering. It then delves into Search-o1 and Agentic RAG for deep research agents, emphasizing iterative retrieval and analysis to overcome knowledge gaps in large reasoning models, contrasting them with simpler retrieval-augmented generation.

Key takeaways

Chapters

0:05 Introduction to Search-Based Model Improvement
1:51 AlphaCode: Solving Competitive Programming Problems
5:58 AlphaCode's Pre-training and Sampling Pipeline
12:03 AlphaCode's Filtering, Clustering, and Evaluation
15:17 AlphaCode's Performance and Variance Analysis
22:21 Impact of Model Size and Sampling Budget
23:30 Pass@k vs. 10@k and Test-Time Compute
37:16 AlphaCode Takeaways and Limitations
40:43 AlphaCode 2: Leveraging Existing LLMs
45:22 AlphaCode 2's Fine-tuning and Sampling Process
48:20 AlphaCode 2's Evaluation and Performance Gains
55:46 Improving Sampling Efficiency and Reasoning

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.