Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents
Watch on YouTube →
Overview
This lecture explores self-improving AI agents, detailing AlphaCode and AlphaCode 2's approaches to competitive programming problem-solving through massive sampling, filtering, and clustering. It then delves into Search-o1 and Agentic RAG for deep research agents, emphasizing iterative retrieval and analysis to overcome knowledge gaps in large reasoning models, contrasting them with simpler retrieval-augmented generation.
Key takeaways
- AlphaCode and AlphaCode 2 significantly advanced AI in competitive programming by employing massive sampling, filtering, and clustering techniques.
- Search-o1 and Agentic RAG improve deep research agents by enabling iterative retrieval, analysis, and integration of external knowledge to overcome LLM knowledge gaps.
- AlphaCode 2 achieved nearly 2x performance gain over AlphaCode by fine-tuning existing LLMs (Gemini Pro) and using a learned scoring model.
- Models often exhibit overconfidence, with probabilities not accurately reflecting correctness, highlighting the need for calibration research.
- Effective deep research agents require more than simple RAG; they need iterative reasoning, document analysis, and the ability to refine searches across multiple turns.
Chapters
0:05
Introduction to Search-Based Model Improvement
- Focus on improving models using search, differentiating between code model sampling and deep research agents.
- AlphaCode and AlphaCode 2 search patterns will be used in homework assignments.
- The core challenge is curating answers from a model's search space.
1:51
AlphaCode: Solving Competitive Programming Problems
- Addresses complex coding problems beyond simple autocompletion, aiming for end-to-end solutions.
- AlphaCode ranked in the top 54% of participants in 10 contests, demonstrating AI generalization.
- Differs from human evaluation tasks by requiring problem understanding and reasoning, not just code completion.
5:58
AlphaCode's Pre-training and Sampling Pipeline
- Pre-trained models on GitHub data (700GB) and code contests, using masked language model loss and next token prediction.
- Fine-tuning involved regularization (GOLD) to assign higher probability to meaningful patterns.
- Generated 1 million diverse sample programs per question in Python and C++ using high sampling temperature.
12:03
AlphaCode's Filtering, Clustering, and Evaluation
- Filtered samples to those passing provided tests, then clustered them for semantic equivalence.
- Used a separate test input generation model to create new test inputs for unseen problems.
- Submitted a curated subset of solutions to the Codeforces platform for evaluation.
15:17
AlphaCode's Performance and Variance Analysis
- Achieved an average ranking of 54.3 across 10 competitions, competitive with 28% of participants.
- Variance in performance across different contest IDs is attributed to problem input distribution and selection stage bottlenecks.
- Introduced 10@k metric, measuring submissions from a large sample pool, distinct from pass@k (unlimited attempts).
22:21
Impact of Model Size and Sampling Budget
- Larger models (e.g., 41B parameters) consistently outperform smaller ones (e.g., 9B).
- Increasing the sampling budget (from 1k to 1 million samples) improves performance (10@k).
- Clustering further enhances performance by ensuring diverse solutions are submitted.
23:30
Pass@k vs. 10@k and Test-Time Compute
- Pass@k measures the percentage of problems solved with k generated samples (unlimited evaluation attempts).
- 10@k measures performance with a fixed number of submissions (10) from a larger sample pool.
- Solve rate scales log-linearly with the sampling budget, even after selection stages.
37:16
AlphaCode Takeaways and Limitations
- Large-scale sampling, filtering, and clustering yield high coverage and novelty.
- Limitations include loss function as a poor proxy for solve rates and weakness in dynamic programming/constructive algorithms.
- Required a large-scale sampling budget, making it computationally expensive.
40:43
AlphaCode 2: Leveraging Existing LLMs
- Hypothesis: Use existing LLMs (Gemini Pro) instead of pre-training.
- Key changes: fine-tuning Gemini Pro, using multiple AlphaCode 2 models for diverse sampling, and employing a scoring model (reward model) for candidate selection.
45:22
AlphaCode 2's Fine-tuning and Sampling Process
- Fine-tuned Gemini Pro on CodeContests V2 and a high-quality scoring dataset.
- Segmented data to fine-tune multiple models for diverse sample generation.
- Split sampling across models, focusing on C++, and randomized temperature/metadata for diversity.
48:20
AlphaCode 2's Evaluation and Performance Gains
- Filtered out 95% of samples that didn't compile or were incorrect, leaving ~50k samples.
- Aggregated top 10 largest clusters, reranked by a scoring model.
- Achieved a 43% solve rate with 1 million samples, nearly double AlphaCode's 25%.
55:46
Improving Sampling Efficiency and Reasoning
- High cost of sampling (95% wasted) suggests need for better sampling or prompting.
- Self-refinement (like homework 1) or RL could reduce sampling needs.
- Decomposing complex problems and embedding reasoning (e.g., chain-of-thought) are key for improving model capabilities.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.