Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling
Watch on YouTube →
Overview
This lecture explores test-time compute scaling for LLMs, demonstrating how techniques beyond parameter tuning can significantly enhance model capabilities. The 'Large Language Monkeys' paper introduces repeated sampling with verifiers to improve performance, while scaling laws reveal an exponential power-law relationship between coverage and sample count. The discussion highlights the critical role of automated verification, the 'generation verification gap' in domains lacking verifiers, and introduces sequential revisions and reward models (outcome-based and process-based) as further methods to optimize inference.
Key takeaways
- Repeated sampling with verifiers (the 'Large Language Monkeys' approach) can significantly boost the performance of smaller LLMs, potentially surpassing larger models.
- A power-law relationship governs the trade-off between test-time compute (number of samples) and problem coverage, predictable via scaling laws.
- The 'generation verification gap' highlights the challenge of improving LLM performance in domains lacking automated verification methods.
- Sequential revisions and process reward models (PRMs) offer alternative and complementary strategies to parallel sampling for enhancing inference quality.
- For easier and medium-difficulty problems, investing in test-time compute can yield greater improvements than further pre-training.
- Archon's architecture search framework demonstrates that intelligently combining LLMs and inference techniques (like fusion and layered critiques) can create highly optimized inference pipelines that exceed the performance of single, large models.
Chapters
- LLM development involves pre-training (compute-intensive, trillions of tokens), fine-tuning (less data), and inference (using the model).
- Focus is on improving inference without changing model parameters or fine-tuning.
- Idea: repeatedly ask the same input problem to an LLM (10x, 100x, etc.).
- A verifier selects the correct response from the generated candidates.
- This method can make smaller models (e.g., Llama 3 8B) outperform larger ones (e.g., GPT-4o) on single attempts.
- Repeated sampling shows improved coverage on agentic benchmarks like SWE-bench.
- DeepSeek-V3 can outperform Claude 3.5 after 1,000 samples on coding problems.
- Unit tests can serve as automated verifiers for coding tasks.
- A predictable power-law relationship exists between coverage and the number of parallel samples drawn.
- Formula: Coverage = 1 - (1-p)^k, where p is single-attempt success rate and k is number of samples.
- This behavior holds across various model sizes (70M to 70B parameters) and domains.
- The observed power-law scaling requires a 'long tail' of hard problems.
- Per-problem scaling follows an exponential law, but across a suite of problems, a power law emerges.
- Empirically, datasets exhibit this long tail where simpler problems are solved easily, and harder ones require more samples.
- Previously, most compute was spent on pre-training (hundreds of millions/billions of dollars).
- A new paradigm allows spending more compute on inference to increase model capability.
- Inference compute can be done offline, allowing agents to continuously improve answers.
- Automated verification is crucial to select correct samples from multiple generations.
- Domains like math (formal proofs) and coding (unit tests) have easier verification.
- AI as a compiler (e.g., generating CUDA from PyTorch) is verifiable by comparing outputs.
- In domains lacking verifiers, there's a large gap between 'best of n' methods and true coverage.
- Majority voting plateaus quickly and misses rare correct answers for hard problems.
- Reward models can score responses but still leave a significant gap, especially for complex datasets like MATH.
- Test-time compute scaling can also involve sequential revisions, where a model iteratively improves its own answer.
- Reward models (outcome-based and process-based) score generated responses or steps.
- Process Reward Models (PRMs) score each step of a generated solution, enabling guided search (e.g., beam search).
- Research on math datasets shows difficulty bins (1-5) based on pass-at-one performance.
- For easier problems, test-time compute can be more favorable than scaling pre-training.
- For the hardest problems, larger models with more pre-training still perform better, even with significant test-time compute.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.