Save this video — free

Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling

Stanford Online · 1:03:21 · Watch on YouTube

Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling Watch on YouTube →

Overview

This lecture explores test-time compute scaling for LLMs, demonstrating how techniques beyond parameter tuning can significantly enhance model capabilities. The 'Large Language Monkeys' paper introduces repeated sampling with verifiers to improve performance, while scaling laws reveal an exponential power-law relationship between coverage and sample count. The discussion highlights the critical role of automated verification, the 'generation verification gap' in domains lacking verifiers, and introduces sequential revisions and reward models (outcome-based and process-based) as further methods to optimize inference.

Key takeaways

Chapters

0:00 LLM Development Stages: Pre-training, Fine-tuning, and Inference
1:52 The 'Large Language Monkeys' Paper: Repeated Sampling for Improvement
5:36 Repeated Sampling on Agentic Benchmarks like SWE-bench
8:35 Inference Scaling Laws: Coverage vs. Number of Samples
12:13 Justification for Power Law: The Long Tail of Hard Problems
18:40 Shifting Compute Budgets: From Pre-training to Inference
20:20 The Necessity of Automated Verification for Repeated Sampling
25:34 The Generation Verification Gap in Domains Without Verifiers
45:04 Beyond Parallel Sampling: Sequential Revisions and Reward Models
57:20 Difficulty-Based Scaling and Compute Allocation

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.