Save this video — free

Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification

Stanford Online · 1:12:59 · Watch on YouTube

Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification Watch on YouTube →

Overview

This lecture explores the evolution of verification techniques for Large Language Models (LLMs), starting with OpenAI's 2021 paper on training verifiers for math problems using the GSM8K dataset. It progresses through OpenAI's 2023 'Verify, Step-by-Step' paper introducing process-based reward models (PRMs) over outcome-based reward models (ORMs), Stanford's 2024 'Math-Shepherd' which automates PRM label collection, and finally Stanford's recent 'Weaver' that uses an ensemble of weak verifiers with weak-to-strong supervision for robust verification without extensive human annotation.

Key takeaways

Chapters

0:00 Introduction to Verification and the Generation-Verification Gap
1:52 OpenAI's 2021 Paper: Training Verifiers for Math Problems
5:55 Training Methodology for the GSM8K Verifier
8:51 Dual Loss Objective and Architecture
11:47 Training Process and Ablation Studies
18:23 Verification vs. Fine-Tuning Performance
23:39 Test-Time Scaling Benefits and Limitations
35:25 OpenAI's 2023 Paper: Verify, Step-by-Step (Process vs. Outcome)
37:16 Training PRMs with Human Annotations
48:44 PRM Performance vs. ORM and Majority Voting
1:02:28 Math-Shepherd: Automated PRM Labeling
1:12:05 Math-Shepherd Verification Process and Results

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.