Save this video — free

Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks

Stanford Online · 1:15:18 · Watch on YouTube

Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks Watch on YouTube →

Overview

Stanford Online's CS329A lecture Part 8 explores agentic evaluations and long-horizon tasks, introducing three benchmarks: METR for task duration and reliability, GDPVal for economic value against industry experts, and DeepScholar-Bench for research synthesis. The discussion highlights AI's progress in completing longer, more complex tasks, but also identifies significant gaps in reliability, context understanding, and comprehensive knowledge retrieval, particularly for real-world, context-heavy applications.

Key takeaways

Chapters

0:00 Introduction to Agentic Evaluations and Long Horizon Tasks
0:42 Student Examples of Agentic Evaluation Tasks
2:18 METR: Measuring Task Duration and Time Horizon
9:51 METR Task Suites: SWAA, HCAST, and RE-Bench
11:50 Human Baseline and Task Difficulty in METR
20:37 METR Trends: Task Length vs. Model Success Rate
22:26 METR Trends: Time Horizon Improvement Over Time
25:10 Drivers of METR Improvement
33:38 METR: Reliability at 80% Success Rate
37:36 Common Failure Modes in AI Agents
42:02 Challenges and Limitations of METR
45:53 GDPVal: Evaluating Economic Value of AI Tasks
58:58 GDPVal Trends: Model Win Rate vs. Industry Professionals
1:04:10 GDPVal Failure Modes and Context Dependency

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.