Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks
Watch on YouTube →
Overview
Stanford Online's CS329A lecture Part 8 explores agentic evaluations and long-horizon tasks, introducing three benchmarks: METR for task duration and reliability, GDPVal for economic value against industry experts, and DeepScholar-Bench for research synthesis. The discussion highlights AI's progress in completing longer, more complex tasks, but also identifies significant gaps in reliability, context understanding, and comprehensive knowledge retrieval, particularly for real-world, context-heavy applications.
Key takeaways
- AI models show exponential growth in task completion time horizon (METR), doubling every ~7 months at 50% reliability, but this linear trend is contradicted by GDPVal's comparison to human experts.
- GDPVal reveals AI win rates against industry experts are modest (e.g., ~48% for Claude Opus 4.1), with models excelling at shorter, digital tasks but struggling with context-heavy, real-world professional work.
- DeepScholar-Bench highlights AI's difficulty in performing deep research synthesis, specifically in finding comprehensive sources, extracting key facts, and ensuring verifiable citations, with current systems scoring below 19%.
- Key limitations for AI agents include low reliability (needing >95% success), poor handling of ambiguous prompts, and a lack of deep contextual understanding, especially for tasks not well-represented in training data.
- While AI shows promise in areas like software engineering and ML research, significant headroom exists for improving performance in complex, context-dependent, and highly reliable real-world applications.
Chapters
- Lecture focuses on agentic evaluations and long-horizon tasks.
- Traditional benchmarks (chatbots, QA) are saturating.
- New metrics measure task duration, complexity, and economic value.
- A private Wiki QA task for a coding agent.
- Multi-hop question answering and knowledge tasks.
- Literature review section generation for academic papers.
- Focuses on the duration of tasks AI models can complete.
- Tasks range from 1-30 seconds (SWAA) to 30 hours (HCAST) and 8 hours (RE-Bench).
- Human professionals' completion times serve as a calibration anchor.
- SWAA: Atomic actions (1-30 seconds).
- HCAST: Software/research engineering tasks (1 minute to 30 hours).
- RE-Bench: Full ML research tasks (up to 8 hours).
- Skilled professionals with ~5 years of experience record completion times.
- Geometric mean used for task difficulty rating.
- Challenge: Experienced professionals may underestimate task difficulty for models.
- SWAA tasks (short) show higher model success rates.
- HCAST tasks (diverse) have variable success rates.
- RE-Bench tasks (long, complex) show lower success rates.
- GPT-2 (2019): ~2-second task completion horizon.
- GPT-4 (2023): ~few minutes.
- Claude 3.7 Sonnet (2025 forecast): ~59 minutes at 50% success rate.
- Improved logical reasoning capabilities.
- Enhanced code generation.
- Better tool usage and error recovery mechanisms.
- 80% success rate significantly reduces task completion horizon compared to 50%.
- Claude 3.7 Sonnet achieves ~15 minutes at 80% success, vs. 59 minutes at 50%.
- Highlights headroom for improving model reliability.
- Poor planning or task decomposition.
- Incorrect tool choice or mental math/reasoning.
- Premature task abandonment or repetitive loops.
- Lower performance on messy tasks with no single correct answer.
- SWE-bench estimates may be biased by model exposure to GitHub repos.
- Model performance resembles contractors due to lack of context on unseen codebases.
- Focuses on real-world, economically valuable tasks.
- Compares model output win rate against industry experts with >10 years experience.
- Covers diverse sectors: manufacturing, finance, healthcare, retail, etc.
- Linear improvement trend observed, unlike METR's exponential trend.
- GPT-4o win rate ~12-30%, Claude Opus 4.1 ~47.6% (2024-2025).
- Models perform better on shorter tasks and digital/text-based tasks.
- Common failures: instruction following, not using reference data, formatting errors.
- Models struggle when context is underspecified in prompts.
- Real work is context-heavy; models act more like contractors without domain knowledge.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.