Save this video — free

Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning

Stanford Online · 1:14:56 · Watch on YouTube

Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning Watch on YouTube →

Overview

This lecture explores advanced AI planning and multi-step reasoning through three papers: LATs, SPRINT, and SWiRL. LATs unifies reasoning, acting, and planning in LLMs using Monte Carlo Tree Search (MCTS) and reflection for improved exploration. SPRINT enables LLMs to identify and execute parallelizable reasoning steps, significantly reducing inference time and improving accuracy. SWiRL focuses on training LLMs for multi-step reasoning and tool use via RL, demonstrating generalization across datasets and tools without direct tool execution during training.

Key takeaways

Chapters

0:00 Introduction to Planning and Multi-Step Reasoning
0:26 Language Agent Tree Search (LATs) Overview
5:39 LATs: Reasoning, Action, and Search Loop
8:53 LATs vs. Math-Shepherd and ReAct
11:57 LATs: Core Intuition and Stages
13:17 LATs: Maze Navigation Example - Selection and Expansion
15:29 LATs: Maze Navigation Example - Execution and Evaluation
18:51 LATs: Simulation and Backpropagation
22:18 LATs: UCT and Backpropagation Details
27:23 LATs: Reflection and Performance
30:19 LATs on WebShop and Summary
38:49 SPRINT: Parallelizing LLM Reasoning
45:30 SPRINT: Synthetic Data Generation for Parallelism
50:04 SPRINT: Fine-tuning and Inference
1:00:30 SPRINT: Training Recipe and Results
1:03:40 SPRINT: Generalization and Task Dependency

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.