Stanford CS329A Self-Improving AI Agents | Part 9 | Future Research Areas
Watch on YouTube →
Overview
Stanford Online's CS329A course concludes by exploring future research directions in self-improving AI agents, focusing on enhancing diversity in reasoning chains (Multi-Agent Fine-tuning), improving verification robustness (DeepSeekMath-V2's self-verification), and breaking data barriers through self-generated tasks (coding agents proposing abduction, deduction, induction). The discussion also highlights the growing importance of efficiency, with a focus on 'intelligence per watt' and the potential for local inference to redistribute demand from cloud-based models.
Key takeaways
- Multi-agent fine-tuning using specialized generation and critic agents significantly improves reasoning diversity and sustains performance gains over iterations, unlike single-agent methods that plateau.
- DeepSeekMath-V2 introduces self-verification for theorem proving by training meta-verifiers to identify proof issues without reference solutions, breaking the verification bottleneck.
- AI models can self-propose tasks (abduction, deduction, induction) for training, achieving state-of-the-art results in coding benchmarks without human-curated data, demonstrating curriculum learning.
- Local inference models are rapidly improving in accuracy (3.1x since 2023) and can handle a majority of common user queries, suggesting a shift from cloud-only AI.
- The 'intelligence per watt' metric is critical for future AI, highlighting the need for energy-efficient models and hardware, especially for local accelerators.
- Key future research areas include understanding synthetic data flywheels, developing robust continual learning mechanisms, and optimizing infrastructure for high-throughput, low-latency test-time scaling.
Chapters
- Recap of topics covered: LLM scaling, evolutionary strategies, tool-use for end-to-end workflows, retrieval and memory, planning, multi-step reasoning.
- Definition of an agent as a generalization of LLMs: goal-oriented, environment interaction, feedback collection, self-correction.
- Current paradigm involves orchestrating LLMs and tools, with LLMs as judges for verifiers, tool calls, and search algorithms.
- Focus on self-improvement through planning, multi-step reasoning, and self-correction.
- Identified research areas: self-improvement, efficiency, generalization across domains.
- Challenge: Current self-improvement loops are often limited to narrow domains like math and coding.
- Problem: Single LLM-generated data lacks diversity, leading to performance plateaus after few iterations.
- Solution: Multi-agent fine-tuning using specialized generation and critic agents to increase diversity.
- Paper proposes using multiple specialized agents (generation and critic) to produce diverse initial solutions.
- Generation agents create diverse answers; critic agents evaluate and refine.
- Iterative process involves debate rounds, summarization of agent responses, and critique.
- Enables diversity at the generation stage, leading to improved performance and sustained diversity over iterations.
- N generation models produce outputs, which are summarized.
- A critique process is applied, and the critique is added to the input for the next round.
- Majority voting on updated answers and subsequent summarization continues the loop.
- Generation models are fine-tuned from the same base model to produce good answers, filtering for majority vote matches.
- Multi-agent fine-tuning shows sustained performance increases over iterations, unlike single-agent fine-tuning which plateaus.
- Llama models show responsiveness to this technique.
- Responses remain diverse over multiple fine-tuning iterations, even on math datasets.
- Generalizes beyond in-domain data, showing higher performance on adjacent domains like GSM8K.
- Challenge: Verifying model outputs, especially reasoning chains, is difficult.
- Traditional RL relies on final answer matching ground truth, which can miss incorrect reasoning.
- DeepSeekMath-V2 addresses theorem proving by enabling self-verification loops.
- LLM-as-judge fails for complex proofs; human experts identify subtle issues.
- Proposes training verifiers/meta-verifiers to identify issues in proofs without reference solutions.
- Architecture includes a generator, a verifier, and a meta-verifier.
- Meta-verifier reviews verifier's analysis for correctness and identifies issues.
- Verifier is trained to score proofs (0.5-1), and the generator produces harder proofs to improve the verifier.
- Iterative optimization loop shows proof scores climbing over iterations on IMO and CNML problems.
- Achieves ~42% proof score on IMO shortlist 2024 with best-of-32 proofs.
- Demonstrates that identifying issues in reasoning chains can push models towards correction.
- Self-verification breaks the verification bottleneck by automating issue identification and improvement.
- Problem: Current AI training relies on human-curated data, which becomes a bottleneck as models surpass human intelligence.
- Paper proposes models that can propose tasks and then solve them, reducing reliance on external data.
- Focus on coding domain, with tasks categorized into abduction, deduction, and induction.
- Proposer generates tasks, solver attempts them, and rewards are based on task difficulty and solver success rate.
- Proposer generates tasks based on coding paradigms (deduction, abduction, induction).
- Tasks are selected based on optimizing for difficulty: not trivial, not impossible.
- Proposer is conditioned on past generated examples to promote diversity.
- Proposed tasks are validated through program integrity checks, safety checks, and output consistency.
- This approach leads to curriculum learning, where task difficulty evolves over time.
- Achieves state-of-the-art on coding benchmarks without human-curated prompt data.
- Outperforms models trained on tens of thousands of expert examples.
- Emergent behaviors include increasing complexity metrics and improving diversity of programs and answers.
- Key takeaways: need for diversity in reasoning chains, better verification, and breaking data barriers.
- Synthetic data generation by models can improve performance and transferability across tasks.
- Larger models benefit more from this 'data flywheel' and generalization.
- Research question: how far can self-improvement push performance in verifiable domains to reduce reliance on labeled data in non-verifiable domains?
- Domains with slow verification (scientific discovery, chip design, chemical experiments) are challenging for RL.
- Subjective domains like creative writing also pose verification difficulties.
- Approaches for slow verification include training separate reward models to predict simulation outcomes.
- Efficiency is crucial; 'intelligence per watt' is a key metric for future AI workloads.
- Current AI inference is cloud-based ('mainframe era'), leading to exploding compute demand and energy consumption.
- Observation: 77% of ChatGPT requests are for practical guidance, information, or writing, addressable by smaller, local models.
- Significant improvements in local inference accelerators (e.g., 126x GPU memory improvement since 2012).
- New metric 'intelligence per watt' combines task accuracy with power efficiency.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.