Stanford CS329A Self-Improving AI Agents | Part 4 | Learning from Feedback with Tools/Code
Watch on YouTube →
Overview
This lecture explores self-improving AI agents by detailing three papers: ReAct, which combines reasoning and action for grounded task completion; RLEF, a framework for improving code generation using execution feedback; and Constitutional AI, where models critique and revise themselves based on human-defined principles to enhance helpfulness and harmlessness. These methods demonstrate how AI can learn from interaction, feedback, and self-reflection to improve performance in diverse applications from question answering to code generation.
Key takeaways
- ReAct enables LLMs to perform grounded tasks by interleaving reasoning (thought) with tool use (action), improving interpretability and reducing hallucination.
- RLEF significantly boosts code generation performance by using execution feedback (test results) in an iterative reinforcement learning loop, allowing models to learn from errors.
- Constitutional AI leverages AI feedback based on human-defined principles to make LLMs more harmless and helpful, bypassing the need for extensive human preference labeling.
- The combination of reasoning and action (ReAct), execution feedback (RLEF), and AI-driven self-critique (Constitutional AI) are key paradigms for building more capable and reliable AI agents.
- While LLMs can be trained to follow principles, achieving an optimal balance between helpfulness and harmlessness remains a critical research challenge, often involving trade-offs.
Chapters
- LLMs excel at NLP but need interaction with tools/code for real-world tasks.
- Lecture covers ReAct (tool calling), RLEF (code LLMs with execution feedback), and Constitutional AI (self-critique).
- Each technique focuses on different feedback sources for model improvement.
- ReAct models human decision-making: think, then act, then observe, then reason.
- Addresses LLM limitations like hallucination and lack of real-world grounding.
- Combines Chain of Thought (reasoning) with tool interaction (acting).
- ReAct uses prompting to elicit verbal reasoning and subsequent actions.
- Improves performance on knowledge-intensive tasks like HotpotQA and FEVER.
- Provides interpretable traces, allowing humans to trust model responses.
- Valid actions can be framed as a classification task from a predefined set.
- Early implementation used PaLM with a frozen LLM and few-shot examples.
- Model decides whether to reason or take action based on current state.
- Standard prompting fails on HotpotQA question about Apple remote's original device.
- Chain of Thought improves but still lacks correctness.
- ReAct uses interleaved search actions and reasoning to find the correct answer.
- Interleaving thought and action allows models to behave better by incorporating observations.
- Modern models like Qwen often have this capability built-in via distillation.
- Models don't inherently 'know' what they know; they rely on interaction and feedback.
- ReAct's loop is interleaved: thought 1, act 1, thought 2, act 2, informing each step.
- Contradictory search results require user-defined strategies like majority voting.
- ReAct reduces hallucination compared to pure internal state reliance.
- ReAct outperforms action-only and often Chain of Thought on HotpotQA and FEVER.
- Combining ReAct with Chain of Thought/self-consistency yields best results.
- Demonstrates value in combining internal knowledge with external search via reasoning.
- WebShop task: agent purchases products based on user instructions.
- ReAct outperforms imitation learning and imitation learning + RL.
- Still lags behind human experts, with errors cascading in multi-step processes.
- Challenges: large action spaces require many demonstrations, higher inference costs.
- Benefits: better performance on QA/fact-checking/decision tasks, lower hallucination, interpretable traces.
- Handling noisy or misleading feedback requires reflection, backtracking, or confidence metrics.
- RLEF framework uses reinforcement learning with execution feedback for code generation.
- Actions are generated code; observations are test results (pass/fail).
- Iterative refinement loop uses public tests for immediate feedback and private tests for reward signals.
- Two phases: exploitation (inference-time feedback) and policy update (execution results).
- Example: Palindrome substring detection code improves through iterations.
- Separation of public/private tests prevents memorization and guides learning.
- RLEF significantly improves solve rates in competitive programming (CodeContests).
- Learning from errors in early turns leads to fewer wrong outputs and targeted repairs.
- Binary feedback is sufficient for simpler problems; complex ones may need error traces.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.