The End of Coding: Andrej Karpathy on Agents, AutoResearch, and the Loopy Era of AI
Watch on YouTube →
Overview
Andrej Karpathy discusses the shift in software development towards AI agents and "claws" that automate coding and research, highlighting the need to optimize token throughput and delegate tasks to multiple agents. He shares his experiences using AI to automate his home and explores the potential of auto-research to accelerate scientific discovery by removing humans from the loop.
Key takeaways
- The shift towards AI agents is fundamentally changing software development, with developers delegating a significant portion of coding tasks to AI.
- Maximizing token throughput is crucial in the age of AI agents, similar to maximizing GPU utilization in traditional machine learning.
- "Claws" represent a new level of AI autonomy, operating persistently and proactively on behalf of users with sophisticated memory systems.
- Auto-research has the potential to accelerate scientific discovery by removing humans from the loop and automating experimentation.
- The current generation of LLMs exhibits "jaggedness," with strong performance in verifiable tasks but limitations in areas requiring nuance and intent.
- Collaboration with untrusted workers can be enabled in auto-research by having a trusted pool of workers to verify the results.
Chapters
0:00
The Rise of AI Agents and the "AI Psychosis"
- Karpathy describes feeling in a constant state of "AI psychosis" due to the rapid advancements in AI capabilities.
- He delegates 80% or more of his coding tasks to AI agents, a significant shift from his previous workflow.
- Karpathy emphasizes that the default workflow of software engineers has fundamentally changed since December due to AI agents.
0:33
Skill Issues and Maximizing Token Throughput
- Karpathy believes limitations in AI projects often stem from skill issues in prompting and utilizing available tools.
- He aims to maximize token throughput, treating unused subscription tokens as a sign of inefficiency.
- Karpathy draws a parallel between maximizing GPU utilization during his PhD and maximizing token throughput now.
0:37
The Future of Coding: Multiple Agents and Looping "Claws"
- Karpathy envisions a future where multiple AI agents collaborate in teams, with "claws" providing persistent, autonomous functionality.
- He defines a "claw" as an entity that operates independently with sophisticated memory systems, looping and acting on the user's behalf without constant supervision.
- Karpathy praises Peter Steinberg's Open Claw for its compelling personality, sophisticated memory system, and WhatsApp portal.
15:16
Automating Home Management with "Dobby" the Elf Claw
- Karpathy created a "claw" named Dobby to automate his home, controlling lights, HVAC, shades, pool, spa, and security system.
- Dobby uses AI to identify smart home subsystems, reverse engineer APIs, and create a centralized dashboard.
- Dobby sends WhatsApp notifications with images when a FedEx truck arrives, demonstrating advanced home automation.
19:04
The Overproduction of Bespoke Apps and the Agentic Web
- Karpathy argues that many custom apps are unnecessary, suggesting that agents should directly use exposed API endpoints.
- He envisions an "agentic web" where agents act on behalf of humans, requiring a reconfiguration of the industry.
- Karpathy believes that the current vibe coding will become trivial in a year or two, with AI easily translating human intent.
25:26
Distractions and Security Concerns Limiting Claw Development
- Karpathy admits to being distracted from fully developing his "claws" due to other projects.
- He expresses security and privacy concerns about granting AI full access to his digital life, limiting its capabilities.
- Karpathy mentions Jensen's tools being too busy as another limitation.
27:01
Auto Research: Removing Humans from the Research Loop
- Karpathy defines auto-research as a way to maximize token throughput by removing himself as the bottleneck in the research process.
- He aims to create autonomous systems that run for extended periods without his involvement, focusing on leverage and minimal input.
- Auto-research involves setting an objective, defining metrics and boundaries, and letting the system run independently.
29:02
Recursive Self-Improvement and Frontier Labs
- Karpathy's interest in training GPT-2 models is a playground for exploring recursive self-improvement of LLMs.
- He was surprised when auto-research found hyperparameter tunings for his model that he had missed.
- Karpathy believes frontier labs are focused on recursively self-improving LLMs, experimenting on smaller models and extrapolating results.
32:23
Automated Scientists and the Refactoring of Research Abstractions
- Karpathy advocates for removing researchers from the loop and automating as much of the research process as possible.
- He suggests a system with a queue of ideas, automated scientists, and workers pulling items to try out.
- Karpathy emphasizes the need to rethink abstractions and reshuffle processes to maximize token throughput.
34:13
Program MDs: Code for Research Organizations
- Karpathy proposes that research organizations can be described by "program MDs" – markdown files outlining roles and connections.
- He suggests tuning the code of research organizations to optimize for risk-taking, stand-up frequency, and other factors.
- Karpathy discusses the idea of a contest where people write different program MDs to see which yields the most improvement on the same hardware.
38:35
The Loop of Metrics and Autonomous Agents
- Karpathy emphasizes the importance of objective metrics for effective auto-research.
- He notes that the current LLM ecosystem is well-suited for tasks with easy-to-evaluate metrics, such as writing efficient CUDA kernels.
- Karpathy cautions that the models are still rough around the edges and that pushing too far ahead can be counterproductive.
42:12
Jaggedness and the Limitations of Generalization
- Karpathy suggests that LLMs are trained via reinforcement learning, struggling with nuance and intent.
- He uses the example of LLMs consistently providing the same old jokes, even as their other capabilities improve.
- Karpathy argues that the joke situation suggests that improvements in code generation do not necessarily translate to broader intelligence.
48:15
Speciation of Intelligences and the Capacity Constraint
- Karpathy suggests that the labs are trying to have a single model that is arbitrarily intelligent in all these different domains and they just stuff it into the parameters.
- Karpathy predicts more speciation in intelligences, with smaller models specializing in specific tasks.
- He questions whether the capacity constraint on available compute infrastructure drives more speciation.
54:59
Open Ground: Collaboration and Untrusted Workers in Auto Research
- Karpathy discusses the need for more collaboration in research, particularly with untrusted workers.
- He envisions a system where anyone can contribute code, with a trusted pool of workers verifying the results.
- Karpathy draws an analogy to blockchain, with commits building on each other and proof of work being experimentation.
- He suggests that a swarm of agents on the internet could collaborate to improve LLMs and potentially outperform frontier labs.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, No Priors: AI, Machine Learning, Tech, & Startups.