How I use LLMs
Watch on YouTube →
Overview
Andrej Karpathy provides a practical guide to using Large Language Models (LLMs), focusing on ChatGPT and its ecosystem. He details core concepts like tokenization, context windows, and model training stages (pre-training and post-training). Karpathy demonstrates various LLM applications, including text generation, tool use (internet search, Python interpreter), advanced data analysis, multimodal capabilities (audio, image, video), and quality-of-life features like memory and custom instructions, emphasizing the importance of choosing the right model and tier for specific tasks.
Key takeaways
- LLMs process information as token streams, enabling multimodality (text, audio, image, video) by representing each modality as tokens.
- Thinking models, trained with reinforcement learning, exhibit improved reasoning and accuracy on complex math and code problems.
- Tool use, such as internet search and Python interpreters, significantly expands LLM capabilities beyond their static training data.
- Advanced features like Claude Artifacts and Cursor demonstrate LLMs' potential to generate functional applications and assist in professional software development.
- Multimodal LLMs can natively process audio and video, moving beyond text-based intermediaries for more natural human-computer interaction.
- Customization features like ChatGPT's memory, custom instructions, and Custom GPTs allow users to personalize LLM behavior and save prompting effort for recurring tasks.
Chapters
- ChatGPT, launched by OpenAI in 2022, popularized text-based LLM interaction.
- The LLM ecosystem now includes competitors like Google's Gemini, Meta's Llama, Anthropic's Claude, and xAI's Grok.
- Leaderboards like Chatbot Arena and Scale AI's leaderboard help track model performance.
- LLMs process text by breaking it down into tokens, which are small chunks of text.
- The Tiktokenizer tool visualizes how text is converted into token sequences and their IDs.
- Conversations are managed as a sequence of tokens within a context window, acting as the model's working memory.
- Pre-training compresses the internet into a model's parameters (e.g., 1 trillion parameters for a 1TB zip file).
- This stage imbues the model with world knowledge but results in a knowledge cutoff date.
- Post-training refines the model's persona, often as an assistant, through human-curated conversation datasets.
- LLMs are suitable for common knowledge queries (e.g., caffeine in an Americano) that are frequently mentioned online.
- Users should verify information, as LLMs provide probabilistic recollections, not guaranteed facts.
- For recent information or niche topics, LLMs may lack knowledge due to their training cutoff.
- Starting a new chat resets the token context window, improving model focus and reducing costs.
- Overloading the context window can distract the model and decrease performance accuracy.
- Keeping the context window concise and relevant is crucial for faster and better results.
- Different LLM providers offer various pricing tiers with access to different model sizes (e.g., GPT-40 mini vs. GPT-40).
- Larger, more capable models are generally more expensive to run and thus cost more for users.
- Users should choose tiers based on their needs, balancing cost with performance for professional or personal use.
- Claude (Anthropic) and Gemini (Google) offer professional plans with advanced models like Claude 3.5 Sonnet and Gemini Pro.
- Grok (xAI) has different versions (e.g., Grok 3), and users should select the most advanced available.
- Experimenting with multiple providers and tiers is recommended to find the best fit for specific tasks.
- Reinforcement learning (RL) trains models on problem-solving, enabling them to discover thinking strategies.
- These 'thinking models' exhibit inner monologues, backtracking, and assumption revisiting, improving accuracy on complex tasks.
- RL-tuned models can take minutes to process queries but offer higher accuracy for math, code, and reasoning problems.
- Standard GPT-40 struggled with a programming bug, offering general debugging advice.
- GPT-40 O1 Pro (a thinking model) took 1 minute to analyze the code and identify the specific parameter mismatch.
- Other models like Claude 3.5 Sonnet, Gemini, and Grok 3 also successfully solved the programming issue, sometimes without needing a dedicated 'thinking' mode.
- LLMs can use internet search tools to access recent information not present in their training data.
- Tools like Perplexity AI and ChatGPT's web search integrate this capability, fetching and summarizing web content.
- Models may automatically detect the need for search or require explicit user instruction to use the tool.
- Search tools are ideal for queries about current market status, recent events, product launches, and trending topics.
- Examples include checking stock market status, finding release dates for TV shows, or understanding recent news.
- Perplexity AI is highlighted for its strong integration of search capabilities for these types of queries.
- Deep research features (e.g., in ChatGPT Pro, Perplexity, Grok) combine extensive internet search with model thinking.
- These tools can spend tens of minutes processing information to generate comprehensive reports on complex topics.
- Examples include researching health actives like Ca-AKG or exploring company funding and size.
- Users can upload documents (PDFs, text files) to LLMs like Claude 3.7 to provide specific context.
- The LLM can then answer questions or summarize information directly from the uploaded files.
- This functionality is useful for understanding research papers, books, or personal documents like blood test results.
- LLMs can utilize a Python interpreter to write and execute code, enabling complex calculations and data manipulation.
- ChatGPT recognizes when to use the interpreter for problems beyond its direct calculation capabilities (e.g., large multiplications).
- Not all LLMs have access to or utilize programming language tools, potentially leading to hallucinations on complex problems.
- This feature allows ChatGPT to act as a junior data analyst, processing and visualizing uploaded data.
- It can generate plots, fit trend lines, and perform extrapolations, but requires user scrutiny of the generated code and assumptions.
- An example showed ChatGPT implicitly assuming a valuation for a missing data point and misrepresenting extrapolation results.
- Claude's Artifacts feature allows users to generate custom web applications directly within the chat interface.
- An example demonstrated creating a flashcard app from text and a conceptual diagram generator using Mermaid.
- These are local browser-based applications, useful for specific tasks like visualization or interactive learning.
- For professional coding, dedicated apps like Cursor (using Claude 3.7 Sonnet) offer integrated LLM assistance.
- These tools have file system context and advanced features like 'Composer' for autonomous agent-like code generation and editing.
- Vibe coding refers to letting the AI composer handle low-level programming tasks with minimal human direction.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.