Deep Dive into LLMs like ChatGPT
Watch on YouTube →
Overview
Andrej Karpathy provides a comprehensive deep dive into Large Language Models (LLMs) like ChatGPT, detailing their construction from pre-training on internet data to supervised fine-tuning (SFT) with curated conversations and reinforcement learning (RL) for emergent reasoning. He explains tokenization, Transformer architecture, inference, computational requirements (GPUs), and the psychological implications like hallucinations, emphasizing LLMs as powerful but imperfect tools requiring careful use and verification.
Key takeaways
- LLMs are trained in three stages: pre-training (internet knowledge), SFT (imitating expert conversations), and RL (discovering reasoning strategies).
- Tokenization is fundamental: LLMs process text as sequences of tokens, impacting their abilities in character-level tasks and complex reasoning.
- Hallucinations are mitigated through knowledge refusal and tool use (web search, code interpreter), but not entirely eliminated.
- RL-trained 'thinking' models show emergent reasoning by discovering novel problem-solving strategies, unlike SFT models which primarily imitate.
- LLMs exhibit 'Swiss cheese' capabilities, being highly competent in many areas but failing randomly on simple tasks, necessitating verification.
- The future involves multimodal LLMs and agents performing long-running tasks, requiring careful supervision and new approaches for context management.
Chapters
- Goal: Provide mental models for understanding LLMs like ChatGPT.
- Covers the entire pipeline of LLM construction for a general audience.
- Discusses strengths, weaknesses, and sharp edges of LLM tools.
- Data collection involves processing vast amounts of text from the internet.
- Hugging Face's 'FineWeb' dataset (44 TB) is an example of curated internet data.
- Common Crawl (2.7 billion web pages indexed) is a primary source for raw data.
- URL filtering removes spam, malware, and adult content domains.
- Text extraction isolates relevant content from raw HTML markup.
- Language filtering prioritizes specific languages (e.g., >65% English for FineWeb).
- Personally Identifiable Information (PII) like addresses and SSNs are filtered out.
- Deduplication and other filtering steps refine the dataset.
- The resulting dataset (e.g., FineWeb) is approximately 44 TB of text.
- Neural networks require text to be represented as a 1D sequence of symbols (tokens).
- Byte Pair Encoding (BPE) algorithm creates a vocabulary of ~100,000 tokens.
- GPT-4 uses a vocabulary of 100,277 tokens (cl100k_base tokenizer).
- The 'tiktoken' tool visualizes text-to-token conversion for models like GPT-4.
- 'Hello world' tokenizes into two tokens: 'hello' (15339) and ' world' (1917).
- Tokenization is case-sensitive and sensitive to whitespace variations.
- The FineWeb dataset is estimated to be around 15 trillion tokens.
- These tokens are the fundamental units fed into the neural network.
- Token IDs are unique identifiers, not inherently meaningful values.
- The goal is to model statistical relationships between tokens in a sequence.
- Training uses windows of tokens (context) to predict the next token.
- Context lengths can range from zero up to a maximum (e.g., 8,000 tokens).
- Input: Sequences of tokens (context).
- Output: Probability distribution over the entire vocabulary for the next token.
- Initially, the network parameters are random, leading to random predictions.
- Modern LLMs use the Transformer architecture, characterized by attention mechanisms.
- A production-grade Transformer (e.g., 8.5 billion parameters) involves layers of computation.
- Tokens are embedded into vector representations before processing through layers (norm, matmul, softmax).
- Neural networks are massive mathematical functions parameterized by billions of weights.
- Training optimizes these parameters (knobs) to align predictions with training data statistics.
- The architecture is designed to be expressive, optimizable, and parallelizable.
- Inference is the process of generating new text from a trained model.
- Starts with a prefix (seed tokens) and iteratively predicts the next token.
- Sampling from the probability distribution introduces stochasticity, leading to varied outputs.
- GPT-2 (Generative Pre-trained Transformer 2) had 1.6 billion parameters.
- Maximum context length was 1,024 tokens, trained on ~100 billion tokens.
- Training cost in 2019 was estimated at $40,000; modern reproduction costs ~$100-600.
- Training involves millions of updates, each improving prediction accuracy.
- The 'loss' metric quantifies model performance; decreasing loss indicates improvement.
- Updates are processed in batches (e.g., 1 million tokens per update).
- Training requires significant compute, typically using clusters of GPUs (e.g., 8x NVIDIA H100s).
- Cloud providers like Lambda rent out GPU instances for training.
- The demand for GPUs has driven Nvidia's market capitalization to $3.4 trillion.
- Companies release trained models (base models) which are essentially token simulators.
- Base models require further post-training to become useful assistants.
- GPT-2 (1.5B parameters) and Llama 3 (405B parameters) are examples of released base models.
- Base models act as sophisticated autocomplete, continuing token sequences based on learned statistics.
- They lack inherent conversational ability or instruction-following capabilities without post-training.
- Knowledge is stored probabilistically in parameters, representing a 'lossy compression' of the internet.
- Prompting can elicit knowledge stored in parameters, but it's often vague and probabilistic.
- Models can exhibit strong memorization of training data (e.g., Wikipedia entries).
- Knowledge cut-off dates (e.g., end of 2023 for Llama 3) limit awareness of recent events.
- Few-shot prompting provides examples within the prompt to guide the model's behavior.
- Models exhibit 'in-context learning,' adapting to patterns presented in the prompt.
- This allows creating assistant-like behavior even with a base model.
- Post-training refines base models for instruction following and conversational interaction.
- Computationally less expensive than pre-training, often taking hours instead of months.
- Focuses on training data consisting of conversations between humans and assistants.
- Conversations are encoded into token sequences using specific protocols.
- Special tokens (e.g., `IM_START`, `IM_END`, `USER`, `ASSISTANT`) signal conversational structure.
- These special tokens are new, untaught tokens introduced during post-training.
- OpenAI's 2022 paper on InstructGPT detailed SFT using human labelers.
- Labelers created prompts and wrote ideal assistant responses based on instructions (helpful, truthful, harmless).
- The InstructGPT dataset itself was not released, but open-source efforts like OpenAssistant followed.
- Current SFT datasets (e.g., UltraChat) often use LLMs to generate synthetic conversations.
- Human involvement is still crucial for editing, curation, and ensuring quality.
- These datasets contain millions of conversations across diverse topics.
- ChatGPT's responses are statistical simulations of human labelers following instructions.
- The model imitates the persona and behavior defined in the SFT data.
- Labelers are often experts, making the simulation a simulation of skilled individuals.
- Hallucinations occur when LLMs fabricate information, often due to imitating confident answers in training data.
- Models lack direct internet access for real-time fact-checking during inference.
- Older models like Falcon 7B instruct are prone to making up facts when uncertain.
- Adding examples to the training set where the correct response is 'I don't know' helps.
- Models can be probed empirically to identify knowledge gaps.
- This teaches the model to express uncertainty rather than fabricating answers.
- LLMs can be equipped with tools like web search to retrieve real-time information.
- Special tokens (`SEARCH_START`, `SEARCH_END`) allow the model to invoke tools.
- Retrieved information is placed in the context window, acting as working memory.
- Knowledge in parameters is a vague recollection; context window knowledge is immediate working memory.
- Providing information directly in the context window yields higher quality results than relying on memory.
- Prompt engineering is key to effectively utilizing LLMs as tools.
- LLMs lack persistent self-identity; they are stateless processes restarted per conversation.
- Responses about identity (e.g., 'I am ChatGPT by OpenAI') are often statistical guesses based on training data prevalence.
- Identity can be explicitly programmed via system messages or hardcoded examples in SFT data.
- LLMs perform computation token-by-token with finite compute per token.
- Complex reasoning must be distributed across many tokens (e.g., step-by-step math).
- Asking for immediate answers in a single token often fails for complex calculations.
- LLMs struggle with precise counting and complex mental arithmetic.
- The Code Interpreter tool allows LLMs to write and execute Python code for calculations.
- Leaning on tools like code interpreters is more reliable than trusting LLM's internal arithmetic.
- Models operate on tokens, not characters, hindering performance on spelling-related tasks.
- Tasks requiring character-level manipulation (e.g., every third character) often fail.
- Using the Code Interpreter tool circumvents tokenization limitations for character manipulation.
- LLMs exhibit 'Swiss cheese' capabilities: highly competent in many areas but randomly failing on simple tasks.
- Examples include miscounting letters (e.g., 'R' in 'strawberry') or basic numerical comparisons.
- These failures stem from tokenization, finite per-token computation, and potential cognitive distractions (e.g., Bible verse patterns).
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.