Save this video — free

Deep Dive into LLMs like ChatGPT

Andrej Karpathy · 3:31:24 · Watch on YouTube

Deep Dive into LLMs like ChatGPT Watch on YouTube →

Overview

Andrej Karpathy provides a comprehensive deep dive into Large Language Models (LLMs) like ChatGPT, detailing their construction from pre-training on internet data to supervised fine-tuning (SFT) with curated conversations and reinforcement learning (RL) for emergent reasoning. He explains tokenization, Transformer architecture, inference, computational requirements (GPUs), and the psychological implications like hallucinations, emphasizing LLMs as powerful but imperfect tools requiring careful use and verification.

Key takeaways

Chapters

0:00 Introduction to LLMs like ChatGPT
1:00 Stage 1: Pre-training - Downloading and Processing the Internet
5:30 Data Filtering and Cleaning for Pre-training
9:01 PII Removal and Data Set Finalization
12:22 Tokenization: Representing Text for Neural Networks
20:00 Exploring Tokenization with Tiktoken
24:00 Pre-training Data Size: 15 Trillion Tokens
25:17 Stage 2: Neural Network Training - Predicting the Next Token
29:13 Neural Network Input/Output and Random Initialization
33:32 Transformer Architecture: The Core of Modern LLMs
38:24 Transformer Internals: Parameters and Mathematical Expressions
43:22 Stage 3: Inference - Generating New Data
51:42 GPT-2: A Foundational Modern LLM (2019)
57:30 Training GPT-2: Monitoring Loss and Updates
1:05:02 Computational Requirements: GPUs and Data Centers
1:11:41 Model Releases: Base Models vs. Assistants
1:17:27 Interacting with Llama 3 Base Model
1:23:58 Eliciting Knowledge and Memorization in Base Models
1:33:28 Few-Shot Prompting and In-Context Learning
1:38:52 Post-training Stage: Turning Base Models into Assistants
1:48:21 Tokenizing Conversations and Special Tokens
1:56:53 InstructGPT: Early Work on Supervised Fine-Tuning (SFT)
2:05:06 Modern SFT Data Sets: Synthetic and Human-Assisted
2:08:20 Understanding ChatGPT's Behavior: Simulating Human Labelers
2:13:53 LLM Psychology: Hallucinations and Their Origins
2:20:47 Mitigating Hallucinations: Knowledge Refusal
2:33:44 Mitigation 2: Tool Use for Factual Accuracy
2:45:37 Context Window vs. Parameter Knowledge
2:49:07 LLM Psychology: Knowledge of Self and Identity
2:58:26 LLM Psychology: Computational Limitations and Reasoning
3:08:35 Tool Use for Math and Counting: Code Interpreter
3:21:52 LLM Psychology: Tokenization and Spelling Deficits
3:27:24 LLM Psychology: 'Swiss Cheese' Capabilities and Random Failures

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.