Save this video — free

Let's reproduce GPT-2 (124M)

Andrej Karpathy · 4:01:26 · Watch on YouTube

Let's reproduce GPT-2 (124M) Watch on YouTube →

Overview

Andrej Karpathy meticulously reproduces the GPT-2 124M model from scratch, detailing each component from tokenization and Transformer architecture to training optimizations. He leverages Hugging Face Transformers for initial weight loading, implements a custom GPT model in PyTorch, and explores performance enhancements like TF32, BF16, Torch Compile, Flash Attention, and distributed training across 8 GPUs. The process culminates in training on the FineWebEdu dataset, achieving competitive HellaSwag scores and demonstrating significant learning efficiency gains over original GPT-2 benchmarks.

Key takeaways

Chapters

0:00 Introduction to Reproducing GPT-2 124M
5:40 Loading the Original GPT-2 Weights via Hugging Face
10:11 Analyzing GPT-2 Model Parameters: Embeddings and Weights
18:50 Sampling from the Loaded GPT-2 Model
20:41 Implementing a Custom GPT Model from Scratch
22:28 GPT-2 Architecture: Decoder-Only Transformer Modifications
24:56 Structuring the Custom GPT Model: NN Module Skeleton
29:13 Implementing the Transformer Block: Pre-Normalization
34:13 Implementing the MLP Block: GELU Activation
39:14 Implementing the Attention Mechanism: Efficient PyTorch Implementation
45:16 Completing the GPT-2 Implementation and Weight Porting
51:47 Implementing the Forward Pass for Generation
55:32 Generating Text with the Custom GPT-2 Model
1:03:30 Sampling Strategy: Top-K Sampling
1:08:33 Initializing a Random GPT Model from Scratch
1:11:59 Device Auto-Detection (CPU, CUDA, MPS)
1:16:40 Preparing the Tiny Shakespeare Dataset
1:19:12 Creating Input and Target Batches for Training
1:27:31 Calculating and Backpropagating the Loss
1:34:03 Optimizer Setup and Initial Training Loop
1:42:39 Implementing a Simple Data Loader for Training
1:50:16 Addressing the Weight Tying Bug
2:02:20 Refining Initialization According to GPT-2 Source Code
2:16:59 Leveraging Hardware: TensorFloat-32 (TF32)
2:46:42 Mixed Precision Training with BFloat16
3:00:25 Accelerating Training with Torch Compile
3:20:26 Integrating Flash Attention for Faster Attention Calculation
3:31:50 Optimizing Batch Sizes for GPU Efficiency
3:45:12 Adopting GPT-3 Hyperparameters for Optimization

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.