Let's reproduce GPT-2 (124M)
Watch on YouTube →
Overview
Andrej Karpathy meticulously reproduces the GPT-2 124M model from scratch, detailing each component from tokenization and Transformer architecture to training optimizations. He leverages Hugging Face Transformers for initial weight loading, implements a custom GPT model in PyTorch, and explores performance enhancements like TF32, BF16, Torch Compile, Flash Attention, and distributed training across 8 GPUs. The process culminates in training on the FineWebEdu dataset, achieving competitive HellaSwag scores and demonstrating significant learning efficiency gains over original GPT-2 benchmarks.
Key takeaways
- Reproducing GPT-2 124M from scratch is feasible with modern tools and compute, achieving competitive results.
- Performance optimizations like TF32, BF16, Torch Compile, and Flash Attention significantly accelerate training.
- Distributed Data Parallel (DDP) with gradient accumulation enables training large models on limited hardware.
- High-quality, filtered datasets like FineWebEdu allow matching or exceeding older benchmarks with less data.
- HellaSwag evaluation provides a useful, albeit imperfect, signal for language model progress.
Chapters
- Goal: reproduce GPT-2 124 million parameter model.
- GPT-2 released in 2019 by OpenAI with a blog post, paper, and GitHub code.
- Focus on the 124M parameter version, part of a series of models with varying sizes.
- Reproducing the model can be done in about an hour for around $10 on cloud compute.
- Original GPT-2 code used TensorFlow; aim is to use PyTorch.
- Hugging Face Transformers provides a PyTorch implementation and converted weights.
- Loading the 'gpt2' model from Hugging Face defaults to the 124M parameter version.
- Extracting the state_dict to inspect raw tensors and parameter shapes.
- Token embedding shape: 50257 (vocab size) x 768 (embedding dimension).
- Positional embedding shape: 1024 (max sequence length) x 768 (embedding dimension).
- Visualizing position embeddings shows learned sinusoidal structure.
- Examining individual Transformer layer weights reveals complex structure.
- Using Hugging Face pipeline to sample text generations from the loaded model.
- Prefix: 'hello I'm a language model'.
- Generated five sequences of 30 tokens each.
- Output shows coherent text, demonstrating successful model loading and generation.
- Goal: implement a GPT model from scratch for full understanding.
- Will mirror Hugging Face Transformers' schema for easier weight loading.
- First task: load OpenAI's GPT-2 124M weights into the custom class.
- Later: initialize from scratch and train to surpass original performance.
- GPT-2 is a decoder-only Transformer, omitting the encoder and cross-attention.
- Key differences from original Transformer: layer norm reshuffling and an additional final layer norm.
- Layer norms are moved before attention/MLP blocks (pre-normalization).
- An extra layer norm is added before the final classifier.
- Main container: `Transformer` NN module dict.
- Includes token embeddings (`wte`) and positional embeddings (`pe`).
- 12 Transformer blocks are stored in `h` as an NN module list.
- Final layer norm (`lnf`) and language model head (`lm_head`) complete the structure.
- Block structure uses pre-normalization: LayerNorm -> Attention -> Residual -> LayerNorm -> MLP -> Residual.
- Residual pathways incorporate normalization, unlike post-normalization versions.
- Attention is the communication/reduce operation; MLP is the map operation per token.
- Transformer is a repeated application of map-reduce blocks.
- MLP consists of two linear projections with a GELU nonlinearity.
- GPT-2 uses the approximate GELU (`nn.GELU(approximate='tanh')`).
- GELU is smoother than ReLU and avoids the 'dying ReLU' problem by always contributing a local gradient.
- Approximate GELU was used historically due to TensorFlow performance; exact GELU is preferred now.
- Multi-head attention is implemented efficiently using tensor gymnastics (transpose, split, view).
- Heads are treated as a batch dimension for parallel computation.
- Key operations: QKV calculation, masking (autoregressive), softmax normalization, attention-value matrix multiplication.
- Final transpose/reshape reassembles concatenated head outputs.
- Custom GPT-2 implementation is less than 100 lines of code.
- Configuration matches GPT-2 124M: block_size=1024, vocab_size=50257, 12 layers, 12 heads, embed_dim=768.
- Code to load weights from Hugging Face state_dict into the custom model.
- Handles weight transposition issues from TensorFlow source and ignores non-parameter buffers.
- Forward function takes token indices (batch x time).
- Combines token and position embeddings.
- Passes through Transformer blocks, final LayerNorm, and LM head.
- Outputs logits (batch x time x vocab_size) for next token prediction.
- Setting model to evaluation mode (`model.eval()`).
- Moving model tensors to GPU (`.to('cuda')`).
- Tokenizing prefix 'hello I'm a language model' using Tiktoken.
- Iterative generation loop: predict next token, append, repeat.
- Extracting logits only from the last token position.
- Applying softmax to get probabilities.
- Using top-K sampling (K=50, default in Hugging Face pipeline) to filter unlikely tokens.
- Renormalizing probabilities and sampling the next token.
- PyTorch initializes layers randomly by default (e.g., Xavier initialization).
- Creating a model with `GPT(GPTConfig())` uses default 124M parameters.
- Generating text from a randomly initialized model produces garbage output.
- This confirms the model architecture is correct, ready for training.
- Code automatically detects available device: CUDA > MPS (Apple Silicon) > CPU.
- Ensures tensors and model are on the same device to avoid errors.
- Forward pass correctly uses the detected device for tensor creation.
- Allows running on systems without GPUs, though slower.
- Using the 'tiny_shakespeare' dataset for initial debugging and training.
- Dataset size: ~1 million characters, ~200k words.
- Tokenizing the first 1000 characters yields ~300 tokens.
- Character-level data is ASCII, 1 byte per character.
- Tokenizing the entire dataset and preparing batches of size B x T.
- Fetching `B*T + 1` tokens to create input (`x`) and target (`y`) tensors.
- Input `x` uses tokens from index 0 to `B*T - 1`.
- Target `y` uses tokens from index 1 to `B*T`, shifted by one position.
- Modifying the forward pass to return logits and loss.
- Using `torch.nn.functional.cross_entropy` for loss calculation.
- Logits and targets are flattened to `(B*T, vocab_size)` and `(B*T,)` respectively.
- Loss is calculated for the entire batch and returned.
- Using AdamW optimizer with default parameters.
- Training loop: zero gradients, forward pass, calculate loss, backward pass, optimizer step.
- Printing step number and loss value.
- Initial loss is ~10.82, expected for random initialization (1/vocab_size probability).
- Overfitting a single batch: loss rapidly decreases to near zero.
- Data loader iterates through tokenized text, creating batches of B x T.
- Fetches `B*T + 1` tokens to create input/target pairs.
- Handles looping back to the beginning of the file when data is exhausted.
- Initial training run on Tiny Shakespeare shows loss decreasing from ~11 to ~6.6 in 50 batches.
- GPT-2 shares weights between token embeddings (wte) and the final LM head.
- Original code did not implement this weight tying.
- Fix: Redirect `wte.weight` to point to `lm_head.weight`.
- This reduces parameter count by ~30% and acts as an inductive bias.
- GPT-2 source code uses specific standard deviations for initialization: 0.02 for weights, 0.0 for biases.
- Token/position embeddings initialized with 0.02 (positional embeddings sometimes 0.01).
- Custom initialization function applies these settings to linear layers and embeddings.
- Residual layer scaling: weights in residual blocks scaled by 1/sqrt(N_layers) as per GPT-2 paper.
- Default PyTorch precision is FP32 (32 bits per float).
- TF32 offers an 8x theoretical speedup on Ampere GPUs by using 19-bit mantissa internally.
- Enabling TF32 requires a single line: `torch.backends.cuda.matmul.allow_tf32 = True`.
- Achieved ~3x speedup (1000ms to 300ms per iteration), limited by memory bandwidth.
- BF16 uses 16 bits, maintaining FP32's range but reducing precision.
- Enabling BF16 via `torch.autocast(dtype=torch.bfloat16)`.
- Activations (logits) become BF16, while parameters remain FP32.
- Achieved ~1.1x speedup (300ms to ~270ms), indicating memory bandwidth is still a bottleneck.
- Torch.compile optimizes PyTorch code by removing Python overhead and fusing operations.
- Achieved ~2.3x speedup (270ms to ~120ms per iteration).
- Fusion reduces memory read/write round trips by keeping data on-chip.
- Sampling and HellaSwag evaluation were temporarily disabled due to compatibility issues with Torch Compile.
- Flash Attention is a kernel fusion algorithm that avoids materializing the full attention matrix.
- Replaces standard attention computation with a single fused kernel (`torch.nn.functional.scaled_dot_product_attention`).
- Achieved ~1.3x speedup (120ms to ~95ms per iteration), a ~27% improvement.
- This optimization highlights the importance of memory access patterns over raw FLOPs.
- GPU kernels perform best with batch sizes that are powers of two (e.g., 64, 128).
- GPT-2's vocabulary size (50257) is an 'ugly' number; padding to 50304 improved performance by ~4%.
- This padding adds computation but aligns with GPU block tile sizes, reducing inefficient boundary kernels.
- Using PyTorch nightly is crucial for these optimizations; older versions show larger gains.
- Using AdamW optimizer with GPT-3's beta parameters (0.9, 0.95) and epsilon (1e8).
- Implementing global gradient clipping at 1.0 to stabilize training.
- Introducing a cosine decay learning rate schedule with linear warmup.
- Skipping gradual batch size increase for simplicity, but noting its potential benefits.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.