Let's build GPT: from scratch, in code, spelled out.
Watch on YouTube →
Overview
Andrej Karpathy provides a comprehensive, code-driven tutorial on building a Transformer-based language model from scratch, mirroring the architecture of GPT. He starts with character-level tokenization of "tiny Shakespeare," progresses through implementing a Bigram model, and then details the self-attention mechanism, multi-head attention, feed-forward networks, residual connections, and layer normalization. The final scaled-up model, trained for approximately 15 minutes on an A100 GPU, achieves a validation loss of 1.48, demonstrating the core components of modern large language models.
Key takeaways
- Andrej Karpathy demonstrates building a GPT-like Transformer language model from scratch using Python and PyTorch, starting with character-level tokenization and progressing to multi-head self-attention, residual connections, and layer normalization.
- The "tiny Shakespeare" dataset is used to train a character-level model, illustrating core concepts like token embeddings, positional encodings, causal masking, and scaled dot-product attention.
- Key architectural components of Transformers, including self-attention, multi-head attention, feed-forward networks, residual connections, and layer normalization, are implemented and explained step-by-step.
- Scaling up the model with more layers, larger embedding dimensions, and increased context significantly improves performance, reducing validation loss from ~2.5 to 1.48.
- The distinction between decoder-only Transformers (like GPT for generation) and encoder-decoder Transformers (for translation) is clarified, highlighting the role of causal masking in autoregressive models.
- The process of pre-training (language modeling on vast data) and fine-tuning (alignment for specific tasks like chat) is outlined, explaining how models like ChatGPT are developed beyond basic text completion.
Chapters
- ChatGPT's ability to generate text sequentially and probabilistically is demonstrated.
- Language models predict the next word/token in a sequence.
- The Transformer architecture, introduced in the "Attention Is All You Need" paper, is the foundation of GPT.
- The goal is to train a character-level Transformer language model.
- The "tiny Shakespeare" dataset, a 1MB file of Shakespeare's works, is used for training.
- The model will predict the next character based on preceding characters.
- Text is converted into a sequence of integers (tokens) using a character-level tokenizer.
- A vocabulary of 65 unique characters (including spaces and punctuation) is identified from the dataset.
- Encoder and decoder mappings between characters and integers are established.
- The entire dataset is encoded into a single large PyTorch tensor of integers.
- The data is split into 90% for training and 10% for validation to monitor overfitting.
- The validation set is held out to ensure the model generalizes rather than memorizes.
- Transformers are trained on fixed-size chunks (blocks) of data, not the entire sequence at once.
- A block size (e.g., 8) defines the maximum context length.
- Each block contains multiple training examples (block_size examples in a block of size block_size + 1).
- Multiple independent chunks of data are processed simultaneously in batches for GPU efficiency.
- A batch size (e.g., 4) determines how many sequences are processed in parallel.
- Input (X) and target (Y) tensors are created with dimensions Batch x Time (e.g., 4x8).
- The simplest language model, a Bigram model, is implemented using PyTorch's nn.Module.
- A token embedding table maps integer tokens to dense vectors.
- Logits (scores for the next token) are generated, and the cross-entropy loss is calculated against targets.
- PyTorch's cross-entropy loss function is used to evaluate model predictions against targets.
- Logits need reshaping (Batch*Time x VocabSize) to match PyTorch's cross-entropy input expectations.
- The initial loss for a random model is high (e.g., 4.87), indicating poor predictions.
- A `generate` function is implemented to extend a given sequence of tokens.
- The model predicts the next token, samples from the probability distribution (softmax), and appends it.
- Initial generation from a random model produces nonsensical output.
- The Adam optimizer is used for training, with a learning rate of 3e-4.
- A standard training loop samples batches, calculates loss, backpropagates gradients, and updates model parameters.
- Training for 10,000 iterations reduces the loss significantly, improving generation quality.
- The Bigram model lacks context; tokens don't communicate.
- Self-attention allows tokens to interact and weigh information from other tokens in the sequence.
- A "masked" aggregation (averaging preceding tokens) is the initial concept before matrix multiplication optimization.
- A mathematical trick using matrix multiplication with a lower triangular matrix enables efficient weighted aggregation.
- This allows calculating weighted sums of past elements in a data-dependent manner.
- The weights determine how much information from each preceding token is incorporated.
- Each token emits Query (Q) and Key (K) vectors; Q dot K calculates affinities (weights).
- Value (V) vectors are also generated; weighted aggregation of V using affinities forms the output.
- A "head size" (e.g., 16) defines the dimensionality of Q, K, and V for a single attention head.
- Attention scores are scaled by 1/sqrt(head_size) to stabilize variance during training.
- A causal mask (lower triangular) prevents future tokens from influencing past tokens in decoder blocks.
- Softmax is applied to affinities to get attention weights, which then weight the aggregation of Value vectors.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.