The spelled-out intro to language modeling: building makemore
Watch on YouTube →
Overview
Andrej Karpathy builds a character-level bigram language model called makemore from scratch, demonstrating how to process text data, count character bigram frequencies, and represent these counts in a 2D PyTorch tensor. He then implements the same model using a neural network, showing how gradient-based optimization can learn the same probability distributions as direct counting, achieving a comparable loss of ~2.45. This neural network approach, while initially more complex, offers greater flexibility for future model enhancements.
Key takeaways
- Makemore demonstrates building character-level language models, starting with bigrams and progressing towards transformers.
- A bigram model predicts the next character based solely on the preceding one, using counts or neural network weights.
- Directly counting bigram frequencies and normalizing yields probabilities, achieving a loss of ~2.45.
- A neural network approach using a single linear layer and softmax also learns these probabilities via gradient descent, reaching a similar loss.
- The neural network framework, though initially more complex for bigrams, offers scalability and flexibility for more advanced models like transformers.
Chapters
- Makemore is a GitHub repository for building language models step-by-step.
- The goal is to model sequences of characters to predict the next character.
- The project will cover models from bigrams to transformers, including GPT-2 equivalents.
- The dataset 'names.txt' contains 32,000 names.
- Names are treated as sequences of characters.
- Analysis reveals word lengths range from 2 to 15 characters.
- A bigram model predicts the next character based only on the previous one.
- Special start and end tokens are added to words for context.
- Character bigram frequencies are counted using a Python dictionary.
- Counts are stored in a 2D PyTorch tensor (N) for efficiency.
- Rows represent the first character, columns represent the second character of a bigram.
- The tensor size is 28x28, accommodating 26 letters plus start/end tokens.
- A mapping (S2I) is created from characters to integer indices (0-27).
- 'a' maps to 0, 'z' to 25, '.' (start/end token) to 26.
- This mapping is crucial for indexing into the PyTorch tensor.
- Matplotlib is used to visualize the 28x28 count tensor.
- The visualization shows the frequency of each character bigram.
- Observations include zero counts for impossible bigrams (e.g., end token as first char).
- The model is simplified to use only one special token ('.') at index 0.
- The tensor size is reduced to 27x27.
- Character mappings are adjusted: '.' maps to 0, 'a' to 1, ..., 'z' to 26.
- The probability distribution for the next character is derived from the counts tensor.
- Probabilities are calculated by normalizing rows of the counts tensor.
- PyTorch's `multinomial` function is used for sampling, with a seeded generator for determinism.
- A loop iteratively samples characters starting from the '.' token.
- The process continues until the '.' token is sampled again, indicating the end of a name.
- Initial generated names are nonsensical due to the bigram model's simplicity.
- A uniform distribution (untrained model) produces random character sequences.
- The trained bigram model produces slightly more name-like, but still poor, results.
- The bigram model's limitations stem from its lack of long-range context.
- Re-normalizing rows in each sampling iteration is inefficient.
- A probability matrix (P) is pre-calculated by normalizing the counts tensor (N).
- Broadcasting rules in PyTorch are used for efficient row-wise normalization.
- Broadcasting allows operations between tensors of different shapes.
- Incorrect broadcasting (e.g., omitting `keepdim=True`) can lead to column-wise normalization instead of row-wise.
- Careful attention to tensor shapes and `keepdim` is crucial to avoid subtle bugs.
- Model quality is assessed using the negative log likelihood (NLL) loss.
- NLL is derived from the product of probabilities assigned to actual bigrams in the dataset.
- A lower NLL indicates a better model that assigns higher probabilities to observed sequences.
- The NLL loss is calculated by summing the log probabilities of observed bigrams.
- The average NLL over the dataset serves as the final loss metric (~2.45).
- Model smoothing (adding fake counts) prevents zero probabilities and infinite loss.
- The problem is reframed for a neural network: predict the next character given the previous one.
- The network takes a character's integer encoding as input and outputs a probability distribution.
- The loss function remains the negative log likelihood.
- The dataset is converted into pairs of (input character index, target character index).
- Inputs and targets are stored as PyTorch tensors.
- The dataset is initially limited to the word 'Emma' for demonstration.
- Integer character indices are converted to one-hot encoded vectors.
- The `torch.nn.functional.one_hot` function is used with `num_classes=27`.
- The resulting one-hot vectors are cast to `float32` for neural network compatibility.
- A single linear layer with 27 input and 27 output neurons is defined.
- Weights (W) are initialized using `torch.randn` from a normal distribution.
- Matrix multiplication (`@`) efficiently computes `W * X` for all inputs in parallel.
- The linear layer outputs 'logits' (log counts).
- Exponentiating logits yields approximate counts.
- Normalizing these counts produces probability distributions for the next character.
- The softmax function (exponentiate logits, then normalize) converts outputs into probabilities.
- Softmax ensures outputs are positive and sum to one, suitable for classification tasks.
- This transformation is applied after the linear layer.
- The forward pass involves one-hot encoding, linear transformation, softmax, and loss calculation.
- The loss is the average negative log likelihood, calculated using `probs.gather` and `torch.log`.
- The initial loss is high (~3.76), indicating poor initial weights.
- Gradients are reset using `zero_grad()` or `None`.
- `loss.backward()` computes gradients of the loss with respect to weights (W).
- Weights are updated using gradient descent: `W.data -= learning_rate * W.grad`.
- The training process is scaled to use all 228,000 bigrams from 'names.txt'.
- Gradient descent iteratively updates weights to minimize the NLL loss.
- The loss converges to a value around 2.45, matching the direct counting method.
- Direct counting provides an explicit, optimal solution for bigram models.
- Gradient-based optimization achieves the same result through iterative learning.
- The neural network approach is more flexible for complex models (e.g., transformers).
- Sampling involves feeding a character index, performing a forward pass to get probabilities, and sampling the next character.
- The process is repeated until the end token is generated.
- The generated samples match those from the direct counting method, confirming model equivalence.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.