Building makemore Part 2: MLP
Watch on YouTube →
Overview
Andrej Karpathy implements a Multi-Layer Perceptron (MLP) for character-level language modeling, building upon the previous bigram model. This MLP uses embeddings to represent characters, a hidden layer for non-linear transformations, and an output layer for predicting the next character, drawing inspiration from the Bengio et al. (2003) paper. The implementation details include efficient tensor manipulation with PyTorch's `view` and `cat`, numerical stability considerations for cross-entropy, and a systematic approach to hyperparameter tuning and training with mini-batches.
Key takeaways
- MLPs with embeddings and hidden layers significantly outperform bigram models by capturing richer contextual information and allowing generalization through continuous representation.
- PyTorch's tensor manipulation capabilities, particularly `view` for efficient reshaping and flexible indexing, are crucial for building neural network components.
- Numerical stability in loss calculation is critical; `F.cross_entropy` handles potential overflows better than manual softmax and log calculations.
- Mini-batch training is essential for practical deep learning, balancing gradient accuracy with computational efficiency.
- A systematic approach involving train/dev/test splits and hyperparameter tuning (learning rate, network size) is necessary to avoid overfitting and achieve optimal performance.
Chapters
- Bigram models use only one previous character for prediction, leading to limited context.
- Increasing context exponentially increases the number of states, making count-based models infeasible.
- Introduces the Multi-Layer Perceptron (MLP) as a solution for character-level language modeling.
- Discusses the influential Bengio et al. (2003) paper on neural language models.
- Explains word embeddings: associating each word with a fixed-size vector in a continuous space.
- Embeddings allow generalization by placing similar words (e.g., synonyms) close in the embedding space.
- Details the MLP architecture: input embeddings, hidden layer, and output layer with softmax.
- Input: sequence of previous characters (e.g., 3 characters).
- Embedding Layer: lookup table (matrix C) mapping character indices to dense vectors.
- Hidden Layer: fully connected layer with non-linearity (e.g., tanh).
- Output Layer: predicts probabilities for all possible next characters (e.g., 27 characters).
- Defines `block_size` as the context length (number of characters used for prediction).
- Creates input `x` (context sequences) and output `y` (target characters) tensors.
- Illustrates data generation with a rolling window approach, padding with '.' characters.
- Initializes an embedding matrix `c` (e.g., 27 characters x 2 dimensions).
- Demonstrates embedding a single integer index by looking up a row in `c`.
- Explains the equivalence of direct indexing and one-hot encoding followed by matrix multiplication.
- Shows PyTorch's flexible indexing capabilities for tensors.
- Can index with lists or tensors of integers to retrieve multiple embedding vectors simultaneously.
- `c[x]` efficiently embeds all integers in the input tensor `x`.
- Input to the hidden layer is the concatenated embeddings of context characters.
- Demonstrates `torch.cat` and `torch.unbind` for concatenating embeddings.
- Introduces `view` as a more efficient way to reshape tensors without copying data.
- Defines weights `w1` and biases `b1` for the hidden layer.
- Calculates hidden layer activations `h = torch.tanh(x @ w1 + b1)`.
- Explains broadcasting for adding the bias vector `b1`.
- Defines weights `w2` and biases `b2` for the output layer.
- Calculates logits: `logits = h @ w2 + b2`.
- Applies `torch.softmax` to logits to obtain a probability distribution over the next character.
- Calculates the loss by indexing into probabilities with the target character `y`.
- Uses negative log likelihood: `-log_prob`.
- Introduces `F.cross_entropy` as a more efficient and numerically stable alternative.
- Sets up the training loop: zero gradients, backward pass (`loss.backward()`), parameter update.
- Updates parameters using `p.data += -learning_rate * p.grad`.
- Demonstrates initial loss reduction from ~17 to ~2.3 on a small subset.
- Explains the inefficiency of training on the entire dataset at once.
- Introduces mini-batching: randomly selecting subsets of data for each training step.
- Uses `torch.randint` to select indices for mini-batches, significantly speeding up training.
- Discusses the importance of choosing an appropriate learning rate.
- Demonstrates a learning rate search by training for a fixed number of steps with exponentially spaced learning rates.
- Plots loss vs. learning rate exponent to identify a good range (e.g., around 10^-1).
- Explains the need for train, development (dev), and test splits to prevent overfitting.
- Training set: optimizes model parameters.
- Dev set: tunes hyperparameters (e.g., hidden layer size, embedding dimension).
- Test set: final evaluation of the chosen model.
- Initial model with 100 hidden neurons and 2D embeddings shows underfitting (train/dev loss similar).
- Increases hidden layer size to 300 neurons, then to 200 neurons with 10D embeddings.
- Observes gradual loss reduction (from ~2.3 to ~2.17) as model capacity increases.
- Visualizes 2D embeddings, showing clustering of vowels and separation of unique characters.
- Demonstrates sampling from the trained model to generate new character sequences.
- Generated samples show more name-like structures compared to the bigram model.
- Achieved a dev loss of ~2.17, surpassing the bigram model's ~2.45.
- Suggests further improvements: increasing context length, tuning optimization details, and exploring paper ideas.
- Provides a Google Colab link for interactive experimentation without local setup.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.