Save this video — free

The spelled-out intro to language modeling: building makemore

Andrej Karpathy · 1:57:45 · Watch on YouTube

The spelled-out intro to language modeling: building makemore Watch on YouTube →

Overview

Andrej Karpathy builds a character-level bigram language model called makemore from scratch, demonstrating how to process text data, count character bigram frequencies, and represent these counts in a 2D PyTorch tensor. He then implements the same model using a neural network, showing how gradient-based optimization can learn the same probability distributions as direct counting, achieving a comparable loss of ~2.45. This neural network approach, while initially more complex, offers greater flexibility for future model enhancements.

Key takeaways

Chapters

0:00 Introduction to Makemore and Character-Level Language Models
3:13 Loading and Analyzing the Names Dataset
4:27 Understanding Bigram Statistics and Data Representation
13:29 Transitioning from Dictionary Counts to a 2D Tensor
15:34 Creating Character-to-Integer Mappings
18:29 Visualizing the Bigram Count Tensor
21:40 Refining the Bigram Model with a Single Special Token
23:30 Sampling Characters from the Bigram Model
30:17 Generating Names with the Bigram Model
35:17 Comparing Bigram Model Performance to Uniform Distribution
37:37 Improving Efficiency: Pre-calculating Probability Matrix
42:43 Understanding and Avoiding Broadcasting Errors
53:27 Evaluating Model Quality with Negative Log Likelihood
57:20 Calculating NLL Loss for the Bigram Model
1:04:11 Introducing the Neural Network Approach to Bigram Modeling
1:06:44 Preparing the Training Data for the Neural Network
1:10:53 One-Hot Encoding Input Characters
1:13:32 Constructing the First Neural Network Layer (Linear Layer)
1:18:37 Interpreting Neural Network Outputs: Logits, Counts, and Probabilities
1:23:28 The Softmax Function for Probability Distribution Output
1:28:37 Implementing the Forward Pass and Calculating Loss
1:30:44 Backpropagation and Gradient Descent for Weight Updates
1:38:33 Training the Neural Network on the Full Dataset
1:45:40 Comparing Counting vs. Neural Network Approaches
1:50:09 Sampling from the Trained Neural Network Model

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.