Save this video — free

Building makemore Part 3: Activations & Gradients, BatchNorm

Andrej Karpathy · 1:55:58 · Watch on YouTube

Building makemore Part 3: Activations & Gradients, BatchNorm Watch on YouTube →

Overview

Andrej Karpathy delves into neural network initialization and activation/gradient behavior, highlighting how poor initialization leads to high initial loss and "hockey stick" loss curves. He demonstrates fixing these issues by adjusting weights and biases to achieve expected initial loss and better activation distributions, ultimately improving training performance and reducing saturation in activation functions like Tanh. The lecture also introduces Batch Normalization as a modern technique to stabilize training in deeper networks by normalizing activations, though it introduces complexities like batch coupling and requires careful handling during inference.

Key takeaways

Chapters

0:00 Introduction: Moving Beyond MLP for Character-Level Language Modeling
2:03 Code Refactoring and Baseline MLP Performance
5:04 Understanding torch.no_grad Decorator
6:34 Initial Model Sampling and Word Generation Quality
7:00 Problem 1: Poor Initialization and High Initial Loss
10:11 Illustrating Extreme Logits and High Loss
14:07 Fixing Initialization: Zeroing Bias and Scaling Weights
18:38 Impact of Improved Initialization on Training
21:40 Problem 2: Saturated Tanh Activations
25:15 Gradient Vanishing in Tanh's Flat Regions
33:43 Dead Neurons and Non-linearities (ReLU, Sigmoid)
39:59 Fixing Tanh Saturation: Scaling Weights
42:29 Impact of Tanh Saturation Fix on Training
46:53 Principled Weight Initialization: Scaling by Fan-in
52:21 Kaiming Initialization and Gain Factors
1:02:20 Implementing Kaiming Initialization for Tanh
1:06:44 Training with Kaiming Initialization
1:07:26 Introduction to Batch Normalization
1:10:47 Batch Normalization Mechanism: Standardization
1:16:39 Batch Normalization: Scale and Shift (Learnable Parameters)
1:19:04 Batch Normalization Impact on a Simple Network
1:23:37 Batch Normalization's Cost: Batch Coupling
1:30:02 Handling Batch Normalization at Inference Time
1:38:51 Running Mean/Variance Update in Batch Normalization
1:41:40 Batch Normalization Implementation Details
1:45:08 Summary of Batch Normalization Layer
1:47:30 ResNet Architecture and BN Placement
1:55:01 PyTorch Linear and BatchNorm Layer Details

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.