Save this video — free

Building makemore Part 5: Building a WaveNet

Andrej Karpathy · 56:22 · Watch on YouTube

Building makemore Part 5: Building a WaveNet Watch on YouTube →

Overview

Andrej Karpathy refactors the 'makemore' character-level language model to resemble a WaveNet architecture, increasing input context from 3 to 8 characters and implementing a hierarchical fusion of information. He introduces custom 'flatten' and 'sequential' modules, refines the Batch Norm 1D layer for multi-dimensional inputs, and demonstrates how this deeper, more structured approach, despite similar parameter counts, shows potential for improved performance, reaching a validation loss of 1.993 with larger embeddings.

Key takeaways

Chapters

0:00 Introduction to Complexifying the Character-Level Language Model
2:18 Review of Starter Code and Existing Layers
11:36 Improving Data Visualization and Loss Curve Plotting
15:17 Refactoring the Forward Pass with Custom Embedding and Flatten Layers
21:39 Introducing a Custom Sequential Container for Model Organization
25:42 Debugging Batch Norm 1D with Single-Example Batches
28:33 Baseline Performance with Increased Context Length
35:41 Debugging Tensor Shapes and Understanding Linear Layer Behavior
40:44 Leveraging PyTorch Linear Layer's Multi-Dimensional Input Handling
46:59 Implementing Hierarchical Fusion with Flatten Consecutive

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.