Save this video — free

Let's build GPT: from scratch, in code, spelled out.

Andrej Karpathy · 1:56:20 · Watch on YouTube

Let's build GPT: from scratch, in code, spelled out. Watch on YouTube →

Overview

Andrej Karpathy provides a comprehensive, code-driven tutorial on building a Transformer-based language model from scratch, mirroring the architecture of GPT. He starts with character-level tokenization of "tiny Shakespeare," progresses through implementing a Bigram model, and then details the self-attention mechanism, multi-head attention, feed-forward networks, residual connections, and layer normalization. The final scaled-up model, trained for approximately 15 minutes on an A100 GPU, achieves a validation loss of 1.48, demonstrating the core components of modern large language models.

Key takeaways

Chapters

0:00 Introduction to ChatGPT and Language Models
3:37 Building a Character-Level Transformer: Tiny Shakespeare Dataset
15:31 Character Tokenization and Vocabulary Creation
20:46 Preparing Training Data: Tensors and Splits
23:50 Batching Data for Transformer Training
30:06 Introducing the Batch Dimension for Efficiency
37:00 Implementing a Bigram Language Model
43:17 Calculating Loss with Cross-Entropy
48:27 Generating Text from the Bigram Model
57:34 Training the Bigram Model with Adam Optimizer
1:02:39 Introducing the Transformer Architecture: Self-Attention
1:10:44 Efficient Weighted Aggregation via Matrix Multiplication
1:43:21 Implementing a Single Head of Self-Attention
1:49:57 Scaled Dot-Product Attention and Causal Masking

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.