Save this video — free

Let's build the GPT Tokenizer

Andrej Karpathy · 2:13:35 · Watch on YouTube

Let's build the GPT Tokenizer Watch on YouTube →

Overview

Andrej Karpathy details the intricacies of tokenization in large language models, explaining its necessity and the "foot guns" it introduces. He covers the Byte Pair Encoding (BPE) algorithm, its implementation, and how it's used in GPT-2, GPT-4, and SentencePiece tokenizers, highlighting issues like non-English language performance, arithmetic difficulties, and "solid gold Magikarp" token anomalies, while also discussing special tokens and vocabulary size considerations.

Key takeaways

Chapters

0:00 Introduction to Tokenization and its Importance
0:20 Character-Level Tokenization vs. Subword Tokenization
7:00 Tokenization Issues: Spelling, Arithmetic, and Language Bias
10:00 Interactive Tokenizer Demo (tiktoken.app)
11:51 Case Sensitivity and Contextual Tokenization
18:45 Python Code Tokenization and Efficiency
20:29 Comparing GPT-2 and GPT-4 Tokenizer Vocabulary Sizes
24:17 Understanding Unicode Code Points
29:00 Why Not Use Raw Unicode Code Points as Tokens?
30:17 UTF-8 Encoding and its Limitations for Tokenization
37:13 Introduction to Byte Pair Encoding (BPE)
45:01 Implementing BPE: Training the Tokenizer
58:56 BPE Training Loop and Vocabulary Size
1:05:23 Tokenizer as a Separate Pre-processing Stage
1:10:43 Implementing BPE Decoding: Tokens to Text
1:20:23 Implementing BPE Encoding: Text to Tokens
1:34:16 BPE Algorithm Basics and Implementation Refinements
1:35:37 GPT-2 Tokenizer: Regex Chunking and Merging Rules
1:59:00 OpenAI's tiktoken Library for Inference
2:10:28 Special Tokens: End-of-Text and Conversation Delimiters

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.