Let's build the GPT Tokenizer
Watch on YouTube →
Overview
Andrej Karpathy details the intricacies of tokenization in large language models, explaining its necessity and the "foot guns" it introduces. He covers the Byte Pair Encoding (BPE) algorithm, its implementation, and how it's used in GPT-2, GPT-4, and SentencePiece tokenizers, highlighting issues like non-English language performance, arithmetic difficulties, and "solid gold Magikarp" token anomalies, while also discussing special tokens and vocabulary size considerations.
Key takeaways
- Tokenization is a critical, often overlooked, factor influencing LLM performance, leading to issues in spelling, arithmetic, and cross-lingual tasks.
- Byte Pair Encoding (BPE) is the dominant algorithm, iteratively merging frequent byte pairs to create subword tokens, balancing vocabulary size and sequence length.
- Special tokens (like `<|endoftext|>`, `<|im_start|>`) are essential for structuring data, delimiting documents/conversations, and enabling fine-tuning.
- Discrepancies between tokenizer training data and LLM training data can create 'dead' tokens, leading to undefined behavior and security vulnerabilities (e.g., 'solid gold Magikarp').
- Libraries like OpenAI's `tiktoken` provide efficient inference, while SentencePiece offers training capabilities but with more complexity and potential pitfalls.
- Optimizing tokenization density (e.g., YAML over JSON) is crucial for managing LLM costs and context window limitations.
Chapters
- Tokenization is a critical but complex part of LLMs, often tracing back to model oddities.
- Naive character-level tokenization is contrasted with more sophisticated schemes like BPE.
- Tokenization translates strings into sequences of integers (tokens) for LLM input.
- Early LLMs used character-level tokenization (e.g., 65 unique characters from Shakespeare).
- State-of-the-art models use subword tokenization (e.g., GPT-2's 50,257 tokens).
- Tokens are the fundamental units LLMs process, impacting attention mechanisms.
- Tokenization can cause LLMs to struggle with spelling and simple string manipulation.
- Non-English languages often perform worse due to less training data and less efficient tokenization.
- Arbitrary number tokenization (e.g., 677 as two tokens) hinders arithmetic capabilities.
- The tiktoken web app visualizes tokenization in real-time.
- GPT-2 tokenizer splits 'tokenization' into two tokens (3642, 1634).
- Spaces are often part of tokens (e.g., 'the' is token 262, ' the' is token 379).
- Case sensitivity impacts tokenization: 'egg' (2 tokens) vs. ' egg' (1 token).
- LLMs must learn semantic equivalence from token patterns despite tokenization differences.
- Non-English languages use more tokens for the same meaning, inflating sequence length.
- Python's indentation (spaces) leads to many individual space tokens (token 220).
- This inefficiency bloats sequence length, consuming context window limits.
- GPT-4 tokenizer improved Python space handling, merging multiple spaces into single tokens.
- GPT-2 tokenizer has ~50k tokens; GPT-4 tokenizer has ~100k tokens.
- Larger vocabulary allows denser representation, effectively doubling context length.
- Increasing vocabulary size impacts embedding table size and output layer computation.
- Python strings are sequences of Unicode code points (integers representing characters).
- Unicode standard defines ~150,000 characters across 161 scripts.
- The `ord()` function in Python retrieves the Unicode code point for a character.
- A Unicode vocabulary (~150,000) is too large for efficient LLM processing.
- The Unicode standard is dynamic, making it an unstable representation.
- LLMs require a fixed, tunable vocabulary size for efficient training and inference.
- UTF-8 encodes Unicode code points into 1-4 byte streams.
- Using raw UTF-8 bytes yields a vocabulary of only 256 tokens.
- A small vocabulary leads to excessively long sequences, exceeding context limits.
- BPE iteratively merges the most frequent adjacent token pairs to create new tokens.
- This process compresses byte sequences into a more manageable vocabulary.
- The goal is to balance vocabulary size, sequence length, and compression ratio.
- Start with raw UTF-8 byte sequences (0-255).
- Iteratively find the most frequent byte pair, merge it, and add a new token ID.
- The `merges` dictionary stores the learned merge rules (e.g., 101+32 -> 256).
- The number of merges determines the final vocabulary size (e.g., 256 bytes + 20 merges = 276 tokens).
- Vocabulary size is a hyperparameter, balancing density and computational cost.
- GPT-4 uses ~100,000 tokens; GPT-2 used ~50,000.
- Tokenizers are trained independently on their own datasets, separate from the LLM.
- The trained tokenizer (vocabulary and merges) translates text to/from token sequences.
- Tokenizer training data composition (languages, code) influences token density.
- Reconstruct byte sequences from token IDs using the learned vocabulary and merges.
- Decode byte sequences into UTF-8 strings.
- Handle invalid UTF-8 sequences using `errors='replace'` to avoid crashes.
- Encode text into UTF-8 bytes, then into integer token IDs.
- Iteratively apply learned BPE merges to compress the sequence.
- The process prioritizes merges with lower indices in the `merges` dictionary.
- BPE iteratively merges frequent pairs, reducing sequence length and increasing vocabulary.
- Edge cases like empty strings or single tokens require special handling.
- Encoding and decoding should ideally be inverse operations for unseen data.
- GPT-2 uses regex to pre-chunk text, preventing merges across certain character types (letters, numbers, punctuation).
- This approach enforces specific merging rules, avoiding semantic-punctuation conflation.
- The `encoder.py` file shows inference logic, but training details remain proprietary.
- tiktoken is OpenAI's official library for tokenization inference.
- It supports various tokenizers like GPT-2 and GPT-4 (CL100K).
- GPT-4's CL100K tokenizer uses a modified regex and merges spaces differently than GPT-2.
- Special tokens (e.g., `<|endoftext|>`) are added to the vocabulary for specific functions.
- `<|endoftext|>` signals document boundaries during training.
- Fine-tuned models (like chat models) use special tokens to delimit messages (e.g., `<|im_start|>`, `<|im_end|>`).
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Andrej Karpathy.