Save this video — free

But what is cross-entropy? | Compression is Intelligence Part 2

3Blue1Brown · 33:51 · Watch on YouTube

But what is cross-entropy? | Compression is Intelligence Part 2 Watch on YouTube →

Overview

3Blue1Brown explains cross-entropy, a core concept in information theory and machine learning, by connecting it to file compression and language model training. The video demonstrates how cross-entropy quantifies the inefficiency of using a compression scheme optimized for one probability distribution (Q) when applied to another (P), mirroring how language models are trained to minimize the difference between their predicted token distributions and the true data distribution. This principle is applied to language tree discovery via file compression and forms the basis for training modern language models, with distillation offering a richer training signal by comparing model distributions.

Key takeaways

Chapters

0:00 Compression as a Tool for Discovering Linguistic Structure
2:34 Introduction to Cross-Entropy and Its Role in Language Models
5:02 Information Content and Optimal Encoding
8:41 Defining Cross-Entropy: Performance of One Code on Another Distribution
21:48 Cross-Entropy as a Measure of Pattern Difference
25:00 Cross-Entropy Loss in Language Model Training

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, 3Blue1Brown.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.