Save this video — free

The matrix math behind transformer neural networks, one step at a time!!!

StatQuest with Josh Starmer · 23:43 · Watch on YouTube

The matrix math behind transformer neural networks, one step at a time!!! Watch on YouTube →

Overview

StatQuest with Josh Starmer breaks down the matrix mathematics underpinning Transformer neural networks, focusing on an encoder-decoder architecture for translation. The explanation details how word embeddings, positional encodings, self-attention (query, key, value matrices, dot products, scaling, softmax), residual connections, and encoder-decoder attention are computed using matrix operations, culminating in the final output probabilities.

Key takeaways

Chapters

0:00 Input Processing: Word Embeddings and Positional Encoding
7:11 Self-Attention Calculation: Queries, Keys, Values, and Similarity
18:23 Decoder Operations: Masked Self-Attention and Encoder-Decoder Attention

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.