Save this video — free

The GELU, SiLU and SwiGLU activation functions, clearly explained!!!

StatQuest with Josh Starmer · 24:18 · Watch on YouTube

The GELU, SiLU and SwiGLU activation functions, clearly explained!!! Watch on YouTube →

Overview

Josh Starmer explains the evolution of activation functions in neural networks, moving from the limitations of sigmoid and ReLU to the more advanced GELU, SiLU, and SwiGLU. These newer functions, derived from concepts like input-dependent dropout, offer smoother curves and better generalization, reducing overfitting compared to ReLU. GELU and SiLU are closely related, while SwiGLU introduces additional trainable parameters for greater flexibility.

Key takeaways

Chapters

0:00 Introduction to GELU, SiLU, and SwiGLU vs. ReLU
5:27 Evolution from Sigmoid to ReLU and the Problem of Overfitting
18:37 Deriving GELU and SiLU from Input-Dependent Dropout

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.