Save this video — free

Reinforcement Learning with Neural Networks: Mathematical Details

StatQuest with Josh Starmer · 25:01 · Watch on YouTube

Reinforcement Learning with Neural Networks: Mathematical Details Watch on YouTube →

Overview

Josh Starmer of StatQuest details the mathematical underpinnings of training a neural network using reinforcement learning via the policy gradients method. He walks through calculating the derivative of cross-entropy with respect to a bias term, using the chain rule to combine derivatives of the sigmoid activation function and the cross-entropy loss, and then updating the bias based on a reward signal derived from the outcome of an action.

Key takeaways

Chapters

0:00 Introduction to Reinforcement Learning for Neural Networks
2:29 Calculating Initial Probabilities and Making a Choice
5:05 Quantifying Error with Cross-Entropy and Derivative Calculation
13:59 Derivatives of Sigmoid and Input to Bias
19:16 Updating the Bias with Rewards and Gradient Descent

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.