Save this video — free

RL @ ECE-UofT - Lecture 05: Deep Bootstraping and epsilon-Greedy Improvement

Ali Bereyhi · 2:42:21 · Watch on YouTube

RL @ ECE-UofT - Lecture 05: Deep Bootstraping and epsilon-Greedy Improvement Watch on YouTube →

Overview

Ali Bereyhi develops n-step temporal-difference learning as a bridge between one-step TD and Monte Carlo, then introduces TD(λ), which geometrically combines returns of different depths to balance estimation bias and variance. He derives eligibility traces as an online, backward-time implementation of TD(λ), then shifts from policy evaluation to control, explaining why greedy improvement can prevent exploration and how ε-greedy action selection addresses that problem.

Key takeaways

Chapters

0:00 Model-Free Reinforcement Learning: Policy Evaluation and Improvement
5:07 Monte Carlo and TD(0): Complete Returns Versus One-Step Targets
6:54 N-Step TD Extends the Bootstrap Horizon
13:00 Why Deeper Bootstrapping Can Improve State-Value Estimates
20:01 N-Step Returns and the Standard Value-Update Rule
25:20 Applying N-Step TD Along a Sampled Trajectory
29:32 N-Step TD Connects TD(0) to Monte Carlo
35:32 Choosing N: The Bias–Variance Trade-Off
40:40 Why One Fixed N Is Hard to Tune
45:21 TD(λ) Geometrically Weights Returns of Different Depths
52:03 λ-Returns Reuse Each Trajectory at Multiple Horizons
56:01 TD(λ) Updates Values Using the λ-Return
58:27 λ=0 and λ=1 Recover TD(0) and Monte Carlo
1:06:47 Forward-View TD(λ) Discounts Deeper Predictions
1:13:02 Credit Assignment: Weighing the Reliability of Future Outcomes
1:21:00 Why TD(λ) Needs a Backward, Online View
1:25:00 Eligibility Traces Track Recently Visited States
1:31:41 Initializing and Updating the Eligibility Table
1:33:42 Repeated Visits Accumulate State Eligibility
1:41:23 Use TD Errors to Update Every Eligible State
1:48:48 Eligibility Traces Make TD(λ) Updates Online
1:54:33 Forward and Backward TD(λ) Views Agree at Episode End
1:58:44 Eligibility-Trace Limits Recover TD and Monte Carlo
2:04:30 Prediction Algorithms Versus Online Control
2:09:05 Control Updates Values and Policy During Interaction
2:16:20 Monte Carlo Control Improves After Each Episode
2:18:22 Greedy Improvement Can Lock a Learner Into a Bad Action
2:25:06 Company A and Company B Show the Exploration Problem
2:29:34 ε-Greedy Improvement Balances Exploration and Exploitation
2:32:47 Action Probabilities Under ε-Greedy Selection
2:37:04 The ε-Greedy Improvement Guarantee
2:39:57 Next Steps: SARSA, Q-Learning, and On- versus Off-Policy Learning

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Ali Bereyhi.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.