Save this video — free

StatQuest: Random Forests Part 2: Missing data and clustering

StatQuest with Josh Starmer · 10:48 · Watch on YouTube

StatQuest: Random Forests Part 2: Missing data and clustering Watch on YouTube →

Overview

Josh Starmer's StatQuest on Random Forests Part 2 details two methods for handling missing data: imputation during model training and classification of new samples. For training data, missing values are initially guessed and iteratively refined using proximity matrices derived from running data through random forest trees. For new samples, two versions are created (one for each potential outcome of the missing feature) and run through the forest to determine which best aligns with the model's predictions.

Key takeaways

Chapters

0:00 Handling Missing Data During Random Forest Training

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.