StatQuest: Random Forests Part 2: Missing data and clustering
Watch on YouTube →
Overview
Josh Starmer's StatQuest on Random Forests Part 2 details two methods for handling missing data: imputation during model training and classification of new samples. For training data, missing values are initially guessed and iteratively refined using proximity matrices derived from running data through random forest trees. For new samples, two versions are created (one for each potential outcome of the missing feature) and run through the forest to determine which best aligns with the model's predictions.
Key takeaways
- Random forests handle missing data in training sets by iteratively refining initial guesses using proximity measures derived from tree leaf node assignments.
- Proximity matrices, calculated by tracking co-occurrence in leaf nodes across trees, can be transformed into distance matrices for visualization (heatmaps, MDS plots).
- For classifying new samples with missing features, two versions of the sample are created (one for each potential value of the missing feature) and run through the forest to determine the most likely classification.
- The iterative refinement of missing values in training data continues until the imputed values stabilize across iterations.
- The proximity of samples in a random forest is defined by how often they end up in the same leaf node across all trees in the forest.
Chapters
- Missing data in the original training set requires an initial guess for categorical (most common value) and numeric (median value) features.
- Proximity matrices are built by running data down random forest trees, counting how often samples land in the same leaf node.
- Proximity values are normalized by the total number of trees and used to refine guesses for missing categorical (weighted average proximity) and numeric (weighted average) values.
- This imputation process is repeated iteratively until missing values converge.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.