Save this video — free

UUtah Data Mining | Fall 2026 | L2 - Statistical Phenomenon

UofU Data Science · 1:19:21 · Watch on YouTube

UUtah Data Mining | Fall 2026 | L2 - Statistical Phenomenon Watch on YouTube →

Overview

The lecture introduces probabilistic thinking for data mining: data is often modeled as independent, identically distributed (IID) samples, but random variation alone can produce surprising patterns. Using uniform hashing and birthdays, it explains why collisions appear around √m samples, then derives the coupon collector’s mHₘ ≈ m ln m expected samples for seeing every one of m outcomes.

Key takeaways

Chapters

0:00 Course Website, Lecture Materials, and Homework 1
8:20 Evaluating Probabilistic Code Through Plots
11:30 IID Samples as a Foundation for Data Analysis
14:00 Sample Size, Estimation, and the Central Limit Theorem
17:10 Uniform Samples from a Finite Universe
21:10 Hash Tables, Uniform Hashing, and Random Salts
27:12 Birthday Paradox Demonstration: A Collision After 38 People
34:00 Birthday Collision Probability and Pair Counting
44:10 Why Hash Collisions Begin Around √m Samples
49:00 Birthday-Model Assumptions and Where Approximations Fail
1:01:10 Coupon Collector: How Many Draws to See Every Outcome?
1:06:30 Breaking Complete Coverage into New-Item Waiting Times
1:09:40 Harmonic Numbers Give the Coupon Collector Expectation
1:12:00 Why the Final Coupon Makes Collection Take m log m
1:15:40 Use Collision and Coverage Baselines to Interpret Data

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, UofU Data Science.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.