Save this video — free

UUtah F2026 | Data Mining | L6 - Similarities

UofU Data Science · 1:20:18 · Watch on YouTube

UUtah F2026 | Data Mining | L6 - Similarities Watch on YouTube →

Overview

The lecture develops similarity as a way to score how alike objects are, then uses Jaccard similarity to compare sets of text shingles. It explains how k-gram choices shape document representations and introduces MinHash, whose key property is that two documents’ hash values collide with probability equal to their Jaccard similarity—a foundation for scaling similarity search with locality-sensitive hashing.

Key takeaways

Chapters

0:00 Project Proposal Requirements and Data Mining Project Scope
9:00 Homework 2: GloVe Embeddings and Approximate Similarity Search
12:00 Similarity Scores Reverse the Direction of Distance
17:00 Converting Similarities into Distances
23:00 Similarity as a Modeling Choice and the Limits of Approximation
27:00 Jaccard Similarity Measures Set Overlap
35:00 Bag-of-Words Vectors Count Terms but Ignore Word Order
41:00 K-Grams and Shingles Preserve Local Word Context
45:00 Text-Shingle Representations Require Explicit Modeling Choices
55:00 Character K-Grams Extend the Method Beyond Word Tokenization
57:00 Comparing Four Dr. Seuss Lines with Jaccard Similarity
1:02:00 Worked Jaccard Scores Show Overlap and Set Size Together
1:04:00 Jaccard Sets Avoid a Fixed Vocabulary
1:09:00 Scaling Similar-Document Search Beyond Brute Force
1:14:00 MinHash Targets Jaccard Similarity Through Collision Probability
1:18:00 The MinHash Minimum Enables One-Pass Document Processing

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, UofU Data Science.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.