Try it free

← All courses

University of Utah CS 5140/6140 Data Mining

Professor Jeff Phillips · University of Utah · 10 lectures with notes

Students in this class: ask your lecturer for the class code, and these lectures will already be in your library when you sign up.

UUtah Data Mining | Fall 2026 | L13 : Streaming Freq Apx

7 Oct 2026

Misra–Gries and Count-Min Sketch estimate frequent items in a stream with small memory and additive error guarantees.

UUtah Data Mining | Fall 2026 | L12: Steaming & Sampling

2 Oct 2026

Reservoir and priority sampling maintain useful uniform or weighted samples from a stream using limited memory.

UUtah Data Mining | Fall 2026 | Choosing k (#clusters)

30 Sep 2026

Use plots, elbow curves, silhouette scores, and—when a likelihood model is justified—BIC to choose a useful, not universally correct, number of clusters.

UUtah Data Mining | Fall 2026 | L10: Spectral Clustering

25 Sep 2026

Spectral clustering uses normalized graph Laplacians and their low-eigenvalue embeddings to find balanced graph partitions.

UUtah Fall 2026 | Data Mining | L9 - k-Means and friends

23 Sep 2026

Lloyd’s algorithm minimizes squared-Euclidean k-means cost locally; k-means++ makes its initialization substantially more reliable.

UUtah Data Mining | Fall 2026 | L8: Hierarchical Agglomerative Clustering

18 Sep 2026

HAC linkage choices determine whether clustering favors compact groups, distant boundaries, or connected shapes.

UUtah Data Mining | Fall 2026 | L7 - LSH & Distribution Dist

16 Sep 2026

LSH uses randomized hash collisions and banding to retrieve approximate neighbors, while distribution distances encode different modeling assumptions.

UUtah F2026 | Data Mining | L6 - Similarities

11 Sep 2026

Jaccard similarity compares sets of text shingles, while MinHash estimates that similarity through randomized hash collisions.

UUtah Fall 2026 | Data Mining | L5 NN Search

9 Sep 2026

HNSW combines layered neighbor graphs and beam search to make approximate nearest-neighbor retrieval practical in high-dimensional vector databases.

UUtah Fall 2026 | Data Mining | L4 - Metric Distances

4 Sep 2026

Choosing a distance function defines similarity in data mining and can change which points an algorithm treats as neighbors.

Keep the lectures you learn from

Paste a video, playlist, or channel URL — get transcripts, AI summaries with clickable timestamps, and search across everything.

Get started — it's free