UUtah | Data Mining | Fall 2026| L3 - Embeddings
Watch on YouTube →
Overview
The lecture traces word representations from hand-built structures and sparse co-occurrence counts to learned, contextual embeddings. It explains how distributional context, pointwise mutual information, Word2Vec and GloVe produce word vectors, why static vectors struggle with ambiguous words such as “bank,” and how ELMo and transformer-based models use context to create more flexible representations.
Key takeaways
- The distributional hypothesis gives a practical route from raw text to meaning: words that appear in similar contexts can be represented by similar vectors.
- Positive PMI assigns a context coordinate a positive value when two words co-occur more often than the independence baseline P(i)P(j) predicts, while setting negative scores to zero.
- Word2Vec and GloVe made it practical to use compact embeddings—often 50–300 dimensions—instead of sparse co-occurrence vectors with roughly 100,000 coordinates.
- Static embeddings cannot distinguish the financial and river meanings of “bank” because both uses share one vocabulary vector; ELMo and BERT address this by conditioning representations on context.
- Transformer attention can select relevant information from long text histories, including earlier sentences, rather than limiting a representation to a small fixed context window.
- Embedding methods are useful beyond language: learned vectors for images, proteins, and genomes allow many of the same similarity and data-mining techniques to be applied.
Chapters
0:00
Course Logistics: Project Groups and the Randomized-Algorithms Homework
- Students are encouraged to form project groups with classmates and consult the project document as details approach.
- The homework is due September 8 and asks students to simulate the birthday paradox and coupon collector process.
- Because simulations are random, students should compare observed rates with theoretical expectations rather than rely only on ordinary unit tests.
- The assignment also emphasizes efficient algorithms and data structures; a poorly structured coupon collector simulation can take hours to run.
9:10
From Text to Ordered, High-Dimensional Word Vectors
- The lecture distinguishes data points that are individual words or tokens from data points that are whole documents.
- The goal is to map each word to a fixed-length vector, such as a 300-dimensional array of real-valued coordinates.
- Coordinate order matters: comparisons pair coordinate 1 with coordinate 1, coordinate 2 with coordinate 2, and so on.
- Once words are represented as vectors, distances, dot products, and other similarity measures can compare them.
15:00
From WordNet and Dictionaries to the Distributional Hypothesis
- Earlier approaches compared words through sound, spelling and character features, dictionary order, or manually assembled synonym resources.
- WordNet organizes roughly 100,000 English words in a semantic hierarchy, such as dogs and cats as related kinds of animals.
- These hand-built systems can encode useful relationships but are awkward representations for data mining and machine learning.
- The distributional hypothesis, associated with linguists including Zellig Harris, says that words used in similar contexts tend to have similar meanings.
21:00
Vector Arithmetic Reveals Word Relationships and Analogies
- In a 300-dimensional embedding, the difference between the vectors for “king” and “queen” can encode a direction associated with gender.
- Adding that direction to the vector for “man” can place the result near “woman,” illustrating the analogy king:queen :: man:woman.
- Other vector differences can capture relations such as royalty; adding a similar direction to “woman” may point toward “queen.”
- These relationships are emergent properties of learned embeddings, not necessarily rules explicitly programmed into the model.
31:30
Corpus Context Windows Turn Usage into Training Evidence
- A corpus is a text collection containing a total of n word occurrences, potentially extending to trillions of tokens.
- To represent “octopus,” the lecture gathers examples of the word in text and defines each context using three words on either side.
- Words such as “arms,” “shipwreck,” and “cave” become evidence about the contexts in which “octopus” appears.
- This operationalizes the distributional hypothesis: words with overlapping context patterns should receive similar representations.
35:30
Building Sparse Co-Occurrence Vectors from Nearby Words
- For each target word j, a co-occurrence vector records how often each vocabulary word i appears in j’s context.
- With a vocabulary of 100,000 words, the resulting vector can have 100,000 dimensions, though most coordinates are zero.
- The lecture defines word frequency from corpus counts and joint frequency from counts of words occurring together in context.
- Sparse storage helps with the many zero entries, but the high dimensionality remains costly for downstream methods.
39:50
Pointwise Mutual Information Scores Informative Word Associations
- Pointwise mutual information compares joint occurrence, P(i,j), with the independent baseline P(i)P(j).
- The score is log(P(i,j) / (P(i)P(j))); positive values indicate that two words co-occur more than independence predicts.
- Positive PMI applies max(0, PMI), setting negative associations to zero and retaining words that occur together unusually often.
- The resulting coordinates weight context words by association strength instead of treating every observed co-occurrence equally.
44:00
Document Vectors, TF-IDF, and the Long-Running Utility of BM25
- TF-IDF uses term frequency and inverse document frequency to emphasize words that matter in a document but are unusual across the collection.
- BM25 extends this kind of weighting for document retrieval and remains competitive on many information-retrieval benchmarks.
- The co-occurrence approach provides comparable context-based representations for individual words, but its 100,000-dimensional vectors are unwieldy.
- These sparse representations establish a useful baseline while motivating lower-dimensional learned embeddings.
46:00
Self-Supervised Learning Compresses Context into Word Embeddings
- Instead of storing a 100,000-dimensional context vector, learned methods aim for compact embeddings, often around 50–300 dimensions.
- Training uses text itself as supervision: examples of a word and its surrounding context provide signals for learning representations.
- Word2Vec, released by Google around 2013, and GloVe, developed at Stanford around 2014, helped establish learned word vectors as a default NLP tool.
- Pretrained embeddings offered vectors for large vocabularies and enabled vector arithmetic and improvements across language-processing tasks.
55:00
Next-Word Prediction and Neural Networks Learn from Text Sequences
- A next-word model uses preceding words to predict the following word; earlier n-gram approaches estimated probabilities from short histories such as two words.
- A neural-network approach represents context words as vectors and learns parameters that help predict or associate the target word.
- The lecture describes recurrent models such as LSTMs, which were used to handle sequences before attention-based architectures became dominant.
- Training can update both word representations and network weights through backpropagation as the model processes many context windows.
1:00:00
Static Embeddings Give Ambiguous Words Only One Vector
- In static Word2Vec- or GloVe-style embeddings, each vocabulary entry receives one fixed vector regardless of its sentence.
- The word “bank” therefore shares one representation whether it means a financial institution or the edge of a river.
- Contexts for the two meanings are mixed during training, limiting how precisely the embedding can represent either usage.
- This limitation motivated methods that produce a word representation from both the word and its context.
1:03:00
ELMo Makes Word Representations Depend on Sentence Context
- ELMo, associated with the Allen Institute for AI and University of Washington researchers, represents a word using a function of the word and its context.
- A contextual model can map “bank” differently in “stored money in the bank” and “swam to the river bank.”
- ELMo used bidirectional sequence context, allowing information from both sides of a word to shape its representation.
- Contextual embeddings improved NLP systems while requiring downstream tasks to use a model-produced representation rather than a single stored vector per word.
1:07:00
BERT and Transformer Models Build Contextual Representations
- BERT uses transformer-based contextual representations; the transformer architecture was introduced separately in the 2017 paper “Attention Is All You Need.”
- Unlike static embeddings, transformer layers update a token’s representation using information from the surrounding sequence.
- The lecture connects these contextual representations to language-model systems such as GPT, whose name includes “Transformer.”
- Different internal layers can encode different information: intermediate layers may capture richer context, while later layers can become more specialized for a training objective.
1:10:00
Attention Selects Relevant Context Across Longer Sequences
- Attention allows a model to assign greater weight to relevant tokens instead of relying only on a small, fixed window such as three words on either side.
- A token can draw useful information from earlier sentences or paragraphs when predicting or representing the next word.
- Transformer layers repeatedly use attention to combine contextual information, supporting longer-range dependencies more directly than older recurrent approaches.
- The lecture notes that practical transformer embedding dimensions can be much larger than the 300 dimensions common in early word vectors.
1:14:00
Using GloVe and BERT Embeddings in Python and Beyond Text
- The lecture shows that a pretrained GloVe vector for a word such as “bank” can be retrieved with a small amount of Python code.
- A BERT-based workflow can produce a contextual representation of “bank” using the surrounding phrase, such as a sentence involving “tax.”
- For a later homework, students will use static word vectors to explore similarity and distance before working with more complex contextual embeddings.
- Vector representations also underpin applications involving images, proteins, and genomes, where learned vectors enable downstream data-mining methods.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, UofU Data Science.