UUtah CS 6340 NLP | Fall 2026 | Pretraining & finetuning
Watch on YouTube →
Overview
Pretraining gives transformer models useful, transferable language patterns before task-specific fine-tuning; the lecture contrasts decoder-only models trained with next-token prediction against encoder-only models trained with masked language modeling. It also explains how attention masks, language-model and classification heads, and data choices shape model capabilities, then connects scaling laws to practical risks in corpus quality, bias, privacy, multilingual coverage, copyright, and benchmark contamination.
Key takeaways
- Decoder-only pretraining uses causal attention and next-token prediction; source and target text can share one sequence, with target tokens attending to the preceding source.
- Encoder-only models such as BERT use bidirectional attention and roughly 15% masked-token prediction, making them well suited to classification, tagging, and retrieval rather than free-form generation.
- A decoder-only model can be fine-tuned with its existing vocabulary-level LM head, whereas an encoder classifier usually needs a newly initialized task-specific head that must be trained.
- Chinchilla scaling recommends increasing model parameters and training tokens in equal proportions as compute grows, rather than scaling model size much faster than data.
- Pretraining-data quality is an engineering and financial constraint: malformed repetitive batches can destabilize training, and a processing bug can waste hundreds of thousands of dollars or more.
- Corpus design involves competing goals: filtering can reduce harmful content but may erase marginalized voices, while multilingual balancing must avoid both underrepresenting low-resource languages and overfitting their limited examples.
Chapters
- Attention lets a token compare its query with keys across long contexts, helping transformers model distant relationships more effectively in practice than RNNs.
- Contextual token representations support word-sense disambiguation, while rotary positional embeddings (RoPE) encode relative token positions.
- Sebastian Raschka's model gallery illustrates newer components such as RMSNorm, gated attention, partial RoPE, and mixture-of-experts layers in models including GLM-5 and Qwen 3.5.
- Transformer models have many parameters: the original 2017 model had about 65 million parameters in its small version, while modern models can reach hundreds of billions.
- Pretraining uses large text or code collections and self-supervised objectives such as next-token prediction to learn general linguistic patterns without manual labels.
- Fine-tuning specializes pretrained weights with labeled examples; simply prompting a model does not change its weights and is not fine-tuning.
- Pretraining followed by fine-tuning is one form of transfer learning: knowledge from a source task helps a target task.
- The lecture organizes model design around transformer type, pretraining objective, and the downstream task head.
- Encoder-only models were prominent for classification and sequence labeling from roughly 2018–2021; generative AI shifted attention toward text generation after 2022.
- The next-token or masked-token objective and the chosen output head determine how pretrained representations are adapted to a task.
- A decoder-only transformer removes the encoder stack and cross-attention, retaining causal self-attention, feed-forward layers, and an LM head.
- For translation, source tokens and target tokens can be concatenated in one sequence with a separator; causal attention lets target positions use earlier source tokens.
- Unlike an encoder, causal attention cannot let source tokens see later source tokens, so source representations are less expressive—but the simpler stack scales efficiently.
- Decoder-only pretraining trains on large, quality-filtered corpora by maximizing the probability of each observed next token with negative log-likelihood loss.
- Teacher forcing shifts tokenized text by one position, adds a beginning-of-sequence token, and computes predictions for all positions in a single causal-attention forward pass.
- Losses are averaged across sequence positions and batch examples, then backpropagated to update the model.
- A pretrained decoder-only model can continue training on a specialized dataset, such as human translations for a low-resource language.
- The existing LM head maps the final token representation to vocabulary logits; softmax produces the next-token distribution used for the same negative log-likelihood objective.
- The model can retain all its existing parameters and update them using task data; Hugging Face provides `AutoModelForCausalLM` for causal language modeling.
- Encoder-only transformers use bidirectional self-attention, making them suitable for tasks that can inspect the full input, such as classification and token tagging.
- Masked language modeling replaces about 15% of input tokens with a mask token and trains the model to recover the original tokens from their context.
- Because encoder attention can see tokens on both sides, ordinary next-token prediction would leak future information; BERT popularized masked-token pretraining in 2018.
- BERT-style inputs can begin with a CLS token and use separator tokens to distinguish paired inputs, such as two sentences for a relationship judgment.
- The final CLS representation summarizes the sequence and can support classification or similarity-based retrieval, including question-to-document matching.
- Fine-tuning replaces the masked-language-model vocabulary projection with a randomly initialized classification matrix sized to the number of classes; that new head must be trained.
- Decoder-only classifiers can instead use the final input-token representation, while encoder-only classifiers commonly use CLS.
- BERT substantially improved on the GLUE benchmark: its reported average rose to about 82 from a GPT-1 baseline around 75.
- On SQuAD reading comprehension and common-sense tasks, BERT also outperformed earlier task-specific baselines.
- Interpretability research found that BERT representations implicitly encode linguistic structure such as part of speech, syntax, and named entities.
- Encoder models such as ModernBERT remain useful for classification and retrieval, where a strong sequence representation may be more suitable than a generative LLM.
- The lecture contrasts decoder-only GPT and Llama models, encoder-only BERT and RoBERTa, and encoder-decoder T5.
- T5 combines span corruption with sequence generation; the lecture treats decoder-only models as the more relevant default for current general-purpose LLMs.
- A companion Colab demonstrates teacher forcing and source-versus-target token handling, and shows how classification can be phrased as generating labels such as “positive” or “negative.”
- Pretraining corpora combine broad web data with sources such as GitHub code and smaller amounts of higher-quality text to provide linguistic patterns and world knowledge.
- The lecture uses AI2's DOLMA as an example of an openly documented corpus effort, including data sources and curation practices.
- Corpus sizes have grown from BERT's roughly 3 billion words to GPT-3's 300 billion tokens and modern corpora containing trillions of tokens.
- Scaling laws estimate language-model loss from model size, available training data, and compute budget, which developers can approximate from hardware resources.
- Kaplan et al. (2020) recommended increasing model size faster than training tokens as compute grew.
- Hoffmann et al. (2022), introducing Chinchilla, argued that model parameters and training tokens should scale in equal proportions, showing that smaller models trained on more data can match larger ones.
- Guest researcher Kyle Lo's “microwave gang” example shows Reddit posts with repetitive “beep” replies triggering loss spikes during pretraining.
- Other data or processing bugs can make training loss rise irrecoverably or erase performance on a benchmark such as MMOU.
- A single data-processing error can waste hundreds of thousands of dollars; the lecture cites possible costs up to $6 million for larger failed runs.
- Web corpora contain social bias and harmful content; removing every toxic example can also reduce opportunities for models to learn to recognize and avoid harmful outputs.
- A “naughty words” filter that removed LGBTQ-related terms could erase educational and community-produced material along with abusive content.
- Data curation also raises labor concerns when workers review disturbing material, including pornography, under potentially harmful working conditions.
- Misinformation is another corpus risk; TruthfulQA evaluates whether models repeat common falsehoods such as the claim that Earth is flat.
- English dominates web text despite thousands of languages worldwide; the lecture notes that many languages lack even a general language-modeling corpus.
- Multilingual C4 (mC4) and newer corpora broaden language coverage, but low-resource languages need oversampling to compete with English.
- Oversampling the same scarce examples too often can cause memorization, while reducing high-resource data too much can hurt performance in those languages.
- Tokenizer training data also affects how efficiently different scripts and languages are represented.
- Research showed that GPT-2 could reproduce personal information from web data when prompted with a suitable prefix, making privacy filtering important.
- Whether training on copyrighted material qualifies as fair use remains contested, particularly if generated books or other outputs substitute for creators' work.
- Public benchmark answers can enter training data through sources such as GitHub, undermining train-test separation and making it hard to know whether a model solved a task or memorized it.
- HTML artifacts and parsing errors also require cleanup, though the lecture presents them as a smaller concern than privacy, copyright, and contamination.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, UofU Data Science.