CMU Introduction To Deep Learning 11-785, Fall 2026: Lecture 14
Watch on YouTube →
Overview
The lecture explains why ordinary recurrent neural networks can represent long-range dependencies yet struggle to preserve useful memories and train across long sequences: recurrent weights and activation Jacobians can make signals and gradients vanish or explode. It motivates Long Short-Term Memory (LSTM) cells, whose gated constant-error-carousel memory is updated according to inputs and context, then introduces the forget, input, and output gates, code-based backpropagation, and the simpler Gated Recurrent Unit (GRU).
Key takeaways
- A recurrent adder can generalize to n-bit addition by reusing a local rule over two bits and a carry state, whereas an unstructured MLP may need to represent 2^(2n) input pairs.
- The magnitude of an ordinary RNN’s recurrent eigenvalues controls long-run state behavior: values above 1 can amplify signals, while values below 1 make past inputs decay.
- In standard RNNs, repeated multiplication by activation Jacobians and weight matrices makes useful memory and backpropagated gradients vulnerable to vanishing or explosion.
- An LSTM cell state uses a gated carry and gated additive update, allowing memory retention to depend on input and context rather than only on repeated application of a recurrent weight matrix.
- LSTM gates do not eliminate every training difficulty: gradients can still explode, so gradient clipping can be used to limit their magnitude.
- A GRU simplifies the LSTM design by combining forgetting and input gating and using a single state representation.
Chapters
- The opening minutes contain informal classroom conversation and equipment setup, including a request to turn on a microphone.
- A brief exchange about CMU clothing precedes the start of the technical lecture.
- At about 7:16, students are asked to turn on the attendance poll.
- The instructor requests direct email feedback about what students have and have not learned during the first half of the semester.
- A convolutional model with a seven-step window suits dependencies limited to the previous seven time points; recurrence carries state across an arbitrarily long sequence.
- Parsing nested code braces illustrates the challenge: an opening brace may need to be remembered until a matching close brace appears, potentially millions of lines later.
- An MLP that maps pairs of n-bit numbers may need to account for 2^(2n) possible input pairs, while a recurrent adder reuses the same local operation at each bit.
- The adder processes two input bits and a carry state—eight possible local combinations—so its learned rule can generalize to numbers of arbitrary length.
- Parity also needs only a compact recurrent state that flips as each bit arrives, rather than a separate MLP mapping for every n-bit input.
- A finite time-delay network with bounded activations has finitely many operations, so bounded inputs produce bounded outputs.
- An RNN reuses its recurrent transformation at every step, so repeated application can amplify or suppress its state.
- Saturating activations may avoid numerical infinity but can still destroy useful distinctions by making the network insensitive to new inputs.
- For a linear scalar recurrence, h_t = w h_(t−1) + c x_t, the response to a single input at time zero is proportional to w^t.
- When |w| exceeds 1, the response grows; when |w| is below 1, it decays toward zero; at |w| = 1, it does not decay in magnitude.
- The expanded state is a sum of responses to inputs at earlier times, making the recurrent weight determine how quickly past inputs fade or amplify.
- For a vector hidden state, repeated recurrence applies powers of the square weight matrix W to earlier states.
- The lecture uses an eigenvalue decomposition to show that the largest eigenvalue magnitude controls long-run growth or decay: above 1 can grow, below 1 decays.
- Complex eigenvalues can create oscillations while the overall envelope still grows or shrinks according to eigenvalue magnitude.
- In the illustrated sigmoid examples, differences between starting inputs largely disappear after roughly 20 recurrent steps as the state saturates.
- Tanh can retain distinctions longer than sigmoid in the examples, but eventually its state also becomes dominated by learned parameters rather than the original input.
- ReLU recurrence can rapidly grow or die away, so none of these standard activations alone supplies reliable, input-dependent long-term memory.
- A deep network is a nested function, so backpropagation applies the chain rule through alternating weight matrices and activation Jacobians.
- For elementwise recurrent activations, the Jacobian is diagonal; ReLU derivatives are 0 or 1, tanh derivatives are at most 1, and sigmoid derivatives are at most 0.25.
- Repeated Jacobians with entries no greater than 1 tend to shrink gradients, while the weight matrices can amplify or suppress them.
- A weight matrix can be decomposed as W = U S Vᵀ, where the orthonormal matrices rotate vectors and the diagonal singular-value matrix S changes their lengths.
- Directions associated with singular values below 1 shrink, while directions above 1 expand; repeated multiplication and intervening rotations can repeatedly redirect gradients.
- Because most singular values in the lecture’s general account are below 1, gradients tend statistically to vanish in many directions, though some may grow.
- A 19-layer MNIST network with 1,024 units per layer illustrates derivatives fading as backpropagation moves toward earlier layers.
- The plotted ReLU, ELU, sigmoid, and tanh examples show gradients becoming faint after only a few layers; ELU lasts longest in the displayed comparisons.
- The lecture notes that gradients need not all vanish: their magnitude can become concentrated in a small number of components while most others flatten.
- Nested code parsing requires carrying information from an opening brace forward until its matching close brace appears.
- Learning the relationship also requires the later closing-brace error to affect parameters responsible for recognizing the earlier opening brace.
- Ordinary RNNs therefore face two linked problems over long sequences: state information can fade forward, and learning signals can fade backward.
- The proposed design removes recurrent weights and nonlinear activations from the direct memory path, reducing parameter-driven decay or explosion along that path.
- Gates inspect the current input and context to decide whether to preserve, erase, or add information to memory.
- Context matters: a brace-like input should update memory differently if it occurs inside a comment block or if the relevant opening brace is already remembered.
- An LSTM separates its cell state C from its exposed hidden state h, giving memory a dedicated path through time.
- The constant error carousel updates C through a gated multiplicative carry and a gated additive candidate, rather than repeatedly applying the ordinary RNN transformation.
- The lecture compares the cell state to a count of open braces: a closing brace can decrement the count, while unrelated inputs leave it intact.
- The forget gate uses a sigmoid output to control how much of the previous cell state is retained.
- A tanh candidate detects potential new content, while the input gate controls how much of that candidate is added to the cell.
- Both gate calculations use the current input and previous hidden state; optional peephole connections also expose the previous cell state to gates.
- After updating C, the LSTM applies tanh to the cell state and scales it with a sigmoid output gate to produce h.
- The output gate controls what part of stored memory is exposed as the hidden state for later time steps and deeper network layers.
- The lecture describes the architecture as more parameter-rich than a basic RNN, with memory retention controlled by learned gates responding to input and context.
- The forward computation is organized as gate affine transforms and activations, cell-state update, output gate, and hidden-state calculation.
- Rather than expand every derivative path by hand, backpropagation reverses the forward operations and applies local derivatives through each step.
- Students are directed to the course handout and slides because LSTM derivative calculations will appear in homework.
- A Gated Recurrent Unit (GRU) combines the forgetting and input decisions and avoids keeping separate raw-memory and transformed-memory copies.
- LSTM’s direct cell path avoids the ordinary RNN’s guaranteed repeated decay mechanism, but gradients can still explode; gradient clipping is one way to control large gradients.
- The closing summary contrasts ordinary RNN memory’s dependence on weights and activations with LSTM memory updates driven by gated input patterns, and previews further topics for the next class.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Carnegie Mellon University Deep Learning.