Lesson 5 of 5
Vanishing and exploding gradients
The gradient from k steps back is a product of k factors built from the same recurrent weights, so it shrinks or grows geometrically. Clipping tames the growth; nothing simple fixes the shrinking, and a plain RNN fails a memory test once the delay grows beyond a short range.
Follow a gradient back through a scalar RNN hop by hop, computing the factor for each hop.
Show that the gradient from steps back is a product of Jacobians, and bound its size with the largest singular value of the recurrent matrix.
Explain why saturated states block the gradient, and why a recurrent weight above 1 does not guarantee growth.
Explain why clipping tames exploding gradients but nothing so simple fixes vanishing ones, and describe truncated backpropagation through time.
Show with a delayed-recall experiment that a plain RNN fails once the delay grows beyond a short range.

The character RNN from the previous lesson learned to beat the bigram model, and its samples look like Shakespeare line by line. But little carries over from one line to the next. Is that the model's size, the training, or something deeper?
Something deeper. Backpropagation through time gets a loss's gradient to every earlier state, but on the way back the gradient is multiplied at every step by nearly the same factor. That is the repeated multiplication of powers, stability and explosion, and it makes the signal from far back either vanish or explode. This lesson follows one gradient back step by step, bounds its size in general, looks at what can be done about it, and ends with an experiment where a plain RNN fails a simple memory test. That failure is the reason for the LSTM of module 8.
Follow one gradient back
Start with the smallest case, where every factor is a single number. In a scalar RNN, , the rule from the previous lesson for the gradient arriving from step becomes
when affects the loss only through (it has no output loss of its own). In words: to go back one step, multiply by the recurrent weight and by the tanh slope at the state being left. Call the factor of that hop.
The diagram below applies this rule over four steps. The RNN has , and and reads the pulse , the fading memory from the first lesson. The loss depends only on the last state, . Boxes hold the states, amber arcs carry the gradient back one hop at a time with each hop's factor written on it, and a chart underneath collects the gradient after 0, 1, 2 and 3 hops on a log scale.
Now the general scalar case. Follow one loss term back steps. Each hop contributes its own factor, and the chain rule multiplies them:
The product runs over the states the gradient leaves on its way back, down to . In the diagram, and : the product is , the gradient at . If the state hovers around some value , every factor is about , and the product of of them is . Some numbers:
- and states near : . After 10 steps the gradient is of its size, after 20 steps .
- and states near : . Slow decay, but decay: and . A weight above 1 is not enough to keep the gradient alive, because the tanh slope is below 1.
- with the state latched at (the first lesson's memory): . A saturated state is a flat part of the tanh, and the gradient through it dies within a few steps, as step 8 of the diagram showed for . The very mechanism that held a memory in the forward pass blocks the learning signal in the backward pass.
The tanh slope is never more than 1, so each factor has size at most . The gradient can only grow if , which needs and states close enough to 0 that the tanh is steep there.
Each step back is one more pane: the same fraction gets through every time, so the signal from far back fades geometrically.
Walk three hops yourself, through states that are not all alike. One hop grows the gradient and one shrinks it hard; see which and why.
A scalar RNN has , and its loss depends only on the last state, , so . The forward pass recorded , and (the inputs that produced them do not matter here).
- Compute the factor of each hop: from to , from to , and from to .
- Compute , and by multiplying the factors in turn.
- Which hop makes the gradient larger, which shrinks it the most, and why?
- Compare with the bound .
To check your work, enter three numbers as a column: , and , to 4 decimals.
Work it on real paper: writing each step is the point. Then check your final answer here and compare your working with the walk-through.
One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.
When the states stay in one region, one number, , decides everything. Try it for a weight above 1.
A scalar RNN has , and during a long stretch of text its state stays near . By roughly what factor is a gradient multiplied on its way back 20 steps through that stretch? Give 3 decimal places.
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
A product of Jacobians
A real RNN has many hidden units, so each factor is a matrix, not a number. The story is the same, with matrices. Between and lie recurrent steps, and each step is a function from the previous state to the next. Its Jacobian, the matrix of partial derivatives , is
Here is the diagonal matrix with the entries of on its diagonal, so is the recurrent matrix with row scaled by the tanh slope of unit . To see it, differentiate with respect to : the chain rule gives the tanh slope times the inner derivative . With one unit, is the number from the previous section.
The chain rule is multiplication of Jacobians, so
with (a diagonal matrix is its own transpose). The second formula is the gradient version. By the chain rule, is the transpose of the Jacobian product times the gradient at , as in Jacobians and the chain rule, and transposing a product reverses its order: . Read it from right to left: the gradient starts at and is multiplied by one factor for every step it travels back, each factor being the previous lesson's rule, "through the tanh slopes, then through ". After steps it has been multiplied by matrices, each built from the same .
A bound from singular values
With matrices we cannot just multiply numbers, but we can bound how much each factor stretches. In symmetric matrices and the SVD the largest singular value was the gain of a matrix: for every . Each backward factor does two things:
- It multiplies the entries of the gradient by the slopes, each between 0 and 1, which cannot make it longer.
- It multiplies by , which has the same singular values as and so stretches by at most .
So each step back can stretch the gradient by at most , and after steps
If , the gradient from steps back is guaranteed to shrink at least geometrically: vanishing is certain. If the bound allows growth but does not force it, because the tanh slopes may pull every factor below 1, as they did for in the diagram.
Go slower: The bound, one factor at a time
Let be any gradient vector and the slopes, each in . The factor is , since multiplying by multiplies entry by entry.
First, the slopes cannot lengthen : , since every .
Second, stretches by at most . If then , an SVD with the same singular values. So for every .
Together, with : . Apply this times, once per factor, starting from : each application multiplies the bound by one more , which gives .
The eigenvalue view
The singular-value bound holds for any states. When the states settle into a steady pattern, we can say more. The factors are then nearly the same matrix every step, and a product of copies behaves like . The eigenvalues of powers, stability and explosion take over: the gradient scales like , where is the spectral radius, the size of the largest eigenvalue ( has the same eigenvalues as ).
Now think about what holding a memory requires. A state that stores something robustly sits at a fixed point that pulls nearby states back toward it, so that small disturbances die out instead of wiping the memory. Pulling nearby states back means the Jacobian shrinks small differences, so , and the gradient, which travels back through the same Jacobians, shrinks too. Storing information robustly and passing a learning signal back through it pull in opposite directions. The scalar latch above is the one-unit version: its factor is also how much a small disturbance of the stored state shrinks at each step, which is what makes the memory robust and the gradient vanish.
Watching it happen
The instrument below runs the matrix version on a random RNN with 8 tanh units. Its recurrent matrix is , where is a fixed random orthogonal matrix (every singular value 1), so is both the spectral radius and the largest singular value of . The chart plots against the number of steps back , from 0 to 80: the spectral norm of the whole product of Jacobians, written with double bars as in symmetric matrices and the SVD, the most it can stretch any gradient. The vertical axis is a log scale, so geometric decay or growth is a straight line. A toggle includes the tanh slopes or leaves them out.
The same experiment in code, with a 64-unit RNN driven by random inputs. G is a random matrix with spectral radius close to 1 (and largest singular value close to 2), and W_hh is G times a scale. The code walks a gradient of length 1 back from the last state with the previous lesson's rule and records its length after every step.
At scales 0.5 and 1 the gradient from 80 steps back is smaller than of where it started. At scale 4 it is larger than . Only in a narrow band around scale 2 does it stay within a few powers of ten, and nothing in training keeps inside that band. Notice also which way the singular-value bound works. At scale 1 the largest singular value is about 2, so the bound allows growth by up to , yet the gradient vanished: the tanh slopes and the directions of the matrices did the rest. A bound below 1 guarantees vanishing; a bound above 1 guarantees nothing.
To recap: going back steps multiplies the gradient by factors built from the same and the tanh slopes; their size is at most ; saturated states make the factors smaller still; and in practice the result is geometric decay or growth.
What vanishing does to learning
Why does a tiny gradient from far back matter, if the gradient from nearby steps is fine? Because every weight update adds them together. Every entry of is a sum of contributions, and each contribution links a loss at some step to a state some number of steps earlier, scaled by roughly .
With , a link one step long counts , a link five steps long counts , and a link 20 steps long counts . The long-range contributions are in the sum, but they are drowned out by the short-range ones. Gradient descent follows the sum, so it learns the short-range patterns and barely sees the long-range ones.
What can be done
The two problems are not symmetric, and neither are their remedies.
Exploding gradients: clip them. This is the clipping of the previous lesson. When the product of factors blows up, clipping by norm caps the length of the whole gradient and keeps its direction, so one bad batch cannot throw the weights far away. It treats the symptom, a step that is too long, and that is enough in practice.
Vanishing gradients: no simple fix. Clipping cannot help, and neither can scaling the gradient up, because the problem is relative: the long-range contributions are tiny compared with the short-range ones in the same sum, and scaling multiplies both. Optimizers such as Adam, which divide each weight's step by the typical size of its gradient, help when a gradient is small but consistent; they cannot pull a faint long-range signal out from under larger short-range contributions to the same weights. Careful initialization (a recurrent matrix close to the identity, or an orthogonal one, whose singular values are all 1) helps at the start of training, but nothing keeps the matrix there. The fix that worked was a change of architecture: give the network a memory path along which the per-step factor is a number the network controls and can hold near 1. That is the cell state of the LSTM, the subject of the next module.
Truncated backpropagation through time. A different, practical limit also cuts long-range learning. A long text cannot be unrolled all at once: the forward pass stores every state, so memory grows with every step. The usual compromise cuts the text into consecutive chunks of steps and treats the two directions differently:
- The forward pass carries the state across chunk boundaries: the last state of one chunk becomes of the next, so the model can still carry information forward indefinitely.
- The backward pass stops at the boundary: the carried-in is treated as a constant, so no gradient crosses into the previous chunk.
The model can use information from further back than steps, but it gets no training signal telling it to keep such information. The training in the previous lesson was simpler still: every window was cut from a random place and started from a zero state, so the model never saw more than 32 characters of context while it trained.
Check that you can tell apart what truncation keeps and what it cuts.
You train on a long text with truncated BPTT: chunks of 50 characters, the last hidden state of each chunk carried into the next chunk as its starting state, and gradients stopped at chunk boundaries. Which statement is true?
Choose one answer, then check.
Where the plain RNN breaks: delayed recall
Everything so far says the long-range signal is weak. A clean experiment shows how weak, with a task that needs memory and nothing else. The delayed-recall task uses 9 tokens: four symbols (ids 0 to 3), four noise tokens (ids 4 to 7) and a cue (id 8). Each sequence is
where is a random symbol, the are random noise tokens, and is the delay. For example, with one sequence might be 2, 5, 7, 4, 8, 2: the symbol 2, three noise tokens, the cue, and the symbol again.
The model is trained exactly like the language model: predict the next token at every position, with teacher forcing and the average cross-entropy. Most targets are noise tokens, which nothing can predict better than a one-in-four guess. The one that matters is the last: the token after the cue is , which was read steps earlier. Guessing gives 25% accuracy at that position; remembering gives 100%.
This is the previous sections in miniature. The gradient that could teach the network to store comes from a single position and has to travel back steps, while the noise positions contribute larger, short-range gradients to the same weights.
Build the delayed-recall task and measure how far back a plain RNN can remember.
make_recall_batch(batch_size, delay, rng): sequences , turned into next-token pairs.recall_accuracy(params, delay, rng, n): the fraction of fresh sequences whose last prediction is .train_recall(delay, ...): train a fresh RNN with the language-model loss at every position.delay_sweep(delays): one trained RNN per delay, and its accuracy.
Provided: init_rnn, rnn_forward, rnn_loss_and_grads and clip_grads from mlref.rnn, Adam from mlref.nn, and the constants K = 4 symbols, N_NOISE = 4 noise tokens, CUE = 8 and V = 9. The tests run the sweep over delays 2, 4, 8, 16, 24 and 32, which takes about 10 seconds.
This one trains a model in your browser, so a run can take a minute or more. The charts update as it goes, and Stop ends it early.
raise NotImplementedError, then press Run tests. Each check says what it expects.In our runs the RNN solves delays 2, 4 and 8 completely. At 16 the outcome depends on the run: with some random seeds it learns the task, with others it gets partway or stays at chance. At 24 and 32 it stays at chance, 25%: after 800 training steps its answers show no trace of . Training longer moves the boundary out only slowly, because each extra step of delay multiplies the useful signal by another factor below 1. Module 8 runs a version of this experiment, set up as a classifier that answers once at the end, with an RNN and an LSTM side by side, and the LSTM keeps learning at delays where the RNN usually fails.
To finish the module, explain the whole chain of reasoning in your own words, from the product of factors to the fix.
Explain to a colleague why a plain RNN trained with gradient descent struggles to learn dependencies more than a few dozen steps back. Cover where the product of Jacobians comes from, what bounds its size, why gradient clipping does not help, and what kind of fix does.
Saved in this browser as you type.