Lesson 6 of 6
Momentum and Adam
Plain gradient descent has no memory and one step size for every parameter. Momentum adds a memory of recent gradients, which speeds up steady directions and damps oscillations. Adam also gives every parameter its own step size. This lesson derives both, bias correction included, and has you implement them.
Explain what momentum does in a long, narrow valley, and compute a momentum run by hand.
Derive the Adam update as running averages of the gradient and its square, and explain why its first step is in every coordinate.
Derive Adam's bias correction from a geometric series.
Implement SGD with momentum and Adam, and compare them with plain gradient descent on a badly scaled problem.

The loop from the last lesson updates every parameter with , where is the gradient of the current batch. That rule has two weaknesses you have already seen. It has no memory: each step uses only the current gradient, so a direction that points the same way step after step never builds up speed, and a direction whose gradient flips sign every step keeps zig-zagging. And it has one step size for everything: the learning rate must be small enough for the steepest direction, which leaves the gentle directions crawling.
The optimizer is the rule that turns gradients into updates, and two changes to plain gradient descent fix these weaknesses. Momentum gives the parameters a memory of recent gradients. Adam adds a separate step size for every parameter. Adam, in a variant called AdamW, is the optimizer behind most large language models.
Momentum: build up speed
Gradient descent met momentum on a quadratic valley and worked out exactly when it helps. Here is the short version, and what changes when the gradients come from mini-batches.
In a long, narrow valley, where the curvature is large across the valley and small along it, the learning rate has to be small enough for the steep direction, so progress along the gentle direction crawls. Momentum (Polyak's "heavy ball", 1964) gives the parameters a velocity . Module 5 wrote the momentum coefficient as ; here it is (mu), because Adam below needs and for something else. Typically , and each step does two things, starting from :
The first line decays the old velocity by the factor and adds the current gradient . The second moves the weights along the velocity instead of along the gradient. Unrolled, the velocity is a running sum of past gradients in which each older gradient counts times less. Three consequences:
- Steady directions speed up. If the gradient is the same every step, the velocity after steps is . That geometric series approaches , so the velocity approaches . With that is times the plain step. Along the floor of the valley, where the gradient points the same way step after step, momentum accelerates.
- Oscillating directions cancel. Across the valley the gradient flips sign every step, so consecutive contributions to the velocity mostly cancel, and the zig-zag is damped. Module 5 showed the price: with coefficient , no direction can shrink faster than a factor per step, so on a valley that plain descent already handles well, heavy momentum is slower.
- Mini-batch noise averages out. The velocity sums about the last gradients (10 of them for ). Each mini-batch gradient is the true gradient plus noise; the true parts add up step after step, while noise from different batches points in different directions and partly cancels. So the velocity is a steadier guide than any single batch gradient.
A worked example, one parameter, , , and a steady gradient :
- , and the parameter moves .
- , and it moves .
- , and it moves .
After three steps it has moved , against for plain gradient descent, and the step size keeps growing toward .
The widget answers one question: when does momentum help, and when does it hurt? It shows the narrow valley from module 5 as a contour map: each closed curve joins points of equal loss, the lime cross marks the minimum at the center, and the blue dot is the start at . The curvature is 10 across the valley (the direction) and 1 along it (the direction), so plain descent is stable only for . Play runs gradient descent from the start and draws its path; the chart under the map plots how far the loss is above its minimum, as with the lowest possible loss, step by step (a straight falling line means the gap shrinks by the same factor every step), and a live note gives each direction's shrink factor per step and, at the end, how many steps the run took. The widget calls the momentum coefficient , as module 5 did; it is the above.
Now run momentum by hand with different numbers.
SGD with momentum uses and , starts from velocity , and sees the same gradient at every step. By how much has the parameter decreased after three steps? (Enter a positive number.)
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
To recap: momentum keeps a decaying sum of past gradients and steps along it. Directions where the gradient keeps its sign speed up by up to , directions where it flips cancel out, and batch noise averages away.
Adam: a step size for every parameter
Momentum still uses one learning rate for every parameter. But gradients in a network differ in size by orders of magnitude from one parameter to another: the output layer's biases may see gradients a thousand times larger than an early layer's weights. One learning rate is then too large for some parameters and too small for others.
Adam (Kingma and Ba; the paper appeared in 2014 and was published at ICLR in 2015), which mini-batches and linear regression previewed, keeps two running averages for every parameter, and divides one by the square root of the other. Here is the full algorithm. With counting steps from 1, and two decay rates between 0 and 1, and every operation entry by entry (so squares each entry of the gradient at step ):
with . The defaults are , and , a tiny number that only prevents division by zero.
Why it works. Take the two averages one at a time.
- is momentum's velocity in another form. Call the momentum velocity for a moment, since now means the average square, and give it the same decay: . It adds each gradient at full weight, while adds it at weight . Both start at , and multiplying the update of by gives exactly the update of , so at every step: an average of recent gradients instead of a sum.
- is the typical size of recent gradients, their root-mean-square over roughly the last steps.
- Their ratio is a pure number, roughly between and , whatever the scale of the gradient: near when recent gradients agree in sign, and near 0 when they are noise around zero.
So every parameter moves by roughly per step when its gradient is consistent, and less when it is not, whether its raw gradients are or . This is why one Adam learning rate, such as or , is a reasonable starting point for wildly different models.
Why the bias correction. Both averages start at zero, so for the first several steps they are too small. With the defaults, after one step and . Dividing by and undoes exactly this: and . The first step is then , entry by entry: every coordinate moves by exactly (up to ), in the direction opposite to its gradient. Without the correction, the first step would be times larger than that, because starts even further below its true value than does.
Go slower: Where comes from
Step 1, unroll the average. Start from and apply the update times:
Step 2, suppose the gradient were constant, for every . A good average should then equal . Instead, pulling out of the sum and writing ,
Step 3, sum the geometric series. . Multiply by : . The weights on the gradients add up to instead of 1, because the missing weight went to the zero we started from.
Step 4, correct. Dividing by makes the weights add up to 1: exactly when the gradient is constant. The same argument with gives . As grows, and the correction fades away: with it is under 1% after 44 steps (), and with after about 4,600.
Now run two Adam steps by hand, where the two parameters' gradients differ by a factor of 20.
Run Adam by hand for two steps on a model with two parameters, starting from , with , , and ignored. The gradients are at step 1 and at step 2.
- Compute , , the corrected and , and . How far did each coordinate move?
- Compute , , the corrected averages (use and ) and .
- At step 1 the two gradients differ in size by a factor of 20. Why did both coordinates move by the same amount? And why did the second coordinate move so little at step 2?
To check part 2 before reading the steps, enter as a column, the first coordinate at the top, to 4 decimal places.
Work it on real paper: writing each step is the point. Then check your final answer here and compare your working with the walk-through.
One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.
To recap: Adam's step for each parameter is about times (average gradient) divided by (typical gradient size). It is close to when recent gradients agree, close to 0 when they conflict, and independent of the gradient's overall scale.
Racing the optimizers
Now implement both optimizers. The code ends by racing plain SGD, momentum and Adam on a badly scaled bowl, . Its curvature is 1 along and 100 along , so the stable learning rate for plain gradient descent is below , the same trap as the widget's valley, ten times worse.
Here is what the race does. With , plain descent multiplies the steep coordinate by each step, so it flips across the valley every step, while the gentle coordinate shrinks by only per step: after 100 steps is still . Momentum at the same learning rate swings across the valley too, but its speed builds along the floor, so much that it overshoots the minimum along the floor, to at step 23. It swings back past the minimum with ever smaller swings ( at step 47, at step 71) and settles. The picture below shows the two kinds of path on a valley like this one, seen from above.
On a badly scaled bowl, plain gradient descent zig-zags across the valley and creeps along it. A heavy ball keeps its speed along the floor, where the gradient keeps pointing the same way, overshoots, and settles at the bottom.
Implement the two optimizers from the lesson. Each takes a list of parameter arrays when it is created, keeps its own state for each one, and has step(grads), which updates every parameter in place from a list of gradients in the same order.
SGDMomentum(params, lr, momentum=0.9): , then .Adam(params, lr=1e-3, beta1=0.9, beta2=0.999, eps=1e-8), with bias correction.
The last lines race plain SGD, momentum and Adam on the badly scaled bowl , starting from . Plain SGD's learning rate is just under its stability limit .
raise NotImplementedError, then press Run tests. Each check says what it expects.