Lesson 6 of 6
Mini-batches and linear regression
A real training loss is an average over many examples, so training estimates its gradient from a small random batch, and Adam gives each weight its own step size. Then fit linear regression both ways, exactly and by descent, and see why the scale of the features decides how fast descent gets there.
Explain why a mini-batch gradient is an unbiased but noisy estimate of the full gradient, and what the batch size changes.
Describe what Adam does differently from plain gradient descent and why it helps with poor conditioning.
Derive the normal equations and solve a small linear regression exactly.
Train linear regression with mini-batch gradient descent and compare it with the closed-form solution.
Predict the learning-rate limit of linear regression from its Hessian, and explain standardizing features as conditioning.

The last lesson ran gradient descent on losses written as one short formula, whose exact gradient costs almost nothing to compute. A real training loss is an average over a dataset, with one term per example. For a small regression that is 50 terms. For a language model it is trillions of tokens, and computing the exact gradient once would mean a pass over all of them for every single step.
This lesson makes two moves. First, it estimates the gradient from a small random sample of examples, which is how every large model is trained, and it names the optimizer most of them use. Second, it fits linear regression, the one model where gradient descent can be checked against an exact answer, and shows why the scale of the features decides how fast descent gets there.
Mini-batch gradients
Why can a random handful of examples stand in for the whole dataset? Because the full gradient is an average, and an average can be estimated from a sample.
A training loss is an average over examples: , where is the loss on example . By the sum rule, its gradient is the average of the per-example gradients, . Computing it exactly costs a pass over all examples.
So instead, pick a random mini-batch of examples and average the gradients over just those:
The hat on marks it as an estimate. A tiny example shows what that estimate does. Suppose a model has one weight and four examples, whose per-example gradients are , , and . The full gradient is their average, . With batches of two examples there are six possible batches:
| batch | its average gradient |
|---|---|
| examples with gradients and | |
| and | |
| and | |
| and | |
| and | |
| and |
No batch gives exactly 3, and two of them are off by . But the average over all six batches is : exactly the full gradient.
That is what unbiased means: averaged over all the batches you might draw, the estimate equals the true gradient. Any single estimate is noisy. Its variance falls like (averaging independent random draws divides the variance of one draw by ), so its typical error, the standard deviation (the square root of the variance), falls like : four times the batch size halves the typical size of the noise. Each step costs work proportional to rather than . With the method is called stochastic gradient descent (SGD); in practice the name is used for any mini-batch size. One pass through the whole dataset is an epoch.
Why does trading exactness for noise pay off? Real datasets are redundant: a thousand examples of a pattern give nearly the same gradient as a million. Many cheap noisy steps make far more progress per unit of compute than a few exact ones. Two side effects follow. The loss curve becomes jagged, so you judge progress by its trend. And with a constant learning rate the weights never settle exactly: near the minimum the gradient is small, but the noise is not, so it keeps them jittering around the minimum. Lowering over the course of training (a schedule) shrinks that jitter.
Mini-batch steps are noisy but head the right way on average; at a constant learning rate they keep jittering near the bottom.
So far. A mini-batch gradient is the average gradient over a random sample. It is right on average and costs of a full pass, and its noise shrinks like . The price is a jagged loss curve and weights that jitter near the minimum unless the learning rate is lowered.
Check what changes, and what does not, when a full-batch run switches to mini-batches.
Your dataset has 50,000 examples. You switch from full-batch gradient descent to mini-batches of 32, keeping the learning rate the same. Which statement is right?
Choose one answer, then check.
Adam: a step size for each weight
Momentum, from the last lesson, eases badly conditioned valleys by averaging gradients over time. The optimizer most language models are trained with adds a second idea: give every weight its own step size, matched to the size of its own gradients.
Adam (Kingma and Ba; the paper appeared in 2014 and was published at ICLR in 2015) keeps two running averages for every parameter: one of its gradient, like momentum, and one of its squared gradient. Each parameter's step is times its averaged gradient divided by the square root of its averaged squared gradient. That square root is roughly the parameter's typical gradient size.
A small example. One weight gets gradients of about on every step, another gets gradients of about . Plain gradient descent moves them and per step, so the first moves ten thousand times faster than the second. Adam divides each averaged gradient by its typical size, and , so both move about per step.
That is a cheap, per-coordinate answer to poor conditioning; it helps most when the badly scaled directions line up with individual parameters. Adam and its variant AdamW are the default optimizers for training language models. You will derive and implement Adam in module 6 and use it to train LSTMs in module 8.
Linear regression, two ways
Linear regression is the one model where you can check gradient descent against an exact answer. It also shows conditioning in its purest form, because its loss is exactly a quadratic bowl.
Its loss is the mean squared error from the gradients lesson, with gradient . At the minimum the gradient is zero. Set it to zero and solve, one move at a time.
-
Start from .
-
Multiply both sides by : .
-
Multiply out the bracket: .
-
Move to the right side. The result is the normal equations:
When the columns of are independent, is invertible and this has exactly one solution. In code, solve it with np.linalg.solve rather than by forming an inverse, as in module 3.
Take the three points from the gradients lesson: feature values with targets , and , whose column of ones carries the bias.
- . Dot the columns of with each other: ones with ones gives , ones with the feature gives , the feature with itself gives . So .
- . Dot each column with : .
- Solve. The determinant is , so the inverse is (swap the two diagonal entries, negate the other two, divide by the determinant, as in module 3), and .
The best line is . Its predictions are , and , so its residuals are , and : the gradient is zero, as it must be at the minimum.
Now gradient descent. The Hessian is the Jacobian of the gradient. The gradient is a matrix times plus a constant, so its Jacobian is that matrix: the Hessian is at every point. The loss is exactly a quadratic bowl, and everything in the last lesson applies without approximation. For the three points, , with eigenvalues about and (the roots of , whose coefficients are the trace and the determinant , as in module 4): a condition number of . Even this tiny problem is a narrow valley.
The scale of the features sets the curvatures. If one feature takes values in the thousands and another in single digits, has one enormous eigenvalue and one small one, the condition number is huge, and gradient descent crawls. Standardizing each feature (subtract its mean, divide by its standard deviation) makes the curvatures similar. It is conditioning, done by hand.
The code below fits a bias and a single feature ranging from 0 to 10: 50 noisy points around the line . It solves the normal equations, prints the Hessian's eigenvalues, runs 2,000 steps of full-batch gradient descent at half the stability limit, and then repeats the fit with the feature standardized. Even this mild mismatch in scale gives a condition number above 100.
The closed form is : bias about , slope about . After 100 steps of descent the slope is too large by about a quarter and the bias is only about half its final value; after 2,000 steps gradient descent agrees with the closed form to four decimals. The flat direction, which mixes the bias with the slope, is what takes so long.
Why standardizing fixes it
Standardizing the feature fixes it completely: both Hessian eigenvalues become exactly 2, the bowl is perfectly round, and a single step at lands on the answer. (The weights differ from the ones above because the feature is now measured in different units; they describe the same line.)
Subtracting the mean matters as much as dividing by the standard deviation. A feature that is always positive, like this , overlaps with the column of ones, so the bias and the slope compete to explain the same thing, and that overlap is the flat, mixed direction. Centering makes the two columns perpendicular. On this data, dividing by the standard deviation alone brings the condition number from 161 to about 26, subtracting the mean alone brings it to about 8.5, and both together bring it to 1.
In short. The normal equations give linear regression's answer in one solve. Gradient descent reaches the same answer, at a pace set by the condition number of , and standardizing the features brings that condition number down toward 1.
Now write both methods yourself: the closed form, and mini-batch descent that should land close to it.
Fit linear regression both ways.
loss_and_grad(X, y, w): the mean squared error and its gradient.normal_equations(X, y): the exact minimizer, from , usingnp.linalg.solve.minibatch_sgd(X, y, lr, epochs, batch_size, seed=0): start from . Each epoch, shuffle the examples, walk through them in batches (keeping the last, smaller batch), take one step per batch using that batch's own average gradient, and record the loss on all the data at the end of the epoch.
The demo fits . That curve is still linear regression, because it is linear in the weights: is just another feature column.
raise NotImplementedError, then press Run tests. Each check says what it expects.The next exercise closes the loop between theory and practice: predict the learning-rate limit from the Hessian, then watch a sweep of learning rates confirm it, and see what standardizing does to the condition number.
Predict a learning-rate limit from the Hessian, then test the prediction.
hessian(X): the Hessian of the mean squared error, .lr_limit(X): of that Hessian.gd_losses(X, y, lr, steps): full-batch gradient descent from , returning the loss before the first step and after every step:steps + 1values.standardize(X): subtract each column's mean and divide by its standard deviation.
In the demo the second feature is ten times larger than the first. It sweeps learning rates at fixed fractions of your predicted limit and plots every loss curve on one chart, then repeats the run after standardizing.
raise NotImplementedError, then press Run tests. Each check says what it expects.Finally, diagnose a colleague's slow model with everything from this lesson and the last.
A colleague's linear model trains painfully slowly with gradient descent. One feature is annual income in dollars (values around 50,000), another is age in years. Every learning rate they try either diverges or crawls. Explain what is happening in terms of the Hessian and the learning-rate limit, and suggest two fixes.
Saved in this browser as you type.