Lesson 2 of 6
Derivatives by experiment
The definition of the derivative doubles as a way to measure one. Forward and central differences, why the central one is far more accurate, and why making the step too small backfires in floating point. This is gradient checking.
Estimate a derivative with forward and central differences.
Use Taylor's formula to show that the forward difference's error is proportional to and the central difference's to .
Explain why a very small step makes the estimate worse in floating point, and choose a sensible step for gradient checking.

In the last lesson you derived in seven lines of algebra. Suppose one of those lines had dropped a minus sign. The result would still look like a reasonable formula, and code built on it would still run. In module 6 you will write backward passes with dozens of such formulas, and a network trained with a wrong gradient often still learns a little, so a falling loss proves nothing.
You need a check that cannot share your mistakes. The definition of the derivative gives one: forget the formula, pick a small step, and measure the slope directly. This lesson makes that measurement precise in three steps: two ways to estimate a slope, where each one's error comes from, and why a step that is too small makes things worse.
Two ways to estimate a slope
The definition says the derivative is the value the secant's slope settles on as the step shrinks. So pick one small and compute the fraction. This is a finite difference, and comparing it with your formula is called gradient checking. There are two common versions:
The forward difference is the slope of the secant from to . The central difference is the slope of the secant from to , which straddles the point; its run is .
Try both on at , where the true derivative is , with . You need three values: , and .
- Forward: , an error of .
- Central: , an error of .
The widget runs the same comparison for every step size. It answers: how does each estimate's error change as shrinks? On screen are the white curve , the blue point at with its amber tangent (slope 12), a coral dot at with the coral forward secant through it, and a violet dot at with the dashed violet line of the central difference. The readouts give both estimates, with their errors underneath. Under the slider, a chart plots each estimate's error against , both on log scales: a label of on either axis means .
Where the error comes from
Both estimates use two evaluations of , yet at the central one is about sixty times more accurate (). Why? The answer comes from extending the tangent line with more terms. Taylor's formula says that for a smooth function
where is the derivative of (the second derivative, which measures curvature: how fast the slope itself changes) and is the third. The first two terms are the tangent line; each further term adds a correction that is one power of smaller. The factors and are what make the right side bend like : differentiate twice with respect to and you get back , and differentiate three times and you get back .
Check it on at , one derivative at a time. The value is . The slope is , so . Differentiating once more gives , so . Differentiating gives . The coefficients are then , and , so the formula gives , which is exactly multiplied out. (The fourth derivative of is 0, so for this function the formula stops there.)
Taylor's formula also explains a number from the last lesson. There, the tangent line's error for at , divided by , settled near . The tangent line keeps the first two terms of the formula, so its error starts with the third, . For , the slope is and the slope of is , so . That term is then .
Go slower: The error of each difference
Forward, first move. Subtract from both sides of Taylor's formula:
Forward, second move. Divide by :
The leading error is : proportional to . For at 2 with the error terms are , the error we measured.
Central, first move. Write Taylor's formula for a step of . The odd powers of change sign:
Central, second move. Subtract this from the formula for , term by term. . . The terms are equal, so they cancel. . So
Central, third move. Divide by :
The leading error is : proportional to . For : , again what we measured. The curvature term cancelled because the central difference is symmetric around .
So halving halves the forward error and quarters the central error. On the widget's log-log chart that shows as the coral curve falling one unit and the violet curve two units for every tenfold drop in .
So far. Taylor's formula says the forward difference's error is about and the central difference's about . The central one wins because it straddles the point, so the curvature terms cancel.
Try one central difference by hand, then compare its error with the formula from the slow box.
Estimate the derivative of at with a central difference and . You need and . Enter the estimate, not the exact derivative, as a decimal to two decimal places.
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
Why a tiny step fails
Both errors shrink as shrinks, which suggests making as small as possible. That is wrong, and the reason is the way computers store numbers.
A standard 64-bit float keeps about 16 significant decimal digits. The gap between and the next number a float can represent is machine epsilon, . When is tiny, and agree in most of their leading digits. Subtracting them cancels those digits and leaves mostly rounding noise, of size roughly . Dividing by then magnifies that noise by . The total error is a truncation part that shrinks as shrinks plus a roundoff part that grows like , and the best is roughly where the two are equal. Take a function whose values and derivatives are of moderate size, around 1, so the constants can be dropped:
- Forward: truncation about , roundoff about . They are equal when . Multiply both sides by : , so the best step is near , and the error there is about .
- Central: truncation about , roundoff about . They are equal when . Multiply both sides by : , so the best step is near , and the error there is about , hundreds of times smaller than the forward difference's best.
The code below tests those predictions on at , whose exact derivative is . It prints both errors for down to and plots them against , where :
Both curves fall, reach a bottom and climb again. The forward error falls tenfold with each row and bottoms out at about at . The central error falls a hundredfold with each row, reaches about at , wobbles at that level down to (roundoff is partly luck), and from on it climbs. On the left of each bottom (large ), truncation dominates; on the right (small ), roundoff does. This is why the course's numerical_gradient helper uses a central difference with , and why a gradient check compares relative error against a small threshold rather than expecting exact agreement.
In short. A central difference has an error proportional to and a forward difference an error proportional to , while rounding adds an error that grows like . So a gradient check uses a central difference with near and compares against a tolerance, never expecting exact agreement.
Now build the tools yourself. The last function finds the best step size by experiment, and the chart it draws should reproduce the bottoms you just saw.
Build the tools for checking derivatives.
sigmoid(x)andsigmoid_prime(x), working on numbers and on numpy arrays. Use .forward_diff(f, x, h)andcentral_diff(f, x, h), the two finite differences from the lesson.best_h(f, df, x, hs): given the exact derivativedf, return the step size inhswhose central difference is most accurate atx.
The code at the bottom prints both errors for at as shrinks, and plots them. Before you run it, predict where each curve bottoms out.
raise NotImplementedError, then press Run tests. Each check says what it expects.