ml.lab
Python sleeps until you run code
05 Derivatives, gradients and Jacobians

Lesson 2 of 6

Derivatives by experiment

The definition of the derivative doubles as a way to measure one. Forward and central differences, why the central one is far more accurate, and why making the step too small backfires in floating point. This is gradient checking.

About 35 minutes
By the end you can
  • Estimate a derivative with forward and central differences.

  • Use Taylor's formula to show that the forward difference's error is proportional to hh and the central difference's to h2h^2.

  • Explain why a very small step makes the estimate worse in floating point, and choose a sensible step for gradient checking.

In the last lesson you derived σ′(x)=σ(x) (1−σ(x))\sigma'(x) = \sigma(x)\,(1 - \sigma(x)) in seven lines of algebra. Suppose one of those lines had dropped a minus sign. The result would still look like a reasonable formula, and code built on it would still run. In module 6 you will write backward passes with dozens of such formulas, and a network trained with a wrong gradient often still learns a little, so a falling loss proves nothing.

You need a check that cannot share your mistakes. The definition of the derivative gives one: forget the formula, pick a small step, and measure the slope directly. This lesson makes that measurement precise in three steps: two ways to estimate a slope, where each one's error comes from, and why a step that is too small makes things worse.

Two ways to estimate a slope

The definition says the derivative is the value the secant's slope settles on as the step shrinks. So pick one small hh and compute the fraction. This is a finite difference, and comparing it with your formula is called gradient checking. There are two common versions:

Dfwd=f(x+h)−f(x)h,Dcen=f(x+h)−f(x−h)2h.D_{\text{fwd}} = \frac{f(x + h) - f(x)}{h}, \qquad D_{\text{cen}} = \frac{f(x + h) - f(x - h)}{2h}.

The forward difference is the slope of the secant from xx to x+hx + h. The central difference is the slope of the secant from x−hx - h to x+hx + h, which straddles the point; its run is 2h2h.

Try both on f(x)=x3f(x) = x^3 at x=2x = 2, where the true derivative is 3x2=3×4=123x^2 = 3 \times 4 = 12, with h=0.1h = 0.1. You need three values: f(2.1)=2.13=9.261f(2.1) = 2.1^3 = 9.261, f(2)=8f(2) = 8 and f(1.9)=1.93=6.859f(1.9) = 1.9^3 = 6.859.

  • Forward: 9.261−80.1=1.2610.1=12.61\frac{9.261 - 8}{0.1} = \frac{1.261}{0.1} = 12.61, an error of 12.61−12=0.6112.61 - 12 = 0.61.
  • Central: 9.261−6.8590.2=2.4020.2=12.01\frac{9.261 - 6.859}{0.2} = \frac{2.402}{0.2} = 12.01, an error of 12.01−12=0.0112.01 - 12 = 0.01.

The widget runs the same comparison for every step size. It answers: how does each estimate's error change as hh shrinks? On screen are the white curve x3x^3, the blue point at x0=2x_0 = 2 with its amber tangent (slope 12), a coral dot at x0+hx_0 + h with the coral forward secant through it, and a violet dot at x0−hx_0 - h with the dashed violet line of the central difference. The readouts give both estimates, with their errors underneath. Under the hh slider, a chart plots each estimate's error against hh, both on log scales: a label of −4-4 on either axis means 10−410^{-4}.

Where the error comes from

Both estimates use two evaluations of ff, yet at h=0.1h = 0.1 the central one is about sixty times more accurate (0.61/0.01=610.61 / 0.01 = 61). Why? The answer comes from extending the tangent line with more terms. Taylor's formula says that for a smooth function

f(x+h)=f(x)+f′(x) h+12f′′(x) h2+16f′′′(x) h3+⋯f(x + h) = f(x) + f'(x)\,h + \tfrac{1}{2} f''(x)\,h^2 + \tfrac{1}{6} f'''(x)\,h^3 + \cdots

where f′′f'' is the derivative of f′f' (the second derivative, which measures curvature: how fast the slope itself changes) and f′′′f''' is the third. The first two terms are the tangent line; each further term adds a correction that is one power of hh smaller. The factors 12\tfrac{1}{2} and 16\tfrac{1}{6} are what make the right side bend like ff: differentiate 12f′′(x) h2\tfrac{1}{2} f''(x)\,h^2 twice with respect to hh and you get back f′′(x)f''(x), and differentiate 16f′′′(x) h3\tfrac{1}{6} f'''(x)\,h^3 three times and you get back f′′′(x)f'''(x).

Check it on x3x^3 at x=2x = 2, one derivative at a time. The value is f(2)=8f(2) = 8. The slope is f′(x)=3x2f'(x) = 3x^2, so f′(2)=12f'(2) = 12. Differentiating 3x23x^2 once more gives f′′(x)=6xf''(x) = 6x, so f′′(2)=12f''(2) = 12. Differentiating 6x6x gives f′′′(x)=6f'''(x) = 6. The coefficients are then 1212, 12×12=6\tfrac{1}{2} \times 12 = 6 and 16×6=1\tfrac{1}{6} \times 6 = 1, so the formula gives 8+12h+6h2+h38 + 12h + 6h^2 + h^3, which is exactly (2+h)3(2 + h)^3 multiplied out. (The fourth derivative of x3x^3 is 0, so for this function the formula stops there.)

Taylor's formula also explains a number from the last lesson. There, the tangent line's error for sin⁡\sin at x=1x = 1, divided by h2h^2, settled near −0.42-0.42. The tangent line keeps the first two terms of the formula, so its error starts with the third, 12f′′(1) h2\tfrac{1}{2} f''(1)\,h^2. For sin⁡\sin, the slope is cos⁡\cos and the slope of cos⁡\cos is −sin⁡-\sin, so f′′=−sin⁡f'' = -\sin. That term is then −12sin⁡(1) h2≈−12×0.8415 h2≈−0.42 h2-\tfrac{1}{2}\sin(1)\,h^2 \approx -\tfrac{1}{2} \times 0.8415\,h^2 \approx -0.42\,h^2.

Go slower: The error of each difference

Forward, first move. Subtract f(x)f(x) from both sides of Taylor's formula:

f(x+h)−f(x)=f′(x) h+12f′′(x) h2+16f′′′(x) h3+⋯f(x + h) - f(x) = f'(x)\,h + \tfrac{1}{2} f''(x)\,h^2 + \tfrac{1}{6} f'''(x)\,h^3 + \cdots

Forward, second move. Divide by hh:

Dfwd=f′(x)+12f′′(x) h+16f′′′(x) h2+⋯D_{\text{fwd}} = f'(x) + \tfrac{1}{2} f''(x)\,h + \tfrac{1}{6} f'''(x)\,h^2 + \cdots

The leading error is 12f′′(x) h\tfrac{1}{2} f''(x)\,h: proportional to hh. For x3x^3 at 2 with h=0.1h = 0.1 the error terms are 12(12)(0.1)+16(6)(0.01)=0.6+0.01=0.61\tfrac{1}{2}(12)(0.1) + \tfrac{1}{6}(6)(0.01) = 0.6 + 0.01 = 0.61, the error we measured.

Central, first move. Write Taylor's formula for a step of −h-h. The odd powers of hh change sign:

f(x−h)=f(x)−f′(x) h+12f′′(x) h2−16f′′′(x) h3+⋯f(x - h) = f(x) - f'(x)\,h + \tfrac{1}{2} f''(x)\,h^2 - \tfrac{1}{6} f'''(x)\,h^3 + \cdots

Central, second move. Subtract this from the formula for f(x+h)f(x + h), term by term. f(x)−f(x)=0f(x) - f(x) = 0. f′(x) h−(−f′(x) h)=2f′(x) hf'(x)\,h - (-f'(x)\,h) = 2f'(x)\,h. The h2h^2 terms are equal, so they cancel. 16f′′′(x) h3−(−16f′′′(x) h3)=13f′′′(x) h3\tfrac{1}{6}f'''(x)\,h^3 - (-\tfrac{1}{6}f'''(x)\,h^3) = \tfrac{1}{3}f'''(x)\,h^3. So

f(x+h)−f(x−h)=2f′(x) h+13f′′′(x) h3+⋯f(x + h) - f(x - h) = 2 f'(x)\,h + \tfrac{1}{3} f'''(x)\,h^3 + \cdots

Central, third move. Divide by 2h2h:

Dcen=f′(x)+16f′′′(x) h2+⋯D_{\text{cen}} = f'(x) + \tfrac{1}{6} f'''(x)\,h^2 + \cdots

The leading error is 16f′′′(x) h2\tfrac{1}{6} f'''(x)\,h^2: proportional to h2h^2. For x3x^3: 16(6)(0.01)=0.01\tfrac{1}{6}(6)(0.01) = 0.01, again what we measured. The curvature term cancelled because the central difference is symmetric around xx.

So halving hh halves the forward error and quarters the central error. On the widget's log-log chart that shows as the coral curve falling one unit and the violet curve two units for every tenfold drop in hh.

So far. Taylor's formula says the forward difference's error is about 12f′′(x) h\tfrac{1}{2}f''(x)\,h and the central difference's about 16f′′′(x) h2\tfrac{1}{6}f'''(x)\,h^2. The central one wins because it straddles the point, so the curvature terms cancel.

Try one central difference by hand, then compare its error with the formula from the slow box.

Work it outA central difference by hand

Estimate the derivative of f(x)=x4f(x) = x^4 at x=1x = 1 with a central difference and h=0.1h = 0.1. You need 1.14=1.46411.1^4 = 1.4641 and 0.94=0.65610.9^4 = 0.6561. Enter the estimate, not the exact derivative, as a decimal to two decimal places.

central difference

Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.

Why a tiny step fails

Both errors shrink as hh shrinks, which suggests making hh as small as possible. That is wrong, and the reason is the way computers store numbers.

A standard 64-bit float keeps about 16 significant decimal digits. The gap between 11 and the next number a float can represent is machine epsilon, ε≈2.2×10−16\varepsilon \approx 2.2 \times 10^{-16}. When hh is tiny, f(x+h)f(x + h) and f(x)f(x) agree in most of their leading digits. Subtracting them cancels those digits and leaves mostly rounding noise, of size roughly ε ∣f(x)∣\varepsilon\,|f(x)|. Dividing by hh then magnifies that noise by 1/h1/h. The total error is a truncation part that shrinks as hh shrinks plus a roundoff part that grows like 1/h1/h, and the best hh is roughly where the two are equal. Take a function whose values and derivatives are of moderate size, around 1, so the constants can be dropped:

  • Forward: truncation about hh, roundoff about ε/h\varepsilon/h. They are equal when h=ε/hh = \varepsilon / h. Multiply both sides by hh: h2=εh^2 = \varepsilon, so the best step is near ε≈1.5×10−8\sqrt{\varepsilon} \approx 1.5 \times 10^{-8}, and the error there is about h=ε≈1.5×10−8h = \sqrt{\varepsilon} \approx 1.5 \times 10^{-8}.
  • Central: truncation about h2h^2, roundoff about ε/h\varepsilon/h. They are equal when h2=ε/hh^2 = \varepsilon / h. Multiply both sides by hh: h3=εh^3 = \varepsilon, so the best step is near ε1/3≈6×10−6\varepsilon^{1/3} \approx 6 \times 10^{-6}, and the error there is about h2=ε2/3≈4×10−11h^2 = \varepsilon^{2/3} \approx 4 \times 10^{-11}, hundreds of times smaller than the forward difference's best.

The code below tests those predictions on exe^x at x=1x = 1, whose exact derivative is e≈2.718e \approx 2.718. It prints both errors for h=10−1h = 10^{-1} down to 10−1410^{-14} and plots them against kk, where h=10−kh = 10^{-k}:

⌘+Enter runs · edit freelyPython sleeps until you run code

Both curves fall, reach a bottom and climb again. The forward error falls tenfold with each row and bottoms out at about 7×10−97 \times 10^{-9} at h=10−8h = 10^{-8}. The central error falls a hundredfold with each row, reaches about 6×10−116 \times 10^{-11} at h=10−5h = 10^{-5}, wobbles at that level down to h=10−7h = 10^{-7} (roundoff is partly luck), and from h=10−8h = 10^{-8} on it climbs. On the left of each bottom (large hh), truncation dominates; on the right (small hh), roundoff does. This is why the course's numerical_gradient helper uses a central difference with h=10−5h = 10^{-5}, and why a gradient check compares relative error against a small threshold rather than expecting exact agreement.

In short. A central difference has an error proportional to h2h^2 and a forward difference an error proportional to hh, while rounding adds an error that grows like 1/h1/h. So a gradient check uses a central difference with hh near 10−510^{-5} and compares against a tolerance, never expecting exact agreement.

Now build the tools yourself. The last function finds the best step size by experiment, and the chart it draws should reproduce the bottoms you just saw.

Code itDerivatives by experiment

Build the tools for checking derivatives.

  • sigmoid(x) and sigmoid_prime(x), working on numbers and on numpy arrays. Use σ′=σ(1−σ)\sigma' = \sigma(1 - \sigma).
  • forward_diff(f, x, h) and central_diff(f, x, h), the two finite differences from the lesson.
  • best_h(f, df, x, hs): given the exact derivative df, return the step size in hs whose central difference is most accurate at x.

The code at the bottom prints both errors for sin⁡\sin at x=1x = 1 as hh shrinks, and plots them. Before you run it, predict where each curve bottoms out.

⌘+Enter runsPython sleeps until you run code
Write your code where the starter says raise NotImplementedError, then press Run tests. Each check says what it expects.
Next: Gradients