07 · Basic Derivative Rules¶
Deriving every function from first principles (Module 6's limit definition) would make ML math unbearably slow. In practice we use a small toolkit of rules, each provable from first principles once and reused forever.
The power rule¶
This generalizes what we proved by hand in Module 6: for \(n=2\), \(\frac{d}{dx}[x^2] = 2x^{2-1} = 2x\) — exactly matching our first-principles derivation. A few more examples: \(\frac{d}{dx}[x^3] = 3x^2\), \(\frac{d}{dx}[x] = 1\cdot x^0 = 1\) (slope of the line \(y=x\) is 1, as expected), and \(\frac{d}{dx}[1] = \frac{d}{dx}[x^0] = 0\) — the derivative of any constant is zero (a constant doesn't change, so its rate of change is zero).
The constant multiple rule¶
Constants "pass through" differentiation. Example: \(\frac{d}{dx}[5x^2] = 5\cdot 2x = 10x\).
The sum rule¶
You can differentiate term by term. Example: for \(f(x) = 3x^2 + 4x + 7\),
This single rule is why the derivative of a sum-of-squared-errors cost function can be computed term by term — differentiate the error contributed by each data point separately, then sum (used constantly from Module 9 onward).
The chain rule (introduction)¶
For a composite function \(h(x) = f(g(x))\) (Module 5's terminology), the chain rule says:
In words: differentiate the "outer" function (treating the inner function as a single blob), then multiply by the derivative of the "inner" function. This is the single most important rule for ML — it is literally the mathematical mechanism behind backpropagation (Level 3), since a neural network is one long composition of functions.
Worked chain rule example¶
Let \(h(x) = (3x+1)^2\). This is a composition: outer function \(f(u) = u^2\), inner function \(g(x) = 3x+1\), with \(u = g(x)\).
- \(f'(u) = 2u\) (power rule), so \(f'(g(x)) = 2(3x+1)\).
- \(g'(x) = 3\) (sum rule + constant multiple rule: derivative of \(3x\) is 3, derivative of constant \(1\) is 0).
Check by expanding first. \((3x+1)^2 = 9x^2 + 6x + 1\). By the power/sum rules directly: \(\frac{d}{dx}[9x^2+6x+1] = 18x + 6\) — matches exactly.
Worked numeric example¶
Let's verify \(h'(x) = 18x+6\) for \(h(x)=(3x+1)^2\) at \(x=2\), both symbolically and via finite differences (Module 6's technique).
Symbolic: \(h'(2) = 18(2)+6 = 42\).
Finite difference with \(h=0.001\): \(h(2.001) = (3(2.001)+1)^2 = (7.003)^2 = 49.042009\), and \(h(2) = 7^2 = 49\), so \(\frac{h(2.001)-h(2)}{0.001} = \frac{0.042009}{0.001} = 42.009\) — very close to the exact \(42\).
import numpy as np
def h(x):
return (3 * x + 1) ** 2
def h_prime_exact(x):
return 18 * x + 6
def h_prime_numeric(x, eps=1e-3):
return (h(x + eps) - h(x)) / eps
x0 = 2.0
print("exact h'(2):", h_prime_exact(x0))
print("numeric h'(2):", h_prime_numeric(x0))
Expected output (matches hand computation):
How It Actually Works¶
A computer algebra system applies the rules from this module (power, sum,
product, chain) not to a formula written as text, but to an expression
tree — a data structure where each node is an operation (+, *, **,
sin, ...) and each leaf is a variable or constant. Differentiating
\(f(x) = (3x^2+1)^5\) means walking this tree: the root node is pow, whose
derivative rule says "multiply by the exponent, reduce the power by one,
and multiply by the derivative of the inner subtree" — i.e. the chain rule
is applied as a local, per-node rewrite, recursively, exactly the way
you apply it by hand, just done systematically over the whole tree instead
of by pattern-matching a page of algebra.
This is mechanically identical to what a deep learning framework's autodiff engine does, with one key difference: a symbolic system builds and returns the new tree representing \(f'(x)\) as a formula (which you could then print, simplify, or evaluate at many points), whereas autodiff (Level 1 Module 08 onward) evaluates the derivative rule at each node numerically, for one specific input, immediately, without ever materializing a symbolic formula. Symbolic differentiation of a deeply composed function like a 50-layer neural network would produce an expression with an astronomically large number of terms (each chain-rule application can roughly double term count); autodiff sidesteps this entirely by carrying only numeric derivative values through the same computational graph, which is why every ML framework uses autodiff, not a symbolic differentiator, to train models.
Exercise¶
- Using the power, constant-multiple, and sum rules, differentiate \(f(x) = 4x^3 - 2x^2 + 7x - 5\) by hand.
- Using the chain rule, differentiate \(g(x) = (2x - 5)^3\) by hand (outer function \(u^3\), inner function \(2x-5\)).
- Expand \((2x-5)^3\) fully and differentiate term-by-term with the power rule; confirm it matches your chain-rule answer from step 2.
- Write
g_prime_numericusing the finite-difference trick and confirm it matches your symbolic answer at \(x=1\).