Lesson 2 of 5 · 28 min
Linear regression and gradient descent
Last lesson ended with a brute-force grid and a warning that it cannot scale. Here is the method that actually trains models from a line fit up to networks with millions of parameters: gradient descent. The idea fits in one sentence. Stand somewhere on the loss landscape, feel which direction is downhill, take a small step that way, and repeat.
The loss as a landscape
We keep the same sensor-calibration data and the same model, y = w * x + b, with six pairs of reading and reference temperature. The loss is the mean squared error:
L(w, b) = (1/n) * sum of e_i^2, where e_i = w * x_i + b - y_i
Think of w and b as two compass coordinates and L as the height of the ground. For this model the ground is a smooth bowl with a single lowest point. At any position, the gradient is the pair of slopes (dL/dw, dL/db). It points straight uphill, and its length says how steep it is. Walking the opposite way reduces the loss.
Deriving the gradient
You only need the chain rule. For one example, the error is e = w * x + b - y, and its square is e^2. Differentiate with respect to w:
d(e^2)/dw = 2 * e * de/dw = 2 * e * x
because e changes by x for every unit change in w. For b, the error changes by 1 per unit, so:
d(e^2)/db = 2 * e * 1 = 2 * e
The loss is an average of these squares, and the derivative of an average is the average of the derivatives:
dL/dw = (2/n) * sum of e_i * x_i
dL/db = (2/n) * sum of e_i
Read them as stories. If your predictions are too high on average (positive errors), dL/db is positive, so you should lower b. The dL/dw formula weights each error by its input, because a mistake on a large reading is more the slope's fault than the offset's. The update rule moves against the gradient by a factor called the learning rate lr:
w = w - lr * dL/dw
b = b - lr * dL/db
Gradient descent in pure Python
Start at w = 0, b = 0, where the loss is 300.8, and use lr = 0.02. Before running, check step 1 by hand: the errors are just -y, the sum of y * x is 391.7, so dL/dw = -(2/6) * 391.7 = -130.57 and w becomes 0 + 0.02 * 130.57 = 2.611. The code should agree.
Step 1 shows w = 2.6113, matching the hand calculation, and the loss drops from 300.8 to 53.0. Then something instructive happens. By step 5, w has overshot to 4.0 (the true value is near 1.96) while b is still only 1.28. After 50 steps the loss is 7.8, after 200 it is 0.88, and only by step 1000 does it reach w = 1.9643, b = 10.1406, loss 0.0140, essentially the best possible. Why so slow? Keep that question; we answer it below.
Too big, too small
The learning rate is the one number you must choose. Run the same loop for four values and compare after 200 steps.
The four rows read: lr = 0.0001 gives loss 95.08 (barely started, b is only 0.52), lr = 0.02 gives 0.881, lr = 0.06 gives 0.01635 (the best of the four), and lr = 0.07 gives w around -1.4e19 and a loss near 3e39. The first is too small: each step is tiny, so you would need hundreds of thousands of steps. The last is too big: each step overshoots the valley and lands higher up the opposite wall, so the error grows every iteration until the numbers overflow.
There is a clean reason for the cliff. Along any direction where the bowl has curvature c (the second derivative), one step multiplies the distance from the minimum by 1 - lr * c. That shrinks only when its magnitude is below 1, which means lr < 2 / c. For this data the steepest curvature is about 32, so the limit is 2 / 32, roughly 0.0626. That is why 0.06 works and 0.07 explodes. Notice that 0.06 is the fastest only because it sits just under the cliff, and the oscillation there is damped, not absent.
The same curvature explains the slow crawl above. The bowl is steep along one direction (about 32) and almost flat along another (about 0.37), a ratio near 87. The learning rate must respect the steep direction, which leaves the flat direction, mostly the offset b, moving at a crawl. The cure is cheap: centre the inputs.
numpy and centred inputs
If we subtract the mean of x before fitting, the two directions decouple and have similar curvature. The code also becomes shorter with numpy, because err * xc and np.mean act on all six examples at once.
Now lr = 0.1, ten times higher than before, is safe, and the fit is complete in about 40 steps: w settles at 1.9629 and b at 17.0167, which is simply the mean of y, because the centred line passes through the average point. Converting back, the intercept in original units is 10.147, and the model predicts 23.89 C for reading 7, close to the grid-search answer of 23.85 from lesson 1. Scaling inputs is not a trick for toy problems. Every serious pipeline does it, including the one you will meet for sensor data on a robot.
Check yourself
You run gradient descent and the loss gets larger on every step, ending in nan. What is the most likely cause and first fix?
Check yourself
Four examples have errors 2, -1, 1 and 0. Using dL/db = (2/n) * sum of errors, what is dL/db?