Lesson 1 of 5 · 20 min
What learning means
A robot reads a number from a sensor and has to decide what it means. For a datasheet part, someone already wrote the rule. For your own cheap sensor on your own robot, nobody did, and the rule has to come from measurements. Machine learning is the craft of getting that rule out of data instead of out of a manual. This lesson strips the idea to three pieces, a model, a loss and a search, and builds all of them in a few lines of plain Python.
A model is a function with knobs
Say an analog temperature sensor gives you a raw reading, and you want degrees Celsius. The simplest plausible relationship is a straight line:
y = w * x + b
Here x is the input (the reading), y is the prediction (the temperature), and w and b are the parameters: a slope and an offset. The formula is the model. The parameters are the knobs. Different knob settings give different lines, and "learning" means turning the knobs until the line fits the world.
Notice what is fixed and what is free. The shape of the function (a straight line) is chosen by you before any data arrives. Only w and b are found from data. A bigger model, like the neural networks later in this course, just has more knobs and a more flexible shape. The recipe stays the same.
The data
You put the sensor next to a trusted thermometer and write down six pairs. To keep the numbers small, the raw reading is divided by 100.
| Reading / 100 (x) | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Reference temperature, C (y) | 12.1 | 14.2 | 15.8 | 18.1 | 20.0 | 21.9 |
The points are almost on a line, but not quite. Real sensors have noise, so no choice of w and b will hit all six exactly. We need a way to say which imperfect line is better.
The loss: a number for how wrong you are
For one example, the error is prediction minus truth: e = (w * x + b) - y. Averaging raw errors would let positives cancel negatives, so we square each one and then average:
loss = (1/n) * sum of (w*x + b - y)^2
This is the mean squared error (MSE). Squaring makes every term non-negative, punishes one large miss more than several small ones, and, most importantly for the next lesson, gives a smooth function that has a derivative everywhere.
Work one example by hand. Try w = 2 and b = 10. The predictions are 12, 14, 16, 18, 20, 22, so the errors are -0.1, -0.2, 0.2, -0.1, 0.0, 0.1. Their squares add up to 0.11, and 0.11 / 6 is about 0.018. A very small loss, so this is a good line. Now the same thing in code, for four guesses.
You should see losses of 300.818, 15.085, 0.018 and 1.052. The guess (2, 10) matches the hand calculation. Look at the last line too: moving b by just one degree, from 10 to 9, multiplies the loss by almost sixty. The loss is a landscape over the knob settings, and you want the lowest point in it.
Learning is search
Once you have a loss, learning is an optimisation problem: find the w and b that make the loss as small as possible. The most naive search tries a whole grid of candidates and keeps the best. It is slow and crude, but it is honest, and it shows the principle without any calculus.
The search tries 81 x 201 = 16,281 combinations and reports w = 1.95, b = 10.2, loss 0.0146. The exact best line for this data (which statisticians can write in closed form) has w = 1.963, b = 10.147 and loss 0.0140, so the grid got close but is limited by its step size. The model can now answer a question it was never asked: a reading of 7 gives about 23.85 C.
The grid has an obvious flaw. Two parameters at 81 and 201 values is 16 thousand evaluations. Ten parameters at 100 values each would be 10^20 evaluations, and a small neural network has thousands. The cost grows exponentially with the number of knobs. The next lesson replaces the grid with a method that uses the slope of the loss to walk downhill, which costs almost nothing per knob.
Fitting is not the same as understanding
A low loss on the data you fitted says only that the line passes near those six points. Two warnings follow from that.
First, extrapolation. The data covers readings 1 to 6. Asking about reading 7 is a small step outside, and a straight line is a fair guess. Asking about reading 50 trusts the straight line far beyond the evidence, and real sensors saturate, so that answer could be nonsense.
Second, a model flexible enough to hit every point exactly would also fit the noise, then fail on the next measurement. We will measure that failure directly in lesson 3 by keeping some data hidden from the search. For now, remember the order: choose a model, define a loss, search for parameters, then check on data the search never saw.
Check yourself
In the model y = w * x + b, which of these is a parameter that learning must find?
Check yourself
A model makes errors of 1, -2 and 2 on three examples. What is the mean squared error?