Lesson 4 of 5 · 28 min
A neuron and a tiny network
The word "neural network" sounds biological and mysterious. The thing itself is arithmetic you already know: multiply, add, and squash through a simple curve. In this lesson you build one artificial neuron, teach it two logic gates, watch it fail on a third, and then fix the failure by stacking neurons into layers. The failure is the interesting part, because it explains why networks have the shape they have.
One neuron
A neuron takes inputs x1 and x2, multiplies each by a weight, adds a bias, and passes the total through an activation function f:
z = w1 * x1 + w2 * x2 + b
out = f(z)
You have seen this already: it is the linear model from lesson 1 with two inputs and an extra function applied to the result. The weights and bias are the parameters to learn. Two activation functions matter for us.
The step function outputs 1 when z is at least 0, and 0 otherwise. It is the original perceptron, a hard yes/no decision. The sigmoid is a smooth version:
sigmoid(z) = 1 / (1 + e^(-z))
It maps any number into the range 0 to 1, with sigmoid(0) = 0.5, sigmoid(2) = 0.881, sigmoid(-2) = 0.119. Its derivative has the neat form s * (1 - s), never larger than 0.25. Smoothness matters because, as lesson 2 showed, gradient descent needs slopes. A step function has a slope of zero almost everywhere, so there is nothing to follow.
A forward pass is just evaluating the neuron. Take w1 = 6, w2 = 6, b = -3. For input (1, 0), z = 6 - 3 = 3 and sigmoid(3) = 0.953. For (0, 0), z = -3 and sigmoid(-3) = 0.047. This neuron says yes when either input is on, which is an OR gate. We will see how it can find such weights on its own.
Teaching a neuron a rule
A logic gate is a perfect toy problem because the "dataset" is a four-row truth table. The perceptron learning rule is the oldest training algorithm. For each example, compute the output, and if it is wrong, nudge the weights toward the right answer:
w = w + (target - out) * x, and b = b + (target - out)
If the neuron wrongly says 0 when the target is 1, the error is +1 and each weight grows by its input, making z larger next time. If it wrongly says 1, the weights shrink. Right answers leave everything alone. The code trains on AND, OR and XOR, and prints the number of mistakes per pass through the table, called an epoch.
AND is learned in 6 epochs and finishes at w1 = 2, w2 = 1, b = -3. Verify it: for (1, 1), z = 2 + 1 - 3 = 0, which the step function maps to 1. For (1, 0), z = -1, giving 0. For (0, 1), z = -2, giving 0. OR is learned in 4 epochs with w1 = 1, w2 = 1, b = -1. XOR never converges: after the first couple of epochs it keeps making 3 or 4 mistakes in every pass, indefinitely. This is not bad luck or a bad learning rate. It is impossible, and the next section shows why.
Why XOR breaks a single neuron
The equation w1 * x1 + w2 * x2 + b = 0 is a straight line in the plane of inputs. A single neuron answers 1 on one side of the line and 0 on the other. AND and OR have a line that separates the 1 outputs from the 0 outputs. XOR does not: its 1s sit at (0, 1) and (1, 0), its 0s at (0, 0) and (1, 1), diagonally opposite, and no line separates the diagonals.
You can prove it with four inequalities, using the step rule that output 1 needs z at least 0. We need b < 0 for (0, 0), then w1 + b >= 0 and w2 + b >= 0 for the two ones, and w1 + w2 + b < 0 for (1, 1). Add the middle two: w1 + w2 + 2b >= 0. Rearrange: w1 + w2 + b >= -b, and since b is negative, -b is positive. That forces w1 + w2 + b to be positive, contradicting the last requirement. No weights exist.
A sigmoid neuron with gradient descent
The perceptron rule only works for the step function. Modern networks use smooth activations and the gradient method from lesson 2. For one neuron, the loss is the mean squared error between out and the target. The chain rule has one extra link compared with lesson 2: the error passes backward through the sigmoid. The derivative of the loss with respect to z is 2 * (out - t) * out * (1 - out), and then the weight gradient is that value times the input, just as before. Here is a single sigmoid neuron learning OR.
The loss falls from 0.2575 at the start to 0.0013 after 2000 epochs. The neuron settles on weights near 6.18 for both inputs and a bias near -2.84. It outputs 0.055 for (0, 0) and 0.966 or above for the other three, which is a confident OR gate. Notice that it never reaches exactly 0 or 1. Sigmoid saturates gradually, and a trained network's outputs are better read as confidence than as a hard answer.
Stacking neurons: a hidden layer
If one line cannot solve XOR, use two lines and combine them. The XOR rule in words is "at least one input on, but not both". That is OR and not AND. So give the network a hidden layer of two neurons, one acting as OR and one as AND, and an output neuron that fires when OR is on and AND is off. The weights below are set by hand, using large values so the sigmoids are nearly step functions.
The table shows XOR outputs of 0.000, 1.000, 1.000 and 0.000. Check one row by hand. For (1, 1), the OR neuron has z = 20 + 20 - 10 = 30 and the AND neuron z = 40 - 30 = 10, so both hidden values are about 1. The output neuron computes 20 - 20 - 10 = -10, and sigmoid(-10) is about 0.00005, which prints as 0.000. The hidden layer has transformed the inputs into a new pair of features, OR and AND, in which the classes are separable by a single line. That is what hidden layers do: they learn a better representation, and the last neuron has an easy job.
Why the activation matters: without a nonlinear function between layers, two layers collapse. W2 * (W1 * x) equals (W2 * W1) * x, which is one linear layer in disguise, and the XOR problem would return. Training such a network uses the same chain rule, applied layer by layer from the output backward. That procedure is called backpropagation, and it is nothing more than lesson 2 repeated.
Check yourself
Why can a single neuron not learn XOR?
Check yourself
What happens if you remove the activation function between the two layers of a network?