A neuron is a
weighted sum
Week 1 was about models that draw lines and models that split. This week starts the machine everything else in this course is built from — and it turns out to be one line of arithmetic you already know.
Your ML lead wants to replace the mispricing rules engine — the one that flags listings priced out of line with their condition — with “a small neural network.” The slide says deep learning. The ask is a GPU budget and a quarter.
You do not yet know enough to say yes or no, and that is the uncomfortable part. Not because the maths is hard, but because nobody has ever told you what the object actually is.
So: here is the whole thing. It fits on one line, you can check it with a calculator, and by the end of this lesson you will know the single question that decides whether that budget is money well spent.
The whole object, in one line
A neuron takes some numbers, multiplies each by a weight, adds them up, adds one more number called the bias, and passes the total through a function that squashes it.
z = (x₁·w₁ + x₂·w₂ + … + xₙ·wₙ) + b
output = activation(z) That is it. The sum of pairwise multiplications is the dot product of the input vector and the weight vector — which is why linear algebra runs through all of this, and why GPUs, built to do exactly that operation millions of times at once, became the hardware of AI. NLP in Action (ch. 5) puts it plainly: the base unit of any neural network is the neuron, and it is a dot product with a decision bolted on.
Make it concrete. Suppose a vision model has already produced five wear scores for a listing photo, each from 0 to 1, and a neuron decides whether the listing goes to a human grader or lists automatically.
stains carries a weight of 3.0. A neuron counts evidence by weight, not by headcount.You have already met this object
Take the neuron above, use a sigmoid as the activation, and write down what it computes: a weighted sum of features, squashed into a probability. That is logistic regression — the straight-line classifier from Lesson 03, under a different name. Build a Large Language Model (From Scratch) (Appendix A) introduces its very first computation graph with exactly this move, and points out that this logistic regression classifier is, read another way, a neural network with a single layer.
So one neuron buys you nothing you did not already have in week one. It draws one straight boundary through your feature space and calls everything on one side a yes. The entire question — the one your GPU budget turns on — is what you get when you stack them.
The problem one line cannot solve
Back to the mispricing flag. Two numbers describe a listing: how its asking price compares to the market median for that model, and how good its condition is. Both centred, so negative means below average and positive means above.
A listing is mispriced when those two disagree. A pristine Acne coat at 40% of market is underpriced — you are giving away margin. A bobbled, faded one at 130% is overpriced — it will sit for six months. Cheap-and-worn is fine. Expensive-and-perfect is fine.
Now try to draw that rule as a single straight line. Mispriced listings sit in the top-left and bottom-right corners; fair ones in the top-right and bottom-left. Any straight line you draw puts one mispriced corner and one fair corner on the same side. There is no angle that works. This is not a hard problem — it is an impossible one, for a single neuron.
What the hidden layer actually bought
Four ways to attack the same 6,000 listings, scored on 3,000 held-out ones. The fourth row is the ceiling: what the true rule itself scores, given that 4% of labels are noisy and listings sitting right on an axis are genuinely ambiguous. Nothing can beat it.
| Model | Features it was given | Held-out accuracy |
|---|---|---|
| 1 neuron | price, condition | 41.1% |
| 1 neuron | price, condition, price × condition | 83.6% |
| 1 hidden layer of 4 neurons | price, condition | 83.5% |
| the true rule, as an oracle | — | 83.9% |
So look at what the four hidden neurons found. Each one is itself just a weighted sum, and two of them are readable in plain English:
hidden 3: +1.83·price +2.66·condition −1.33 fires when BOTH are high
hidden 4: −2.62·price −1.91·condition −1.39 fires when BOTH are low Both feed the output neuron with a negative weight. Hidden 3 fires at +0.95 for the expensive-and-pristine corner and pushes the answer down to “fair.” Hidden 4 fires at +0.94 for the cheap-and-worn corner and does the same. In the two mispriced corners, neither fires — both read about −1 — nothing pushes the answer down, and it comes out at 0.97 and 0.99.
The network did not learn “mispriced.” It learned two agreement detectors, and defined mispriced as neither of them fired. That is a hand-built feature, built by gradient descent instead of by a person in a planning meeting.
Which is the whole deal: a hidden layer is feature engineering that trains itself. Lesson 02 said features are the product. Lesson 03 said tree models reach interactions by splitting twice. A neural network reaches them by stacking weighted sums — and you pay for that convenience in data, in compute, and in your ability to explain what it did.
Why the squashing function is not optional
One detail decides whether any of this works. Stack two neurons with no activation function between them and you get a weighted sum of a weighted sum — which is just another weighted sum. A hundred layers of that collapse back into one straight line. The nonlinearity is the entire reason depth buys anything.
The step function in the original perceptron squashes to 0 or 1, but it has a flat gradient everywhere, so there is no signal telling the weights which way to move. NLP in Action (ch. 5) makes the swap explicit: backpropagation needs an activation that is “nonlinear and continuously differentiable,” which is why the sigmoid, tanh and ReLU exist. That is also why the next lesson is about gradients — in this one, how the weights were found is a black box on purpose.
The part your ML lead’s slide left out
Three honest costs, all visible in the experiment above.
Capacity is guesswork. Two hidden neurons is the textbook minimum for this problem. Run it on four different random starts and it scores 79.2%, 79.9%, 65.6% and 79.2% — sometimes it finds the structure, sometimes it lands in a bad spot and stays there. Four neurons hits 83.5% on every seed. Nothing told us that in advance; somebody tried it.
On tables, the boring thing often wins. Row 2 of that table — one human noticing that price and condition interact — matched the network for a fraction of the cost, and produced a model you can read. On tabular data with tens of columns, gradient-boosted trees (Lesson 03) remain the thing to beat, and they usually win.
Neural networks earn their budget where features cannot be hand-built. Nobody can write down the column that distinguishes a genuine Bottega weave from a convincing copy, or encode what a listing’s description means. Pixels, text, audio: there the learned-feature machine is not a luxury, it is the only option — and that is the road to embeddings, transformers and LLMs over the next four weeks.
The question to put to the slide, then, is not “is a neural network better.” It is: can a person write down the feature we are missing? If yes, that is a sprint, not a quarter. If no, you are in the right building.
What to ask in the next standup
Which feature would a person have to hand-build for a linear model to solve this?
If someone can name it in the meeting, build it and skip the quarter. A neural network is worth its cost precisely when the honest answer is “nobody can write that down.”What is our non-neural baseline, and by how much are we beating it?
Logistic regression and gradient-boosted trees, from Lesson 03, measured on the same split. Lesson 01’s rule still holds: a number with nothing to compare it to is not a result.How many layers and neurons, and how did we choose?
“It’s what the tutorial used” is a real answer and a fine starting point — as long as it is said out loud rather than presented as design. Expect the honest version to be “we tried four sizes.”How much does this change between random seeds?
Our two-neuron net swung 14 points across four starts. Ask for the spread, not the best run — same discipline as the k-fold spread in Lesson 04.When it flags a listing, what will we tell the seller?
A tree gives you a reason. A network gives you a number. Decide before launch whether this decision needs an explanation — for pricing nudges, probably not; for suspending an account or failing an authentication, almost certainly yes.
20-minute hands-on
Pure numpy — no PyTorch, no framework, nothing hidden. Paste into a free Google Colab notebook; it runs in about thirteen seconds on a CPU and prints every number in this lesson. The training loop is deliberately short and deliberately unexplained: that is Lesson 07.
import numpy as np
sig = lambda z: 1 / (1 + np.exp(-z))
# ---- PART 1 --- one neuron, run by hand ---------------------------------
FLAGS = ["pilling", "fading", "stains", "seam_wear", "odour"]
weights = np.array([1.4, 0.6, 3.0, 1.3, 2.2]) # learned earlier; here we just run them
bias = -2.0
for name, x in {"A visibly worn parka": np.array([0.85, 0.55, 0.20, 0.70, 0.15]),
"B near-pristine coat": np.array([0.05, 0.10, 0.00, 0.05, 0.00]),
"C one mark, else clean": np.array([0.10, 0.05, 0.70, 0.05, 0.10])}.items():
z = x @ weights + bias # <- the whole neuron is this line
print(f"{name:<24} z={z:+.3f} sigmoid={sig(z):.3f} "
f"{'REVIEW' if z >= 0 else 'auto-list'}")
# ---- PART 2 --- the mispricing flag -------------------------------------
def listings(n, seed):
"""price and condition, both centred. mispriced = the two DISAGREE."""
r = np.random.default_rng(seed)
p, c = r.uniform(-1, 1, n), r.uniform(-1, 1, n)
mis = (np.sign(p) != np.sign(c)).astype(float)
amb = (np.abs(p) < .15) | (np.abs(c) < .15) # genuinely borderline listings
mis[amb] = r.integers(0, 2, amb.sum())
flip = r.uniform(size=n) < .04 # 4% of labels are simply wrong
mis[flip] = 1 - mis[flip]
return np.c_[p, c], mis
def train(X, y, hidden, epochs=6000, lr=3.0, seed=0):
"""Full-batch gradient descent. HOW this works is Lesson 07 — today it is a box."""
r, (n, d) = np.random.default_rng(seed), X.shape
if not hidden: # one neuron = logistic regression
w, b = r.normal(0, 1, d), 0.0
for _ in range(epochs):
g = (sig(X @ w + b) - y) / n
w -= lr * (X.T @ g); b -= lr * g.sum()
return (lambda Z: sig(Z @ w + b)), (w, b)
W1, b1 = r.normal(0, 1, (d, hidden)), r.normal(0, .5, hidden)
W2, b2 = r.normal(0, 1, (hidden, 1)), np.zeros(1)
for _ in range(epochs):
H = np.tanh(X @ W1 + b1) # the hidden layer
o = sig(H @ W2 + b2).ravel() # the output neuron
d2 = (o - y)[:, None] / n
d1 = (d2 @ W2.T) * (1 - H**2) # backprop through tanh
W2 -= lr * (H.T @ d2); b2 -= lr * d2.sum(0)
W1 -= lr * (X.T @ d1); b1 -= lr * d1.sum(0)
return (lambda Z: sig(np.tanh(Z @ W1 + b1) @ W2 + b2).ravel()), (W1, b1, W2, b2)
Xtr, ytr = listings(6000, 6)
Xte, yte = listings(3000, 77)
acc = lambda f, X: ((f(X) >= .5) == yte).mean() * 100
truth = (np.sign(Xte[:, 0]) != np.sign(Xte[:, 1])).astype(float)
print(f"\nceiling — the true rule itself scores {(truth == yte).mean()*100:.1f}%")
one, p1 = train(Xtr, ytr, hidden=0, seed=1)
print(f"1 neuron, [price, condition] {acc(one, Xte):.1f}% "
f"weights: price {p1[0][0]:+.3f} condition {p1[0][1]:+.3f}")
I = lambda X: np.c_[X, X[:, 0] * X[:, 1]] # the hand-built interaction column
hand, ph = train(I(Xtr), ytr, hidden=0, seed=1)
print(f"1 neuron, + hand-built price x condition {acc(hand, I(Xte)):.1f}% "
f"weight on the product: {ph[0][2]:+.2f}")
net, (W1, b1, W2, b2) = train(Xtr, ytr, hidden=4, seed=0)
print(f"4 hidden neurons, raw columns only {acc(net, Xte):.1f}%")
for j in range(4):
print(f" hidden {j+1}: {W1[0,j]:+.2f}*price {W1[1,j]:+.2f}*condition {b1[j]:+.2f}"
f" -> output weight {W2[j,0]:+.2f}")
corners = np.array([[.7, .7], [-.7, -.7], [.7, -.7], [-.7, .7]])
H = np.tanh(corners @ W1 + b1)
for nm, h, o in zip(["expensive + pristine (fair)", "cheap + worn (fair)",
"expensive + worn (OVER)", "cheap + pristine (UNDER)"], H, net(corners)):
print(f" {nm} hidden=({', '.join(f'{v:+.2f}' for v in h)}) -> {o:.2f}")
for sd in range(4): # how much does the size matter?
f, _ = train(Xtr, ytr, hidden=2, seed=sd)
print(f" (2 hidden neurons, seed {sd}: {acc(f, Xte):.1f}%)", end="")
print() Four things to try. Change hidden=4 to 8 and then 32 — accuracy stops improving at the ceiling, which is what “more neurons” buys you once the structure is already found. Replace np.tanh with lambda v: np.maximum(0, v) and the matching gradient (H > 0) to swap in ReLU, the activation every modern network actually uses. Then delete the activation entirely — use H = X @ W1 + b1 — and watch a two-layer network collapse back to 41%, because without the squash it is one layer. Finally, add a third feature that is pure noise and see how much accuracy it costs: that is Lesson 04 showing up again, in a new machine.
A neuron is a dot product, a bias and a squash — logistic regression wearing a different hat. One of them draws one straight line and can do nothing more. Stack a layer of them and the model starts inventing the interaction columns a human would otherwise have to notice, name and build by hand. That is the whole trade, and the whole bill.
Vocabulary from this lesson
These lessons were written with the help of AI (Claude) and reviewed and edited by me.