Top of the class,
bottom in the job
Every model can be made to score better. The question that decides whether your team is making progress or fooling itself is: better on what?
Slide three is the model scoreboard for “Sell faster”, the listing model from the last two sprints. It only goes one way:
sprint 4 0.75 → sprint 5 0.80 → sprint 6 0.83 → sprint 7 0.92
“We deepened the trees and added seller history.” Everyone is pleased. You sign off.
Meanwhile the operations dashboard, which measures whether the listings the model flagged actually sold, has not moved since sprint 4. Not down. Just flat, for three months, while the scoreboard climbed 17 points.
Both numbers are correct. They are measuring different things, and only one of them is your business.
A model can always memorise harder
Last lesson ended on a loose thread: a decision tree with no depth limit scored 0.582, worse than the plain linear model. Here is why.
A tree with unlimited depth keeps splitting until each leaf holds a handful of listings — eventually, one listing. At that point it is not a model of the resale market. It is a lookup table of the 21,000 listings it was shown, and it answers every one of them perfectly.
Ask it about a listing it has never seen and it finds the nearest memorised pocket, which was built out of one or two items and the random noise attached to them. That is overfitting: learning the training set rather than the thing the training set is a sample of.
Real-World Machine Learning (ch. 4) names the trap in the sprint review exactly — double-dipping the training data: if the same data is used both to fit a model and to judge it, the judgement comes out too optimistic, and you end up choosing a model that disappoints on new data.
Not wrong. Optimistic — and optimistic in a way that gets worse precisely as the model gets more complex, which is to say precisely as your team works harder.
The gap is the measurement
Train score minus new-data score is not a curiosity. It is the overfitting, quantified, and it is the number worth putting on the slide:
| Tree depth | Leaves | On training data | On new listings | Gap |
|---|---|---|---|---|
| 1 — one question | 2 | 0.600 | 0.604 | −0.004 |
| 3 | 8 | 0.726 | 0.722 | 0.004 |
| 6 — the right answer | 64 | 0.769 | 0.749 | 0.020 |
| 10 | 578 | 0.830 | 0.723 | 0.107 |
| 14 | 1,815 | 0.915 | 0.673 | 0.242 |
| 20 | 3,817 | 0.987 | 0.621 | 0.366 |
| none — grow freely | 5,217 | 1.000 | 0.590 | 0.410 |
0.92 is the depth-14 row.A single yes/no question outperforms the model that scored perfectly. If you remember one pairing from this lesson, make it that one.
So hold some data back — and then distrust it
The fix is the obvious one. Keep 20–30% of the listings out of training entirely, fit on the rest, and score on the part the model has never seen. RWML calls this the holdout method, and it is the minimum bar: no team should report a model number that was computed on data the model was fitted to.
But the holdout has a weakness that costs real money, and it is the reason that sprint-over-sprint scoreboards mislead even after you fix the first problem. The holdout estimate is unbiased but noisy. Which 30% of listings happened to land in the test set changes the answer.
Here is that noise, measured. Same 2,000 listings, same model, same code — only the random seed of the split changes, twenty times:
k-fold: hold out every row, exactly once
Cross-validation removes the luck by refusing to pick one split. Shuffle the listings into k equal piles. Hold out pile 1, train on the other nine, score pile 1. Then hold out pile 2, and so on. Average the ten scores.
Every listing is tested on exactly once, by a model that never saw it. You pay for it in compute — ten fits instead of one — and for anything short of a large neural network, that is the cheapest insurance in the whole stack. RWML’s guidance: “Use at least 10 folds (or more) when you can.”
What this costs in a real quarter
One last measurement, because it is the one that changes how you run a review. On this data a depth-5 tree is genuinely better than a depth-6 tree on new listings — by 0.013 AUC. Small, but real and repeatable.
Score both on a single random holdout, 200 different times. The worse model looks better in 23% of them.
So when the scoreboard moves from 0.70 to 0.72 after a sprint of work, roughly one time in four that is not the work. It is the split. The team will nevertheless write a changelog entry explaining why their change caused it, and the next sprint will build on a conclusion that was never there.
What to ask in the next standup
Was this number computed on data the model was trained on?
The only question on this list that is not optional. If the answer is yes, or is unclear, the number is not evidence of anything yet.What does it score on training data, and what does it score held out?
Ask for both, always, as a pair. The gap between them is the overfitting, and it belongs on the slide beside the headline.Is that a single split or cross-validated — and what is the spread across folds?
A mean without a spread cannot settle an argument. Two models closer together than the fold-to-fold variation are tied, whatever the decimal says.How many examples are in the test set?
150 held-out rows gave an estimate that wobbled by 0.036; 9,000 gave 0.004. On a small evaluation set, most of what you are reading is sampling noise. This applies just as hard to a 40-prompt LLM eval — Lesson 21.How many times have we looked at this test set?
Tune against a holdout for three months and you have fitted to it by hand. That is why teams keep a final set that is opened once, before launch, and not again.
15-minute hands-on
Paste into a free Google Colab notebook. It rebuilds Lesson 03’s resale catalogue, then does the thing you can never do at work: generates 200,000 listings from the same underlying truth, so you can see the answer the holdout is trying to guess. Every number in this lesson is printed by this script.
import numpy as np, pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split, cross_val_score, KFold
from sklearn.metrics import roc_auc_score
F = ["brand_tier", "discount", "condition", "photo_count"]
def catalogue(n, seed):
"""Lesson 03's resale catalogue. We control the truth here, so we can look at
the future. In real life you cannot — and that is the entire problem."""
rng = np.random.default_rng(seed)
brand_tier = rng.uniform(0, 1, n)
discount = rng.uniform(0, 60, n)
condition = rng.integers(1, 6, n)
photo_count = rng.integers(1, 10, n)
hs = 0.055 * discount - 1.4
lux = 1.5 - 0.0062 * (discount - 18.0) ** 2
logit = (-0.55 + 0.42*(condition-3) + 0.11*(photo_count-5) - 0.9*brand_tier
+ (1-brand_tier)*hs + brand_tier*lux + rng.normal(0, 0.35, n))
sold = (rng.uniform(size=n) < 1/(1+np.exp(-logit))).astype(int)
return pd.DataFrame(dict(brand_tier=brand_tier, discount=discount,
condition=condition, photo_count=photo_count, sold=sold))
df = catalogue(30_000, 11) # the catalogue you have today
future = catalogue(200_000, 2027) # next quarter's listings — the oracle
tr, te = train_test_split(df, test_size=0.3, random_state=11, stratify=df.sold)
print(f"{'depth':>6} {'leaves':>7} {'train':>7} {'holdout':>8} {'10-fold':>8} {'NEW':>7}")
for d in [1, 2, 3, 4, 5, 6, 8, 10, 14, 20, None]:
m = DecisionTreeClassifier(max_depth=d, random_state=11).fit(tr[F], tr.sold)
cv = cross_val_score(DecisionTreeClassifier(max_depth=d, random_state=11), tr[F],
tr.sold, cv=KFold(10, shuffle=True, random_state=7),
scoring="roc_auc").mean()
print(f"{str(d):>6} {m.get_n_leaves():>7} "
f"{roc_auc_score(tr.sold, m.predict_proba(tr[F])[:,1]):>7.3f} "
f"{roc_auc_score(te.sold, m.predict_proba(te[F])[:,1]):>8.3f} {cv:>8.3f} "
f"{roc_auc_score(future.sold, m.predict_proba(future[F])[:,1]):>7.3f}")
# How much does ONE holdout wobble? 2,000 listings, same model, 20 different splits.
small = catalogue(2_000, 11)
hold, cvs = [], []
for s in range(20):
a, b = train_test_split(small, test_size=0.3, random_state=s, stratify=small.sold)
m = DecisionTreeClassifier(max_depth=5, random_state=11).fit(a[F], a.sold)
hold.append(roc_auc_score(b.sold, m.predict_proba(b[F])[:,1]))
cvs.append(cross_val_score(DecisionTreeClassifier(max_depth=5, random_state=11),
small[F], small.sold, cv=KFold(10, shuffle=True, random_state=s),
scoring="roc_auc").mean())
hold, cvs = np.array(hold), np.array(cvs)
truth = roc_auc_score(future.sold, DecisionTreeClassifier(max_depth=5, random_state=11)
.fit(small[F], small.sold).predict_proba(future[F])[:,1])
print(f"\n2,000 listings, identical model, 20 different random splits:")
print(f" one 30% holdout : {hold.min():.3f} – {hold.max():.3f} (spread {np.ptp(hold):.3f})")
print(f" 10-fold CV : {cvs.min():.3f} – {cvs.max():.3f} (spread {np.ptp(cvs):.3f})")
print(f" truth : {truth:.3f}")
# The ten fold scores behind ONE of those cross-validated numbers
folds = cross_val_score(DecisionTreeClassifier(max_depth=5, random_state=11), small[F],
small.sold, cv=KFold(10, shuffle=True, random_state=0),
scoring="roc_auc")
print(" ten folds:", " ".join(f"{f:.3f}" for f in folds), f"-> mean {folds.mean():.3f}")
# Depth 5 really is better than depth 6. How often does one holdout disagree?
wrong = sum(
roc_auc_score(b.sold, DecisionTreeClassifier(max_depth=6, random_state=11).fit(a[F],a.sold).predict_proba(b[F])[:,1])
> roc_auc_score(b.sold, DecisionTreeClassifier(max_depth=5, random_state=11).fit(a[F],a.sold).predict_proba(b[F])[:,1])
for a, b in (train_test_split(small, test_size=0.3, random_state=s, stratify=small.sold)
for s in range(200)))
print(f"\ndepth 5 beats depth 6 on new listings by 0.013 AUC — yet on a single random")
print(f"holdout, depth 6 looks better {wrong}/200 = {wrong/2:.0f}% of the time.") Two experiments worth running after it. Change catalogue(2_000, 11) to catalogue(500, 11) and watch the holdout spread roughly double — small evaluation sets are where teams fool themselves hardest. Then change the rng.normal(0, 0.35, n) noise term to 0.9 and re-run the depth sweep: the noisier the world, the shallower the tree you can afford. How much complexity you can support is a property of the data, not of the algorithm.
Any model can be made to score better on data it has already seen — that direction has no floor. The only number that means anything is the one computed on examples the model has never been shown, and the only way to stop that number lying to you is to compute it many times and look at the spread.
Vocabulary from this lesson
These lessons were written with the help of AI (Claude) and reviewed and edited by me.