Lesson 04 · Machine learning foundationsRoadmap to be an AI engineer

Top of the class,
bottom in the job

Every model can be made to score better. The question that decides whether your team is making progress or fooling itself is: better on what?

Tuesday · 11:20 · sprint review

Slide three is the model scoreboard for “Sell faster”, the listing model from the last two sprints. It only goes one way:

sprint 4  0.75  →  sprint 5  0.80  →  sprint 6  0.83  →  sprint 7  0.92

“We deepened the trees and added seller history.” Everyone is pleased. You sign off.

Meanwhile the operations dashboard, which measures whether the listings the model flagged actually sold, has not moved since sprint 4. Not down. Just flat, for three months, while the scoreboard climbed 17 points.

Both numbers are correct. They are measuring different things, and only one of them is your business.

A model can always memorise harder

Last lesson ended on a loose thread: a decision tree with no depth limit scored 0.582, worse than the plain linear model. Here is why.

A tree with unlimited depth keeps splitting until each leaf holds a handful of listings — eventually, one listing. At that point it is not a model of the resale market. It is a lookup table of the 21,000 listings it was shown, and it answers every one of them perfectly.

Ask it about a listing it has never seen and it finds the nearest memorised pocket, which was built out of one or two items and the random noise attached to them. That is overfitting: learning the training set rather than the thing the training set is a sample of.

Real-World Machine Learning (ch. 4) names the trap in the sprint review exactly — double-dipping the training data: if the same data is used both to fit a model and to judge it, the judgement comes out too optimistic, and you end up choosing a model that disappoints on new data.

Not wrong. Optimistic — and optimistic in a way that gets worse precisely as the model gets more complex, which is to say precisely as your team works harder.

The scissors
One dataset of 30,000 resale listings, one algorithm, eleven settings of a single knob: how deep the tree may grow. Score is AUC — the chance the model ranks a listing that sold above one that didn’t. 0.50 is a coin flip.
1.000.900.800.700.60the gap= 0.410scored on the data it trained onon new listingspeaks here — 0.749cross-validation stops herethe training score stops here1234568101420none2 leaves645,217 leavesmaximum tree depth  →  more complex
Score on the training dataScore on 200,000 new listings — the truth10-fold cross-validation estimate
The amber line never turns down. That is the whole problem: the training score rewards complexity without limit, so a team optimising it will always conclude that more is better. The blue line is what the business actually gets, and it peaks at depth 6 and then falls off a cliff. Follow the training score to its maximum and you ship the far-right model: perfect on the data it has seen, 0.590 on new listings — worse than the two-leaf tree on the far left. The dashed line is the point of this lesson: cross-validation, computed without ever touching new data, sits almost on top of the truth at every single depth.

The gap is the measurement

Train score minus new-data score is not a curiosity. It is the overfitting, quantified, and it is the number worth putting on the slide:

Tree depthLeavesOn training dataOn new listingsGap
1 — one question20.6000.604−0.004
380.7260.7220.004
6 — the right answer640.7690.7490.020
105780.8300.7230.107
141,8150.9150.6730.242
203,8170.9870.6210.366
none — grow freely5,2171.0000.5900.410
Read down the last column. A small gap means the model learned something general. A large gap means it memorised, and the headline score on the left is borrowed against a future that will not pay it back. The sprint review’s 0.92 is the depth-14 row.
The free-growing tree, on its own training data
1.000
flawless · 5,217 leaves
The same tree, on new listings
0.590
near enough a coin flip
A tree that asks one question
0.604
beats it · 2 leaves

A single yes/no question outperforms the model that scored perfectly. If you remember one pairing from this lesson, make it that one.

So hold some data back — and then distrust it

The fix is the obvious one. Keep 20–30% of the listings out of training entirely, fit on the rest, and score on the part the model has never seen. RWML calls this the holdout method, and it is the minimum bar: no team should report a model number that was computed on data the model was fitted to.

But the holdout has a weakness that costs real money, and it is the reason that sprint-over-sprint scoreboards mislead even after you fix the first problem. The holdout estimate is unbiased but noisy. Which 30% of listings happened to land in the test set changes the answer.

Here is that noise, measured. Same 2,000 listings, same model, same code — only the random seed of the split changes, twenty times:

Twenty runs of the identical model
2,000 listings, depth-5 tree. Nothing changes between runs except which rows fall into the test set.
truth: 0.714measured on 200,000new listingsOne 30% holdout0.6520.722spread 0.07110-fold cross-validation0.6840.710spread 0.0260.640.660.680.700.720.74
Nothing about the model changed. A team that ran this on a Monday would report 0.652; the same team on Tuesday, 0.722. That is a 7-point swing — larger than most of the real improvements anyone ships in a quarter — invented entirely by the random number generator. The cross-validated estimate, made from exactly the same 2,000 rows, lands in a third of the space.

k-fold: hold out every row, exactly once

Cross-validation removes the luck by refusing to pick one split. Shuffle the listings into k equal piles. Hold out pile 1, train on the other nine, score pile 1. Then hold out pile 2, and so on. Average the ten scores.

Every listing is tested on exactly once, by a model that never saw it. You pay for it in compute — ten fits instead of one — and for anything short of a large neural network, that is the cheapest insurance in the whole stack. RWML’s guidance: “Use at least 10 folds (or more) when you can.”

Ten rounds, ten scores, one answer
The actual folds from the exercise below. Amber = the pile held out and scored in that round; grey = the piles trained on.
1
2
3
4
5
6
7
8
9
10
AUC
round 1
0.755
round 2
0.729
round 3
0.663
round 4
0.648
round 5
0.679
round 6
0.728
round 7
0.745
round 8
0.702
round 9
0.706
round 10
0.637
average
0.699
Look at the right-hand column before you look at the average. The individual folds run from 0.637 to 0.755 — and any one of them is what you would have reported had you done a single split and stopped. The average, 0.699, is the number worth saying out loud. The spread tells you how much to trust it: if two models are a tenth of that spread apart, they are tied.

What this costs in a real quarter

One last measurement, because it is the one that changes how you run a review. On this data a depth-5 tree is genuinely better than a depth-6 tree on new listings — by 0.013 AUC. Small, but real and repeatable.

Score both on a single random holdout, 200 different times. The worse model looks better in 23% of them.

So when the scoreboard moves from 0.70 to 0.72 after a sprint of work, roughly one time in four that is not the work. It is the split. The team will nevertheless write a changelog entry explaining why their change caused it, and the next sprint will build on a conclusion that was never there.

What to ask in the next standup

  1. Was this number computed on data the model was trained on?
    The only question on this list that is not optional. If the answer is yes, or is unclear, the number is not evidence of anything yet.
  2. What does it score on training data, and what does it score held out?
    Ask for both, always, as a pair. The gap between them is the overfitting, and it belongs on the slide beside the headline.
  3. Is that a single split or cross-validated — and what is the spread across folds?
    A mean without a spread cannot settle an argument. Two models closer together than the fold-to-fold variation are tied, whatever the decimal says.
  4. How many examples are in the test set?
    150 held-out rows gave an estimate that wobbled by 0.036; 9,000 gave 0.004. On a small evaluation set, most of what you are reading is sampling noise. This applies just as hard to a 40-prompt LLM eval — Lesson 21.
  5. How many times have we looked at this test set?
    Tune against a holdout for three months and you have fitted to it by hand. That is why teams keep a final set that is opened once, before launch, and not again.

15-minute hands-on

Paste into a free Google Colab notebook. It rebuilds Lesson 03’s resale catalogue, then does the thing you can never do at work: generates 200,000 listings from the same underlying truth, so you can see the answer the holdout is trying to guess. Every number in this lesson is printed by this script.

import numpy as np, pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split, cross_val_score, KFold
from sklearn.metrics import roc_auc_score

F = ["brand_tier", "discount", "condition", "photo_count"]

def catalogue(n, seed):
    """Lesson 03's resale catalogue. We control the truth here, so we can look at
       the future. In real life you cannot — and that is the entire problem."""
    rng = np.random.default_rng(seed)
    brand_tier  = rng.uniform(0, 1, n)
    discount    = rng.uniform(0, 60, n)
    condition   = rng.integers(1, 6, n)
    photo_count = rng.integers(1, 10, n)
    hs  = 0.055 * discount - 1.4
    lux = 1.5 - 0.0062 * (discount - 18.0) ** 2
    logit = (-0.55 + 0.42*(condition-3) + 0.11*(photo_count-5) - 0.9*brand_tier
             + (1-brand_tier)*hs + brand_tier*lux + rng.normal(0, 0.35, n))
    sold = (rng.uniform(size=n) < 1/(1+np.exp(-logit))).astype(int)
    return pd.DataFrame(dict(brand_tier=brand_tier, discount=discount,
                             condition=condition, photo_count=photo_count, sold=sold))

df     = catalogue(30_000, 11)      # the catalogue you have today
future = catalogue(200_000, 2027)   # next quarter's listings — the oracle
tr, te = train_test_split(df, test_size=0.3, random_state=11, stratify=df.sold)

print(f"{'depth':>6} {'leaves':>7} {'train':>7} {'holdout':>8} {'10-fold':>8} {'NEW':>7}")
for d in [1, 2, 3, 4, 5, 6, 8, 10, 14, 20, None]:
    m = DecisionTreeClassifier(max_depth=d, random_state=11).fit(tr[F], tr.sold)
    cv = cross_val_score(DecisionTreeClassifier(max_depth=d, random_state=11), tr[F],
                         tr.sold, cv=KFold(10, shuffle=True, random_state=7),
                         scoring="roc_auc").mean()
    print(f"{str(d):>6} {m.get_n_leaves():>7} "
          f"{roc_auc_score(tr.sold, m.predict_proba(tr[F])[:,1]):>7.3f} "
          f"{roc_auc_score(te.sold, m.predict_proba(te[F])[:,1]):>8.3f} {cv:>8.3f} "
          f"{roc_auc_score(future.sold, m.predict_proba(future[F])[:,1]):>7.3f}")

# How much does ONE holdout wobble? 2,000 listings, same model, 20 different splits.
small = catalogue(2_000, 11)
hold, cvs = [], []
for s in range(20):
    a, b = train_test_split(small, test_size=0.3, random_state=s, stratify=small.sold)
    m = DecisionTreeClassifier(max_depth=5, random_state=11).fit(a[F], a.sold)
    hold.append(roc_auc_score(b.sold, m.predict_proba(b[F])[:,1]))
    cvs.append(cross_val_score(DecisionTreeClassifier(max_depth=5, random_state=11),
               small[F], small.sold, cv=KFold(10, shuffle=True, random_state=s),
               scoring="roc_auc").mean())
hold, cvs = np.array(hold), np.array(cvs)
truth = roc_auc_score(future.sold, DecisionTreeClassifier(max_depth=5, random_state=11)
                      .fit(small[F], small.sold).predict_proba(future[F])[:,1])
print(f"\n2,000 listings, identical model, 20 different random splits:")
print(f"  one 30% holdout : {hold.min():.3f} – {hold.max():.3f}   (spread {np.ptp(hold):.3f})")
print(f"  10-fold CV      : {cvs.min():.3f} – {cvs.max():.3f}   (spread {np.ptp(cvs):.3f})")
print(f"  truth           : {truth:.3f}")

# The ten fold scores behind ONE of those cross-validated numbers
folds = cross_val_score(DecisionTreeClassifier(max_depth=5, random_state=11), small[F],
                        small.sold, cv=KFold(10, shuffle=True, random_state=0),
                        scoring="roc_auc")
print("  ten folds:", " ".join(f"{f:.3f}" for f in folds), f"-> mean {folds.mean():.3f}")

# Depth 5 really is better than depth 6. How often does one holdout disagree?
wrong = sum(
    roc_auc_score(b.sold, DecisionTreeClassifier(max_depth=6, random_state=11).fit(a[F],a.sold).predict_proba(b[F])[:,1])
  > roc_auc_score(b.sold, DecisionTreeClassifier(max_depth=5, random_state=11).fit(a[F],a.sold).predict_proba(b[F])[:,1])
    for a, b in (train_test_split(small, test_size=0.3, random_state=s, stratify=small.sold)
                 for s in range(200)))
print(f"\ndepth 5 beats depth 6 on new listings by 0.013 AUC — yet on a single random")
print(f"holdout, depth 6 looks better {wrong}/200 = {wrong/2:.0f}% of the time.")

Two experiments worth running after it. Change catalogue(2_000, 11) to catalogue(500, 11) and watch the holdout spread roughly double — small evaluation sets are where teams fool themselves hardest. Then change the rng.normal(0, 0.35, n) noise term to 0.9 and re-run the depth sweep: the noisier the world, the shallower the tree you can afford. How much complexity you can support is a property of the data, not of the algorithm.

Any model can be made to score better on data it has already seen — that direction has no floor. The only number that means anything is the one computed on examples the model has never been shown, and the only way to stop that number lying to you is to compute it many times and look at the spread.

Vocabulary from this lesson

overfittingunderfittinggeneralisationmodel optimismdouble-dippingtraining setholdout methodtest setvalidation setcross-validationk-foldfoldleave-one-outbias–varianceregularisationhyperparametergrid searchcapacity

Lesson 04 of the “Roadmap to be an AI engineer” series · builds on §05 Evaluation of the fundamentals guide · read Real-World ML ch. 4, §4.1Next: The model that saw tomorrow

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Back to top