Lesson 05 · Machine learning foundationsRoadmap to be an AI engineer

The model that
saw tomorrow

Lesson 04 was about a score that was too optimistic. This one is about a score that was never real — and the three ordinary pipeline habits that manufacture it.

Thursday · 09:40 · model demo

“Sell faster v2.” The model ranks every incoming listing by its chance of selling within 30 days, so ops can push re-shoots and price advice at the bottom of the list first.

The team has done everything you asked for after last month. Held-out test set. Cross-validated. Spread across folds reported.

v1 (four features)  0.74  →  v2 (+ seller history, + listing lifecycle)  0.95

Nobody in the room is lying and no rule has been broken. You ship it.

Five weeks later the ops lead mentions, carefully, that the priority queue does not seem to beat sorting by newest. You pull the production numbers. The model is scoring 0.56 — a coin flip with a good haircut.

If the number surprises you, it is a leak

Before the mechanism, the reflex. Real-World Machine Learning (ch. 6) opens its NYC taxi example with a rule of thumb worth memorising: when cross-validated accuracy is better than you expected, assume the model is cheating somewhere, because real data is endlessly inventive at fooling you.

Their tip-prediction model was brilliant because one feature dominated it: payment_type. Cash passengers tipped as often as anyone — but when a passenger paid cash, the driver never recorded the tip. So in the data, cash meant no tip, every time, with no exceptions. The model had not learned about tipping. It had learned how the data was written down.

That is data leakage: information reaches the model during training that will not exist, or will not mean the same thing, at the moment of the real prediction. The model does not get the answer wrong. It gets a different, easier question right.

Leakage is not one bug. It is a family, and the three members below cover nearly everything you will meet. All of them were in “Sell faster v2.”

Leak one: a column that exists because the sale already happened

The new feature was days_live — how long the listing had been on site. Plainly useful: stale listings sell less.

Except of a listing that sold, days_live is the number of days it took to sell. The listing closes on the day it sells. An item that did not sell keeps ageing: 40 days, 90, 150.

So the column does not describe the listing. It describes the outcome, written in a different unit. The model discovered in about four splits that small days_live means sold, and reported 0.95.

And at the instant the prediction is actually needed — a seller has just hit publish — every listing in the world has been live for zero days. The feature the model leaned its whole weight on is a constant.

What the slide said, and what production got
25,315 resale listings over two years, random 25% holdout. Score is AUC: the chance the model ranks a listing that sold above one that didn’t. 0.50 is a coin flip. “Next quarter” is 12,000 genuinely later listings, scored the way production scores them.
the four honest features0.7390.630+ days_liveleak 1 · target leakage0.9530.555−0.398 the day it shipped+ seller sell-rateleak 2 · built over all rows0.7830.5610.500.600.700.800.901.00AUC  →  bettercoin flip
What the holdout saidWhat next quarter actually gave
Read the amber bars downward and the blue bars downward and notice they run in opposite directions. Each leak buys a better number on the slide and pays for it in production. The honest model is the worst on the left and the best on the right — and its blue bar, 0.630, is the only number in this chart that was ever worth planning around.

The smell test that takes four minutes

You do not need to audit a pipeline to catch this one. Score every feature on its own, as if it were the entire model, and look at the ranking.

Each feature scored alone against “did it sell?”
Single-column AUC on the same holdout. A legitimate feature is a weak signal; that is why you need several of them.
days_live0.945discount0.670condition0.570brand_tier0.548photo_count0.5480.500.600.700.800.901.00AUC of that one column, used alone
The whole five-feature model scored 0.953. One column, by itself, scores 0.945. The other four together are worth 0.008. When a single feature nearly equals the model, it is not a strong feature — it is the answer in disguise. Run this sweep on any model before you believe its headline, and run it again every time someone adds a column.

Leak two: a lookup table that read the test set

Lesson 02 ended on a deliberate trap, and here it is. The team added seller_sell_rate: for each seller, the share of their listings that sold. One groupby, one map, four lines.

It was computed over the whole table — training rows and test rows together. So for a seller with six listings, that seller’s test listing contributed roughly a sixth of its own feature value. The answer was folded into the question, thinly, 25,000 times.

To show how little real signal is needed for this to work, the seller IDs in the simulation below are pure noise: they have no effect whatsoever on whether anything sells. Here is what that column of nothing did.

How the seller sell-rate was computedHoldout saysNext quarter gives
Feature not used at all0.7390.630
From the training rows only — correct0.7070.557
From the whole table — one word different0.7830.561
Computed correctly, the noise column is an honest nuisance: it costs 0.032 on the holdout, because the model wastes capacity on randomness. Computed over the whole table, the same noise column becomes the best-scoring model on the slide — better than the honest baseline it is actually worse than. The leak did not amplify a weak signal. It manufactured signal out of a column that contained none. That is why leakage is more dangerous than overfitting: overfitting flatters a real model, leakage invents one.

The rule this gives you is mechanical and worth writing into your team’s definition of done: anything fitted to data — an average, an encoding, a scaler, a vocabulary, a vector index — is fitted on training rows only, then applied to the others. In scikit-learn that means every such step lives inside a Pipeline so that cross-validation refits it per fold. In a warehouse it means the feature job knows a cutoff date. Machine Learning Engineering in Action (ch. 5) lists both “culling features not available at prediction time” and “correlation analysis to target (leakage prevention)” as standing items in the analysis phase, before any modelling starts.

Leak three: the same jacket on both sides of the split

This is the one specific to a resale marketplace, and the one most teams never look for.

Items get relisted. A seller drops the price and runs it again; ops re-shoots it and it goes back up. In the simulation below, 60% of items appear more than once — which is roughly what a real second-hand catalogue looks like.

Now split the rows at random. The second listing of a jacket lands in training; the third lands in test. Same brand, same condition, same photos, same hidden desirability that no column records. The model does not generalise to it. It recognises it.

73% of the test rows were a relist of an item the model had already trained on. The holdout was not measuring prediction. It was measuring recall.

Four items, twelve listings, three ways to split them
Listings in time order, left to right. Amber = held out and scored. Grey = trained on.
item
A
B
A
C
B
D
A
C
D
B
D
D
leaked?
rows at random
A
B
A
C
B
D
A
C
D
B
D
D
3 of 3
by time
A
B
A
C
B
D
A
C
D
B
D
D
3 of 3
by item (and time)
A
B
A
C
B
D
A
C
D
B
D
D
0 of 5
The right-hand column counts the held-out listings whose item also appears in training. A random split leaks every one of them. A time-ordered split leaks every one of them too — relists come later, so a jacket first listed before the cut and relisted after it straddles the boundary. Only grouping by item closes it. The two devices do different jobs and you usually need both: time keeps the future out of the past; grouping keeps the same thing off both sides.
How the 25,315 listings were splitScore it reportsTest rows whose item was trained on
Random rows — the default everywhere0.73973%
By time — last six months held out0.68930%
By item — no item on both sides0.6530%
By time and by item0.6450%
Same data, same features, same algorithm, same code. Only the definition of “held out” changes, and the answer moves by 0.094 AUC — four times the entire improvement Lesson 04’s team celebrated over four sprints. The bottom row is the honest estimate of how this model treats a jacket it has never seen, from a seller it has never seen, next quarter. Every row above it is a different way of grading the model on its own homework.

Putting the clock back into the split

Everything above has one cause: the training data knew something about the test data. Time is the cleanest fence you can build against that, because the real prediction always happens at a moment, with only the past behind it.

Real-World Machine Learning states the rule plainly in ch. 4, in its cautions for cross-validation on real data: when features are tied to time, nothing from the future may be used to predict the past, so arrange the holdout set or the k-folds so that every training row comes from before every test row.

In practice that turns Lesson 04’s shuffled k-fold into a staircase. Train on months 1–12, validate on 13–15. Train on 1–15, validate on 16–18. Train on 1–18, validate on 19–21. Each fold is a rehearsal of the thing you will actually do: stand at a date, use only what existed then, and predict forward. The industry calls it walk-forward or out-of-time validation; TimeSeriesSplit in scikit-learn does it in one line.

This also settles the question Lesson 04 left open about three-way splits. Order them in time, not at random: train on the oldest stretch, validate on the middle one while you tune, and keep a test set of the most recent weeks sealed until you are done choosing. A test set from the newest data answers the only question a launch cares about — how will this behave on what comes next — and it answers it once.

Two costs are worth saying out loud so nobody is surprised. An honest split gives a lower number, which is politically awkward and arithmetically correct. And a time-ordered holdout uses less data for training, so part of its drop is scarcity rather than truth. The fix is not to go back to shuffling. It is to retrain on everything up to today once the split has told you what to ship.

The demo on Thursday
0.95
three leaks, stacked
Production, five weeks later
0.56
a coin flip
An honest split, measured first
0.645
available on day one, free

The honest number was always computable. Nobody had to wait five weeks and an ops complaint to learn it — the only thing standing between the team and 0.645 was four hours of asking where each column came from.

What to ask in the next standup

  1. For each feature, what value would this have had at the exact moment we need the prediction?
    The one question that catches leak one, and it is answered feature by feature, out loud, with a real timestamp in mind — not “yes, we checked.” RWML’s rule for admitting anything as a feature at all: “The value of the feature must be known at the time predictions are needed.”
  2. Which single feature, scored on its own, comes closest to the full model?
    Four minutes of work. If one column is within a few points of the whole model, stop and find out why before anyone writes a slide. Nothing legitimate is nearly as good alone as everything together.
  3. Was every average, encoding, scaler and index fitted on training rows only?
    Leak two hides here, and it hides from code review because the leaky version is shorter. The answer should be a Pipeline or a dated feature job, not a recollection.
  4. Can the same item, seller or customer appear on both sides of the split?
    Relists, repeat sellers, returned-and-relisted stock, duplicate photos of one garment. Ask what the unit of generalisation is — a listing, or an item? — and group the split by that.
  5. Does our test set come after our training set in time?
    If the prediction is about the future, the evaluation has to be too. And once it is live, put the production score on the same chart as the offline score. A gap that opens between them is the leak you have not found yet — Lesson 29 is about what to do when that gap appears later on its own.

20-minute hands-on

Paste into a free Google Colab notebook. It builds a resale catalogue where items carry a hidden desirability the model cannot see and 60% of them get relisted, then reproduces all three leaks. It runs in about fifteen seconds. Every number in this lesson is printed by it.

import numpy as np, pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, GroupShuffleSplit
from sklearn.metrics import roc_auc_score

F = ["brand_tier", "discount", "condition", "photo_count"]

def catalogue(n_items, seed, day0, today):
    """Resale listings. Items have a hidden desirability the model cannot see,
       sellers are pure noise, and unsold items get relisted weeks later."""
    rng   = np.random.default_rng(seed)
    qual  = rng.normal(0, 1.8, n_items)          # the thing no column records
    bt    = rng.uniform(0, 1, n_items)
    disc  = rng.uniform(0, 60, n_items)
    cond  = rng.integers(1, 6, n_items)
    phot  = rng.integers(1, 10, n_items)
    sell  = rng.integers(0, n_items // 3, n_items)
    start = rng.integers(day0, day0 + 500, n_items)
    rows = []
    for i in range(n_items):
        d = start[i]
        for _ in range(1 + rng.choice(4, p=[.40, .25, .20, .15])):   # relists
            logit = (-1.6 + 0.30*(cond[i]-3) + 0.09*(phot[i]-5) - 0.9*bt[i]
                     + 0.055*disc[i] + qual[i] + rng.normal(0, .35))
            sold = int(rng.uniform() < 1/(1+np.exp(-logit)))
            # a SOLD listing closes the day it sells; an unsold one keeps ageing
            live = (int(round(rng.gamma(2., 9.))) + 1 if sold
                    else max(min(today - d, int(rng.integers(14, 150))), 1))
            rows.append((i, int(sell[i]), d, bt[i], disc[i], cond[i], phot[i], live, sold))
            d += int(rng.integers(25, 70))
    return pd.DataFrame(rows, columns=["item_id","seller_id","day"] + F + ["days_live","sold"])

HIST = catalogue(12_000, 11,   0, today=730)     # two years of history
NEW  = catalogue( 6_000, 2027, 760, today=1_000) # next quarter — you never have this

fit = lambda a, c: RandomForestClassifier(n_estimators=120, random_state=11,
                                          n_jobs=-1).fit(a[c], a.sold)
auc = lambda m, b, c: roc_auc_score(b.sold, m.predict_proba(b[c])[:, 1])

tr, te = train_test_split(HIST, test_size=.25, random_state=11, stratify=HIST.sold)
honest = fit(tr, F)
print(f"honest baseline   holdout {auc(honest,te,F):.3f}   next quarter {auc(honest,NEW,F):.3f}")

# ---- LEAK 1: days_live exists only because the sale already happened ----------
L, live = F + ["days_live"], NEW.copy()
live["days_live"] = 0          # a brand-new listing has been live zero days
leak1 = fit(tr, L)
print(f"+ days_live       holdout {auc(leak1,te,L):.3f}   next quarter {auc(leak1,live,L):.3f}")
for f in L:                     # the four-minute smell test
    s = roc_auc_score(te.sold, te[f]); print(f"     {f:<12} alone {max(s,1-s):.3f}")

# ---- LEAK 2: a lookup built over rows the model is about to be tested on ------
for lab, base in (("from TRAINING rows", tr), ("from the WHOLE table", HIST)):
    a, b, nw = tr.copy(), te.copy(), NEW.copy()
    lk, g = base.groupby("seller_id").sold.mean(), base.sold.mean()
    for part in (a, b, nw):
        part["seller_rate"] = part.seller_id.map(lk).fillna(g)
    C = F + ["seller_rate"]; m = fit(a, C)
    print(f"{lab:<22} holdout {auc(m,b,C):.3f}   next quarter {auc(m,nw,C):.3f}")

# ---- LEAK 3: the same item on both sides of the split -------------------------
ov = set(tr.item_id) & set(te.item_id)
print(f"random rows   {auc(honest,te,F):.3f}  "
      f"({te.item_id.isin(ov).mean()*100:.0f}% of test rows are a relist of a trained item)")
gi, gj = next(GroupShuffleSplit(1, test_size=.25, random_state=11)
              .split(HIST, groups=HIST.item_id))
print(f"by item       {auc(fit(HIST.iloc[gi],F), HIST.iloc[gj], F):.3f}")
cut   = HIST.day.quantile(.75)
a2,b2 = HIST[HIST.day < cut], HIST[HIST.day >= cut]
ov2   = set(a2.item_id) & set(b2.item_id)
print(f"by time       {auc(fit(a2,F), b2, F):.3f}  "
      f"({b2.item_id.isin(ov2).mean()*100:.0f}% overlap — relists straddle the cut)")
a3 = HIST[(HIST.day < cut) & ~HIST.item_id.isin(ov2)]
print(f"time AND item {auc(fit(a3,F), b2, F):.3f}  <- the number worth planning with")

Three things to try afterwards. Change the relist mix to p=[1,0,0,0] so nothing is ever relisted, and watch the random-split score fall into line with the grouped one — that difference is the leak, isolated. Raise qual’s spread from 1.8 to 3.0 to make the hidden desirability stronger, and the group leak grows with it, because the more your outcome depends on something you never measured, the more a repeated row is worth memorising. Then drop days_live but keep a price_drops column counting how many times the seller cut the price, and decide for yourself whether it is a feature or a leak. The answer depends entirely on when your pipeline computes it — and that is the whole lesson in one column.

A leak is not a flaw in the model. It is a promise your training data makes that your product cannot keep. Every feature faces one question, and it is about a moment rather than a method: at the instant we need this prediction, does this number exist yet — and does it mean the same thing it meant in the warehouse?

Vocabulary from this lesson

data leakagetarget leakagetrain–test contaminationpreprocessing leakagegroup leakagetemporal leakagepoint-in-time correctnessfeature availabilitycutoff datetime-ordered splitout-of-time validationwalk-forwardTimeSeriesSplitGroupKFoldPipelinetraining–serving skewbacktestoffline–online gap

Lesson 05 of the “Roadmap to be an AI engineer” series · builds on §03 Data & features of the fundamentals guide · read Real-World ML ch. 6 §6.1.2 and ch. 4 §4.1Next: A neuron is a weighted sum

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Back to top