The model that
saw tomorrow
Lesson 04 was about a score that was too optimistic. This one is about a score that was never real — and the three ordinary pipeline habits that manufacture it.
“Sell faster v2.” The model ranks every incoming listing by its chance of selling within 30 days, so ops can push re-shoots and price advice at the bottom of the list first.
The team has done everything you asked for after last month. Held-out test set. Cross-validated. Spread across folds reported.
v1 (four features) 0.74 → v2 (+ seller history, + listing lifecycle) 0.95
Nobody in the room is lying and no rule has been broken. You ship it.
Five weeks later the ops lead mentions, carefully, that the priority queue does not seem to beat sorting by newest. You pull the production numbers. The model is scoring 0.56 — a coin flip with a good haircut.
If the number surprises you, it is a leak
Before the mechanism, the reflex. Real-World Machine Learning (ch. 6) opens its NYC taxi example with a rule of thumb worth memorising: when cross-validated accuracy is better than you expected, assume the model is cheating somewhere, because real data is endlessly inventive at fooling you.
Their tip-prediction model was brilliant because one feature dominated it: payment_type. Cash passengers tipped as often as anyone — but when a passenger paid cash, the driver never recorded the tip. So in the data, cash meant no tip, every time, with no exceptions. The model had not learned about tipping. It had learned how the data was written down.
That is data leakage: information reaches the model during training that will not exist, or will not mean the same thing, at the moment of the real prediction. The model does not get the answer wrong. It gets a different, easier question right.
Leakage is not one bug. It is a family, and the three members below cover nearly everything you will meet. All of them were in “Sell faster v2.”
Leak one: a column that exists because the sale already happened
The new feature was days_live — how long the listing had been on site. Plainly useful: stale listings sell less.
Except of a listing that sold, days_live is the number of days it took to sell. The listing closes on the day it sells. An item that did not sell keeps ageing: 40 days, 90, 150.
So the column does not describe the listing. It describes the outcome, written in a different unit. The model discovered in about four splits that small days_live means sold, and reported 0.95.
And at the instant the prediction is actually needed — a seller has just hit publish — every listing in the world has been live for zero days. The feature the model leaned its whole weight on is a constant.
The smell test that takes four minutes
You do not need to audit a pipeline to catch this one. Score every feature on its own, as if it were the entire model, and look at the ranking.
Leak two: a lookup table that read the test set
Lesson 02 ended on a deliberate trap, and here it is. The team added seller_sell_rate: for each seller, the share of their listings that sold. One groupby, one map, four lines.
It was computed over the whole table — training rows and test rows together. So for a seller with six listings, that seller’s test listing contributed roughly a sixth of its own feature value. The answer was folded into the question, thinly, 25,000 times.
To show how little real signal is needed for this to work, the seller IDs in the simulation below are pure noise: they have no effect whatsoever on whether anything sells. Here is what that column of nothing did.
| How the seller sell-rate was computed | Holdout says | Next quarter gives |
|---|---|---|
| Feature not used at all | 0.739 | 0.630 |
| From the training rows only — correct | 0.707 | 0.557 |
| From the whole table — one word different | 0.783 | 0.561 |
The rule this gives you is mechanical and worth writing into your team’s definition of done: anything fitted to data — an average, an encoding, a scaler, a vocabulary, a vector index — is fitted on training rows only, then applied to the others. In scikit-learn that means every such step lives inside a Pipeline so that cross-validation refits it per fold. In a warehouse it means the feature job knows a cutoff date. Machine Learning Engineering in Action (ch. 5) lists both “culling features not available at prediction time” and “correlation analysis to target (leakage prevention)” as standing items in the analysis phase, before any modelling starts.
Leak three: the same jacket on both sides of the split
This is the one specific to a resale marketplace, and the one most teams never look for.
Items get relisted. A seller drops the price and runs it again; ops re-shoots it and it goes back up. In the simulation below, 60% of items appear more than once — which is roughly what a real second-hand catalogue looks like.
Now split the rows at random. The second listing of a jacket lands in training; the third lands in test. Same brand, same condition, same photos, same hidden desirability that no column records. The model does not generalise to it. It recognises it.
73% of the test rows were a relist of an item the model had already trained on. The holdout was not measuring prediction. It was measuring recall.
| How the 25,315 listings were split | Score it reports | Test rows whose item was trained on |
|---|---|---|
| Random rows — the default everywhere | 0.739 | 73% |
| By time — last six months held out | 0.689 | 30% |
| By item — no item on both sides | 0.653 | 0% |
| By time and by item | 0.645 | 0% |
Putting the clock back into the split
Everything above has one cause: the training data knew something about the test data. Time is the cleanest fence you can build against that, because the real prediction always happens at a moment, with only the past behind it.
Real-World Machine Learning states the rule plainly in ch. 4, in its cautions for cross-validation on real data: when features are tied to time, nothing from the future may be used to predict the past, so arrange the holdout set or the k-folds so that every training row comes from before every test row.
In practice that turns Lesson 04’s shuffled k-fold into a staircase. Train on months 1–12, validate on 13–15. Train on 1–15, validate on 16–18. Train on 1–18, validate on 19–21. Each fold is a rehearsal of the thing you will actually do: stand at a date, use only what existed then, and predict forward. The industry calls it walk-forward or out-of-time validation; TimeSeriesSplit in scikit-learn does it in one line.
This also settles the question Lesson 04 left open about three-way splits. Order them in time, not at random: train on the oldest stretch, validate on the middle one while you tune, and keep a test set of the most recent weeks sealed until you are done choosing. A test set from the newest data answers the only question a launch cares about — how will this behave on what comes next — and it answers it once.
Two costs are worth saying out loud so nobody is surprised. An honest split gives a lower number, which is politically awkward and arithmetically correct. And a time-ordered holdout uses less data for training, so part of its drop is scarcity rather than truth. The fix is not to go back to shuffling. It is to retrain on everything up to today once the split has told you what to ship.
The honest number was always computable. Nobody had to wait five weeks and an ops complaint to learn it — the only thing standing between the team and 0.645 was four hours of asking where each column came from.
What to ask in the next standup
For each feature, what value would this have had at the exact moment we need the prediction?
The one question that catches leak one, and it is answered feature by feature, out loud, with a real timestamp in mind — not “yes, we checked.” RWML’s rule for admitting anything as a feature at all: “The value of the feature must be known at the time predictions are needed.”Which single feature, scored on its own, comes closest to the full model?
Four minutes of work. If one column is within a few points of the whole model, stop and find out why before anyone writes a slide. Nothing legitimate is nearly as good alone as everything together.Was every average, encoding, scaler and index fitted on training rows only?
Leak two hides here, and it hides from code review because the leaky version is shorter. The answer should be aPipelineor a dated feature job, not a recollection.Can the same item, seller or customer appear on both sides of the split?
Relists, repeat sellers, returned-and-relisted stock, duplicate photos of one garment. Ask what the unit of generalisation is — a listing, or an item? — and group the split by that.Does our test set come after our training set in time?
If the prediction is about the future, the evaluation has to be too. And once it is live, put the production score on the same chart as the offline score. A gap that opens between them is the leak you have not found yet — Lesson 29 is about what to do when that gap appears later on its own.
20-minute hands-on
Paste into a free Google Colab notebook. It builds a resale catalogue where items carry a hidden desirability the model cannot see and 60% of them get relisted, then reproduces all three leaks. It runs in about fifteen seconds. Every number in this lesson is printed by it.
import numpy as np, pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, GroupShuffleSplit
from sklearn.metrics import roc_auc_score
F = ["brand_tier", "discount", "condition", "photo_count"]
def catalogue(n_items, seed, day0, today):
"""Resale listings. Items have a hidden desirability the model cannot see,
sellers are pure noise, and unsold items get relisted weeks later."""
rng = np.random.default_rng(seed)
qual = rng.normal(0, 1.8, n_items) # the thing no column records
bt = rng.uniform(0, 1, n_items)
disc = rng.uniform(0, 60, n_items)
cond = rng.integers(1, 6, n_items)
phot = rng.integers(1, 10, n_items)
sell = rng.integers(0, n_items // 3, n_items)
start = rng.integers(day0, day0 + 500, n_items)
rows = []
for i in range(n_items):
d = start[i]
for _ in range(1 + rng.choice(4, p=[.40, .25, .20, .15])): # relists
logit = (-1.6 + 0.30*(cond[i]-3) + 0.09*(phot[i]-5) - 0.9*bt[i]
+ 0.055*disc[i] + qual[i] + rng.normal(0, .35))
sold = int(rng.uniform() < 1/(1+np.exp(-logit)))
# a SOLD listing closes the day it sells; an unsold one keeps ageing
live = (int(round(rng.gamma(2., 9.))) + 1 if sold
else max(min(today - d, int(rng.integers(14, 150))), 1))
rows.append((i, int(sell[i]), d, bt[i], disc[i], cond[i], phot[i], live, sold))
d += int(rng.integers(25, 70))
return pd.DataFrame(rows, columns=["item_id","seller_id","day"] + F + ["days_live","sold"])
HIST = catalogue(12_000, 11, 0, today=730) # two years of history
NEW = catalogue( 6_000, 2027, 760, today=1_000) # next quarter — you never have this
fit = lambda a, c: RandomForestClassifier(n_estimators=120, random_state=11,
n_jobs=-1).fit(a[c], a.sold)
auc = lambda m, b, c: roc_auc_score(b.sold, m.predict_proba(b[c])[:, 1])
tr, te = train_test_split(HIST, test_size=.25, random_state=11, stratify=HIST.sold)
honest = fit(tr, F)
print(f"honest baseline holdout {auc(honest,te,F):.3f} next quarter {auc(honest,NEW,F):.3f}")
# ---- LEAK 1: days_live exists only because the sale already happened ----------
L, live = F + ["days_live"], NEW.copy()
live["days_live"] = 0 # a brand-new listing has been live zero days
leak1 = fit(tr, L)
print(f"+ days_live holdout {auc(leak1,te,L):.3f} next quarter {auc(leak1,live,L):.3f}")
for f in L: # the four-minute smell test
s = roc_auc_score(te.sold, te[f]); print(f" {f:<12} alone {max(s,1-s):.3f}")
# ---- LEAK 2: a lookup built over rows the model is about to be tested on ------
for lab, base in (("from TRAINING rows", tr), ("from the WHOLE table", HIST)):
a, b, nw = tr.copy(), te.copy(), NEW.copy()
lk, g = base.groupby("seller_id").sold.mean(), base.sold.mean()
for part in (a, b, nw):
part["seller_rate"] = part.seller_id.map(lk).fillna(g)
C = F + ["seller_rate"]; m = fit(a, C)
print(f"{lab:<22} holdout {auc(m,b,C):.3f} next quarter {auc(m,nw,C):.3f}")
# ---- LEAK 3: the same item on both sides of the split -------------------------
ov = set(tr.item_id) & set(te.item_id)
print(f"random rows {auc(honest,te,F):.3f} "
f"({te.item_id.isin(ov).mean()*100:.0f}% of test rows are a relist of a trained item)")
gi, gj = next(GroupShuffleSplit(1, test_size=.25, random_state=11)
.split(HIST, groups=HIST.item_id))
print(f"by item {auc(fit(HIST.iloc[gi],F), HIST.iloc[gj], F):.3f}")
cut = HIST.day.quantile(.75)
a2,b2 = HIST[HIST.day < cut], HIST[HIST.day >= cut]
ov2 = set(a2.item_id) & set(b2.item_id)
print(f"by time {auc(fit(a2,F), b2, F):.3f} "
f"({b2.item_id.isin(ov2).mean()*100:.0f}% overlap — relists straddle the cut)")
a3 = HIST[(HIST.day < cut) & ~HIST.item_id.isin(ov2)]
print(f"time AND item {auc(fit(a3,F), b2, F):.3f} <- the number worth planning with") Three things to try afterwards. Change the relist mix to p=[1,0,0,0] so nothing is ever relisted, and watch the random-split score fall into line with the grouped one — that difference is the leak, isolated. Raise qual’s spread from 1.8 to 3.0 to make the hidden desirability stronger, and the group leak grows with it, because the more your outcome depends on something you never measured, the more a repeated row is worth memorising. Then drop days_live but keep a price_drops column counting how many times the seller cut the price, and decide for yourself whether it is a feature or a leak. The answer depends entirely on when your pipeline computes it — and that is the whole lesson in one column.
A leak is not a flaw in the model. It is a promise your training data makes that your product cannot keep. Every feature faces one question, and it is about a moment rather than a method: at the instant we need this prediction, does this number exist yet — and does it mean the same thing it meant in the warehouse?
Vocabulary from this lesson
These lessons were written with the help of AI (Claude) and reviewed and edited by me.