Features are the product
The algorithm is a commodity you can swap in an afternoon. What you chose to measure is not — and it decides almost everything.
The team ships a price-suggestion model to help sellers list faster. It trains on exactly what the listings table holds: brand_id, category_id, size, colour, condition_grade.
Day one, a seller uploads a Hermès silk scarf in excellent condition. The model suggests €41. Another lists a Zara cotton tee and is told €28. Sellers stop trusting the suggestion within a week.
Nothing is broken. The model is working perfectly on the information it was given — and it was never given anything that says a Hermès scarf and a Zara tee occupy different universes.
The model only sees what you measured
To the model, brand_id = 847 is not “Hermès”. It is an arbitrary integer. Unless the training set contains enough Hermès items for the model to independently discover that this particular integer travels with high prices, that column carries almost no signal — and in a long-tail resale catalogue, most brands appear a handful of times.
A feature is not a column you happen to have. It is a measurement you decided to make. Real-World Machine Learning (ch. 5) defines feature engineering as applying transformations to raw input so the result relates more directly to the thing you are predicting. The canonical example in the book is a finance dataset holding balance and debt: the model struggles with both as raw columns, and succeeds the moment you hand it debt ÷ balance.
Your equivalent is not brand_id. It is the brand’s median sold price — one number, computed from your own sales history, that tells the model in a single value what tier of the market this item lives in.
Six values, five of them arbitrary identifiers. Nothing here states that this is an expensive brand, a winter garment, or a near-mint item.
Every value now points at the target. “€385 brand, outerwear, near-mint, listed in season” is a sentence a human pricer would recognise.
What that is worth, measured
Same training data, same algorithm, same held-out test set. The only thing that changes is which features the model is allowed to see. Error is mean absolute error in euros — on average, how far the suggested price lands from what the item actually sold for. Lower is better.
Five things features buy you
RWML lists why this work pays, and each one has a resale-shaped version:
- Relate the data to the target.
price ÷ brand_medianinstead of a brand ID. - Bring in external sources. Retail RRP, a brand tier list, seasonality — none of it lives in your listings table.
- Use unstructured data. The listing title, the description, the photos. Turn them into counts and vectors and the model can read them.
- Create interpretable features. “Days since the seller’s last sale” is something you can act on. Hidden layer 7, neuron 212 is not.
- Generate many, then select. Invent fifty candidates cheaply, then let feature selection keep the ten that earn their place.
The trap underneath
Features are computed twice: once over the training set before fitting, and once over live input before predicting. Those two computations must be identical. If brand_median_sold is calculated over all history during training but over the last 30 days in production, the model is being fed a different measurement than the one it learned from, and quality quietly degrades with nobody touching the model.
This is training–serving skew, and it is the most common way a model that tested well disappoints in production. It is also where your data engineering work becomes load-bearing: the same transformation, defined once, run in both places.
brand_median_sold computed over the full history, offline, in a batch job.What to ask in the next standup
What are the top five features by importance, and what does each one mean in plain English?
If the team cannot say what the strongest signal means, nobody can tell whether it is real or a leak.Who from merchandising or authentication helped choose these features?
Feature engineering is where domain expertise enters. If only engineers were in the room, the expertise is still outside the model.Is every feature computable at the moment of prediction, from data we have then?
Guards against both skew and leakage — the subject of Lesson 05.What does the baseline score, and how much did the features buy over it?
Carried over from Lesson 01. Features without a baseline are an unmeasured expense.
15-minute hands-on
Paste into a free Google Colab notebook. It builds a synthetic resale catalogue where price is genuinely driven by brand tier, condition and season — then trains the same model twice, once on raw IDs and once with three engineered features.
import numpy as np, pandas as pd
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error
rng = np.random.default_rng(7)
n, N_BRANDS, N_CATS = 40_000, 900, 40
# A long-tailed catalogue: a few huge brands, hundreds appearing a handful of times.
w = 1.0 / np.arange(1, N_BRANDS + 1) ** 1.1
w /= w.sum()
brand_id = rng.choice(N_BRANDS, size=n, p=w)
category_id = rng.integers(0, N_CATS, n)
condition = rng.integers(1, 6, n) # 1 poor .. 5 mint
is_winter = rng.integers(0, 2, n)
photo_count = rng.integers(1, 10, n)
# The hidden truth the model must discover. Note the IDs themselves carry no
# order: brand 12 is not cheaper than brand 800.
brand_tier = np.exp(rng.normal(3.2, 1.0, N_BRANDS))
cat_mult = np.exp(rng.normal(0.0, 0.6, N_CATS))
price = (brand_tier[brand_id] * cat_mult[category_id]
* (0.55 + 0.11 * condition) * (1 + 0.18 * is_winter)
+ 1.5 * photo_count + rng.normal(0, 6, n)).clip(5)
df = pd.DataFrame(dict(brand_id=brand_id, category_id=category_id, condition=condition,
is_winter=is_winter, photo_count=photo_count, price=price))
tr, te = train_test_split(df, test_size=0.25, random_state=7)
tr, te = tr.copy(), te.copy()
# THE FEATURES: median sold price per brand and per category, learned from the
# TRAINING split only, then mapped onto both. Unseen brands fall back to global.
global_median = tr["price"].median()
for src, name in (("brand_id", "brand_median_sold"),
("category_id", "category_median_sold")):
lookup = tr.groupby(src)["price"].median()
for part in (tr, te):
part[name] = part[src].map(lookup).fillna(global_median)
def run(cols, label):
m = HistGradientBoostingRegressor(random_state=7).fit(tr[cols], tr["price"])
mae = mean_absolute_error(te["price"], m.predict(te[cols]))
print(f"{label:<34} MAE EUR {mae:6.2f}")
print(f"{'baseline: always the median':<34} MAE EUR "
f"{mean_absolute_error(te['price'], np.full(len(te), global_median)):6.2f}")
run(["brand_id", "category_id", "condition"], "raw IDs only")
run(["brand_id", "category_id", "condition",
"is_winter", "photo_count"], "+ every other raw column")
run(["brand_median_sold", "category_id", "condition",
"is_winter", "photo_count"], "+ brand median sold")
run(["brand_median_sold", "category_median_sold", "condition",
"is_winter", "photo_count"], "+ category median sold") You should see roughly 29.56 / 18.81 / 18.55 / 9.47 / 8.40 — the five bars above. Two things to take from it. Adding every remaining raw column moved the error by 1%, while one derived column cut it almost in half: the model was never short of data, it was short of a measurement that meant something.
And notice that lookup is built from tr alone. Compute those medians over the whole dataset instead — one word changed — and the test set has quietly helped build its own features. The score improves and the number becomes a lie. Try it, watch the MAE fall, and remember the feeling: that is data leakage, and it is Lesson 05.
A model can only reason about what you chose to measure. Swapping algorithms moves the result a few percent; adding the right feature moves it by half. The features are where your knowledge of second-hand fashion actually enters the machine.
Vocabulary from this lesson
These lessons were written with the help of AI (Claude) and reviewed and edited by me.