Lesson 02 · Classic MLRoadmap to be an AI engineer

Features are the product

The algorithm is a commodity you can swap in an afternoon. What you chose to measure is not — and it decides almost everything.

Tuesday · 11:20 · #pricing

The team ships a price-suggestion model to help sellers list faster. It trains on exactly what the listings table holds: brand_id, category_id, size, colour, condition_grade.

Day one, a seller uploads a Hermès silk scarf in excellent condition. The model suggests €41. Another lists a Zara cotton tee and is told €28. Sellers stop trusting the suggestion within a week.

Nothing is broken. The model is working perfectly on the information it was given — and it was never given anything that says a Hermès scarf and a Zara tee occupy different universes.

The model only sees what you measured

To the model, brand_id = 847 is not “Hermès”. It is an arbitrary integer. Unless the training set contains enough Hermès items for the model to independently discover that this particular integer travels with high prices, that column carries almost no signal — and in a long-tail resale catalogue, most brands appear a handful of times.

A feature is not a column you happen to have. It is a measurement you decided to make. Real-World Machine Learning (ch. 5) defines feature engineering as applying transformations to raw input so the result relates more directly to the thing you are predicting. The canonical example in the book is a finance dataset holding balance and debt: the model struggles with both as raw columns, and succeeds the moment you hand it debt ÷ balance.

Your equivalent is not brand_id. It is the brand’s median sold price — one number, computed from your own sales history, that tells the model in a single value what tier of the market this item lives in.

One Burberry trench coat, two ways of describing it
Left: the raw listing row. Right: the same item after feature engineering. Same coat, same database — a different amount of knowledge handed to the model.
Raw columnswhat the table holds
brand_id312
category_id88
size“M”
colour“beige”
condition_grade“B”
created_at2026-09-28

Six values, five of them arbitrary identifiers. Nothing here states that this is an expensive brand, a winter garment, or a near-mint item.

Engineered featureswhat you chose to measure
brand_median_sold€385
category_median_sold€120
condition_rank4 / 5
is_outerweartrue
listed_in_seasontrue
photo_count7

Every value now points at the target. “€385 brand, outerwear, near-mint, listed in season” is a sentence a human pricer would recognise.

This is where your domain knowledge enters the system. RWML §5.1.3 calls feature engineering “a mechanism that imbues domain expertise into a machine-learning model” — the only place in the pipeline where what your merchandisers know becomes something the model can use.

What that is worth, measured

Same training data, same algorithm, same held-out test set. The only thing that changes is which features the model is allowed to see. Error is mean absolute error in euros — on average, how far the suggested price lands from what the item actually sold for. Lower is better.

Five feature sets, one algorithm
Gradient-boosted trees, 40,000 sold listings, held-out test set. Every bar on one shared scale, €0 to €35. These are the exact numbers the exercise below prints.
Baselinealways the median price
€29.56
Raw IDs onlybrand_id, category_id, condition
€18.81
+ every other raw columnseason, photo count
€18.55
+ brand median soldone derived column
€9.47
+ category median soldtwo derived columns
€8.40
€0€35 — worse →
Baseline from Lesson 01Raw columnsEngineered features
Look at rows three and four. Handing the model every remaining raw column improved it by 1%. Handing it one column derived from your own sales history cut the error by 49%. No algorithm changed anywhere on this chart — only what the model was allowed to see.
Every extra raw column
−1%
more data, no knowledge
One derived column
−49%
brand median sold
Algorithm changes
0
same model throughout

Five things features buy you

RWML lists why this work pays, and each one has a resale-shaped version:

  1. Relate the data to the target.price ÷ brand_median instead of a brand ID.
  2. Bring in external sources. Retail RRP, a brand tier list, seasonality — none of it lives in your listings table.
  3. Use unstructured data. The listing title, the description, the photos. Turn them into counts and vectors and the model can read them.
  4. Create interpretable features. “Days since the seller’s last sale” is something you can act on. Hidden layer 7, neuron 212 is not.
  5. Generate many, then select. Invent fifty candidates cheaply, then let feature selection keep the ten that earn their place.

The trap underneath

Features are computed twice: once over the training set before fitting, and once over live input before predicting. Those two computations must be identical. If brand_median_sold is calculated over all history during training but over the last 30 days in production, the model is being fed a different measurement than the one it learned from, and quality quietly degrades with nobody touching the model.

This is training–serving skew, and it is the most common way a model that tested well disappoints in production. It is also where your data engineering work becomes load-bearing: the same transformation, defined once, run in both places.

One definition, two paths — they must agree
Training40,000 sold listings. brand_median_sold computed over the full history, offline, in a batch job.
≠?
The definitionWritten once, shared by both paths. A library function, a dbt model, a feature store — the mechanism matters less than there being exactly one.
≠?
ServingOne seller, one item, 200 ms budget. The same number must come out, or the model sees an input it never trained on.
When someone proposes a clever new feature, the second question after “does it help?” is “can we compute it at prediction time, from data we have at that moment?” If not, it is not a feature — it is a research finding.

What to ask in the next standup

  1. What are the top five features by importance, and what does each one mean in plain English?
    If the team cannot say what the strongest signal means, nobody can tell whether it is real or a leak.
  2. Who from merchandising or authentication helped choose these features?
    Feature engineering is where domain expertise enters. If only engineers were in the room, the expertise is still outside the model.
  3. Is every feature computable at the moment of prediction, from data we have then?
    Guards against both skew and leakage — the subject of Lesson 05.
  4. What does the baseline score, and how much did the features buy over it?
    Carried over from Lesson 01. Features without a baseline are an unmeasured expense.

15-minute hands-on

Paste into a free Google Colab notebook. It builds a synthetic resale catalogue where price is genuinely driven by brand tier, condition and season — then trains the same model twice, once on raw IDs and once with three engineered features.

import numpy as np, pandas as pd
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error

rng = np.random.default_rng(7)
n, N_BRANDS, N_CATS = 40_000, 900, 40

# A long-tailed catalogue: a few huge brands, hundreds appearing a handful of times.
w = 1.0 / np.arange(1, N_BRANDS + 1) ** 1.1
w /= w.sum()
brand_id    = rng.choice(N_BRANDS, size=n, p=w)
category_id = rng.integers(0, N_CATS, n)
condition   = rng.integers(1, 6, n)          # 1 poor .. 5 mint
is_winter   = rng.integers(0, 2, n)
photo_count = rng.integers(1, 10, n)

# The hidden truth the model must discover. Note the IDs themselves carry no
# order: brand 12 is not cheaper than brand 800.
brand_tier = np.exp(rng.normal(3.2, 1.0, N_BRANDS))
cat_mult   = np.exp(rng.normal(0.0, 0.6, N_CATS))
price = (brand_tier[brand_id] * cat_mult[category_id]
         * (0.55 + 0.11 * condition) * (1 + 0.18 * is_winter)
         + 1.5 * photo_count + rng.normal(0, 6, n)).clip(5)

df = pd.DataFrame(dict(brand_id=brand_id, category_id=category_id, condition=condition,
                       is_winter=is_winter, photo_count=photo_count, price=price))
tr, te = train_test_split(df, test_size=0.25, random_state=7)
tr, te = tr.copy(), te.copy()

# THE FEATURES: median sold price per brand and per category, learned from the
# TRAINING split only, then mapped onto both. Unseen brands fall back to global.
global_median = tr["price"].median()
for src, name in (("brand_id", "brand_median_sold"),
                  ("category_id", "category_median_sold")):
    lookup = tr.groupby(src)["price"].median()
    for part in (tr, te):
        part[name] = part[src].map(lookup).fillna(global_median)

def run(cols, label):
    m = HistGradientBoostingRegressor(random_state=7).fit(tr[cols], tr["price"])
    mae = mean_absolute_error(te["price"], m.predict(te[cols]))
    print(f"{label:<34} MAE EUR {mae:6.2f}")

print(f"{'baseline: always the median':<34} MAE EUR "
      f"{mean_absolute_error(te['price'], np.full(len(te), global_median)):6.2f}")
run(["brand_id", "category_id", "condition"],                      "raw IDs only")
run(["brand_id", "category_id", "condition",
     "is_winter", "photo_count"],                                  "+ every other raw column")
run(["brand_median_sold", "category_id", "condition",
     "is_winter", "photo_count"],                                  "+ brand median sold")
run(["brand_median_sold", "category_median_sold", "condition",
     "is_winter", "photo_count"],                                  "+ category median sold")

You should see roughly 29.56 / 18.81 / 18.55 / 9.47 / 8.40 — the five bars above. Two things to take from it. Adding every remaining raw column moved the error by 1%, while one derived column cut it almost in half: the model was never short of data, it was short of a measurement that meant something.

And notice that lookup is built from tr alone. Compute those medians over the whole dataset instead — one word changed — and the test set has quietly helped build its own features. The score improves and the number becomes a lie. Try it, watch the MAE fall, and remember the feeling: that is data leakage, and it is Lesson 05.

A model can only reason about what you chose to measure. Swapping algorithms moves the result a few percent; adding the right feature moves it by half. The features are where your knowledge of second-hand fashion actually enters the machine.

Vocabulary from this lesson

featurefeature engineeringderived featurecategorical vs numericone-hot encodingtarget encodingfeature importancefeature selectiondomain expertisemean absolute errortraining–serving skewfeature store

Lesson 02 of the “Roadmap to be an AI engineer” series · builds on §03 Data & features of the fundamentals guide · read Real-World ML ch. 2 and 5Next: Lines and trees

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Back to top