Lesson 03 · Machine learning foundationsRoadmap to be an AI engineer

Lines and trees

A linear model gives every feature one number that applies to everybody. Sometimes that is exactly right. Sometimes it is the reason your pricing advice is backwards for half the catalogue.

Thursday · 09:15 · #seller-tools

The team ships “Sell faster”: a model that estimates the chance a listing sells within 30 days, and nudges the seller to cut the price when the odds look poor. It is a logistic regression — fast, explainable, the sensible first choice.

Six weeks later, merchandising brings a complaint. Sellers of Chanel, Gucci and Prada were told to drop their prices. They did. Their items stopped selling altogether. Meanwhile the model was telling Zara sellers that discounting would not help them much — and those are the listings where it works best.

Nobody shipped a bug. The model was asked a question it structurally could not answer.

One coefficient, two opposite populations

A linear model learns a fixed equation with one weight per feature. Here is the whole model, honestly:

log-odds(sells in 30 days) =
     0.50 × condition
   + 0.27 × photo_count
   − 0.29 × brand_tier
   − 0.26 × discount     ← one number, applied to every listing

That last coefficient has to be a single value. But the data contains two populations that disagree about it:

High-street items — a Zara coat, an H&M dress. Cut the price, the odds go up. Straightforwardly.

Luxury items — a Chanel flap bag. A modest cut helps. A 50% cut does not read as “bargain”, it reads as “probably fake, or probably damaged”. Buyers scroll past. The odds collapse.

The relationship is not monotonic, and it changes direction depending on another feature. Fitting one straight line through both populations gives you the average of two opposite truths: −0.26, mildly negative, wrong for everyone.

This is the distinction Real-World Machine Learning (ch. 3) draws between parametric and nonparametric models. A parametric model assumes the shape of the relationship in advance; the data only fills in the numbers. A nonparametric model lets the shape itself be learned — “the form and complexity of f adapts to the complexity of the data.”

The same question, three answers
Chance a listing sells within 30 days (%), by brand tier (rows) and discount (columns). Condition and photo count held fixed. Darker = more likely to sell.
What actually happensthe sell rate in the data
5%
15%
25%
35%
45%
55%
0.9luxury
26
47
43
17
2
0
0.7
23
42
41
23
5
0
0.5mid
21
36
40
29
12
2
0.3
19
31
38
37
27
14
0.1high street
17
27
37
45
51
53
What the linear model predictsone coefficient per feature
5%
15%
25%
35%
45%
55%
0.9luxury
25
23
20
18
16
14
0.7
29
26
24
21
19
16
0.5mid
34
30
27
24
22
19
0.3
38
35
31
28
25
23
0.1high street
43
39
36
32
29
26
What boosted trees predictsplits learned from the data
5%
15%
25%
35%
45%
55%
0.9luxury
26
42
41
19
2
0
0.7
26
39
39
28
6
1
0.5mid
26
34
37
29
9
3
0.3
25
30
36
27
26
22
0.1high street
24
28
36
36
56
65
0%
65% — sellsevery cell is labelled, so the grids read without colour
Read the top-left panel diagonally. The best discount depends on the brand tier: deep cuts work at the bottom of the market and destroy the listing at the top. That is a feature interaction — and the middle panel shows what a linear model can do about it, which is nothing. Its surface is a smooth tilted plane, because a plane is the only shape it has. The right-hand panel is gradient-boosted trees, given exactly the same four columns, recovering the ridge on their own.

A tree asks questions instead of weighing features

A decision tree does not fit an equation. It repeatedly splits the data: find the single question that best separates sold from unsold, ask it, then repeat inside each half. Nothing about the shape is decided in advance — depth and structure are learned.

Here is the real tree our code below produces, three questions deep. Follow the right-hand branch:

Three questions, and the interaction falls out
Actual splits chosen by a depth-3 decision tree on 21,000 listings. Nobody told it to look at brand tier and discount together.
condition ≤ 3.5 ?the strongest single split
no — near-mint items →
discount > 37.9% ?a deep cut
yes — and now, crucially …
brand_tier ≤ 0.27 ?high street, or premium?
yes — high streetno — premium / luxury
deep discount, high street59%sell within 30 days · n = 854
deep discount, luxury12%sell within 30 days · n = 2,260
Identical condition, identical discount. The only difference is the brand tier, and the sell rate moves from 59% to 12%. A linear model cannot express that split without someone first guessing the interaction and hand-building it as a feature. The tree found it in three questions, from the raw columns.

One tree is not enough

A single tree that is allowed to keep splitting will carve the training data into tiny pockets and memorise it. In the scoreboard below, the unlimited-depth tree scores worse than the linear model — it learned the noise. (That is overfitting, and it is Lesson 04.)

The fix is not a better tree, it is many of them:

Random forest — grow hundreds of trees, each on a random subsample of rows and features, then average their predictions (or let them vote). Each tree overfits differently, and the errors largely cancel. RWML calls this the ensemble effect.

Boosting — grow trees in sequence, each one trained to fix what the previous ones got wrong. RWML’s appendix describes the mechanism: each new tree gives more weight to the examples the model still gets wrong (for regression, the ones with the largest residuals), which lowers the error on new data. Gradient boosting (XGBoost, LightGBM, scikit-learn’s HistGradientBoosting) is the default winner on tabular business data, and is almost certainly what your team should try first on a table.

Seven models, one dataset, four raw columns
30,000 listings, 30% held out. Score is AUC — the chance the model ranks a listing that sold above one that didn’t. 0.50 is a coin flip, so the axis starts there; 1.00 would be perfect. These are the exact numbers the exercise below prints.
Baselineeveryone gets the same score
0.500
Logistic regressionthe four raw columns
0.683
One decision treeunlimited depth
0.582
One decision treedepth capped at 4
0.737
Random forest300 trees, averaged
0.764
Gradient boostingtrees that fix each other
0.766
Logistic regression+ the interaction, hand-built
0.773
0.50 — coin flip0.80
BaselineLinear modelsTree models
Note the last bar. The linear model with the interaction written in by hand beats every tree on this data. That is the honest version of the lesson: trees are not smarter, they are less dependent on you already knowing the answer. If you know the shape, encode it and keep the simple model. In the real world you usually don’t.

What the mistake was worth

On the segment that caused the complaint — luxury items already discounted more than 40% — the two models are not slightly different. They are in different businesses.

What actually sold
1.4%
772 listings, held-out set
Logistic regression said
16.9%
12× too optimistic
Gradient boosting said
2.5%
close enough to act on

And in the other direction: on luxury items at a modest 10–28% discount, 46% actually sold. The linear model predicted 23% and told those sellers to cut further; boosting predicted 42%. Both of the model’s errors pushed sellers the same wrong way.

Which one to reach for

Linear modelsTrees & ensembles
Shape of the relationshipAssumed in advance — a straight line, a planeLearned from the data; adapts to its complexity
InteractionsOnly if you build them as features by handFound automatically, that’s what a split is
InterpretabilityOne coefficient per feature — a sentenceOne tree is readable; 300 are not. Use feature importance and SHAP
Feature scalingRequired — normalise firstNot required; splits don’t care about units
Beyond the training rangeExtrapolates — keeps going up the lineCannot extrapolate; flat past the edge of what it saw
Data neededWorks on small dataWants thousands of rows to split safely
Speed & costMicroseconds; trivial to serveMilliseconds; still cheap, bigger artefact
Reach for it whenYou understand the mechanism, or you must defend every prediction to a regulatorTabular business data and you want the best number — your default on a table
Neither column is “the advanced one”. RWML’s framing: nonlinear models can buy accuracy on complicated real-world data, but nothing guarantees they beat a linear model, because they overfit more easily; linear models, in turn, are easier to explain and quicker to compute and scale.

What to ask in the next standup

  1. Is the relationship we’re modelling monotonic — does more of this always mean more of that?
    Ask merchandising, not the data team. If the answer is “it depends on the brand”, a plain linear model is already wrong.
  2. What does a linear model score, and how much do the trees buy over it?
    Lesson 01’s baseline, one level up. A boosted model that beats logistic regression by 0.004 AUC is not worth the operational weight.
  3. If the trees won, which interaction did they find — and should it become a feature?
    The tree is telling you something true about the business. Write it down; it may be worth more than the model.
  4. Will this model ever be asked about an item outside the range it trained on?
    Trees go flat past the edge of their data. A €9,000 Birkin scored by a model trained up to €2,000 gets the €2,000 answer, confidently.

15-minute hands-on

Paste into a free Google Colab notebook. It builds a resale catalogue where discounting genuinely helps high-street items and genuinely hurts luxury ones, then runs every model in the scoreboard above on the same four columns.

import numpy as np, pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

rng = np.random.default_rng(11)
n = 30_000
brand_tier  = rng.uniform(0, 1, n)     # 0 = high street .. 1 = luxury
discount    = rng.uniform(0, 60, n)    # % below the brand's usual resale price
condition   = rng.integers(1, 6, n)
photo_count = rng.integers(1, 10, n)

# THE HIDDEN TRUTH. High street: more discount, more likely to sell.
# Luxury: a modest cut helps, a deep one reads as "probably fake" and kills it.
hs  = 0.055 * discount - 1.4
lux = 1.5 - 0.0062 * (discount - 18.0) ** 2
logit = (-0.55 + 0.42 * (condition - 3) + 0.11 * (photo_count - 5)
         - 0.9 * brand_tier
         + (1 - brand_tier) * hs + brand_tier * lux
         + rng.normal(0, 0.35, n))
sold = (rng.uniform(size=n) < 1 / (1 + np.exp(-logit))).astype(int)

df = pd.DataFrame(dict(brand_tier=brand_tier, discount=discount, condition=condition,
                       photo_count=photo_count, sold=sold))
F = ["brand_tier", "discount", "condition", "photo_count"]
tr, te = train_test_split(df, test_size=0.3, random_state=11, stratify=df.sold)
tr, te = tr.copy(), te.copy()

def score(model, cols, label):
    model.fit(tr[cols], tr.sold)
    auc = roc_auc_score(te.sold, model.predict_proba(te[cols])[:, 1])
    print(f"{label:<40} AUC {auc:.3f}")
    return model

print(f"{'baseline: same score for everyone':<40} AUC 0.500")
lr = score(make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)), F,
           "logistic regression (raw columns)")
score(DecisionTreeClassifier(random_state=11), F,              "one tree, unlimited depth")
score(DecisionTreeClassifier(max_depth=4, random_state=11), F, "one tree, depth 4")
score(RandomForestClassifier(n_estimators=300, min_samples_leaf=20,
                             random_state=11, n_jobs=-1), F,   "random forest, 300 trees")
gb = score(HistGradientBoostingClassifier(random_state=11), F, "gradient boosting")

# Now hand the LINEAR model the interaction it could never find by itself.
for part in (tr, te):
    part["lux_x_disc"]  = part.brand_tier * part.discount
    part["lux_x_disc2"] = part.brand_tier * part.discount ** 2
score(make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)),
      F + ["lux_x_disc", "lux_x_disc2"], "logistic + hand-built interaction")

# The one coefficient the linear model had to pick for discount:
print("\ncoefficients:", dict(zip(F, lr[-1].coef_[0].round(3))))

# The segment that caused the complaint.
seg = te[(te.brand_tier > 0.75) & (te.discount > 40)]
print(f"\nluxury, discounted >40%  (n={len(seg)})")
print(f"  actually sold      {seg.sold.mean():.1%}")
print(f"  logistic predicted {lr.predict_proba(seg[F])[:, 1].mean():.1%}")
print(f"  boosting predicted {gb.predict_proba(seg[F])[:, 1].mean():.1%}")

# And the tree that found the interaction on its own:
print(export_text(DecisionTreeClassifier(max_depth=3, random_state=11).fit(tr[F], tr.sold),
                  feature_names=F, show_weights=True))

Three things to look for. The discount coefficient comes out negative — the linear model’s single best compromise between two populations that want opposite advice. The unlimited-depth tree scores 0.582, below the linear model, because it memorised the training set. And in the printed tree, find the two leaves that share a condition and a discount band but split on brand_tier: that is the interaction, discovered without anyone suggesting it.

Then try the experiment that matters: delete the lux line so discounting helps everyone equally, and re-run. The linear model catches up and the trees stop being worth their complexity. The shape of the problem, not the sophistication of the algorithm, is what decides.

A linear model gives each feature one coefficient for the whole catalogue. When the right answer depends on which item you are looking at, that single number is an average of contradictions. Trees split instead of weigh — so they find the “it depends” you did not know to look for.

Vocabulary from this lesson

parametricnonparametriclinear modellogistic regressioncoefficientdecision boundarydecision treesplitleafinteractionmonotonicensemblebaggingrandom forestboostingresidualAUCextrapolationnormalisation

Lesson 03 of the “Roadmap to be an AI engineer” series · builds on §04 Models of the fundamentals guide · read Real-World ML ch. 3 and the algorithm appendixNext: Top of the class, bottom in the job

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Back to top