Roadmap to be an AI engineer · Lesson 01 · Classic ML

The model that was 97% accurate

Why the most quoted number in machine learning is often the least useful one, and the three numbers an AI engineer looks at instead.

Accuracy · baseline vs model
97.0 / 97.8%
almost no difference
Recall · fakes caught
0% / 80%
the real difference
Precision · flags that were right
60%
the price of catching them
Accuracy asks how often we are right. The business needs to know what we catch, and at what cost.
BaselinesClass imbalanceConfusion matrixPrecision & recallThresholds

One fake-listing detector, measured four ways. Illustrative numbers for 10,000 listings, 300 of them fake.

Thursday, 16:40 · #trust-and-safety

The team demos a new fake-listing detector for the marketplace. The slide says “97% accuracy on held-out data.” Everyone nods.

Three weeks after launch, support is still drowning in counterfeit complaints. Someone checks the logs: the model has flagged zero listings. It never needed to. Only 3% of listings are fake, so a model that calls everything genuine is right 97% of the time. Nobody lied. The wrong number was simply put on the slide.

What happened

What “accurate” actually counted

Take 10,000 new listings: 300 fake, 9,700 genuine. Accuracy is correct predictions ÷ all predictions, and it treats every listing as equally important.

On an imbalanced problem (fraud, churn, defects, anything rare) the majority class dominates that fraction. A model can score brilliantly while ignoring the minority class completely, and that is usually the class the business built the model for. The fix is to look at the four kinds of outcome separately, in a confusion matrix.

Model A“everything is genuine”

said fakesaid genuinereally fake0caught300missedreally genuine0false alarm9,700correct pass

Model Ba real classifier

said fakesaid genuinereally fake240caught60missedreally genuine160false alarm9,540correct pass

Rows are the truth, columns are what the model said. Model A scores 97.0% accuracy, Model B 97.8%: less than a point apart on the slide, yet A catches nothing and B catches four out of five fakes.

The concept

The three numbers that replace accuracy

Recall: of the fakes that really exist, how many did we catch? 240 ÷ 300 = 80%. Precision: of the listings we flagged, how many were really fake? 240 ÷ 400 = 60%. Baseline: what does the dumbest possible model score? Model A is the baseline, and it belonged on the slide next to the result.

Accuracy
A 97.0%
B 97.8%
Recall
A 0%
B 80%
Precision
A undefined, never flags
B 60%
Model A, the baselineModel B

Every bar uses one shared scale. Accuracy hides the difference; recall exposes it.

The trade-off

Precision and recall pull against each other

A classifier usually outputs a probability, and someone picks a threshold above which a listing is flagged. Lower it and you catch more fakes but flag more honest sellers; raise it and the opposite happens. No setting maximises both.

So the threshold is a business decision disguised as a technical one: what does a missed counterfeit cost, and what does a false alarm cost? Real-World Machine Learning (ch. 4) covers the tools for exploring the trade-off, including the ROC curve. Machine Learning Engineering in Action (ch. 3–4) argues that agreeing these costs is part of scoping, before any model is trained.

Low threshold · 0.2Catches nearly every fake; many honest sellers sent to review. Fits when review is cheap and fakes are expensive.
Middle · 0.5The library default. Rarely right for your costs; just the one nobody chose.
High threshold · 0.8Few false alarms, safe to auto-remove. Many fakes slip through.
Many teams run two thresholds: a high one for automatic action and a lower one for human review

For your next standup

Three questions to ask

  1. 1What does the baseline score?If nobody can say what “predict the majority class” or the current rules score, the headline number has no meaning.
  2. 2What are precision and recall on the class we care about, at the threshold we’ll ship?Ask for the confusion matrix, not a single score. Four numbers, one glance.
  3. 3Who chose the threshold, and what does each kind of mistake cost us?If the answer is “0.5, the default”, the business trade-off hasn’t been made yet.

15-minute hands-on

See it for yourself

Open a free Google Colab notebook (or any Python with scikit-learn) and paste this. It builds a dataset where 3% of rows are “fraud”, then compares a do-nothing baseline with a real model.

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report

# 10,000 rows, 3% positive class
X, y = make_classification(n_samples=10_000, n_features=12, n_informative=5,
                           weights=[0.97], flip_y=0, random_state=7)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25,
                                          stratify=y, random_state=7)

models = {
    "baseline: always genuine": DummyClassifier(strategy="most_frequent"),
    "logistic regression":      LogisticRegression(max_iter=1000,
                                                   class_weight="balanced"),
}
for name, m in models.items():
    m.fit(X_tr, y_tr)
    pred = m.predict(X_te)
    print(f"\n=== {name} - accuracy {accuracy_score(y_te, pred):.1%}")
    print(classification_report(y_te, pred, digits=2, zero_division=0))

Read the row labelled 1 in each report. The baseline scores 97% accuracy with recall 0. The logistic regression catches over 90% of the positives, and its accuracy drops to the low 80s, because class_weight="balanced" tells it a missed fraud matters more than a false alarm. Change "balanced" to None, run it again, and watch recall fall while accuracy climbs. You have just moved along this lesson’s trade-off by hand.

Accuracy answers “how often are we right?” The business needs “how often do we catch what matters, and at what cost?” AI engineering starts the moment you ask the second question.

Vocabulary from this lesson

accuracybaselineclass imbalanceconfusion matrixtrue / false positiveprecisionrecallthresholdheld-out test setROC curveclass weighting

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Lesson 01 of “Roadmap to be an AI engineer” · builds on §02 Classic ML · read Real-World ML ch. 1 and 4Next: Features are the product
Back to top