Roadmap to be an AI engineer · Lesson 01 · Classic ML
The model that was 97% accurate
Why the most quoted number in machine learning is often the least useful one, and the three numbers an AI engineer looks at instead.
One fake-listing detector, measured four ways. Illustrative numbers for 10,000 listings, 300 of them fake.
Thursday, 16:40 · #trust-and-safety
The team demos a new fake-listing detector for the marketplace. The slide says “97% accuracy on held-out data.” Everyone nods.
Three weeks after launch, support is still drowning in counterfeit complaints. Someone checks the logs: the model has flagged zero listings. It never needed to. Only 3% of listings are fake, so a model that calls everything genuine is right 97% of the time. Nobody lied. The wrong number was simply put on the slide.
What happened
What “accurate” actually counted
Take 10,000 new listings: 300 fake, 9,700 genuine. Accuracy is correct predictions ÷ all predictions, and it treats every listing as equally important.
On an imbalanced problem (fraud, churn, defects, anything rare) the majority class dominates that fraction. A model can score brilliantly while ignoring the minority class completely, and that is usually the class the business built the model for. The fix is to look at the four kinds of outcome separately, in a confusion matrix.
Model A“everything is genuine”
Model Ba real classifier
Rows are the truth, columns are what the model said. Model A scores 97.0% accuracy, Model B 97.8%: less than a point apart on the slide, yet A catches nothing and B catches four out of five fakes.
The concept
The three numbers that replace accuracy
Recall: of the fakes that really exist, how many did we catch? 240 ÷ 300 = 80%. Precision: of the listings we flagged, how many were really fake? 240 ÷ 400 = 60%. Baseline: what does the dumbest possible model score? Model A is the baseline, and it belonged on the slide next to the result.
Every bar uses one shared scale. Accuracy hides the difference; recall exposes it.
The trade-off
Precision and recall pull against each other
A classifier usually outputs a probability, and someone picks a threshold above which a listing is flagged. Lower it and you catch more fakes but flag more honest sellers; raise it and the opposite happens. No setting maximises both.
So the threshold is a business decision disguised as a technical one: what does a missed counterfeit cost, and what does a false alarm cost? Real-World Machine Learning (ch. 4) covers the tools for exploring the trade-off, including the ROC curve. Machine Learning Engineering in Action (ch. 3–4) argues that agreeing these costs is part of scoping, before any model is trained.
For your next standup
Three questions to ask
- 1
What does the baseline score?
If nobody can say what “predict the majority class” or the current rules score, the headline number has no meaning. - 2
What are precision and recall on the class we care about, at the threshold we’ll ship?
Ask for the confusion matrix, not a single score. Four numbers, one glance. - 3
Who chose the threshold, and what does each kind of mistake cost us?
If the answer is “0.5, the default”, the business trade-off hasn’t been made yet.
15-minute hands-on
See it for yourself
Open a free Google Colab notebook (or any Python with scikit-learn) and paste this. It builds a dataset where 3% of rows are “fraud”, then compares a do-nothing baseline with a real model.
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report
# 10,000 rows, 3% positive class
X, y = make_classification(n_samples=10_000, n_features=12, n_informative=5,
weights=[0.97], flip_y=0, random_state=7)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25,
stratify=y, random_state=7)
models = {
"baseline: always genuine": DummyClassifier(strategy="most_frequent"),
"logistic regression": LogisticRegression(max_iter=1000,
class_weight="balanced"),
}
for name, m in models.items():
m.fit(X_tr, y_tr)
pred = m.predict(X_te)
print(f"\n=== {name} - accuracy {accuracy_score(y_te, pred):.1%}")
print(classification_report(y_te, pred, digits=2, zero_division=0)) Read the row labelled 1 in each report. The baseline scores 97% accuracy with recall 0. The logistic regression catches over 90% of the positives, and its accuracy drops to the low 80s, because class_weight="balanced" tells it a missed fraud matters more than a false alarm. Change "balanced" to None, run it again, and watch recall fall while accuracy climbs. You have just moved along this lesson’s trade-off by hand.
Accuracy answers “how often are we right?” The business needs “how often do we catch what matters, and at what cost?” AI engineering starts the moment you ask the second question.
Vocabulary from this lesson
These lessons were written with the help of AI (Claude) and reviewed and edited by me.