Roadmap To Be An AI Engineer

A field guide from zero: the fundamentals, a step-by-step learning route, hands-on projects to build, and a series of short lessons where each idea is explained through something that actually goes wrong at a real online shop.

Scope & datadecision, metric, labels, features
Build the modelprompt, RAG, fine-tune, train
Evaluateheld-out data, evals, baselines
Deploy & monitorserve, watch drift, retrain
Drift in production sends you back to the data: it is a loop, not a line
Data engineeringEvaluationMLOpsSafety & securityCost & latency

The AI engineering lifecycle. A model that was right at launch quietly becomes wrong as the world changes. Everything on this page hangs on this picture.

Start here

What AI engineering is

Data engineering moves data from where it is created to where it can be trusted. AI engineering starts where that ends: it turns trusted data and a model, one you trained or one you rent, into a product feature that behaves reliably in front of real users, measurably, safely and at a cost the business can carry.

The work runs in the loop shown above. Machine Learning Engineering in Action (Wilson) argues that most failed ML projects die in scoping and in monitoring, not in the modelling. Those are exactly the two places where a technical leader has the most leverage.

01 · The first architecture decision

The build ladder: rent, ground, adapt, or train

For language tasks, the most consequential choice is how much of the model you own. Each rung up costs more data, more compute and rarer skills. Climb only when the rung below has been measured and found insufficient.

1 · Prompt a modelHours. No training data. Start every feature here.
2 · RAGDays. Your documents fetched at question time. Adds knowledge, not skill.
3 · Fine-tuneWeeks. 1k–100k examples. Changes behaviour and format.
4 · PretrainMonths. Billions of tokens. Almost never the answer.…but building a tiny one is the best way to learn.
cheaper, fastermore cost, data, skill and ownership

RAG fixes “the model doesn’t know our facts”; fine-tuning fixes “the model doesn’t behave the way we need”. Raschka’s Build a Large Language Model (From Scratch) walks rungs 3 and 4 by hand. For tabular predictions such as fraud, churn or demand, a classic ML model on your own data usually beats any LLM on cost and accuracy.

02 · The foundation under everything

Classic machine learning

Every model, from a two-line logistic regression to a frontier LLM, is the same idea: a function with adjustable parameters, tuned on examples so its outputs match known answers, then trusted on inputs it has never seen. Learn it on small tabular problems first; the vocabulary carries straight over to LLMs.

Supervised

Learn from labelled examples

Fraud or not, sale price. Classification predicts a category, regression a number.

Unsupervised

Find structure without labels

Customer segments, unusual transactions, compressed representations.

Self-supervised

Labels come from the data

Hide the next word and predict it. This is how LLMs are pretrained.

Real-World Machine Learning (ch. 1–4) frames the workflow you will practise again and again. Four ideas carry most of the weight:

  • Train, validation and test split. Tune on one slice, choose on another, report on a third nobody touched.
  • Overfitting. A model that memorises its training data looks brilliant there and fails in production. Cross-validation estimates real-world behaviour.
  • Baselines. Always compare against “predict the most common answer” or today’s rules. A model that can’t beat a dumb baseline is a cost, not an asset.
  • The right metric. Accuracy lies on imbalanced problems. Precision, recall and the cost of each mistake decide whether to ship. This is Lesson 01.

03 · Where data engineering meets AI

Data and features decide the result

A feature is one measurable input the model sees: days since the seller’s last sale, the price relative to the brand median, the number of photos. Real-World ML (ch. 5 and 7) shows that features built from domain knowledge often beat a fancier algorithm.

Features are computed by pipelines, so they must be identical at training and prediction time (training–serving skew) and may only use information available at the moment of prediction (data leakage). Fundamentals of Data Engineering (ch. 9) treats ML as one more consumer of the serving stage, which is the right mental model.

Why a leader cares

When an ML project disappoints, the cause is far more often missing labels, leaky features or a skewed pipeline than a weak algorithm. Ask about the data before you ask about the model.

04 · The engine inside modern AI

Neural networks and the training loop

A neural network is layers of simple units, each computing a weighted sum of its inputs and passing it through a non-linear function. NLP in Action (ch. 5) builds one from a single neuron; Appendix A of Raschka’s book is a compact PyTorch primer. Training is the same loop whether the model has a thousand parameters or a trillion:

Forward passrun a batch, get predictions
Lossone number: how wrong were we?
Backpropagationeach weight’s share of the error
Optimiser stepnudge every weight downhill
Repeat for many epochs until the loss on held-out data stops improving

For LLMs the loss is cross-entropy on the next token, and Raschka’s training loop uses the AdamW optimiser. Hyperparameters such as learning rate and batch size are the knobs you set; weights are what training learns.

05 · Language as numbers

Tokens and embeddings

Models only do arithmetic, so text must become numbers. NLP in Action tells this history in order, and each step is still in use: bag-of-words counts (ch. 2), TF-IDF weighting behind classic keyword search (ch. 3), topic vectors (ch. 4), and word vectors (ch. 6), where words with similar meaning sit close together and closeness is measured with cosine similarity.

LLMs add two upgrades (Raschka ch. 2): text is split into tokens, sub-word pieces of roughly three-quarters of a word, and each token is mapped to a learned embedding vector plus its position.

Text“Vintage Levi’s”
TokensVint | age | Lev | i | ‘s
Embeddings
Transformer × Nattention + feed-forward
Next token
jeans
jacket
501
Pick one token, append it, repeat: that is all generation is

Token splits are illustrative; each model has its own tokeniser. Two practical consequences: prices and context limits are counted in tokens, and the same embedding idea powers semantic search and RAG.

06 · The architecture behind LLMs

Transformers and attention

Older language models read one word at a time with recurrent networks (NLP in Action ch. 8–9) and struggled to link words far apart. The transformer replaced that with self-attention: for each token, the model weighs how much every earlier token should influence it. In “the jacket fits, but it runs small”, attention is how it gets linked to jacket.

Raschka builds this in chapter 3 up to causal multi-head attention (a token may only look backwards; several attention patterns run in parallel), then stacks it into a GPT block in chapter 4.

The one sentence to remember

An LLM is a stack of attention blocks trained to predict the next token. Chat, coding and reasoning all emerge from doing that very well on enough data.

07 · From random weights to assistant

How LLMs are made

Raschka organises the process into three stages, and knowing them lets you judge vendor claims and fine-tuning proposals.

Stage 1 · Builddata sampling, attention, GPT architecture (ch. 2–4)
Stage 2 · Pretrainnext-token prediction on huge text → foundation model (ch. 5)
Stage 3 · Fine-tuneas a classifier (ch. 6) or to follow instructions (ch. 7)
Instruction fine-tuning
Training on instruction and response pairs, turning a text-completer into something that answers requests.
LoRA / PEFT
Freeze the big model and train small add-on matrices (Raschka appendix E). Makes fine-tuning affordable on one GPU.
Preference tuning
An optional further step (such as DPO or RLHF) that trains toward responses people prefer.
Open vs closed weights
Open-weight models can be downloaded and self-hosted; closed ones are reached only by API. A data-residency and cost decision as much as a quality one.

08 · What most AI engineers build today

Building with LLMs: prompts, RAG, tools and agents

Most day-to-day work sits on rungs 1 and 2: wiring a hosted model into a product with good instructions, the right context and a way to act. Semantic search with vectors is covered in NLP in Action (ch. 3–6); prompting, tool use and agents are newer than the books, so those lessons use current documentation.

Offline: indexing, a data pipeline

Documentsdocs, policies, tickets
Chunksplit into passages
Embedpassage → vector
Vector indexe.g. Postgres + pgvector

Online: every user question

Questionfrom the user
Retrieve top-kmost similar passages
Promptquestion + passages
LLManswer with citations
Most RAG failures are retrieval failures: measure retrieval separately from answer quality

Retrieval-augmented generation. The top row is a data pipeline in everything but name, which is why data engineers move into AI work so naturally.

Prompt

Prompt and context engineering

Instructions, examples, output format and what fits in the context window. Prompts are code: versioned, reviewed, tested.

RAG

Retrieval-augmented generation

Fetch relevant passages at question time. Keeps answers current and citable without retraining.

Tools

Tool use

The model asks for an action (“look up order 88421”); your code runs it and returns the result.

Agents

Agents

An LLM in a loop of planning and tool calls. Powerful, harder to make reliable: errors compound step by step.

09 · Knowing it works

Evaluation separates demos from products

A demo works on the five examples its builder tried. A product works on the ten thousand its users try. Real-World ML (ch. 4) gives the classic toolkit: held-out data, cross-validation, confusion matrices, ROC curves. Raschka (ch. 7) adds using a second model as a judge to score free-text answers.

  • Golden dataset. A fixed set of real inputs with known-good outputs, built with the business. Fifty well-chosen cases beat five thousand random ones.
  • Automatic metrics. Exact match, precision and recall, retrieval hit rate. Run them on every change, like unit tests.
  • LLM-as-judge. A model grades answers against a rubric. Calibrate it against human ratings before trusting it.
  • Human review and online metrics. Spot checks, A/B tests and business outcomes, the measures that ultimately count.
The single most useful question

If a team cannot show its eval set and the score before and after a change, it does not yet know whether its AI feature works.

10 · Making it run reliably

Production ML: MLOps, serving and drift

Machine Learning Engineering in Action carries this section, and its message is uncomfortable: the model is a small part of the system.

Scoping
Agree the decision, the metric, the budget and the “good enough” bar before anyone trains anything (Wilson ch. 3–4).
Experiment tracking
Record every run’s data, code, parameters and metrics (MLflow in the book) so results are reproducible.
Model registry
A versioned store of models marking which one is live, so you can roll back without rebuilding (ch. 16).
Batch vs online
Score everything nightly into a table, or answer each request in milliseconds. Batch is simpler and cheaper.
Feature store
One place that computes features identically for training and serving.
Drift
Feature, label and concept drift degrade a model over time (ch. 12). Watch input, prediction and business metrics.
Passive vs active retraining
Retrain on a schedule and keep the new model only if it wins on fresh data, or retrain when monitoring trips a threshold. Wilson favours the simpler passive approach where possible.

11 · The disciplines beneath everything

Safety, security, privacy and cost

Security

Prompt injection

Text in a document or email that tries to instruct the model. Limit what tools can do and confirm side effects with a human.

Safety

Hallucination and guardrails

Fluent, confident errors. Ground answers in sources, require citations, validate structured output.

Privacy

GDPR and the EU AI Act

Know which personal data enters prompts, training sets and vendor logs, and minimise it. The use case itself now carries compliance weight.

Cost

Cost and latency

Tokens are the new bytes scanned. Model choice, prompt length and caching can move unit cost 10–100×.

12 · What you’d actually learn

The core toolbox

Depth beats breadth. Python, one ML library and one LLM API take you further than a tour of twenty frameworks.

1

Python, NumPy, pandas

The language of all AI work. Tensors are just arrays with more dimensions.

2

scikit-learn

Classic ML in one consistent API: splits, pipelines, models, metrics.

3

PyTorch

Raschka’s library throughout. Learn tensors, autograd and a training loop.

4

Hugging Face and an LLM API

Open models and tokenisers, plus one hosted API for prompting, tools and structured output.

5

Vector search

pgvector inside Postgres is the pragmatic start. Specialised databases only at scale.

6

MLOps basics

MLflow, FastAPI and Docker, GitHub Actions, and Airflow from the data roadmap.

13 · Recognise it in a standup

Vocabulary that pays off immediately

Training vs inference
Learning the weights (expensive, occasional) vs using them (cheap per call, constant). Most of the bill is inference.
Precision vs recall
Of what we flagged, how much was right, vs of what was there, how much we caught.
Overfitting
Memorising training data instead of learning the pattern.
Data leakage
Future information or the answer itself sneaking into features.
Token
The unit an LLM reads and writes, about three-quarters of a word.
Context window
How many tokens a model can consider at once.
Embedding
A vector representing meaning; similar things sit close together.
Temperature
Randomness in picking the next token. Low for extraction, higher for creative text.
Hallucination
Fluent output not supported by facts or sources.
Fine-tuning vs RAG
Change how the model behaves vs change what it knows at question time.
Eval
A repeatable test suite for model behaviour: the CI of AI features.
Drift
Production data moving away from training data.
Agent
An LLM that loops through planning and tool calls to finish a task.

14 · Step-by-step learning route

Thirty short lessons, one a day

Each lesson follows the same shape: a real situation, one visual, the questions to ask in standup, a 15-minute hands-on exercise and the vocabulary it unlocks. The last column says which book chapter to read for depth.

#LessonConceptRead
Week 1 · Machine learning foundations → project P1
01The model that was 97% accuratereadBaselines, precision, recallRWML 1, 4
02Features are the productreadFeature engineeringRWML 2, 5
03Lines and treesreadLinear models vs trees and boostingRWML 3
04Top of the class, bottom in the jobreadOverfitting, cross-validationRWML 4
05The model that saw tomorrowreadData leakage, time-based splitsRWML 4, 6
Week 2 · From numbers to neurons → project P2
06A neuron is a weighted sumreadNeural networksNLPiA 5 · LLM App. A
07Walking downhill in the darkLoss, gradient descentLLM 5
08Counting wordsBag of words, TF-IDF, BM25NLPiA 2–3
09King minus man plus womanWord vectorsNLPiA 6
10Search that understands “cheap”Cosine similarity, semantic searchNLPiA 3–4
Week 3 · Inside an LLM → project P3
11Why the bill counts tokensTokenisation, BPELLM 2
12What “it” refers toSelf-attentionLLM 3
13One GPT block, end to endTransformer architectureLLM 4
14Reading the internet oncePretrainingLLM 5
15Same prompt, different answerTemperature, top-kLLM 5
Week 4 · Adapting models → projects P3, P4
16Teaching a GPT to spot spamClassification fine-tuningLLM 6
17From autocomplete to assistantInstruction fine-tuningLLM 7
18Fine-tuning on one GPULoRA, PEFTLLM App. E
19Prompts are codePrompt and context engineeringcurrent docs
20An open-book examRAG end to endNLPiA 3–6
Week 5 · Evaluation and building blocks → project P4
21The demo that worked five timesEvals, golden sets, LLM-as-judgeLLM 7 · RWML 4
22Where to cut the documentChunking, vector indexescurrent docs
23The model that can press buttonsTool use, agentscurrent docs
24The email that gave ordersPrompt injection, guardrailscurrent docs
25Tokens are the new bytes scannedCost, latency, cachingcurrent docs
Week 6 · Production → project P5
26The project that should never have startedScoping ML workMLEiA 3–4
27Which model is live?Experiment tracking, registryMLEiA 5–7, 16
28Nightly or instantBatch vs online inferenceMLEiA 16 · FoDE 9
29The model that got worse by itselfDrift, retrainingMLEiA 12
30Where the pipelines meetFeature stores, the AI platformFoDE 9 · MLEiA 16

RWML = Real-World Machine Learning · NLPiA = Natural Language Processing in Action · LLM = Build a Large Language Model (From Scratch) · MLEiA = Machine Learning Engineering in Action · FoDE = Fundamentals of Data Engineering.

15 · Hands-on projects to build

Five projects for GitHub

Each project closes a week, reuses the previous one, and exercises data engineering as well as AI. A good README states the question, the baseline, the eval result and what you would do next.

  1. P1

    taxi-tip-model · end of week 1

    Predict whether a New York taxi passenger tips, the worked example of Real-World ML ch. 6, using the BigQuery public dataset.

    Partitioned extract to Parquet, scikit-learn pipeline, baseline vs logistic regression vs gradient boosting, confusion matrix in the README.

  2. P2

    review-sentiment · end of week 2

    Classify reviews (Real-World ML ch. 8) three ways: TF-IDF with logistic regression, averaged word vectors, and a small PyTorch network.

    Shows what each representation buys you. Add a tiny semantic-search tool over the same reviews.

  3. P3

    tiny-gpt · weeks 3–4

    Follow Raschka chapter by chapter: tokeniser, attention, GPT model, pretraining on a small public-domain text, then fine-tune a spam classifier with LoRA.

    Turns “transformer” from a buzzword into code you wrote. Runs on a laptop or a free Colab GPU.

  4. P4

    docs-rag · weeks 4–5

    A question-answering service over public documentation you know well, with no confidential data.

    Chunking pipeline into Postgres and pgvector, an LLM API for cited answers, a 40-question golden set, and a report on retrieval hit rate, quality and cost per answer.

  5. P5

    ml-in-production · week 6

    Take P1 to production standards: MLflow tracking and registry, FastAPI and Docker serving, an Airflow DAG for scheduled retraining, a drift report and GitHub Actions tests.

    The repo that proves you can sit with the team and review their pipelines line by line.

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Roadmap to be an AI engineer · a field guide for technical leadersTheory drawn from five reference books, paraphrased and cited by chapter
Back to top