Roadmap To Be An AI Engineer
A field guide from zero: the fundamentals, a step-by-step learning route, hands-on projects to build, and a series of short lessons where each idea is explained through something that actually goes wrong at a real online shop.
The AI engineering lifecycle. A model that was right at launch quietly becomes wrong as the world changes. Everything on this page hangs on this picture.
Start here
What AI engineering is
Data engineering moves data from where it is created to where it can be trusted. AI engineering starts where that ends: it turns trusted data and a model, one you trained or one you rent, into a product feature that behaves reliably in front of real users, measurably, safely and at a cost the business can carry.
The work runs in the loop shown above. Machine Learning Engineering in Action (Wilson) argues that most failed ML projects die in scoping and in monitoring, not in the modelling. Those are exactly the two places where a technical leader has the most leverage.
01 · The first architecture decision
The build ladder: rent, ground, adapt, or train
For language tasks, the most consequential choice is how much of the model you own. Each rung up costs more data, more compute and rarer skills. Climb only when the rung below has been measured and found insufficient.
RAG fixes “the model doesn’t know our facts”; fine-tuning fixes “the model doesn’t behave the way we need”. Raschka’s Build a Large Language Model (From Scratch) walks rungs 3 and 4 by hand. For tabular predictions such as fraud, churn or demand, a classic ML model on your own data usually beats any LLM on cost and accuracy.
02 · The foundation under everything
Classic machine learning
Every model, from a two-line logistic regression to a frontier LLM, is the same idea: a function with adjustable parameters, tuned on examples so its outputs match known answers, then trusted on inputs it has never seen. Learn it on small tabular problems first; the vocabulary carries straight over to LLMs.
Learn from labelled examples
Fraud or not, sale price. Classification predicts a category, regression a number.
Find structure without labels
Customer segments, unusual transactions, compressed representations.
Labels come from the data
Hide the next word and predict it. This is how LLMs are pretrained.
Real-World Machine Learning (ch. 1–4) frames the workflow you will practise again and again. Four ideas carry most of the weight:
- Train, validation and test split. Tune on one slice, choose on another, report on a third nobody touched.
- Overfitting. A model that memorises its training data looks brilliant there and fails in production. Cross-validation estimates real-world behaviour.
- Baselines. Always compare against “predict the most common answer” or today’s rules. A model that can’t beat a dumb baseline is a cost, not an asset.
- The right metric. Accuracy lies on imbalanced problems. Precision, recall and the cost of each mistake decide whether to ship. This is Lesson 01.
03 · Where data engineering meets AI
Data and features decide the result
A feature is one measurable input the model sees: days since the seller’s last sale, the price relative to the brand median, the number of photos. Real-World ML (ch. 5 and 7) shows that features built from domain knowledge often beat a fancier algorithm.
Features are computed by pipelines, so they must be identical at training and prediction time (training–serving skew) and may only use information available at the moment of prediction (data leakage). Fundamentals of Data Engineering (ch. 9) treats ML as one more consumer of the serving stage, which is the right mental model.
When an ML project disappoints, the cause is far more often missing labels, leaky features or a skewed pipeline than a weak algorithm. Ask about the data before you ask about the model.
04 · The engine inside modern AI
Neural networks and the training loop
A neural network is layers of simple units, each computing a weighted sum of its inputs and passing it through a non-linear function. NLP in Action (ch. 5) builds one from a single neuron; Appendix A of Raschka’s book is a compact PyTorch primer. Training is the same loop whether the model has a thousand parameters or a trillion:
For LLMs the loss is cross-entropy on the next token, and Raschka’s training loop uses the AdamW optimiser. Hyperparameters such as learning rate and batch size are the knobs you set; weights are what training learns.
05 · Language as numbers
Tokens and embeddings
Models only do arithmetic, so text must become numbers. NLP in Action tells this history in order, and each step is still in use: bag-of-words counts (ch. 2), TF-IDF weighting behind classic keyword search (ch. 3), topic vectors (ch. 4), and word vectors (ch. 6), where words with similar meaning sit close together and closeness is measured with cosine similarity.
LLMs add two upgrades (Raschka ch. 2): text is split into tokens, sub-word pieces of roughly three-quarters of a word, and each token is mapped to a learned embedding vector plus its position.
Token splits are illustrative; each model has its own tokeniser. Two practical consequences: prices and context limits are counted in tokens, and the same embedding idea powers semantic search and RAG.
06 · The architecture behind LLMs
Transformers and attention
Older language models read one word at a time with recurrent networks (NLP in Action ch. 8–9) and struggled to link words far apart. The transformer replaced that with self-attention: for each token, the model weighs how much every earlier token should influence it. In “the jacket fits, but it runs small”, attention is how it gets linked to jacket.
Raschka builds this in chapter 3 up to causal multi-head attention (a token may only look backwards; several attention patterns run in parallel), then stacks it into a GPT block in chapter 4.
An LLM is a stack of attention blocks trained to predict the next token. Chat, coding and reasoning all emerge from doing that very well on enough data.
07 · From random weights to assistant
How LLMs are made
Raschka organises the process into three stages, and knowing them lets you judge vendor claims and fine-tuning proposals.
- Instruction fine-tuning
- Training on instruction and response pairs, turning a text-completer into something that answers requests.
- LoRA / PEFT
- Freeze the big model and train small add-on matrices (Raschka appendix E). Makes fine-tuning affordable on one GPU.
- Preference tuning
- An optional further step (such as DPO or RLHF) that trains toward responses people prefer.
- Open vs closed weights
- Open-weight models can be downloaded and self-hosted; closed ones are reached only by API. A data-residency and cost decision as much as a quality one.
08 · What most AI engineers build today
Building with LLMs: prompts, RAG, tools and agents
Most day-to-day work sits on rungs 1 and 2: wiring a hosted model into a product with good instructions, the right context and a way to act. Semantic search with vectors is covered in NLP in Action (ch. 3–6); prompting, tool use and agents are newer than the books, so those lessons use current documentation.
Offline: indexing, a data pipeline
Online: every user question
Retrieval-augmented generation. The top row is a data pipeline in everything but name, which is why data engineers move into AI work so naturally.
Prompt and context engineering
Instructions, examples, output format and what fits in the context window. Prompts are code: versioned, reviewed, tested.
Retrieval-augmented generation
Fetch relevant passages at question time. Keeps answers current and citable without retraining.
Tool use
The model asks for an action (“look up order 88421”); your code runs it and returns the result.
Agents
An LLM in a loop of planning and tool calls. Powerful, harder to make reliable: errors compound step by step.
09 · Knowing it works
Evaluation separates demos from products
A demo works on the five examples its builder tried. A product works on the ten thousand its users try. Real-World ML (ch. 4) gives the classic toolkit: held-out data, cross-validation, confusion matrices, ROC curves. Raschka (ch. 7) adds using a second model as a judge to score free-text answers.
- Golden dataset. A fixed set of real inputs with known-good outputs, built with the business. Fifty well-chosen cases beat five thousand random ones.
- Automatic metrics. Exact match, precision and recall, retrieval hit rate. Run them on every change, like unit tests.
- LLM-as-judge. A model grades answers against a rubric. Calibrate it against human ratings before trusting it.
- Human review and online metrics. Spot checks, A/B tests and business outcomes, the measures that ultimately count.
If a team cannot show its eval set and the score before and after a change, it does not yet know whether its AI feature works.
10 · Making it run reliably
Production ML: MLOps, serving and drift
Machine Learning Engineering in Action carries this section, and its message is uncomfortable: the model is a small part of the system.
- Scoping
- Agree the decision, the metric, the budget and the “good enough” bar before anyone trains anything (Wilson ch. 3–4).
- Experiment tracking
- Record every run’s data, code, parameters and metrics (MLflow in the book) so results are reproducible.
- Model registry
- A versioned store of models marking which one is live, so you can roll back without rebuilding (ch. 16).
- Batch vs online
- Score everything nightly into a table, or answer each request in milliseconds. Batch is simpler and cheaper.
- Feature store
- One place that computes features identically for training and serving.
- Drift
- Feature, label and concept drift degrade a model over time (ch. 12). Watch input, prediction and business metrics.
- Passive vs active retraining
- Retrain on a schedule and keep the new model only if it wins on fresh data, or retrain when monitoring trips a threshold. Wilson favours the simpler passive approach where possible.
11 · The disciplines beneath everything
Safety, security, privacy and cost
Prompt injection
Text in a document or email that tries to instruct the model. Limit what tools can do and confirm side effects with a human.
Hallucination and guardrails
Fluent, confident errors. Ground answers in sources, require citations, validate structured output.
GDPR and the EU AI Act
Know which personal data enters prompts, training sets and vendor logs, and minimise it. The use case itself now carries compliance weight.
Cost and latency
Tokens are the new bytes scanned. Model choice, prompt length and caching can move unit cost 10–100×.
12 · What you’d actually learn
The core toolbox
Depth beats breadth. Python, one ML library and one LLM API take you further than a tour of twenty frameworks.
Python, NumPy, pandas
The language of all AI work. Tensors are just arrays with more dimensions.
scikit-learn
Classic ML in one consistent API: splits, pipelines, models, metrics.
PyTorch
Raschka’s library throughout. Learn tensors, autograd and a training loop.
Hugging Face and an LLM API
Open models and tokenisers, plus one hosted API for prompting, tools and structured output.
Vector search
pgvector inside Postgres is the pragmatic start. Specialised databases only at scale.
MLOps basics
MLflow, FastAPI and Docker, GitHub Actions, and Airflow from the data roadmap.
13 · Recognise it in a standup
Vocabulary that pays off immediately
- Training vs inference
- Learning the weights (expensive, occasional) vs using them (cheap per call, constant). Most of the bill is inference.
- Precision vs recall
- Of what we flagged, how much was right, vs of what was there, how much we caught.
- Overfitting
- Memorising training data instead of learning the pattern.
- Data leakage
- Future information or the answer itself sneaking into features.
- Token
- The unit an LLM reads and writes, about three-quarters of a word.
- Context window
- How many tokens a model can consider at once.
- Embedding
- A vector representing meaning; similar things sit close together.
- Temperature
- Randomness in picking the next token. Low for extraction, higher for creative text.
- Hallucination
- Fluent output not supported by facts or sources.
- Fine-tuning vs RAG
- Change how the model behaves vs change what it knows at question time.
- Eval
- A repeatable test suite for model behaviour: the CI of AI features.
- Drift
- Production data moving away from training data.
- Agent
- An LLM that loops through planning and tool calls to finish a task.
14 · Step-by-step learning route
Thirty short lessons, one a day
Each lesson follows the same shape: a real situation, one visual, the questions to ask in standup, a 15-minute hands-on exercise and the vocabulary it unlocks. The last column says which book chapter to read for depth.
| # | Lesson | Concept | Read |
|---|---|---|---|
| Week 1 · Machine learning foundations → project P1 | |||
| 01 | The model that was 97% accurateread | Baselines, precision, recall | RWML 1, 4 |
| 02 | Features are the productread | Feature engineering | RWML 2, 5 |
| 03 | Lines and treesread | Linear models vs trees and boosting | RWML 3 |
| 04 | Top of the class, bottom in the jobread | Overfitting, cross-validation | RWML 4 |
| 05 | The model that saw tomorrowread | Data leakage, time-based splits | RWML 4, 6 |
| Week 2 · From numbers to neurons → project P2 | |||
| 06 | A neuron is a weighted sumread | Neural networks | NLPiA 5 · LLM App. A |
| 07 | Walking downhill in the dark | Loss, gradient descent | LLM 5 |
| 08 | Counting words | Bag of words, TF-IDF, BM25 | NLPiA 2–3 |
| 09 | King minus man plus woman | Word vectors | NLPiA 6 |
| 10 | Search that understands “cheap” | Cosine similarity, semantic search | NLPiA 3–4 |
| Week 3 · Inside an LLM → project P3 | |||
| 11 | Why the bill counts tokens | Tokenisation, BPE | LLM 2 |
| 12 | What “it” refers to | Self-attention | LLM 3 |
| 13 | One GPT block, end to end | Transformer architecture | LLM 4 |
| 14 | Reading the internet once | Pretraining | LLM 5 |
| 15 | Same prompt, different answer | Temperature, top-k | LLM 5 |
| Week 4 · Adapting models → projects P3, P4 | |||
| 16 | Teaching a GPT to spot spam | Classification fine-tuning | LLM 6 |
| 17 | From autocomplete to assistant | Instruction fine-tuning | LLM 7 |
| 18 | Fine-tuning on one GPU | LoRA, PEFT | LLM App. E |
| 19 | Prompts are code | Prompt and context engineering | current docs |
| 20 | An open-book exam | RAG end to end | NLPiA 3–6 |
| Week 5 · Evaluation and building blocks → project P4 | |||
| 21 | The demo that worked five times | Evals, golden sets, LLM-as-judge | LLM 7 · RWML 4 |
| 22 | Where to cut the document | Chunking, vector indexes | current docs |
| 23 | The model that can press buttons | Tool use, agents | current docs |
| 24 | The email that gave orders | Prompt injection, guardrails | current docs |
| 25 | Tokens are the new bytes scanned | Cost, latency, caching | current docs |
| Week 6 · Production → project P5 | |||
| 26 | The project that should never have started | Scoping ML work | MLEiA 3–4 |
| 27 | Which model is live? | Experiment tracking, registry | MLEiA 5–7, 16 |
| 28 | Nightly or instant | Batch vs online inference | MLEiA 16 · FoDE 9 |
| 29 | The model that got worse by itself | Drift, retraining | MLEiA 12 |
| 30 | Where the pipelines meet | Feature stores, the AI platform | FoDE 9 · MLEiA 16 |
RWML = Real-World Machine Learning · NLPiA = Natural Language Processing in Action · LLM = Build a Large Language Model (From Scratch) · MLEiA = Machine Learning Engineering in Action · FoDE = Fundamentals of Data Engineering.
15 · Hands-on projects to build
Five projects for GitHub
Each project closes a week, reuses the previous one, and exercises data engineering as well as AI. A good README states the question, the baseline, the eval result and what you would do next.
- P1
taxi-tip-model · end of week 1
Predict whether a New York taxi passenger tips, the worked example of Real-World ML ch. 6, using the BigQuery public dataset.
Partitioned extract to Parquet, scikit-learn pipeline, baseline vs logistic regression vs gradient boosting, confusion matrix in the README.
- P2
review-sentiment · end of week 2
Classify reviews (Real-World ML ch. 8) three ways: TF-IDF with logistic regression, averaged word vectors, and a small PyTorch network.
Shows what each representation buys you. Add a tiny semantic-search tool over the same reviews.
- P3
tiny-gpt · weeks 3–4
Follow Raschka chapter by chapter: tokeniser, attention, GPT model, pretraining on a small public-domain text, then fine-tune a spam classifier with LoRA.
Turns “transformer” from a buzzword into code you wrote. Runs on a laptop or a free Colab GPU.
- P4
docs-rag · weeks 4–5
A question-answering service over public documentation you know well, with no confidential data.
Chunking pipeline into Postgres and pgvector, an LLM API for cited answers, a 40-question golden set, and a report on retrieval hit rate, quality and cost per answer.
- P5
ml-in-production · week 6
Take P1 to production standards: MLflow tracking and registry, FastAPI and Docker serving, an Airflow DAG for scheduled retraining, a drift report and GitHub Actions tests.
The repo that proves you can sit with the team and review their pipelines line by line.
These lessons were written with the help of AI (Claude) and reviewed and edited by me.