Roadmap To Be A DevOps Engineer / Lesson 01
From Laptop to Live
A Friday fix copied over SFTP takes every product page down. The four faults a pipeline removes.
1The situation
17:10 Fri — the fix. A developer fixes a bug on the shop: the "under €50" filter hides items priced at exactly €50. Three lines of code. It works on his laptop. He opens an SFTP client, drags the changed files onto the production server, and closes his laptop.
17:24 Fri — the outage. Every product page returns a 500. The fix was written against a newer library version than the one installed on the server. It worked on his machine because his machine had the newer library.
17:46 Fri — the recovery. Nobody knows exactly which files were replaced, so the team restores last night’s full backup. Twenty-two minutes of lost checkout traffic — and the original bug is back.
2Four failures, four practices
| Fault | What happened | The practice that removes it |
|---|---|---|
| 01 | Laptop and server had different library versions, and nothing checked | containers & infrastructure as code |
| 02 | Nothing stood between the developer and production | continuous integration |
| 03 | The previous version existed only as a backup, not as a deployable thing | versioned artifacts & rollback |
| 04 | Customers noticed the outage 14 minutes before the team did | monitoring & alerting |
Every later tool in the roadmap — Docker, CI pipelines, Terraform, Kubernetes, Prometheus — is an answer to one of these four faults. If you lose the thread of why a tool exists, ask which fault it removes.
3What a pipeline actually changes
The pipeline does not make the developer more careful. It inserts a build that pins the same library version everywhere, a test gate the change must pass, and a registry that keeps the previous version deployable. Rollback stops being an archaeology project and becomes one command.
The one sentence to keep: a deploy pipeline turns "hope it works" into "prove it works, and keep the last thing that did."
4The scoreboard: four DORA metrics
| Metric | Before | After | What it means |
|---|---|---|---|
| Deploy frequency | ~1 / week | 14 / week | how often a change can reach customers |
| Lead time for change | 3 days | 25 min | commit → live; what stakeholders feel as "speed" |
| Change failure rate | 30% | 6% | share of deploys needing a fix or rollback |
| Time to restore | 22 min | 4 min | broken → healthy; what customers actually feel |
Note the last two. Intuition says shipping more often is riskier; the research says the opposite, and the reason is batch size — a deploy carrying 3 lines has a small surface for failure, and when it breaks you already know which 3 lines to blame.
5Hands-on (15 minutes)
-
A project with one function and one test
# price_filter.py def under(price, limit): return price <= limit # test_price_filter.py from price_filter import under def test_boundary(): assert under(50, 50) is True -
Write the gate — this script is a CI pipeline; GitHub Actions and Jenkins are this idea plus scheduling, logs and permissions.
#!/usr/bin/env bash set -euo pipefail # stop at the first failure echo "▸ test"; pytest -q echo "▸ build"; tar -czf build-$(git rev-parse --short HEAD).tgz *.py echo "▸ deploy"; echo "would ship build-$(git rev-parse --short HEAD).tgz" -
Break it on purpose — change the function to
price < limitand rerun. The test fails,set -estops the script, the deploy line never runs. That refusal is the whole value of CI. -
Notice what is still missing — the build ran on your machine with your Python version. Fault 01 is still open. That is Docker, and it is where the roadmap goes next.
6Three questions to ask your team this week
- "If we had to undo the last deploy right now, what would we do and how long would it take?" A good answer names a command; a worried answer names a backup.
- "What has to pass before code reaches customers, and can a person skip it?" This maps your real gates, not the ones on the wiki.
- "How do we find out that production is broken?" If the answer is "support tickets", monitoring is the highest-value thing to build next.
Next: Fault 01 up close — why "works on my machine" happens, and what a container image actually is (a filesystem plus a start command, not a small virtual machine).
Written with the help of AI (Claude) and reviewed by Rayhanul Islam.