Roadmap to be a DevOps engineer
A field guide from zero: the fundamentals, a step-by-step learning route, hands-on projects to build, and a series of short lessons where each idea is explained through something that actually goes wrong at a real online shop.
The DevOps lifecycle. Dev builds it, Ops runs it, and DevOps automates the road between them. Everything on this page hangs on this picture.
The fundamentals
DevOps is the practice of getting an idea from a developer’s laptop into customers’ hands safely and repeatedly, and then keeping it healthy. It merges building software and running it, and automates the road between them. These are the ideas a DevOps engineer works with every day.
§1 Culture and shared ownership
The people who build software also share responsibility for running it: “you build it, you run it”. Outages are blamed on processes, not people.
- Blameless post-incident reviews
- Runbooks, so knowledge is not trapped in one head
- One accountable name per responsibility
- Developers in the on-call rota
You can buy every tool below and still fail if teams do not cooperate. Culture is the part a CTO shapes most directly.
§2 Version control (Git)
Git records every change — who, what, when and why — and is the single source of truth that the whole loop triggers from.
- Branches and pull requests: review built into the workflow
- Branch protection and required checks
- Tags and semantic versions
- Configuration and infrastructure live here too
No version control, no DevOps: builds, tests and deploys all start from a change landing in Git.
§3 CI/CD, the automated pipeline
Continuous integration builds and tests every push within minutes; continuous delivery carries what passed towards production.
- push → build → test → scan → stage → production
- A failing step stops the change: the automation is the safety net
- Build once, promote the same artifact
- GitHub Actions, GitLab CI, Jenkins, Argo CD
The single biggest lever on the four DORA numbers: from monthly releases with held breath to many quiet deploys a day.
§4 Infrastructure as code
Servers, networks and databases are described in text files, reviewed in pull requests, and a tool makes reality match them.
- Plan before apply; state tracks what exists
- The same definition for staging and production
- Drift: what someone clicked versus what the code says
- Terraform, OpenTofu, CloudFormation, Pulumi
Disaster recovery becomes a re-run instead of a scramble, and the setup is documented by definition.
§5 Containers and orchestration
Docker packages an app with its exact dependencies; Kubernetes runs many copies across machines, restarts them and scales them.
- Images, layers, registries, tags vs digests
- Requests, limits and health probes
- Rolling updates and autoscaling
- Docker, Kubernetes, Helm, EKS, GKE
Containers end “works on my machine”; orchestration keeps systems up and sized to demand without manual servers.
§6 Cloud and networking
Compute, storage and networking rented on demand, plus the plumbing that decides how traffic reaches services safely.
- shopper → DNS → load balancer → app servers → database
- Private networks, firewalls, least-privilege IAM
- Cost, sizing and budgets
- Where personal data physically lives (GDPR)
The cloud is both your biggest infrastructure cost and your biggest source of flexibility.
§7 Configuration and automation
Configuration management keeps every server in a known state; small Bash or Python scripts remove repetitive manual work.
- Config injected at start-up, never baked into images
- Secrets in a secret manager, never in files or Git
- Hunting toil: recurring manual work
- Ansible, Bash, Python, Vault
Every manual step removed is a step that cannot be fumbled at 2am, and a person freed for better work.
§8 Observability and monitoring
Metrics, logs and traces tell you whether production is healthy and, when it is not, why. Alerting tells a human in time.
- The four golden signals: latency, traffic, errors, saturation
- Dashboards that answer one question each
- Alerts that fire on real problems, not noise
- Prometheus, Grafana, OpenSearch, Sentry
The difference between a customer telling you checkout is down and fixing it before anyone noticed.
§9 Security and DevSecOps
Security checks woven into the everyday pipeline instead of a yearly audit: “shift left”, where problems are cheap.
- Dependency and image scanning on every change
- Least-privilege access, rotated secrets
- Supply chain: pinned versions, SBOMs
- Trivy, Dependabot, Snyk, SAST/DAST
For a business with customer accounts and payments under GDPR, a breach is existential.
§10 Reliability engineering (SRE)
Reliability measured and budgeted: an indicator that matters, a target, and an error budget for the gap.
- SLI, e.g. the share of checkouts that succeed
- SLO, e.g. 99.9 % per month; error budget 0.1 %
- On-call rotas, incident response, postmortems
- Backups that have actually been restored
It replaces “ship or be careful?” arguments with a number.
§11 The four numbers it all serves
The DORA research found four measures that separate high-performing teams from the rest.
- Deployment frequency: how often you can ship (want it up)
- Lead time for change: commit to live (want it down)
- Change failure rate: share of deploys needing a fix (down)
- Time to restore: broken to healthy (down)
Shipping more often is safer, not riskier: a three-line change has a three-line blast radius.
§12 Words you will hear in a standup
- CI / CD
- Build and test every change / carry it towards production
- Artifact
- The built, versioned output of a pipeline
- Digest
- A content hash naming one exact image; unlike a tag, it never moves
- Deploy vs release
- Code running on servers vs behaviour visible to users
- Blast radius
- How much breaks when something goes wrong
- Drift
- Real infrastructure no longer matching its code
- SLO / error budget
- The reliability target and the failure it allows
- Toil
- Manual, repetitive work that grows with traffic
- Runbook
- The written steps for a known situation, usable at 3am
The learning route, from zero
Roughly forty weeks at 8 to 10 hours a week. Each stop ends with something you can show, so you always know when to move on.
-
Foundations and the delivery model
Week 1
The delivery loop, the four DORA metrics, what a pipeline changes, build vs deploy vs release, who owns what. Lessons 01–05.
Ready when you can explain to a non-engineer why deploying more often reduces risk.
-
Linux and networking
Weeks 2 to 6
The command line, processes, ps/top/df/ss, systemd services and journalctl, users and permissions, SSH keys, DNS, TCP, TLS, HTTP and ports. Bash, and enough Python to parse a log. Lessons 06–12.
Ready when a fresh VM runs your app as a systemd service that survives a reboot, and you can say why it is slow (Project 1).
-
Git and the collaboration model
Weeks 7 to 8
Commits, branches, pull requests, protected branches, CODEOWNERS, semantic versioning, revert vs reset.
Ready when main is protected and every change arrived through a pull request.
-
CI/CD pipelines
Weeks 9 to 12
GitHub Actions workflows, the test pyramid, build once and promote, secrets in pipelines, rolling / blue-green / canary, feature flags.
Ready when a deliberately broken test blocks a merge and every artifact is named by its commit.
-
Containers
Weeks 13 to 16
Dockerfiles, layer caching, slim and multi-stage builds, registries, tags vs digests, volumes, Compose, non-root containers, image scanning.
Ready when a stranger can run your stack with one command and your pipeline deploys it (Projects 2 and 3).
-
Kubernetes
Weeks 17 to 22
Pods, deployments, services, ingress, ConfigMaps and Secrets, probes, requests and limits, rolling updates, autoscaling, CrashLoopBackOff — and when not to use Kubernetes.
Ready when you can do a zero-downtime rolling update and explain an OOMKill (Project 4).
-
Infrastructure as code
Weeks 23 to 27
Terraform or OpenTofu: providers, resources, state and locking, modules, plan in pull requests; Ansible for what runs on the machines.
Ready when destroy followed by apply gives you the same environment (Project 5).
-
One cloud, properly
Weeks 28 to 32
Compute choices, object storage, managed databases with restorable backups, VPCs and subnets, IAM, billing, budgets and EU data residency.
Ready when you can read your own bill line by line and no key has administrator rights.
-
Observability
Weeks 33 to 36
Logs vs metrics vs traces, the golden signals, Prometheus and PromQL basics, Grafana, alert design, SLOs, OpenTelemetry.
Ready when your alert fires before you would have noticed by hand (Project 6).
-
Reliability, security and on-call
Weeks 37 to 40
Incident roles, blameless postmortems, supply-chain security, secrets rotation, backups with RTO/RPO, sustainable on-call, gentle chaos testing.
Ready with a tested restore, a written postmortem and seven projects on GitHub (Project 7).
Hands-on projects
Build these in order. Each one reuses the same small shop app and the previous project, so by the end you have taken one app from a hand-started process to a monitored, rebuildable system. Never commit a secret, never use real personal data, and set a budget alert before creating anything in the cloud.
Project 1: a server you can rebuild and diagnose
Ubuntu in a local VM, Python, systemd, Bash
- Create a VM and connect with an SSH key, not a password.
- Run a small web app as a dedicated non-root user under a systemd unit with a restart limit.
- Make the journal persistent, reboot, and prove the app came back.
- Kill it with kill -9 and watch the restart counter move.
- Write RUNBOOK.md: the first four commands and what normal looks like on this box.
Show it off with a setup.sh that rebuilds everything on a fresh VM, and the runbook.
Project 2: containerise the shop
Docker, Compose, PostgreSQL, a small Flask or FastAPI app
- Add a /products endpoint backed by Postgres and a /healthz check.
- Write a Dockerfile from a slim base that installs dependencies before copying code.
- Run as a non-root user; keep config in a git-ignored .env with a committed .env.example.
- Compose the app and database with a named volume and a health check.
- Time a rebuild after a one-line change, reorder the Dockerfile, and record the difference.
Show it off with “clone, copy the env file, docker compose up” working in under fifteen minutes.
Project 3: the pipeline — test, build, push, deploy
GitHub Actions, GitHub Container Registry, the Project 1 VM
- Add pytest tests, including a boundary case, and make them a required check.
- On merge, build an image tagged with the commit SHA and push it.
- Scan the image and fail on critical vulnerabilities.
- Deploy to the VM through a protected production environment; the SSH key is an encrypted secret.
- Write rollback.sh for any previous SHA and time it.
Show it off with a screenshot of a blocked merge and a rollback under two minutes.
Project 4: the shop on Kubernetes
kind, kubectl, optionally Helm
- Deploy three replicas with a Service, ConfigMap and Secret.
- Add readiness and liveness probes, requests and limits.
- Roll out a new version while a curl loop counts failed requests.
- Break it five ways — bad tag, crash on start, OOMKill, failing probe, missing secret — and fix each.
- Roll back with kubectl rollout undo.
Show it off with a zero-failure rollout and one line per failure in the README.
Project 5: rebuild from code
Terraform or OpenTofu, one cloud provider, remote state
- Set a budget alert and MFA before creating anything; pick an EU region.
- Remote, locked, encrypted state.
- A VPC, a VM running your image, and a private managed database.
- Plan on every pull request, posted as a comment.
- Destroy and apply again, and time the rebuild; then change something by hand and watch the drift.
Show it off with the rebuild time and a bill that stayed under budget.
Project 6: see it break
Prometheus, Grafana, Alertmanager, a metrics library
- Expose request count, errors and latency from the app.
- Build one dashboard with the four golden signals.
- Write a checkout SLO and two alerts that say what to check first.
- Inject 10 % checkout errors and a filling disk; record when each alert fired.
- Keep personal data out of logs and metric labels.
Show it off with the gap between the alert and the moment you would have noticed.
Project 7: game day and postmortem
Everything above
- Restore a backup into a fresh database and measure RTO and RPO.
- Have someone break one thing without telling you what.
- Respond with a timestamped timeline: mitigate first, diagnose second.
- Write a blameless postmortem with three owned action items.
- Ship one of the action items.
Show it off with the postmortem: it is the most useful thing in an interview.
The lessons
Short lessons, each built around one real-life scene at a small online resale shop: something breaks, someone asks why, and the lesson walks through the mechanism with diagrams, the numbers that matter and a hands-on exercise. Read them in order; later lessons build on earlier ones.
Phase 0: Foundations
Why any of this exists: the delivery model, its vocabulary, and who owns what.
- 01From Laptop to LiveThe pipelineA Friday fix copied over SFTP takes every product page down. The four faults a pipeline removes.
- 02Works on My MachineContainersThumbnails come out sideways and a green CI gate guards the wrong machine. What a container image really is.
- 03The Three EnvironmentsEnvironmentsA staging job emails 9,412 real customers. What dev, staging and production are each for.
- 04Build vs Deploy vs ReleaseDeploy ≠ releaseOne word, three meanings, and a rollback that took a fix with it. Two undo buttons instead of one.
- 05Who Does WhatOwnershipA capacity gap nobody held and an on-call rota of one. Five responsibilities, one name each.
Phase 1: Linux & networking
You cannot debug what you cannot inspect: the Linux box and the network around it.
New lessons are added regularly. Start with the fundamentals, follow the route, build the projects, and read one lesson whenever you finish a stop on the route. Next up: Lesson 08, Permissions and Users.
These lessons were written with the help of AI (Claude) and reviewed and edited by me.