Roadmap to be a DevOps engineer

A field guide from zero: the fundamentals, a step-by-step learning route, hands-on projects to build, and a series of short lessons where each idea is explained through something that actually goes wrong at a real online shop.

The fundamentals

DevOps is the practice of getting an idea from a developer’s laptop into customers’ hands safely and repeatedly, and then keeping it healthy. It merges building software and running it, and automates the road between them. These are the ideas a DevOps engineer works with every day.

§1 Culture and shared ownership

The people who build software also share responsibility for running it: “you build it, you run it”. Outages are blamed on processes, not people.

  • Blameless post-incident reviews
  • Runbooks, so knowledge is not trapped in one head
  • One accountable name per responsibility
  • Developers in the on-call rota

You can buy every tool below and still fail if teams do not cooperate. Culture is the part a CTO shapes most directly.

§2 Version control (Git)

Git records every change — who, what, when and why — and is the single source of truth that the whole loop triggers from.

  • Branches and pull requests: review built into the workflow
  • Branch protection and required checks
  • Tags and semantic versions
  • Configuration and infrastructure live here too

No version control, no DevOps: builds, tests and deploys all start from a change landing in Git.

§3 CI/CD, the automated pipeline

Continuous integration builds and tests every push within minutes; continuous delivery carries what passed towards production.

  • push → build → test → scan → stage → production
  • A failing step stops the change: the automation is the safety net
  • Build once, promote the same artifact
  • GitHub Actions, GitLab CI, Jenkins, Argo CD

The single biggest lever on the four DORA numbers: from monthly releases with held breath to many quiet deploys a day.

§4 Infrastructure as code

Servers, networks and databases are described in text files, reviewed in pull requests, and a tool makes reality match them.

  • Plan before apply; state tracks what exists
  • The same definition for staging and production
  • Drift: what someone clicked versus what the code says
  • Terraform, OpenTofu, CloudFormation, Pulumi

Disaster recovery becomes a re-run instead of a scramble, and the setup is documented by definition.

§5 Containers and orchestration

Docker packages an app with its exact dependencies; Kubernetes runs many copies across machines, restarts them and scales them.

  • Images, layers, registries, tags vs digests
  • Requests, limits and health probes
  • Rolling updates and autoscaling
  • Docker, Kubernetes, Helm, EKS, GKE

Containers end “works on my machine”; orchestration keeps systems up and sized to demand without manual servers.

§6 Cloud and networking

Compute, storage and networking rented on demand, plus the plumbing that decides how traffic reaches services safely.

  • shopper → DNS → load balancer → app servers → database
  • Private networks, firewalls, least-privilege IAM
  • Cost, sizing and budgets
  • Where personal data physically lives (GDPR)

The cloud is both your biggest infrastructure cost and your biggest source of flexibility.

§7 Configuration and automation

Configuration management keeps every server in a known state; small Bash or Python scripts remove repetitive manual work.

  • Config injected at start-up, never baked into images
  • Secrets in a secret manager, never in files or Git
  • Hunting toil: recurring manual work
  • Ansible, Bash, Python, Vault

Every manual step removed is a step that cannot be fumbled at 2am, and a person freed for better work.

§8 Observability and monitoring

Metrics, logs and traces tell you whether production is healthy and, when it is not, why. Alerting tells a human in time.

  • The four golden signals: latency, traffic, errors, saturation
  • Dashboards that answer one question each
  • Alerts that fire on real problems, not noise
  • Prometheus, Grafana, OpenSearch, Sentry

The difference between a customer telling you checkout is down and fixing it before anyone noticed.

§9 Security and DevSecOps

Security checks woven into the everyday pipeline instead of a yearly audit: “shift left”, where problems are cheap.

  • Dependency and image scanning on every change
  • Least-privilege access, rotated secrets
  • Supply chain: pinned versions, SBOMs
  • Trivy, Dependabot, Snyk, SAST/DAST

For a business with customer accounts and payments under GDPR, a breach is existential.

§10 Reliability engineering (SRE)

Reliability measured and budgeted: an indicator that matters, a target, and an error budget for the gap.

  • SLI, e.g. the share of checkouts that succeed
  • SLO, e.g. 99.9 % per month; error budget 0.1 %
  • On-call rotas, incident response, postmortems
  • Backups that have actually been restored

It replaces “ship or be careful?” arguments with a number.

§11 The four numbers it all serves

The DORA research found four measures that separate high-performing teams from the rest.

  • Deployment frequency: how often you can ship (want it up)
  • Lead time for change: commit to live (want it down)
  • Change failure rate: share of deploys needing a fix (down)
  • Time to restore: broken to healthy (down)

Shipping more often is safer, not riskier: a three-line change has a three-line blast radius.

§12 Words you will hear in a standup

CI / CD
Build and test every change / carry it towards production
Artifact
The built, versioned output of a pipeline
Digest
A content hash naming one exact image; unlike a tag, it never moves
Deploy vs release
Code running on servers vs behaviour visible to users
Blast radius
How much breaks when something goes wrong
Drift
Real infrastructure no longer matching its code
SLO / error budget
The reliability target and the failure it allows
Toil
Manual, repetitive work that grows with traffic
Runbook
The written steps for a known situation, usable at 3am

The learning route, from zero

Roughly forty weeks at 8 to 10 hours a week. Each stop ends with something you can show, so you always know when to move on.

  1. Foundations and the delivery model

    Week 1

    The delivery loop, the four DORA metrics, what a pipeline changes, build vs deploy vs release, who owns what. Lessons 01–05.

    Ready when you can explain to a non-engineer why deploying more often reduces risk.

  2. Linux and networking

    Weeks 2 to 6

    The command line, processes, ps/top/df/ss, systemd services and journalctl, users and permissions, SSH keys, DNS, TCP, TLS, HTTP and ports. Bash, and enough Python to parse a log. Lessons 06–12.

    Ready when a fresh VM runs your app as a systemd service that survives a reboot, and you can say why it is slow (Project 1).

  3. Git and the collaboration model

    Weeks 7 to 8

    Commits, branches, pull requests, protected branches, CODEOWNERS, semantic versioning, revert vs reset.

    Ready when main is protected and every change arrived through a pull request.

  4. CI/CD pipelines

    Weeks 9 to 12

    GitHub Actions workflows, the test pyramid, build once and promote, secrets in pipelines, rolling / blue-green / canary, feature flags.

    Ready when a deliberately broken test blocks a merge and every artifact is named by its commit.

  5. Containers

    Weeks 13 to 16

    Dockerfiles, layer caching, slim and multi-stage builds, registries, tags vs digests, volumes, Compose, non-root containers, image scanning.

    Ready when a stranger can run your stack with one command and your pipeline deploys it (Projects 2 and 3).

  6. Kubernetes

    Weeks 17 to 22

    Pods, deployments, services, ingress, ConfigMaps and Secrets, probes, requests and limits, rolling updates, autoscaling, CrashLoopBackOff — and when not to use Kubernetes.

    Ready when you can do a zero-downtime rolling update and explain an OOMKill (Project 4).

  7. Infrastructure as code

    Weeks 23 to 27

    Terraform or OpenTofu: providers, resources, state and locking, modules, plan in pull requests; Ansible for what runs on the machines.

    Ready when destroy followed by apply gives you the same environment (Project 5).

  8. One cloud, properly

    Weeks 28 to 32

    Compute choices, object storage, managed databases with restorable backups, VPCs and subnets, IAM, billing, budgets and EU data residency.

    Ready when you can read your own bill line by line and no key has administrator rights.

  9. Observability

    Weeks 33 to 36

    Logs vs metrics vs traces, the golden signals, Prometheus and PromQL basics, Grafana, alert design, SLOs, OpenTelemetry.

    Ready when your alert fires before you would have noticed by hand (Project 6).

  10. Reliability, security and on-call

    Weeks 37 to 40

    Incident roles, blameless postmortems, supply-chain security, secrets rotation, backups with RTO/RPO, sustainable on-call, gentle chaos testing.

    Ready with a tested restore, a written postmortem and seven projects on GitHub (Project 7).

Hands-on projects

Build these in order. Each one reuses the same small shop app and the previous project, so by the end you have taken one app from a hand-started process to a monitored, rebuildable system. Never commit a secret, never use real personal data, and set a budget alert before creating anything in the cloud.

Project 1: a server you can rebuild and diagnose

Ubuntu in a local VM, Python, systemd, Bash

  1. Create a VM and connect with an SSH key, not a password.
  2. Run a small web app as a dedicated non-root user under a systemd unit with a restart limit.
  3. Make the journal persistent, reboot, and prove the app came back.
  4. Kill it with kill -9 and watch the restart counter move.
  5. Write RUNBOOK.md: the first four commands and what normal looks like on this box.

Show it off with a setup.sh that rebuilds everything on a fresh VM, and the runbook.

Project 2: containerise the shop

Docker, Compose, PostgreSQL, a small Flask or FastAPI app

  1. Add a /products endpoint backed by Postgres and a /healthz check.
  2. Write a Dockerfile from a slim base that installs dependencies before copying code.
  3. Run as a non-root user; keep config in a git-ignored .env with a committed .env.example.
  4. Compose the app and database with a named volume and a health check.
  5. Time a rebuild after a one-line change, reorder the Dockerfile, and record the difference.

Show it off with “clone, copy the env file, docker compose up” working in under fifteen minutes.

Project 3: the pipeline — test, build, push, deploy

GitHub Actions, GitHub Container Registry, the Project 1 VM

  1. Add pytest tests, including a boundary case, and make them a required check.
  2. On merge, build an image tagged with the commit SHA and push it.
  3. Scan the image and fail on critical vulnerabilities.
  4. Deploy to the VM through a protected production environment; the SSH key is an encrypted secret.
  5. Write rollback.sh for any previous SHA and time it.

Show it off with a screenshot of a blocked merge and a rollback under two minutes.

Project 4: the shop on Kubernetes

kind, kubectl, optionally Helm

  1. Deploy three replicas with a Service, ConfigMap and Secret.
  2. Add readiness and liveness probes, requests and limits.
  3. Roll out a new version while a curl loop counts failed requests.
  4. Break it five ways — bad tag, crash on start, OOMKill, failing probe, missing secret — and fix each.
  5. Roll back with kubectl rollout undo.

Show it off with a zero-failure rollout and one line per failure in the README.

Project 5: rebuild from code

Terraform or OpenTofu, one cloud provider, remote state

  1. Set a budget alert and MFA before creating anything; pick an EU region.
  2. Remote, locked, encrypted state.
  3. A VPC, a VM running your image, and a private managed database.
  4. Plan on every pull request, posted as a comment.
  5. Destroy and apply again, and time the rebuild; then change something by hand and watch the drift.

Show it off with the rebuild time and a bill that stayed under budget.

Project 6: see it break

Prometheus, Grafana, Alertmanager, a metrics library

  1. Expose request count, errors and latency from the app.
  2. Build one dashboard with the four golden signals.
  3. Write a checkout SLO and two alerts that say what to check first.
  4. Inject 10 % checkout errors and a filling disk; record when each alert fired.
  5. Keep personal data out of logs and metric labels.

Show it off with the gap between the alert and the moment you would have noticed.

Project 7: game day and postmortem

Everything above

  1. Restore a backup into a fresh database and measure RTO and RPO.
  2. Have someone break one thing without telling you what.
  3. Respond with a timestamped timeline: mitigate first, diagnose second.
  4. Write a blameless postmortem with three owned action items.
  5. Ship one of the action items.

Show it off with the postmortem: it is the most useful thing in an interview.

New lessons are added regularly. Start with the fundamentals, follow the route, build the projects, and read one lesson whenever you finish a stop on the route. Next up: Lesson 08, Permissions and Users.

These lessons were written with the help of AI (Claude) and reviewed and edited by me.

Back to top