Roadmap To Be A DevOps Engineer / Lesson 02
Works on My Machine
Thumbnails come out sideways and a green CI gate guards the wrong machine. What a container image really is.
Fault addressed: 01 — environment drift (from Lesson 01)
1The situation
The same resale shop. The Friday outage got a post-mortem and a promise to build a pipeline. Then the following week happened.
Mon 09:00 — a new developer starts. She follows the README: install Python, PostgreSQL, the image library, run the migrations. Two version numbers in the README are a year old. She is productive Wednesday afternoon — eleven hours of setup, two of them someone else’s time.
Wed 15:40 — thumbnails come out sideways. Seller photos render upright on every laptop
and rotated 90° in production. The server has libvips 8.12, the laptops have 8.15, and
the two disagree about applying the camera’s EXIF rotation tag. Nobody wrote a bug. The
environments simply differ.
Thu 11:20 — tests pass and production still breaks. The new CI job runs the tests on the runner’s Python 3.11. Production runs 3.9. A dict-ordering assumption is fine in one and wrong in the other. The gate is real, but it is guarding a different machine.
One root cause: the environment is an undeclared dependency. The code is in Git and reviewed; the thing the code runs inside — OS packages, library versions, runtime, env vars — is assembled by hand, differently, on every machine, and written down nowhere but a drifting README.
2Mental model: it is not a small virtual machine
A VM emulates a computer, so each one boots its own kernel. A container does not: it is a normal process on the host, using the host’s one kernel, told it can only see its own slice of filesystem, process list and network.
Kernel shipped 3× (~3.6 GB, ~40 s per boot) versus shipped 0× (~90 MB, ~0.3 s to start). Everything people like about containers and everything they must respect about them follows from that one row: a container is not as strong an isolation boundary as a VM, and it cannot run a different kernel.
3The mechanism: what is actually inside an image
An image is two things in a tarball:
- Read-only layers — a stack of filesystem diffs. For a small Flask app:
debian-slim rootfs74 MB →apt install libvips22 MB →pip install -r requirements.txt38 MB →COPY app/0.4 MB. Total 134 MB. - A manifest (JSON) —
WORKDIR,ENV,USER,EXPOSE,CMD, plus the digestsha256:9f2c…. The digest is a hash of the bytes, so the image on your laptop is byte-for-byte what runs in production.
There is no kernel in it, no boot loader, no init system unless you put one there.
A container is what the runtime makes from that image: the layers mounted read-only, a thin writable layer on top, and the manifest’s command started as one host process. On exit the writable layer is discarded and the image is untouched. That is why anything you must keep — uploads, the database — lives in a volume or a managed service.
sha256:9f2c… on your laptop is byte-for-byte what runs in
production. The container is the running instance — image plus a scratch layer plus a
process — and the scratch layer’s lifetime is the container’s lifetime. Layer sizes are
from a small Flask app; heights are drawn roughly to scale.4The cost model: layers are also the build cache
Each Dockerfile instruction produces one layer, and a cached layer is reused only if that instruction and every instruction before it are unchanged. A changed layer invalidates every layer above it, so line order decides your build time.
| Rebuild after a one-line code change | Time |
|---|---|
Order A — COPY . . first |
41 s |
| Order B — lockfile copied and installed first | 1.6 s |
Same app, four lines reordered, a 26× faster inner loop. (Small Flask app, nine pip dependencies; absolute numbers differ, the ratio usually does not.)
The second cost — the base image. Approximate compressed sizes, pulled on every new node and cold CI runner:
| Base | Size |
|---|---|
python:3.12 |
~1020 MB |
python:3.12-slim |
~155 MB — the sane default |
python:3.12-alpine |
~53 MB |
scratch + static binary |
~12 MB |
Smaller is not automatically better: Alpine uses musl instead of glibc, so many pre-built
Python wheels do not apply and packages compile from source — slower builds, occasional
subtle differences. Start from -slim.
scratch for scale. Smaller is not automatically better:
Alpine uses musl instead of glibc, so many pre-built Python wheels do not apply and
packages compile from source — slower builds, occasional subtle differences.
-slim is the default worth starting from.5Four things a container is not
| # | The belief | What is actually true |
|---|---|---|
| 01 | "It’s a lightweight VM." | A host process with a restricted view. No kernel of its own — a kernel exploit or a bad --privileged reaches the host. Use VMs as the boundary between untrusted tenants. |
| 02 | "The image includes my data." | The image is read-only and identical everywhere; data lives in the writable layer, which dies with the container. A container you are afraid to kill is a design smell. |
| 03 | ":latest means the newest version." |
latest is just a tag someone chose to move. Pin a version tag in dev and a digest in production, or you re-introduce the drift you adopted containers to remove. |
| 04 | "Containerising made us portable." | The runtime is portable. Config, secrets, DNS names and the database are not in the image and still differ per environment — Lesson 03’s subject. |
The one sentence to keep: an image is a versioned, hashed filesystem plus a start command — so the environment stops being something you set up and becomes something you build, review and roll back like code.
6Hands-on (15 minutes)
Needs Docker or Podman. Every step makes one claim from this lesson observable.
-
The runtime is pinned, not yours
python3 --version docker run --rm python:3.9-slim python -V # 3.9.x docker run --rm python:3.12-slim python -V # 3.12.x -
There is no kernel inside — the distro differs,
uname -ris identical:uname -r docker run --rm python:3.12-slim uname -r docker run --rm python:3.12-slim cat /etc/os-release | head -2 -
Build a real image
# app.py from flask import Flask app = Flask(__name__) @app.get("/") def home(): return "shop ok\n" # requirements.txt flask==3.0.3 # Dockerfile FROM python:3.12-slim WORKDIR /srv COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY app.py . ENV FLASK_APP=app.py EXPOSE 8000 CMD ["flask", "run", "--host=0.0.0.0", "--port=8000"] docker build -t shop:v1 . docker run --rm -p 8000:8000 shop:v1 curl localhost:8000 # → shop ok -
Look at the layers — read
SIZEagainst the layer list above:docker history shop:v1 docker images shop:v1 -
Feel the cache (the important step)
touch app.py && time docker build -t shop:v2 . # fast, install CACHED # move `COPY app.py .` above the requirements lines, then: touch app.py && time docker build -t shop:v3 . # slow, pip reinstalls -
Watch the writable layer disappear
docker run --rm shop:v1 sh -c 'echo hello > /srv/note.txt; cat /srv/note.txt' docker run --rm shop:v1 sh -c 'cat /srv/note.txt' # No such file
7Three questions to ask your team this week
- "Is the exact runtime our tests run against the same one production runs?" The answer should be one image digest used by CI and production. "Both are on 3.12" is not the same answer.
- "How long does it take a new developer to get the app running locally?" Half a day of README archaeology is a recurring cost. With a committed Dockerfile or Compose file, the target is one command and under fifteen minutes.
- "Do we deploy tags or digests, and does
:latestappear anywhere in production?" A moving tag quietly reintroduces drift while looking like containerisation. Five-minute audit, large payoff.
Next: Lesson 03 — The Three Environments. What dev, staging and production are each for, why the same image must run in all three, and where the things that legitimately differ (config, secrets, data) belong.
Written with the help of AI (Claude) and reviewed by Rayhanul Islam.