Roadmap To Be A DevOps Engineer / Lesson 03
The Three Environments
A staging job emails 9,412 real customers. What dev, staging and production are each for.
Thread picked up from Lesson 02: the image is portable — config, secrets and data are not.
1The situation
The resale shop now builds one image in CI, so the drift from Lesson 02 is gone. Then the environments — as opposed to the runtime — had their turn.
Tue 10:15 — the rehearsal that rehearsed nothing. A change passes on staging, ships, and breaks checkout in four minutes. Staging runs one container; production runs three behind a load balancer. The new code kept the basket in process memory, so on staging every request found its basket and in production two out of three did not.
Wed 16:40 — the email that was not a test. Someone triggers the abandoned-basket job on staging. Staging holds a two-week-old copy of the production database and the production SMTP credentials, because copying both was the fastest way to make staging realistic. 9,412 real customers get a reminder about a basket they emptied a fortnight ago.
Thu 09:05 — the hotfix that skipped the queue. Staging has been red for nine days (a migration applied there by hand, so its schema exists nowhere else). A one-line VAT fix is needed, staging cannot validate it, and it goes straight to production with a shrug. The gate is now optional, and everyone has learned that it can be.
Thu 14:30 — the config in the image. Last month’s Dockerfile gained
ENV STRIPE_KEY=sk_live_…. It works. It also means the live key is a layer in the registry,
readable by anyone who can docker pull — and testing against test Stripe now needs a
different image.
One root cause: three environments, no agreement about what each is for. Each drifted toward whatever was locally convenient.
2What each environment is for
An environment is not a size of server. It is a question someone needs answered before real customers meet a change.
| Development | Staging | Production | |
|---|---|---|---|
| Question | Does my change do what I think? | Does the release work on a production-shaped system? | Is it working for real people right now? |
| Optimised for | Iteration speed | Fidelity of shape | Reliability & observability |
| Data | Tiny seeded fixture | Production-like volume, never production identities | The real thing, backed up and audited |
| Blast radius | One laptop | Zero outward (test keys, email sink) | Everything |
| Fails when | Setup takes half a day | Treated as a second dev box | Anyone can change it by hand |
Staging’s job is not to test the code — CI already did that, on the same image, faster. Staging tests everything that only happens when you deploy: the migration against a real-sized table, the rolling restart behind a load balancer, the config that must exist, the secret that must be mounted, the third-party call that must be reachable. Those four are exactly what the shop’s bad week was made of.
3The mechanism: one artifact, three configs
Build once, promote the same bytes, inject everything environment-specific at start-up. If a change must be rebuilt to move to the next environment, you are no longer testing the thing you will ship.
4Parity: what must match, what must not
| Property | Identical? | Why |
|---|---|---|
| Image digest | yes | Anything else means staging validated a different program. Free. |
| Topology (replicas, LB, TLS) | yes | Catches the whole "works on one instance" class — the shop’s Tuesday. Cheap in containers. |
| Migration path | yes | Same tool, same command, run by the pipeline. A migration that only ever ran on an empty table is untested. |
| Config keys | yes | A variable present in one environment and not the other is a Friday incident waiting. |
| OS / runtime | yes | Free with the image (Lesson 02). |
| Machine size | no | A quarter of production is fine. Load testing is a separate exercise. |
| Data volume | roughly | Enough rows that a missing index or slow migration shows up. |
| Data content | must differ | No real names, emails, addresses, payment tokens. GDPR, and Wednesday. |
| Outbound integrations | must differ | Test keys, email sink, no partner webhooks. Staging must be unable to reach the outside world even when the code tries. |
5Why the gate pays for itself
The same defect — one missing environment variable — costs wildly different amounts depending on which environment notices it. The shop’s own log for one Thursday:
| Caught | Engineer-minutes | What it consisted of |
|---|---|---|
| On the laptop | 5 | Reload fails, developer adds the variable. |
| In CI | 12 | Red pipeline: 4 read, 3 fix, 5 re-run. Still one person, before review. |
| On staging | 45 | 15 find, 10 fix in two places, 20 redeploy + verify. Two people, release slips an hour. |
| In production | 380 | 25 detect · 40 roll back · 60 repair 31 orders written without a VAT rate · 90 customer emails and two refunds · 75 re-release · 90 post-mortem. |
Illustrative, not an industry statistic — the point is the shape: each boundary a defect crosses multiplies its cost by roughly three to eight, because each one adds people, coordination and damage to undo. That ratio is why cheap gates go first and why nobody argues about a 12-minute CI run.
6Four ways teams break this
| # | The shape | What it does to you |
|---|---|---|
| 01 | Staging as a second dev box | Hand-edited, permanently red, drifted. A gate that exists but is skipped is worse than none — it hides the fact that nothing is checking. |
| 02 | Production data copied into staging | Turns a low-security environment into a personal-data processor with live credentials attached. Ends up in a regulator’s inbox, not a post-mortem. |
| 03 | Config baked into the image | Every environment needs its own build, so you can no longer promote one artifact. Secrets become registry layers. Rolling back config means rolling back code. |
| 04 | No staging at all | Legitimate — only if the deploy mechanism is tested another way: canary or blue/green with automatic rollback, plus feature flags. Without that machinery it is deploying and hoping. |
The one sentence to keep: one artifact, promoted unchanged, configured from the outside — dev proves the code, staging proves the release mechanism on a production-shaped system, production is the only place real data lives.
7Hands-on (20 minutes)
Needs Docker. Each step makes one claim observable.
-
An app that reads its environment (
app.py) — returnAPP_ENV,DATABASE_URL, the first 8 chars ofSTRIPE_KEYandsocket.gethostname(); add a/basket/addendpoint that appends to a module-level dict (deliberately in process memory). Dockerfile as in Lesson 02 — no config, no secrets in it. -
Two environments, one image
docker run --rm -d -p 8001:8000 --env-file dev.env --name dev shop:v1 docker run --rm -d -p 8002:8000 --env-file staging.env --name staging shop:v1 curl -s localhost:8001/ ; curl -s localhost:8002/ docker inspect --format '{{.Image}}' dev staging # same sha256 -
Watch the single-replica illusion break (the shop’s Tuesday, in 30 seconds)
docker run --rm -d -p 8003:8000 --env-file staging.env --name staging2 shop:v1 for p in 8002 8003 8002 8003; do curl -s localhost:$p/basket/add; echo; done # items: 1, 1, 2, 2 — two independent baskets, not one basket of fourFix the design, not the count: move the basket to Redis or the DB, re-run. Any replica can now serve any request — the property that makes rolling deploys possible at all.
-
See why a baked secret is not a secret
printf 'FROM shop:v1\nENV STRIPE_KEY=sk_live_REALKEY123\n' > Dockerfile.bad docker build -f Dockerfile.bad -t shop:leaky . docker history --no-trunc shop:leaky | grep -o 'sk_live_[A-Za-z0-9]*' docker inspect --format '{{json .Config.Env}}' shop:leaky docker rmi shop:leaky -
Fail loudly at start-up — turn the 45-minute row into the 12-minute row:
REQUIRED = ["APP_ENV","DATABASE_URL","STRIPE_KEY"] missing = [k for k in REQUIRED if not os.environ.get(k)] if missing: raise SystemExit(f"refusing to start, missing config: {missing}")Teardown:
docker rm -f dev staging staging2
8Three questions to ask your team this week
- "Is the artifact running in production the exact one staging approved, or was it rebuilt?" The good answer names one digest and shows it in both deploy logs. "Same commit" is a weaker claim — a rebuild can pick up a new base image or a floating dependency.
- "What can staging reach on the outside, and what real data does it hold?" Listen for test payment keys, an email sink, an anonymised restore. If the answer is "a copy of production and the live SMTP key", that is this week’s work.
- "Last time staging was broken, what happened to the release?" Reveals whether the gate is real. If "we shipped anyway", the follow-up is not about discipline but cost: what would make staging boringly reliable, and who owns it?
Next: Lesson 04 — Build vs Deploy vs Release. Why "it’s deployed" and "customers have it" are separate decisions, with separate rollback buttons.
Written with the help of AI (Claude) and reviewed by Rayhanul Islam.