Roadmap To Be A DevOps Engineer / Lesson 04
Build vs Deploy vs Release
One word, three meanings, and a rollback that took a fix with it. Two undo buttons instead of one.
Thread picked up from Lesson 03: the same artifact is now in production — but "in production" and "customers have it" are not the same event.
1The situation
The shop builds one image and promotes it unchanged, configured from the outside. The artifact discipline is fixed. The vocabulary is not.
Mon 11:02 — "it’s deployed." Ana means: merged, CI green, image in the registry. Kai on ops hears running in production. Marketing hears customers can use it and sends the Monday newsletter — 22,140 recipients — pointing at a filter that renders an empty page. One word, three meanings, one hour of apologies.
Wed 20:40 — the rollback that took a fix with it. New checkout copy draws 30 complaints in ten minutes. On-call does the only undo they have: redeploy the previous image. The complaints stop — and so does the payment-timeout fix that travelled in the same artifact. Retries quietly start failing again and 14 orders die overnight before anyone connects the two. The undo was blunt because it acted on the artifact, not the behaviour.
Thu 09:15 — the release nobody performed. A ticket has read done for three weeks. The code is in production behind a flag and nobody ever turned the flag on: the release had no owner, no date and no line in any tool. Three weeks of build cost, zero customer value. In this direction deployed-but-not-released is not an outage — it is waste no dashboard shows.
Fri 17:05 — a hundred percent at once. Image resizing breaks for every visitor for six minutes: 2,410 sessions. The identical defect exposed to 5 % first would have reached 121 — and the undo would have been a value written to a file, not six minutes of waiting for containers.
One root cause, four symptoms: three events wearing one word.
2Three verbs, three owners
Build turns a commit into an immutable artifact — nothing in the world changes, the artifact simply exists. Deploy puts that artifact into a running, reachable state in an environment — an infrastructure act, with restarts, migrations and health checks. Release makes a behaviour visible to some set of users — a product act, expressed as a traffic weight or a flag.
| Build | Deploy | Release | |
|---|---|---|---|
| Produces | an immutable artifact sha256:9f2c1d… |
that artifact running and reachable in one environment | that behaviour visible to N % of users |
| Triggered by | a merged commit | a promotion decision | a product decision |
| Owned by | CI, no human in the loop | the pipeline, watched by on-call | product, with an on-call veto |
| Evidence it happened | a digest in the registry | containers healthy at that digest | flag state or traffic weight |
| Undo action | none — you never roll back a build, you build another | redeploy the previous digest | write the flag back |
| Undo takes | — | ~6 min | ~5 s |
| Undo granularity | — | everything in the artifact | exactly one behaviour |
| If it goes wrong | red pipeline, nobody outside notices | restarts, migrations, capacity — can hurt everyone at once | wrong behaviour, but only for the exposed slice |
3The timeline (Fig. 1 in the artifact)
The deploy was one infrastructure event. The release was three product decisions taken over ninety minutes. Collapse the two and you lose the shaded band — the only window in which a change can be verified in production with nobody at risk.
4Where the fork actually lives (Fig. 2)
A release does not add code to production. Both paths are already there, shipped together in the same artifact; a value that lives outside the artifact decides which one a session gets.
Two consequences: dormant code is not free (a flag parked at 100 % for a month is debt with a name; deleting it is a follow-up deploy), and schema changes do not work this way.
The exception: migrations
A flag can be flipped back; a dropped column cannot. Database changes are deploy-time acts with no cheap undo — hence expand, then contract. Deploy one: add the new column, write to both, read from the old. Release the behaviour. Only in a later deploy, once the release has been at 100 % long enough that you would not go back, remove the old column. Both halves in one deploy is how a rollback becomes data loss.
5What the separation buys (Figs. 3 & 4)
Time to put the old behaviour back — log scale, so each step is a multiple:
| Undo mechanism | Level | Time |
|---|---|---|
| Flip a feature flag | release | 5 s |
| Shift traffic weight (blue/green swap) | release | 40 s |
| Redeploy the previous digest | deploy | 6 min |
| Fix forward through CI | deploy | 28 min |
Two orders of magnitude between the top and bottom halves — the difference between "customers noticed for five seconds" and "customers noticed for six minutes." A team whose only undo is the bottom half will always be slower than the incident.
Blast radius of the same six-minute defect (the shop’s Friday, drawn twice):
| Release decision | Sessions that saw the bug | Undo |
|---|---|---|
| Released to 5 % first | 121 | flag off, 5 s |
| Released to 100 % at once | 2,410 | redeploy, 6 min |
A canary does not make code safer. It makes being wrong twenty times cheaper.
The sentence to keep: build produces an artifact, deploy makes it run, release makes it visible — and the third has its own owner, its own timing and its own undo button.
6A vocabulary contract
| Ambiguous | Says which event happened |
|---|---|
| "it’s deployed" | "built — sha256:9f2c1d is in the registry" |
| "it’s live" | "deployed to prod, released to 0 % of users" |
| "roll it back" | "flag off" — or "redeploy 9f2c1d, which also reverts the payment fix" |
| "we ship on Fridays" | "we deploy daily; we release when product flips the flag" |
| "done" | "released to 100 %, flag removed in the next deploy" |
Footnote on the Lesson 01 metrics: deployment frequency counts deploys, not releases. A team can honestly deploy ten times a day and release once a month. That is not gaming the metric — it is the point of separating them.
7Hands-on (20 minutes) — needs Docker
1. One artifact, both code paths. app.py ships old and new together and reads the flag from a
path that is not in the image:
import hashlib, json, socket
from http.server import BaseHTTPRequestHandler, HTTPServer
FLAGS = "/etc/shop/flags.json"
def percent(name):
try:
with open(FLAGS) as f: return json.load(f).get(name, {}).get("percent", 0)
except FileNotFoundError: return 0 # no flag store = released to nobody
def exposed(name, session):
bucket = int(hashlib.sha256(session.encode()).hexdigest(), 16) % 100
return bucket < percent(name) # same session, same answer, always
class H(BaseHTTPRequestHandler):
def do_GET(self):
session = self.path.strip("/") or "anon"
body = "NEW size filter" if exposed("size_filter", session) else "old size filter"
self.send_response(200); self.end_headers()
self.wfile.write(f"{body} | host={socket.gethostname()}\n".encode())
def log_message(self, *a): pass
HTTPServer(("", 8000), H).serve_forever()
printf 'FROM python:3.12-slim\nCOPY app.py /app.py\nCMD ["python","/app.py"]\n' > Dockerfile
docker build -q -t shop:v2 .
2. Deploy it with the release at zero.
mkdir -p flags && echo '{"size_filter":{"percent":0}}' > flags/flags.json
docker run -d -p 8000:8000 -v "$PWD/flags:/etc/shop:ro" --name shop shop:v2
for s in a b c d e; do curl -s localhost:8000/$s; done # all five say "old size filter"
The new filter is in production. Nobody can see it — the shaded band in Fig. 1, and the state Thursday’s ticket sat in for three weeks.
3. Release, without deploying.
docker inspect --format '{{.State.StartedAt}} {{.Image}}' shop # note both
time (echo '{"size_filter":{"percent":5}}' > flags/flags.json)
for i in $(seq 1 200); do curl -s localhost:8000/s$i; done | grep -c NEW # ~10 of 200
docker inspect --format '{{.State.StartedAt}} {{.Image}}' shop # identical
Same digest, same start time, no restart — and what customers experience just changed.
4. Widen it, then time both undo buttons.
echo '{"size_filter":{"percent":50}}' > flags/flags.json
for i in $(seq 1 200); do curl -s localhost:8000/s$i; done | grep -c NEW # ~100
# undo the RELEASE
time (echo '{"size_filter":{"percent":0}}' > flags/flags.json; curl -s localhost:8000/s1)
# undo the DEPLOY (same mechanics as rolling back to a previous digest)
time (docker rm -f shop >/dev/null && \
docker run -d -p 8000:8000 -v "$PWD/flags:/etc/shop:ro" --name shop shop:v2 >/dev/null && \
until curl -sf localhost:8000/s1 >/dev/null; do sleep 0.2; done)
On a laptop the ratio is roughly 100×; on a real cluster with three replicas, health checks and a load balancer it is closer to the 6 min above.
5. Feel why the deploy undo is blunt. Put two unrelated changes in one artifact, then roll the deploy back because only one is wrong:
sed -i 's#self.end_headers()#self.send_header("X-Payment-Retry","on"); self.end_headers()#' app.py
sed -i 's#host={socket.gethostname()}#host={socket.gethostname()} vat=19.00#' app.py
docker build -q -t shop:v3 .
docker rm -f shop && docker run -d -p 8000:8000 -v "$PWD/flags:/etc/shop:ro" --name shop shop:v3
curl -si localhost:8000/s1 | grep -E 'X-Payment-Retry|vat' # both changes present
# the retry header turns out to be wrong. roll back the deploy:
docker rm -f shop && docker run -d -p 8000:8000 -v "$PWD/flags:/etc/shop:ro" --name shop shop:v2
curl -si localhost:8000/s1 | grep -E 'X-Payment-Retry|vat' # BOTH are gone
docker rm -f shop # teardown
The VAT fix was fine and it went anyway, because it travelled in the same artifact. That is Wednesday — and the reason a flag is worth its plumbing.
8Three questions to ask your team this week
- "When a ticket says done, which of the three actually happened?" A good answer separates them without prompting: image built, running in prod, exposed to N %. If done means merged, nobody in the room knows what customers currently have.
- "For the last user-visible change, how long between deploy and release, and who made the release call?" "Same second, automatically" is legitimate at low volume — but then your only undo is a redeploy and your blast radius is always 100 %. The follow-up is the blast-radius table: what would it take to make the first exposure 5 %?
- "If we had to undo the last change right now, would we flip something or redeploy — and what else comes back with the redeploy?" The second half catches Wednesday. If nobody can list what else is in the artifact, the rollback is a guess.
Next: Lesson 05 — Who Does What: dev, ops, platform and SRE, and where the boundaries actually fall in a small team.
Written with the help of AI (Claude) and reviewed by Rayhanul Islam.