Roadmap To Be A DevOps Engineer / Lesson 05
Who Does What
A capacity gap nobody held and an on-call rota of one. Five responsibilities, one name each.
Thread picked up from Lesson 04: four lessons have named four decisions — build, deploy, release, undo — without once saying who is allowed to make each one.
1The situation
The shop has the machinery now: one image, promoted unchanged, config injected at start-up, a flag store outside the artifact. Ana, Kai and Mira are three developers with one production server. What they do not have is a sentence saying which of them holds what.
Tue 19:00 — the cell nobody held. Marketing’s autumn campaign needs capacity. Ana assumes Kai will scale the containers; Kai assumes it was in the ticket; the ticket says nothing. Traffic lands against two containers and 40 % of requests time out for 22 minutes. Nobody was careless. The cell was empty — nobody owned the runtime, so nobody failed to own it.
Thu 03:14 — the rota of one. The pager wakes Kai for the fourth time in nine days. He wrote none of the four changes. He is on call every night of every month, because the rota has one name in it. Ana and Mira have never seen 03:14 — which is why the third page looks exactly like the first: the person who could prevent it never feels it.
Fri 16:40 — the flag with no veto. The campaign page sits behind a flag at 10 %. Ana, who wants the campaign to land, pushes it to 100 % while error rates are still elevated. Nobody stopped her — not out of agreement, but because the right to say "not while it’s degraded" had no owner. Lesson 04 said release is owned by product with an on-call veto. There was no veto because there was no on-call, only Kai.
Six weeks later — the hire that made it slower. They hire Priya, "the DevOps engineer", and route everything infrastructural through her. Lead time for a routine config change goes from 6 hours to 3.2 working days. Priya is not slow; she is one queue serving three people. In week six, 11 of her 14 tickets are one-line changes the author could have made. The shop did not gain a capability. It gained a handover.
Four incidents, one shape: a responsibility that was never explicitly placed. Empty in the first, concentrated on one head in the second, missing a counterweight in the third, and in the fourth placed where the people who needed it could not reach.
2Five responsibilities, not four job titles
"Dev", "ops", "platform", "SRE" are labels for bundles, and the bundles differ at every company — which is why arguing about them is unproductive. The work does not differ. Every system that reaches customers contains these five, named or not.
| # | Responsibility | The decision it owns | Evidence it is really held | How it fails when vacant |
|---|---|---|---|---|
| 1 | Write the change | What the system will do differently | A merged commit with a name on it | Orphan code: shipped, unowned |
| 2 | Own the pipeline | How a commit becomes an artifact, and what gates it | A green pipeline someone is responsible for repairing | Pipeline rot — all wait, none fix |
| 3 | Own the runtime | Servers, cluster, network, capacity the artifact lands on | A headroom number, and an upgrade that happened | Tuesday 19:00 |
| 4 | Decide exposure | Who sees the new behaviour, and when (L04’s release) | A flag value with a date and a name | The 3-week "done" ticket, or Friday’s 100 % |
| 5 | Carry the pager | Answering at 03:14 — and authority to act without asking | A rota with a name per night, and a runbook | The rota of one |
One accountable name per responsibility, per service. Many contribute; exactly one is accountable.
The three rules
- R1 — Every responsibility has exactly one accountable name. Merging is fine. Vacancy is Tuesday.
- R2 — 1 and 5 live in the same rota. Whoever writes the change must be in the rota that carries its pager — not the same night, the same rota. Break this and the feedback loop that makes software reliable is severed: the author never learns what 03:14 costs.
- R3 — 4 and 5 may share a team, but not a single head under pressure. Exposure needs someone who wants the change to land; the pager needs someone who can refuse. Put both in one person at 16:40 on a Friday and the refusal never happens.
3The grid at three, twelve and forty people (Fig. 1)
| Three people | Twelve people | Forty people | |
|---|---|---|---|
| 1 · Write the change | Ana · Mira · Kai | Checkout team (6) + Catalog (4) | 6 stream teams, ~5 each |
| 2 · Own the pipeline | Kai (evenings only) | Platform eng (1) | Platform team (4) — paved road |
| 3 · Own the runtime | Kai (same head, same hours) | Platform eng (1) | Platform team (4) — opt-out allowed |
| 4 · Decide exposure | Ana — no veto exists | PM per stream, on-call veto | PM per stream, on-call veto |
| 5 · Carry the pager | Kai — alone, 30 nights/mo | Rota of 6 — 5.1 nights/mo | Rota of 8 per team — 3.8 nights/mo |
Moving right, only two things really happen: the pipeline and runtime cells detach from a person and become a product, and the pager cell widens from one name to a rota. Rows 2 + 3 merge safely at every size. Rows 4 + 5 must never collapse into one head.
At three people the grid cannot be satisfied — there are not enough humans. You cannot fix that column by reorganising; you can only name the merges and add a counterweight to each:
- Kai holds pipeline + runtime alone → write the runbook before you need it, and have Ana perform one deploy a month from it.
- Kai is the whole rota → put three names on a calendar anyway, even if two can only escalate.
- Ana decides exposure with no veto → agree one written rule: no exposure increase while error rate is elevated, and none above 10 % after 16:00 on a Friday. A rule on paper is a veto that works while everyone is asleep.
4The two shapes that go wrong (Fig. 2)
The wall is the old failure: ops carries consequences it cannot prevent. The queue is the modern one, and it is sneakier — it looks like specialisation and it hires well. But a team that owns the pipeline and the right to refuse becomes a gate, and every stream team’s lead time becomes its queue depth. The problem is not that a platform team exists; it is that the road runs through them instead of being built by them.
Platform as a product, not a gate: the platform team’s customers are the developers, and its success metric is adoption, not tickets closed. A paved road — pipeline template, standard runtime, a way to get a secret — that a stream team uses themselves, with an opt-out for the team that genuinely needs something else. If a stream team cannot deploy, scale or get a secret without a human in another team saying yes, you have a gate however you have labelled it.
Two heuristics worth stealing: "you build it, you run it" (Vogels, Amazon, 2006) is R2 in four words; Team Topologies (Skelton & Pais) gives the vocabulary for the right-hand column — stream-aligned teams owning a slice end to end, a platform team serving them, enabling teams that teach rather than do.
5What the boundaries cost (Figs. 3 & 4)
One config change, two org shapes:
| Hands on the work | Waiting in a queue | Lead time | Flow efficiency | |
|---|---|---|---|---|
| One cross-functional team | 4 h 10 m | 2 h 05 m | 6 h 15 m | 67 % |
| Three teams, three queues | 4 h 55 m (+45 m re-acquiring context ×3) | 21 h 00 m | 25 h 55 m | 19 % |
The work grew by 18 %. The waiting grew by 10×. This is why "make the handover smoother" never pays and removing the handover always does: you are not optimising the work, you are deleting the wait. Measure it in your own tracker as touch time ÷ lead time; under 25 % means your org chart, not your engineers, is the constraint.
What rota size actually changes — the service generates 73 out-of-hours pages a year, and that number is in every row:
| Rota size | Nights on call / person / month | Weeks on call / year | Pages / person / year |
|---|---|---|---|
| 1 | 30.4 | 52.0 | 73.0 |
| 3 | 10.1 | 17.3 | 24.3 |
| 6 | 5.1 | 8.7 | 12.2 |
| 8 | 3.8 | 6.5 | 9.1 |
Rota size divides the pain; it never reduces it. Below six people a rota cannot give anyone a clear week — the practical floor, and the real reason a 12-person org can hold the grid honestly while a 3-person one cannot. And note the arithmetic: halving the page rate on a rota of 3 gives each person 12.2 — identical to doubling the team. One costs an afternoon reviewing alerts; the other costs three salaries. (Lesson 56.)
The sentence to keep: five responsibilities — write, pipeline, runtime, exposure, pager — each with exactly one accountable name; author and pager in the same rota; exposure and pager never in the same head.
6Say it so it can be checked
| Sounds organised | Names a holder |
|---|---|
| "Kai handles infrastructure" | "Kai is accountable for pipeline and runtime; Ana can deploy from the runbook" |
| "we’re all on call" | "the rota is Kai / Ana / Mira, one week each, escalation to Kai" |
| "product decides when it goes out" | "Ana sets the flag; on-call may veto any increase while error rate is elevated" |
| "we hired a DevOps engineer" | "Priya owns the pipeline as a product; teams deploy themselves and she measures adoption" |
| "that’s a platform ticket" | "that is self-service; if it is not, that is a platform bug and it has a number" |
Footnote to Lesson 01’s metrics: lead time for change is an org-chart measurement at least as much as a tooling one. Fig. 3 moves it 4× without touching a pipeline step. When lead time will not come down, look for the queue before you look for the slow test.
7Hands-on (20 minutes) — Python 3 and git, no Docker
An ownership model you cannot run is a diagram, and a diagram drifts. Run it and it is a test.
1. Write the shop down as data. Individuals go in people; anything else is a team.
mkdir -p ~/ownership && cd ~/ownership
cat > ownership.json <<'JSON'
{
"service": "shop-checkout",
"org": "resale shop, 3 people",
"people": ["ana", "mira", "kai"],
"responsibilities": {
"write_the_change": "ana",
"own_the_pipeline": "kai",
"own_the_runtime": "kai",
"decide_exposure": "ana",
"carry_the_pager": "kai"
},
"authors": ["ana", "mira", "kai"],
"rota": ["kai"]
}
JSON
2. Write the three rules as a checker.
cat > owners.py <<'PY'
#!/usr/bin/env python3
import json, sys
RESP = ["write_the_change", "own_the_pipeline", "own_the_runtime",
"decide_exposure", "carry_the_pager"]
org = json.load(open("ownership.json"))
holders = org["responsibilities"]
people = set(org.get("people", []))
authors = set(org.get("authors", []))
rota = list(dict.fromkeys(org.get("rota", [])))
problems = []
# R1 - every responsibility has exactly one accountable name
for r in RESP:
if not holders.get(r):
problems.append(f"VACANT {r}: nobody is accountable")
# R2 - authors must sit in the rota that carries their pager
outside = sorted(authors - set(rota))
if outside:
problems.append("WALL writes but never paged: " + ", ".join(outside))
# R3 - exposure and the pager must not be one head
exp, pag = holders.get("decide_exposure"), holders.get("carry_the_pager")
if exp and exp == pag and exp in people:
problems.append(f"NO VETO {exp} both decides exposure and carries the pager")
# bus factor on the infrastructure cells
infra = {holders.get("own_the_pipeline"), holders.get("own_the_runtime")} - {None}
if len(infra) == 1 and next(iter(infra)) in people:
problems.append(f"BUS=1 pipeline and runtime both held by {next(iter(infra))}")
# rota arithmetic - Fig. 4
n = len(rota)
if n == 0:
problems.append("NO ROTA there is no rota, only a habit")
elif n < 6:
problems.append(f"THIN ROTA {n} on the rota = {30.4/n:.1f} nights on call each, per month")
print(f"\n{org['service']} -- {org['org']}\n")
for r in RESP:
print(f" {r:<18} {holders.get(r) or '-- NOBODY --'}")
print(f" {'rota':<18} {', '.join(rota) or '-- NOBODY --'}"
f" ({30.4/n:.1f} nights/month each)" if n else "")
print()
for p in problems:
print(" " + p)
print(f"\n {len(problems)} finding(s)\n")
sys.exit(1 if problems else 0)
PY
python3 owners.py
Three findings, each a dated incident above. WALL is Thursday 03:14 (Ana and Mira write, Kai is
paged). BUS=1 is Tuesday 19:00 (one head holds the runtime, and that head was busy). THIN ROTA
is the rota of one. NO VETO stays quiet only because Ana and Kai happen to be different people.
3. Provoke R3 on purpose — the tidy-looking move a tired team makes when the product person goes on holiday:
python3 - <<'PY'
import json
d = json.load(open("ownership.json"))
d["responsibilities"]["decide_exposure"] = "kai"
json.dump(d, open("ownership.json","w"), indent=2)
PY
python3 owners.py | grep 'NO VETO'
One person now decides both whether the change is exposed and whether they get woken by it. At 16:40 on a Friday those interests point in opposite directions. Put it back.
4. Grow to twelve and re-run. Nothing in the checker changes — only the names, and a rota of six.
cat > ownership.json <<'JSON'
{
"service": "shop-checkout",
"org": "resale shop, 12 people",
"people": ["ana", "mira", "raj", "lena", "tom", "sofia", "priya"],
"responsibilities": {
"write_the_change": "checkout-team",
"own_the_pipeline": "platform",
"own_the_runtime": "platform",
"decide_exposure": "checkout-pm",
"carry_the_pager": "checkout-team"
},
"authors": ["ana", "mira", "raj", "lena", "tom", "sofia"],
"rota": ["ana", "mira", "raj", "lena", "tom", "sofia"]
}
JSON
python3 owners.py ; echo "exit=$?"
Zero findings, exit 0 — which means this file belongs in CI. An ownership model that fails the build when someone leaves is the only kind that stays true.
5. Price the rota, then price the alternative.
cat > rota.py <<'PY'
PAGES_PER_YEAR = 73 # out-of-hours pages the service generates
print(f"{'rota':>5}{'nights/mo':>12}{'weeks/yr':>10}{'pages/person/yr':>18}")
for n in (1, 3, 6, 8, 12):
print(f"{n:>5}{30.4/n:>12.1f}{52/n:>10.1f}{PAGES_PER_YEAR/n:>18.1f}")
PY
python3 rota.py
# now fix the three noisiest alerts instead of hiring:
sed -i 's/PAGES_PER_YEAR = 73/PAGES_PER_YEAR = 36/' rota.py && python3 rota.py
A rota of 3 at 36 pages a year gives each person 12.0 — the same relief as a rota of 6 at 73.
6. Make the boundary enforceable in git. CODEOWNERS is the grid written where a tool acts on
it. Note the last line: the flag directory — responsibility 4 — requires both product and on-call.
That is R3’s veto, encoded.
git init -q shop && cd shop
mkdir -p app .github/workflows infra flags
cat > CODEOWNERS <<'TXT'
/app/ @checkout-team
/.github/workflows/ @platform
/infra/ @platform
/flags/ @checkout-pm @oncall
TXT
printf 'x\n' > app/cart.py; printf 'y\n' > flags/flags.json; printf 'z\n' > infra/main.tf
git add -A && git -c user.email=a@b -c user.name=a commit -qm init
printf 'x2\n' > app/cart.py; printf 'y2\n' > flags/flags.json; printf 'w\n' > scripts.sh
git add -A && git -c user.email=a@b -c user.name=a commit -qm change
python3 - <<'PY'
import subprocess
rules = []
for line in open("CODEOWNERS"):
line = line.split("#")[0].split()
if line: rules.append((line[0], line[1:]))
changed = subprocess.run(["git","diff","--name-only","HEAD~1"],
capture_output=True, text=True).stdout.split()
for f in changed:
hit = [o for p, o in rules if ("/" + f).startswith(p)]
print(f"{f:<24} {' '.join(hit[-1]) if hit else '** UNOWNED SURFACE **'}")
PY
scripts.sh comes back unowned — a file at the repo root no rule covers. That is R1 catching a
vacancy at the only moment it is cheap to fix.
8Three questions to ask your team this week
- "For our busiest service, name the one person or team accountable for each of the five: write, pipeline, runtime, exposure, pager." Time the answers — any cell that takes more than five seconds is the cell that will fail next. Then ask which cells hold the same name: a merge you chose is a decision; a merge you discover is an incident waiting. Rows 2 and 3 sharing a name is normal; rows 4 and 5 sharing a head is not.
- "Who was woken last month, how many times, and did any of them write the change that woke them?" The second half is R2. A rota in which the authors never appear predicts that next month’s pages look like last month’s. If the answer is "one person, every time", that is Thursday 03:14 — and the cheapest fix is the noisiest alert, not a hire.
- "When a team needs a new environment, a secret or a scaling change — do they do it themselves or file a ticket? How long does the ticket sit before anyone touches it?" That waiting number is Fig. 3’s queue, usually 3–10× the work itself. If the answer is "ticket", the follow-up is not "how do we close tickets faster" but "which three of these should be self-service by next quarter, and who measures whether teams actually use it?"
Next: Phase 1 begins — Lesson 06, The Nine Commands: ps, top, df, du, tail, grep, systemctl, journalctl, ss.
Written with the help of AI (Claude) and reviewed by Rayhanul Islam.