Roadmap To Be A DevOps Engineer / Lesson 07
Where a Service Actually Lives
A 02:40 reboot and 3 h 44 m of downtime. What turns a process into a service.
Lesson 06 was how to look at a box. This one is about the thing that keeps the app alive
when nobody is looking at all — and about the state it quietly ends up in, failed.
1Tuesday, 02:40 — the box reboots
Unattended upgrades install a kernel patch and reboot the shop’s production VM at 02:40. The kernel is back in 38 seconds. Nginx comes back with it. The shop does not. The uptime check emails a shared mailbox nobody reads at 3 a.m. A customer emails at 06:12; Kai fixes it at 06:24. 3 h 44 m of downtime, ended by a four-second command.
Lesson 06’s Tier 1 finds nothing wrong: no python3 in ps, top idle, df at 24 %. And that
is the point — nothing broke. The app was started in March with python3 app.py & inside a
screen session. A process started from a login shell belongs to that login session, and a reboot
ends every login session. Nothing on the box had ever been told to start the shop. For six months
the app ran only because the machine happened not to restart.
What turns a process into a service — three properties & does not give you
| # | Property | The command that provides it |
|---|---|---|
| 1 | Starts at boot, with no human present | systemctl enable |
| 2 | Restarts on failure — with a policy and a limit | Restart= + StartLimitBurst |
| 3 | Logs somewhere findable by name | journalctl -u <svc> |
Several things provide these (supervisord, runit, Docker’s --restart, and later Kubernetes,
which is these three properties spread across many machines). systemd is the one already on the
box, free, and the vocabulary every answer on the internet is written in.
Fig. 1 — the unit lifecycle
inactive → activating → active (running). Then the process exits, and the fork matters:
- exit 0 with
Restart=on-failure→ systemd does nothing. Unit goes inactive and stays there, silently. Trap: an app that exits 0 on a bad config never restarts. UseRestart=alwaysif exit 0 is never correct for your app. - non-zero exit, fatal signal (incl. the OOM killer), or timeout → the restart policy fires →
wait
RestartSec→ count the attempt → under the limit, back toactivating. - over
StartLimitBurstattempts insideStartLimitIntervalSec→failed, and it stays down. No further restarts, ever, until a person runssystemctl reset-failed shop && systemctl start shop.
Four of the five states announce themselves. failed does not — it is a resting state reached
automatically that looks exactly like a unit that was never started. Hence the single most
valuable command here: systemctl list-units --failed = everything on the machine that has
already given up.
systemctl list-units --failed: it is the list of everything on the machine that has
already given up. Run it on your production box today.2The whole thing is a text file (Fig. 2)
# /etc/systemd/system/shop.service [1]
[Unit]
Description=Resale shop web app
After=network-online.target postgresql.service [2]
Wants=network-online.target
StartLimitIntervalSec=60 [3]
StartLimitBurst=5
[Service]
User=shop [4]
WorkingDirectory=/srv/shop
EnvironmentFile=/etc/shop/env [5]
ExecStart=/srv/shop/.venv/bin/python -u /srv/shop/app.py [6]
Restart=on-failure [7]
RestartSec=5
TimeoutStopSec=20
[Install]
WantedBy=multi-user.target [8]
/etc/systemd/systemis yours. Package units live in/usr/lib/systemd/system; a same-named file in/etcwins. Never edit the packaged one.After=is ordering;Wants=is the dependency.After=alone only says "if postgres starts, start after it".network-online, notnetwork— the IP must exist first.StartLimit*belong in[Unit], not[Service]— they moved in systemd 229. Left in[Service]they are ignored with an "Unknown key" warning almost nobody reads.User=shop— the app does not run as root. One line, and the blast radius of a bug shrinks enormously (Lesson 08). Paths must be absolute.EnvironmentFile=— this is where Lesson 03’s "config injected at start-up" physically lands. One image, three environments, this file is the only difference.chmod 600, root-owned, never in git.ExecStart=is not a shell. No pipes, no>, no globs, no$VARexpansion as you expect. Use/bin/sh -cexplicitly if you need them.-ukeeps Python unbuffered or the logs arrive minutes late.Restart=is a policy with a limit — see §3.[Install]is the sectionsystemctl enablereads.enablejust makes a symlink inmulti-user.target.wants/. No[Install]= enable fails = nothing starts at boot.
Two commands everyone forgets: systemctl daemon-reload (makes systemd re-read the file —
skip it and you debug the old version for twenty minutes) and systemctl enable ≠
systemctl start. start runs it now, enable makes it run at boot. The 02:40 outage is exactly
what "started but never enabled" looks like six months later. The only honest test is
systemctl is-enabled shop, followed at some point by a real reboot.
Three habits: systemctl cat shop (the unit systemd is actually using, overrides included);
systemctl edit shop (writes a drop-in at …/shop.service.d/override.conf with only your
changes — survives package upgrades, gives a reviewer a diff); systemd-analyze verify <file>
before you reload.
daemon-reload is what makes systemd re-read the file — edit a unit without
it and you will spend twenty minutes debugging the old version. And
enable is a different verb from start: start runs it now,
enable makes it run at boot. The shop’s 02:40 outage is exactly what “started
but never enabled” looks like six months later — which is why the only honest test of
“it comes back” is systemctl is-enabled shop, followed at some point by an
actual reboot.3Restart= is a policy, and the limit is the interesting part (Fig. 3)
Measured, with a 20-line stand-in for systemd’s restart logic, same instantly-crashing app, one setting changed:
| Setting | What happened | End state |
|---|---|---|
RestartSec=0.1 (the systemd default, 100 ms) |
5 start attempts inside 1.6 s | failed — down permanently, 1.6 s after healthy |
RestartSec=3 |
9 restarts in 26 s; the 10 s window never held more than 4 attempts | never failed — restarts forever, shop down 94 % of the window |
The counter is a sliding window, not a total. Restart quickly and five attempts land in one
window, so systemd concludes the service is broken and stops. Restart slowly and they never share
a window, so systemd never concludes anything. Raising RestartSec makes systemd more
persistent, not less — the opposite of what people intend. The window, not the delay, decides.
Shape three, off that scale — the slow death: a leak that kills the app every six hours
restarts it four times a day, forever. Never failed, never alerted, ~120 small outages a month.
ss, and later the
golden signals — and use systemctl show shop -p NRestarts as the
second signal, never the only one.Restart= values
Restart= |
clean exit (0) | non-zero exit | signal / timeout | systemctl stop |
use for |
|---|---|---|---|---|---|
no (default) |
no | no | no | no | one-shot jobs, migrations, backups — where failure should stay visible |
on-failure |
no | yes | yes | no | the sane default for a web app |
on-abnormal |
no | no | yes | no | when your app exits 3 to mean "my config is wrong, do not retry" |
always |
yes | yes | yes | no | a daemon that must never be down; the only one that catches exit-0 bugs — and hides them best |
systemctl stop never restarts anything, for any policy: a manual stop is a statement of intent.
systemd ≥ 254 adds RestartSteps= and RestartMaxDelaySec= for real exponential backoff.
The rule: a restart policy buys time, never a fix, and it is not free — every automatic restart
drops whatever was in flight. So every restart must leave a trace a human reads:
systemctl show shop -p NRestarts is a number; put it on a dashboard and alert when it moves.
A service that restarted 112 times last month is an incident that has learned to hide.
The unit state is a terrible health signal. Row 1 says failed while the shop is down; row 2
says active most of the time while the shop is equally down. Alert on the port and the endpoint
(Lesson 06’s ss, later the golden signals); use NRestarts as the second signal, never the only one.
4"Where are the logs?" should be a one-word answer
Under a unit, stdout/stderr go to the journal tagged with the unit’s name: journalctl -u shop.
-u shop -b— this unit, this boot.-b -1= the previous boot (how you read what happened before a 02:40 reboot).-p err --since '15 min ago'— turns ten thousand lines into the four that matter.-f— live tail; run it while you restart the service in another window.--disk-usage— the journal is capacity-managed by default:SystemMaxUse=is 10 % of the filesystem, capped at 4 GB, keeping 15 % free. Lesson 06’s 28 GBapp.logcould not happen here; instead a DEBUG flood evicts your own history.journalctl --vacuum-time=14dreclaims on demand.
The trap: journald’s default is Storage=auto = persistent only if /var/log/journal
exists. Where it does not, the journal lives in /run (RAM) and is erased by every reboot —
so on a default box the logs explaining why the machine rebooted are deleted by the reboot.
journalctl --list-boots # one line back = volatile; many = persistent
mkdir -p /var/log/journal
systemd-tmpfiles --create --prefix /var/log/journal
systemctl restart systemd-journald
Also form the habit of journalctl -k — the kernel ring buffer, where the OOM killer writes. An
OOM-killed process looks, in the app’s own logs, exactly like one that stopped mid-sentence.
5Say it so it can be checked
| Sounds like an answer | Can be checked by someone else |
|---|---|
| "the app is running" | "shop.service active (running) since 02:41, NRestarts=0" |
| "it’ll come back if the box reboots" | "systemctl is-enabled shop → enabled; last proved by a reboot on 12 Sep" |
| "it keeps crashing" | "4 restarts in 40 min, each preceded by Errno 111 to postgres" |
| "the logs are gone" | "--list-boots shows 1 boot: journal is volatile, /var/log/journal missing" |
| "nothing else is broken" | "systemctl list-units --failed → 0 loaded units listed" |
Footnote to Lesson 05: the 02:40 outage was not a skills problem — it was an inventory
problem. Nobody owned the question "what runs on this box, and which of it survives a reboot?"
That question has a two-command answer (systemctl list-unit-files --state=enabled,
systemctl list-units --failed) and until somebody’s name is against asking it, the answer drifts.
Phase 6 removes the question entirely by putting the unit file in git next to the app (Lesson 39).
6Hands-on (20 min)
Part A — needs a real Linux host (cheap VM, Raspberry Pi, Multipass, or WSL2 with
systemd=true in /etc/wsl.conf). If systemctl --user status says Failed to connect to bus,
you are in a container without systemd and Part A will not work there.
- A1 — write
/srv/shop/app.py(ahttp.serverone-liner) and the unit above, thensystemd-analyze verify→daemon-reload→enable --now. - A2 — read the three things that matter:
systemctl status shop(the since, not the dot),systemctl is-enabled shop,systemctl show shop -p NRestarts,ss -tlnp | grep 8080. - A3 —
kill -9 $(systemctl show shop -p MainPID --value), watch it come back, watchNRestartsgo to 1. Follow along withjournalctl -u shop -f. - A4 — make it die at start-up (
sed -i '1i import sys; sys.exit(1)'),restart, wait 12 s:Active: failed (Result: exit-code), "Start request repeated too quickly", andstartis refused. Recovery is two commands:reset-failedand thenstart. - A5 — the only honest test:
sudo reboot, thensystemctl is-active shopandsystemctl list-units --failed.
Part B — runs anywhere. Build crashy.sh (an app that always exits 1) and supervise.sh
(20 lines implementing RESTART_SEC, LIMIT_BURST, LIMIT_INTERVAL as a sliding window). Then:
RESTART_SEC=0.1 ./supervise.sh→ 5 attempts, start-limit hit, total elapsed 1.6 s.RESTART_SEC=3 timeout 26 ./supervise.sh→ 9 attempts, the window count plateaus at 4, it never gives up.LIMIT_INTERVAL=30on the same slow run → it gives up after five. The window decides.
B4 — write down the two numbers for your service: how long it takes to start, and how long it
can be down before a customer notices. RestartSec > the first.
StartLimitBurst × (RestartSec + start time) = how long systemd will keep trying — that must be
smaller than the second.
Runbook additions
systemctl list-units --failed · systemctl status <svc> (read the since) ·
systemctl is-enabled <svc> · systemctl show <svc> -p NRestarts · journalctl -u <svc> -b -p err
Policy that would have prevented 02:40: nothing on a production box may be started with &,
nohup, screen or tmux. Every one of those is an outage with a date on it — the next reboot —
and the reboot will not be scheduled by you.
7Three questions to ask your team this week
- "If our production box rebooted right now, what would not come back — and how would we find out?" The honest answer is usually "from a customer". The check is one command; the proof is one scheduled reboot. A team that has never rebooted production on purpose does not know whether it can.
- "How many times did each of our services restart itself last month?" If nobody knows, the
restart policy is not resilience, it is concealment.
NRestartsis free and already counted. - "Is anything on our servers running because a person once typed a command?" Cron jobs, a
screensession, a script from a migration two years ago. Invisible, undocumented, reboot-fatal — and every unit file the audit produces is a line of Phase 6’s IaC already written.
Next — Lesson 08: Permissions and users — why the app should not run as root.
Written with the help of AI (Claude) and reviewed by Rayhanul Islam.