Roadmap To Be A DevOps Engineer / Lesson 07

Lesson 07 systemd Fundamentals §7 Automation / §10 Reliability About 10 min read

Where a Service Actually Lives

A 02:40 reboot and 3 h 44 m of downtime. What turns a process into a service.

Lesson 06 was how to look at a box. This one is about the thing that keeps the app alive when nobody is looking at all — and about the state it quietly ends up in, failed.

1Tuesday, 02:40 — the box reboots

Unattended upgrades install a kernel patch and reboot the shop’s production VM at 02:40. The kernel is back in 38 seconds. Nginx comes back with it. The shop does not. The uptime check emails a shared mailbox nobody reads at 3 a.m. A customer emails at 06:12; Kai fixes it at 06:24. 3 h 44 m of downtime, ended by a four-second command.

Lesson 06’s Tier 1 finds nothing wrong: no python3 in ps, top idle, df at 24 %. And that is the point — nothing broke. The app was started in March with python3 app.py & inside a screen session. A process started from a login shell belongs to that login session, and a reboot ends every login session. Nothing on the box had ever been told to start the shop. For six months the app ran only because the machine happened not to restart.

What turns a process into a service — three properties & does not give you

# Property The command that provides it
1 Starts at boot, with no human present systemctl enable
2 Restarts on failure — with a policy and a limit Restart= + StartLimitBurst
3 Logs somewhere findable by name journalctl -u <svc>

Several things provide these (supervisord, runit, Docker’s --restart, and later Kubernetes, which is these three properties spread across many machines). systemd is the one already on the box, free, and the vocabulary every answer on the internet is written in.

Fig. 1 — the unit lifecycle

inactive → activating → active (running). Then the process exits, and the fork matters:

  • exit 0 with Restart=on-failure → systemd does nothing. Unit goes inactive and stays there, silently. Trap: an app that exits 0 on a bad config never restarts. Use Restart=always if exit 0 is never correct for your app.
  • non-zero exit, fatal signal (incl. the OOM killer), or timeout → the restart policy fires → wait RestartSec → count the attempt → under the limit, back to activating.
  • over StartLimitBurst attempts inside StartLimitIntervalSec → failed, and it stays down. No further restarts, ever, until a person runs systemctl reset-failed shop && systemctl start shop.

Four of the five states announce themselves. failed does not — it is a resting state reached automatically that looks exactly like a unit that was never started. Hence the single most valuable command here: systemctl list-units --failed = everything on the machine that has already given up.

One service, five states. Four of them are loud. under the limit → start it again Inactive (dead) nothing is running Activating runs ExecStart= Active (running) the only state you want Deactivating systemctl stop only boot, or systemctl start The process exits exit status 0 non-zero, signal, timeout A clean exit is not a failure With Restart=on-failure, systemd does nothing at all: the unit goes inactive and stays there, silently. The restart policy fires Restart=on-failure covers a non-zero exit, any fatal signal (SIGSEGV, SIGKILL, the OOM killer) and timeouts. Trap: an app that exits 0 on a bad config never restarts under on-failure. Use Restart=always if exit 0 is never correct. Wait RestartSec, then count More than StartLimitBurst (default 5) start attempts inside StartLimitIntervalSec (10 s)? yes — give up Failed — and it stays down No further restarts, ever. The only way out is a person typing: systemctl reset-failed shop && systemctl start shop This is the state nobody knows exists.
The lifecycle of one unit — and the exit that stays down. Four of the five states announce themselves. failed does not — it is a resting state, reached automatically, that no amount of waiting will leave. A unit in failed looks exactly like a unit that was never started, which is why the single most valuable command in this lesson is systemctl list-units --failed: it is the list of everything on the machine that has already given up. Run it on your production box today.

2The whole thing is a text file (Fig. 2)

# /etc/systemd/system/shop.service          [1]
[Unit]
Description=Resale shop web app
After=network-online.target postgresql.service   [2]
Wants=network-online.target
StartLimitIntervalSec=60                    [3]
StartLimitBurst=5

[Service]
User=shop                                   [4]
WorkingDirectory=/srv/shop
EnvironmentFile=/etc/shop/env               [5]
ExecStart=/srv/shop/.venv/bin/python -u /srv/shop/app.py   [6]
Restart=on-failure                          [7]
RestartSec=5
TimeoutStopSec=20

[Install]
WantedBy=multi-user.target                  [8]
  1. /etc/systemd/system is yours. Package units live in /usr/lib/systemd/system; a same-named file in /etc wins. Never edit the packaged one.
  2. After= is ordering; Wants= is the dependency. After= alone only says "if postgres starts, start after it". network-online, not network — the IP must exist first.
  3. StartLimit* belong in [Unit], not [Service] — they moved in systemd 229. Left in [Service] they are ignored with an "Unknown key" warning almost nobody reads.
  4. User=shop — the app does not run as root. One line, and the blast radius of a bug shrinks enormously (Lesson 08). Paths must be absolute.
  5. EnvironmentFile= — this is where Lesson 03’s "config injected at start-up" physically lands. One image, three environments, this file is the only difference. chmod 600, root-owned, never in git.
  6. ExecStart= is not a shell. No pipes, no >, no globs, no $VAR expansion as you expect. Use /bin/sh -c explicitly if you need them. -u keeps Python unbuffered or the logs arrive minutes late.
  7. Restart= is a policy with a limit — see §3.
  8. [Install] is the section systemctl enable reads. enable just makes a symlink in multi-user.target.wants/. No [Install] = enable fails = nothing starts at boot.

Two commands everyone forgets: systemctl daemon-reload (makes systemd re-read the file — skip it and you debug the old version for twenty minutes) and systemctl enable ≠ systemctl start. start runs it now, enable makes it run at boot. The 02:40 outage is exactly what "started but never enabled" looks like six months later. The only honest test is systemctl is-enabled shop, followed at some point by a real reboot.

Three habits: systemctl cat shop (the unit systemd is actually using, overrides included); systemctl edit shop (writes a drop-in at …/shop.service.d/override.conf with only your changes — survives package upgrades, gives a reviewer a diff); systemd-analyze verify <file> before you reload.

$ sudo systemctl edit --full --force shop.service $ sudo systemctl daemon-reload && systemctl enable --now shop # /etc/systemd/system/shop.service [Unit] Description=Resale shop web app After=network-online.target postgresql.service Wants=network-online.target StartLimitIntervalSec=60 StartLimitBurst=5 [Service] User=shop WorkingDirectory=/srv/shop EnvironmentFile=/etc/shop/env ExecStart=/srv/shop/.venv/bin/python -u /srv/shop/app.py Restart=on-failure RestartSec=5 TimeoutStopSec=20 [Install] WantedBy=multi-user.target [1] [2] [3] [4] [5] [6] [7] [8] [1] /etc/systemd/system is yours Package units live in /usr/lib/systemd/system. A file of the same name in /etc wins. Never edit the packaged one. [2] After= is ordering. Wants= is the dependency. After= alone only says “if postgres starts, start after it”. network-online, not network: the IP has to exist first. [3] These two belong in [Unit], not [Service] They moved in systemd 229. Left in [Service] they are ignored with an “Unknown key” warning almost nobody reads. [4] User=shop — the app does not run as root One line, and the blast radius of a bug shrinks enormously. Lesson 08 is about the rest. Paths must be absolute. [5] EnvironmentFile= — where Lesson 03’s config lands One image, three environments: this file is the only thing that differs. chmod 600, root-owned, and never in git. [6] ExecStart= is not a shell No pipes, no >, no globs, no $VAR expansion as you expect. -u keeps Python unbuffered, or the logs arrive minutes late. [7] Restart= is a policy, and it has a limit RestartSec=5 spaces the attempts out — which makes systemd give up LESS often, not more — which is Fig. 3, below. [8] [Install] is the section systemctl enable reads enable just makes a symlink in multi-user.target.wants/. No [Install] section = enable fails = nothing starts at boot.
One unit file, annotated. The two commands above the file are not optional and are the two everyone forgets. daemon-reload is what makes systemd re-read the file — edit a unit without it and you will spend twenty minutes debugging the old version. And enable is a different verb from start: start runs it now, enable makes it run at boot. The shop’s 02:40 outage is exactly what “started but never enabled” looks like six months later — which is why the only honest test of “it comes back” is systemctl is-enabled shop, followed at some point by an actual reboot.

3Restart= is a policy, and the limit is the interesting part (Fig. 3)

Measured, with a 20-line stand-in for systemd’s restart logic, same instantly-crashing app, one setting changed:

Setting What happened End state
RestartSec=0.1 (the systemd default, 100 ms) 5 start attempts inside 1.6 s failed — down permanently, 1.6 s after healthy
RestartSec=3 9 restarts in 26 s; the 10 s window never held more than 4 attempts never failed — restarts forever, shop down 94 % of the window

The counter is a sliding window, not a total. Restart quickly and five attempts land in one window, so systemd concludes the service is broken and stops. Restart slowly and they never share a window, so systemd never concludes anything. Raising RestartSec makes systemd more persistent, not less — the opposite of what people intend. The window, not the delay, decides.

Shape three, off that scale — the slow death: a leak that kills the app every six hours restarts it four times a day, forever. Never failed, never alerted, ~120 small outages a month.

a start attempt not serving measured, not modelled — see the hands-on RestartSec=0.1 the systemd default (100 ms) 5 starts inside 1.6 s → start limit hit unit state: failed — and nothing happens again, ever Down until a human notices. An alert on unit state fires here — if one exists. RestartSec=3 “stop it hammering the database” 9 starts in 26 s · the 10 s window never holds more than 4 · limit never hit Unit state: never failed, not once — and the shop is down 94 % of this window. 0 5 10 15 20 25 30 s The counter is a sliding window, not a total StartLimitBurst counts start attempts inside the last StartLimitIntervalSec. Restart quickly and five attempts land inside the window together, so systemd concludes the service is broken and stops. Restart slowly and they never share a window, so systemd never concludes anything. Raising RestartSec makes systemd MORE persistent, which is rarely what was intended. Shape three, off this scale: the slow death A leak that kills the app every six hours restarts it four times a day. Never failed, never alerted, ~120 outages a month.
The same crash, two values of RestartSec. The setting that sounds more cautious is the one that lets the outage run forever. And notice what both rows have in common: the unit state is a terrible health signal. In the first row it says failed while the shop is down; in the second it says active most of the time while the shop is equally down. Alert on the port and the endpoint — Lesson 06’s ss, and later the golden signals — and use systemctl show shop -p NRestarts as the second signal, never the only one.

Restart= values

Restart= clean exit (0) non-zero exit signal / timeout systemctl stop use for
no (default) no no no no one-shot jobs, migrations, backups — where failure should stay visible
on-failure no yes yes no the sane default for a web app
on-abnormal no no yes no when your app exits 3 to mean "my config is wrong, do not retry"
always yes yes yes no a daemon that must never be down; the only one that catches exit-0 bugs — and hides them best

systemctl stop never restarts anything, for any policy: a manual stop is a statement of intent. systemd ≥ 254 adds RestartSteps= and RestartMaxDelaySec= for real exponential backoff.

The rule: a restart policy buys time, never a fix, and it is not free — every automatic restart drops whatever was in flight. So every restart must leave a trace a human reads: systemctl show shop -p NRestarts is a number; put it on a dashboard and alert when it moves. A service that restarted 112 times last month is an incident that has learned to hide.

The unit state is a terrible health signal. Row 1 says failed while the shop is down; row 2 says active most of the time while the shop is equally down. Alert on the port and the endpoint (Lesson 06’s ss, later the golden signals); use NRestarts as the second signal, never the only one.

4"Where are the logs?" should be a one-word answer

Under a unit, stdout/stderr go to the journal tagged with the unit’s name: journalctl -u shop.

  • -u shop -b — this unit, this boot. -b -1 = the previous boot (how you read what happened before a 02:40 reboot).
  • -p err --since '15 min ago' — turns ten thousand lines into the four that matter.
  • -f — live tail; run it while you restart the service in another window.
  • --disk-usage — the journal is capacity-managed by default: SystemMaxUse= is 10 % of the filesystem, capped at 4 GB, keeping 15 % free. Lesson 06’s 28 GB app.log could not happen here; instead a DEBUG flood evicts your own history. journalctl --vacuum-time=14d reclaims on demand.

The trap: journald’s default is Storage=auto = persistent only if /var/log/journal exists. Where it does not, the journal lives in /run (RAM) and is erased by every reboot — so on a default box the logs explaining why the machine rebooted are deleted by the reboot.

journalctl --list-boots          # one line back = volatile; many = persistent
mkdir -p /var/log/journal
systemd-tmpfiles --create --prefix /var/log/journal
systemctl restart systemd-journald

Also form the habit of journalctl -k — the kernel ring buffer, where the OOM killer writes. An OOM-killed process looks, in the app’s own logs, exactly like one that stopped mid-sentence.

5Say it so it can be checked

Sounds like an answer Can be checked by someone else
"the app is running" "shop.service active (running) since 02:41, NRestarts=0"
"it’ll come back if the box reboots" "systemctl is-enabled shop → enabled; last proved by a reboot on 12 Sep"
"it keeps crashing" "4 restarts in 40 min, each preceded by Errno 111 to postgres"
"the logs are gone" "--list-boots shows 1 boot: journal is volatile, /var/log/journal missing"
"nothing else is broken" "systemctl list-units --failed → 0 loaded units listed"

Footnote to Lesson 05: the 02:40 outage was not a skills problem — it was an inventory problem. Nobody owned the question "what runs on this box, and which of it survives a reboot?" That question has a two-command answer (systemctl list-unit-files --state=enabled, systemctl list-units --failed) and until somebody’s name is against asking it, the answer drifts. Phase 6 removes the question entirely by putting the unit file in git next to the app (Lesson 39).

6Hands-on (20 min)

Part A — needs a real Linux host (cheap VM, Raspberry Pi, Multipass, or WSL2 with systemd=true in /etc/wsl.conf). If systemctl --user status says Failed to connect to bus, you are in a container without systemd and Part A will not work there.

  1. A1 — write /srv/shop/app.py (a http.server one-liner) and the unit above, then systemd-analyze verify → daemon-reload → enable --now.
  2. A2 — read the three things that matter: systemctl status shop (the since, not the dot), systemctl is-enabled shop, systemctl show shop -p NRestarts, ss -tlnp | grep 8080.
  3. A3 — kill -9 $(systemctl show shop -p MainPID --value), watch it come back, watch NRestarts go to 1. Follow along with journalctl -u shop -f.
  4. A4 — make it die at start-up (sed -i '1i import sys; sys.exit(1)'), restart, wait 12 s: Active: failed (Result: exit-code), "Start request repeated too quickly", and start is refused. Recovery is two commands: reset-failed and then start.
  5. A5 — the only honest test: sudo reboot, then systemctl is-active shop and systemctl list-units --failed.

Part B — runs anywhere. Build crashy.sh (an app that always exits 1) and supervise.sh (20 lines implementing RESTART_SEC, LIMIT_BURST, LIMIT_INTERVAL as a sliding window). Then:

  • RESTART_SEC=0.1 ./supervise.sh → 5 attempts, start-limit hit, total elapsed 1.6 s.
  • RESTART_SEC=3 timeout 26 ./supervise.sh → 9 attempts, the window count plateaus at 4, it never gives up.
  • LIMIT_INTERVAL=30 on the same slow run → it gives up after five. The window decides.

B4 — write down the two numbers for your service: how long it takes to start, and how long it can be down before a customer notices. RestartSec > the first. StartLimitBurst × (RestartSec + start time) = how long systemd will keep trying — that must be smaller than the second.

Runbook additions

systemctl list-units --failed · systemctl status <svc> (read the since) · systemctl is-enabled <svc> · systemctl show <svc> -p NRestarts · journalctl -u <svc> -b -p err

Policy that would have prevented 02:40: nothing on a production box may be started with &, nohup, screen or tmux. Every one of those is an outage with a date on it — the next reboot — and the reboot will not be scheduled by you.

7Three questions to ask your team this week

  1. "If our production box rebooted right now, what would not come back — and how would we find out?" The honest answer is usually "from a customer". The check is one command; the proof is one scheduled reboot. A team that has never rebooted production on purpose does not know whether it can.
  2. "How many times did each of our services restart itself last month?" If nobody knows, the restart policy is not resilience, it is concealment. NRestarts is free and already counted.
  3. "Is anything on our servers running because a person once typed a command?" Cron jobs, a screen session, a script from a migration two years ago. Invisible, undocumented, reboot-fatal — and every unit file the audit produces is a line of Phase 6’s IaC already written.

Next — Lesson 08: Permissions and users — why the app should not run as root.

Written with the help of AI (Claude) and reviewed by Rayhanul Islam.

Back to top