systemd Service Failures: Failed Units, Restart Loops and Dependency Deadlocks
Almost every Linux service outage ends up in front of systemctl status. This guide shows how to read what it actually tells you, how to separate a crashing process from a broken dependency, and how to stop a restart loop from hiding its own root cause.
Read the status block in the right order
The output of systemctl status contains four facts that, taken in order, narrow almost any failure to a single cause. Read them deliberately rather than skipping to the log tail.
systemctl status nginx.service --no-pager -l
- Loaded: tells you whether the unit file was found and whether it is enabled.
Loaded: bad settingornot-foundis a unit-file problem, not a service problem. - Drop-In: lists override fragments under
/etc/systemd/system/<unit>.d/. Overrides are the single most common reason a unit behaves differently from its vendor file. - Active: the state and the reason.
failed (Result: exit-code),failed (Result: timeout)andinactive (dead)are three completely different investigations. - Process / Main PID: the exit status.
status=1/FAILUREis the application's own exit code;status=203/EXECmeans systemd could not execute the binary at all.
The exit codes in the 200 range come from systemd itself, not from your program, and each names a specific failure:
| Code | Meaning | Usual cause |
|---|---|---|
| 200/CHDIR | Could not change to WorkingDirectory | Path missing or not readable by User= |
| 203/EXEC | Could not execute the binary | Wrong ExecStart path, missing interpreter, or no execute bit |
| 204/FDS | Failed to set up file descriptors | Socket activation misconfigured |
| 208/STDIN | Standard input could not be opened | StandardInput= points at a missing path |
| 209/STDOUT | Standard output could not be opened | Log path or directory missing |
| 217/USER | User= does not exist | Service account removed or never created |
| 226/NAMESPACE | Namespace or sandbox setup failed | ProtectSystem=, PrivateTmp= or ReadWritePaths= conflict |
A 203 is never an application bug. Check the path exists, is executable, and is reachable by the configured User= before reading another log line.
Separate the unit's logs from everything else
Unscoped journalctl output is unusable during an incident. Scope it to the unit and to the current boot, and include the priority so warnings are not lost in debug noise.
# this boot, this unit, no pager journalctl -u nginx.service -b --no-pager # the last failure only: start at the last restart journalctl -u nginx.service --since "10 min ago" -o short-precise # include everything the unit's children logged too journalctl _SYSTEMD_UNIT=nginx.service + UNIT=nginx.service -b # warnings and worse, across all units, ordered journalctl -b -p warning --no-pager
The -o short-precise format adds microseconds, which matters when you are trying to establish whether the database died before or after the application that depends on it.
If a service writes to its own file rather than the journal, systemd sees only the exit code and the journal looks empty. Confirm where output goes before concluding there is nothing to read:
systemctl show nginx.service -p StandardOutput -p StandardError systemctl cat nginx.service | grep -i 'Standard\|Log'
Restart loops: when the fix hides the fault
Restart=always with a short RestartSec= turns a clean crash into a storm that overwrites its own evidence. systemd eventually gives up and reports a rate-limit message that is easy to misread as the cause:
nginx.service: Start request repeated too quickly. nginx.service: Failed with result 'exit-code'. nginx.service: Scheduled restart job, restart counter is at 5.
"Start request repeated too quickly" is the consequence. The real error is in the first of the five attempts, not the last. Retrieve it directly:
journalctl -u nginx.service -b | grep -n -m1 -A20 'Starting\|Started'
The rate limiter is governed by two directives, and raising them is almost always the wrong response during diagnosis — lowering the restart rate so you can read the logs is the right one:
[Unit] StartLimitIntervalSec=60 StartLimitBurst=5 [Service] Restart=on-failure RestartSec=10s
To investigate, stop the loop entirely and run the ExecStart command by hand as the service user. This surfaces errors that systemd swallows:
systemctl stop nginx.service systemctl show nginx.service -p ExecStart -p User -p Environment sudo -u nginx /usr/sbin/nginx -g 'daemon off;'
Clear the counter once you have the answer with systemctl reset-failed nginx.service.
Dependency problems look like service problems
When the Active line reads failed (Result: dependency), the unit never executed. Something it required failed first, and chasing the application is wasted effort. systemd distinguishes two concepts that are routinely conflated:
| Directive | Controls | Failure behaviour |
|---|---|---|
Requires= | Requirement | If the dependency fails, this unit is stopped too |
Wants= | Weak requirement | Dependency failure is tolerated |
After= / Before= | Ordering only | No requirement at all — ordering without Requires= does not guarantee the unit ran |
BindsTo= | Strict binding | This unit stops whenever the dependency stops, even cleanly |
The classic bug is After=postgresql.service with no Requires=. The application starts in the right order but happily starts when the database is dead.
Map the real graph rather than guessing it:
# what this unit pulls in, recursively systemctl list-dependencies myapp.service # what depends on this unit systemctl list-dependencies --reverse postgresql.service # which units failed this boot systemctl --failed --no-pager # ordering cycles systemd broke on its own journalctl -b | grep -i 'ordering cycle\|breaking ordering'
An ordering cycle is serious: systemd resolves it by deleting one edge of its choosing, so the unit it drops varies between boots and produces an intermittent failure that is maddening to reproduce.
Timeouts, Type= mismatches and the hang that is not a hang
Result: timeout usually means the Type= is wrong for how the process actually behaves. systemd waits for a readiness signal that will never arrive.
| Type= | systemd considers the unit started when… | Use for |
|---|---|---|
simple | the process is forked (immediately) | Foreground processes |
exec | the binary has been executed successfully | Foreground processes, stricter than simple |
forking | the parent exits | Classic daemons that background themselves |
notify | the process calls sd_notify(READY=1) | systemd-aware daemons |
oneshot | the process exits successfully | Scripts and migrations |
A daemon that forks but is declared Type=simple appears to start and then immediately registers as dead. A foreground process declared Type=forking hangs until TimeoutStartSec= expires — 90 seconds by default — then is killed. Both look like application faults and are neither.
Shutdown has the mirror problem. A unit that ignores SIGTERM stalls the whole shutdown until TimeoutStopSec= expires:
systemd-analyze blame | head -20 systemd-analyze critical-chain myapp.service systemctl show myapp.service -p TimeoutStartUSec -p TimeoutStopUSec -p Type
If the service needs longer than the default to come up legitimately — a large database replaying WAL, for example — raise TimeoutStartSec= deliberately rather than setting it to infinity, which simply converts a noisy failure into a silent hang.
Validate before you reload
Editing unit files by hand in /usr/lib/systemd/system/ is a mistake: a package update overwrites it. Use an override, which systemd merges over the vendor file:
systemctl edit myapp.service # creates /etc/systemd/system/myapp.service.d/override.conf systemctl cat myapp.service # shows the merged, effective result systemd-analyze verify myapp.service # parses and reports errors before you apply systemctl daemon-reload systemctl restart myapp.service
One override trap deserves naming: list-valued directives such as ExecStart= and Environment= append rather than replace. To replace one you must first clear it with an empty assignment:
[Service] ExecStart= ExecStart=/usr/local/bin/myapp --config /etc/myapp/prod.yaml
Omitting the blank ExecStart= yields "Service has more than one ExecStart= setting, which is only allowed for Type=oneshot" — a confusing error with a trivial cause.
A diagnosis checklist that works under pressure
- Run
systemctl --failedfirst. The unit you were told about may be a downstream casualty. - Read
Loaded,Drop-In,Activeand the exit code before any log. - If the result is
dependency, stop and investigate the dependency instead. - If the code is in the 200s, it is a unit-file or permission fault, not an application fault.
- If it is a restart loop, read the first attempt in the journal, not the last.
- If the result is
timeout, verifyType=matches the process's actual behaviour. - Reproduce by running
ExecStartmanually asUser=with the unit stopped. - After any edit:
systemd-analyze verify, thendaemon-reload, then restart, thenreset-failed.
dependency in the Active line means the unit never ran at all. Use journalctl -u … --no-pager -b scoped to the boot, and always check systemd-analyze verify after editing a unit.