RCWRCW IT TrainingFree hands-on labs & simulators← Back to home
Linux · Detailed troubleshooting guide

systemd Service Failures: Failed Units, Restart Loops and Dependency Deadlocks

Almost every Linux service outage ends up in front of systemctl status. This guide shows how to read what it actually tells you, how to separate a crashing process from a broken dependency, and how to stop a restart loop from hiding its own root cause.

Published October 6, 2026 · By , Enterprise Infrastructure Architect

Read the status block in the right order

The output of systemctl status contains four facts that, taken in order, narrow almost any failure to a single cause. Read them deliberately rather than skipping to the log tail.

systemctl status nginx.service --no-pager -l
  • Loaded: tells you whether the unit file was found and whether it is enabled. Loaded: bad setting or not-found is a unit-file problem, not a service problem.
  • Drop-In: lists override fragments under /etc/systemd/system/<unit>.d/. Overrides are the single most common reason a unit behaves differently from its vendor file.
  • Active: the state and the reason. failed (Result: exit-code), failed (Result: timeout) and inactive (dead) are three completely different investigations.
  • Process / Main PID: the exit status. status=1/FAILURE is the application's own exit code; status=203/EXEC means systemd could not execute the binary at all.

The exit codes in the 200 range come from systemd itself, not from your program, and each names a specific failure:

CodeMeaningUsual cause
200/CHDIRCould not change to WorkingDirectoryPath missing or not readable by User=
203/EXECCould not execute the binaryWrong ExecStart path, missing interpreter, or no execute bit
204/FDSFailed to set up file descriptorsSocket activation misconfigured
208/STDINStandard input could not be openedStandardInput= points at a missing path
209/STDOUTStandard output could not be openedLog path or directory missing
217/USERUser= does not existService account removed or never created
226/NAMESPACENamespace or sandbox setup failedProtectSystem=, PrivateTmp= or ReadWritePaths= conflict

A 203 is never an application bug. Check the path exists, is executable, and is reachable by the configured User= before reading another log line.

Separate the unit's logs from everything else

Unscoped journalctl output is unusable during an incident. Scope it to the unit and to the current boot, and include the priority so warnings are not lost in debug noise.

# this boot, this unit, no pager
journalctl -u nginx.service -b --no-pager

# the last failure only: start at the last restart
journalctl -u nginx.service --since "10 min ago" -o short-precise

# include everything the unit's children logged too
journalctl _SYSTEMD_UNIT=nginx.service + UNIT=nginx.service -b

# warnings and worse, across all units, ordered
journalctl -b -p warning --no-pager

The -o short-precise format adds microseconds, which matters when you are trying to establish whether the database died before or after the application that depends on it.

If a service writes to its own file rather than the journal, systemd sees only the exit code and the journal looks empty. Confirm where output goes before concluding there is nothing to read:

systemctl show nginx.service -p StandardOutput -p StandardError
systemctl cat nginx.service | grep -i 'Standard\|Log'

Restart loops: when the fix hides the fault

Restart=always with a short RestartSec= turns a clean crash into a storm that overwrites its own evidence. systemd eventually gives up and reports a rate-limit message that is easy to misread as the cause:

nginx.service: Start request repeated too quickly.
nginx.service: Failed with result 'exit-code'.
nginx.service: Scheduled restart job, restart counter is at 5.

"Start request repeated too quickly" is the consequence. The real error is in the first of the five attempts, not the last. Retrieve it directly:

journalctl -u nginx.service -b | grep -n -m1 -A20 'Starting\|Started'

The rate limiter is governed by two directives, and raising them is almost always the wrong response during diagnosis — lowering the restart rate so you can read the logs is the right one:

[Unit]
StartLimitIntervalSec=60
StartLimitBurst=5

[Service]
Restart=on-failure
RestartSec=10s

To investigate, stop the loop entirely and run the ExecStart command by hand as the service user. This surfaces errors that systemd swallows:

systemctl stop nginx.service
systemctl show nginx.service -p ExecStart -p User -p Environment
sudo -u nginx /usr/sbin/nginx -g 'daemon off;'

Clear the counter once you have the answer with systemctl reset-failed nginx.service.

Dependency problems look like service problems

When the Active line reads failed (Result: dependency), the unit never executed. Something it required failed first, and chasing the application is wasted effort. systemd distinguishes two concepts that are routinely conflated:

DirectiveControlsFailure behaviour
Requires=RequirementIf the dependency fails, this unit is stopped too
Wants=Weak requirementDependency failure is tolerated
After= / Before=Ordering onlyNo requirement at all — ordering without Requires= does not guarantee the unit ran
BindsTo=Strict bindingThis unit stops whenever the dependency stops, even cleanly

The classic bug is After=postgresql.service with no Requires=. The application starts in the right order but happily starts when the database is dead.

Map the real graph rather than guessing it:

# what this unit pulls in, recursively
systemctl list-dependencies myapp.service

# what depends on this unit
systemctl list-dependencies --reverse postgresql.service

# which units failed this boot
systemctl --failed --no-pager

# ordering cycles systemd broke on its own
journalctl -b | grep -i 'ordering cycle\|breaking ordering'

An ordering cycle is serious: systemd resolves it by deleting one edge of its choosing, so the unit it drops varies between boots and produces an intermittent failure that is maddening to reproduce.

Timeouts, Type= mismatches and the hang that is not a hang

Result: timeout usually means the Type= is wrong for how the process actually behaves. systemd waits for a readiness signal that will never arrive.

Type=systemd considers the unit started when…Use for
simplethe process is forked (immediately)Foreground processes
execthe binary has been executed successfullyForeground processes, stricter than simple
forkingthe parent exitsClassic daemons that background themselves
notifythe process calls sd_notify(READY=1)systemd-aware daemons
oneshotthe process exits successfullyScripts and migrations

A daemon that forks but is declared Type=simple appears to start and then immediately registers as dead. A foreground process declared Type=forking hangs until TimeoutStartSec= expires — 90 seconds by default — then is killed. Both look like application faults and are neither.

Shutdown has the mirror problem. A unit that ignores SIGTERM stalls the whole shutdown until TimeoutStopSec= expires:

systemd-analyze blame | head -20
systemd-analyze critical-chain myapp.service
systemctl show myapp.service -p TimeoutStartUSec -p TimeoutStopUSec -p Type

If the service needs longer than the default to come up legitimately — a large database replaying WAL, for example — raise TimeoutStartSec= deliberately rather than setting it to infinity, which simply converts a noisy failure into a silent hang.

Validate before you reload

Editing unit files by hand in /usr/lib/systemd/system/ is a mistake: a package update overwrites it. Use an override, which systemd merges over the vendor file:

systemctl edit myapp.service          # creates /etc/systemd/system/myapp.service.d/override.conf
systemctl cat myapp.service           # shows the merged, effective result
systemd-analyze verify myapp.service  # parses and reports errors before you apply
systemctl daemon-reload
systemctl restart myapp.service

One override trap deserves naming: list-valued directives such as ExecStart= and Environment= append rather than replace. To replace one you must first clear it with an empty assignment:

[Service]
ExecStart=
ExecStart=/usr/local/bin/myapp --config /etc/myapp/prod.yaml

Omitting the blank ExecStart= yields "Service has more than one ExecStart= setting, which is only allowed for Type=oneshot" — a confusing error with a trivial cause.

A diagnosis checklist that works under pressure

  1. Run systemctl --failed first. The unit you were told about may be a downstream casualty.
  2. Read Loaded, Drop-In, Active and the exit code before any log.
  3. If the result is dependency, stop and investigate the dependency instead.
  4. If the code is in the 200s, it is a unit-file or permission fault, not an application fault.
  5. If it is a restart loop, read the first attempt in the journal, not the last.
  6. If the result is timeout, verify Type= matches the process's actual behaviour.
  7. Reproduce by running ExecStart manually as User= with the unit stopped.
  8. After any edit: systemd-analyze verify, then daemon-reload, then restart, then reset-failed.
Key takeaway: Read the exit code before the logs: a non-zero status points at the application, a signal points at the kernel or a timeout, and dependency in the Active line means the unit never ran at all. Use journalctl -u … --no-pager -b scoped to the boot, and always check systemd-analyze verify after editing a unit.