Linux Memory Pressure: Reading the OOM Killer, Swap Thrash and cgroup Limits
A Linux box with almost no free memory is usually healthy, and one with plenty free can still be thrashing. This guide covers the metrics that actually indicate memory pressure, and how to read an OOM dump to find out what really died and why.
free is not the metric you think it is
Linux deliberately uses otherwise-idle RAM for page cache. A healthy, well-used server shows very little free memory, and that is correct behaviour rather than a warning sign.
free -h # total used free shared buff/cache available # Mem: 31Gi 9.2Gi 412Mi 1.1Gi 21Gi 20Gi
Here free is 412 MiB but available is 20 GiB. The kernel can reclaim most of the page cache on demand. available is the only column that answers "can I start another process".
grep -E 'MemTotal|MemFree|MemAvailable|Dirty|Writeback|SwapTotal|SwapFree' /proc/meminfo
Alert on MemAvailable as a percentage of MemTotal. Alerting on MemFree generates constant false positives on exactly the machines that are performing well.
Pressure Stall Information: the signal that is actually actionable
PSI reports the proportion of time tasks were stalled waiting on memory. Unlike utilisation, it measures harm directly, and it is available on any kernel 4.20 or later.
cat /proc/pressure/memory # some avg10=0.00 avg60=0.00 avg300=0.00 total=0 # full avg10=0.00 avg60=0.00 avg300=0.00 total=0
- some — at least one task was stalled waiting for memory.
- full — every runnable task was stalled; nothing useful progressed.
A sustained some avg60 above roughly 10 means reclaim is costing real throughput. Any sustained full above zero means the machine is spending measurable wall-clock time doing nothing but reclaim, which is the quantitative definition of thrashing.
cat /sys/fs/cgroup/<slice>/memory.pressure # per-cgroup, same format
Reading an OOM kill dump
When the kernel kills a process, it logs a full report. Most people read the last line and stop, which is where the misdiagnosis starts.
dmesg -T | grep -i -A5 'killed process' journalctl -k --since "1 hour ago" | grep -i 'oom\|killed process'
Out of memory: Killed process 12843 (java) total-vm:9183424kB, anon-rss:6291456kB, file-rss:0kB, shmem-rss:0kB, UID:1001 pgtables:13420kB oom_score_adj:0
- anon-rss is the figure that matters — anonymous resident memory, which cannot be reclaimed, only swapped or freed.
- total-vm is virtual address space and is frequently enormous and harmless. A JVM reserving 9 GiB of address space while using 6 GiB resident is normal.
- oom_score_adj shows whether the process had been biased for or against selection.
The critical field appears just above the kill line and tells you which memory ran out:
oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null), cpuset=docker-9f2c.scope,mems_allowed=0, oom_memcg=/system.slice/docker-9f2c.scope
| Constraint | Meaning | Where to fix it |
|---|---|---|
CONSTRAINT_NONE | The host genuinely exhausted memory | Host sizing, or the leaking process |
CONSTRAINT_MEMCG | A cgroup hit its own limit; the host was fine | Container or unit memory limit |
CONSTRAINT_CPUSET | Memory exhausted within a restricted cpuset | NUMA or cpuset placement |
CONSTRAINT_MEMORY_POLICY | A NUMA node was exhausted | NUMA balancing or node pinning |
A CONSTRAINT_MEMCG kill on a host showing 60 percent free memory is not a host capacity problem, and adding RAM will change nothing.
cgroup v2 limits and the kill you can prevent
Under cgroup v2, three files govern behaviour, and the middle one is the one most people never set:
| File | Effect |
|---|---|
memory.max | Hard limit. Exceeding it triggers a cgroup OOM kill. |
memory.high | Soft limit. The kernel throttles and reclaims aggressively instead of killing. |
memory.min | Reclaim protection — memory guaranteed not to be reclaimed. |
systemctl show myapp.service -p MemoryMax -p MemoryHigh -p MemoryCurrent cat /sys/fs/cgroup/system.slice/myapp.service/memory.events # low 0 # high 142 # max 3 # oom 1 # oom_kill 1
memory.events is the cleanest evidence available. A rising high counter with zero oom_kill means the workload is being throttled but surviving — uncomfortable, not fatal. Setting MemoryHigh= slightly below MemoryMax= converts abrupt kills into gradual slowdowns you can alert on.
Swap, swappiness and when swap is correct
Swap is not a substitute for RAM and its presence does not cause slowness — using it heavily does. The useful distinction is between cold pages being swapped out once (beneficial) and pages being swapped in and out continuously (thrashing).
vmstat 1 10 # si and so columns: sustained non-zero values are thrash # a one-off burst during a backup is not sar -W 1 5 # pswpin/s and pswpout/s swapon --show
Find the processes actually swapped out rather than guessing:
for p in /proc/[0-9]*; do
s=$(awk '/^VmSwap/{print $2}' $p/status 2>/dev/null)
[ -n "$s" ] && [ "$s" -gt 0 ] && echo "$s KB $(cat $p/comm)"
done | sort -rn | head -10
vm.swappiness sets the kernel's relative preference for reclaiming anonymous pages versus page cache. It is a ratio, not an on/off switch:
sysctl vm.swappiness sysctl -w vm.swappiness=10 # test echo 'vm.swappiness=10' > /etc/sysctl.d/99-swap.conf # persist
- 60 — distribution default, reasonable for general workloads.
- 10 — databases and latency-sensitive services; prefer dropping cache over swapping.
- 1 — swap only under genuine pressure.
- 0 — does not disable swap; it makes the kernel avoid it until an allocation would otherwise fail, which can make OOM kills more likely and more abrupt.
Overcommit, and why raising it rarely helps
Linux allows processes to reserve more address space than exists, because most reservations are never touched. Three modes are available:
| vm.overcommit_memory | Behaviour |
|---|---|
0 (default) | Heuristic — obvious overcommits refused, most allowed |
1 | Always allow. Required by Redis and some in-memory stores for fork-based persistence |
2 | Strict accounting against swap + RAM × overcommit_ratio; allocations fail rather than OOM later |
sysctl vm.overcommit_memory vm.overcommit_ratio grep -E 'Commit' /proc/meminfo
Mode 2 converts an unpredictable OOM kill into a predictable malloc() failure, which well-written software handles gracefully and poorly-written software crashes on anyway. It is worth considering on single-purpose hosts and rarely worth it on mixed ones.
Protecting the processes that matter
The OOM killer selects a victim by score, roughly proportional to memory use. You can bias the choice, and should for anything whose death makes the incident worse — an SSH daemon, a monitoring agent, a cluster heartbeat:
cat /proc/$(pidof sshd)/oom_score cat /proc/$(pidof sshd)/oom_score_adj # bias strongly against selection (-1000 .. 1000) echo -900 > /proc/$(pidof sshd)/oom_score_adj
In a unit file, make it durable and explicit:
[Service] OOMScoreAdjust=-900 MemoryHigh=2G MemoryMax=3G
Never set -1000 on a process that can itself leak: it becomes unkillable by the OOM path and the kernel will work through every other process on the box before touching it.
Triage order
- Check
MemAvailable, notMemFree. - Check
/proc/pressure/memory—fullabove zero means real stalling. - Search
dmesg -Tfor OOM dumps and read the constraint line. CONSTRAINT_MEMCGmeans a container limit, not a host shortage.- Check
memory.eventsfor the cgroup:highversusoom_killdistinguishes throttling from killing. - Check
vmstatsi/so for sustained swap traffic, not one-off bursts. - Identify growth with
smem -rs ussor per-processVmRSSover time before blaming the kernel. - Set
MemoryHigh=andOOMScoreAdjust=deliberately rather than removing limits.
oom_kill_constraint=CONSTRAINT_MEMCG means a container hit its own limit while the host had memory to spare. Fix the limit or the leak; raising overcommit only moves the failure.