The OOM Killer Explained: Why Linux Killed Your Process
Your process disappeared. Nothing in its own logs, no crash, no error. It was simply gone.
Check the kernel log:
sudo dmesg | grep -i -E 'out of memory|oom-kill'
sudo journalctl -k | grep -i 'killed process'
Out of memory: Killed process 4823 (postgres) total-vm:8419288kB,
anon-rss:7891244kB, file-rss:0kB, shmem-rss:8192kB, UID:26
pgtables:16284kB oom_score_adj:0
The kernel killed it. Here is why that is a thing Linux does.
Overcommit
Linux grants memory allocations it cannot necessarily honour.
When a process calls malloc(), the kernel maps address space and returns success. No physical memory is committed. Pages are allocated when the process actually writes to them, which is why a program can allocate 100GB on a 16GB machine and appear to succeed.
This is deliberate and mostly sensible. Programs routinely request far more than they use: a fork() duplicates the address space of a process that then immediately exec()s something small, and allocators reserve generously. Refusing those requests would waste enormous amounts of memory.
The consequence is that the kernel can promise more than it has. When processes collectively try to use what they were promised, there is no allocation to fail, because the allocations already succeeded. The kernel’s only remaining option is to kill something.
cat /proc/sys/vm/overcommit_memory
# 0 = heuristic (default), 1 = always allow, 2 = strict accounting
Mode 2 with overcommit_ratio refuses allocations beyond a limit, trading OOM kills for allocation failures. It is appropriate for some server workloads and breaks others, because many programs handle malloc returning NULL badly or not at all.
How the victim is chosen
Every process gets a score, roughly proportional to its resident memory as a fraction of total, then adjusted:
cat /proc/$(pgrep -f postgres | head -1)/oom_score
cat /proc/$(pgrep -f postgres | head -1)/oom_score_adj
oom_score_adj ranges from -1000 to 1000 and is added to the calculation. -1000 makes a process effectively immune; 1000 makes it the preferred target.
The practical result is that the largest process usually dies, which is frequently the wrong one. A memory leak in a small script pushes the system over the edge, and your database gets killed because it is the biggest thing running. The kernel is optimising for reclaiming the most memory fastest, not for identifying the culprit.
Reading the log entry properly
The OOM message includes a table of every process it considered:
[ pid ] uid tgid total_vm rss pgtables_bytes swapents oom_score_adj name
[ 1234] 0 1234 52341 1823 118784 0 0 sshd
[ 4823] 26 4823 2104822 1972811 16674816 0 0 postgres
[ 5102] 1000 5102 891234 823411 7241728 0 0 node
rss is in pages, typically 4KB each. Multiply by 4096 for bytes.
Read the whole table, not just the killed line. It shows what memory looked like at the moment of failure, which is the evidence you need to determine whether the killed process was the cause or a bystander. If postgres was killed at 7.5GB but node was at 3.1GB and had been 200MB an hour earlier, the leak is in node.
Protecting a process
# protect a running process
echo -1000 | sudo tee /proc/4823/oom_score_adj
# make something a preferred victim
echo 500 | sudo tee /proc/5102/oom_score_adj
For a systemd service, set it in the unit so it survives restarts:
[Service]
OOMScoreAdjust=-800
sudo systemctl edit postgresql
sudo systemctl daemon-reload
Protecting something means something else dies. Marking a database immune does not create memory; it redirects the kill to the next candidate, which might be sshd. Protect deliberately and sparingly, and never protect everything.
Constraining rather than protecting
Better than choosing who dies is preventing the situation. cgroups let you cap a service’s memory so it fails within its own limit instead of destabilising the machine:
[Service]
MemoryMax=2G
MemoryHigh=1.5G
MemoryHigh applies throttling pressure as the limit approaches, giving the process a chance to release memory. MemoryMax is hard: exceeding it triggers an OOM kill inside that cgroup only, so the misbehaving service dies and nothing else is affected.
This is the correct fix for a service with a known leak. It converts a system-wide outage into one service restarting.
For containers, the same mechanism:
docker run -m 2g --memory-swap 2g myimage
Our Docker run builder covers the flags.
Swap
Swap does not prevent OOM kills. It changes when they happen.
Swap gives the kernel somewhere to put pages that are not being used, which is genuinely useful: a long-running process with allocated-but-idle memory can have it swapped out to make room for page cache, improving overall performance.
What swap does not do is create memory. If the working set, the pages actively being touched, exceeds physical RAM, the system thrashes: constantly swapping pages in and out, with disk I/O saturated and everything unresponsive.
Thrashing is frequently worse than an OOM kill. An OOM kill loses one process; thrashing makes the entire machine unusable, sometimes for long enough that you cannot SSH in to fix it.
# is the machine thrashing
vmstat 1
# high si and so columns sustained means swapping heavily
Our swap and zram guides cover sizing. zram is genuinely good on desktops: compressed swap in RAM, several times faster than disk, and it meaningfully delays the point at which things go wrong.
earlyoom, for desktops
The kernel OOM killer triggers only when allocation actually fails. By then the machine has usually been thrashing for minutes and is completely unresponsive, which is the common experience of “Linux froze and I had to hard reboot”.
earlyoom watches available memory from user space and kills something while the system is still responsive:
sudo apt install earlyoom
sudo systemctl enable --now earlyoom
# /etc/default/earlyoom
EARLYOOM_ARGS="-m 5 -s 5 --avoid '^(sshd|systemd)$' --prefer '^(chrome|firefox)$'"
Kill something when under 5 percent memory and 5 percent swap, never sshd, prefer a browser tab. Losing a browser is a vastly better outcome than a frozen desktop.
Fedora enables this by default and it is a clear improvement.
Diagnosing the actual cause
# memory by process, largest first
ps aux --sort=-%mem | head -15
# a real breakdown, not the misleading free output
free -h
cat /proc/meminfo | grep -E 'MemTotal|MemAvailable|Cached|Buffers'
# watch it develop
watch -n 5 'free -h; echo; ps aux --sort=-%mem | head -5'
MemAvailable is the number that matters, not free. Linux uses spare memory for page cache, so free being near zero is normal and healthy. MemAvailable estimates what could be made available under pressure. Our free command guide covers why the output confuses people.
For a leak that develops over hours, log it and look at the trend:
while true; do
echo "$(date +%s) $(ps -o rss= -p PID)" >> /tmp/mem.log
sleep 60
done
A steadily climbing RSS with no plateau is a leak. A sawtooth is normal garbage collection.
Frequently Asked Questions
Why did Linux kill my process instead of failing the allocation?
Linux overcommits memory by default, meaning it grants allocation requests larger than available RAM on the assumption that most processes never touch everything they request. When processes actually use more than exists, there is no allocation to fail because the allocations already succeeded, so the kernel kills something to reclaim memory.
How does the OOM killer choose which process to kill?
It scores every process roughly in proportion to its memory use, adjusted by the oom_score_adj value, then kills the highest scorer. This usually means the largest process dies, which is frequently your database or application server rather than whatever actually caused the memory pressure.
How do I stop the OOM killer from killing a specific process?
Write a negative value to /proc/PID/oom_score_adj, with -1000 making a process effectively immune. For a systemd service, set OOMScoreAdjust in the unit file so it applies on every start. Protecting a process means something else dies instead, so use it deliberately.
Where do I find evidence that the OOM killer ran?
Check dmesg or the journal for lines containing Out of memory or oom-kill. The entry names the killed process, its memory use, and includes a table of every process considered. If a process vanished with no log of its own, this is the first place to look.
Does adding swap prevent the OOM killer?
It delays it rather than preventing it. Swap gives the kernel somewhere to put inactive pages, which buys time and works well for genuinely idle memory. If the working set truly exceeds physical memory, swap turns an OOM kill into severe thrashing, which is often worse because the machine becomes unresponsive rather than losing one process.
What does earlyoom do that the kernel OOM killer does not?
The kernel OOM killer only triggers when allocation genuinely fails, by which point the system has usually been thrashing and unresponsive for a long time. earlyoom watches available memory from user space and kills something while the machine is still responsive, which produces a much better experience on a desktop.