Diagnose and Fix Linux OOM Killer on Production Servers

Introduction: Linux OOM killer on production servers

When a Linux server runs critically low on memory the kernel invokes the OOM killer to reclaim memory by terminating processes. On production servers an unexpected OOM killer event can cause application downtime, data loss, and noisy on call pages for operations teams. This guide targets system administrators and engineers who need practical diagnostics and fixes that work under pressure.

We will cover how to read dmesg and kernel logs, trace the terminated process, change oom_score_adj and oom_score, use cgroups and systemd memory limits, tune vm.overcommit and swap behavior, and add monitoring and alerts to prevent recurrence. Each section provides commands and playbook style steps you can apply in production.

How the OOM killer works in Linux

The Linux kernel tracks memory usage across the system. When available memory plus reclaimable pages cannot satisfy allocation requests, the kernel tries reclaim and swap. If that fails and the system can not make progress the OOM killer selects one or more processes to kill based on heuristics. Those heuristics include process memory usage, oom_score, oom_score_adj, and the kernel estimation of process badness.

Understanding these heuristics matters for targeted mitigation. By adjusting oom_score_adj and by changing how much memory processes are allowed to consume under cgroups you can steer the kernel away from killing critical services. You can also tune the memory subsystem so the kernel has more headroom before invoking the OOM killer.

Finding OOM events in logs and dmesg

The kernel writes OOM events to dmesg and to the system journal. Start with dmesg and journald to find the killer notice and the list of processes involved. Look for lines that include the phrase Out of memory or OOM killed, the PID and process name, and the oom victim details. These entries show the memory footprint at the time and the kernel signal used.

See also  Linux Journald Log Management: Inspect, Limit, Vacuum

Commands and files to check include the following list. Run them on the affected host or on retained log archives to identify the event timestamp and scope.

  • sudo dmesg | grep -i oom
  • sudo journalctl -k | grep -i oom
  • journalctl -u yourservice.service –since “YYYY-MM-DD HH:MM”
  • /var/log/kern.log and /var/log/messages depending on distro

Trace the killed process and memory accounting

Once you identify the PID and name from logs you can inspect procfs snapshots if you collected them. If the process is gone you can still reconstruct its footprint from logs. For live investigation use tools such as ps, top, smem, and pmap to inspect resident set size and virtual size. smem provides per process proportional set size that is helpful for shared memory heavy processes.

Check cgroup membership and memory.stat values for the unit if you run services under systemd or cgroups. Look at /sys/fs/cgroup/memory/… to find which slice or unit hit its limit. On systemd systems use systemctl status and systemd-cgtop for quick overviews. This helps tell whether a single process or a group exhausted available allocations.

Adjust oom_score_adj and oom_score safely

oom_score_adj provides a simple way to bias OOM selection. A higher positive value increases likelihood of being killed, a negative value protects a process. For critical services set a negative oom_score_adj, but test carefully because protected processes can cause other services to be killed instead. Change the score at runtime in procfs or set it in unit files for persistent behavior.

Examples you can run as root include echoing values to /proc/PID/oom_score_adj, and adding directives in systemd unit files under Service with MemoryAccounting and an ExecStartPre that sets oom_score_adj. Use conservative values such as -100 for high importance services and +100 for batch jobs that are safe to kill.

See also  Optimize NVMe SSD Performance on Linux Servers
Linux OOM killer

Configure cgroups and systemd memory limits

Cgroups let you limit memory per service or slice so a runaway process cannot take down the whole host. On systemd systems enable MemoryAccounting and set MemoryMax for the service to a safe value. For containers use cgroup v1 or v2 tools to set limits per container instance. Limits prevent global OOM events by failing allocations inside the group instead of the whole kernel invoking the killer.

Practical steps include the following commands and unit file changes. Test limits during low traffic before applying in production to avoid service thrashing.

  • In unit file: MemoryAccounting=yes and MemoryMax=512M
  • Runtime: systemctl set-property yourservice MemoryMax=512M
  • For containers: echo 536870912 > /sys/fs/cgroup/memory/yourgroup/memory.limit_in_bytes

Tune vm.overcommit and swap behavior

Kernel overcommit controls how aggressive allocations are allowed before actual memory is backed. vm.overcommit_memory values change allocation policy, and vm.overcommit_ratio affects allowed commit when policy is heuristic. For production servers avoid pure conservative settings that break legitimate workloads, instead tune to match application behavior. For example set vm.overcommit_memory to 2 with an appropriate vm.overcommit_ratio for systems with predictable memory usage.

Swap is an important safety net. Ensure you have enough swap to absorb transient spikes, or enable zswap where appropriate. If you add swap be aware latency costs for I O sensitive workloads. A combination of modest swap and proper cgroup limits reduces the chance of kernel level OOM events.

Immediate mitigation and monitoring

If you encounter an ongoing OOM situation apply quick mitigations such as lowering oom_score_adj on critical services, adding swap, killing noncritical batch jobs, or moving workloads off the host. Use ps and systemctl to identify and stop batch or cron processes that can be sacrificed. These actions buy time while you perform root cause analysis.

For long term stability put monitoring and alerts in place. Instrument memory usage per host, per cgroup, and per process where possible. Integrate alerts for headroom thresholds so teams can react before the kernel steps in. Automate remediation for known patterns, such as restarting leak prone workers or scaling out application nodes.

  • Monitor metrics: available memory, page faults, swap usage, memory.max and memory.current for cgroups
  • Alert at 70 and 90 percent of steady state working set to allow action before OOM
See also  Diagnose and Speed Linux Boot: systemd Unit Tuning

FAQ

Q: How do I know if the OOM killer killed my database process? A: Check dmesg and journalctl for OOM lines that list the database process name and PID, then correlate with service logs and timestamps to confirm.

Q: Can I completely disable the OOM killer? A: Disabling the OOM killer is not recommended on production hosts. It can lead to system hang when memory is exhausted. Better approaches are cgroup limits and careful tuning of overcommit and swap.

Q: Will setting oom_score_adj protect a process permanently? A: Setting a negative oom_score_adj reduces the chance of selection but it is not absolute. If the system has no other options the kernel may still kill a protected process, so use limits and monitoring as additional protections.

Q: What if my containers are causing OOM events? A: Use cgroups to limit container memory, enable MemoryAccounting, and set MemoryMax for container units. Ensure orchestration tooling respects these limits and scales workloads appropriately.

Conclusion

Diagnosing and fixing Linux OOM killer events on production servers requires a mix of fast incident response and longer term system design. Start by locating OOM log entries in dmesg and the system journal, then trace the terminated process footprint and cgroup accounting. Short term mitigations such as adjusting oom_score_adj, adding swap, and stopping noncritical processes give you breathing room to investigate. For durable fixes use cgroups and systemd memory limits to bound aggressive processes, tune vm.overcommit and swap to match your workload profile, and add monitoring so you can detect headroom exhaustion before the kernel acts.

Avoid one off protections that transfer the problem to other services. Instead create predictable memory budgets per service and automate scaling or restarts for known failure modes. With consistent logging, alerting, and conservative limits you can reduce on call noise and prevent surprise terminations, while maintaining service stability and performance across your production fleet.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top