Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Tangential, but does anyone know why in 2026 and on Debian 13, my machine still hangs when some process exhausts RAM?

Is there really no higher-priority kernel process to prevent total freeze of the system and send a SIGKILL to the culprit process when such a scenario happens?



This is called an oomkiller. The kernel has one but it kicks in very late and the kernel prefers to do page trashing instead of killing processes.

systemd-oomd should be integrated in systemd, you can configure it to your liking and see if it improves your problem.


I wish there was an easy way to configure it to say "target user processes first, specifically java (or these days python)" as in my experience they are always the culprits. Processes owned by system accounts or root should be the last ones killed.


Similarly, in the past I have wished for the ability to exempt a process from the oomkiller. I've run servers where the top memory user was also the server's entire reason for existence, and if that process gets killed the server may as well be down. It would literally have been better for any other process to get killed, but it was always the application process because of the memory usage.


systemd-oomd works reasonably well and there is source code. Perhaps claude can help add more detailed policy support to it.

Just found this comment:

https://news.ycombinator.com/item?id=49663299


I believe that you mean: https://en.wikipedia.org/wiki/Thrashing_(computer_science)

Chris Siebenmann discusses when the OOM killer triggers: https://utcc.utoronto.ca/~cks/space/blog/linux/OOMKillerWhen

Chris disables systemd-oomd after it obliterates his X session with no explanation: https://utcc.utoronto.ca/~cks/space/blog/linux/SystemdOomdNo...


> First off, this is exactly how systemd-oomd is supposed to behave under memory pressure. The documentation is specific on this; systemd-oomd itself says:

> > [...] If the configured limits are exceeded, systemd-oomd will select a cgroup to terminate, and send SIGKILL to all processes in it. [...]

> By having the user@.service template be enrolled in systemd-oomd, Fedora made the cgroup that systemd-oomd would select to be killed be all of your processes (across all of your sessions, if you have more than one). ...

Maybe *Fedora* has fixed or improved in the last 4 years. Or maybe they don't run Fedora.


One thing Fedora does now, is use zram.

In my experience it works really well. I wonder why my computer is a bit sluggish, and find out I have several gigs in zram.

If that was in swap on a disk, it would be really painful.


If it were swap on disk fronted by zswap, it'd be even better ;)


I'm a Fedora developer and I can assure you that Fedora's behaviour when it runs out of memory is still terrible.


Adding to this, glad systemd-oomd finally added solid rulesets in 261

Made it far easier to target any containers that got too hot rather than ever risk anything higher priority.


Set `/sys/kernel/mm/lru_gen/min_ttl_ms` at boot (see https://docs.kernel.org/admin-guide/mm/multigen_lru.html).

User-space OOM killers never really worked for me and imo are not a proper solution anyway. This option instead lets you make the kernel OOM killer actually work for desktop use.

Currently have it set to `1000` and it works very well for me (don't remember the last time I had a full system freeze due to OOM).


Don't worry, it's not just you: https://lkml.org/lkml/2019/8/4/15

It's because linux is a toy OS. Specifically, it overcommits memory in the hope/assumption that it won't all be used at once, but doesn't have a way to gracefully degrade when applications collectively want to use more memory(+swap) than it actually has. You can turn off overcommit, but applications are designed with the overcommitting feature in mind, so your experience might not be as good as you were hoping for.

Making a massive swap space helps a little bit. It's better to just never let your actual memory usage go above 85% to 90%. It's fine to go above if you're trying to optimize a server with a specific set of processes to wring every last bit of efficiency out of it, but not for general desktop computing.

If it really bothers you OpenBSD (edit: thanks for the reminder TimTheTinker) and Illumos don't allow overcommit at all and Windows handles this situation much more gracefully, so WSL is an option too. If you don't mind Oracle (i do), solaris also doesn't allow overcommit.


> It's because linux is a toy OS. Specifically, it overcommits memory

…by default. It can be disabled via a sysctl:

* https://www.kernel.org/doc/Documentation/vm/overcommit-accou...


You're right. I'm glad the very next sentence in my comment landed.

The problem with turning it off is that the system and applications have been architected assuming that it will be on, so things like fork/execing a memory heavy processes or allocating memory inside a cgroup (which still pretends overcommit is enabled and there's still no way to disable that assumption) that used to work fine might break with no good way to get them to work again. This comment (and siblings) have more specifics: https://news.ycombinator.com/item?id=27794237#27795199


Illumos also doesn't overcommit, if you want the more modern OS descended from Solaris.


can't believe i forgot to mention them! Updating the parent for visibility


or use the system's oom ?


That's what's breaking the system and causing freezes. You can tune it a bit to minimize when it happens, but not get rid of the issue entirely.


All I want is the oomkiller to always kill firefox. Somehow that's very difficult to achieve.


The oom killer is a monkey with a gun. It’s a last resort when you were going to crash (or worse, hang) anyway.

You should have a much better “plan A” to avoid this situation.


... which then kills sshd, locking you out of being able to get in and do any recovery.


the system will restart sshd?


the system will kill whatever it damn well pleases. You can tune it with priorities to ask it to try to not kill that, and once it kills your sshd once, you'll probably configure it to exempt sshd from being OOM killed at all. That doesn't fix the memory pressure or hanging or stalling, but you'll at least be able to log into the box still instead of dragging out a serial cable.


are you using systemd or some other init ?

both. freebsd has init scripts. ubuntu has systemd.

this guide should help make it less likely for sshd or your login shell to get killed: https://ianlpaterson.com/blog/oom-lockout-ssh-survival-harde...


> does anyone know why

In a nutshell, overcommit. It's more or less broken by design but it's also incredibly practical so pretty much everyone does it.

Couple that with the fact that it's difficult bordering on impossible to correctly determine the culprit. If you've got 16 GB RAM and the user launches 3 processes each of which attempts to use 8 GB who should you kill?


Notably windows doesn't use overcommit, and degrades much more gracefully under memory pressure. The biggest tradeoff is the amount of disk space consumed by a page file that also has to reserve space for unused pages that have been allocated but never been swapped in. On linux you can turn overcommit off, but there's too much software written around the assumption that overcommit is on


Is that still the case today? Notably (IIUC) overcommit is required for certain security measures. I believe it was chromium that I noticed mmaping somewhere north of 1 TB of memory on startup so that it can do (again IIUC) something akin to ASLR internally.


On Windows you can achieve something like manual overcommit by calling VirtualAlloc with just MEM_RESERVE. That gives you a continuous space in your process's virtual address space, without actually backing it with any physical pages. Kind of like what a malloc does on linux

But where linux would automagically back those pages once you use them, Windows requires you to actually ask for those pages to be backed by something (physical memory or page file) by calling VirtualAlloc with MEM_COMMIT on the range you actually want to use


At a glance that seems like a much more sensible design. I guess it's dead in the water for posix on account of fork being CoW? This is quite the rabbit hole. I wonder if programming languages ought to be designed in such a way to accommodate a preemptive signal indicating allocation failure in place of a page fault? Rather than malloc returning null or etc.


> I guess it's dead in the water for posix on account of fork being CoW?

If that's the only issue, it's avoidable for most use cases. Lots of processes that fork are doing fork/exec to run a helper program. If they know they will do that and that they will be a large process, it's often useful to setup a fork/exec helper in early application startup.

However, there are some applications that use CoW more intentionally. Lock -> fork -> (unlock in parent / persist coherent snapshot in child) is a common pattern; I believe redis uses thst pattern and I've seen it mentioned in discussions about MMO servers. I believe postgres uses fork and CoW for transaction isolation ... but postgres also runs on Windows so there must be another way or I don't understand.

For the persist case, you could imagine some sort of flag to fork to allow overcommit and maybe even to let CoW requests stall in the parent rather than fail... the child is expected to do its work and exit in a limited time.


Windows also does a neat trick Linux lacks: automatically adding more swap, up to a limit. Systems with loads of RAM barely lose any storage to swap, but once they do get hit, they can get many gigabytes of swap space without user interaction. I believe macOS does it too, of course.

I'm sure there are many reasons why Linux can't do that by default, but it's a real shame.


> Windows also does a neat trick Linux lacks: automatically adding more swap, up to a limit.

Given that one can have swap files, and can also use LVM LVs for swap, and given that userspace OOM killers that work way better than the built in one -for some workloads- exist, I see no reason why you couldn't have this on Linux. This comment [0] mentions a project that claims to do just that -and seems to use swapfiles to do it-, but I don't have any experience with it.

FWIW, I did find the README by the original author [1] far more informative than the one written by the new maintainer.

[0] <https://news.ycombinator.com/item?id=49656129>

[1] <https://pqxx.org/development/swapspace/>


I assume that's primarily because the stance of most distros would require something like that to be strictly opt-in. I don't know if it's possible to trigger a service based on overall swap usage? But given that the oom killer exists I don't see why it couldn't be trivially repurposed to add swap files on the fly.


> Couple that with the fact that it's difficult bordering on impossible to correctly determine the culprit. If you've got 16 GB RAM and the user launches 3 processes each of which attempts to use 8 GB who should you kill?

In all cases, yes, in some cases no, you can make some heuristics for common use cases

For example, if I have 3 process hogs, on desktop I'd rather have my dev containers be killed, than anything I'm using.

Or on server, I'd rather have anything else but SSH/VPN software killed, because that's needed to debug the problem.


With some workarounds, I've put chrome and slack into the same RAM-limited cgroup - no more whole system freezes. From my anecdotal evidence, this also answers the question "who should you kill" :)


OS are designed to fully exploit available resources, Linux tries its best before triggering an OOM kill.

I recommend using the earlyoom if you want more aggresive oom kill:

https://github.com/rfjakob/earlyoom

The README contains a lot of interesting information.


A strange behavior I sometimes run into with earlyoom is that I try to start up some buggy software of mine and it seemingly never starts.

It took a long evening to figure out that it gets earlyoom'd immediately because it tries to allocate too much. Previously the very familiar hitching and freezing was a very easy sign of what kind of issue I was dealing with


Just mentioning this in case it's helpful:

If you know ahead of time which programs / processes are at risk of unacceptably high memory usage, check out "ulimit".


It should eventually kill it. But having swap means it will try to use it as RAM which might delay/freeze system hard


It does not if you switch swapp off and use zram instead. I am typing right now on such a setup wityh 16 GiB ram and it occasionally, once a week or so, kills my firefox due to oom.

If you are you using disk swap - not sure why would if you have a SSD, but I once heard some justification for doing that - then install early OOM.



I always install earlyoom for that reason


Linux and the software running on it are generally very bad at handling out-of-memory situations.


This has finally been fixed in the latest Ubuntu version (26), it now force closes the culprit.


Ah nice. I was dealing with that in one of our environments where a security update ended up causing apt to use more memory than usual so the oom killer nuked our elasticsearch process to "free up some memory".

And since that happened on all nodes around the same time, it took out the entire cluster. If you are not familiar, random, uncontrolled node restarts in any kind of multi node database or search product are a great way to trigger outages. So, not great.

I've had quite a few encounters with the oomkiller killing processes that were important and didn't need killing. Or as I like to phrase it "killing the one reason this server exists".

These days the way to size a server is "have enough memory to run whatever you need running + at least half a GB for whatever apt might randomly demand at any point". And guess what, memory tends to be expensive in cloud environments so people tend to get vms with as little as half a GB of ram.


How did they fix it?


I couldn't find any information about that. Do you have something to back it up (news articles, changelog entries), or was it just your subjective experience?

not tangential at all...try setting up a swap partition


I had a 1GB Debian VM which started freezing (requiring a hard reboot) after a routine aptitude upgrade to apply security patches. It was indeed caused by low memory, but not out of memory as there was still enough swap space remaining.

The culprit turned out to be the kernel itself, and rolling back to a 6.1 series kernel made the problem go away. I see that Linus's love for vibe-coding is already paying dividends.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: