Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: The core of the problem is that systemd starts to use all physical RAM, then all swap, and then gets killed by the out-of-memory killer. This cycle takes about 20 seconds. The behavior can be triggered by many actions, e.g. trying to log-in graphically, or running systemctl daemon-reexec. Version-Release number of selected component (if applicable): systemd v241-10.git511646b.f30 How reproducible: the problem occurs every time when a triggering action is run. Most actions that somehow involve systemd seem to be triggering actions. Steps to Reproduce: type either systemctl daemon-reecxec or try to login graphically Actual results: system becomes slow due to swapping, then systemd gets killed by the oom killer Expected results: system should continue to work normally Additional info: The system used to work for some time without problems. The problem started suddenly.
Please attach strace to pid1: sudo strace -o logfile -p1 and run one of the triggering actions and attach the log file here.
Created attachment 1612691 [details] strace output
I have a similar problem. My strace log is attached above. If I ssh into the system, systemd starts running 100% CPU (a forked process by systemd as the pid is not 1) and allocating over 50% of the ram (64 GB on this computer). After a while, it goes away. Logging into the system is very slow, 20 sec. to get in. This is a NIS client if that matters.
Created attachment 1612799 [details] strace of the subprocess that consumes all CPU and all RAM
Created attachment 1612800 [details] Strace of PID 1
(In reply to Zbigniew Jędrzejewski-Szmek from comment #1) > Please attach strace to pid1: sudo strace -o logfile -p1 > and run one of the triggering actions and attach the log file here. I attached traces of both PID 1 and of the systemd subprocess that consumes all RAM and all CPU. I'm afraid none of them will be helpful, though.
I don't see anything bad in the pid1 trace in attachment 1612800 [details]. First brk() is called with 0x5562241bd000, the last with 0x55622419c000, so only about 135kB of heap are allocated. The other trace is empty. Jussi's trace does not contain any memory related syscalls. Are you using nss_nis?
Yes, I am using NIS.
Here is what it looks like from top: PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND 23031 root 20 0 52.2g 17.0g 4740 R 98.2 26.9 0:31.10 (systemd)
Yeah, that looks very much like the long-standing issue with nss_nis. *** This bug has been marked as a duplicate of bug 1705641 ***