Bug 1746844 - systemd consumes all RAM until killed by OOM killer
Summary: systemd consumes all RAM until killed by OOM killer
Keywords:
Status: CLOSED DUPLICATE of bug 1705641
Alias: None
Product: Fedora
Classification: Fedora
Component: systemd
Version: 30
Hardware: x86_64
OS: Linux
unspecified
unspecified
Target Milestone: ---
Assignee: systemd-maint
QA Contact: Fedora Extras Quality Assurance
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2019-08-29 11:37 UTC by Johannes Niediek
Modified: 2019-10-24 19:22 UTC (History)
8 users (show)

Fixed In Version:
Clone Of:
Environment:
Last Closed: 2019-10-24 19:22:26 UTC
Type: Bug
Embargoed:


Attachments (Terms of Use)
strace output (95.96 KB, text/plain)
2019-09-07 16:03 UTC, Jussi Eloranta
no flags Details
strace of the subprocess that consumes all CPU and all RAM (26 bytes, text/plain)
2019-09-08 08:28 UTC, Johannes Niediek
no flags Details
Strace of PID 1 (2.97 MB, application/zip)
2019-09-08 08:29 UTC, Johannes Niediek
no flags Details

Description Johannes Niediek 2019-08-29 11:37:51 UTC
Description of problem: The core of the problem is that systemd starts to use all physical RAM, then all swap, and then gets killed by the out-of-memory killer. This cycle takes about 20 seconds. The behavior can be triggered by many actions, e.g. trying to log-in graphically, or running systemctl daemon-reexec.


Version-Release number of selected component (if applicable):
systemd v241-10.git511646b.f30

How reproducible:
the problem occurs every time when a triggering action is run. Most actions that somehow involve systemd seem to be triggering actions.

Steps to Reproduce:
type either systemctl daemon-reecxec or try to login graphically

Actual results:
system becomes slow due to swapping, then systemd gets killed by the oom killer

Expected results:
system should continue to work normally


Additional info:
The system used to work for some time without problems. The problem started suddenly.

Comment 1 Zbigniew Jędrzejewski-Szmek 2019-09-03 16:14:10 UTC
Please attach strace to pid1: sudo strace -o logfile -p1
and run one of the triggering actions and attach the log file here.

Comment 2 Jussi Eloranta 2019-09-07 16:03:34 UTC
Created attachment 1612691 [details]
strace output

Comment 3 Jussi Eloranta 2019-09-07 16:06:14 UTC
I have a similar problem. My strace log is attached above.

If I ssh into the system, systemd starts running 100% CPU (a forked process by systemd as the pid is not 1) and allocating over 50% of the ram (64 GB on this computer). After a while, it goes away. Logging into the system is very slow, 20 sec. to get in. This is a NIS client if that matters.

Comment 4 Johannes Niediek 2019-09-08 08:28:45 UTC
Created attachment 1612799 [details]
strace of the subprocess that consumes all CPU and all RAM

Comment 5 Johannes Niediek 2019-09-08 08:29:59 UTC
Created attachment 1612800 [details]
Strace of PID 1

Comment 6 Johannes Niediek 2019-09-08 08:30:52 UTC
(In reply to Zbigniew Jędrzejewski-Szmek from comment #1)
> Please attach strace to pid1: sudo strace -o logfile -p1
> and run one of the triggering actions and attach the log file here.

I attached traces of both PID 1 and of the systemd subprocess that consumes all RAM and all CPU. I'm afraid none of them will be helpful, though.

Comment 7 Zbigniew Jędrzejewski-Szmek 2019-10-24 12:35:11 UTC
I don't see anything bad in the pid1 trace in attachment 1612800 [details]. First brk() is called with 0x5562241bd000,
the last with 0x55622419c000, so only about 135kB of heap are allocated. The other trace is empty.
Jussi's trace does not contain any memory related syscalls.

Are you using nss_nis?

Comment 8 Jussi Eloranta 2019-10-24 14:51:13 UTC
Yes, I am using NIS.

Comment 9 Jussi Eloranta 2019-10-24 14:54:44 UTC
Here is what it looks like from top:

  PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND    
23031 root      20   0   52.2g  17.0g   4740 R  98.2  26.9   0:31.10 (systemd)

Comment 10 Zbigniew Jędrzejewski-Szmek 2019-10-24 19:22:26 UTC
Yeah, that looks very much like the long-standing issue with nss_nis.

*** This bug has been marked as a duplicate of bug 1705641 ***


Note You need to log in before you can comment on or make changes to this bug.