Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1932832

Summary: OCP Worker Node High Load crio, systemd and python Process and a lot of Zombies
Product: OpenShift Container Platform Reporter: mchebbi <mchebbi>
Component: NodeAssignee: Sascha Grunert <sgrunert>
Node sub component: CRI-O QA Contact: Weinan Liu <weinliu>
Status: CLOSED DUPLICATE Docs Contact:
Severity: medium    
Priority: medium CC: aos-bugs, miminar, sgrunert, weinliu
Version: 4.6Keywords: Reopened
Target Milestone: ---   
Target Release: 4.8.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-05-27 07:50:32 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description mchebbi@redhat.com 2021-02-25 11:50:30 UTC
Hello,
customer is experiencing a weired situation with ocp 4.6. he gets OCP Worker with High Load crio, systemd and python Process and a lot of Zombies.

url of gathered data: shorturl.at/bqyXY

[root@lnxdev-01 ~]# oc get nodes
NAME                      STATUS   ROLES        AGE    VERSION
kubmadev-01.bpa.bund.de   Ready    master       273d   v1.19.0+e49167a
kubmadev-02.bpa.bund.de   Ready    master       273d   v1.19.0+e49167a
kubmadev-03.bpa.bund.de   Ready    master       273d   v1.19.0+e49167a
kubwodev-01.bpa.bund.de   Ready    sdi,worker   273d   v1.19.0+e49167a
kubwodev-02.bpa.bund.de   Ready    sdi,worker   273d   v1.19.0+e49167a
kubwodev-03.bpa.bund.de   Ready    sdi,worker   273d   v1.19.0+e49167a
kubwodev-04.bpa.bund.de   Ready    worker       273d   v1.19.0+e49167a
kubwodev-05.bpa.bund.de   Ready    worker       194d   v1.19.0+e49167a
kubwodev-06.bpa.bund.de   Ready    worker       194d   v1.19.0+e49167a

Openshift Version 4.6.16. Only SAP Data Intelligence 3.1.1 runs on the Cluster. kubwodev-04-06 are OCS Nodes.

The response Time from the node kubwodev-03 is very high. A ssh session take 1,5 - 2 minutes:
[root@lnxdev-01 ~]# time ssh core.bund.de
Red Hat Enterprise Linux CoreOS 46.82.202101301821-0
  Part of OpenShift 4.6, RHCOS is a Kubernetes native operating system
  managed by the Machine Config Operator (`clusteroperator/machine-config`).

WARNING: Direct SSH access to machines is not recommended; instead,
make configuration changes via `machineconfig` objects:
  https://docs.openshift.com/container-platform/4.6/architecture/architecture-rhcos.html

---
Last login: Mon Feb 22 10:27:09 2021 from 10.40.150.60
Failed to list units: Connection timed out
[core@kubwodev-03 ~]$ logout
Connection to kubwodev-03.bpa.bund.de closed.
real    2m6.501s
user    0m0.015s
sys     0m0.010s
[root@lnxdev-01 ~]#

The system load is between 6 and 8.

on this node we have 410 - 420 zombies like this:

[core@kubwodev-03 ~]$ ps -waux | grep " Z "
core      657907  0.0  0.0      0     0 ?        Z    09:39   0:00 [sshd] <defunct>
sshd      663780  0.0  0.0      0     0 ?        Z    09:42   0:00 [sshd] <defunct>
sshd      664477  0.0  0.0      0     0 ?        Z    09:43   0:00 [sshd] <defunct>
core      714402  0.0  0.0  12788  1052 pts/4    S+   10:06   0:00 grep --color=auto  Z
root     2393754  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2393771  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2393780  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2393782  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2393908  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394022  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394027  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394357  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394398  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394399  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394548  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394674  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394811  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394871  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394891  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394908  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394917  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2394965  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>
root     2395153  0.0  0.0      0     0 ?        Z    Feb20   0:00 [conmon] <defunct>

and high load processes:
 PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
   1760 root      20   0 3899932 284484  35020 S 314.2   0.3  10819:46 crio
2522139 1972      20   0  565736  28432   9400 S  99.7   0.0   4562:29 python3.6
      1 root      20   0 1001604 761708   9020 R  99.0   0.8   3679:46 systemd
  82670 12000     20   0 6954496   4.1g 525692 S  36.8   4.3 601:06.50 hdbnameserver
   1810 root      20   0   12.2g   4.7g  65072 S  17.2   5.0   2126:30 kubelet
 229129 polkitd   20   0   11.5g 969876  22208 S  15.9   1.0  62:12.00 java
  86051 polkitd   20   0 8490076   4.2g 149548 S   5.6   4.4 405:00.00 prometheus
   3393 nfsnobo+  20   0  722488  49188  10596 S   4.6   0.0 174:23.54 node_exporter
    952 root      20   0  509000 328656 245028 S   3.3   0.3 118:05.80 systemd-journal
   1462 openvsw+  10 -10 1247220  97308  34684 S   1.3   0.1  98:00.46 ovs-vswitchd
 209869 polkitd   20   0  230204  61060   8028 S   1.0   0.1  21:25.43 python
 719937 8797      20   0 2359672  77840  47172 S   1.0   0.1   1:42.03 vsystem
1057827 core      20   0   71496   6852   4640 R   1.0   0.0   0:00.13 top

thanks in advance for your help and support.

Comment 1 Peter Hunt 2021-02-25 19:26:53 UTC
This seems like systemd getting overwhelmed. conmon double forks and reparents to systemd, and when systemd is under load, it takes a while for it to cleanup its zombie children. Once the load reduces, it should get around to it.

To speed up the process, the recommendation made in the case "The reservation may be increased..." is a good one.

Comment 2 Peter Hunt 2021-03-03 19:50:22 UTC
I don't think this is a Node bug. Increasing system reserved will help the situation. Please reopen if you disagree

Comment 8 Sascha Grunert 2021-04-28 08:38:34 UTC
I think this issue is related to https://bugzilla.redhat.com/show_bug.cgi?id=1952137, so I'll continue there with my investigations.

Comment 9 mchebbi@redhat.com 2021-04-28 19:49:31 UTC
(In reply to Sascha Grunert from comment #8)
> I think this issue is related to
> https://bugzilla.redhat.com/show_bug.cgi?id=1952137, so I'll continue there
> with my investigations.

Hello,
Thanks for your feedback.
I will wait for your updates.

Comment 10 mchebbi@redhat.com 2021-05-04 08:34:48 UTC
Hello,

Customer another outage this morning. it's a 2-3 week cycle until the problem resurfaces. he is expecting another outage in the second half on May. 

do you have any suggestions to prevent it from happening or fix it?


can we update the customer version to 2.6.28 to test the fix [1]

[1]-https://bugzilla.redhat.com/show_bug.cgi?id=1952137

waiting for your feedback.

Comment 11 Sascha Grunert 2021-05-04 09:10:19 UTC
(In reply to mchebbi from comment #10)
> Hello,
> 
> Customer another outage this morning. it's a 2-3 week cycle until the
> problem resurfaces. he is expecting another outage in the second half on
> May. 
> 
> do you have any suggestions to prevent it from happening or fix it?
> 
> 
> can we update the customer version to 2.6.28 to test the fix [1]
> 
> [1]-https://bugzilla.redhat.com/show_bug.cgi?id=1952137
> 
> waiting for your feedback.

Hey, yes please update to the latest 4.6.28 release if it becomes available.

Comment 12 Sascha Grunert 2021-05-10 09:21:33 UTC
4.6.28 should be released on Wed 2021-05-12

Comment 14 Weinan Liu 2021-05-13 07:04:59 UTC
(In reply to Sascha Grunert from comment #12)
> 4.6.28 should be released on Wed 2021-05-12

May I ask if this is a duplicate to https://bugzilla.redhat.com/show_bug.cgi?id=1952137? 
And I do not see any PR to the bz

The Target Release is 4.8

Comment 15 Sascha Grunert 2021-05-20 08:40:23 UTC
(In reply to Weinan Liu from comment #14)
> (In reply to Sascha Grunert from comment #12)
> > 4.6.28 should be released on Wed 2021-05-12
> 
> May I ask if this is a duplicate to
> https://bugzilla.redhat.com/show_bug.cgi?id=1952137?

Yes I think it's a duplicate.


> The Target Release is 4.8

Does this mean that the customer switched to 4.8 as next version?

Comment 17 Weinan Liu 2021-05-24 09:58:25 UTC
(In reply to Sascha Grunert from comment #15)
> (In reply to Weinan Liu from comment #14)
> > (In reply to Sascha Grunert from comment #12)
> > > 4.6.28 should be released on Wed 2021-05-12
> > 
> > May I ask if this is a duplicate to
> > https://bugzilla.redhat.com/show_bug.cgi?id=1952137?
> 
> Yes I think it's a duplicate.
> 
> 
> > The Target Release is 4.8
> 
> Does this mean that the customer switched to 4.8 as next version?

@Sascha, as per comment #16, I guess it not. I see 1952137 is targeting to 4.8 too, if it's duplicated issue, shall we make this one (1932832) duplicated to 1952137?

Comment 18 Sascha Grunert 2021-05-27 07:50:32 UTC
Yes, let's close this one and track the issue in 1952137

*** This bug has been marked as a duplicate of bug 1952137 ***