Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1812555

Summary: "(xyz) container exited with code 255" causing failures across many components in CI
Product: OpenShift Container Platform Reporter: bpeterse
Component: NodeAssignee: Ryan Phillips <rphillips>
Status: CLOSED NOTABUG QA Contact: Sunil Choudhary <schoudha>
Severity: medium Docs Contact:
Priority: unspecified    
Version: 4.3.zCC: aconstan, aos-bugs, dwalsh, jokerman, lmohanty, nagrawal, tsweeney
Target Milestone: ---Keywords: Upgrades
Target Release: 4.5.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-05-18 15:29:46 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description bpeterse 2020-03-11 15:12:24 UTC
Description of problem:

This is a CI flake, an example can be seen here:

https://search.svc.ci.openshift.org/?search=console-operator+container+exited+with+code+255&maxAge=48h&context=2&type=all

I see from this job:

https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910

many, many components failing for the same "container exited with code 255" reason. For example:

```
sdn container exited with code 255, 
configmap-cabundle-injector-controller container exited with code 255 
service-serving-cert-signer-controller container exited with code 255 
configmap-cabundle-injector-controller container exited with code 255
service-serving-cert-signer-controller container exited with code 255 
kube-apiserver-operator container exited with code 255
kube-scheduler-operator-container container exited with code 255
tuned container exited with code 255
kube-rbac-proxy container exited with code 255
dns-node-resolver container exited with code 255
```

How reproducible:

It appears to be a flake, thus causing build failures. 

Additional Info:

Originally reported as part of this bug: https://bugzilla.redhat.com/show_bug.cgi?id=1809259

Comment 3 Ryan Phillips 2020-03-12 14:01:12 UTC
`
ovs-vswitchd is not running
ovsdb-server is not running
`

in the worker logs right before the pods exits with 255 errors.

Comment 4 Ryan Phillips 2020-03-12 14:03:48 UTC
I suspect this is a networking issue and not a crio issue.

Comment 5 Alexander Constantinescu 2020-03-19 13:24:30 UTC
Hi 

I am going to send this back to the node team. 

Looking at the job: https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910

We do see all networking related pods stop, but they are being interrupted, ex: https://storage.googleapis.com/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910/artifacts/e2e-aws-upgrade/pods/openshift-sdn_sdn-92xh2_sdn_previous.log

The final line states: "interrupt: Gracefully shutting down ..."

Which comes from here: https://github.com/openshift/sdn/blob/master/pkg/openshift-sdn/cmd.go#L59

I have looked at the master-logs (that SDN pod was running on ip-10-0-128-41.us-west-1.compute.internal), and I found this at the moment the SDN receives the interrupt:

Master logs: https://storage.googleapis.com/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910/artifacts/e2e-aws-upgrade/nodes/masters-journal

Mar 11 13:44:25 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:25.734222    1976 prober.go:125] Liveness probe for "sdn-92xh2_openshift-sdn(cbd87496-6398-11ea-bb8b-062f2d5e4cd9):sdn" succeeded
Mar 11 13:44:25 ip-10-0-128-41 rpm-ostree[1123]: Txn UpdateDeployment on /org/projectatomic/rpmostree1/rhcos successful
Mar 11 13:44:25 ip-10-0-128-41 rpm-ostree[1123]: client(id:cli dbus:1.1522 unit:machine-config-daemon-host.service uid:0) vanished; remaining=0
Mar 11 13:44:25 ip-10-0-128-41 rpm-ostree[1123]: In idle state; will auto-exit in 63 seconds
Mar 11 13:44:26 ip-10-0-128-41 podman[4221]: 2020-03-11 13:44:26.098084971 +0000 UTC m=+0.084018203 container remove fb0a8cb0063921477a94f49f7f67215cabf146b2c979b7e9444c3c6ebdd6462c (image=quay.io/openshift-release-dev/ocp-v4.0-art-dev@sha256:219ab1660d30a18e7ac8b8ea796bd2eb915bbc3421ca8f7b2cde28b3a40c51dc, name=ostree-container-pivot-fee1e033-3854-4772-8b1b-8e9567933889)
Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.106722    1976 manager.go:1068] Destroyed container: "/system.slice/machine-config-daemon-host.service" (aliases: [], namespace: "")
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Started Machine Config Daemon Initial.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: machine-config-daemon-host.service: Consumed 37.334s CPU time
Mar 11 13:44:26 ip-10-0-128-41 logger[4235]: rendered-master-1c9b7441dff267bacfe0f47bdf964b6e
Mar 11 13:44:26 ip-10-0-128-41 root[4236]: machine-config-daemon[40883]: initiating reboot: Node will reboot into config rendered-master-1c9b7441dff267bacfe0f47bdf964b6e
Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.124760    1976 factory.go:116] Using factory "raw" for container "/system.slice/machine-config-daemon-reboot.service"
Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.125133    1976 manager.go:1011] Added container: "/system.slice/machine-config-daemon-reboot.service" (aliases: [], namespace: "")
Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.125834    1976 container.go:464] Start housekeeping for container "/system.slice/machine-config-daemon-reboot.service"
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Started machine-config-daemon: Node will reboot into config rendered-master-1c9b7441dff267bacfe0f47bdf964b6e.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping Kubernetes Kubelet...
Mar 11 13:44:26 ip-10-0-128-41 sh[4238]: Warning: The unit file, source configuration file or drop-ins of kubelet.service changed on disk. Run 'systemctl daemon-reload' to reload units.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopped Kubernetes Kubelet.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: kubelet.service: Consumed 7min 11.505s CPU time
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container cbde811c33c80f381d7750c261d038d394358660a171898c8a7a99025291934c.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-367f34bb3b2772b54d9837fa7dca1ed774cc0c65c072de98d097ce1aca3aa71d.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container ca077a21a4da6c82ba03f98d62d258a7689a60ba7088977ecfc5757cee46b3de.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-4b098e87758071a2b6ecdf7807c0fad477f637ca00cb84afd4903ce494bf2b04.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-64d93badd15526cad7f3d284792fe39a545b9d404f0eb458683df1ccb15fd792.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping NFS status monitor for NFSv2/3 locking....
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-632966634c057bed450b8fecc450fd2d27ca98de9b17b1b53eec9bc6049a4273.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-0f630627b1cb3571e37f4594eda57061eb83ecc9ba308cd7b899ab93189f067d.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-2a1cc038ff192a74f983a3892615ef1a4b942224444ca62535651723f7ee1aad.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-ce20f60f80404ece134bda543c295f95aa10583634105f805e2bdb1d301cd5ae.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Closed LVM2 poll daemon socket.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container 936f822e6584047f2a9e7172ee78d1d4b561453c506a648fc3c18a9fdfb946f6.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container 6127f38a34f09b3d43c068016b113575ce823384bd3fac63ec210af00a562eb1.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-0aaf31bf411b13e25dc4f612ddd460750d72ba46f03dbea2e3d2e2b5cb03f51b.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopped target Multi-User System.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping irqbalance daemon...
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-62e75a7f1d16e8c229b1c55b75a55dedd7fe6aa6e11a206f7bd43426f3e51ab7.scope.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container 4b098e87758071a2b6ecdf7807c0fad477f637ca00cb84afd4903ce494bf2b04.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping Restore /run/initramfs on shutdown...
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container a34f6b7aafa81a66c03823aebbbb341dc4c86a3cea02341c7ec92118ff0df016.
Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopped target Timers.

The first line show a successful SDN readinessProbe upon which the machines reboots and stops the kubelet which is the origin of the interrupt seen in the SDN.

We should thus find out why the kubelet and node reboots and kills all pods. 

-Alex

Comment 6 Ryan Phillips 2020-05-14 19:17:31 UTC
"interrupt: Gracefully shutting down ..." is showing up in search ci for failing runs 0.36%. I believe we can close this issue.

Comment 7 Red Hat Bugzilla 2023-09-14 05:54:11 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days