Bug 1812555
| Summary: | "(xyz) container exited with code 255" causing failures across many components in CI | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | bpeterse |
| Component: | Node | Assignee: | Ryan Phillips <rphillips> |
| Status: | CLOSED NOTABUG | QA Contact: | Sunil Choudhary <schoudha> |
| Severity: | medium | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 4.3.z | CC: | aconstan, aos-bugs, dwalsh, jokerman, lmohanty, nagrawal, tsweeney |
| Target Milestone: | --- | Keywords: | Upgrades |
| Target Release: | 4.5.0 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2020-05-18 15:29:46 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
bpeterse
2020-03-11 15:12:24 UTC
Additional related builds: - https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20915 - https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20909 show the same, many components with "container exited with code 255" ` ovs-vswitchd is not running ovsdb-server is not running ` in the worker logs right before the pods exits with 255 errors. I suspect this is a networking issue and not a crio issue. Hi I am going to send this back to the node team. Looking at the job: https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910 We do see all networking related pods stop, but they are being interrupted, ex: https://storage.googleapis.com/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910/artifacts/e2e-aws-upgrade/pods/openshift-sdn_sdn-92xh2_sdn_previous.log The final line states: "interrupt: Gracefully shutting down ..." Which comes from here: https://github.com/openshift/sdn/blob/master/pkg/openshift-sdn/cmd.go#L59 I have looked at the master-logs (that SDN pod was running on ip-10-0-128-41.us-west-1.compute.internal), and I found this at the moment the SDN receives the interrupt: Master logs: https://storage.googleapis.com/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910/artifacts/e2e-aws-upgrade/nodes/masters-journal Mar 11 13:44:25 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:25.734222 1976 prober.go:125] Liveness probe for "sdn-92xh2_openshift-sdn(cbd87496-6398-11ea-bb8b-062f2d5e4cd9):sdn" succeeded Mar 11 13:44:25 ip-10-0-128-41 rpm-ostree[1123]: Txn UpdateDeployment on /org/projectatomic/rpmostree1/rhcos successful Mar 11 13:44:25 ip-10-0-128-41 rpm-ostree[1123]: client(id:cli dbus:1.1522 unit:machine-config-daemon-host.service uid:0) vanished; remaining=0 Mar 11 13:44:25 ip-10-0-128-41 rpm-ostree[1123]: In idle state; will auto-exit in 63 seconds Mar 11 13:44:26 ip-10-0-128-41 podman[4221]: 2020-03-11 13:44:26.098084971 +0000 UTC m=+0.084018203 container remove fb0a8cb0063921477a94f49f7f67215cabf146b2c979b7e9444c3c6ebdd6462c (image=quay.io/openshift-release-dev/ocp-v4.0-art-dev@sha256:219ab1660d30a18e7ac8b8ea796bd2eb915bbc3421ca8f7b2cde28b3a40c51dc, name=ostree-container-pivot-fee1e033-3854-4772-8b1b-8e9567933889) Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.106722 1976 manager.go:1068] Destroyed container: "/system.slice/machine-config-daemon-host.service" (aliases: [], namespace: "") Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Started Machine Config Daemon Initial. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: machine-config-daemon-host.service: Consumed 37.334s CPU time Mar 11 13:44:26 ip-10-0-128-41 logger[4235]: rendered-master-1c9b7441dff267bacfe0f47bdf964b6e Mar 11 13:44:26 ip-10-0-128-41 root[4236]: machine-config-daemon[40883]: initiating reboot: Node will reboot into config rendered-master-1c9b7441dff267bacfe0f47bdf964b6e Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.124760 1976 factory.go:116] Using factory "raw" for container "/system.slice/machine-config-daemon-reboot.service" Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.125133 1976 manager.go:1011] Added container: "/system.slice/machine-config-daemon-reboot.service" (aliases: [], namespace: "") Mar 11 13:44:26 ip-10-0-128-41 hyperkube[1976]: I0311 13:44:26.125834 1976 container.go:464] Start housekeeping for container "/system.slice/machine-config-daemon-reboot.service" Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Started machine-config-daemon: Node will reboot into config rendered-master-1c9b7441dff267bacfe0f47bdf964b6e. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping Kubernetes Kubelet... Mar 11 13:44:26 ip-10-0-128-41 sh[4238]: Warning: The unit file, source configuration file or drop-ins of kubelet.service changed on disk. Run 'systemctl daemon-reload' to reload units. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopped Kubernetes Kubelet. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: kubelet.service: Consumed 7min 11.505s CPU time Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container cbde811c33c80f381d7750c261d038d394358660a171898c8a7a99025291934c. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-367f34bb3b2772b54d9837fa7dca1ed774cc0c65c072de98d097ce1aca3aa71d.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container ca077a21a4da6c82ba03f98d62d258a7689a60ba7088977ecfc5757cee46b3de. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-4b098e87758071a2b6ecdf7807c0fad477f637ca00cb84afd4903ce494bf2b04.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-64d93badd15526cad7f3d284792fe39a545b9d404f0eb458683df1ccb15fd792.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping NFS status monitor for NFSv2/3 locking.... Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-632966634c057bed450b8fecc450fd2d27ca98de9b17b1b53eec9bc6049a4273.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-0f630627b1cb3571e37f4594eda57061eb83ecc9ba308cd7b899ab93189f067d.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-2a1cc038ff192a74f983a3892615ef1a4b942224444ca62535651723f7ee1aad.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-ce20f60f80404ece134bda543c295f95aa10583634105f805e2bdb1d301cd5ae.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Closed LVM2 poll daemon socket. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container 936f822e6584047f2a9e7172ee78d1d4b561453c506a648fc3c18a9fdfb946f6. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container 6127f38a34f09b3d43c068016b113575ce823384bd3fac63ec210af00a562eb1. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-0aaf31bf411b13e25dc4f612ddd460750d72ba46f03dbea2e3d2e2b5cb03f51b.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopped target Multi-User System. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping irqbalance daemon... Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping crio-conmon-62e75a7f1d16e8c229b1c55b75a55dedd7fe6aa6e11a206f7bd43426f3e51ab7.scope. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container 4b098e87758071a2b6ecdf7807c0fad477f637ca00cb84afd4903ce494bf2b04. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping Restore /run/initramfs on shutdown... Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopping libcontainer container a34f6b7aafa81a66c03823aebbbb341dc4c86a3cea02341c7ec92118ff0df016. Mar 11 13:44:26 ip-10-0-128-41 systemd[1]: Stopped target Timers. The first line show a successful SDN readinessProbe upon which the machines reboots and stops the kubelet which is the origin of the interrupt seen in the SDN. We should thus find out why the kubelet and node reboots and kills all pods. -Alex "interrupt: Gracefully shutting down ..." is showing up in search ci for failing runs 0.36%. I believe we can close this issue. The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days |