Bug 1809259
| Summary: | console-operator container exited with code 255 | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | Lalatendu Mohanty <lmohanty> |
| Component: | Management Console | Assignee: | bpeterse |
| Status: | CLOSED NOTABUG | QA Contact: | Yadan Pei <yapei> |
| Severity: | high | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 4.3.z | CC: | aos-bugs, bpeterse, jokerman, rphillips, wking |
| Target Milestone: | --- | Keywords: | Upgrades |
| Target Release: | 4.5.0 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2020-03-11 17:37:11 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Lalatendu Mohanty
2020-03-02 18:02:51 UTC
In the same CI job I also see following errors, not sure if we need a different bugzilla for this. Feb 28 23:53:34.598 E ns/openshift-sdn pod/sdn-k4266 node/ip-10-0-143-11.us-west-1.compute.internal container=sdn container exited with code 255 (Error): ing endpoints for openshift-multus/multus-admission-controller-monitor-service:metrics to [10.129.0.7:9091 10.130.0.6:9091]\nI0228 23:53:21.575938 4767 proxy.go:334] hybrid proxy: mainProxy.syncProxyRules complete\nI0228 23:53:21.649475 4767 proxier.go:367] userspace proxy: processing 0 service events\nI0228 23:53:21.649494 4767 proxier.go:346] userspace syncProxyRules took 73.533087ms\nI0228 23:53:21.649504 4767 proxy.go:337] hybrid proxy: unidlingProxy.syncProxyRules complete\nI0228 23:53:21.649515 4767 proxy.go:331] hybrid proxy: syncProxyRules start\nI0228 23:53:21.828778 4767 proxy.go:334] hybrid proxy: mainProxy.syncProxyRules complete\nI0228 23:53:21.925144 4767 proxier.go:367] userspace proxy: processing 0 service events\nI0228 23:53:21.925170 4767 proxier.go:346] userspace syncProxyRules took 96.368179ms\nI0228 23:53:21.925185 4767 proxy.go:337] hybrid proxy: unidlingProxy.syncProxyRules complete\nI0228 23:53:23.668407 4767 roundrobin.go:310] LoadBalancerRR: Setting endpoints for openshift-sdn/sdn:metrics to [10.0.134.214:9101 10.0.140.119:9101 10.0.143.11:9101 10.0.151.113:9101 10.0.151.69:9101]\nI0228 23:53:23.668474 4767 roundrobin.go:240] Delete endpoint 10.0.131.136:9101 for service "openshift-sdn/sdn:metrics"\nI0228 23:53:23.668539 4767 proxy.go:331] hybrid proxy: syncProxyRules start\nI0228 23:53:23.714500 4767 ovs.go:169] Error executing ovs-ofctl: 2020-02-28T23:53:23Z|00001|vconn_stream|ERR|connection dropped mid-packet\novs-ofctl: OpenFlow receive failed (Protocol error)\nI0228 23:53:23.881775 4767 proxy.go:334] hybrid proxy: mainProxy.syncProxyRules complete\nI0228 23:53:23.969167 4767 proxier.go:367] userspace proxy: processing 0 service events\nI0228 23:53:23.969200 4767 proxier.go:346] userspace syncProxyRules took 87.394869ms\nI0228 23:53:23.969214 4767 proxy.go:337] hybrid proxy: unidlingProxy.syncProxyRules complete\nF0228 23:53:33.803527 4767 healthcheck.go:82] SDN healthcheck detected OVS server change, restarting: timed out waiting for the condition\n To help with impact analysis or updates we need to find answers to the following questions. It is fine if we do not answer some of these questions at this point of time, but we should try to get answers. What symptoms (in Telemetry, Insights, etc.) does a cluster experiencing this bug exhibit? What kind of clusters are impacted because of the bug? What cluster functionality is degraded while hitting the bug? Does the upgrade complete? What is the expected rate of the failure (%) for vulnerable clusters which attempt the update? What is the observed rate of failure we see in CI? Can this bug cause data loss? Data loss = API server data loss or CRD state information loss etc. Is it possible to recover the cluster from the bug? Is recovery automatic without intervention? I.e. is the condition transient? Is recovery possible with the only intervention being 'oc adm upgrade …' to a new release image with a fix? Is there a manual workaround that exists to recover from the bug? What are manual steps? Approximate time estimation for fixing this bug? Is this a regression? From which version does this regress?? Does this bug cause user application downtime that would have otherwise not been triggered by an upgrade? Similar errors in other CI jobs Feb 28 20:21:35.567 E ns/openshift-console pod/console-749b95899d-2vqtf node/ip-10-0-152-187.us-east-2.compute.internal container=console container exited with code 2 (Error): 2020/02/28 20:07:18 cmd/main: cookies are secure!\n2020/02/28 20:07:18 cmd/main: Binding to 0.0.0.0:8443...\n2020/02/28 20:07:18 cmd/main: using TLS\n Feb 28 20:21:36.235 I ns/openshift-image-registry pod/node-ca-79tqn Pulling image "quay.io/openshift-release-dev/ocp-v4.0-art-dev@sha256:2b98ff923d24a240219a4047217e2b3b65b0ffb60b420394818259216e09d0fc" https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/19495 The previous link has many, many issues like this: ``` Feb 28 20:19:28.658 W ns/openshift-apiserver pod/apiserver-qxrbd MountVolume.SetUp failed for volume "config" : couldn't propagate object cache: timed out waiting for the condition Feb 28 20:20:30.856 W ns/openshift-monitoring pod/node-exporter-tg27m MountVolume.SetUp failed for volume "node-exporter-tls" : couldn't propagate object cache: timed out waiting for the condition Feb 28 20:37:39.846 W ns/openshift-monitoring pod/kube-state-metrics-78cb884b99-t4kv5 MountVolume.SetUp failed for volume "kube-state-metrics-token-hmtq5" : couldn't propagate object cache: timed out waiting for the condition Feb 28 20:37:39.853 W ns/openshift-monitoring pod/kube-state-metrics-78cb884b99-t4kv5 MountVolume.SetUp failed for volume "kube-state-metrics-tls" : couldn't propagate object cache: timed out waiting for the condition Feb 28 20:37:40.341 W ns/openshift-console pod/console-7688c4dc8f-gkstd MountVolume.SetUp failed for volume "console-oauth-config" : couldn't propagate object cache: timed out waiting for the condition Feb 28 20:37:41.835 W ns/openshift-console pod/console-7688c4dc8f-gkstd MountVolume.SetUp failed for volume "trusted-ca-bundle" : couldn't propagate object cache: timed out waiting for the condition (2 times) Feb 28 20:37:48.851 W ns/openshift-monitoring pod/prometheus-k8s-0 MountVolume.SetUp failed for volume "config" : couldn't propagate object cache: timed out waiting for the condition Feb 28 20:37:48.859 W ns/openshift-monitoring pod/prometheus-k8s-0 MountVolume.SetUp failed for volume "prometheus-trusted-ca-bundle" : couldn't propagate object cache: timed out waiting for the condition ``` Which im sure leads to the probe errors: ``` Feb 28 20:37:54.973 W ns/openshift-marketplace pod/certified-operators-9cd6ddbc6-2mlpd Readiness probe errored: rpc error: code = Unknown desc = command error: command timed out, stdout: , stderr: , exit code -1 Feb 28 20:38:02.979 W ns/openshift-marketplace pod/community-operators-76574c78d4-5qhtz Liveness probe errored: rpc error: code = Unknown desc = command error: command timed out, stdout: , stderr: , exit code -1 (2 times) Feb 28 20:45:31.125 W ns/openshift-marketplace pod/certified-operators-9cd6ddbc6-q6f9l Liveness probe failed: command timed out Feb 28 20:48:20.611 W ns/openshift-marketplace pod/community-operators-dcc96ddc5-4k9kl Readiness probe failed: command timed out ``` Related to the first search link shared: https://search.svc.ci.openshift.org/?search=console-operator+container+exited+with+code+255&maxAge=48h&context=2&type=all I see from this job: https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910 many, many components failing for the same "container exited with code 255" reason. For example: ``` sdn container exited with code 255, configmap-cabundle-injector-controller container exited with code 255 service-serving-cert-signer-controller container exited with code 255 configmap-cabundle-injector-controller container exited with code 255 service-serving-cert-signer-controller container exited with code 255 kube-apiserver-operator container exited with code 255 kube-scheduler-operator-container container exited with code 255 tuned container exited with code 255 kube-rbac-proxy container exited with code 255 dns-node-resolver container exited with code 255 ``` I think there are several bugs here that are not console specific. Additional related builds: - https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20915 - https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20909 show the same, many components with "container exited with code 255" There are numerous pods that exit 255, verified in the worker logs. Prior to these exits networking is dumping the following log:
11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:17.321Z|00153|connmgr|INFO|br0<->unix#230: 4 flow_mods in the last 0 s (4 deletes)
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:17.342Z|00154|bridge|INFO|bridge br0: deleted interface vethecffba5b on port 9
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:26.605Z|00026|jsonrpc|WARN|unix#166: receive error: Connection reset by peer
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:26.605Z|00027|reconnect|WARN|unix#166: connection dropped (Connection reset by peer)
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: ovs-vswitchd is not running.
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: ovsdb-server is not running.
12:56:11 is the start of the not running errors... I see them continue through the end of the log at 13:40:43.
https://gcsweb-ci.svc.ci.openshift.org/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20909/artifacts/e2e-aws-upgrade/nodes/
https://storage.googleapis.com/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20909/artifacts/e2e-aws-upgrade/nodes/workers-journal
Closing this out to replace with more specific bugs. See the following: - "(xyz) container exited with code 255" causing failures across many components in CI https://bugzilla.redhat.com/show_bug.cgi?id=1812555 - MountVolume.SetUp failed for volume (xyz) causing failures across many pods in CI https://bugzilla.redhat.com/show_bug.cgi?id=1812550 The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days |