Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1809259

Summary: console-operator container exited with code 255
Product: OpenShift Container Platform Reporter: Lalatendu Mohanty <lmohanty>
Component: Management ConsoleAssignee: bpeterse
Status: CLOSED NOTABUG QA Contact: Yadan Pei <yapei>
Severity: high Docs Contact:
Priority: unspecified    
Version: 4.3.zCC: aos-bugs, bpeterse, jokerman, rphillips, wking
Target Milestone: ---Keywords: Upgrades
Target Release: 4.5.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-03-11 17:37:11 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Lalatendu Mohanty 2020-03-02 18:02:51 UTC
Description of problem:
console-operator container exited with code 255 


Feb 28 23:47:37.566 E ns/openshift-console-operator pod/console-operator-5d786798c8-ths7n node/ip-10-0-151-113.us-west-1.compute.internal container=console-operator container exited with code 255 (Error): 

version: 14666 (18094)\nW0228 23:39:48.594564       1 reflector.go:289] github.com/openshift/client-go/config/informers/externalversions/factory.go:101: watch of *v1.Infrastructure ended with: too old resource version: 14507 (18097)\nW0228 23:40:28.859115       1 reflector.go:289] k8s.io/client-go/informers/factory.go:133: watch of *v1.Deployment ended with: too old resource version: 14202 (16411)\nW0228 23:40:29.933799       1 reflector.go:289] k8s.io/client-go/informers/factory.go:133: watch of *v1.ConfigMap ended with: too old resource version: 15407 (18172)\nW0228 23:41:04.939707       1 reflector.go:289] k8s.io/client-go/informers/factory.go:133: watch of *v1.Secret ended with: too old resource version: 15351 (16839)\nW0228 23:41:13.952004       1 reflector.go:289] k8s.io/client-go/informers/factory.go:133: watch of *v1.ConfigMap ended with: too old resource version: 15407 (18380)\nW0228 23:41:20.948095       1 reflector.go:289] k8s.io/client-go/informers/factory.go:133: watch of *v1.ConfigMap ended with: too old resource version: 15407 (18413)\nW0228 23:45:50.644820       1 reflector.go:289] github.com/openshift/client-go/config/informers/externalversions/factory.go:101: watch of *v1.ClusterOperator ended with: too old resource version: 21395 (21422)\nW0228 23:45:53.367641       1 reflector.go:289] github.com/openshift/client-go/config/informers/externalversions/factory.go:101: watch of *v1.ClusterOperator ended with: too old resource version: 21422 (21473)\nI0228 23:47:37.372785       1 observer_polling.go:78] Observed change: file:/var/run/configmaps/config/controller-config.yaml (current: "bee4a3cf6f26d7215852d2753f99948ec8356d257c972163bd17e64f4d73ed82", lastKnown: "41adc4f67c9486b2d108e39abc8b009b458b16fccaa6c860984b7a3410299dff")\nW0228 23:47:37.373484       1 builder.go:108] Restart triggered because of file /var/run/configmaps/config/controller-config.yaml was modified\nF0228 23:47:37.373674       1 leaderelection.go:66] leaderelection lost\nI0228 23:47:37.380316       1 controller.go:70] Shutting down Console\n

Version-Release number of selected component (if applicable):
4.3.nightly

How reproducible:

There are 106 occurrence of this issue in CI in last two days

Steps to Reproduce:
1.
2.
3.

Actual results:


Expected results:


Additional info:

Comment 2 Lalatendu Mohanty 2020-03-02 18:07:24 UTC
In the same CI job I also see following errors, not sure if we need a different bugzilla for this. 

Feb 28 23:53:34.598 E ns/openshift-sdn pod/sdn-k4266 node/ip-10-0-143-11.us-west-1.compute.internal container=sdn container exited with code 255 (Error): ing endpoints for openshift-multus/multus-admission-controller-monitor-service:metrics to [10.129.0.7:9091 10.130.0.6:9091]\nI0228 23:53:21.575938    

4767 proxy.go:334] hybrid proxy: mainProxy.syncProxyRules complete\nI0228 23:53:21.649475    4767 proxier.go:367] userspace proxy: processing 0 service events\nI0228 23:53:21.649494    4767 proxier.go:346] userspace syncProxyRules took 73.533087ms\nI0228 23:53:21.649504    4767 proxy.go:337] hybrid proxy: unidlingProxy.syncProxyRules complete\nI0228 23:53:21.649515    4767 proxy.go:331] hybrid proxy: syncProxyRules start\nI0228 23:53:21.828778    4767 proxy.go:334] hybrid proxy: mainProxy.syncProxyRules complete\nI0228 23:53:21.925144    4767 proxier.go:367] userspace proxy: processing 0 service events\nI0228 23:53:21.925170    4767 proxier.go:346] userspace syncProxyRules took 96.368179ms\nI0228 23:53:21.925185    4767 proxy.go:337] hybrid proxy: unidlingProxy.syncProxyRules complete\nI0228 23:53:23.668407    4767 roundrobin.go:310] LoadBalancerRR: Setting endpoints for openshift-sdn/sdn:metrics to [10.0.134.214:9101 10.0.140.119:9101 10.0.143.11:9101 10.0.151.113:9101 10.0.151.69:9101]\nI0228 23:53:23.668474    4767 roundrobin.go:240] Delete endpoint 10.0.131.136:9101 for service "openshift-sdn/sdn:metrics"\nI0228 23:53:23.668539    4767 proxy.go:331] hybrid proxy: syncProxyRules start\nI0228 23:53:23.714500    4767 ovs.go:169] Error executing ovs-ofctl: 2020-02-28T23:53:23Z|00001|vconn_stream|ERR|connection dropped mid-packet\novs-ofctl: OpenFlow receive failed (Protocol error)\nI0228 23:53:23.881775    4767 proxy.go:334] hybrid proxy: mainProxy.syncProxyRules complete\nI0228 23:53:23.969167    4767 proxier.go:367] userspace proxy: processing 0 service events\nI0228 23:53:23.969200    4767 proxier.go:346] userspace syncProxyRules took 87.394869ms\nI0228 23:53:23.969214    4767 proxy.go:337] hybrid proxy: unidlingProxy.syncProxyRules complete\nF0228 23:53:33.803527    4767 healthcheck.go:82] SDN healthcheck detected OVS server change, restarting: timed out waiting for the condition\n

Comment 3 Lalatendu Mohanty 2020-03-02 18:12:13 UTC
To help with impact analysis or updates we need to find answers to the following questions. It is fine if we do not answer some of these questions at this point of time, but we should try to get answers.

What symptoms (in Telemetry, Insights, etc.) does a cluster experiencing this bug exhibit?
What kind of clusters are impacted because of the bug? 
What cluster functionality is degraded while hitting the bug?
Does the upgrade complete?
What is the expected rate of the failure (%) for vulnerable clusters which attempt the update?
What is the observed rate of failure we see in CI?
Can this bug cause data loss? Data loss = API server data loss or CRD state information loss etc. 
Is it possible to recover the cluster from the bug?
Is recovery automatic without intervention?  I.e. is the condition transient?
Is recovery possible with the only intervention being 'oc adm upgrade …' to a new release image with a fix?
Is there a manual workaround that exists to recover from the bug? What are manual steps? 
Approximate time estimation for fixing this bug?
Is this a regression? From which version does this regress??
Does this bug cause user application downtime that would have otherwise not been triggered by an upgrade?

Comment 4 Lalatendu Mohanty 2020-03-02 18:16:21 UTC
Similar errors in other CI jobs

Feb 28 20:21:35.567 E ns/openshift-console pod/console-749b95899d-2vqtf node/ip-10-0-152-187.us-east-2.compute.internal container=console container exited with code 2 (Error): 2020/02/28 20:07:18 cmd/main: cookies are secure!\n2020/02/28 20:07:18 cmd/main: Binding to 0.0.0.0:8443...\n2020/02/28 20:07:18 cmd/main: using TLS\n
Feb 28 20:21:36.235 I ns/openshift-image-registry pod/node-ca-79tqn Pulling image "quay.io/openshift-release-dev/ocp-v4.0-art-dev@sha256:2b98ff923d24a240219a4047217e2b3b65b0ffb60b420394818259216e09d0fc" 

https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/19495

Comment 5 bpeterse 2020-03-11 13:39:22 UTC
The previous link has many, many issues like this:

```
 Feb 28 20:19:28.658 W ns/openshift-apiserver pod/apiserver-qxrbd MountVolume.SetUp failed for volume "config" : couldn't propagate object cache: timed out waiting for the condition 
 Feb 28 20:20:30.856 W ns/openshift-monitoring pod/node-exporter-tg27m MountVolume.SetUp failed for volume "node-exporter-tls" : couldn't propagate object cache: timed out waiting for the condition 
 Feb 28 20:37:39.846 W ns/openshift-monitoring pod/kube-state-metrics-78cb884b99-t4kv5 MountVolume.SetUp failed for volume "kube-state-metrics-token-hmtq5" : couldn't propagate object cache: timed out waiting for the condition 
 Feb 28 20:37:39.853 W ns/openshift-monitoring pod/kube-state-metrics-78cb884b99-t4kv5 MountVolume.SetUp failed for volume "kube-state-metrics-tls" : couldn't propagate object cache: timed out waiting for the condition 
 Feb 28 20:37:40.341 W ns/openshift-console pod/console-7688c4dc8f-gkstd MountVolume.SetUp failed for volume "console-oauth-config" : couldn't propagate object cache: timed out waiting for the condition 
 Feb 28 20:37:41.835 W ns/openshift-console pod/console-7688c4dc8f-gkstd MountVolume.SetUp failed for volume "trusted-ca-bundle" : couldn't propagate object cache: timed out waiting for the condition (2 times) 
 Feb 28 20:37:48.851 W ns/openshift-monitoring pod/prometheus-k8s-0 MountVolume.SetUp failed for volume "config" : couldn't propagate object cache: timed out waiting for the condition 
 Feb 28 20:37:48.859 W ns/openshift-monitoring pod/prometheus-k8s-0 MountVolume.SetUp failed for volume "prometheus-trusted-ca-bundle" : couldn't propagate object cache: timed out waiting for the condition 

```
Which im sure leads to the probe errors:

```
 Feb 28 20:37:54.973 W ns/openshift-marketplace pod/certified-operators-9cd6ddbc6-2mlpd Readiness probe errored: rpc error: code = Unknown desc = command error: command timed out, stdout: , stderr: , exit code -1 
 Feb 28 20:38:02.979 W ns/openshift-marketplace pod/community-operators-76574c78d4-5qhtz Liveness probe errored: rpc error: code = Unknown desc = command error: command timed out, stdout: , stderr: , exit code -1 (2 times) 
 Feb 28 20:45:31.125 W ns/openshift-marketplace pod/certified-operators-9cd6ddbc6-q6f9l Liveness probe failed: command timed out 
 Feb 28 20:48:20.611 W ns/openshift-marketplace pod/community-operators-dcc96ddc5-4k9kl Readiness probe failed: command timed out 

```

Comment 6 bpeterse 2020-03-11 14:40:19 UTC
Related to the first search link shared:

https://search.svc.ci.openshift.org/?search=console-operator+container+exited+with+code+255&maxAge=48h&context=2&type=all

I see from this job:

https://prow.svc.ci.openshift.org/view/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20910

many, many components failing for the same "container exited with code 255" reason. For example:

```
sdn container exited with code 255, 
configmap-cabundle-injector-controller container exited with code 255 
service-serving-cert-signer-controller container exited with code 255 
configmap-cabundle-injector-controller container exited with code 255
service-serving-cert-signer-controller container exited with code 255 
kube-apiserver-operator container exited with code 255
kube-scheduler-operator-container container exited with code 255
tuned container exited with code 255
kube-rbac-proxy container exited with code 255
dns-node-resolver container exited with code 255
```

I think there are several bugs here that are not console specific.

Comment 8 Ryan Phillips 2020-03-11 16:33:32 UTC
There are numerous pods that exit 255, verified in the worker logs. Prior to these exits networking is dumping the following log:

    11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:17.321Z|00153|connmgr|INFO|br0<->unix#230: 4 flow_mods in the last 0 s (4 deletes)
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:17.342Z|00154|bridge|INFO|bridge br0: deleted interface vethecffba5b on port 9
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:26.605Z|00026|jsonrpc|WARN|unix#166: receive error: Connection reset by peer
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: 2020-03-11T13:23:26.605Z|00027|reconnect|WARN|unix#166: connection dropped (Connection reset by peer)
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: ovs-vswitchd is not running.
Mar 11 13:26:08 ip-10-0-149-202 hyperkube[2125]: ovsdb-server is not running.

12:56:11 is the start of the not running errors... I see them continue through the end of the log at 13:40:43.

https://gcsweb-ci.svc.ci.openshift.org/gcs/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20909/artifacts/e2e-aws-upgrade/nodes/
https://storage.googleapis.com/origin-ci-test/logs/release-openshift-origin-installer-e2e-aws-upgrade/20909/artifacts/e2e-aws-upgrade/nodes/workers-journal

Comment 9 bpeterse 2020-03-11 17:37:11 UTC
Closing this out to replace with more specific bugs.  See the following:
- "(xyz) container exited with code 255" causing failures across many components in CI https://bugzilla.redhat.com/show_bug.cgi?id=1812555
- MountVolume.SetUp failed for volume (xyz) causing failures across many pods in CI https://bugzilla.redhat.com/show_bug.cgi?id=1812550

Comment 10 Red Hat Bugzilla 2023-09-14 05:53:42 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days