Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem (please be detailed as possible and provide log snippests): From operator log: http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-ai3c33-t1/jnk-ai3c33-t1_20200630T113618/logs/failed_testcase_ocs_logs_1593517309/test_deployment_ocs_logs/ocs_must_gather/quay-io-rhceph-dev-ocs-must-gather-sha256-4db6344cc517f57df8ca261314aeba25d5348d3f033a14e90267b75b905efc8a/ceph/namespaces/openshift-storage/pods/ocs-operator-7fdf85b7b4-jhm9t/ocs-operator/ocs-operator/logs/current.log I see some errors: 2020-06-30T13:02:45.731471005Z {"level":"error","ts":"2020-06-30T13:02:45.731Z","logger":"controller_storagecluster","msg":"Failed to update status","Request.Namespace":"openshift-storage","Request.Name":"ocs-storagecluster","error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again","stacktrace":"github.com/go-logr/zapr.(*zapLogger).Error\n\t/go/src/github.com/go-logr/zapr/zapr.go:128\ngithub.com/openshift/ocs-operator/pkg/controller/storagecluster.(*ReconcileStorageCluster).Reconcile\n\t/go/src/github.com/openshift/ocs-operator/pkg/controller/storagecluster/reconcile.go:366\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:216\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:192\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:171\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker-fm\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:157\nk8s.io/apimachinery/pkg/util/wait.JitterUntil.func1\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:152\nk8s.io/apimachinery/pkg/util/wait.JitterUntil\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:153\nk8s.io/apimachinery/pkg/util/wait.Until\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:88"} 2020-06-30T13:02:45.731550981Z {"level":"error","ts":"2020-06-30T13:02:45.731Z","logger":"controller-runtime.controller","msg":"Reconciler error","controller":"storagecluster-controller","request":"openshift-storage/ocs-storagecluster","error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again","stacktrace":"github.com/go-logr/zapr.(*zapLogger).Error\n\t/go/src/github.com/go-logr/zapr/zapr.go:128\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:218\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:192\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:171\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker-fm\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:157\nk8s.io/apimachinery/pkg/util/wait.JitterUntil.func1\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:152\nk8s.io/apimachinery/pkg/util/wait.JitterUntil\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:153\nk8s.io/apimachinery/pkg/util/wait.Until\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:88"} I see operator pod is in this state: ocs-operator-7fdf85b7b4-jhm9t 0/1 Running 0 47m 10.129.2.12 ip-10-0-138-54.us-east-2.compute.internal <none> <none> Version of all relevant components (if applicable): OCP: 4.3.0-0.nightly-2020-06-29-084049 OCS: 4.4.1 - 4.4.1-465.ci RC2 build Does this issue impact your ability to continue to work with the product (please explain in detail what is the user impact)? YES Is there any workaround available to the best of your knowledge? NO Rate from 1 - 5 the complexity of the scenario you performed that caused this bug (1 - very simple, 5 - very complex)? 1 Can this issue reproducible? Trying now: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9312/ Can this issue reproduce from the UI? Haven't tried If this is a regression, please provide more details to justify this: Yes Steps to Reproduce: 1. Install OCP 4.3 nightly 2. Install OCS 4.4.1 build Actual results: Installation of OCS is not completed Expected results: OCS operator to finish installation Additional info: Jenkins job: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9306/consoleFull Must gather logs: http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-ai3c33-t1/jnk-ai3c33-t1_20200630T113618/logs/failed_testcase_ocs_logs_1593517309/test_deployment_ocs_logs/
Update here - second time I was able to deploy OCS on top of OCP 4.3 - but then the deployment failed on creating credentials on S3 as we hit the limit of 400 buckets under our account. Execution here: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9312/console When I connected to the cluster I see that it finished the deployment here: $ oc get csv -n openshift-storage NAME DISPLAY VERSION REPLACES PHASE lib-bucket-provisioner.v1.0.0 lib-bucket-provisioner 1.0.0 Succeeded ocs-operator.v4.4.1-465.ci OpenShift Container Storage 4.4.1-465.ci Succeeded We are cleaning the buckets and will re-run once more: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-trigger-aws-ipi-3az-rhcos-3m-3w-tier1/45
I see that other job we scheduled passed deployment without any issue here: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9319/console As well we got results from tier1. Maybe it was some temporary issue but I would still like to see if someone can see from the logs what is the root cause of the failure from the first deployment. @Travis, can you please have a look at operator logs if you see why the operator didn't get to succeeded status in first attempt? What the error I pasted in the 1st comment: "error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again is about? Thanks
@Petr, so it seems to be intermitten, not consistent? @José could you have a look to see if anything stands out?
Yes Michael 1st out of 3 runs we saw this issue. But worth to at least try to check the logs and maybe Jose will see what was the root cause of first failure. Thanks
I'm seeing nothing immediately obvious in the ocs-operator logs. Supposedly it thinks the CephCluster isn't reporting status but the CephCluster CR says it's HEALTH_OK. I found the following in the Rook-Ceph log: E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-0" E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-1" E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2" E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"] But the ceph osd tree command shows a good result: ID CLASS WEIGHT TYPE NAME STATUS REWEIGHT PRI-AFF -1 6.00000 root default -5 6.00000 region us-east-2 -10 2.00000 zone us-east-2a -9 2.00000 host ocs-deviceset-2-0-mgs6w 2 ssd 2.00000 osd.2 up 1.00000 1.00000 -14 2.00000 zone us-east-2b -13 2.00000 host ocs-deviceset-0-0-gjx95 1 ssd 2.00000 osd.1 up 1.00000 1.00000 -4 2.00000 zone us-east-2c -3 2.00000 host ocs-deviceset-1-0-7gz7p 0 ssd 2.00000 osd.0 up 1.00000 1.00000 And indeed all the OSD Pods are Ready. Unfortunately, this was probably from a version of must-gather before we were collecting the StorageCluster YAML, so I can't inspect its conditions to see what else might have happened. I'm not sure what else to look at, unfortunately.
@Seb, does the above give you any clue?
I don't see anything suspicious, I think the clusterdisruption-controller logs are harmless. Rohan, please confirm? Messages like: clusterdisruption-controller: could not find the CrushFindResult.Location["zone"] for "rook-ceph-osd-0" E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2" E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"] However, just like José shown the osd tree seems correct.
> Hit this issue on vSphere with ( OCP 44 ) Versions: ----------- openshift installer (4.4.0-0.nightly-2020-06-30-230655) ocs-operator.v4.4.1-465.ci > operator is in Installing Phase $ oc get csv NAME DISPLAY VERSION REPLACES PHASE lib-bucket-provisioner.v1.0.0 lib-bucket-provisioner 1.0.0 Succeeded ocs-operator.v4.4.1-465.ci OpenShift Container Storage 4.4.1-465.ci Installing > pods are in running state $ oc get pods NAME READY STATUS RESTARTS AGE csi-cephfsplugin-provisioner-6fdd566b4d-hp96g 5/5 Running 0 75m csi-cephfsplugin-provisioner-6fdd566b4d-hpbgx 5/5 Running 0 75m csi-cephfsplugin-q99mk 3/3 Running 0 75m csi-cephfsplugin-rchlg 3/3 Running 0 75m csi-cephfsplugin-vvngq 3/3 Running 0 75m csi-rbdplugin-jczkm 3/3 Running 0 75m csi-rbdplugin-provisioner-7d784b9d7c-slp7m 5/5 Running 0 75m csi-rbdplugin-provisioner-7d784b9d7c-th694 5/5 Running 0 75m csi-rbdplugin-sr2d6 3/3 Running 0 75m csi-rbdplugin-xvprd 3/3 Running 0 75m lib-bucket-provisioner-5b9cb4f848-r5k4x 1/1 Running 0 77m noobaa-core-0 1/1 Running 0 69m noobaa-db-0 1/1 Running 0 69m noobaa-operator-6f9db4df44-xwzrb 1/1 Running 0 76m ocs-operator-7fdf85b7b4-kcrzt 0/1 Running 0 76m rook-ceph-crashcollector-compute-0-7fbb7b4b76-vzv85 1/1 Running 0 69m rook-ceph-crashcollector-compute-1-6d4656d5f5-snk5p 1/1 Running 0 69m rook-ceph-crashcollector-compute-2-88cb7fddb-ph7fs 1/1 Running 0 69m rook-ceph-drain-canary-compute-0-6c7798b9b8-db88k 1/1 Running 0 69m rook-ceph-drain-canary-compute-1-7499bc7d6f-rrdm9 1/1 Running 0 69m rook-ceph-drain-canary-compute-2-6b8dd759db-brt9v 1/1 Running 0 69m rook-ceph-mds-ocs-storagecluster-cephfilesystem-a-8cb7fcf874ltf 1/1 Running 0 69m rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-6948457fldg6g 1/1 Running 0 69m rook-ceph-mgr-a-6fc677b769-mxstx 1/1 Running 0 70m rook-ceph-mon-a-c765d5889-tmnhc 1/1 Running 0 72m rook-ceph-mon-b-6cf94f4c8b-8tqvq 1/1 Running 0 71m rook-ceph-mon-c-66c6bf9b6c-pwpvr 1/1 Running 0 71m rook-ceph-operator-68557c549b-5mnnl 1/1 Running 0 76m rook-ceph-osd-0-5bfd85dbfb-xbsxh 1/1 Running 0 69m rook-ceph-osd-1-547cd58874-wppsv 1/1 Running 0 69m rook-ceph-osd-2-8b85f5b55-bhjvn 1/1 Running 0 69m rook-ceph-osd-prepare-ocs-deviceset-0-0-22rdt-5sgr6 0/1 Completed 0 70m rook-ceph-osd-prepare-ocs-deviceset-1-0-x5429-8nfcn 0/1 Completed 0 70m rook-ceph-osd-prepare-ocs-deviceset-2-0-vztrz-zmrjf 0/1 Completed 0 70m rook-ceph-rgw-ocs-storagecluster-cephobjectstore-a-65f697dxsbsf 1/1 Running 0 68m rook-ceph-tools-b7789c95-sd26k 1/1 Running 0 68m > Strangely there is no backingstore and bucketclass eventhough RGW pod is running $ oc get backingstore -A No resources found. $ oc get bucketclass -A No resources found. $ > $ oc describe csv ocs-operator.v4.4.1-465.ci Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal RequirementsUnknown 84m operator-lifecycle-manager requirements not yet checked Normal RequirementsNotMet 84m (x2 over 84m) operator-lifecycle-manager one or more requirements couldn't be found Normal InstallWaiting 83m operator-lifecycle-manager installing: waiting for deployment rook-ceph-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available... Normal InstallSucceeded 83m operator-lifecycle-manager install strategy completed with no errors Warning ComponentUnhealthy 83m (x2 over 83m) operator-lifecycle-manager installing: waiting for deployment ocs-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available... Normal InstallWaiting 82m (x3 over 84m) operator-lifecycle-manager installing: waiting for deployment ocs-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available... Normal AllRequirementsMet 77m (x6 over 84m) operator-lifecycle-manager all requirements found, attempting install Normal InstallSucceeded 77m (x5 over 84m) operator-lifecycle-manager waiting for install components to report healthy Normal NeedsReinstall 24m (x22 over 83m) operator-lifecycle-manager installing: waiting for deployment ocs-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available... Warning InstallCheckFailed 3m51s (x24 over 77m) operator-lifecycle-manager install timeout > operator pod logs has Reconciler errors {"level":"error","ts":"2020-07-01T08:02:10.817Z","logger":"controller-runtime.controller","msg":"Reconciler error","controller":"storagecluster-controller","request":"openshift-storage/ocs-storagecluster","error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again","stacktrace":"github.com/go-logr/zapr.(*zapLogger).Error\n\t/go/src/github.com/go-logr/zapr/zapr.go:128\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:218\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:192\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:171\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker-fm\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:157\nk8s.io/apimachinery/pkg/util/wait.JitterUntil.func1\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:152\nk8s.io/apimachinery/pkg/util/wait.JitterUntil\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:153\nk8s.io/apimachinery/pkg/util/wait.Until\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:88"} Jenkins Job: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9332/consoleFull Must gather: http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-pr2377-b2203/jnk-pr2377-b2203_20200701T060451/logs/failed_testcase_ocs_logs_1593583814/test_deployment_ocs_logs/ > kept the cluster for live debugging kubeconfig: http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-pr2377-b2203/jnk-pr2377-b2203_20200701T060451/openshift-cluster-dir/auth/kubeconfig
@seb can you please take a look at the live cluster with provided kubeconfig by Vijay? Thanks
> E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-0" > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-1" > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2" > E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"] It is not this, this is from when the OSD prepare job is in-progress. We can be sure that this issue went away since there is a poddisruptionbudget for every OSD credit to @Umanga for discovering that this is because the NooBaa resource has not intialized with error: cannot create an auth token for operator, error: credentials not found See NooBaa CR's status: status: accounts: admin: secretRef: name: noobaa-admin namespace: openshift-storage actualImage: quay.io/rhceph-dev/mcg-core@sha256:3b492d7c5f8904f8e63bc332238fa0c78a78fcc9dc316dee37f6823917ef7a26 conditions: - lastHeartbeatTime: "2020-07-01T07:11:44Z" lastTransitionTime: "2020-07-01T07:11:44Z" message: 'cannot create an auth token for operator, error: credentials not found' reason: TemporaryError status: "False" type: Available - lastHeartbeatTime: "2020-07-01T07:11:44Z" lastTransitionTime: "2020-07-01T07:11:44Z" message: 'cannot create an auth token for operator, error: credentials not found' reason: TemporaryError status: "True" type: Progressing - lastHeartbeatTime: "2020-07-01T07:11:44Z" lastTransitionTime: "2020-07-01T07:11:44Z" message: 'cannot create an auth token for operator, error: credentials not found' reason: TemporaryError status: "False" type: Degraded - lastHeartbeatTime: "2020-07-01T07:11:44Z" lastTransitionTime: "2020-07-01T07:11:44Z" message: 'cannot create an auth token for operator, error: credentials not found' reason: TemporaryError status: "False" type: Upgradeable observedGeneration: 1 phase: Configuring readme: "\n\n\tNooBaa operator is still working to reconcile this system.\n\tCheck out the system status.phase, status.conditions, and events with:\n\n\t\tkubectl -n openshift-storage describe noobaa\n\t\tkubectl -n openshift-storage get noobaa -o yaml\n\t\tkubectl -n openshift-storage get events --sort-by=metadata.creationTimestamp\n\n\tYou can wait for a specific condition with:\n\n\t\tkubectl -n openshift-storage wait noobaa/noobaa --for condition=available --timeout -1s\n\n\tNooBaa Core Version: 5.4.0\n\tNooBaa Operator Version: 2.2.0\n" services: serviceMgmt: externalDNS: - https://noobaa-mgmt-openshift-storage.apps.jnk-pr2377-b2203.qe.rh-ocs.com internalDNS: - https://noobaa-mgmt.openshift-storage.svc:443 internalIP: - https://172.30.42.129:443 nodePorts: - https://10.46.27.140:30370 podPorts: - https://10.128.2.23:8443 serviceS3: externalDNS: - https://s3-openshift-storage.apps.jnk-pr2377-b2203.qe.rh-ocs.com internalDNS: - https://s3.openshift-storage.svc:443 internalIP: - https://172.30.95.215:443
Have another potential repro from a job Vijay started https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9332 ocs-operator is not ready almost 3 hours after deployment started but the error seems different (Readiness probe failed: stat: cannot stat '/tmp/operator-sdk-ready': No such file or directory): (venv) [Elad@localhost ocs-ci]$ oc describe pod ocs-operator-7fdf85b7b4-kcrzt -n openshift-storage Name: ocs-operator-7fdf85b7b4-kcrzt Namespace: openshift-storage Priority: 0 Node: compute-2/10.46.27.140 Start Time: Wed, 01 Jul 2020 10:04:11 +0300 Labels: name=ocs-operator pod-template-hash=7fdf85b7b4 Annotations: alm-examples: [ { "apiVersion": "ocs.openshift.io/v1", "kind": "StorageCluster", "metadata": { "name": "example-storagecluster", "namespace": "openshift-storage" }, "spec": { "manageNodes": false, "monPVCTemplate": { "spec": { "accessModes": [ "ReadWriteOnce" ], "resources": { "requests": { "storage": "10Gi" } }, "storageClassName": "gp2" } }, "storageDeviceSets": [ { "count": 3, "dataPVCTemplate": { "spec": { "accessModes": [ "ReadWriteOnce" ], "resources": { "requests": { "storage": "1Ti" } }, "storageClassName": "gp2", "volumeMode": "Block" } }, "name": "example-deviceset", "placement": {}, "portable": true, "resources": {} } ] } } ] capabilities: Full Lifecycle categories: Storage containerImage: quay.io/ocs-dev/ocs-operator:4.4.0 createdAt: 2020-06-27 06:30:53 description: Red Hat OpenShift Container Storage provides hyperconverged storage for applications within an OpenShift cluster. k8s.v1.cni.cncf.io/networks-status: [{ "name": "openshift-sdn", "interface": "eth0", "ips": [ "10.128.2.13" ], "dns": {}, "default-route": [ "10.128.2.1" ] }] olm.operatorGroup: openshift-storage-operatorgroup olm.operatorNamespace: openshift-storage olm.skipRange: >=0.0.1 <4.4.1-465.ci olm.targetNamespaces: openshift-storage openshift.io/scc: restricted operatorframework.io/cluster-monitoring: true operatorframework.io/suggested-namespace: openshift-storage operators.operatorframework.io/internal-objects: ["cephclusters.ceph.rook.io", "cephblockpools.ceph.rook.io", "cephobjectstores.ceph.rook.io", "cephobjectstoreusers.ceph.rook.io", "cephnf... repository: https://github.com/openshift/ocs-operator support: Red Hat Status: Running IP: 10.128.2.13 IPs: IP: 10.128.2.13 Controlled By: ReplicaSet/ocs-operator-7fdf85b7b4 Containers: ocs-operator: Container ID: cri-o://d959cd860d4d21b467987339aff1f0daaaeedfb4613cbc72ddd48cbd77194a8c Image: quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45 Image ID: quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45 Port: 60000/TCP Host Port: 0/TCP Command: ocs-operator State: Running Started: Wed, 01 Jul 2020 10:04:31 +0300 Ready: False Restart Count: 0 Readiness: exec [stat /tmp/operator-sdk-ready] delay=4s timeout=1s period=10s #success=1 #failure=1 Environment: WATCH_NAMESPACE: (v1:metadata.annotations['olm.targetNamespaces']) POD_NAME: ocs-operator-7fdf85b7b4-kcrzt (v1:metadata.name) OPERATOR_NAME: ocs-operator ROOK_CEPH_IMAGE: quay.io/rhceph-dev/rook-ceph@sha256:498808453dafc8c987adc948589f175155bff4b530da73871716bfc9a4ac43e7 CEPH_IMAGE: quay.io/rhceph-dev/rhceph@sha256:20cf789235e23ddaf38e109b391d1496bb88011239d16862c4c106d0e05fea9e NOOBAA_CORE_IMAGE: quay.io/rhceph-dev/mcg-core@sha256:3b492d7c5f8904f8e63bc332238fa0c78a78fcc9dc316dee37f6823917ef7a26 NOOBAA_DB_IMAGE: registry.redhat.io/rhscl/mongodb-36-rhel7@sha256:c260fda42b5113d0183f85d7c9e873101268d91f1e74dcb38cfe76269b4ed47a Mounts: /var/run/secrets/kubernetes.io/serviceaccount from ocs-operator-token-966qk (ro) Conditions: Type Status Initialized True Ready False ContainersReady False PodScheduled True Volumes: ocs-operator-token-966qk: Type: Secret (a volume populated by a Secret) SecretName: ocs-operator-token-966qk Optional: false QoS Class: BestEffort Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute for 300s node.kubernetes.io/unreachable:NoExecute for 300s node.ocs.openshift.io/storage=true:NoSchedule Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled <unknown> default-scheduler Successfully assigned openshift-storage/ocs-operator-7fdf85b7b4-kcrzt to compute-2 Normal Pulling 163m kubelet, compute-2 Pulling image "quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45" Normal Pulled 163m kubelet, compute-2 Successfully pulled image "quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45" Normal Created 163m kubelet, compute-2 Created container ocs-operator Normal Started 163m kubelet, compute-2 Started container ocs-operator Warning Unhealthy 3m33s (x954 over 162m) kubelet, compute-2 Readiness probe failed: stat: cannot stat '/tmp/operator-sdk-ready': No such file or directory
comment #10 and #13 are of the same repro, same cluster
(In reply to Rohan CJ from comment #12) > > > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-0" > > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-1" > > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2" > > E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"] > > It is not this, this is from when the OSD prepare job is in-progress. We can > be sure that this issue went away since there is a poddisruptionbudget for > every OSD > > credit to @Umanga for discovering that this is because the NooBaa resource > has not intialized with error: cannot create an auth token for operator, > error: credentials not found > See NooBaa CR's status: > > status: > accounts: > admin: > secretRef: > name: noobaa-admin > namespace: openshift-storage > actualImage: > quay.io/rhceph-dev/mcg-core@sha256: > 3b492d7c5f8904f8e63bc332238fa0c78a78fcc9dc316dee37f6823917ef7a26 > conditions: > - lastHeartbeatTime: "2020-07-01T07:11:44Z" > lastTransitionTime: "2020-07-01T07:11:44Z" > message: 'cannot create an auth token for operator, error: credentials > not found' > reason: TemporaryError > status: "False" > type: Available > - lastHeartbeatTime: "2020-07-01T07:11:44Z" > lastTransitionTime: "2020-07-01T07:11:44Z" > message: 'cannot create an auth token for operator, error: credentials > not found' So the issue for first execution for which I opened bug looks like this was related to bucket limit issue I described in comment: https://bugzilla.redhat.com/show_bug.cgi?id=1852488#c3 If so than it's not a bug. About another issue hit by Vijay, it's probably different issue and if that's the case this should be closed as NOT A BUG and for Vijay's issue we should report the new one I think.
Looks like questions on the rook side have already been answered, removing needsinfo.
Needs some NooBaa eyes :)
we are on it since yesterday :)
re=adding lost acks
3 PRs have been added: - https://github.com/noobaa/noobaa-core/pull/6068 is on noobaa-core master ==> has apparently been backported to 5.4 via https://github.com/noobaa/noobaa-core/pull/6071 (adding) - https://github.com/noobaa/noobaa-core/pull/6067 is on noobaa-core 5.4 (corresponding to OCS 4.4) - https://github.com/noobaa/noobaa-operator/pull/356 is on noobaa-operator master ==> has apparently been backported to 2.2 via https://github.com/noobaa/noobaa-operator/pull/357 (adding) So are we all set?! It seems like....
Deployments passed in RC4 ( OCS operator v4.4.1-476.ci ) Below are different combination of Jobs > OCP 4.3 + OCS 4.4 ( vSphere ) https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9458/ > OCP 4.5 + OCS 44 ( AWS ) https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9453/ > OCP 4.4 + OCP 4.4 ( AWS ) https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9441/ Moving to Verified.
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2020:2830
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days