Bug 1852488 - CSV is stuck in installing for OCP 4.4.1 RC2 build on OCP 4.3
Summary: CSV is stuck in installing for OCP 4.4.1 RC2 build on OCP 4.3
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat OpenShift Container Storage
Classification: Red Hat Storage
Component: Multi-Cloud Object Gateway
Version: 4.4
Hardware: Unspecified
OS: Unspecified
urgent
urgent
Target Milestone: ---
: OCS 4.4.1
Assignee: Nimrod Becker
QA Contact: Raz Tamir
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2020-06-30 14:22 UTC by Petr Balogh
Modified: 2023-09-14 06:03 UTC (History)
14 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2020-07-07 06:10:29 UTC
Embargoed:


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Github noobaa noobaa-core pull 6067 0 None closed Backport to 5.4: Add a lock on the load semaphore during make_changes 2021-01-19 13:56:01 UTC
Github noobaa noobaa-core pull 6068 0 None closed Add and expose the system state via API (currently only initialization states) 2021-01-19 13:56:01 UTC
Github noobaa noobaa-core pull 6071 0 None closed Bsckport to 5.4: Add and expose the system state via API (currently only initializati… 2021-01-19 13:56:01 UTC
Github noobaa noobaa-operator pull 356 0 None closed Refactor system creation flow. 2021-01-19 13:56:01 UTC
Github noobaa noobaa-operator pull 357 0 None closed Backport to 2.2 2021-01-19 13:56:02 UTC
Red Hat Product Errata RHBA-2020:2830 0 None None None 2020-07-07 06:10:37 UTC

Description Petr Balogh 2020-06-30 14:22:18 UTC
Description of problem (please be detailed as possible and provide log
snippests):

From operator log:
http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-ai3c33-t1/jnk-ai3c33-t1_20200630T113618/logs/failed_testcase_ocs_logs_1593517309/test_deployment_ocs_logs/ocs_must_gather/quay-io-rhceph-dev-ocs-must-gather-sha256-4db6344cc517f57df8ca261314aeba25d5348d3f033a14e90267b75b905efc8a/ceph/namespaces/openshift-storage/pods/ocs-operator-7fdf85b7b4-jhm9t/ocs-operator/ocs-operator/logs/current.log 

I see some errors:
2020-06-30T13:02:45.731471005Z {"level":"error","ts":"2020-06-30T13:02:45.731Z","logger":"controller_storagecluster","msg":"Failed to update status","Request.Namespace":"openshift-storage","Request.Name":"ocs-storagecluster","error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again","stacktrace":"github.com/go-logr/zapr.(*zapLogger).Error\n\t/go/src/github.com/go-logr/zapr/zapr.go:128\ngithub.com/openshift/ocs-operator/pkg/controller/storagecluster.(*ReconcileStorageCluster).Reconcile\n\t/go/src/github.com/openshift/ocs-operator/pkg/controller/storagecluster/reconcile.go:366\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:216\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:192\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:171\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker-fm\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:157\nk8s.io/apimachinery/pkg/util/wait.JitterUntil.func1\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:152\nk8s.io/apimachinery/pkg/util/wait.JitterUntil\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:153\nk8s.io/apimachinery/pkg/util/wait.Until\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:88"}
2020-06-30T13:02:45.731550981Z {"level":"error","ts":"2020-06-30T13:02:45.731Z","logger":"controller-runtime.controller","msg":"Reconciler error","controller":"storagecluster-controller","request":"openshift-storage/ocs-storagecluster","error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again","stacktrace":"github.com/go-logr/zapr.(*zapLogger).Error\n\t/go/src/github.com/go-logr/zapr/zapr.go:128\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:218\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:192\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:171\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker-fm\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:157\nk8s.io/apimachinery/pkg/util/wait.JitterUntil.func1\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:152\nk8s.io/apimachinery/pkg/util/wait.JitterUntil\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:153\nk8s.io/apimachinery/pkg/util/wait.Until\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:88"}

I see operator pod is in this state:
ocs-operator-7fdf85b7b4-jhm9t                                     0/1     Running     0          47m   10.129.2.12   ip-10-0-138-54.us-east-2.compute.internal   <none>           <none>


Version of all relevant components (if applicable):
OCP: 4.3.0-0.nightly-2020-06-29-084049
OCS: 4.4.1 - 4.4.1-465.ci RC2 build

Does this issue impact your ability to continue to work with the product
(please explain in detail what is the user impact)?
YES


Is there any workaround available to the best of your knowledge?
NO

Rate from 1 - 5 the complexity of the scenario you performed that caused this
bug (1 - very simple, 5 - very complex)?
1

Can this issue reproducible?
Trying now:
https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9312/

Can this issue reproduce from the UI?
Haven't tried

If this is a regression, please provide more details to justify this:
Yes

Steps to Reproduce:
1. Install OCP 4.3 nightly
2. Install OCS 4.4.1 build


Actual results:
Installation of OCS is not completed

Expected results:
OCS operator to finish installation

Additional info:
Jenkins job:
https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9306/consoleFull

Must gather logs:
http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-ai3c33-t1/jnk-ai3c33-t1_20200630T113618/logs/failed_testcase_ocs_logs_1593517309/test_deployment_ocs_logs/

Comment 3 Petr Balogh 2020-06-30 15:59:57 UTC
Update here - second time I was able to deploy OCS on top of OCP 4.3 - but then the deployment failed on creating credentials on S3 as we hit the limit of 400 buckets under our account.


Execution here:
https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9312/console

When I connected to the cluster I see that it finished the deployment here:
$ oc get csv -n openshift-storage
NAME                            DISPLAY                       VERSION        REPLACES   PHASE
lib-bucket-provisioner.v1.0.0   lib-bucket-provisioner        1.0.0                     Succeeded
ocs-operator.v4.4.1-465.ci      OpenShift Container Storage   4.4.1-465.ci              Succeeded


We are cleaning the buckets and will re-run once more:
https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-trigger-aws-ipi-3az-rhcos-3m-3w-tier1/45

Comment 4 Petr Balogh 2020-06-30 21:13:18 UTC
I see that other job we scheduled passed deployment without any issue here:
https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9319/console

As well we got results from tier1. Maybe it was some temporary issue but I would still like to see if someone can see from the logs what is the root cause of the failure from the first deployment.

@Travis, can you please have a look at operator logs if you see why the operator didn't get to succeeded status in first attempt?

What the error I pasted in the 1st comment:
"error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again
is about?

Thanks

Comment 5 Michael Adam 2020-06-30 21:59:52 UTC
@Petr, so it seems to be intermitten, not consistent?

@José could you have a look to see if anything stands out?

Comment 6 Petr Balogh 2020-06-30 22:45:32 UTC
Yes Michael 1st out of 3 runs we saw this issue.
But worth to at least try to check the logs and maybe Jose will see what was the root cause of first failure.

Thanks

Comment 7 Jose A. Rivera 2020-07-01 03:40:09 UTC
I'm seeing nothing immediately obvious in the ocs-operator logs. Supposedly it thinks the CephCluster isn't reporting status but the CephCluster CR says it's HEALTH_OK. I found the following in the Rook-Ceph log:

E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-0"
E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-1"
E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2"
E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"]

But the ceph osd tree command shows a good result:

ID  CLASS WEIGHT  TYPE NAME                                STATUS REWEIGHT PRI-AFF 
 -1       6.00000 root default                                                     
 -5       6.00000     region us-east-2                                             
-10       2.00000         zone us-east-2a                                          
 -9       2.00000             host ocs-deviceset-2-0-mgs6w                         
  2   ssd 2.00000                 osd.2                        up  1.00000 1.00000 
-14       2.00000         zone us-east-2b                                          
-13       2.00000             host ocs-deviceset-0-0-gjx95                         
  1   ssd 2.00000                 osd.1                        up  1.00000 1.00000 
 -4       2.00000         zone us-east-2c                                          
 -3       2.00000             host ocs-deviceset-1-0-7gz7p                         
  0   ssd 2.00000                 osd.0                        up  1.00000 1.00000 

And indeed all the OSD Pods are Ready.

Unfortunately, this was probably from a version of must-gather before we were collecting the StorageCluster YAML, so I can't inspect its conditions to see what else might have happened. I'm not sure what else to look at, unfortunately.

Comment 8 Michael Adam 2020-07-01 07:10:50 UTC
@Seb, does the above give you any clue?

Comment 9 Sébastien Han 2020-07-01 07:35:31 UTC
I don't see anything suspicious, I think the clusterdisruption-controller logs are harmless.
Rohan, please confirm?

Messages like:

clusterdisruption-controller: could not find the CrushFindResult.Location["zone"] for "rook-ceph-osd-0"
E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2"
E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"]

However, just like José shown the osd tree seems correct.

Comment 10 Vijay Avuthu 2020-07-01 08:36:43 UTC
> Hit this issue on vSphere with ( OCP 44 )

Versions:
-----------

openshift installer (4.4.0-0.nightly-2020-06-30-230655)

ocs-operator.v4.4.1-465.ci

> operator is in Installing Phase

$ oc get csv
NAME                            DISPLAY                       VERSION        REPLACES   PHASE
lib-bucket-provisioner.v1.0.0   lib-bucket-provisioner        1.0.0                     Succeeded
ocs-operator.v4.4.1-465.ci      OpenShift Container Storage   4.4.1-465.ci              Installing


> pods are in running state

$ oc get pods
NAME                                                              READY   STATUS      RESTARTS   AGE
csi-cephfsplugin-provisioner-6fdd566b4d-hp96g                     5/5     Running     0          75m
csi-cephfsplugin-provisioner-6fdd566b4d-hpbgx                     5/5     Running     0          75m
csi-cephfsplugin-q99mk                                            3/3     Running     0          75m
csi-cephfsplugin-rchlg                                            3/3     Running     0          75m
csi-cephfsplugin-vvngq                                            3/3     Running     0          75m
csi-rbdplugin-jczkm                                               3/3     Running     0          75m
csi-rbdplugin-provisioner-7d784b9d7c-slp7m                        5/5     Running     0          75m
csi-rbdplugin-provisioner-7d784b9d7c-th694                        5/5     Running     0          75m
csi-rbdplugin-sr2d6                                               3/3     Running     0          75m
csi-rbdplugin-xvprd                                               3/3     Running     0          75m
lib-bucket-provisioner-5b9cb4f848-r5k4x                           1/1     Running     0          77m
noobaa-core-0                                                     1/1     Running     0          69m
noobaa-db-0                                                       1/1     Running     0          69m
noobaa-operator-6f9db4df44-xwzrb                                  1/1     Running     0          76m
ocs-operator-7fdf85b7b4-kcrzt                                     0/1     Running     0          76m
rook-ceph-crashcollector-compute-0-7fbb7b4b76-vzv85               1/1     Running     0          69m
rook-ceph-crashcollector-compute-1-6d4656d5f5-snk5p               1/1     Running     0          69m
rook-ceph-crashcollector-compute-2-88cb7fddb-ph7fs                1/1     Running     0          69m
rook-ceph-drain-canary-compute-0-6c7798b9b8-db88k                 1/1     Running     0          69m
rook-ceph-drain-canary-compute-1-7499bc7d6f-rrdm9                 1/1     Running     0          69m
rook-ceph-drain-canary-compute-2-6b8dd759db-brt9v                 1/1     Running     0          69m
rook-ceph-mds-ocs-storagecluster-cephfilesystem-a-8cb7fcf874ltf   1/1     Running     0          69m
rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-6948457fldg6g   1/1     Running     0          69m
rook-ceph-mgr-a-6fc677b769-mxstx                                  1/1     Running     0          70m
rook-ceph-mon-a-c765d5889-tmnhc                                   1/1     Running     0          72m
rook-ceph-mon-b-6cf94f4c8b-8tqvq                                  1/1     Running     0          71m
rook-ceph-mon-c-66c6bf9b6c-pwpvr                                  1/1     Running     0          71m
rook-ceph-operator-68557c549b-5mnnl                               1/1     Running     0          76m
rook-ceph-osd-0-5bfd85dbfb-xbsxh                                  1/1     Running     0          69m
rook-ceph-osd-1-547cd58874-wppsv                                  1/1     Running     0          69m
rook-ceph-osd-2-8b85f5b55-bhjvn                                   1/1     Running     0          69m
rook-ceph-osd-prepare-ocs-deviceset-0-0-22rdt-5sgr6               0/1     Completed   0          70m
rook-ceph-osd-prepare-ocs-deviceset-1-0-x5429-8nfcn               0/1     Completed   0          70m
rook-ceph-osd-prepare-ocs-deviceset-2-0-vztrz-zmrjf               0/1     Completed   0          70m
rook-ceph-rgw-ocs-storagecluster-cephobjectstore-a-65f697dxsbsf   1/1     Running     0          68m
rook-ceph-tools-b7789c95-sd26k                                    1/1     Running     0          68m


> Strangely there is no backingstore and bucketclass eventhough RGW pod is running

$ oc get backingstore -A
No resources found.
$ oc get bucketclass -A
No resources found.
$

> $ oc describe csv ocs-operator.v4.4.1-465.ci

Events:
  Type     Reason               Age                   From                        Message
  ----     ------               ----                  ----                        -------
  Normal   RequirementsUnknown  84m                   operator-lifecycle-manager  requirements not yet checked
  Normal   RequirementsNotMet   84m (x2 over 84m)     operator-lifecycle-manager  one or more requirements couldn't be found
  Normal   InstallWaiting       83m                   operator-lifecycle-manager  installing: waiting for deployment rook-ceph-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available...
  Normal   InstallSucceeded     83m                   operator-lifecycle-manager  install strategy completed with no errors
  Warning  ComponentUnhealthy   83m (x2 over 83m)     operator-lifecycle-manager  installing: waiting for deployment ocs-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available...
  Normal   InstallWaiting       82m (x3 over 84m)     operator-lifecycle-manager  installing: waiting for deployment ocs-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available...
  Normal   AllRequirementsMet   77m (x6 over 84m)     operator-lifecycle-manager  all requirements found, attempting install
  Normal   InstallSucceeded     77m (x5 over 84m)     operator-lifecycle-manager  waiting for install components to report healthy
  Normal   NeedsReinstall       24m (x22 over 83m)    operator-lifecycle-manager  installing: waiting for deployment ocs-operator to become ready: Waiting for rollout to finish: 0 of 1 updated replicas are available...
  Warning  InstallCheckFailed   3m51s (x24 over 77m)  operator-lifecycle-manager  install timeout

> operator pod logs has Reconciler errors

{"level":"error","ts":"2020-07-01T08:02:10.817Z","logger":"controller-runtime.controller","msg":"Reconciler error","controller":"storagecluster-controller","request":"openshift-storage/ocs-storagecluster","error":"Operation cannot be fulfilled on storageclusters.ocs.openshift.io \"ocs-storagecluster\": the object has been modified; please apply your changes to the latest version and try again","stacktrace":"github.com/go-logr/zapr.(*zapLogger).Error\n\t/go/src/github.com/go-logr/zapr/zapr.go:128\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:218\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:192\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:171\nsigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker-fm\n\t/go/src/sigs.k8s.io/controller-runtime/pkg/internal/controller/controller.go:157\nk8s.io/apimachinery/pkg/util/wait.JitterUntil.func1\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:152\nk8s.io/apimachinery/pkg/util/wait.JitterUntil\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:153\nk8s.io/apimachinery/pkg/util/wait.Until\n\t/go/src/k8s.io/apimachinery/pkg/util/wait/wait.go:88"}



Jenkins Job: https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9332/consoleFull
Must gather: http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-pr2377-b2203/jnk-pr2377-b2203_20200701T060451/logs/failed_testcase_ocs_logs_1593583814/test_deployment_ocs_logs/

> kept the cluster for live debugging

kubeconfig: http://magna002.ceph.redhat.com/ocsci-jenkins/openshift-clusters/jnk-pr2377-b2203/jnk-pr2377-b2203_20200701T060451/openshift-cluster-dir/auth/kubeconfig

Comment 11 Petr Balogh 2020-07-01 09:35:24 UTC
@seb can you please take a look at the live cluster with provided kubeconfig by Vijay? Thanks

Comment 12 Rohan CJ 2020-07-01 09:48:49 UTC

> E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-0"
> E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-1"
> E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2"
> E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"]

It is not this, this is from when the OSD prepare job is in-progress. We can be sure that this issue went away since there is a poddisruptionbudget for every OSD

credit to @Umanga for discovering that this is because the NooBaa resource has not intialized with error: cannot create an auth token for operator, error: credentials not found
See NooBaa CR's status:

  status:
    accounts:
      admin:
        secretRef:
          name: noobaa-admin
          namespace: openshift-storage
    actualImage: quay.io/rhceph-dev/mcg-core@sha256:3b492d7c5f8904f8e63bc332238fa0c78a78fcc9dc316dee37f6823917ef7a26
    conditions:
    - lastHeartbeatTime: "2020-07-01T07:11:44Z"
      lastTransitionTime: "2020-07-01T07:11:44Z"
      message: 'cannot create an auth token for operator, error: credentials not found'
      reason: TemporaryError
      status: "False"
      type: Available
    - lastHeartbeatTime: "2020-07-01T07:11:44Z"
      lastTransitionTime: "2020-07-01T07:11:44Z"
      message: 'cannot create an auth token for operator, error: credentials not found'
      reason: TemporaryError
      status: "True"
      type: Progressing
    - lastHeartbeatTime: "2020-07-01T07:11:44Z"
      lastTransitionTime: "2020-07-01T07:11:44Z"
      message: 'cannot create an auth token for operator, error: credentials not found'
      reason: TemporaryError
      status: "False"
      type: Degraded
    - lastHeartbeatTime: "2020-07-01T07:11:44Z"
      lastTransitionTime: "2020-07-01T07:11:44Z"
      message: 'cannot create an auth token for operator, error: credentials not found'
      reason: TemporaryError
      status: "False"
      type: Upgradeable
    observedGeneration: 1
    phase: Configuring
    readme: "\n\n\tNooBaa operator is still working to reconcile this system.\n\tCheck
      out the system status.phase, status.conditions, and events with:\n\n\t\tkubectl
      -n openshift-storage describe noobaa\n\t\tkubectl -n openshift-storage get noobaa
      -o yaml\n\t\tkubectl -n openshift-storage get events --sort-by=metadata.creationTimestamp\n\n\tYou
      can wait for a specific condition with:\n\n\t\tkubectl -n openshift-storage
      wait noobaa/noobaa --for condition=available --timeout -1s\n\n\tNooBaa Core
      Version:     5.4.0\n\tNooBaa Operator Version: 2.2.0\n"
    services:
      serviceMgmt:
        externalDNS:
        - https://noobaa-mgmt-openshift-storage.apps.jnk-pr2377-b2203.qe.rh-ocs.com
        internalDNS:
        - https://noobaa-mgmt.openshift-storage.svc:443
        internalIP:
        - https://172.30.42.129:443
        nodePorts:
        - https://10.46.27.140:30370
        podPorts:
        - https://10.128.2.23:8443
      serviceS3:
        externalDNS:
        - https://s3-openshift-storage.apps.jnk-pr2377-b2203.qe.rh-ocs.com
        internalDNS:
        - https://s3.openshift-storage.svc:443
        internalIP:
        - https://172.30.95.215:443

Comment 13 Elad 2020-07-01 09:51:52 UTC
Have another potential repro from a job Vijay started https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9332

ocs-operator is not ready almost 3 hours after deployment started but the error seems different (Readiness probe failed: stat: cannot stat '/tmp/operator-sdk-ready': No such file or directory):


(venv) [Elad@localhost ocs-ci]$ oc describe pod ocs-operator-7fdf85b7b4-kcrzt -n openshift-storage
Name:         ocs-operator-7fdf85b7b4-kcrzt
Namespace:    openshift-storage
Priority:     0
Node:         compute-2/10.46.27.140
Start Time:   Wed, 01 Jul 2020 10:04:11 +0300
Labels:       name=ocs-operator
              pod-template-hash=7fdf85b7b4
Annotations:  alm-examples:
                
                [
                    {
                        "apiVersion": "ocs.openshift.io/v1",
                        "kind": "StorageCluster",
                        "metadata": {
                            "name": "example-storagecluster",
                            "namespace": "openshift-storage"
                        },
                        "spec": {
                            "manageNodes": false,
                            "monPVCTemplate": {
                                "spec": {
                                    "accessModes": [
                                        "ReadWriteOnce"
                                    ],
                                    "resources": {
                                        "requests": {
                                            "storage": "10Gi"
                                        }
                                    },
                                    "storageClassName": "gp2"
                                }
                            },
                            "storageDeviceSets": [
                                {
                                    "count": 3,
                                    "dataPVCTemplate": {
                                        "spec": {
                                            "accessModes": [
                                                "ReadWriteOnce"
                                            ],
                                            "resources": {
                                                "requests": {
                                                    "storage": "1Ti"
                                                }
                                            },
                                            "storageClassName": "gp2",
                                            "volumeMode": "Block"
                                        }
                                    },
                                    "name": "example-deviceset",
                                    "placement": {},
                                    "portable": true,
                                    "resources": {}
                                }
                            ]
                        }
                    }
                ]
              capabilities: Full Lifecycle
              categories: Storage
              containerImage: quay.io/ocs-dev/ocs-operator:4.4.0
              createdAt: 2020-06-27 06:30:53
              description: Red Hat OpenShift Container Storage provides hyperconverged storage for applications within an OpenShift cluster.
              k8s.v1.cni.cncf.io/networks-status:
                [{
                    "name": "openshift-sdn",
                    "interface": "eth0",
                    "ips": [
                        "10.128.2.13"
                    ],
                    "dns": {},
                    "default-route": [
                        "10.128.2.1"
                    ]
                }]
              olm.operatorGroup: openshift-storage-operatorgroup
              olm.operatorNamespace: openshift-storage
              olm.skipRange: >=0.0.1 <4.4.1-465.ci
              olm.targetNamespaces: openshift-storage
              openshift.io/scc: restricted
              operatorframework.io/cluster-monitoring: true
              operatorframework.io/suggested-namespace: openshift-storage
              operators.operatorframework.io/internal-objects:
                ["cephclusters.ceph.rook.io", "cephblockpools.ceph.rook.io", "cephobjectstores.ceph.rook.io", "cephobjectstoreusers.ceph.rook.io", "cephnf...
              repository: https://github.com/openshift/ocs-operator
              support: Red Hat
Status:       Running
IP:           10.128.2.13
IPs:
  IP:           10.128.2.13
Controlled By:  ReplicaSet/ocs-operator-7fdf85b7b4
Containers:
  ocs-operator:
    Container ID:  cri-o://d959cd860d4d21b467987339aff1f0daaaeedfb4613cbc72ddd48cbd77194a8c
    Image:         quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45
    Image ID:      quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45
    Port:          60000/TCP
    Host Port:     0/TCP
    Command:
      ocs-operator
    State:          Running
      Started:      Wed, 01 Jul 2020 10:04:31 +0300
    Ready:          False
    Restart Count:  0
    Readiness:      exec [stat /tmp/operator-sdk-ready] delay=4s timeout=1s period=10s #success=1 #failure=1
    Environment:
      WATCH_NAMESPACE:     (v1:metadata.annotations['olm.targetNamespaces'])
      POD_NAME:           ocs-operator-7fdf85b7b4-kcrzt (v1:metadata.name)
      OPERATOR_NAME:      ocs-operator
      ROOK_CEPH_IMAGE:    quay.io/rhceph-dev/rook-ceph@sha256:498808453dafc8c987adc948589f175155bff4b530da73871716bfc9a4ac43e7
      CEPH_IMAGE:         quay.io/rhceph-dev/rhceph@sha256:20cf789235e23ddaf38e109b391d1496bb88011239d16862c4c106d0e05fea9e
      NOOBAA_CORE_IMAGE:  quay.io/rhceph-dev/mcg-core@sha256:3b492d7c5f8904f8e63bc332238fa0c78a78fcc9dc316dee37f6823917ef7a26
      NOOBAA_DB_IMAGE:    registry.redhat.io/rhscl/mongodb-36-rhel7@sha256:c260fda42b5113d0183f85d7c9e873101268d91f1e74dcb38cfe76269b4ed47a
    Mounts:
      /var/run/secrets/kubernetes.io/serviceaccount from ocs-operator-token-966qk (ro)
Conditions:
  Type              Status
  Initialized       True 
  Ready             False 
  ContainersReady   False 
  PodScheduled      True 
Volumes:
  ocs-operator-token-966qk:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  ocs-operator-token-966qk
    Optional:    false
QoS Class:       BestEffort
Node-Selectors:  <none>
Tolerations:     node.kubernetes.io/not-ready:NoExecute for 300s
                 node.kubernetes.io/unreachable:NoExecute for 300s
                 node.ocs.openshift.io/storage=true:NoSchedule
Events:
  Type     Reason     Age                     From                Message
  ----     ------     ----                    ----                -------
  Normal   Scheduled  <unknown>               default-scheduler   Successfully assigned openshift-storage/ocs-operator-7fdf85b7b4-kcrzt to compute-2
  Normal   Pulling    163m                    kubelet, compute-2  Pulling image "quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45"
  Normal   Pulled     163m                    kubelet, compute-2  Successfully pulled image "quay.io/rhceph-dev/ocs-operator@sha256:d9141225c09540da58d8b206ed21835dc4dbdf9eb058c66ffb95542ee3444c45"
  Normal   Created    163m                    kubelet, compute-2  Created container ocs-operator
  Normal   Started    163m                    kubelet, compute-2  Started container ocs-operator
  Warning  Unhealthy  3m33s (x954 over 162m)  kubelet, compute-2  Readiness probe failed: stat: cannot stat '/tmp/operator-sdk-ready': No such file or directory

Comment 14 Elad 2020-07-01 09:58:17 UTC
comment #10 and #13 are of the same repro, same cluster

Comment 15 Petr Balogh 2020-07-01 10:10:08 UTC
(In reply to Rohan CJ from comment #12)
> 
> > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-0"
> > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-1"
> > E | clusterdisruption-controller: failureDomain type "zone" not associated with "rook-ceph-osd-2"
> > E | clusterdisruption-controller: failure domain type "zone" not associated with osds: ["rook-ceph-osd-0,rook-ceph-osd-1,rook-ceph-osd-2"]
> 
> It is not this, this is from when the OSD prepare job is in-progress. We can
> be sure that this issue went away since there is a poddisruptionbudget for
> every OSD
> 
> credit to @Umanga for discovering that this is because the NooBaa resource
> has not intialized with error: cannot create an auth token for operator,
> error: credentials not found
> See NooBaa CR's status:
> 
>   status:
>     accounts:
>       admin:
>         secretRef:
>           name: noobaa-admin
>           namespace: openshift-storage
>     actualImage:
> quay.io/rhceph-dev/mcg-core@sha256:
> 3b492d7c5f8904f8e63bc332238fa0c78a78fcc9dc316dee37f6823917ef7a26
>     conditions:
>     - lastHeartbeatTime: "2020-07-01T07:11:44Z"
>       lastTransitionTime: "2020-07-01T07:11:44Z"
>       message: 'cannot create an auth token for operator, error: credentials
> not found'
>       reason: TemporaryError
>       status: "False"
>       type: Available
>     - lastHeartbeatTime: "2020-07-01T07:11:44Z"
>       lastTransitionTime: "2020-07-01T07:11:44Z"
>       message: 'cannot create an auth token for operator, error: credentials
> not found'


So the issue for first execution for which I opened bug looks like this was related to bucket limit issue I described in comment: https://bugzilla.redhat.com/show_bug.cgi?id=1852488#c3

If so than it's not a bug. 

About another issue hit by Vijay, it's probably different issue and if that's the case this should be closed as NOT A BUG and for Vijay's issue we should report the new one I think.

Comment 16 Travis Nielsen 2020-07-01 16:14:43 UTC
Looks like questions on the rook side have already been answered, removing needsinfo.

Comment 17 Rohan CJ 2020-07-02 06:12:29 UTC
Needs some NooBaa eyes :)

Comment 18 Nimrod Becker 2020-07-02 06:15:34 UTC
we are on it since yesterday :)

Comment 19 Michael Adam 2020-07-02 22:32:47 UTC
re=adding lost acks

Comment 22 Michael Adam 2020-07-03 10:50:44 UTC
3 PRs have been added:

- https://github.com/noobaa/noobaa-core/pull/6068 is on noobaa-core master
  ==> has apparently been backported to 5.4 via https://github.com/noobaa/noobaa-core/pull/6071 (adding)

- https://github.com/noobaa/noobaa-core/pull/6067 is on noobaa-core 5.4 (corresponding to OCS 4.4)

- https://github.com/noobaa/noobaa-operator/pull/356 is on noobaa-operator master
  ==> has apparently been backported to 2.2 via https://github.com/noobaa/noobaa-operator/pull/357 (adding)

So are we all set?!
It seems like....

Comment 25 Vijay Avuthu 2020-07-06 07:15:29 UTC
Deployments passed in RC4 ( OCS operator	v4.4.1-476.ci )

Below are different combination of Jobs


> OCP 4.3 + OCS 4.4 ( vSphere )

https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9458/

> OCP 4.5 + OCS 44 ( AWS )

https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9453/

> OCP 4.4 + OCP 4.4 ( AWS )

https://ocs4-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/job/qe-deploy-ocs-cluster/9441/

Moving to Verified.

Comment 27 errata-xmlrpc 2020-07-07 06:10:29 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2020:2830

Comment 28 Red Hat Bugzilla 2023-09-14 06:03:12 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days


Note You need to log in before you can comment on or make changes to this bug.