Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem (please be detailed as possible and provide log snippests): On Fusion HCI Provider cluster must gather stuck and finished with "timed out waiting for the condition" Version of all relevant components (if applicable): Does this issue impact your ability to continue to work with the product (please explain in detail what is the user impact)? Is there any workaround available to the best of your knowledge? Rate from 1 - 5 the complexity of the scenario you performed that caused this bug (1 - very simple, 5 - very complex)? Can this issue reproducible? Can this issue reproduce from the UI? If this is a regression, please provide more details to justify this: Steps to Reproduce: 1.Install Fusion HCI PRovider and Client Setup on BM 2.Attempt to collect must gather on Provider and Client 3. Actual results: Mustgather collection stuck dulling gathering logs Expected results: Must gather should complete successfully and add collect logs for all component Additional info: While debugging on issue, it is concluded that it has been stuck in nooba resource collection and after skipping nooba resource the must gather works with image http://quay.io/ypadia/odf-must-gather:kaustav
Hi all, - The initial assumption was, this [0] command is getting stuck and whole must gather command timed out and thought the cause might be due to this Pod being in Pending state in ODF Provider mode, however this was a wrong conclusion - I ran MG on three clusters, including on the one where the issue was observed but didn't hit it now, all MGs are uploaded here [1]. All these clusters are running ODF Provider mode - As per build mails quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a is the RC Must Gather image ================================= => On RackM09 Fusion HCI $ oc adm must-gather --image=quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a [must-gather ] OUT Using must-gather plug-in image: quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: ClusterID: fbe10030-3052-47c4-8fbf-ce703933ed8c ClusterVersion: Stable at "4.14.0" ClusterOperators: All healthy and stable [must-gather ] OUT namespace/openshift-must-gather-m6q4m created [must-gather ] OUT clusterrolebinding.rbac.authorization.k8s.io/must-gather-pb62n created [must-gather ] OUT pod for plug-in image quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a created [must-gather-vwj6x] POD 2023-11-30T05:48:11.114724732Z collection started at: 05:48:11 AM [...] [must-gather-vwj6x] OUT sent 77,332 bytes received 269,507,195 bytes 1,679,654.37 bytes/sec [must-gather-vwj6x] OUT total size is 1,581,597,985 speedup is 5.87 [must-gather ] OUT namespace/openshift-must-gather-m6q4m deleted [must-gather ] OUT clusterrolebinding.rbac.authorization.k8s.io/must-gather-pb62n deleted Reprinting Cluster State: When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: ClusterID: fbe10030-3052-47c4-8fbf-ce703933ed8c ClusterVersion: Stable at "4.14.0" ClusterOperators: All healthy and stable ===================================== => On IBM BM $ oc adm must-gather --image=quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a [must-gather ] OUT Using must-gather plug-in image: quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc446 1ee3d0a When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: ClusterID: ac75ed1b-b68d-4353-846b-a68d930eb44c ClusterVersion: Stable at "4.14.1" ClusterOperators: All healthy and stable [must-gather ] OUT namespace/openshift-must-gather-892q5 created [must-gather ] OUT clusterrolebinding.rbac.authorization.k8s.io/must-gather-vjmk9 created [must-gather ] OUT pod for plug-in image quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a cre ated [must-gather-zm58t] POD 2023-11-30T06:06:56.183488423Z collection started at: 06:06:56 AM [...] [must-gather-zm58t] OUT sent 38,880 bytes received 128,074,907 bytes 1,415,621.96 bytes/sec [must-gather-zm58t] OUT total size is 1,031,190,290 speedup is 8.05 [must-gather ] OUT namespace/openshift-must-gather-892q5 deleted [must-gather ] OUT clusterrolebinding.rbac.authorization.k8s.io/must-gather-vjmk9 deleted Reprinting Cluster State: When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: ClusterID: ac75ed1b-b68d-4353-846b-a68d930eb44c ClusterVersion: Stable at "4.14.1" ClusterOperators: All healthy and stable ================================ => On AWS OCP cluster $ oc adm must-gather --image=quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a [must-gather ] OUT Using must-gather plug-in image: quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc446 1ee3d0a When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: ClusterID: d1344bdd-ce33-4b55-a5b7-0ca41d8b8e46 ClusterVersion: Stable at "4.14.2" ClusterOperators: All healthy and stable [must-gather ] OUT namespace/openshift-must-gather-99fx7 created [must-gather ] OUT clusterrolebinding.rbac.authorization.k8s.io/must-gather-msg4x created [must-gather ] OUT pod for plug-in image quay.io/rhceph-dev/odf4-odf-must-gather-rhel9:9cf58af3fb532862936a2b62d571ad9109bc32d1f4d576c032c0bc4461ee3d0a cre ated [...] [must-gather-zlfnd] OUT sent 26,974 bytes received 34,609,230 bytes 1,539,386.84 bytes/sec [must-gather-zlfnd] OUT total size is 272,753,793 speedup is 7.87 [must-gather ] OUT namespace/openshift-must-gather-99fx7 deleted [must-gather ] OUT clusterrolebinding.rbac.authorization.k8s.io/must-gather-msg4x deleted Reprinting Cluster State: When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: ClusterID: d1344bdd-ce33-4b55-a5b7-0ca41d8b8e46 ClusterVersion: Stable at "4.14.2" ClusterOperators: All healthy and stable ================================= Conclusions: 1. This is not reproducible and I propose it NOT to be a blocker 2. Even if we want to introduce a timeout at the line where the issue was observed, unless this is reproducible we couldn't check whether the fix is working or not [0]: https://github.com/red-hat-storage/odf-must-gather/blob/release-4.14/must-gather/collection-scripts/gather_noobaa_resources#L36 [1]: https://drive.google.com/drive/folders/1pR6huSLVvimLPiNLhQiqx0nAVpbeDfoO?usp=sharing
Today I check on Fusion HCI cluster in Rack04 and QE setup. The issue is not reproduced and Collected the must gather with below command without stucking oc adm must-gather --image=quay.io/rhceph-dev/ocs-must-gather:4.14-fusion-hci