Bug 1845976
| Summary: | OCS 4.5 Independent mode: must-gather commands fails to collect ceph command outputs from external cluster | ||||||
|---|---|---|---|---|---|---|---|
| Product: | [Red Hat Storage] Red Hat OpenShift Container Storage | Reporter: | Neha Berry <nberry> | ||||
| Component: | must-gather | Assignee: | Pulkit Kundra <pkundra> | ||||
| Status: | CLOSED ERRATA | QA Contact: | Neha Berry <nberry> | ||||
| Severity: | high | Docs Contact: | |||||
| Priority: | unspecified | ||||||
| Version: | 4.5 | CC: | assingh, bkunal, ebenahar, jarrpa, jijoy, madam, muagarwa, ocs-bugs, pkundra, sabose, sagrawal, shan, tnielsen | ||||
| Target Milestone: | --- | Keywords: | AutomationBackLog | ||||
| Target Release: | OCS 4.6.0 | ||||||
| Hardware: | Unspecified | ||||||
| OS: | Unspecified | ||||||
| Whiteboard: | |||||||
| Fixed In Version: | 4.6.0-137.ci | Doc Type: | No Doc Update | ||||
| Doc Text: | Story Points: | --- | |||||
| Clone Of: | Environment: | ||||||
| Last Closed: | 2020-12-17 06:22:30 UTC | Type: | Bug | ||||
| Regression: | --- | Mount Type: | --- | ||||
| Documentation: | --- | CRM: | |||||
| Verified Versions: | Category: | --- | |||||
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |||||
| Cloudforms Team: | --- | Target Upstream Version: | |||||
| Embargoed: | |||||||
| Attachments: |
|
||||||
|
Description
Neha Berry
2020-06-10 14:13:21 UTC
I'm not sure we'd collect logs from an external RHCS cluster via ocs-must-gather. We have tools to collect logs on RHCS side. (In reply to Yaniv Kaul from comment #2) > I'm not sure we'd collect logs from an external RHCS cluster via > ocs-must-gather. We have tools to collect logs on RHCS side. Not the logs, only output of ceph commands. Pulkit, can you check if this is possible - i.e running ceph commands on external cluster from toolbox What's the decision - is it going to be worked on for 4.5 or not? You forgot to also remove the ocs-4.5.0? flag. :) Doing so now. If not for OCS 4.5, can we plan to consider the fix for z-stream of OCS 4.5 ? (In reply to Neha Berry from comment #10) > If not for OCS 4.5, can we plan to consider the fix for z-stream of OCS 4.5 ? Yes, should be possible as there's a patch available already Sahina, no, the change is too big, it's too risky. (In reply to leseb from comment #14) > Sahina, no, the change is too big, it's too risky. I think this means we should move it from 4.5.z to 4.6.0 ? (In reply to Michael Adam from comment #15) > (In reply to leseb from comment #14) > > Sahina, no, the change is too big, it's too risky. > > I think this means we should move it from 4.5.z to 4.6.0 ? Done Not putting "fixed in version" because it is there in 4.6 for a long time now. @Neha Can you share the /etc/ceph/ceph.conf from the toolbox pod? Also does the /etc/ceph/keyring match what you expect? Something must be wrong in the toolbox config that is preventing the ceph connection. Neha, can you please give it a try now? (In reply to Mudit Agarwal from comment #25) > Neha, can you please give it a try now? Hi Mudit, What Sidhant tested was just to confirm that we can run ceph commands in the manually created toolbox, in case the toolbox has proper ceph admin key. It is not the solution to OCS must-gather. There is no fix yet to try again. Me, Sidhant got into the Operators meeting and discussed the scenario. The toolbox created during must-gather or the one created via [1] lacks this admin key, hence it is unable to connect to the RHCS cluster to run ceph commands. The error message we get is: $ oc rsh rook-ceph-tools-9858c9845-6z5q8 sh-4.4$ ceph -s [errno 5] RADOS I/O error (error connecting to the cluster) sh-4.4$ The key in secret "rook-ceph-mon", which is part of toolbox pod created during must-gather, does not have admin rights. @Travis confirmed that he will look into the issue as to how to get proper key to toolbox. [1] - oc patch ocsinitialization ocsinit -n openshift-storage --type json --patch '[{ "op": "replace", "path": "/spec/enableCephTools", "value": true }]' Ok, we see now that the expected must-gather commands all require the admin keyring and do not work with the lower-privileged keyring that was provided to the cluster. With the fix to the toolbox that was included in 4.6, it was only to properly use whatever keyring was provided for the external cluster. It didn't mean that the toolbox was expected to have privileges to run every Ceph command. By design, the external cluster provides a lower-privileged key to connect with and the must-gather commands will fail as long as no admin key is provided. As Yaniv originally indicated in the bug, we don't expect to gather the Ceph status of the external cluster. We will need to rely on the RHCS admin to provide information about the external cluster. OCS isn't the admin of the external cluster, so we can't expect to gather admin-privileged info. @Pulkit Either must-gather shouldn't call the admin ceph commands on the external cluster, or we need to ignore the errors. (In reply to Travis Nielsen from comment #27) > Ok, we see now that the expected must-gather commands all require the admin > keyring and do not work with the lower-privileged keyring that was provided > to the cluster. With the fix to the toolbox that was included in 4.6, it was > only to properly use whatever keyring was provided for the external cluster. > It didn't mean that the toolbox was expected to have privileges to run every > Ceph command. > > By design, the external cluster provides a lower-privileged key to connect > with and the must-gather commands will fail as long as no admin key is > provided. > > As Yaniv originally indicated in the bug, we don't expect to gather the Ceph > status of the external cluster. We will need to rely on the RHCS admin to > provide information about the external cluster. OCS isn't the admin of the > external cluster, so we can't expect to gather admin-privileged info. > > @Pulkit Either must-gather shouldn't call the admin ceph commands on the > external cluster, or we need to ignore the errors. Hi Travis, After trying it manually, it does seem difficult to gain access to RHCS admin key as the uploaded JSON doesnt have the key (as you said) But, in case some of our PVCs are pending or we are facing OCS related issues, we would still want some information from the RHCS side. @bipin should we have a KCS article in place on how to collect cpeh commands after adding the ceph admin key to toolbox ? (provided RHCS admin provides the key to support team? ). Just thinking out loud. Let me know if this doesnt make sense at all. Neha, Lets gather must-gather and sos-report from RHCS node for External Mode. -Bipin Kunal Created attachment 1723253 [details] terminal output from must-gather Ack. I will raise a new troubleshooting doc BZ to collect sosreport from RHCS side Verified the fix in OCS version 4.6.0-137.ci and OCP 4.6.0-0.nightly-2020-10-17-040148 must-gather is skipping the collection of ceph commands and creation of a toolbox(must-gather-helper) pod in the openshift-storage namespace Snip from terminal ========================= must-gather-hck57] POD collecting dump of noobaa-db-0 pod from openshift-storage [must-gather-hck57] POD collecting dump of noobaa-operator-6499b55c9b-x6hrg pod from openshift-storage >> [must-gather-hck57] POD Skipping the ceph collection as External Storage is enabled [must-gather-hck57] OUT waiting for gather to complete [must-gather-hck57] OUT downloading gather output Based on the fix, moving the BZ to verified state. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (Moderate: Red Hat OpenShift Container Storage 4.6.0 security, bug fix, enhancement update), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHSA-2020:5605 |