Bug 2253429
| Summary: | ODF Monitoring is missing some of the metric values 4.14 | ||
|---|---|---|---|
| Product: | [Red Hat Storage] Red Hat OpenShift Data Foundation | Reporter: | Daniel Osypenko <dosypenk> |
| Component: | ceph-monitoring | Assignee: | Divyansh Kamboj <dkamboj> |
| Status: | CLOSED ERRATA | QA Contact: | Daniel Osypenko <dosypenk> |
| Severity: | high | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 4.14 | CC: | athakkar, branto, dkamboj, ebenahar, edonnell, fbalak, jolmomar, kdreyer, kramdoss, muagarwa, nthomas, odf-bz-bot, sheggodu, tnielsen |
| Target Milestone: | --- | Keywords: | Regression |
| Target Release: | ODF 4.14.6 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | 4.14.6-1 | Doc Type: | No Doc Update |
| Doc Text: | Story Points: | --- | |
| Clone Of: | 2221488 | Environment: | |
| Last Closed: | 2024-04-01 09:17:35 UTC | Type: | --- |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | 2258861 | ||
| Bug Blocks: | 2242324, 2244409 | ||
|
Description
Daniel Osypenko
2023-12-07 11:59:10 UTC
*** Bug 2253428 has been marked as a duplicate of this bug. *** Avan, please don't change the bug state to ON_QA until the discussion is concluded and bug has all the acks. Also, this bug kept moving from engineering to QA and vice versa. Can we please setup a meeting to discuss and close it because the turnaround time to have such discussion via bug page is too much. Filip, if the fix is not working in 4.14.z then we need a separate bug for 4.14.z I have reproduced an issue on post-upgrade deployment of the IBM cloud cluster and it very good represents the failure history of the tests test_ceph_metrics_available and test_ceph_rbd_metrics_available and may explain why we did not see it on life deployment. Issue happens only when we have OCP 4.15 on both ODF 4.14 and 4.15. Issue happens only with ceph metrics (not with rbd metrics). Issue happened on platforms: IBM cloud, Azure, AWS and GCP. Tested: Upgrade from OCP 4.14 & ODF 4.14 to OCP 4.14 & ODF 4.14 and pos-upgrade check Before upgrade data from the metrics were available After upgrade data are not available for 167 metrics Along with this ocs-storagecluster is in progressing, Data resiliency in progressing, one worker node is not ready (VM is running, no errors on odf-operator-controller-manager, no errors on ocs-metrics-exporter rather than health report issues) Issue also observed recently on vSphere OCP 4.15 & ODF 4.15 NOT-post-upgrade oc get pods -n openshift-storage NAME READY STATUS RESTARTS AGE csi-addons-controller-manager-58746c86d6-qjggl 2/2 Running 0 110m csi-cephfsplugin-gzgxr 2/2 Running 2 31h csi-cephfsplugin-provisioner-85f5789c76-pn2vh 5/5 Running 0 3h14m csi-cephfsplugin-provisioner-85f5789c76-pnltj 5/5 Running 0 166m csi-cephfsplugin-thprw 2/2 Running 2 31h csi-cephfsplugin-zjlj7 2/2 Running 0 31h csi-rbdplugin-6cblq 3/3 Running 3 3h27m csi-rbdplugin-6r674 3/3 Running 0 3h26m csi-rbdplugin-provisioner-5fb5cc859b-g94br 6/6 Running 0 3h14m csi-rbdplugin-provisioner-5fb5cc859b-qts6t 6/6 Running 0 166m csi-rbdplugin-qlc7x 3/3 Running 3 3h27m noobaa-core-0 1/1 Running 0 166m noobaa-db-pg-0 1/1 Running 0 166m noobaa-endpoint-845d6d9998-lz296 1/1 Running 0 166m noobaa-operator-5bcf546c-mmlpr 2/2 Running 0 166m ocs-metrics-exporter-64755696fb-766qm 1/1 Running 0 3h14m ocs-operator-78c8fb9446-4g9x8 1/1 Running 2 (177m ago) 3h14m odf-console-76b8fd5784-wptm4 1/1 Running 0 166m odf-operator-controller-manager-7bff4bf5cf-5ldz9 2/2 Running 0 166m rook-ceph-crashcollector-dosypenk-281-i-fd2hc-worker-1-9kvcf5r5 1/1 Running 0 3h14m rook-ceph-crashcollector-dosypenk-281-i-fd2hc-worker-2-8zwczqwf 1/1 Running 0 166m rook-ceph-exporter-dosypenk-281-i-fd2hc-worker-1-9kv6q-689sx5vx 1/1 Running 0 3h14m rook-ceph-exporter-dosypenk-281-i-fd2hc-worker-2-8zwns-874khbrp 1/1 Running 0 166m rook-ceph-mds-ocs-storagecluster-cephfilesystem-a-7f6766b5774n4 2/2 Running 11 (150m ago) 3h14m rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5b6dd7cb2hfsz 2/2 Running 5 (148m ago) 164m rook-ceph-mgr-a-7bcc5c969-xzsm2 2/2 Running 0 166m rook-ceph-mon-a-d7d7bf65f-hjk48 2/2 Running 0 3h27m rook-ceph-mon-b-8f64c96cf-gqsll 2/2 Running 0 167m rook-ceph-mon-c-6f6f79d5dc-ztv72 0/2 Pending 0 5m30s rook-ceph-operator-7b7b6b8d5c-c26t5 1/1 Running 0 3h14m rook-ceph-osd-0-66948789f4-klprs 2/2 Running 0 3h12m rook-ceph-osd-1-79b6766cff-5nltx 0/2 Pending 0 163m rook-ceph-osd-2-747c74d944-xzzh6 2/2 Running 0 3h27m rook-ceph-tools-57fd4d4d68-9kjgw 1/1 Running 0 3h14m ocs must-gather stuck. adding an OCP must-gather and partially OCS must-gather ocp must-gather https://drive.google.com/file/d/16MisoUMeBZJ--Ilju5wTsBptwXUyDq21/view?usp=sharing ocs must-gather https://drive.google.com/file/d/1iviZ_tPOPlfEy-leX1yo05XF0mHnUrng/view?usp=sharing root cause is similar to https://bugzilla.redhat.com/show_bug.cgi?id=2258861. Divyansh with link the backport PR(4.14.z) and move the BZ to POST OCP 4.14.0-0.nightly-2024-03-11-023324 OCS 4.14.6-1 tests passed: test_ceph_metrics_available - https://ocs4-jenkins-csb-odf-qe.apps.ocp-c1.prod.psi.redhat.com/job/qe-deploy-ocs-cluster/34898/consoleFull test_ceph_rbd_metrics_available - https://ocs4-jenkins-csb-odf-qe.apps.ocp-c1.prod.psi.redhat.com/job/qe-deploy-ocs-cluster/34896/console *** Bug 2262307 has been marked as a duplicate of this bug. *** Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (Red Hat OpenShift Data Foundation 4.14.6 Bug Fix Update), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2024:1579 The needinfo request[s] on this closed bug have been removed as they have been unresolved for 120 days |