Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1944792

Summary: [RFE] Implement Ceph health status in unified observability
Product: Red Hat OpenStack Reporter: Alexon Oliveira <alolivei>
Component: collectd-sensubilityAssignee: Jamie Parker <jamparke>
Status: CLOSED WONTFIX QA Contact: Leonid Natapov <lnatapov>
Severity: low Docs Contact:
Priority: low    
Version: 16.1 (Train)CC: cgussobo, dsilvaju, jjoyce, jschluet, lhh, lmadsen, mburns, mgarciac, shrjoshi
Target Milestone: ---Keywords: FutureFeature, Triaged
Target Release: ---   
Hardware: x86_64   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2023-05-17 19:50:43 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1984193    
Bug Blocks:    

Description Alexon Oliveira 2021-03-30 16:57:07 UTC
Description of problem:

Currently, for STF 1.0 in OSP 13, and also for the later versions (1.1 and 1.2 in OSP 16.x), there is no metric for the Ceph Health status available in the STF Prometheus in order one to be able to create alerts. There are some BZ filled to integrate Ceph metrics altogether with STF, but we do not know if this metric (regarding Ceph Health status) is going to be included and it would be good to have it.

Version-Release number of selected component (if applicable):

STF 1.0 in OSP 13
STF 1.1 / 1.2 in OSP 16.x

Actual results:

No Ceph Health status metrics available

Expected results:

Get Ceph Health status metrics available to create alerts and monitoring for the currently version of STF or the next ones (i.e. If Ceph Health status != "OK", then an alarm will be triggered)

Comment 1 Leif Madsen 2021-07-20 22:59:25 UTC
I believe this will be relatively straight forward (using collectd ceph plugin to create the metrics for a dashboard; this issue is _not_ for replacing the existing Ceph Dashboard shipped with ceph-ansible) once 1984193 is resolved.

Comment 2 Leif Madsen 2021-07-22 20:52:01 UTC
Actually I misspoke because I didn't quite realize this was about capturing the health status as reported by Ceph itself (HEALTH_OK, HEALTH_WARN, etc).

This might be something we can implement in to sensubility, or this would be something added to the ceph plugin in collectd.

I'm going to move this issue to OSP actually because I don't think there is much (or anything) to do on the STF side of things. This is really a data collector issue.

Comment 9 Leif Madsen 2023-05-17 19:50:43 UTC
I am closing this feature for RHOSP 17.1 due to capacity issues and future alignment. For RHOSP 18, we will track this separately as part of an overall observability approach.

For storage monitoring within the Ceph domain for RHOSP 17.1 and earlier, please utilize the built-in monitoring system provided by a Ceph deployment, as documented at https://access.redhat.com/documentation/en-us/red_hat_openstack_platform/17.0/html/deploying_red_hat_ceph_storage_and_red_hat_openstack_platform_together_with_director/assembly_adding-rhcs-dashboard-to-overcloud_deployingcontainerizedrhcs#doc-wrapper