Bug 1614005 - It's not clear what Weeks Remaining panel on Host dashboard mean
Summary: It's not clear what Weeks Remaining panel on Host dashboard mean
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: web-admin-tendrl-monitoring-integration
Version: rhgs-3.4
Hardware: Unspecified
OS: Unspecified
unspecified
unspecified
Target Milestone: ---
: RHGS 3.4.z Batch Update 1
Assignee: gowtham
QA Contact: Elena Bondarenko
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2018-08-08 19:43 UTC by Martin Bukatovic
Modified: 2018-10-31 08:45 UTC (History)
12 users (show)

Fixed In Version: tendrl-monitoring-integration-1.6.3-13.el7rhgs
Doc Type: Bug Fix
Doc Text:
The Weeks Remaining panel in the Host and Cluster dashboard of Grafana did not display complete and accurate metrics about the remaining weeks for volumes utilization due to inaccurate predictions. To avoid projecting any false or inaccurate metrics, the Weeks Remaining panel is now removed from the Host and Cluster Dashboard of Grafana.
Clone Of:
Environment:
Last Closed: 2018-10-31 08:45:18 UTC
Embargoed:


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Github Tendrl monitoring-integration issues 564 0 None None None 2018-09-11 11:39:21 UTC
Red Hat Product Errata RHBA-2018:3427 0 None None None 2018-10-31 08:45:49 UTC

Description Martin Bukatovic 2018-08-08 19:43:14 UTC
Description of problem
======================

When a machine hosts bricks of multiple volumes, what does the value reported
here mean? 

I mean, imagine you are saving incoming data on a gluster volume, but other
volumes maintains it's utilization stable. Now you see some number reported
on Weeks Remaining panel and wonder: does it take into account the fact that
only bricks of the volume with utilization going up will be able to receive
the data? Or does it ignore this limitation?

Version-Release number of selected component
============================================

tendrl-monitoring-integration-1.6.3-7.el7rhgs.noarch

Steps to Reproduce
==================

1. Instal RHGS WA using tendrl-ansible
2. Import Trusted storage pool with at least two volumes, so that each
   storage machine hosts bricks of both volumes
3. Run a workload, uploading data on one volume, while keep the other volume
   idle
4. Go to Host dashboard and check Weeks Remaining panel

Actual results
==============

The description of the panel:

> The Weeks Remaining panel displays the estimated time remaining in
> weeks till host capacity reaches full capacity based on the forecasted
> Weekly Growth Rate.

Doesn't explain the way Weeks Remaining prediction is computed when host
contains bricks of multiple volumes. Does it take into account that
extrapolated growth wrt the volume?

Expected results
================

It's clear how the data are aggregated, when storage machine hosts bricks of
multiple volumes.

Additional question
===================

Which of the solution is correct (wrt what would gluster administrators like
to see reported here) remains not specified.

Comment 1 Martin Bukatovic 2018-08-09 14:58:03 UTC
The same problem applies for "Weeks Remaining" panel on Cluster dashboard.

Comment 3 Ankush Behl 2018-09-05 06:48:08 UTC
(In reply to Martin Bukatovic from comment #1)
> The same problem applies for "Weeks Remaining" panel on Cluster dashboard.

@martin - can you explain why it doesn't make sense on cluster dashboard.

I agree it doesn't make sense on the host dashboard because we are showing weeks remaining for the bricks that are there on the host. We will remove it from host dashboard.

Comment 4 Martin Bukatovic 2018-09-06 18:35:08 UTC
(In reply to Ankush Behl from comment #3)
> (In reply to Martin Bukatovic from comment #1)
> > The same problem applies for "Weeks Remaining" panel on Cluster dashboard.
> 
> @martin - can you explain why it doesn't make sense on cluster dashboard.

Imagine you are saving incoming data on a gluster volume, but other
volumes maintains it's utilization stable. Now you see some number reported
on Weeks Remaining panel for the whole cluster and wonder: does it take into
account the fact that only volume with utilization going up will be able to
receive the data?

Let's demonstrate this on particular example. You have a cluster with volumes:

 * volume A, which is filling up, and WA reports just 2 weeks as remaining
 * volume B, which is not filling up, maintains stable utilization, and WA
   doesn't try to guess this because there is not enough data

Now, what should "Weeks remaining" for the whole cluster be? If you just
sum up the values for remaining free space and incoming data, the resulting
number would not make sense as the data which are uploaded on volume A would
not be magically stored on volume B instead when we run out of free space on
volume A. So I guess this is not the case. But then the question is, how is
this number actually calculated?

> I agree it doesn't make sense on the host dashboard because we are showing
> weeks remaining for the bricks that are there on the host. We will remove it
> from host dashboard.

Ack, I have no problem with that. That said, we need to ask Anand.

Comment 5 Martin Bukatovic 2018-09-07 13:29:21 UTC
Raising needinfo for 1st part of comment 4 (weeks remaining on Cluster
Dashboard).

Comment 6 Anand Paladugu 2018-09-09 17:09:14 UTC
Weeks remaining on whole cluster does not make sense, unless the assumption is that all volumes and bricks are filling up at the same rate ...  It has to be on a volume by volume basis.

Comment 7 Anand Paladugu 2018-09-09 17:10:20 UTC
Plus from a customer perspective, consumption is from a volume perspective.

Comment 8 Ju Lim 2018-09-10 11:26:58 UTC
Action Item: Remove Weeks Remaining panel on Cluster and Host Dashboard.

Comment 9 Ju Lim 2018-09-10 11:28:14 UTC
remove weeks remaining and forecasted weekly from host and cluster dashboard, stretch remaining panels on the row.

Comment 10 Nishanth Thomas 2018-09-10 15:50:27 UTC
Providing the acks based on https://bugzilla.redhat.com/show_bug.cgi?id=1614005#c9

Comment 13 gowtham 2018-09-11 10:13:43 UTC
PR is under review: https://github.com/Tendrl/monitoring-integration/pull/565

Comment 15 Elena Bondarenko 2018-09-25 12:48:37 UTC
There's no Weeks Remaining panel on either Host or Cluster dashboard.

Comment 19 Anmol Sachan 2018-10-23 09:37:36 UTC
Looks good.

Comment 20 gowtham 2018-10-23 11:23:18 UTC
Looks good to me

Comment 23 errata-xmlrpc 2018-10-31 08:45:18 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2018:3427


Note You need to log in before you can comment on or make changes to this bug.