Bug 1646735 - Clients seemed to lose connectivity to gluster volume after rebalance started.
Summary: Clients seemed to lose connectivity to gluster volume after rebalance started.
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: core
Version: rhgs-3.4
Hardware: Unspecified
OS: Unspecified
high
high
Target Milestone: ---
: RHGS 3.4.z Batch Update 4
Assignee: Milind Changire
QA Contact: Sayalee
URL:
Whiteboard: rpc-ping-timeout
Depends On:
Blocks: 1649191 1657798
TreeView+ depends on / blocked
 
Reported: 2018-11-05 22:41 UTC by Andrew Robinson
Modified: 2022-03-13 15:58 UTC (History)
15 users (show)

Fixed In Version: glusterfs-3.12.2-41
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2019-03-27 03:43:39 UTC
Embargoed:


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Red Hat Product Errata RHBA-2019:0658 0 None None None 2019-03-27 03:44:55 UTC

Description Andrew Robinson 2018-11-05 22:41:33 UTC
Description of problem: 

The webserve volume is a 13x(2+1) Distributed-Replicate volume. It was running out of space with 10 replica sets, so three more were added. A volume rebalance was started to make use of the three new replica sets. When the rebalance was started, the CPU load on the gluster nodes and the gluster clients went very high. The clients were no longer able to access the data from the webserve volume. 

The volume rebalanced was stopped, but even after the rebalance status showed no activity, the high CPU loads continued. The gluster nodes were rebooted (serially), but that didn't help. The customer said the clients were rebooted (though we didn't see that in the one client sosreport). 

Eventually, the customer switched the clients to older storage. This is not a good look for Red Hat.


With the clients disconnected, the customer will run the rebalance and let it complete. There is another volume on the cluster, so the customer does not want to be too aggressive with the rebalance. 


sosreports were collected from the gluster nodes and one client. We need help from engineering on determining the root cause of the client disconnects and timeouts. This customer will be very wary running a rebalance again.



Version-Release number of selected component (if applicable):

RHGS 3.4 on gluster nodes and clients


How reproducible:

We don't know if this is reproducible or a one-time problem.


Steps to Reproduce:
1.
2.
3.

Actual results:

High CPU load balances on the gluster nodes and clients. Clients unable to access data from the gluster volume.


Expected results:


Normal operation during a rebalance with a "normal" rebal-throttle.


Additional info:


Sosreports collected from the three gluster nodes and on client.

Ping timeouts observed on the client logs.

Comment 4 Amar Tumballi 2018-11-06 04:12:06 UTC
I see client ping-timeouts and connection timeout errors! Looks like the issue is similar to the previously reported ping-timeout issues.

https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to see what are the options we need to set on the volumes for mitigating some of these.

Comment 7 Raghavendra G 2018-11-07 02:58:49 UTC
(In reply to Amar Tumballi from comment #4)
> I see client ping-timeouts and connection timeout errors! Looks like the
> issue is similar to the previously reported ping-timeout issues.
> 
> https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to
> see what are the options we need to set on the volumes for mitigating some
> of these.

The ping-timeouts could be a side-effect of high CPU usage and may not be the root cause. If bricks are unresponsive due to high CPU usage, its likely that they can't respond back to pings from clients resulting in ping timer expiry on clients.

Comment 33 errata-xmlrpc 2019-03-27 03:43:39 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2019:0658


Note You need to log in before you can comment on or make changes to this bug.