Bug 1646735

Summary: Clients seemed to lose connectivity to gluster volume after rebalance started.
Product: [Red Hat Storage] Red Hat Gluster Storage Reporter: Andrew Robinson <anrobins>
Component: coreAssignee: Milind Changire <mchangir>
Status: CLOSED ERRATA QA Contact: Sayalee <saraut>
Severity: high Docs Contact:
Priority: high    
Version: rhgs-3.4CC: amukherj, anrobins, atumball, bkunal, ccalhoun, kdhananj, mchangir, moagrawa, nchilaka, olim, rgowdapp, rhs-bugs, sankarshan, saraut, storage-qa-internal
Target Milestone: ---Keywords: Question, ZStream
Target Release: RHGS 3.4.z Batch Update 4   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard: rpc-ping-timeout
Fixed In Version: glusterfs-3.12.2-41 Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2019-03-27 03:43:39 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 1649191, 1657798    

Description Andrew Robinson 2018-11-05 22:41:33 UTC
Description of problem: 

The webserve volume is a 13x(2+1) Distributed-Replicate volume. It was running out of space with 10 replica sets, so three more were added. A volume rebalance was started to make use of the three new replica sets. When the rebalance was started, the CPU load on the gluster nodes and the gluster clients went very high. The clients were no longer able to access the data from the webserve volume. 

The volume rebalanced was stopped, but even after the rebalance status showed no activity, the high CPU loads continued. The gluster nodes were rebooted (serially), but that didn't help. The customer said the clients were rebooted (though we didn't see that in the one client sosreport). 

Eventually, the customer switched the clients to older storage. This is not a good look for Red Hat.


With the clients disconnected, the customer will run the rebalance and let it complete. There is another volume on the cluster, so the customer does not want to be too aggressive with the rebalance. 


sosreports were collected from the gluster nodes and one client. We need help from engineering on determining the root cause of the client disconnects and timeouts. This customer will be very wary running a rebalance again.



Version-Release number of selected component (if applicable):

RHGS 3.4 on gluster nodes and clients


How reproducible:

We don't know if this is reproducible or a one-time problem.


Steps to Reproduce:
1.
2.
3.

Actual results:

High CPU load balances on the gluster nodes and clients. Clients unable to access data from the gluster volume.


Expected results:


Normal operation during a rebalance with a "normal" rebal-throttle.


Additional info:


Sosreports collected from the three gluster nodes and on client.

Ping timeouts observed on the client logs.

Comment 4 Amar Tumballi 2018-11-06 04:12:06 UTC
I see client ping-timeouts and connection timeout errors! Looks like the issue is similar to the previously reported ping-timeout issues.

https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to see what are the options we need to set on the volumes for mitigating some of these.

Comment 7 Raghavendra G 2018-11-07 02:58:49 UTC
(In reply to Amar Tumballi from comment #4)
> I see client ping-timeouts and connection timeout errors! Looks like the
> issue is similar to the previously reported ping-timeout issues.
> 
> https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to
> see what are the options we need to set on the volumes for mitigating some
> of these.

The ping-timeouts could be a side-effect of high CPU usage and may not be the root cause. If bricks are unresponsive due to high CPU usage, its likely that they can't respond back to pings from clients resulting in ping timer expiry on clients.

Comment 33 errata-xmlrpc 2019-03-27 03:43:39 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2019:0658