Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: The webserve volume is a 13x(2+1) Distributed-Replicate volume. It was running out of space with 10 replica sets, so three more were added. A volume rebalance was started to make use of the three new replica sets. When the rebalance was started, the CPU load on the gluster nodes and the gluster clients went very high. The clients were no longer able to access the data from the webserve volume. The volume rebalanced was stopped, but even after the rebalance status showed no activity, the high CPU loads continued. The gluster nodes were rebooted (serially), but that didn't help. The customer said the clients were rebooted (though we didn't see that in the one client sosreport). Eventually, the customer switched the clients to older storage. This is not a good look for Red Hat. With the clients disconnected, the customer will run the rebalance and let it complete. There is another volume on the cluster, so the customer does not want to be too aggressive with the rebalance. sosreports were collected from the gluster nodes and one client. We need help from engineering on determining the root cause of the client disconnects and timeouts. This customer will be very wary running a rebalance again. Version-Release number of selected component (if applicable): RHGS 3.4 on gluster nodes and clients How reproducible: We don't know if this is reproducible or a one-time problem. Steps to Reproduce: 1. 2. 3. Actual results: High CPU load balances on the gluster nodes and clients. Clients unable to access data from the gluster volume. Expected results: Normal operation during a rebalance with a "normal" rebal-throttle. Additional info: Sosreports collected from the three gluster nodes and on client. Ping timeouts observed on the client logs.
I see client ping-timeouts and connection timeout errors! Looks like the issue is similar to the previously reported ping-timeout issues. https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to see what are the options we need to set on the volumes for mitigating some of these.
(In reply to Amar Tumballi from comment #4) > I see client ping-timeouts and connection timeout errors! Looks like the > issue is similar to the previously reported ping-timeout issues. > > https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to > see what are the options we need to set on the volumes for mitigating some > of these. The ping-timeouts could be a side-effect of high CPU usage and may not be the root cause. If bricks are unresponsive due to high CPU usage, its likely that they can't respond back to pings from clients resulting in ping timer expiry on clients.
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2019:0658