Bug 1646735
| Summary: | Clients seemed to lose connectivity to gluster volume after rebalance started. | ||
|---|---|---|---|
| Product: | [Red Hat Storage] Red Hat Gluster Storage | Reporter: | Andrew Robinson <anrobins> |
| Component: | core | Assignee: | Milind Changire <mchangir> |
| Status: | CLOSED ERRATA | QA Contact: | Sayalee <saraut> |
| Severity: | high | Docs Contact: | |
| Priority: | high | ||
| Version: | rhgs-3.4 | CC: | amukherj, anrobins, atumball, bkunal, ccalhoun, kdhananj, mchangir, moagrawa, nchilaka, olim, rgowdapp, rhs-bugs, sankarshan, saraut, storage-qa-internal |
| Target Milestone: | --- | Keywords: | Question, ZStream |
| Target Release: | RHGS 3.4.z Batch Update 4 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | rpc-ping-timeout | ||
| Fixed In Version: | glusterfs-3.12.2-41 | Doc Type: | If docs needed, set a value |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2019-03-27 03:43:39 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | |||
| Bug Blocks: | 1649191, 1657798 | ||
|
Description
Andrew Robinson
2018-11-05 22:41:33 UTC
I see client ping-timeouts and connection timeout errors! Looks like the issue is similar to the previously reported ping-timeout issues. https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to see what are the options we need to set on the volumes for mitigating some of these. (In reply to Amar Tumballi from comment #4) > I see client ping-timeouts and connection timeout errors! Looks like the > issue is similar to the previously reported ping-timeout issues. > > https://red.ht/2PLmJSg captures other rpc ping timeout issues. We need to > see what are the options we need to set on the volumes for mitigating some > of these. The ping-timeouts could be a side-effect of high CPU usage and may not be the root cause. If bricks are unresponsive due to high CPU usage, its likely that they can't respond back to pings from clients resulting in ping timer expiry on clients. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2019:0658 |