Bug 1824179 - [RHEL 8.2] Rebalance status is showing failed on the node when the node is powered off and turned back on
Summary: [RHEL 8.2] Rebalance status is showing failed on the node when the node is po...
Keywords:
Status: CLOSED DUPLICATE of bug 1832306
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: disperse
Version: rhgs-3.5
Hardware: x86_64
OS: Linux
unspecified
high
Target Milestone: ---
: ---
Assignee: Ashish Pandey
QA Contact: Bala Konda Reddy M
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2020-04-15 13:49 UTC by Bala Konda Reddy M
Modified: 2023-09-14 05:55 UTC (History)
8 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2020-09-22 13:07:58 UTC
Embargoed:


Attachments (Terms of Use)
Logs-from-server-200(not-restarted) (14.62 MB, application/gzip)
2020-04-27 07:07 UTC, Susant Kumar Palai
no flags Details

Description Bala Konda Reddy M 2020-04-15 13:49:09 UTC
Description of problem:
On a brick-mux enabled three nodes cluster, Performed add-brick on a distributed-disperse volume and started rebalance, while rebalance is in-progress powered off one of the node(n3) for few hours to try to replicate BZ1818836. 
Once the node is up and running peer status is showing disconnected from node n1 for the node which went down and seeing the rebalance status as failed. The gluster vol rebalance <vol> status is not uniform across the nodes due to peer is in disconnected state from one node.


Version-Release number of selected component (if applicable):
glusterfs-6.0-32

How reproducible:
2/2

Steps to Reproduce:
1. On a three node cluster, enabled brick mux
2. Created a distributed-disperse volume 3X(4+2) and two replicated volumes (1X3) 
3. Mounted on 8 clients and started running IO and lookups 
4. Performed add-brick on distributed-disperse volume and started rebalance
5. While rebalance is in-progress, powered off a node(say n3) for 7-9 hours approx.
6. Powered on the node n3

Actual results:
Once the node n3 is up, rebalance status of n3 is showing as failed.
gluster peer status is showing as disconnected on one of the node for the node n3


Expected results:
peers should be connected once the node is up and rebalance should continue and should not go into failed state


Additional info:

------------------------8<-------------------------- 
rebalance log snippet from the node where it got into failed

641: end-volume
642:
+------------------------------------------------------------------------------+
[2020-04-14 13:35:00.895100] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-22: Connected to t-shelby-client-22, attached to remote volume '/bricks/brick9/ron-add'.
[2020-04-14 13:35:00.896018] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-21: Connected to t-shelby-client-21, attached to remote volume '/bricks/brick9/ron-add'.
[2020-04-14 13:35:00.897724] I [rpc-clnt.c:2035:rpc_clnt_reconfig] 0-t-shelby-client-23: changing port to 49153 (from 0)
[2020-04-14 13:35:00.897792] I [socket.c:871:__socket_shutdown] 0-t-shelby-client-23: intentional socket shutdown(34)
[2020-04-14 13:35:00.907263] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-23: Connected to t-shelby-client-23, attached to remote volume '/bricks/brick9/ron-add'.
[2020-04-14 13:35:00.907295] I [MSGID: 122062] [ec.c:334:ec_up] 0-t-shelby-disperse-3: Going UP
[2020-04-14 13:35:01.001390] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-7: Connected to t-shelby-client-7, attached to remote volume '/bricks/brick4/t-shelby4'.
[2020-04-14 13:35:01.001457] I [MSGID: 122062] [ec.c:334:ec_up] 0-t-shelby-disperse-1: Going UP
[2020-04-14 13:35:01.559149] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-13: Connected to t-shelby-client-13, attached to remote volume '/bricks/brick6/t-shelby6'.
[2020-04-14 13:35:01.559216] I [MSGID: 122062] [ec.c:334:ec_up] 0-t-shelby-disperse-2: Going UP
[2020-04-14 13:35:02.682000] I [dht-rebalance.c:4640:gf_defrag_start_crawl] 0-t-shelby-dht: gf_defrag_start_crawl using commit hash 0
[2020-04-14 13:35:02.884174] W [MSGID: 122053] [ec-common.c:329:ec_check_status] 0-t-shelby-disperse-2: Operation failed on 1 of 6 subvolumes.(up=111111, mask=101111, remaining=000000, good=101111, bad=010000, FOP : 'SETXATTR' failed on '/' with gfid 00000000-0000-0000-0000-000000000001)
[2020-04-14 13:35:02.885954] I [MSGID: 109081] [dht-common.c:5872:dht_setxattr] 0-t-shelby-dht: fixing the layout of /
[2020-04-14 13:35:03.465077] W [MSGID: 122040] [ec-common.c:1232:ec_prepare_update_cbk] 0-t-shelby-disperse-2: Failed to get size and version :  FOP : 'XATTROP' failed on '(null)' with gfid 00000000-0000-0000-0000-000000000001 [Input/output error]
[2020-04-14 13:35:03.465163] E [MSGID: 109039] [dht-common.c:4260:dht_find_local_subvol_cbk] 0-t-shelby-dht: failed to get node-uuid [Input/output error]
[2020-04-14 13:35:03.465245] E [MSGID: 0] [dht-rebalance.c:4242:dht_init_local_subvols_and_nodeuuids] 0-t-shelby-dht: local subvolume determination failed with error: 5 [Input/output error]
[2020-04-14 13:35:03.466206] I [MSGID: 109028] [dht-rebalance.c:5060:gf_defrag_status_get] 0-t-shelby-dht: Rebalance is failed. Time taken is 2.00 secs
[2020-04-14 13:35:03.466239] I [MSGID: 109028] [dht-rebalance.c:5066:gf_defrag_status_get] 0-t-shelby-dht: Files migrated: 0, size: 0, lookups: 0, failures: 0, skipped: 0
[2020-04-14 13:35:03.468534] W [glusterfsd.c:1581:cleanup_and_exit] (-->/lib64/libpthread.so.0(+0x82de) [0x7f407adc32de] -->/usr/sbin/glusterfs(glusterfs_sigwaiter+0xfd) [0x5592cd27786d] -->/usr/sbin/glusterfs(cleanup_and_exit+0x58) [0x5592cd2776b8] ) 0-: received signum (15), shutting down

------------------------8<--------------------------

Comment 9 Susant Kumar Palai 2020-04-27 07:07:17 UTC
Created attachment 1682045 [details]
Logs-from-server-200(not-restarted)

Comment 34 Red Hat Bugzilla 2023-09-14 05:55:30 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days


Note You need to log in before you can comment on or make changes to this bug.