Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: On a brick-mux enabled three nodes cluster, Performed add-brick on a distributed-disperse volume and started rebalance, while rebalance is in-progress powered off one of the node(n3) for few hours to try to replicate BZ1818836. Once the node is up and running peer status is showing disconnected from node n1 for the node which went down and seeing the rebalance status as failed. The gluster vol rebalance <vol> status is not uniform across the nodes due to peer is in disconnected state from one node. Version-Release number of selected component (if applicable): glusterfs-6.0-32 How reproducible: 2/2 Steps to Reproduce: 1. On a three node cluster, enabled brick mux 2. Created a distributed-disperse volume 3X(4+2) and two replicated volumes (1X3) 3. Mounted on 8 clients and started running IO and lookups 4. Performed add-brick on distributed-disperse volume and started rebalance 5. While rebalance is in-progress, powered off a node(say n3) for 7-9 hours approx. 6. Powered on the node n3 Actual results: Once the node n3 is up, rebalance status of n3 is showing as failed. gluster peer status is showing as disconnected on one of the node for the node n3 Expected results: peers should be connected once the node is up and rebalance should continue and should not go into failed state Additional info: ------------------------8<-------------------------- rebalance log snippet from the node where it got into failed 641: end-volume 642: +------------------------------------------------------------------------------+ [2020-04-14 13:35:00.895100] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-22: Connected to t-shelby-client-22, attached to remote volume '/bricks/brick9/ron-add'. [2020-04-14 13:35:00.896018] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-21: Connected to t-shelby-client-21, attached to remote volume '/bricks/brick9/ron-add'. [2020-04-14 13:35:00.897724] I [rpc-clnt.c:2035:rpc_clnt_reconfig] 0-t-shelby-client-23: changing port to 49153 (from 0) [2020-04-14 13:35:00.897792] I [socket.c:871:__socket_shutdown] 0-t-shelby-client-23: intentional socket shutdown(34) [2020-04-14 13:35:00.907263] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-23: Connected to t-shelby-client-23, attached to remote volume '/bricks/brick9/ron-add'. [2020-04-14 13:35:00.907295] I [MSGID: 122062] [ec.c:334:ec_up] 0-t-shelby-disperse-3: Going UP [2020-04-14 13:35:01.001390] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-7: Connected to t-shelby-client-7, attached to remote volume '/bricks/brick4/t-shelby4'. [2020-04-14 13:35:01.001457] I [MSGID: 122062] [ec.c:334:ec_up] 0-t-shelby-disperse-1: Going UP [2020-04-14 13:35:01.559149] I [MSGID: 114046] [client-handshake.c:1105:client_setvolume_cbk] 0-t-shelby-client-13: Connected to t-shelby-client-13, attached to remote volume '/bricks/brick6/t-shelby6'. [2020-04-14 13:35:01.559216] I [MSGID: 122062] [ec.c:334:ec_up] 0-t-shelby-disperse-2: Going UP [2020-04-14 13:35:02.682000] I [dht-rebalance.c:4640:gf_defrag_start_crawl] 0-t-shelby-dht: gf_defrag_start_crawl using commit hash 0 [2020-04-14 13:35:02.884174] W [MSGID: 122053] [ec-common.c:329:ec_check_status] 0-t-shelby-disperse-2: Operation failed on 1 of 6 subvolumes.(up=111111, mask=101111, remaining=000000, good=101111, bad=010000, FOP : 'SETXATTR' failed on '/' with gfid 00000000-0000-0000-0000-000000000001) [2020-04-14 13:35:02.885954] I [MSGID: 109081] [dht-common.c:5872:dht_setxattr] 0-t-shelby-dht: fixing the layout of / [2020-04-14 13:35:03.465077] W [MSGID: 122040] [ec-common.c:1232:ec_prepare_update_cbk] 0-t-shelby-disperse-2: Failed to get size and version : FOP : 'XATTROP' failed on '(null)' with gfid 00000000-0000-0000-0000-000000000001 [Input/output error] [2020-04-14 13:35:03.465163] E [MSGID: 109039] [dht-common.c:4260:dht_find_local_subvol_cbk] 0-t-shelby-dht: failed to get node-uuid [Input/output error] [2020-04-14 13:35:03.465245] E [MSGID: 0] [dht-rebalance.c:4242:dht_init_local_subvols_and_nodeuuids] 0-t-shelby-dht: local subvolume determination failed with error: 5 [Input/output error] [2020-04-14 13:35:03.466206] I [MSGID: 109028] [dht-rebalance.c:5060:gf_defrag_status_get] 0-t-shelby-dht: Rebalance is failed. Time taken is 2.00 secs [2020-04-14 13:35:03.466239] I [MSGID: 109028] [dht-rebalance.c:5066:gf_defrag_status_get] 0-t-shelby-dht: Files migrated: 0, size: 0, lookups: 0, failures: 0, skipped: 0 [2020-04-14 13:35:03.468534] W [glusterfsd.c:1581:cleanup_and_exit] (-->/lib64/libpthread.so.0(+0x82de) [0x7f407adc32de] -->/usr/sbin/glusterfs(glusterfs_sigwaiter+0xfd) [0x5592cd27786d] -->/usr/sbin/glusterfs(cleanup_and_exit+0x58) [0x5592cd2776b8] ) 0-: received signum (15), shutting down ------------------------8<--------------------------
Created attachment 1682045 [details] Logs-from-server-200(not-restarted)
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days