Bug 1144431 - glusterd crash seen when remove-brick was run followed by add-brick and rebalance
Summary: glusterd crash seen when remove-brick was run followed by add-brick and rebal...
Keywords:
Status: CLOSED DUPLICATE of bug 1236986
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: glusterd
Version: rhgs-3.0
Hardware: Unspecified
OS: Unspecified
unspecified
high
Target Milestone: ---
: ---
Assignee: Atin Mukherjee
QA Contact: SATHEESARAN
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2014-09-19 11:18 UTC by Shruti Sampat
Modified: 2019-07-11 08:12 UTC (History)
10 users (show)

Fixed In Version:
Doc Type: Bug Fix
Doc Text:
Clone Of:
Environment:
Last Closed: 2015-07-16 03:59:23 UTC
Embargoed:


Attachments (Terms of Use)

Description Shruti Sampat 2014-09-19 11:18:57 UTC
Description of problem:
------------------------

In a cluster of 4 nodes, that was being managed by RHSC, remove-brick operation was executed on 2 volumes (one distributed and another distributed-replicate). The bricks that were removed were again added with force option, followed by rebalance. glusterd process was found to have crashed on one of the nodes.

Version-Release number of selected component (if applicable):
--------------------------------------------------------------

glusterfs-3.6.0.28-1.el6rhs.x86_64

How reproducible:
------------------

Saw it once.

Steps to Reproduce:
--------------------

1. Create a distribute volume of 2 bricks, and a 2x2 distribute-replicate volume.
2. Run remove-brick operation on both volumes, removing one of the bricks from the distribute volume, and a pair of bricks from the distributed-replicate volume.
3. Add the same bricks that were removed with force.
4. Run rebalance on both volumes.

Actual results:
----------------

I was trying to see the rebalance status using the UI, when the status failed to appear, as glusterd had crashed.

Expected results:
------------------

glusterd is not expected to crash.

Additional info:
-----------------

The nodes had been running for about 3-4 days, and rebalance/remove-brick/add-brick had been performed on these volumes multiple times.

Find core file attached.

Comment 3 Vijay Bellur 2014-09-19 12:24:46 UTC
Relevant logs and backtrace from sosreport of 10.70.37.138:

[2014-09-19 09:49:36.924467] E [glusterd-utils.c:10457:glusterd_volume_rebalance_use_rsp_dict] (-->/usr/lib64/libgfrpc.so.0(rpc_clnt_handle_reply+0xa5) [0x3bdd20e9c5] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_big_locked_cbk+0x60) [0x7f4bcdabd970] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(__glusterd_commit_op_cbk+0x6c4) [0x7f4bcdac0084]))) 0-: Assertion failed: (GD_OP_REBALANCE == op) || (GD_OP_DEFRAG_BRICK_VOLUME == op)
[2014-09-19 09:49:36.924704] E [mem-pool.c:242:__gf_free] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_volume_rebalance_use_rsp_dict+0xdc) [0x7f4bcdaa5e8c] (-->/usr/lib64/libglusterfs.so.0(dict_get_str+0x6e) [0x3bdca1ac4e] (-->/usr/lib64/libglusterfs.so.0(data_destroy+0x55) [0x3bdca1a9d5]))) 0-: Assertion failed: GF_MEM_HEADER_MAGIC == *(uint32_t *)ptr
[2014-09-19 09:49:36.924822] E [mem-pool.c:265:__gf_free] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_volume_rebalance_use_rsp_dict+0xdc) [0x7f4bcdaa5e8c] (-->/usr/lib64/libglusterfs.so.0(dict_get_str+0x6e) [0x3bdca1ac4e] (-->/usr/lib64/libglusterfs.so.0(data_destroy+0x55) [0x3bdca1a9d5]))) 0-: Assertion failed: GF_MEM_TRAILER_MAGIC == *(uint32_t *)((char *)free_ptr + req_size)
pending frames:
frame : type(0) op(0)
frame : type(0) op(0)
frame : type(0) op(0)
patchset: git://git.gluster.com/glusterfs.git
signal received: 11
time of crash: 
2014-09-19 09:49:36
configuration details:
argp 1
backtrace 1
dlfcn 1
libpthread 1
llistxattr 1
setfsid 1
spinlock 1
epoll.h 1
xattr.h 1
st_atim.tv_nsec 1
package-string: glusterfs 3.6.0.28
/usr/lib64/libglusterfs.so.0(_gf_msg_backtrace_nomem+0xb6)[0x3bdca1ff06]
/usr/lib64/libglusterfs.so.0(gf_print_trace+0x33f)[0x3bdca3a59f]
/lib64/libc.so.6[0x3bdb6326b0]
/lib64/libpthread.so.0(pthread_spin_lock+0x0)[0x3bdbe0c380]
/usr/lib64/libglusterfs.so.0(__gf_free+0x14a)[0x3bdca4d50a]
/usr/lib64/libglusterfs.so.0(data_destroy+0x55)[0x3bdca1a9d5]
/usr/lib64/libglusterfs.so.0(dict_get_str+0x6e)[0x3bdca1ac4e]
/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_volume_rebalance_use_rsp_dict+0xdc)[0x7f4bcdaa5e8c]
/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(__glusterd_commit_op_cbk+0x6c4)[0x7f4bcdac0084]
/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_big_locked_cbk+0x60)[0x7f4bcdabd970]
/usr/lib64/libgfrpc.so.0(rpc_clnt_handle_reply+0xa5)[0x3bdd20e9c5]
/usr/lib64/libgfrpc.so.0(rpc_clnt_notify+0x13f)[0x3bdd20fe4f]
/usr/lib64/libgfrpc.so.0(rpc_transport_notify+0x28)[0x3bdd20b668]
/usr/lib64/glusterfs/3.6.0.28/rpc-transport/socket.so(+0x9275)[0x7f4bcd0f9275]
/usr/lib64/glusterfs/3.6.0.28/rpc-transport/socket.so(+0xac5d)[0x7f4bcd0fac5d]
/usr/lib64/libglusterfs.so.0[0x3bdca76367]
/usr/sbin/glusterd(main+0x603)[0x407e93]
/lib64/libc.so.6(__libc_start_main+0xfd)[0x3bdb61ed5d]
/usr/sbin/glusterd[0x4049a9]

Comment 6 Nagaprasad Sathyanarayana 2014-09-22 05:45:11 UTC
As agreed by engineering leads, deferring this from RHS 3.0 release.

Comment 7 SATHEESARAN 2015-01-20 10:18:20 UTC
This bug falls under MUST_FIX list for RHS 3.0.4, and the fix could be verified with regression test cases within stipulated time planned for the release. 

Providing qa_ack

Comment 13 Atin Mukherjee 2015-06-09 07:15:42 UTC
It seems like we need one more fix (#BZ 1229139) to rule out all such possible corruptions. A patch http://review.gluster.org/#/c/11120/ for the same is posted in upstream for review.


Note You need to log in before you can comment on or make changes to this bug.