Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: ------------------------ In a cluster of 4 nodes, that was being managed by RHSC, remove-brick operation was executed on 2 volumes (one distributed and another distributed-replicate). The bricks that were removed were again added with force option, followed by rebalance. glusterd process was found to have crashed on one of the nodes. Version-Release number of selected component (if applicable): -------------------------------------------------------------- glusterfs-3.6.0.28-1.el6rhs.x86_64 How reproducible: ------------------ Saw it once. Steps to Reproduce: -------------------- 1. Create a distribute volume of 2 bricks, and a 2x2 distribute-replicate volume. 2. Run remove-brick operation on both volumes, removing one of the bricks from the distribute volume, and a pair of bricks from the distributed-replicate volume. 3. Add the same bricks that were removed with force. 4. Run rebalance on both volumes. Actual results: ---------------- I was trying to see the rebalance status using the UI, when the status failed to appear, as glusterd had crashed. Expected results: ------------------ glusterd is not expected to crash. Additional info: ----------------- The nodes had been running for about 3-4 days, and rebalance/remove-brick/add-brick had been performed on these volumes multiple times. Find core file attached.
Relevant logs and backtrace from sosreport of 10.70.37.138: [2014-09-19 09:49:36.924467] E [glusterd-utils.c:10457:glusterd_volume_rebalance_use_rsp_dict] (-->/usr/lib64/libgfrpc.so.0(rpc_clnt_handle_reply+0xa5) [0x3bdd20e9c5] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_big_locked_cbk+0x60) [0x7f4bcdabd970] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(__glusterd_commit_op_cbk+0x6c4) [0x7f4bcdac0084]))) 0-: Assertion failed: (GD_OP_REBALANCE == op) || (GD_OP_DEFRAG_BRICK_VOLUME == op) [2014-09-19 09:49:36.924704] E [mem-pool.c:242:__gf_free] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_volume_rebalance_use_rsp_dict+0xdc) [0x7f4bcdaa5e8c] (-->/usr/lib64/libglusterfs.so.0(dict_get_str+0x6e) [0x3bdca1ac4e] (-->/usr/lib64/libglusterfs.so.0(data_destroy+0x55) [0x3bdca1a9d5]))) 0-: Assertion failed: GF_MEM_HEADER_MAGIC == *(uint32_t *)ptr [2014-09-19 09:49:36.924822] E [mem-pool.c:265:__gf_free] (-->/usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_volume_rebalance_use_rsp_dict+0xdc) [0x7f4bcdaa5e8c] (-->/usr/lib64/libglusterfs.so.0(dict_get_str+0x6e) [0x3bdca1ac4e] (-->/usr/lib64/libglusterfs.so.0(data_destroy+0x55) [0x3bdca1a9d5]))) 0-: Assertion failed: GF_MEM_TRAILER_MAGIC == *(uint32_t *)((char *)free_ptr + req_size) pending frames: frame : type(0) op(0) frame : type(0) op(0) frame : type(0) op(0) patchset: git://git.gluster.com/glusterfs.git signal received: 11 time of crash: 2014-09-19 09:49:36 configuration details: argp 1 backtrace 1 dlfcn 1 libpthread 1 llistxattr 1 setfsid 1 spinlock 1 epoll.h 1 xattr.h 1 st_atim.tv_nsec 1 package-string: glusterfs 3.6.0.28 /usr/lib64/libglusterfs.so.0(_gf_msg_backtrace_nomem+0xb6)[0x3bdca1ff06] /usr/lib64/libglusterfs.so.0(gf_print_trace+0x33f)[0x3bdca3a59f] /lib64/libc.so.6[0x3bdb6326b0] /lib64/libpthread.so.0(pthread_spin_lock+0x0)[0x3bdbe0c380] /usr/lib64/libglusterfs.so.0(__gf_free+0x14a)[0x3bdca4d50a] /usr/lib64/libglusterfs.so.0(data_destroy+0x55)[0x3bdca1a9d5] /usr/lib64/libglusterfs.so.0(dict_get_str+0x6e)[0x3bdca1ac4e] /usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_volume_rebalance_use_rsp_dict+0xdc)[0x7f4bcdaa5e8c] /usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(__glusterd_commit_op_cbk+0x6c4)[0x7f4bcdac0084] /usr/lib64/glusterfs/3.6.0.28/xlator/mgmt/glusterd.so(glusterd_big_locked_cbk+0x60)[0x7f4bcdabd970] /usr/lib64/libgfrpc.so.0(rpc_clnt_handle_reply+0xa5)[0x3bdd20e9c5] /usr/lib64/libgfrpc.so.0(rpc_clnt_notify+0x13f)[0x3bdd20fe4f] /usr/lib64/libgfrpc.so.0(rpc_transport_notify+0x28)[0x3bdd20b668] /usr/lib64/glusterfs/3.6.0.28/rpc-transport/socket.so(+0x9275)[0x7f4bcd0f9275] /usr/lib64/glusterfs/3.6.0.28/rpc-transport/socket.so(+0xac5d)[0x7f4bcd0fac5d] /usr/lib64/libglusterfs.so.0[0x3bdca76367] /usr/sbin/glusterd(main+0x603)[0x407e93] /lib64/libc.so.6(__libc_start_main+0xfd)[0x3bdb61ed5d] /usr/sbin/glusterd[0x4049a9]
As agreed by engineering leads, deferring this from RHS 3.0 release.
This bug falls under MUST_FIX list for RHS 3.0.4, and the fix could be verified with regression test cases within stipulated time planned for the release. Providing qa_ack
It seems like we need one more fix (#BZ 1229139) to rule out all such possible corruptions. A patch http://review.gluster.org/#/c/11120/ for the same is posted in upstream for review.