Bug 1204329 - CIFS: [USS]: snapd got crashed after deleting,creating and accessing 256 snapshots in a loop
Summary: CIFS: [USS]: snapd got crashed after deleting,creating and accessing 256 snap...
Keywords:
Status: CLOSED DEFERRED
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: snapshot
Version: rhgs-3.0
Hardware: x86_64
OS: Linux
unspecified
urgent
Target Milestone: ---
: ---
Assignee: rjoseph
QA Contact: storage-qa-internal@redhat.com
URL:
Whiteboard:
Depends On: 1201820
Blocks:
TreeView+ depends on / blocked
 
Reported: 2015-03-21 02:46 UTC by ssamanta
Modified: 2016-09-17 13:06 UTC (History)
6 users (show)

Fixed In Version:
Doc Type: Bug Fix
Doc Text:
Clone Of:
Environment:
Last Closed: 2015-03-23 06:56:47 UTC
Embargoed:


Attachments (Terms of Use)

Description ssamanta 2015-03-21 02:46:21 UTC
Description of problem:
snapd got crashed after deleting (the preexisting 256 snapshots), create 256 snapshots and accessed 256 snaps in a loop.

Version-Release number of selected component (if applicable):
[root@gqas005 ~]# rpm -qa | grep gluster
gluster-nagios-common-0.1.4-1.el6rhs.noarch
glusterfs-fuse-3.6.0.53-1.el6rhs.x86_64
glusterfs-libs-3.6.0.53-1.el6rhs.x86_64
glusterfs-api-3.6.0.53-1.el6rhs.x86_64
glusterfs-cli-3.6.0.53-1.el6rhs.x86_64
glusterfs-geo-replication-3.6.0.53-1.el6rhs.x86_64
gluster-nagios-addons-0.1.14-1.el6rhs.x86_64
rhs-tests-rhs-tests-beaker-rhs-gluster-qe-libs-dev-bturner-2.37-0.noarch
vdsm-gluster-4.14.7.3-1.el6rhs.noarch
glusterfs-3.6.0.53-1.el6rhs.x86_64
glusterfs-server-3.6.0.53-1.el6rhs.x86_64
samba-vfs-glusterfs-4.1.17-4.el6rhs.x86_64
glusterfs-rdma-3.6.0.53-1.el6rhs.x86_64
glusterfs-debuginfo-3.6.0.53-1.el6rhs.x86_64
[root@gqas005 ~]#

How reproducible:
Tried once

Steps to Reproduce:
1.Create 6*2 dist-rep volume and start it
2.Mount the volume and set the volume options for SMB and USS
3.Create 256 snapshots and access from CIFS client
4.In a loop delete 256 preexisting snapshots, Create 256 snapshots, Access the 256 snapshots through a script
5. snapd got crashed 

Actual results:
snapd got crashed

Expected results:
snapd should not crash

Additional info:
http://rhsqe-repo.lab.eng.blr.redhat.com/sosreports/sosreport-gqas005-20150320222245-f0de.tar.xz

core: gqas005.sbu.lab.eng.bos.redhat.com(/tmp/core_glusterfsd.12844)

Dev can use the below systems for futher debugging. The RHS node got crashed 
is gqas005.sbu.lab.eng.bos.redhat.com. 

[root@gqas005 tmp]# cat /etc/glusterfs/glusterd.vol 
volume management
    type mgmt/glusterd
    option working-directory /var/lib/glusterd
    option transport-type socket,rdma
    option transport.socket.keepalive-time 10
    option transport.socket.keepalive-interval 2
    option transport.socket.read-fail-log off
    option ping-timeout 0
    option rpc-auth-allow-insecure on
#   option base-port 49152
end-volume
[root@gqas005 tmp]# 


[root@gqas005 ~]# gluster peer status
Number of Peers: 3

Hostname: gqas012.sbu.lab.eng.bos.redhat.com
Uuid: a5d382c1-691b-4da9-95da-e7eb4ae23aec
State: Peer in Cluster (Connected)

Hostname: gqas009.sbu.lab.eng.bos.redhat.com
Uuid: f5cd1139-11c8-45e3-8105-f74eb1f26e60
State: Peer in Cluster (Connected)

Hostname: gqas006.sbu.lab.eng.bos.redhat.com
Uuid: 4d942257-82cc-4196-ab25-0b40afd06718
State: Peer in Cluster (Connected)
[root@gqas005 ~]# gluster volume status testvol1
Status of volume: testvol1
Gluster process                             TCP Port  RDMA Port  Online  Pid
------------------------------------------------------------------------------
Brick gqas005.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick1/b1                        N/A       N/A        N       N/A  
Brick gqas006.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick2/b2                        49915     0          Y       863  
Brick gqas009.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick3/b3                        49915     0          Y       5410 
Brick gqas012.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick4/b4                        49915     0          Y       8128 
Brick gqas005.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick5/b5                        49940     0          Y       12815
Brick gqas006.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick6/b6                        49916     0          Y       876  
Brick gqas009.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick7/b7                        49916     0          Y       5423 
Brick gqas012.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick8/b8                        49916     0          Y       8141 
Brick gqas005.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick9/b9                        49941     0          Y       12829
Brick gqas006.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick10/b10                      49917     0          Y       889  
Brick gqas009.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick11/b11                      49917     0          Y       5437 
Brick gqas012.sbu.lab.eng.bos.redhat.com:/v
ar/run/gluster/snaps/18bbcb60efcb48bbbf0c6c
9799e1832c/brick12/b12                      49917     0          Y       8154 
Snapshot Daemon on localhost                N/A       N/A        N       12844
NFS Server on localhost                     2049      0          Y       12853
Self-heal Daemon on localhost               N/A       N/A        Y       12861
Quota Daemon on localhost                   N/A       N/A        Y       12873
Snapshot Daemon on gqas006.sbu.lab.eng.bos.
redhat.com                                  49924     0          Y       903  
NFS Server on gqas006.sbu.lab.eng.bos.redha
t.com                                       2049      0          Y       912  
Self-heal Daemon on gqas006.sbu.lab.eng.bos
.redhat.com                                 N/A       N/A        Y       920  
Quota Daemon on gqas006.sbu.lab.eng.bos.red
hat.com                                     N/A       N/A        Y       928  
Snapshot Daemon on gqas009.sbu.lab.eng.bos.
redhat.com                                  49924     0          Y       5451 
NFS Server on gqas009.sbu.lab.eng.bos.redha
t.com                                       2049      0          Y       5461 
Self-heal Daemon on gqas009.sbu.lab.eng.bos
.redhat.com                                 N/A       N/A        Y       5470 
Quota Daemon on gqas009.sbu.lab.eng.bos.red
hat.com                                     N/A       N/A        Y       5481 
Snapshot Daemon on gqas012.sbu.lab.eng.bos.
redhat.com                                  49924     0          Y       8168 
NFS Server on gqas012.sbu.lab.eng.bos.redha
t.com                                       2049      0          Y       8176 
Self-heal Daemon on gqas012.sbu.lab.eng.bos
.redhat.com                                 N/A       N/A        Y       8184 
Quota Daemon on gqas012.sbu.lab.eng.bos.red
hat.com                                     N/A       N/A        Y       8192 
 
Task Status of Volume testvol1
------------------------------------------------------------------------------
There are no active volume tasks
 
[root@gqas005 ~]# df -h /
Filesystem            Size  Used Avail Use% Mounted on
/dev/mapper/vg_gqas005-lv_root
                       50G  8.5G   39G  19% /
[root@gqas005 ~]# free -g
             total       used       free     shared    buffers     cached
Mem:            47          8         38          0          0          6
-/+ buffers/cache:          2         45
Swap:           23          0         23
[root@gqas005 ~]# gluster volume info
 
Volume Name: testvol1
Type: Distributed-Replicate
Volume ID: 3aea9bb7-aa02-407d-860d-c31051e54931
Status: Started
Snap Volume: no
Number of Bricks: 6 x 2 = 12
Transport-type: tcp
Bricks:
Brick1: gqas005.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick1/b1
Brick2: gqas006.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick2/b2
Brick3: gqas009.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick3/b3
Brick4: gqas012.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick4/b4
Brick5: gqas005.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick5/b5
Brick6: gqas006.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick6/b6
Brick7: gqas009.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick7/b7
Brick8: gqas012.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick8/b8
Brick9: gqas005.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick9/b9
Brick10: gqas006.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick10/b10
Brick11: gqas009.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick11/b11
Brick12: gqas012.sbu.lab.eng.bos.redhat.com:/var/run/gluster/snaps/18bbcb60efcb48bbbf0c6c9799e1832c/brick12/b12
Options Reconfigured:
features.quota-deem-statfs: enable
features.quota: on
performance.readdir-ahead: on
features.uss: on
features.show-snapshot-directory: enable
storage.batch-fsync-delay-usec: 0
server.allow-insecure: enable
performance.stat-prefetch: off
auto-delete: disable
snap-max-soft-limit: 90
snap-max-hard-limit: 256



(gdb) bt
#0  0x0000003b2c03f07e in __dentry_grep (table=0x7f1ff0207070, 
    parent=0x7f15759c229c, name=0x7f2113676320 "classes.conf") at inode.c:664
#1  0x0000003b2c03f233 in inode_grep (table=0x7f1ff0207070, 
    parent=0x7f15759c229c, name=0x7f2113676320 "classes.conf") at inode.c:689
#2  0x0000003b2cc11a17 in glfs_resolve_component (fs=<value optimized out>, 
    subvol=0x7f1ff021b6a0, parent=0x7f15759c229c, 
    component=0x7f2113676320 "classes.conf", iatt=0x7f21192bc680, 
    force_lookup=1) at glfs-resolve.c:258
#3  0x0000003b2cc11d77 in glfs_resolve_at (fs=0x7f21136cd2e0, 
    subvol=0x7f1ff021b6a0, at=<value optimized out>, 
    origpath=<value optimized out>, loc=0x7f21192bc820, iatt=0x7f21192bc7b0, 
    follow=0, reval=0) at glfs-resolve.c:376
#4  0x0000003b2cc13f46 in glfs_h_lookupat (fs=0x7f21136cd2e0, 
    parent=<value optimized out>, path=0x7f211417aee9 "classes.conf", 
    stat=0x7f21192bc8e0) at glfs-handleops.c:99
#5  0x00007f211a88aeb9 in svs_lookup_entry (this=0x7f2114005da0, 
    loc=0x7f21241c5f18, buf=0x7f21192bc9e0, postparent=0x7f21192bca50, 
    parent=0x7f2118f4705c, parent_ctx=<value optimized out>, 
    op_errno=0x7f21192bcc7c) at snapview-server.c:295
#6  0x00007f211a88bb72 in svs_get_handle (this=0x7f2114005da0, 
    loc=0x7f21241c5f18, inode_ctx=<value optimized out>, 
    op_errno=0x7f21192bcc7c) at snapview-server.c:1657
#7  0x00007f211a88d504 in svs_revalidate (this=0x7f2114005da0, 
---Type <return> to continue, or q <return> to quit---
    loc=0x7f21241c5f18, parent=0x7f2118f4705c, inode_ctx=0x7f2110645c50, 
    parent_ctx=0x7f21104dda30, buf=0x7f21192bccf0, postparent=0x7f21192bcc80, 
    op_errno=0x7f21192bcc7c) at snapview-server.c:429
#8  0x00007f211a88d99c in svs_lookup (frame=0x7f212473d02c, 
    this=0x7f2114005da0, loc=0x7f21241c5f18, xdata=0x0)
    at snapview-server.c:567
#9  0x0000003b2c02800c in default_lookup_resume (frame=0x7f212473d2dc, 
    this=0x7f2114008830, loc=0x7f21241c5f18, xdata=0x0) at defaults.c:1683
#10 0x0000003b2c043812 in call_resume_wind (stub=0x7f21241c5ed8)
    at call-stub.c:2478
#11 call_resume (stub=0x7f21241c5ed8) at call-stub.c:2841
#12 0x00007f211a680348 in iot_worker (data=0x7f211401b190) at io-threads.c:214
#13 0x00000036602079d1 in start_thread (arg=0x7f21192bd700)
    at pthread_create.c:301
#14 0x000000365fee88fd in clone ()
    at ../sysdeps/unix/sysv/linux/x86_64/clone.S:115
(gdb) 
(gdb) 
(gdb)

Comment 1 Poornima G 2015-03-23 05:29:46 UTC
From the core it looks like:
- The activation of snap233 failed(snapd logs are not present in sosreports to confirm the same)
- Delete snap on snap233 was initiated and glfs_fini() was being executed as a part of cleanup.
- A stubbed lookup on "/.snaps/snap233/etc/cups/classes.conf" was resumed, since glfs_fini() is in progress for snap233, any further fops on that snap will lead to crash.

Snapview-server should take care of not calling any fops on the fs object that is destroyed or being destroyed (i.e. glfs_fini is called on the fs).

This is similar to the BZ https://bugzilla.redhat.com/show_bug.cgi?id=1201820.

Comment 2 Vivek Agarwal 2015-03-23 06:56:47 UTC

*** This bug has been marked as a duplicate of bug 1201820 ***


Note You need to log in before you can comment on or make changes to this bug.