Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: ======================== By default the "features.barrier" option is set to "disable". When we restore volume to a snap created, the "features.barrier" option is set to "enable". Version-Release number of selected component (if applicable): =============================================================== glusterfs 3.6.0 built on May 10 2014 13:57:11 How reproducible: ================ Often Steps to Reproduce: ===================== 1. Create a replicate volume 1 x 2. Start the volume. 2. Perform I/0's from fuse mount. 3. Create a snapshot of the volume. 4. Perform I/0's from mount 5. Restore to the snap taken. Actual results: ====================== Before restore: gluster volume info ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ root@fan [May-15-2014-12:56:06] >gluster v info Volume Name: vol_rep Type: Replicate Volume ID: 00cdfcfa-7b92-48e4-9eb7-082f3c768d6d Status: Started Snap Volume: no Number of Bricks: 1 x 2 = 2 Transport-type: tcp Bricks: Brick1: fan:/rhs/bricks/vol_rep_b1 Brick2: mia:/rhs/bricks/vol_rep_b2 Options Reconfigured: features.barrier: disable root@fan [May-15-2014-12:56:57] >gluster snapshot restore vol_rep_snap3 Snapshot restore: vol_rep_snap3: Snap restored successfully root@fan [May-15-2014-12:57:00] > root@fan [May-15-2014-12:57:01] >gluster v info Volume Name: vol_rep Type: Replicate Volume ID: 00cdfcfa-7b92-48e4-9eb7-082f3c768d6d Status: Stopped Snap Volume: no Number of Bricks: 1 x 2 = 2 Transport-type: tcp Bricks: Brick1: fan:/var/run/gluster/snaps/210ddce0774742409476bdb00060e2bf/brick1/vol_rep_b1 Brick2: mia:/var/run/gluster/snaps/210ddce0774742409476bdb00060e2bf/brick2/vol_rep_b2 Options Reconfigured: features.barrier: enable Expected results: =================== Value of the options should not be changed.
A patch which fixes this issue has been posted upstream. Once that patch is merged upstream, I'll send the relevant patch downstream.
Patch has been posted downstream, https://code.engineering.redhat.com/gerrit/#/c/25998/ is the link to that.
https://code.engineering.redhat.com/gerrit/#/c/25998/
To give more clarity on this bug, I'll update few cases here. **1) Why features.barrier is missing after restore : Following are the phases of brick-ops for snap creation a) Pre-commit So, During pre-commit we send a brick op in which we will enable the barrier, Hence features.barrier is *enabled* during this phase. b) Commit During this phase we actually generate all the volfiles for the snapshotted volume. Here during generating the brick volfiles if we generally copy the information of original volume then features.barrier is enabled and that in-turn is reflected in snapshotted volume. Because of this we internally remove the barrier key for snapshotted volume c) Post-commit Post-commit we send another brick op during which we actually disable the barrier. As mentioned in this case, We ourselves are internally going to enable the barrier and hence it is our responsibility to remove that key during restore. If we retain the features.barrier as it is, then we might face the succeeding snapshot create failure and I/O might be blocked on that particular mount point. ------------------------------------------------------------------------ **2) Why we are not retaining the features.barrier During creation of new volume features.barrier is disabled by default. Hence we thought of doing the same for restore. Because of which the volume info of freshly created and freshly restored volume does not reflect the features.barrier as disabled. ----------------------------------------------------------------------- **3) This bug can be verified by following method Method 1) : a) Create and start a volume (vol1) b) Take a snapshot of that particular volume (snap1) c) Stop the volume vol1 d) restore the taken snapshot snap1 e) Try to create a snapshot of restore snap volume (snap1-alpha) If you are able to take the snapshot then features.barrier is disabled during restore. If you fail to take a snapshot of volume with failure *reconfiguring barrier failed* then its a failed-QA Method 2) : a) Create and start a volume (vol1) b) Mount volume "vol1" (mnt1) c) Take a snapshot of vol1 (snap1) d) stop the volume vol1 e) restore the taken snapshot snap1 f) Start the volume vol1 g) Perform I/O on the same mount point "mnt1" If you are able to perform I/O then features.barrier is disabled during restore. ------------------------------------------------------------------------
Verified 1st case with build: glusterfs-3.6.0.14-1.el6rhs.x86_64 Once the snapshot is restored and volume is started, able to create another snapshot. [root@inception ~]# gluster v i vol2 Volume Name: vol2 Type: Distributed-Replicate Volume ID: cd61bb53-42dd-4f38-a09f-a59b1ecbbba5 Status: Started Snap Volume: no Number of Bricks: 2 x 2 = 4 Transport-type: tcp Bricks: Brick1: inception.lab.eng.blr.redhat.com:/brick2/b2 Brick2: rhs-arch-srv2.lab.eng.blr.redhat.com:/brick2/b2 Brick3: rhs-arch-srv3.lab.eng.blr.redhat.com:/brick2/b2 Brick4: rhs-arch-srv4.lab.eng.blr.redhat.com:/brick2/b2 Options Reconfigured: features.barrier: disable nfs.drc: off [root@inception ~]# gluster snapshot create s2 vol2 snapshot create: success: Snap s2 created successfully [root@inception ~]# gluster volume stop vol2 Stopping volume will make its data inaccessible. Do you want to continue? (y/n) y volume stop: vol2: success [root@inception ~]# gluster snapshot restore s2 Snapshot restore: s2: Snap restored successfully [root@inception ~]# gluster v i vol2 Volume Name: vol2 Type: Distributed-Replicate Volume ID: cd61bb53-42dd-4f38-a09f-a59b1ecbbba5 Status: Stopped Snap Volume: no Number of Bricks: 2 x 2 = 4 Transport-type: tcp Bricks: Brick1: inception.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick1/b2 Brick2: rhs-arch-srv2.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick2/b2 Brick3: rhs-arch-srv3.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick3/b2 Brick4: rhs-arch-srv4.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick4/b2 Options Reconfigured: nfs.drc: off [root@inception ~]# gluster volume start vol2; gluster snapshot create s3 vol2 volume start: vol2: success snapshot create: success: Snap s3 created successfully [root@inception ~]#
Verified 2nd case with build: glusterfs-3.6.0.14-1.el6rhs.x86_64 Once the snapshot is restored and volume is started, able to run IO on the same mount point. Moving the bug to verified state as per comment 5, comment 6 and the observation in 2nd case
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. http://rhn.redhat.com/errata/RHEA-2014-1278.html