Bug 1098196 - [SNAPSHOT] : restoring a snapshot is setting the "features.barrier" option to "enable"
Summary: [SNAPSHOT] : restoring a snapshot is setting the "features.barrier" option to...
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: snapshot
Version: rhgs-3.0
Hardware: Unspecified
OS: Unspecified
high
high
Target Milestone: ---
: RHGS 3.0.0
Assignee: Sachin Pandit
QA Contact: Rahul Hinduja
URL:
Whiteboard: SNAPSHOT
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2014-05-15 13:11 UTC by spandura
Modified: 2016-09-17 12:58 UTC (History)
9 users (show)

Fixed In Version: glusterfs-3.6.0.13-1
Doc Type: Bug Fix
Doc Text:
Clone Of:
: 1098487 (view as bug list)
Environment:
Last Closed: 2014-09-22 19:38:03 UTC
Embargoed:


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Red Hat Product Errata RHEA-2014:1278 0 normal SHIPPED_LIVE Red Hat Storage Server 3.0 bug fix and enhancement update 2014-09-22 23:26:55 UTC

Description spandura 2014-05-15 13:11:12 UTC
Description of problem:
========================
By default the "features.barrier" option is set to "disable". When we restore volume to a snap created, the "features.barrier" option is set to "enable". 

Version-Release number of selected component (if applicable):
===============================================================
glusterfs 3.6.0 built on May 10 2014 13:57:11

How reproducible:
================
Often

Steps to Reproduce:
=====================
1. Create a replicate volume 1 x 2. Start the volume. 

2. Perform I/0's from fuse mount. 

3. Create a snapshot of the volume. 

4. Perform I/0's from mount

5. Restore to the snap taken.

Actual results:
======================
Before restore: gluster volume info
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
root@fan [May-15-2014-12:56:06] >gluster v info
 
Volume Name: vol_rep
Type: Replicate
Volume ID: 00cdfcfa-7b92-48e4-9eb7-082f3c768d6d
Status: Started
Snap Volume: no
Number of Bricks: 1 x 2 = 2
Transport-type: tcp
Bricks:
Brick1: fan:/rhs/bricks/vol_rep_b1
Brick2: mia:/rhs/bricks/vol_rep_b2
Options Reconfigured:
features.barrier: disable

root@fan [May-15-2014-12:56:57] >gluster snapshot restore vol_rep_snap3
Snapshot restore: vol_rep_snap3: Snap restored successfully
root@fan [May-15-2014-12:57:00] >
root@fan [May-15-2014-12:57:01] >gluster v info
 
Volume Name: vol_rep
Type: Replicate
Volume ID: 00cdfcfa-7b92-48e4-9eb7-082f3c768d6d
Status: Stopped
Snap Volume: no
Number of Bricks: 1 x 2 = 2
Transport-type: tcp
Bricks:
Brick1: fan:/var/run/gluster/snaps/210ddce0774742409476bdb00060e2bf/brick1/vol_rep_b1
Brick2: mia:/var/run/gluster/snaps/210ddce0774742409476bdb00060e2bf/brick2/vol_rep_b2
Options Reconfigured:
features.barrier: enable


Expected results:
===================
Value of the options should not be changed.

Comment 2 Sachin Pandit 2014-05-27 11:08:20 UTC
A patch which fixes this issue has been posted upstream. Once that patch
is merged upstream, I'll send the relevant patch downstream.

Comment 3 Sachin Pandit 2014-05-30 05:29:39 UTC
Patch has been posted downstream, https://code.engineering.redhat.com/gerrit/#/c/25998/ is the link to that.

Comment 5 Sachin Pandit 2014-06-09 09:09:33 UTC
To give more clarity on this bug, I'll update few cases here.

**1) Why features.barrier is missing after restore :

Following are the phases of brick-ops for snap creation
a) Pre-commit
So, During pre-commit we send a brick op in which we will enable the barrier,
Hence features.barrier is *enabled* during this phase.

b) Commit
During this phase we actually generate all the volfiles for the snapshotted
volume. Here during generating the brick volfiles if we generally copy the
information of original volume then features.barrier is enabled and that
in-turn is reflected in snapshotted volume. Because of this we internally 
remove the barrier key for snapshotted volume

c) Post-commit
Post-commit we send another brick op during which we actually disable the barrier.

As mentioned in this case, We ourselves are internally going to enable the barrier and hence it is our responsibility to remove that key during restore.
If we retain the features.barrier as it is, then we might face the succeeding
snapshot create failure and I/O might be blocked on that particular mount point. 
------------------------------------------------------------------------

**2) Why we are not retaining the features.barrier

During creation of new volume features.barrier is disabled by default.
Hence we thought of doing the same for restore.
Because of which the volume info of freshly created and freshly restored
volume does not reflect the features.barrier as disabled.
-----------------------------------------------------------------------

**3) This bug can be verified by following method

Method 1) : 
            a) Create and start a volume (vol1)
            b) Take a snapshot of that particular volume (snap1)
            c) Stop the volume vol1
            d) restore the taken snapshot snap1
            e) Try to create a snapshot of restore snap volume (snap1-alpha)
If you are able to take the snapshot then features.barrier is disabled during restore. If you fail to take a snapshot of volume with failure *reconfiguring
barrier failed* then its a failed-QA

Method 2) :
            a) Create and start a volume (vol1)
            b) Mount volume "vol1" (mnt1)
            c) Take a snapshot of vol1 (snap1)
            d) stop the volume vol1
            e) restore the taken snapshot snap1
            f) Start the volume vol1
            g) Perform I/O on the same mount point "mnt1"
If you are able to perform I/O then features.barrier is disabled during
restore.
------------------------------------------------------------------------

Comment 6 Rahul Hinduja 2014-06-09 10:05:19 UTC
Verified 1st case with build: glusterfs-3.6.0.14-1.el6rhs.x86_64

Once the snapshot is restored and volume is started, able to create another snapshot. 

[root@inception ~]# gluster v i vol2
 
Volume Name: vol2
Type: Distributed-Replicate
Volume ID: cd61bb53-42dd-4f38-a09f-a59b1ecbbba5
Status: Started
Snap Volume: no
Number of Bricks: 2 x 2 = 4
Transport-type: tcp
Bricks:
Brick1: inception.lab.eng.blr.redhat.com:/brick2/b2
Brick2: rhs-arch-srv2.lab.eng.blr.redhat.com:/brick2/b2
Brick3: rhs-arch-srv3.lab.eng.blr.redhat.com:/brick2/b2
Brick4: rhs-arch-srv4.lab.eng.blr.redhat.com:/brick2/b2
Options Reconfigured:
features.barrier: disable
nfs.drc: off
[root@inception ~]# gluster snapshot create s2 vol2
snapshot create: success: Snap s2 created successfully
[root@inception ~]# gluster volume stop vol2
Stopping volume will make its data inaccessible. Do you want to continue? (y/n) y
volume stop: vol2: success
[root@inception ~]# gluster snapshot restore s2
Snapshot restore: s2: Snap restored successfully
[root@inception ~]# gluster v i vol2
 
Volume Name: vol2
Type: Distributed-Replicate
Volume ID: cd61bb53-42dd-4f38-a09f-a59b1ecbbba5
Status: Stopped
Snap Volume: no
Number of Bricks: 2 x 2 = 4
Transport-type: tcp
Bricks:
Brick1: inception.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick1/b2
Brick2: rhs-arch-srv2.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick2/b2
Brick3: rhs-arch-srv3.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick3/b2
Brick4: rhs-arch-srv4.lab.eng.blr.redhat.com:/var/run/gluster/snaps/92377579d48840f9a4c7d45254b96a6d/brick4/b2
Options Reconfigured:
nfs.drc: off
[root@inception ~]# gluster volume start vol2; gluster snapshot create s3 vol2
volume start: vol2: success
snapshot create: success: Snap s3 created successfully
[root@inception ~]#

Comment 7 Rahul Hinduja 2014-06-09 10:17:52 UTC
Verified 2nd case with build: glusterfs-3.6.0.14-1.el6rhs.x86_64

Once the snapshot is restored and volume is started, able to run IO on the same mount point.

Moving the bug to verified state as per comment 5, comment 6 and the observation in 2nd case

Comment 9 errata-xmlrpc 2014-09-22 19:38:03 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

http://rhn.redhat.com/errata/RHEA-2014-1278.html


Note You need to log in before you can comment on or make changes to this bug.