Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1273421

Summary: Need to add deps on kernel] vdsm iscsi failover taking too long during controller maintenance
Product: Red Hat Enterprise Virtualization Manager Reporter: rhev-integ
Component: vdsmAssignee: Nir Soffer <nsoffer>
Status: CLOSED ERRATA QA Contact: Aharon Canan <acanan>
Severity: high Docs Contact:
Priority: high    
Version: 3.5.0CC: acanan, agk, amureini, bazulay, bdonahue, bmarzins, bmcclain, cshao, dfediuck, dwysocha, ecohen, fdeutsch, fsimonce, gouyang, heinzm, huiwa, iheim, janfrode, jbuchta, jhunsaker, juwu, leiwang, lpeer, lsurette, mkalinin, msnitzer, nbarcet, nsoffer, ovirt-maint, prajnoha, prockai, rbalakri, rhodain, scohen, slevine, srevivo, tnisan, yaniwang, ycui, yeylon, ylavi, zkabelac
Target Milestone: ovirt-3.5.6Keywords: ZStream
Target Release: 3.5.6   
Hardware: x86_64   
OS: Linux   
Whiteboard: storage
Fixed In Version: Doc Type: Bug Fix
Doc Text:
Previously, the multipath device configuration was overwritten by the iSCSI daemon using the default timeout 120 seconds, and resulted in long delays before VDSM attempts to try the next path. With this update, using an updated kernel, the iSCSI daemon now uses the multipath device timeout instead of overwriting it with a longer timeout time. There is now minimal delay during failover.
Story Points: ---
Clone Of: 980139 Environment:
Last Closed: 2015-12-01 20:41:12 UTC Type: ---
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: Storage RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 980139    
Bug Blocks:    

Comment 1 Allon Mureinik 2015-10-26 13:42:52 UTC
Nir, can you please add some doctext on the impact this has on the customer?

Comment 2 Nir Soffer 2015-10-26 14:49:54 UTC
Added

Comment 4 Aharon Canan 2015-11-05 12:30:03 UTC
Following my discussion with Nir, the only thing to verify here is that we require kernel 3.10.0-229.17.1.el7

Verified using vt18.1

[root@camel-vdsb ~]# rpm -q --requires vdsm-4.16.28-1.el7ev.x86_64 |grep kernel 
kernel >= 3.10.0-229.17.1.el7

Comment 5 Julie 2015-11-11 04:03:25 UTC
Hi Nir and Aharon,
   I'm not sure if I fully understand this bug. Please have a look at my updated text and let me know if you have any feedback. 

Cheers,
Julie

Comment 6 Nir Soffer 2015-11-12 07:40:47 UTC
(In reply to Julie from comment #5)
> Hi Nir and Aharon,
>    I'm not sure if I fully understand this bug. Please have a look at my
> updated text and let me know if you have any feedback. 

Previously, the multipath device configuration was overwritten by the iSCSI daemon using the default timeout 120 seconds, and resulted in long delays before VDSM attempts to re-establish connections with the Manager. 

re-establishing connections with the Manager is very creative but
completely bogus, I wonder how it sneak it into this text :-)

What happens is this:

1. vdsm or another process run by vdsm (e.g. lvm, multipath) try
   to write or read from storage
2. the request should have time out after 5 seconds (according to 
   multipath configuration, but the timeout was overridden to 120 
   seconds
3. The request times out after 120 seconds
4. Multipath try the next path (we may have several paths to the 
   same storage device
5. This request also times out after 120 seconds

This cause delays of minutes during various operations, instead of 
seconds when devices are configured properly.

With this update, using an updated kernel, the iSCSI daemon now uses the multipath device timeout instead of overwriting it with a longer timeout time. There is now minimal delays in re-establishing connections during failover.

I would remove the "re-establishing connections" part, I don't know
if this is technically correct regarding the scsi layer.

The important part is "failover" - when one path to storage is failing,
multipath detect the failure and try the next path, or fail the operation
because all paths are faulty.

Ben, would like to correct my description if needed?

Comment 7 Ben Marzinski 2015-11-12 16:39:31 UTC
The only clarification I have that step 5 is hopefully not not happening serially. Multipathd will be running path checks (by default every 20 seconds on working paths).  Most of the path checker functions (but not all) run asynchronously, which means that the path checker thread is free to check other paths while one is still waiting for a reply. This means that within 20 seconds after you lose your connection to your storage, multipathd should have tried to check your path.  Once this check times out after 120 seconds, multipathd will mark the path as failed, and multipath won't try to failover to it.

So, assuming your device uses an asynchronous checker, the maximum time till multipath notices all the paths are down should be somewhere around 140 seconds in this case.  If your device needs to use a synchronous checker, then it will work as you have outlined above.

Comment 8 Julie 2015-11-17 04:01:24 UTC
Thanks for your feedback, Ben & Nir.
Doc text updated.

Cheers,
Julie

Comment 10 errata-xmlrpc 2015-12-01 20:41:12 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://rhn.redhat.com/errata/RHBA-2015-2530.html