Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1215179

Summary: VRRP failback after failover does not work
Product: Red Hat OpenStack Reporter: wdaniel
Component: openstack-neutronAssignee: Nir Magnezi <nmagnezi>
Status: CLOSED CURRENTRELEASE QA Contact: Eran Kuris <ekuris>
Severity: high Docs Contact:
Priority: high    
Version: 6.0 (Juno)CC: amuller, benglish, chenders, chrisw, ihrachys, jraju, jwaterwo, lpeer, nyechiel, tfreger, yeylon
Target Milestone: asyncKeywords: ZStream
Target Release: 6.0 (Juno)   
Hardware: x86_64   
OS: Linux   
Whiteboard:
Fixed In Version: openstack-neutron-2014.2.3-10.el7ost Doc Type: Bug Fix
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2015-12-03 21:34:08 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1215177    
Bug Blocks:    

Description wdaniel 2015-04-24 13:50:18 UTC
Description of problem:

We are using VRRP for L3 high availability across 3 management nodes.  We are testing VRRP failover by rebooting management node 1 and continually pinging a Floating IP that is routing through that node that will be rebooted.  After rebooting node 1, the floating IP failover happens successfully after a few seconds and is reachable through node management node 2. However, when node 1 (the node that was rebooted) comes back online the floating IP will stop being reachable after ~1 minute.  Doing some more investigation it appears that the virtual router on node 2 was still receiving ping requests for the floating IP but was not sending any responses.  In looking at the router namespace, the router on node 1 had the floating IPs in its namespace but was not receiving any of the ping requests.  The router on node 2 did not have any floating IPs in its namespace.  I will attach terminal output which illustrates this more clearly.   After 10-20 minutes though, the floating IP becomes reachable again and is routing through node 1 successfully.

Version-Release number of selected component (if applicable):

openstack-neutron-2014.2.1-6.el7ost.noarch

How reproducible:

Very

Steps to Reproduce:
1. Set up VRRP for L3 HA
2. Reboot one of the controllers
3. Ping floating IP that was originally routed through the rebooted node

Actual results:

Once the node finishes rebooting the floating IP becomes unavailable for some time

Expected results:

Floating IP failback transitions smoothly without any interruption in service

Comment 23 Eran Kuris 2015-10-11 08:12:22 UTC
Verified on New OpenStack-6.0-RHEL-7 Puddle: 2015-10-09.2 
# rpm -qa |grep neutron openstack-neutron-openvswitch-2014.2.3-19.el7ost.noarch
openstack-neutron-2014.2.3-19.el7ost.noarch
openstack-neutron-common-2014.2.3-19.el7ost.noarch
python-neutronclient-2.3.9-1.el7ost.noarch
python-neutron-2014.2.3-19.el7ost.noarch

--- 10.35.184.133 ping statistics ---
156 packets transmitted, 156 received, 0% packet loss, time 155008ms
rtt min/avg/max/mdev = 0.342/0.470/2.520/0.222 ms

fixed

Comment 25 Lon Hohberger 2015-12-03 21:34:08 UTC
This was resolved by a previous release:

https://rhn.redhat.com/errata/RHSA-2015-1909.html