Description of problem:
We are using VRRP for L3 high availability across 3 management nodes. We are testing VRRP failover by rebooting management node 1 and continually pinging a Floating IP that is routing through that node that will be rebooted. After rebooting node 1, the floating IP failover happens successfully after a few seconds and is reachable through node management node 2. However, when node 1 (the node that was rebooted) comes back online the floating IP will stop being reachable after ~1 minute. Doing some more investigation it appears that the virtual router on node 2 was still receiving ping requests for the floating IP but was not sending any responses. In looking at the router namespace, the router on node 1 had the floating IPs in its namespace but was not receiving any of the ping requests. The router on node 2 did not have any floating IPs in its namespace. I will attach terminal output which illustrates this more clearly. After 10-20 minutes though, the floating IP becomes reachable again and is routing through node 1 successfully.
Version-Release number of selected component (if applicable):
openstack-neutron-2014.2.1-6.el7ost.noarch
How reproducible:
Very
Steps to Reproduce:
1. Set up VRRP for L3 HA
2. Reboot one of the controllers
3. Ping floating IP that was originally routed through the rebooted node
Actual results:
Once the node finishes rebooting the floating IP becomes unavailable for some time
Expected results:
Floating IP failback transitions smoothly without any interruption in service