Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1574950

Summary: HA router can become master but won't be reported in database
Product: Red Hat OpenStack Reporter: Jakub Libosvar <jlibosva>
Component: openstack-neutronAssignee: Slawek Kaplonski <skaplons>
Status: CLOSED WONTFIX QA Contact: Toni Freger <tfreger>
Severity: medium Docs Contact:
Priority: low    
Version: 13.0 (Queens)CC: agurenko, apevec, beagles, chrisw, emacchi, jlibosva, jstransk, lbezdick, mbracho, mbultel, mburns, morazi, skaplons, srevivo, yprokule
Target Milestone: zstreamKeywords: FutureFeature, Triaged, ZStream
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: 1563443 Environment:
Last Closed: 2022-05-01 16:40:50 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1563443    
Bug Blocks:    

Description Jakub Libosvar 2018-05-04 12:05:11 UTC
In case there is an issue with filesystem on the node running l3-agent and HA router changed its status but because of the issue can't write to state file, it can become a master (having FIPs, sending garsp) but won't be reported as master in statefile and neutron database.

Also the issue is swallowed and reported message about failing parsing output of ip monitor, which is not true.

Comment 4 Assaf Muller 2018-05-18 13:17:28 UTC
Jakub and I spoke about this issue and I wanted to provide our thoughts. In the case where a router replica transitions from standby to active (but also in other cases), it might happen that the keepalived-state-change-monitor encounters an error (for example in this case as a result of a permissions issue in /var/lib/neutron), but generally speaking under any error condition, we thought that keepalived-state-change-monitor should update the L3 agent that an error has occurred. Then the L3 agent would put that router replica in 'ERROR' state and update neutron-server, which would update the DB and API responses. This would allow the operator to know that an error happened for that particular router replica and that they should investigate. Bonus points if we also have keepalived-state-change-monitor send the actual error message to the agent. We'd then update the RPC format between the agent and the server and add a DB field like 'error_message' which we could display to the operator.