In case there is an issue with filesystem on the node running l3-agent and HA router changed its status but because of the issue can't write to state file, it can become a master (having FIPs, sending garsp) but won't be reported as master in statefile and neutron database.
Also the issue is swallowed and reported message about failing parsing output of ip monitor, which is not true.
Jakub and I spoke about this issue and I wanted to provide our thoughts. In the case where a router replica transitions from standby to active (but also in other cases), it might happen that the keepalived-state-change-monitor encounters an error (for example in this case as a result of a permissions issue in /var/lib/neutron), but generally speaking under any error condition, we thought that keepalived-state-change-monitor should update the L3 agent that an error has occurred. Then the L3 agent would put that router replica in 'ERROR' state and update neutron-server, which would update the DB and API responses. This would allow the operator to know that an error happened for that particular router replica and that they should investigate. Bonus points if we also have keepalived-state-change-monitor send the actual error message to the agent. We'd then update the RPC format between the agent and the server and add a DB field like 'error_message' which we could display to the operator.