Bug 1448467
| Summary: | Clean-up of a managed bundle or guest node resource causes unnecessary recovery | |||
|---|---|---|---|---|
| Product: | Red Hat Enterprise Linux 7 | Reporter: | Klaus Wenninger <kwenning> | |
| Component: | pacemaker | Assignee: | Ken Gaillot <kgaillot> | |
| Status: | CLOSED ERRATA | QA Contact: | Ofer Blaut <oblaut> | |
| Severity: | urgent | Docs Contact: | ||
| Priority: | urgent | |||
| Version: | 7.4 | CC: | abeekhof, aherr, cfeist, chorn, cluster-maint, cluster-qe, ctowsley, fdinitto, jruemker, kgaillot, kwenning, mnovacek, phagara, rmarigny, royoung | |
| Target Milestone: | rc | Keywords: | FutureFeature, ZStream | |
| Target Release: | 7.7 | |||
| Hardware: | x86_64 | |||
| OS: | Linux | |||
| Whiteboard: | ||||
| Fixed In Version: | pacemaker-1.1.20-1.el7 | Doc Type: | Bug Fix | |
| Doc Text: |
Cause: After cleaning failure history of a managed guest node resource or bundle container, Pacemaker would schedule a reprobe of the resource and recovery of its Pacemaker Remote connection, processed in parallel.
Consequence: The Pacemaker Remote connection recovery would force recovery of the guest node resource or bundle container, even if not needed.
Fix: The connection recovery is now scheduled after getting the reprobe result.
Result: If the reprobe result finds everything OK, the recovery will be avoided.
|
Story Points: | --- | |
| Clone Of: | 1303742 | |||
| : | 1646349 1646350 (view as bug list) | Environment: | ||
| Last Closed: | 2019-08-06 12:53:38 UTC | Type: | Bug | |
| Regression: | --- | Mount Type: | --- | |
| Documentation: | --- | CRM: | ||
| Verified Versions: | Category: | --- | ||
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | ||
| Cloudforms Team: | --- | Target Upstream Version: | ||
| Embargoed: | ||||
| Bug Depends On: | 1303742 | |||
| Bug Blocks: | 1203710, 1646349, 1646350, 1647927, 1707454 | |||
|
Comment 2
Klaus Wenninger
2017-05-05 13:31:27 UTC
This bug seems to be the clone from 7.4 which I do not understand. Why exactly do we have this same bug in 7.5? In 7.4, we were able to fix the issue as long as the remote connection resource is unmanaged first. For 7.5, we are going to see if we can avoid the issue even if the resource is still managed. There's a chance this will not be practical or will need to be bumped to 7.6, but we are going to investigate it. Thanks Ken for jumping in on that... I guess that answers your question. For further details refer to solution #2 in the description as stated in comment #2. + Andrew Beekhof (11 minutes ago) 7abc8ec: PE: Implement probing of container remote nodes (HEAD -> master) + Andrew Beekhof (3 days ago) 73ffa11: PE: Revert e21a4d00 since probing remote connections is no longer a problem needs some reinvestigation in the light of the above. Should make things easier ... moved to rhel-7.5 because of effort constraints I'd be surprised if this wasnt now fixed upstream (In reply to Andrew Beekhof from comment #9) > I'd be surprised if this wasnt now fixed upstream Guess the relevant changes should be in rhel-7.5 as of the current build (pacemaker-1.1.18-5.el7). Haven't dug into the logs but with that version the issue isn't solved out of the box at least. *** Bug 1638580 has been marked as a duplicate of this bug. *** (In reply to Ken Gaillot from comment #12) > *** Bug 1638580 has been marked as a duplicate of this bug. *** Except bz #1638580 included a working fix Fixed in upstream 1.1 branch by commits a07ff46, fade228, and af4f6a1 Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2019:2129 |