Note: This bug is displayed in read-only format because
the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.
The sosreports didn't capture anything from the timeframe in comment 0, but it looks like the issue was reproduced around Jul 27 00:20.
Reassigning to kronosnet, as we've seen these CPG try-again messages affect pacemaker on RHEL 8 in two cases thus far and we suspect it's fixed in knet-1.23.
Jul 27 00:28:28.185 controller-0.redhat.local pacemaker-controld [181374] (crm_cs_flush) info: Sent 0 CPG messages (9 still queued): Try again (rc=6)
Jul 27 00:28:28.475 controller-0.redhat.local pacemaker-controld [181374] (crm_cs_flush) info: Sent 0 CPG messages (9 still queued): Try again (rc=6)
Jul 27 00:28:28.765 controller-0.redhat.local pacemaker-controld [181374] (crm_cs_flush) info: Sent 0 CPG messages (9 still queued): Try again (rc=6)
Thus far in all three cases (including this one), it seems that the sort of "signature" is that we get a "[QUORUM] Sync members" message but there's a long delay before the "[QUORUM] Members" and "Completed service synchronization" messages.
Jul 27 00:19:31 [181356] controller-0.redhat.local corosync notice [CFG ] Node 6 was shut down by sysadmin
Jul 27 00:19:31 [181356] controller-0.redhat.local corosync notice [QUORUM] Sync members[8]: 1 2 3 4 5 7 8 9
Jul 27 00:19:31 [181356] controller-0.redhat.local corosync notice [QUORUM] Sync left[1]: 6
Jul 27 00:19:31 [181356] controller-0.redhat.local corosync notice [TOTEM ] A new membership (1.3e) was formed. Members left: 6
Jul 27 00:19:31 [181356] controller-0.redhat.local corosync notice [QUORUM] Members[8]: 1 2 3 4 5 7 8 9
Jul 27 00:19:31 [181356] controller-0.redhat.local corosync notice [MAIN ] Completed service synchronization, ready to provide service.
Jul 27 00:19:38 [181356] controller-0.redhat.local corosync info [KNET ] link: host: 6 link: 0 is down
Jul 27 00:19:38 [181356] controller-0.redhat.local corosync info [KNET ] host: host: 6 (passive) best link: 0 (pri: 1)
Jul 27 00:19:38 [181356] controller-0.redhat.local corosync warning [KNET ] host: host: 6 has no active links
Jul 27 00:19:50 [181356] controller-0.redhat.local corosync info [KNET ] rx: host: 6 link: 0 is up
Jul 27 00:19:50 [181356] controller-0.redhat.local corosync info [KNET ] host: host: 6 (passive) best link: 0 (pri: 1)
Jul 27 00:19:54 [181356] controller-0.redhat.local corosync notice [QUORUM] Sync members[9]: 1 2 3 4 5 6 7 8 9
Jul 27 00:19:54 [181356] controller-0.redhat.local corosync notice [QUORUM] Sync joined[1]: 6
Jul 27 00:19:54 [181356] controller-0.redhat.local corosync notice [TOTEM ] A new membership (1.43) was formed. Members joined: 6
Jul 27 00:30:31 [181356] controller-0.redhat.local corosync notice [QUORUM] Members[9]: 1 2 3 4 5 6 7 8 9
Jul 27 00:30:31 [181356] controller-0.redhat.local corosync notice [MAIN ] Completed service synchronization, ready to provide service.
We'll likely want this cloned for RHEL 8.
Description of problem: we observed this during a osp17 minor update, the workflow is: - start from a three nodes cluster (controller-0,1,2) - pcs cluster stop of "controller-0" - run dnf update and update containers if applicable - pcs cluster start we randomly see this failing with: 2022-07-26 17:16:52 | 2022-07-26 17:16:52.611660 | 52540017-b3fc-350b-359f-000000000098 | TASK | Start pacemaker cluster 2022-07-26 17:21:54 | [0;31m2022-07-26 17:21:54.232126 | 52540017-b3fc-350b-359f-000000000098 | FATAL | Start pacemaker cluster | controller-0 | error={"changed": false, "msg": "Failed to set the state `online` on the cluster\n"}[0m so the cluster doesn't come up and we hit a 300s timeout in tripleo. This doesn't seem to be related to network disruptions or similar. In var/log/messages we see something like: Jul 26 17:16:54 controller-0 pacemakerd[221857]: notice: Starting Pacemaker 2.1.2-4.el9 Jul 26 17:16:54 controller-0 pacemakerd[221857]: notice: Pacemaker daemon successfully started and accepting connections Jul 26 17:16:54 controller-0 pacemaker-execd[221860]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:16:54 controller-0 pacemaker-execd[221860]: notice: Starting Pacemaker local executor Jul 26 17:16:54 controller-0 pacemaker-execd[221860]: notice: Pacemaker local executor successfully started and accepting connections Jul 26 17:16:54 controller-0 pacemaker-execd[221860]: notice: OCF resource agent search path is /usr/lib/ocf/resource.d Jul 26 17:16:54 controller-0 pacemaker-fenced[221859]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:16:54 controller-0 pacemaker-fenced[221859]: notice: Starting Pacemaker fencer Jul 26 17:16:54 controller-0 pacemaker-fenced[221859]: notice: Connecting to corosync cluster infrastructure Jul 26 17:16:54 controller-0 pacemaker-attrd[221861]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:16:54 controller-0 pacemaker-attrd[221861]: notice: Starting Pacemaker node attribute manager Jul 26 17:16:54 controller-0 pacemaker-controld[221863]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:16:54 controller-0 pacemaker-controld[221863]: notice: Starting Pacemaker controller Jul 26 17:16:54 controller-0 pacemaker-based[221858]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:16:54 controller-0 pacemaker-based[221858]: notice: Starting Pacemaker CIB manager Jul 26 17:16:54 controller-0 pacemaker-schedulerd[221862]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:16:54 controller-0 pacemaker-schedulerd[221862]: notice: Starting Pacemaker scheduler Jul 26 17:16:54 controller-0 pacemaker-schedulerd[221862]: notice: Pacemaker scheduler successfully started and accepting connections Jul 26 17:16:54 controller-0 pacemaker-fenced[221859]: notice: Node controller-0 state is now member Jul 26 17:16:54 controller-0 pacemaker-based[221858]: notice: Connecting to corosync cluster infrastructure Jul 26 17:16:54 controller-0 pacemaker-based[221858]: notice: Node controller-0 state is now member Jul 26 17:16:54 controller-0 pacemaker-based[221858]: notice: Pacemaker CIB manager successfully started and accepting connections Jul 26 17:16:55 controller-0 pacemaker-attrd[221861]: notice: Connecting to corosync cluster infrastructure Jul 26 17:16:55 controller-0 pacemaker-controld[221863]: notice: Connecting to corosync cluster infrastructure Jul 26 17:16:55 controller-0 pacemaker-fenced[221859]: notice: Pacemaker fencer successfully started and accepting connections Jul 26 17:16:55 controller-0 pacemaker-attrd[221861]: notice: Node controller-0 state is now member Jul 26 17:16:55 controller-0 pacemaker-fenced[221859]: warning: Blind faith: not fencing unseen nodes Jul 26 17:16:55 controller-0 pacemaker-attrd[221861]: notice: Pacemaker node attribute manager successfully started and accepting connections Jul 26 17:16:55 controller-0 pacemaker-attrd[221861]: notice: Setting #attrd-protocol[controller-0]: (unset) -> 3 Jul 26 17:16:55 controller-0 pacemaker-attrd[221861]: notice: Recorded local node as attribute writer (was unset) Jul 26 17:16:55 controller-0 pacemaker-controld[221863]: warning: No quorum Jul 26 17:16:55 controller-0 pacemaker-controld[221863]: notice: Node controller-0 state is now member Jul 26 17:16:55 controller-0 pacemaker-controld[221863]: notice: Pacemaker controller successfully started and accepting connections Jul 26 17:16:55 controller-0 pacemaker-controld[221863]: notice: State transition S_STARTING -> S_PENDING Jul 26 17:16:56 controller-0 pacemaker-controld[221863]: notice: Fencer successfully connected Jul 26 17:17:16 controller-0 pacemaker-controld[221863]: warning: Input I_DC_TIMEOUT received in state S_PENDING from crm_timer_popped Jul 26 17:19:16 controller-0 pacemaker-controld[221863]: notice: State transition S_ELECTION -> S_INTEGRATION Jul 26 17:19:16 controller-0 pacemaker-controld[221863]: notice: Cluster does not have watchdog fencing device Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: notice: Feature update failed: Timer expired Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: error: Input I_ERROR received in state S_INTEGRATION from feature_update_callback Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: warning: State transition S_INTEGRATION -> S_RECOVERY Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: warning: Fast-tracking shutdown in response to errors Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: warning: Not voting in election, we're in state S_RECOVERY Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: error: Input I_TERMINATE received in state S_RECOVERY from do_recover Jul 26 17:19:46 controller-0 pacemaker-controld[221863]: notice: Disconnected from the executor Jul 26 17:19:50 controller-0 pacemakerd[221857]: notice: pacemaker-controld[221863] is unresponsive to ipc after 1 tries Jul 26 17:19:57 controller-0 pacemakerd[221857]: notice: pacemaker-controld[221863] is unresponsive to ipc after 2 tries Jul 26 17:20:04 controller-0 pacemakerd[221857]: notice: pacemaker-controld[221863] is unresponsive to ipc after 3 tries Jul 26 17:20:11 controller-0 pacemakerd[221857]: notice: pacemaker-controld[221863] is unresponsive to ipc after 4 tries Jul 26 17:20:18 controller-0 pacemakerd[221857]: error: pacemaker-controld[221863] is unresponsive to ipc after 5 tries but we found the pid so have it killed that we can restart Jul 26 17:20:18 controller-0 pacemakerd[221857]: notice: Stopping pacemaker-controld Jul 26 17:20:18 controller-0 pacemakerd[221857]: warning: pacemaker-controld[221863] terminated with signal 9 (Killed) Jul 26 17:20:18 controller-0 pacemakerd[221857]: notice: Respawning failed child process: pacemaker-controld Jul 26 17:20:18 controller-0 pacemaker-controld[244584]: notice: Additional logging available in /var/log/pacemaker/pacemaker.log Jul 26 17:20:18 controller-0 pacemaker-controld[244584]: notice: Starting Pacemaker controller Jul 26 17:20:18 controller-0 pacemaker-controld[244584]: notice: Connecting to corosync cluster infrastructure Jul 26 17:20:23 controller-0 pacemakerd[221857]: notice: pacemaker-controld[244584] is unresponsive to ipc after 1 tries Jul 26 17:20:29 controller-0 pacemakerd[221857]: notice: pacemaker-controld[244584] is unresponsive to ipc after 2 tries Jul 26 17:20:35 controller-0 pacemakerd[221857]: notice: pacemaker-controld[244584] is unresponsive to ipc after 3 tries Jul 26 17:20:41 controller-0 pacemakerd[221857]: notice: pacemaker-controld[244584] is unresponsive to ipc after 4 tries Jul 26 17:20:47 controller-0 pacemakerd[221857]: error: pacemaker-controld[244584] is unresponsive to ipc after 5 tries but we found the pid so have it killed that we can restart controller-1: Jul 26 17:20:18 controller-1 pacemaker-controld[2495]: notice: do_shutdown of peer controller-0 is complete Jul 26 17:20:18 controller-1 pacemaker-controld[2495]: notice: State transition S_IDLE -> S_INTEGRATION Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: warning: Deletion of transient attributes for node controller-0 (via CIB call 173) failed: Timer expired Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Node update 178 failed: Timer expired (-62) Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Node update 179 failed: Timer expired (-62) Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Quorum update 180 failed: Timer expired (-62) Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Input I_ERROR received in state S_POLICY_ENGINE from cib_quorum_update_complete Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: warning: State transition S_POLICY_ENGINE -> S_RECOVERY Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: warning: Fast-tracking shutdown in response to errors Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: warning: Not voting in election, we're in state S_RECOVERY Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Input I_ERROR received in state S_RECOVERY from crmd_node_update_complete Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Input I_ERROR received in state S_RECOVERY from node_list_update_callback Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: Input I_TERMINATE received in state S_RECOVERY from do_recover Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Stopped 0 recurring operations at shutdown (9 remaining) Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action galera-bundle-2:7 (galera-bundle-2_monitor_30000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action haproxy-bundle-podman-1:73 (haproxy-bundle-podman-1_monitor_60000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action openstack-cinder-volume-podman-0:85 (openstack-cinder-volume-podman-0_monitor_60000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action rabbitmq-bundle-2:10 (rabbitmq-bundle-2_monitor_30000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action ip-172.17.4.44:69 (ip-172.17.4.44_monitor_10000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action rabbitmq-bundle-podman-2:74 (rabbitmq-bundle-podman-2_monitor_60000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action ip-10.0.0.102:68 (ip-10.0.0.102_monitor_10000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action galera-bundle-podman-2:76 (galera-bundle-podman-2_monitor_60000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Recurring action ip-192.168.24.48:83 (ip-192.168.24.48_monitor_10000) incomplete at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: error: 9 resources were active at shutdown Jul 26 17:21:49 controller-1 pacemaker-controld[2495]: notice: Disconnected from the executor Jul 26 17:21:52 controller-1 pacemakerd[2356]: notice: pacemaker-controld[2495] is unresponsive to ipc after 1 tries Jul 26 17:21:59 controller-1 pacemakerd[2356]: notice: pacemaker-controld[2495] is unresponsive to ipc after 2 tries Jul 26 17:22:06 controller-1 pacemakerd[2356]: notice: pacemaker-controld[2495] is unresponsive to ipc after 3 tries Jul 26 17:22:13 controller-1 pacemakerd[2356]: notice: pacemaker-controld[2495] is unresponsive to ipc after 4 tries Jul 26 17:22:20 controller-1 pacemakerd[2356]: error: pacemaker-controld[2495] is unresponsive to ipc after 5 tries but we found the pid so have it killed that we can restart controller-2: Jul 26 17:18:40 controller-2 pacemaker-controld[2419]: warning: Resource update 109 failed: (rc=-62) Timer expired Jul 26 17:22:20 controller-2 pacemaker-controld[2419]: notice: Our peer on the DC (controller-1) is dead Jul 26 17:22:20 controller-2 pacemaker-controld[2419]: notice: State transition S_NOT_DC -> S_ELECTION Jul 26 17:23:00 controller-2 pacemaker-controld[2419]: warning: Deletion of transient attributes for node controller-1 (via CIB call 110) failed: Timer expired Jul 26 17:24:20 controller-2 pacemaker-controld[2419]: notice: State transition S_ELECTION -> S_INTEGRATION Jul 26 17:24:20 controller-2 pacemaker-controld[2419]: notice: Cluster does not have watchdog fencing device Jul 26 17:24:40 controller-2 python3[246757]: ansible-command Invoked with chdir=/var/log/extra executable=/bin/bash warn=False _raw_params=exec >/var/log/extra/pcs_cpu_throttle.txt 2>&1#012if type pcs &>/dev/null; then#012 echo "+ high CPU throttling events"#012 grep throttle_check_thresholds /var/log/pacemaker/pacemaker.log#012fi#012 _uses_shell=True stdin_add_newline=True strip_empty_ends=True argv=None creates=None removes=None stdin=None Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Feature update failed: Timer expired Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: error: Input I_ERROR received in state S_INTEGRATION from feature_update_callback Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: warning: State transition S_INTEGRATION -> S_RECOVERY Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: warning: Fast-tracking shutdown in response to errors Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: warning: Not voting in election, we're in state S_RECOVERY Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: error: Input I_TERMINATE received in state S_RECOVERY from do_recover Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Stopped 0 recurring operations at shutdown (8 remaining) Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action rabbitmq-bundle-podman-0:78 (rabbitmq-bundle-podman-0_monitor_60000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action galera-bundle-podman-0:75 (galera-bundle-podman-0_monitor_60000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action openstack-cinder-backup-podman-0:81 (openstack-cinder-backup-podman-0_monitor_60000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action galera-bundle-0:9 (galera-bundle-0_monitor_30000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action ip-172.17.1.76:67 (ip-172.17.1.76_monitor_10000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action haproxy-bundle-podman-2:74 (haproxy-bundle-podman-2_monitor_60000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action ip-172.17.3.17:83 (ip-172.17.3.17_monitor_10000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Recurring action rabbitmq-bundle-0:10 (rabbitmq-bundle-0_monitor_30000) incomplete at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: error: 8 resources were active at shutdown Jul 26 17:25:00 controller-2 pacemaker-controld[2419]: notice: Disconnected from the executor Jul 26 17:25:07 controller-2 pacemakerd[2365]: notice: pacemaker-controld[2419] is unresponsive to ipc after 1 tries Jul 26 17:25:14 controller-2 pacemakerd[2365]: notice: pacemaker-controld[2419] is unresponsive to ipc after 2 tries Jul 26 17:25:21 controller-2 pacemakerd[2365]: notice: pacemaker-controld[2419] is unresponsive to ipc after 3 tries Jul 26 17:25:26 controller-2 pacemaker-controld[2419]: notice: Disconnected from Corosync Jul 26 17:25:26 controller-2 pacemaker-controld[2419]: notice: Disconnected from the CIB manager Jul 26 17:25:26 controller-2 pacemakerd[2365]: notice: pacemaker-controld[2419] is unresponsive to ipc after 4 tries Version-Release number of selected component (if applicable): pacemaker-2.1.2-4.el9.x86_64 pacemaker-cli-2.1.2-4.el9.x86_64 pacemaker-cluster-libs-2.1.2-4.el9.x86_64 pacemaker-libs-2.1.2-4.el9.x86_64 pacemaker-remote-2.1.2-4.el9.x86_64 pacemaker-schemas-2.1.2-4.el9.noarch How reproducible: randomly but often enough.