Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1579469

Summary: pacemaker remote setup can fail from time to time(?)
Product: Red Hat OpenStack Reporter: Michele Baldessari <michele>
Component: puppet-pacemakerAssignee: Michele Baldessari <michele>
Status: CLOSED DUPLICATE QA Contact: Marian Krcmarik <mkrcmari>
Severity: high Docs Contact:
Priority: high    
Version: 14.0 (Rocky)CC: abeekhof, cluster-maint, jjoyce, jschluet, kgaillot, slinaber, tvignaud
Target Milestone: Upstream M3Keywords: Triaged
Target Release: 14.0 (Rocky)   
Hardware: All   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: No Doc Update
Doc Text:
undefined
Story Points: ---
Clone Of: Environment:
Last Closed: 2018-07-20 20:24:30 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Michele Baldessari 2018-05-17 17:48:19 UTC
Description of problem:
On a 1-controller (1 corosync node) + 1-compute node (1 pcmk remote node) we have seen the pacemaker remote connection failing from time to time. By failing I really mean taking a lot of time (many minutes of TLS handshake failed) and eventually coming up. Funnily enough we have never seen this in proper 3-node clusters with 2 or more remotes, but right now I am trying to test scaling up pacemaker nodes so I am starting with a small setup.

Version-Release number of selected component (if applicable):
pacemaker-1.1.18-11

How reproducible:
Not always

So here is what we saw in the logs:
1) From /var/log/messages on ctrl-0:
May 17 04:29:02 overcloud-controller-0 puppet-user[18271]: (/Stage[main]/Pacemaker::Corosync/File[etc-pacemaker-authkey]/ensure) defined content as '{md5}4a0c7c34ecef8b47616dc8ff675859cf'
...
May 17 04:30:50 overcloud-controller-0 crmd[20526]:  notice: Transition 1 (Complete=0, Pending=0, Fired=0, Skipped=0, Incomplete=0, Source=/var/lib/pacemaker/pengine/pe-input-1.bz2): Complete
May 17 04:30:50 overcloud-controller-0 crmd[20526]:  notice: State transition S_TRANSITION_ENGINE -> S_IDLE
May 17 04:30:50 overcloud-controller-0 crmd[20526]:  notice: State transition S_IDLE -> S_POLICY_ENGINE
May 17 04:30:50 overcloud-controller-0 pengine[20525]:  notice:  * Start      overcloud-novacomputeiha-0     ( overcloud-controller-0 ) 
May 17 04:30:50 overcloud-controller-0 pengine[20525]:  notice: Calculated transition 2, saving inputs in /var/lib/pacemaker/pengine/pe-input-2.bz2
May 17 04:30:50 overcloud-controller-0 crmd[20526]:  notice: Initiating monitor operation overcloud-novacomputeiha-0_monitor_0 locally on overcloud-controller-0
May 17 04:30:50 overcloud-controller-0 crmd[20526]:  notice: Result of probe operation for overcloud-novacomputeiha-0 on overcloud-controller-0: 7 (not running)
May 17 04:30:50 overcloud-controller-0 crmd[20526]:  notice: Initiating start operation overcloud-novacomputeiha-0_start_0 locally on overcloud-controller-0
May 17 04:30:51 overcloud-controller-0 crmd[20526]: warning: Disconnecting after TLS handshake with remote LRMD overcloud-novacomputeiha-0:3121 failed
May 17 04:30:52 overcloud-controller-0 crmd[20526]: warning: Disconnecting after TLS handshake with remote LRMD overcloud-novacomputeiha-0:3121 failed

2) Debug log from ctrl-0 corosync.log:
May 17 08:30:50 [20526] overcloud-controller-0.redhat.local       crmd:    debug: crm_remote_tcp_connect_async: Got canonical name overcloud-novacomputeiha-0.redhat.local for overcloud-novacomputeiha-0
May 17 08:30:50 [20526] overcloud-controller-0.redhat.local       crmd:     info: crm_remote_tcp_connect_async: Attempting TCP connection to 172.17.1.21:3121
May 17 08:30:50 [20526] overcloud-controller-0.redhat.local       crmd:    debug: handle_remote_ra_exec:        began remote lrmd connect, waiting for connect event.
May 17 08:30:50 [20521] overcloud-controller-0.redhat.local        cib:     info: cib_process_request:  Forwarding cib_modify operation for section status to all (origin=local/crmd/37)
May 17 08:30:50 [20521] overcloud-controller-0.redhat.local        cib:     info: cib_perform_op:       Diff: --- 0.7.2 2
May 17 08:30:50 [20521] overcloud-controller-0.redhat.local        cib:     info: cib_perform_op:       Diff: +++ 0.7.3 (null)
May 17 08:30:50 [20521] overcloud-controller-0.redhat.local        cib:     info: cib_perform_op:       +  /cib:  @num_updates=3
May 17 08:30:50 [20521] overcloud-controller-0.redhat.local        cib:     info: cib_perform_op:       +  /cib/status/node_state[@id='1']/lrm[@id='1']/lrm_resources/lrm_resource[@id='overcloud-novacomputeiha-0']/lrm_rsc_op[@id='overcloud-novacomputeiha-0_last_0']:  @operation_key=overcloud-novacomputeiha-0_start_0, @operation=start, @transition-key=3:2:0:76e06a29-850e-4878-9939-6a80fbace42d, @transition-magic=-1:193;3:2:0:76e06a29-850e-4878-9939-6a80fbace42d, @call-id=-1, @rc-code=193, @op-status=-1
May 17 08:30:50 [20521] overcloud-controller-0.redhat.local        cib:     info: cib_process_request:  Completed cib_modify operation for section status: OK (rc=0, origin=overcloud-controller-0/crmd/37, version=0.7.3)
May 17 08:30:50 [20522] overcloud-controller-0.redhat.local stonith-ng:    debug: xml_patch_version_check:      Can apply patch 0.7.3 to 0.7.2
May 17 08:30:50 [20526] overcloud-controller-0.redhat.local       crmd:    debug: te_update_diff:       Processing (cib_modify) diff: 0.7.2 -> 0.7.3 (S_TRANSITION_ENGINE)
May 17 08:30:51 [20526] overcloud-controller-0.redhat.local       crmd:  warning: lrmd_tcp_connect_cb:  Disconnecting after TLS handshake with remote LRMD overcloud-novacomputeiha-0:3121 failed
May 17 08:30:51 [20526] overcloud-controller-0.redhat.local       crmd:     info: lrmd_tls_connection_destroy:  TLS connection destroyed
May 17 08:30:51 [20526] overcloud-controller-0.redhat.local       crmd:    debug: remote_lrm_op_callback:       remote connection event - event_type:disconnect node:overcloud-novacomputeiha-0 action:none rc:ok op_status:complete
May 17 08:30:51 [20526] overcloud-controller-0.redhat.local       crmd:    debug: remote_lrm_op_callback:       Event did not match start action
May 17 08:30:51 [20526] overcloud-controller-0.redhat.local       crmd:    debug: remote_lrm_op_callback:       remote connection event - event_type:connect node:overcloud-novacomputeiha-0 action:none rc:ok op_status:complete
May 17 08:30:52 [20526] overcloud-controller-0.redhat.local       crmd:    debug: crm_remote_tcp_connect_async: Got canonical name overcloud-novacomputeiha-0.redhat.local for overcloud-novacomputeiha-0
May 17 08:30:52 [20526] overcloud-controller-0.redhat.local       crmd:     info: crm_remote_tcp_connect_async: Attempting TCP connection to 172.17.1.21:3121
May 17 08:30:52 [20526] overcloud-controller-0.redhat.local       crmd:    debug: set_key:      using cached LRMD key
May 17 08:30:52 [20526] overcloud-controller-0.redhat.local       crmd:  warning: lrmd_tcp_connect_cb:  Disconnecting after TLS handshake with remote LRMD overcloud-novacomputeiha-0:3121 failed
May 17 08:30:52 [20526] overcloud-controller-0.redhat.local       crmd:     info: lrmd_tls_connection_destroy:  TLS connection destroyed

3) From the compute-0/remote node journal
Usual TZ skew, but authkey in /var/log/messages has been defined before and md5 matches
May 17 04:28:22 overcloud-novacomputeiha-0 puppet-user[15122]: (/Stage[main]/Pacemaker::Remote/File[etc-pacemaker-authkey]/ensure) defined content as '{md5}4a0c7c34ecef8b47616dc8ff675859cf'
...
-- Logs begin at Thu 2018-05-17 08:22:35 UTC, end at Thu 2018-05-17 08:56:43 UTC. --
May 17 08:28:22 overcloud-novacomputeiha-0.redhat.local systemd[1]: Started Pacemaker Remote Service.
May 17 08:28:22 overcloud-novacomputeiha-0.redhat.local systemd[1]: Starting Pacemaker Remote Service...
May 17 08:28:22 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:   notice: Additional logging available in /var/log/pacemaker.log
May 17 08:28:22 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:   notice: Starting TLS listener on port 3121
May 17 08:28:22 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:   notice: Listening on address ::
May 17 08:30:50 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:   notice: LRMD client connection established. 0x55a4fd922380 id: aba6cc98-f31e-492b-af35-e9eb3dfb609c
May 17 08:30:51 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:    error: Remote lrmd tls handshake failed
May 17 08:30:51 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:   notice: LRMD client disconnecting remote client - name: <unknown> id: aba6cc98-f31e-492b-af35-e9eb3dfb609c
May 17 08:30:52 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:   notice: LRMD client connection established. 0x55a4fd922380 id: aeba9849-4300-494c-9c5c-f14a58d8e396
May 17 08:30:52 overcloud-novacomputeiha-0.redhat.local pacemaker_remoted[16407]:    error: Remote lrmd tls handshake failed

4) compute-0 /var/log/pacemaker.log
Set r/w permissions for uid=189, gid=189 on /var/log/pacemaker.log
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: crm_log_init:      Changed active directory to /var/lib/pacemaker/cores
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: lrmd
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_init_remote_tls_server:       Starting TLS listener on port 3121
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: bind_and_listen:   Listening on address ::
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: cib_ro
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: cib_rw
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: cib_shm
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: attrd
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: stonith-ng
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: qb_ipcs_us_publish:        server name: crmd
May 17 08:28:22 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: main:      Starting
May 17 08:30:50 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: crm_remote_accept: New remote connection from ::ffff:172.17.1.13
May 17 08:30:50 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_listen:        LRMD client connection established. 0x55a4fd922380 id: aba6cc98-f31e-492b-af35-e9eb3dfb609c
May 17 08:30:51 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: set_key:   clearing lrmd key cache
May 17 08:30:51 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    error: lrmd_remote_client_msg:    Remote lrmd tls handshake failed
May 17 08:30:51 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_client_destroy:        LRMD client disconnecting remote client - name: <unknown> id: aba6cc98-f31e-492b-af35-e9eb3dfb609c
May 17 08:30:51 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: crm_client_destroy:        Destroying 0 events
May 17 08:30:52 [16407] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: crm_remote_accept: New remote connection from ::ffff:172.17.1.13

I chatted with Andrew about it and he threw this patch around:
diff --git a/lrmd/tls_backend.c b/lrmd/tls_backend.c
index edfb02d..d1356f9 100644
--- a/lrmd/tls_backend.c
+++ b/lrmd/tls_backend.c
@@ -67,10 +67,10 @@ lrmd_remote_client_msg(gpointer data)
             rc = gnutls_handshake(*client->remote->tls_session);
 
             if (rc < 0 && rc != GNUTLS_E_AGAIN) {
-                crm_err("Remote lrmd tls handshake failed");
+                crm_err("Remote lrmd tls handshake failed (%d)", rc);
                 return -1;
             }
-        } while (rc == GNUTLS_E_INTERRUPTED);
+        } while (rc == GNUTLS_E_INTERRUPTED || rc == GNUTLS_E_AGAIN);
 
         if (rc == 0) {
             crm_debug("Remote lrmd tls handshake completed");

I built a package with that patch (pacemaker-1.1.18-11.vandamme.0.el7) and it eventually failed as well (took me 5 times as opposed to 1 or 2, but probably just bad luck there). On the remote node we now get the error code:
May 17 17:14:47 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: crm_remote_accept: New remote connection from ::ffff:172.17.1.16
May 17 17:14:47 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_listen:        LRMD client connection established. 0x55e51c65c300 id: 6e109e0e-6cac-4230-b63f-149f9a87f8fe
May 17 17:14:48 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: set_key:   clearing lrmd key cache
May 17 17:14:48 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    error: lrmd_remote_client_msg:    Remote lrmd tls handshake failed (-24)
May 17 17:14:48 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_client_destroy:        LRMD client disconnecting remote client - name: <unknown> id: 6e109e0e-6cac-4230-b63f-149f9a87f8fe

The -24 suggests "-24	GNUTLS_E_DECRYPTION_FAILED	Decryption has failed." (https://gnutls.org/manual/gnutls.html#Error-codes). Which seems to imply that the authkey was wrong or something? But 1) and 3) would imply that the authkey was set up before with the same md5 (and also it would be weird to start working all of a sudden?)

In fact after about 5mins it eventually just works (without intervention):
May 17 17:19:34 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: crm_remote_accept: New remote connection from ::ffff:172.17.1.16
May 17 17:19:34 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_listen:        LRMD client connection established. 0x55e51c659250 id: d32c9553-1c85-427c-94e3-8983cbf69ed9
May 17 17:19:35 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: set_key:   using cached LRMD key
May 17 17:19:35 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    error: lrmd_remote_client_msg:    Remote lrmd tls handshake failed (-24)
May 17 17:19:35 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_client_destroy:        LRMD client disconnecting remote client - name: <unknown> id: d32c9553-1c85-427c-94e3-8983cbf69ed9
May 17 17:19:35 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: crm_client_destroy:        Destroying 0 events
May 17 17:20:36 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:     info: crm_remote_accept: New remote connection from ::ffff:172.17.1.16
May 17 17:20:36 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:   notice: lrmd_remote_listen:        LRMD client connection established. 0x55e51c659250 id: ae9bcbed-0547-4d57-ac99-c823433f642f
May 17 17:20:36 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: set_key:   clearing lrmd key cache
May 17 17:20:36 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: lrmd_remote_client_msg:    Remote lrmd tls handshake completed


Now unless something I can't think of is changing authkey behind my back, this is totally puzzling to me. Help or ideas would be welcome

Comment 3 Andrew Beekhof 2018-05-18 04:41:10 UTC
I was reading the man page for gnutls_handshake() today and the patch is necessary.

Obviously its not the whole story though, no idea what to do about transient GNUTLS_E_DECRYPTION_FAILED errors :(

TLS 1 vs. 2?
Perhaps a connection blip while director fiddles with the network?

Comment 4 Ken Gaillot 2018-05-18 17:32:16 UTC
The patch mentioned in the Description would make the server block on a client handshake, which would be problematic. It shouldn't have any bearing on the issue at hand; the current code also retries on GNUTLS_E_AGAIN, but only once data is available to be read on the file descriptor (i.e. the same effect but nonblocking).

The main clue I see is when it eventually succeeds:

pacemaker_remoted:    debug: set_key:   using cached LRMD key
May 17 17:19:35 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    error: lrmd_remote_client_msg:    Remote lrmd tls handshake failed (-24)

pacemaker_remoted:    debug: set_key:   clearing lrmd key cache
May 17 17:20:36 [16415] overcloud-novacomputeiha-0.redhat.local pacemaker_remoted:    debug: lrmd_remote_client_msg:    Remote lrmd tls handshake completed

i.e. bad key data was somehow cached, and when our 60-second cache threshold expired, the key was re-read from disk, and worked.

Looking at the Pacemaker code, I don't see any serious issues. I'll see if I can reproduce the issue, but in the meantime, you could try this to see if we're getting a read error from disk:

diff --git a/lib/lrmd/lrmd_client.c b/lib/lrmd/lrmd_client.c
index 3fd6479..d46ce30 100644
--- a/lib/lrmd/lrmd_client.c
+++ b/lib/lrmd/lrmd_client.c
@@ -1038,7 +1038,10 @@ set_key(gnutls_datum_t * key, const char *location)
             key->data = gnutls_realloc(key->data, buf_len);
         }
         next = fgetc(stream);
-        if (next == EOF && feof(stream)) {
+        if (next == EOF) {
+            if (!feof(stream)) {
+                crm_err("Error reading executor key; copy in memory may be corrupted");
+            }
             break;
         }

Comment 5 Michele Baldessari 2018-05-19 12:21:04 UTC
(In reply to Ken Gaillot from comment #4)
> The main clue I see is when it eventually succeeds:
> 
> pacemaker_remoted:    debug: set_key:   using cached LRMD key
> May 17 17:19:35 [16415] overcloud-novacomputeiha-0.redhat.local
> pacemaker_remoted:    error: lrmd_remote_client_msg:    Remote lrmd tls
> handshake failed (-24)
> 
> pacemaker_remoted:    debug: set_key:   clearing lrmd key cache
> May 17 17:20:36 [16415] overcloud-novacomputeiha-0.redhat.local
> pacemaker_remoted:    debug: lrmd_remote_client_msg:    Remote lrmd tls
> handshake completed
> 
> i.e. bad key data was somehow cached, and when our 60-second cache threshold
> expired, the key was re-read from disk, and worked.

Ah thanks Ken! I think this is the clue that I needed. Right now in puppet we set a relationship so that the authkey is created before the cluster is created. I think with 1-node clusters what happens is that an authkey gets generated by the the 'pcs cluster auth', pacemaker reads it and caches it and uses for a bit (while failing) and succeeds only when it is reread from disk (at which point we have the 'proper' puppet generated one).

I'll move this BZ to my side of the pond. Andrew/Ken thanks a lot again for your help! (if the above makes no sense, do scream)

Ken, maybe it's worth to bring in this one in any case as it is fairly useful no matter what? (It helped me rule out networking glitches in my env):
--- a/lrmd/tls_backend.c
+++ b/lrmd/tls_backend.c
@@ -67,10 +67,10 @@ lrmd_remote_client_msg(gpointer data)
             rc = gnutls_handshake(*client->remote->tls_session);
 
             if (rc < 0 && rc != GNUTLS_E_AGAIN) {
-                crm_err("Remote lrmd tls handshake failed");
+                crm_err("Remote lrmd tls handshake failed (%d)", rc);
                 return -1;

Comment 6 Michele Baldessari 2018-05-19 12:41:17 UTC
hohum that is probably not the whole story though. I just saw a failure (but the stack got deleted before I could collect logs) even though I force 'authkey' to be created before pcsd starts and I am moderately sure that pcs does not touch an existing authkey any longer (we had an oldish bug about that). Still investigating.

Comment 7 Ken Gaillot 2018-05-21 13:33:04 UTC
(In reply to Michele Baldessari from comment #5)
> Ken, maybe it's worth to bring in this one in any case as it is fairly
> useful no matter what? (It helped me rule out networking glitches in my env):
> --- a/lrmd/tls_backend.c
> +++ b/lrmd/tls_backend.c
> @@ -67,10 +67,10 @@ lrmd_remote_client_msg(gpointer data)
>              rc = gnutls_handshake(*client->remote->tls_session);
>  
>              if (rc < 0 && rc != GNUTLS_E_AGAIN) {
> -                crm_err("Remote lrmd tls handshake failed");
> +                crm_err("Remote lrmd tls handshake failed (%d)", rc);
>                  return -1;

Yes, I'm definitely improving the log messages, and also avoiding some memory overallocation I noticed. That will likely land in 7.6.

Comment 8 Michele Baldessari 2018-05-28 09:41:18 UTC
Ok so I finally got to the bottom of this one. Currently puppet creates the authkey file with the following ordering constraint:
File['etc-pacemaker-authkey'] -> Exec["Create Cluster ${cluster_name}"]

Before the Exec["Create Cluster ${cluster_name}"] the following operations usually happen:
1) pcs cluster auth...
2) (optional pcsd restart)
3) pcs cluster setup

Now I am not sure which one exactly is at the root of this but I can clearly see that a /remote/cluster_stop API call gets done. That will then call the following:
    I, [2018-05-24T09:40:39.075982 #19327]  INFO -- : Running: /usr/sbin/corosync-cmapctl totem.cluster_name
    I, [2018-05-24T09:40:39.076036 #19327]  INFO -- : CIB USER: hacluster, groups:
    D, [2018-05-24T09:40:39.079019 #19327] DEBUG -- : []
    D, [2018-05-24T09:40:39.079073 #19327] DEBUG -- : ["Failed to initialize the cmap API. Error CS_ERR_LIBRARY\n"]
    D, [2018-05-24T09:40:39.079122 #19327] DEBUG -- : Duration: 0.002975586s
    I, [2018-05-24T09:40:39.079165 #19327]  INFO -- : Return Value: 1
    W, [2018-05-24T09:40:39.079248 #19327]  WARN -- : Cannot read config 'corosync.conf' from '/etc/corosync/corosync.conf': No such file
    W, [2018-05-24T09:40:39.079293 #19327]  WARN -- : Cannot read config 'corosync.conf' from '/etc/corosync/corosync.conf': No such file or directory - /etc/corosync/corosync.conf
    D, [2018-05-24T09:40:39.079575 #19327] DEBUG -- : permission check action=full username=hacluster groups=
    D, [2018-05-24T09:40:39.079598 #19327] DEBUG -- : permission granted for superuser
    I, [2018-05-24T09:40:39.079629 #19327]  INFO -- : Running: /usr/sbin/pcs cluster destroy
    I, [2018-05-24T09:40:39.079644 #19327]  INFO -- : CIB USER: hacluster, groups:
    D, [2018-05-24T09:40:39.874478 #19327] DEBUG -- : ["Shutting down pacemaker/corosync services...\n", "Killing any remaining services...\n", "Removing all cluster configuration files...\n"]

And the removing all cluster configuration files is what removing the authkey file.

Moving the authkey file creation after the cluster setup and before the cluster create, solves this race fully (I did 25 successful runs, I could usually hit the race in 5/6 runs tops)

Comment 12 Michele Baldessari 2018-07-20 20:24:30 UTC
Actually I ended up fixing this on rhos13 via 1598038

*** This bug has been marked as a duplicate of bug 1598038 ***