Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
The FDP team is no longer accepting new bugs in Bugzilla. Please report your issues under FDP project in Jira. Thanks.

Bug 1779381

Summary: Failure to create 3-node OVSDB raft cluster
Product: Red Hat Enterprise Linux Fast Datapath Reporter: Russell Bryant <rbryant>
Component: openvswitch2.12Assignee: Mark Michelson <mmichels>
Status: CLOSED WONTFIX QA Contact: Jianlin Shi <jishi>
Severity: urgent Docs Contact:
Priority: urgent    
Version: RHEL 8.0CC: ctrautma, dcbw, jhsiao, qding, ralongi, tredaelli
Target Milestone: ---   
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2023-04-11 17:21:07 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 1779460    

Description Russell Bryant 2019-12-03 20:55:39 UTC
Description of problem:

I am seeing this issue with an OpenShift 4.3 cluster that has been modified to use IPv6.  I don't think the use of IPv6 is related to this problem, though it could be.

As of this PR (https://github.com/openshift/cluster-network-operator/pull/376), OpenShift configures a 3-node ovsdb raft cluster for the OVN northbound and southbound databases.

The cluster setup is slightly modified in our IPv6 fork.  The current state of that can be found here: https://github.com/openshift-kni/cluster-network-operator/blob/4.3-ipv6/bindata/network/ovn-kubernetes/ovnkube-master.yaml

When I first saw this problem occur, I thought that I just needed to ensure that the node that was the initial cluster leader started first.  Unfortunately, that change has not resolved the issue.  That attepmted workaround was here: https://github.com/openshift-kni/cluster-network-operator/commit/35ba4c6d7ac4a58540f5845aa20e8cb23d390402

I sometimes observe the nodes fail to form a single ovsdb cluster.  When the problem occurs, I will see one DB log that looks somewhat normal . The other two will be full of logs like this:

2019-12-03T18:07:21Z|00017|socket_util|ERR|Dropped 55 log messages in last 55 seconds (most recently, 1 seconds ago) due to excessive rate
2019-12-03T18:07:21Z|00018|socket_util|ERR|9643:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]: bind: Cannot assign requested address

The IP address in the log message is the one from the Node that seems to be behaving normally (the initial cluster leader).

When I check the status of the cluster, they all indicate that they are in their own independent raft cluster, as opposed to members of the same cluster:

$ for n in $(oc get pods -n openshift-ovn-kubernetes | grep -v NAME | grep ovnkube-master | cut -f1 -d' ') ; do echo "**** $n ****" ; oc exec -n openshift-ovn-kubernetes  -it $n -c nbdb -- ovs-appctl -t /var/run/openvswitch/ovnnb_db.ctl cluster/status OVN_Northbound  ; done
**** ovnkube-master-k4qxk ****
cea4
Name: OVN_Northbound
Cluster ID: 5343 (534327f0-8e98-4458-a0c2-119e71dc6dd5)
Server ID: cea4 (cea4b502-4527-430d-b3fa-e9f29ee33d14)
Address: ssl:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]:9643
Status: cluster member
Role: leader
Term: 2
Leader: self
Vote: self

Election timer: 1000
Log: [2, 5]
Entries not yet committed: 0
Entries not yet applied: 0
Connections:
Servers:
    cea4 (cea4 at ssl:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]:9643) (self) next_index=3 match_index=4
**** ovnkube-master-stwg5 ****
8ad1
Name: OVN_Northbound
Cluster ID: a2c7 (a2c7305b-69ac-4e40-bab9-fddedad7b626)
Server ID: 8ad1 (8ad106f2-743b-413d-97d8-c1274f1a51d0)
Address: ssl:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]:9643
Status: cluster member
Role: leader
Term: 2
Leader: self
Vote: self

Election timer: 1000
Log: [2, 4]
Entries not yet committed: 0
Entries not yet applied: 0
Connections:
Servers:
    8ad1 (8ad1 at ssl:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]:9643) (self) next_index=3 match_index=3
**** ovnkube-master-xrf7g ****
7120
Name: OVN_Northbound
Cluster ID: c847 (c84754fc-519e-4f21-b71b-5f5c1d71231c)
Server ID: 7120 (71201422-5a48-475e-8952-233b9c6b773c)
Address: ssl:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]:9643
Status: cluster member
Role: leader
Term: 2
Leader: self
Vote: self

Election timer: 1000
Log: [2, 1222]
Entries not yet committed: 0
Entries not yet applied: 0
Connections:
Servers:
    7120 (7120 at ssl:[2600:1f16:b88:d00:66dc:a75b:c974:7bc]:9643) (self) next_index=519 match_index=1221


Here are the commands used to start the DB on each of the 3 nodes:

**** ovnkube-master-k4qxk-nbdb.txt ****
+ exec /usr/share/openvswitch/scripts/ovn-ctl --db-nb-cluster-local-port=9643 --db-nb-cluster-remote-port=9643 '--db-nb-cluster-local-addr=[2600:1f16:b88:d01:1f88:49cc:6efb:ea8e]' '--db-nb-cluster-remote-addr=[2600:1f16:b88:d00:66dc:a75b:c974:7bc]' --no-monitor --db-nb-cluster-local-proto=ssl --db-nb-cluster-remote-proto=ssl --ovn-nb-db-ssl-key=/ovn-cert/tls.key --ovn-nb-db-ssl-cert=/ovn-cert/tls.crt --ovn-nb-db-ssl-ca-cert=/ovn-ca/ca-bundle.crt '--ovn-nb-log=-vconsole:info -vfile:off' run_nb_ovsdb

**** ovnkube-master-stwg5-nbdb.txt ****
+ exec /usr/share/openvswitch/scripts/ovn-ctl --db-nb-cluster-local-port=9643 --db-nb-cluster-remote-port=9643 '--db-nb-cluster-local-addr=[2600:1f16:b88:d02:7f95:11a6:58c4:d5e]' '--db-nb-cluster-remote-addr=[2600:1f16:b88:d00:66dc:a75b:c
974:7bc]' --no-monitor --db-nb-cluster-local-proto=ssl --db-nb-cluster-remote-proto=ssl --ovn-nb-db-ssl-key=/ovn-cert/tls.key --ovn-nb-db-ssl-cert=/ovn-cert/tls.crt --ovn-nb-db-ssl-ca-cert=/ovn-ca/ca-bundle.crt '--ovn-nb-log=-vconsole:inf
o -vfile:off' run_nb_ovsdb

**** ovnkube-master-xrf7g-nbdb.txt ****
+ exec /usr/share/openvswitch/scripts/ovn-ctl --db-nb-cluster-local-port=9643 '--db-nb-cluster-local-addr=[2600:1f16:b88:d00:66dc:a75b:c974:7bc]' --no-monitor --db-nb-cluster-local-proto=ssl --ovn-nb-db-ssl-key=/ovn-cert/tls.key --ovn-nb-
db-ssl-cert=/ovn-cert/tls.crt --ovn-nb-db-ssl-ca-cert=/ovn-ca/ca-bundle.crt '--ovn-nb-log=-vconsole:info -vfile:off' run_nb_ovsdb




Version-Release number of selected component (if applicable):

openvswitch2.12-2.12.0-4.el7fdp.bz1773598.2.x86_64.rpm
ovn2.11-2.11.1-20.el7fdn.x86_64.rpm


How reproducible:

It happens regularly, but not every time I install a cluster.  I haven't figured out what conditions cause it.

Comment 2 Mark Michelson 2023-04-11 17:21:07 UTC
Closing this old issue since
1) It is reported against an unsupported version of ovsdb
2) There have been hundreds of commits made since this issue was created, and there have been no new issues raised about this. This is likely fixed now.
3) This was reported internally, not by customers. If this truly still is an issue with later versions of ovsdb, we can create a new issue to address it.