Bug 1765247
| Summary: | OSP14 update has a cut in control plane and loose HA of ovndb-servers. | ||
|---|---|---|---|
| Product: | Red Hat OpenStack | Reporter: | Sofer Athlan-Guyot <sathlang> |
| Component: | openstack-tripleo-heat-templates | Assignee: | Sofer Athlan-Guyot <sathlang> |
| Status: | CLOSED NEXTRELEASE | QA Contact: | Sasha Smolyak <ssmolyak> |
| Severity: | urgent | Docs Contact: | |
| Priority: | urgent | ||
| Version: | 14.0 (Rocky) | CC: | dalvarez, jlibosva, mburns, morazi, rsafrono, shdunne |
| Target Milestone: | z4 | Keywords: | Regression, Triaged, ZStream |
| Target Release: | 14.0 (Rocky) | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | openstack-tripleo-heat-templates-9.3.1-0.20190513171775.el7ost | Doc Type: | If docs needed, set a value |
| Doc Text: | Story Points: | --- | |
| Clone Of: | 1760405 | Environment: | |
| Last Closed: | 2020-01-13 13:34:56 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Sofer Athlan-Guyot
2019-10-24 15:37:27 UTC
Clone of https://bugzilla.redhat.com/show_bug.cgi?id=1760405 for osp14. Hi, the "regression" was introduced by a new ovn version. Note that it doesn't prevent the update from finishing, but it let the cluster in a not too good state, requiring a pcs resource cleanup from the user at the end of the update. Not yet in a puddle (latest - 2019-12-06.1, has 774.el7.ost) Verified on 14.0-RHEL-7/2019-12-13.1 with openstack-tripleo-heat-templates-9.3.1-0.20190513171775.el7ost.noarch Performed a minor update from the latest z-release (14z4 or 2019-11-01.1) to the latest OSP14 puddle (2019-12-13.1) using this CI job: https://rhos-qe-jenkins.rhev-ci-vms.eng.rdu2.redhat.com/view/DFG/view/network/view/networking-ovn/job/DFG-network-networking-ovn-update-14_director-rhel-virthost-3cont_2comp_2net-ipv4-geneve-composable/46/ The issue did not occur. After the update OVN DB services are HA, no errors. Full list of resources after the minor update: [heat-admin@controller-0 ~]$ sudo pcs status Cluster name: tripleo_cluster Stack: corosync Current DC: controller-2 (version 1.1.20-5.el7_7.2-3c4c782f70) - partition with quorum Last updated: Sun Dec 15 07:49:31 2019 Last change: Sun Dec 15 03:39:05 2019 by redis-bundle-1 via crm_attribute on controller-1 15 nodes configured 46 resources configured Online: [ controller-0 controller-1 controller-2 ] GuestOnline: [ galera-bundle-0@controller-0 galera-bundle-1@controller-1 galera-bundle-2@controller-2 ovn-dbs-bundle-0@controller-0 ovn-dbs-bundle-1@controller-1 ovn-dbs-bundle-2@controller-2 rabbitmq-bundle-0@controller-0 rabbitmq-bundle-1@controller-1 rabbitmq-bundle-2@controller-2 redis-bundle-0@controller-0 redis-bundle-1@controller-1 redis-bundle-2@controller-2 ] Full list of resources: Docker container set: rabbitmq-bundle [192.168.24.1:8787/rhosp14/openstack-rabbitmq:pcmklatest] rabbitmq-bundle-0 (ocf::heartbeat:rabbitmq-cluster): Started controller-0 rabbitmq-bundle-1 (ocf::heartbeat:rabbitmq-cluster): Started controller-1 rabbitmq-bundle-2 (ocf::heartbeat:rabbitmq-cluster): Started controller-2 Docker container set: galera-bundle [192.168.24.1:8787/rhosp14/openstack-mariadb:pcmklatest] galera-bundle-0 (ocf::heartbeat:galera): Master controller-0 galera-bundle-1 (ocf::heartbeat:galera): Master controller-1 galera-bundle-2 (ocf::heartbeat:galera): Master controller-2 Docker container set: redis-bundle [192.168.24.1:8787/rhosp14/openstack-redis:pcmklatest] redis-bundle-0 (ocf::heartbeat:redis): Slave controller-0 redis-bundle-1 (ocf::heartbeat:redis): Master controller-1 redis-bundle-2 (ocf::heartbeat:redis): Slave controller-2 ip-192.168.24.25 (ocf::heartbeat:IPaddr2): Started controller-2 ip-10.0.0.111 (ocf::heartbeat:IPaddr2): Started controller-1 ip-172.17.1.23 (ocf::heartbeat:IPaddr2): Started controller-2 ip-172.17.1.12 (ocf::heartbeat:IPaddr2): Started controller-1 ip-172.17.3.11 (ocf::heartbeat:IPaddr2): Started controller-2 ip-172.17.4.13 (ocf::heartbeat:IPaddr2): Started controller-1 Docker container set: haproxy-bundle [192.168.24.1:8787/rhosp14/openstack-haproxy:pcmklatest] haproxy-bundle-docker-0 (ocf::heartbeat:docker): Started controller-0 haproxy-bundle-docker-1 (ocf::heartbeat:docker): Started controller-1 haproxy-bundle-docker-2 (ocf::heartbeat:docker): Started controller-2 Docker container set: ovn-dbs-bundle [192.168.24.1:8787/rhosp14/openstack-ovn-northd:pcmklatest] ovn-dbs-bundle-0 (ocf::ovn:ovndb-servers): Slave controller-0 ovn-dbs-bundle-1 (ocf::ovn:ovndb-servers): Master controller-1 ovn-dbs-bundle-2 (ocf::ovn:ovndb-servers): Slave controller-2 Docker container: openstack-cinder-volume [192.168.24.1:8787/rhosp14/openstack-cinder-volume:pcmklatest] openstack-cinder-volume-docker-0 (ocf::heartbeat:docker): Started controller-2 Daemon Status: corosync: active/enabled pacemaker: active/enabled pcsd: active/enabled This hasn't been released in OSP14 and is fixed starting OSP15. |