Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1867886

Summary: upgrade from 4.4.6>4.5 fails: unexpected on-disk state validating against rendered-master
Product: OpenShift Container Platform Reporter: Ke Wang <kewang>
Component: Machine Config OperatorAssignee: Antonio Murdaca <amurdaca>
Status: CLOSED WORKSFORME QA Contact: Michael Nguyen <mnguyen>
Severity: high Docs Contact:
Priority: high    
Version: 4.4CC: agarcial, amurdaca, aos-bugs, bleanhar, cglombek, huirwang, jerzhang, jiajliu, jialiu, jima, jokerman, kewang, kgarriso, lmohanty, openshift-bugzilla-robot, skumari, smilner, walters, wking, zhsun
Target Milestone: ---   
Target Release: 4.4.z   
Hardware: x86_64   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: 1846690 Environment:
Last Closed: 2020-09-29 17:44:31 UTC Type: ---
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1842906, 1846690    
Bug Blocks:    

Comment 1 Lalatendu Mohanty 2020-08-26 16:56:16 UTC
We're asking the following questions to evaluate whether or not this bug warrants blocking an upgrade edge from either the previous X.Y or X.Y.Z. The ultimate goal is to avoid delivering an update which introduces new risk or reduces cluster functionality in any way. Sample answers are provided to give more context and the UpgradeBlocker flag has been added to this bug. It will be removed if the assessment indicates that this should not block upgrade edges. The expectation is that the assignee answers these questions.

Who is impacted?  If we have to block upgrade edges based on this issue, which edges would need blocking?
  example: Customers upgrading from 4.y.Z to 4.y+1.z running on GCP with thousands of namespaces, approximately 5% of the subscribed fleet
  example: All customers upgrading from 4.y.z to 4.y+1.z fail approximately 10% of the time
What is the impact?  Is it serious enough to warrant blocking edges?
  example: Up to 2 minute disruption in edge routing
  example: Up to 90seconds of API downtime
  example: etcd loses quorum and you have to restore from backup
How involved is remediation (even moderately serious impacts might be acceptable if they are easy to mitigate)?
  example: Issue resolves itself after five minutes
  example: Admin uses oc to fix things
  example: Admin must SSH to hosts, restore from backups, or other non standard admin activities
Is this a regression (if all previous versions were also vulnerable, updating to the new, vulnerable version does not increase exposure)?
  example: No, it’s always been like this we just never noticed
  example: Yes, from 4.y.z to 4.y+1.z Or 4.y.z to 4.y.z+1

Comment 3 Sinny Kumari 2020-09-14 10:48:28 UTC
This bug doesn't seem to be related to the original cloned bug https://bugzilla.redhat.com/show_bug.cgi?id=1842906 .
Original bug was happening in 4.5 and onward releases because we removed BindsTo=ignition-firstboot-complete.service from machine-config-daemon-firstboot.service see, https://github.com/openshift/machine-config-operator/commit/75dbab9c54c6cb3470075af1da1b139ecea02d38 ). As a result during upgrade to 4.5 in progress machine-config-daemon-firstboot.service was running again.

This is not the case in 4.4, we have BindsTo=ignition-firstboot-complete.service (https://github.com/openshift/machine-config-operator/blob/release-4.4/templates/common/_base/units/machine-config-daemon-firstboot.service#L8 ) which ensures that machine-config-daemon-firstboot.service will run only during firstboot.

Also, from the journal log mentioned in comment#0, it shows that reboot was made 2 times. First one is when firstboot completed and second one would be possibly result of node upgrade from 4.3 to 4.4.

Log says reason of cluster upgrade failures as "cluster operator openshift-apiserver has not yet successfully rolled out", it would be worth checking if other operators are doing fine. Also, this bug would require must-gahther for further analysis.

Comment 4 Ke Wang 2020-09-25 15:34:49 UTC
Not seen this bug in latest upgrading 4.3.z to 4.4.z nightly CI test with ipi install on aws, will have a try manually, if unable to be reproduced, will close the bug.

Comment 5 Ke Wang 2020-09-27 10:34:57 UTC
$ oc get nodes;oc get co
NAME                                             STATUS   ROLES    AGE   VERSION
kewang-6g7w2-m-0.c.openshift-qe.internal         Ready    master   29m   v1.16.2+417b9fd
kewang-6g7w2-m-1.c.openshift-qe.internal         Ready    master   30m   v1.16.2+417b9fd
kewang-6g7w2-m-2.c.openshift-qe.internal         Ready    master   29m   v1.16.2+417b9fd
kewang-6g7w2-w-a-7pj6t.c.openshift-qe.internal   Ready    worker   19m   v1.16.2+417b9fd
kewang-6g7w2-w-b-v2csq.c.openshift-qe.internal   Ready    worker   19m   v1.16.2+417b9fd
kewang-6g7w2-w-c-vxxpf.c.openshift-qe.internal   Ready    worker   19m   v1.16.2+417b9fd
NAME                                       VERSION   AVAILABLE   PROGRESSING   DEGRADED   SINCE
authentication                             4.3.38    True        False         False      11m
cloud-credential                           4.3.38    True        False         False      30m
cluster-autoscaler                         4.3.38    True        False         False      18m
console                                    4.3.38    True        False         False      15m
dns                                        4.3.38    True        False         False      27m
image-registry                             4.3.38    True        False         False      17m
ingress                                    4.3.38    True        False         False      17m
insights                                   4.3.38    True        False         False      24m
kube-apiserver                             4.3.38    True        False         False      26m
kube-controller-manager                    4.3.38    True        False         False      26m
kube-scheduler                             4.3.38    True        False         False      26m
machine-api                                4.3.38    True        False         False      24m
machine-config                             4.3.38    True        False         False      27m
marketplace                                4.3.38    True        False         False      19m
monitoring                                 4.3.38    True        False         False      11m
network                                    4.3.38    True        False         False      29m
node-tuning                                4.3.38    True        False         False      17m
openshift-apiserver                        4.3.38    True        False         False      24m
openshift-controller-manager               4.3.38    True        False         False      27m
openshift-samples                          4.3.38    True        False         False      18m
operator-lifecycle-manager                 4.3.38    True        False         False      19m
operator-lifecycle-manager-catalog         4.3.38    True        False         False      19m
operator-lifecycle-manager-packageserver   4.3.38    True        False         False      19m
service-ca                                 4.3.38    True        False         False      28m
service-catalog-apiserver                  4.3.38    True        False         False      25m
service-catalog-controller-manager         4.3.38    True        False         False      19m
storage                                    4.3.38    True        False         False      19m

$ oc patch clusterversion/version --patch '{"spec":{"upstream":"https://openshift-release.svc.ci.openshift.org/graph"}}' --type=merge
clusterversion.config.openshift.io/version patched

$ oc adm upgrade --to=4.4.0-0.nightly-2020-09-26-141123 --allow-explicit-upgrade=true --force=true

$ oc get clusterversion
NAME      VERSION   AVAILABLE   PROGRESSING   SINCE   STATUS
version   4.3.38    True        True          7m56s   Working towards 4.4.0-0.nightly-2020-09-26-141123: 24% complete

$ oc get clusterversion
NAME      VERSION                             AVAILABLE   PROGRESSING   SINCE   STATUS
version   4.4.0-0.nightly-2020-09-26-141123   True        False         42m     Cluster version is 4.4.0-0.nightly-2020-09-26-141123

Not seen the bug when upgrading 4.3.31 to 4.4.z nightly.

Comment 7 Ke Wang 2020-09-27 10:39:45 UTC
Sorry, not seen the bug when upgrading 4.3.38 to 4.4.z nightly, not 4.3.31.

Comment 8 Kirsten Garrison 2020-09-29 17:44:31 UTC
Thanks for the update @Ke Wang. I'll close this then, and please reopen if you encounter later.

Comment 9 W. Trevor King 2021-04-05 17:47:40 UTC
Removing UpgradeBlocker from this older bug, to remove it from the suspect queue described in [1].  If you feel like this bug still needs to be a suspect, please add keyword again.

[1]: https://github.com/openshift/enhancements/pull/475

Comment 10 Red Hat Bugzilla 2023-09-15 00:46:16 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days