Bug 1867886
| Summary: | upgrade from 4.4.6>4.5 fails: unexpected on-disk state validating against rendered-master | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | Ke Wang <kewang> |
| Component: | Machine Config Operator | Assignee: | Antonio Murdaca <amurdaca> |
| Status: | CLOSED WORKSFORME | QA Contact: | Michael Nguyen <mnguyen> |
| Severity: | high | Docs Contact: | |
| Priority: | high | ||
| Version: | 4.4 | CC: | agarcial, amurdaca, aos-bugs, bleanhar, cglombek, huirwang, jerzhang, jiajliu, jialiu, jima, jokerman, kewang, kgarriso, lmohanty, openshift-bugzilla-robot, skumari, smilner, walters, wking, zhsun |
| Target Milestone: | --- | ||
| Target Release: | 4.4.z | ||
| Hardware: | x86_64 | ||
| OS: | Linux | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | 1846690 | Environment: | |
| Last Closed: | 2020-09-29 17:44:31 UTC | Type: | --- |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | 1842906, 1846690 | ||
| Bug Blocks: | |||
|
Comment 1
Lalatendu Mohanty
2020-08-26 16:56:16 UTC
This bug doesn't seem to be related to the original cloned bug https://bugzilla.redhat.com/show_bug.cgi?id=1842906 . Original bug was happening in 4.5 and onward releases because we removed BindsTo=ignition-firstboot-complete.service from machine-config-daemon-firstboot.service see, https://github.com/openshift/machine-config-operator/commit/75dbab9c54c6cb3470075af1da1b139ecea02d38 ). As a result during upgrade to 4.5 in progress machine-config-daemon-firstboot.service was running again. This is not the case in 4.4, we have BindsTo=ignition-firstboot-complete.service (https://github.com/openshift/machine-config-operator/blob/release-4.4/templates/common/_base/units/machine-config-daemon-firstboot.service#L8 ) which ensures that machine-config-daemon-firstboot.service will run only during firstboot. Also, from the journal log mentioned in comment#0, it shows that reboot was made 2 times. First one is when firstboot completed and second one would be possibly result of node upgrade from 4.3 to 4.4. Log says reason of cluster upgrade failures as "cluster operator openshift-apiserver has not yet successfully rolled out", it would be worth checking if other operators are doing fine. Also, this bug would require must-gahther for further analysis. Not seen this bug in latest upgrading 4.3.z to 4.4.z nightly CI test with ipi install on aws, will have a try manually, if unable to be reproduced, will close the bug. $ oc get nodes;oc get co
NAME STATUS ROLES AGE VERSION
kewang-6g7w2-m-0.c.openshift-qe.internal Ready master 29m v1.16.2+417b9fd
kewang-6g7w2-m-1.c.openshift-qe.internal Ready master 30m v1.16.2+417b9fd
kewang-6g7w2-m-2.c.openshift-qe.internal Ready master 29m v1.16.2+417b9fd
kewang-6g7w2-w-a-7pj6t.c.openshift-qe.internal Ready worker 19m v1.16.2+417b9fd
kewang-6g7w2-w-b-v2csq.c.openshift-qe.internal Ready worker 19m v1.16.2+417b9fd
kewang-6g7w2-w-c-vxxpf.c.openshift-qe.internal Ready worker 19m v1.16.2+417b9fd
NAME VERSION AVAILABLE PROGRESSING DEGRADED SINCE
authentication 4.3.38 True False False 11m
cloud-credential 4.3.38 True False False 30m
cluster-autoscaler 4.3.38 True False False 18m
console 4.3.38 True False False 15m
dns 4.3.38 True False False 27m
image-registry 4.3.38 True False False 17m
ingress 4.3.38 True False False 17m
insights 4.3.38 True False False 24m
kube-apiserver 4.3.38 True False False 26m
kube-controller-manager 4.3.38 True False False 26m
kube-scheduler 4.3.38 True False False 26m
machine-api 4.3.38 True False False 24m
machine-config 4.3.38 True False False 27m
marketplace 4.3.38 True False False 19m
monitoring 4.3.38 True False False 11m
network 4.3.38 True False False 29m
node-tuning 4.3.38 True False False 17m
openshift-apiserver 4.3.38 True False False 24m
openshift-controller-manager 4.3.38 True False False 27m
openshift-samples 4.3.38 True False False 18m
operator-lifecycle-manager 4.3.38 True False False 19m
operator-lifecycle-manager-catalog 4.3.38 True False False 19m
operator-lifecycle-manager-packageserver 4.3.38 True False False 19m
service-ca 4.3.38 True False False 28m
service-catalog-apiserver 4.3.38 True False False 25m
service-catalog-controller-manager 4.3.38 True False False 19m
storage 4.3.38 True False False 19m
$ oc patch clusterversion/version --patch '{"spec":{"upstream":"https://openshift-release.svc.ci.openshift.org/graph"}}' --type=merge
clusterversion.config.openshift.io/version patched
$ oc adm upgrade --to=4.4.0-0.nightly-2020-09-26-141123 --allow-explicit-upgrade=true --force=true
$ oc get clusterversion
NAME VERSION AVAILABLE PROGRESSING SINCE STATUS
version 4.3.38 True True 7m56s Working towards 4.4.0-0.nightly-2020-09-26-141123: 24% complete
$ oc get clusterversion
NAME VERSION AVAILABLE PROGRESSING SINCE STATUS
version 4.4.0-0.nightly-2020-09-26-141123 True False 42m Cluster version is 4.4.0-0.nightly-2020-09-26-141123
Not seen the bug when upgrading 4.3.31 to 4.4.z nightly.
Sorry, not seen the bug when upgrading 4.3.38 to 4.4.z nightly, not 4.3.31. Thanks for the update @Ke Wang. I'll close this then, and please reopen if you encounter later. Removing UpgradeBlocker from this older bug, to remove it from the suspect queue described in [1]. If you feel like this bug still needs to be a suspect, please add keyword again. [1]: https://github.com/openshift/enhancements/pull/475 The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days |