Bug 1911841
| Summary: | [IPI Baremetal] After restoring to previous state the cluster operator machine-config is degraded | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | Ori Michaeli <omichael> |
| Component: | Machine Config Operator | Assignee: | Yu Qi Zhang <jerzhang> |
| Status: | CLOSED NOTABUG | QA Contact: | Michael Nguyen <mnguyen> |
| Severity: | unspecified | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 4.6 | CC: | augol, jerzhang, mkrejci, omichael, xxia, yanyang |
| Target Milestone: | --- | Keywords: | TestBlocker |
| Target Release: | --- | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2021-03-15 17:14:09 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Ori Michaeli
2020-12-31 18:54:06 UTC
Pretty sure this is because MCO started using ignition version 3.2 in 4.7 and MCD in 4.6 only understands up to 3.1. https://github.com/openshift/machine-config-operator/pull/2248 While I don't think y-stream downgrades are supported really, I was able to break out of this by: - `oc get mc -oyaml` on the MC currentConfig for each node type (currentConfig annotation on the Nodes) - delete all rendered MCs (one will be regenerated per node type) - edit currentConfig MCs from step 1, replacing ignition version 3.2.0 with 3.1.0 - recreate the currentConfig MCs with `oc create -f` - log into each node and delete /etc/machine-config-daemon/currentConfig Should get things moving. Adding the Keywords because this bug blocks: the testing of https://issues.redhat.com/browse/API-1055 ; And the testing of downgrade cases of upgrade subteam. Could you say more about why you are testing a downgrade path? The MCO has been operating under the understanding that we do not support downgrade paths. There might be some reasoning here that we are not understanding. Thank you for providing more context. I think you meant to set needinfo on the reporter. > Could you say more about why you are testing a downgrade path? The upgrade QE guys (like above Yang Yang) have downgrade test case. They say, though downgrade is not officially supported, Dev requires QE should have a basic check for downgrade. So that the cluster function can be ensured to work, when its upgrading hits problem and makes it go into urgent situation. In addition, while doing the basic checking for downgrade, QE hit / reported many issues which made the cluster malfunction, like bug 1907812, bug 1913620, bug 1916586 etc. But they were all fixed. This is why testing downgrade. (In reply to Michelle Krejci from comment #5) > Could you say more about why you are testing a downgrade path? As Xingxing commented, we are testing disaster recovery as part of updates/upgrades testing on IPI BM. Hi, this is Jerry from the MCO team. Regarding major y stream downgrades, the MCO has never guarenteed its ability, and in this case like Seth mentions the MCO in 4.6 does not have the ignition version bump and we are unfortunately unlikely to backport that functionality, given the priority of other work. In terms of workaround, like Seth mentions, its possible to set all ignition 3.2 machineconfigs to 3.1 manually, and it should get past that error (3.1->3.2 should not have changed anything unless you are using LUKS encryption, which would not be backwards compatible). Apologies to the disruption of the QE process. If you believe "MCO downgradeability" should be a supported flow, please raise the issue as a new epic. The MCO today does not consider this a bug. Closing this as NOTABUG for now. If we would like to discuss this further, perhaps Jira is a better place to continue |