Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: Rollback works by deleting resources from the target cluster with label "migrated-by-migplan: <migplan_uid>". This label is attached to all k8s resources migrated by Velero in one of our plugins. Our general rollback strategy is: for namespace in migplan: for GVK in GVKs: toDelete = GVK.getResources(label=migrated-by-migplan=<my_uid>) for resource in toDelete: delete(resource) DeploymentConfigs (DCs) make this strategy non-viable as a universal deletion method, since DCs create Pods (deployer, hooks) on the target cluster after we migrate them, and deletion of the DC doesn't trigger deletion of "completed" deployer and hooks Pods. This can also make rollback hang if PVCs were attached to the DC hooks Pod, since the PVC will not be deleted until the Pod is deleted, and we never send a delete request for the Pod. More info: https://github.com/konveyor/mig-controller/issues/998 Version-Release number of selected component (if applicable): 1.4.4 How reproducible: Any time you migrate a DC that creates deployer / hooks Pods (e.g. mediawiki) Steps to Reproduce: 1. Deploy mediawiki demo app to src cluster relying on DC 2. Migrate mediawiki 3. Observe deployer/hooks Pod get created on target post-migration, note that these Pods don't have the migrated-by-migplan label. 4. Attempt rollback, it may get stuck if the hooks Pod interacted with PVCs. Actual results: Rollback can get stuck waiting on PVC to get removed since "completed" DC Pods prevent PVC deletion. Expected results: Rollback wipes out everything that the migration created.
I forgot I had filed a BZ for this, so the initial PR was merged without the Bug label. Here's the Original PR: https://github.com/konveyor/mig-controller/pull/1127 Cherry-pick to 1.5.0 PR: https://github.com/konveyor/mig-controller/pull/1131
Verified using MTC 1.5.0 openshift-migration-rhel7-operator@sha256:aeda48329c452082e03052d7bd2fc71e9762469540cdb60eb4d4004418f2a0c7 - name: MIG_CONTROLLER_REPO value: openshift-migration-controller-rhel8@sha256 - name: MIG_CONTROLLER_TAG value: 7f109e3ac03acf8f1524033bbbc10b987eeaad81856a3bc03851e1b571e8a9c7 We run a migration using DCs with hookpods, the result was this one (python2_virtual_env) [fedora@preserve-appmigration-workmachine mtc-e2e-qev2]$ oc get pods NAME READY STATUS RESTARTS AGE cakephp-mysql-persistent-1-build 0/1 Error 0 35m cakephp-mysql-persistent-1-clkn4 1/1 Running 0 34m cakephp-mysql-persistent-1-deploy 0/1 Completed 0 35m cakephp-mysql-persistent-1-hook-pre 0/1 Completed 0 35m <---- mysql-1-deploy 0/1 Completed 0 35m mysql-1-vk2gs 1/1 Running 0 35m Then we executed a rollback. After the rollback the result was this one: $ oc get pods No resources found in ocp-cakephp namespace. All pods were deleted including the hook pods. Moved to VERIFIED status.
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (Migration Toolkit for Containers (MTC) image release advisory 1.5.0), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHEA-2021:2929