Bug 2024554
| Summary: | "Inventory" container stops with "panic: runtime error: invalid memory address or nil pointer dereference" message | ||
|---|---|---|---|
| Product: | Migration Toolkit for Virtualization | Reporter: | Tzahi Ashkenazi <tashkena> |
| Component: | Inventory | Assignee: | Fabien Dupont <fdupont> |
| Status: | CLOSED ERRATA | QA Contact: | Tzahi Ashkenazi <tashkena> |
| Severity: | medium | Docs Contact: | Avital Pinnick <apinnick> |
| Priority: | medium | ||
| Version: | 2.2.0 | CC: | dvaanunu, fbladilo, jortel, mguetta |
| Target Milestone: | --- | ||
| Target Release: | 2.2.0 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2021-12-09 19:21:13 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
I haven't reproduced this behavior in my lab and it was up for 4 days. I suspect that VMware's unavailability. We need to fix it, but I don't think it's a blocker. I've targeted the BZ to MTV 2.3.0. The issue was reproduced on a PSI-based cluster (mgn03) using MTV 2.2.0-96 while running a Warm migration. forklift-controller pod was in a crash state till the migration plan was deleted. Attachments: forklift-controller logs Looks like the panic bug is reproduced more often with MTV 2.2.0-96 It was reproduced today on mgn03 and mig01 - both PSI clusters on mig01 the issue was reproduced while running a cold migration. The migration plan cannot end and MTV is not usable till the plan is deleted (via CLI) This panic happens frequently enough to be painful and a fix would be beneficial in 2.2. Changing target release to 2.2.0. Apparently the panic error mentioned in comment #2 is different than the one reported in the bug description. The one mentioned in comment #2 is more frequent, and has another cause. Testing warm migration of 2 VMs using MTV-100 (BM) and MTV-102 (PSI) forklift-controller main container crashed with the same error: E1130 08:51:29.000023 1 runtime.go:78] Observed a panic: "invalid memory address or nil pointer dereference" (runtime error: invalid memory address or nil pointer dereference) goroutine 557 [running]: k8s.io/apimachinery/pkg/util/runtime.logPanic(0x2828220, 0x45abd80) /remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/runtime/runtime.go:74 +0xa6 k8s.io/apimachinery/pkg/util/runtime.HandleCrash(0x0, 0x0, 0x0) /remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/runtime/runtime.go:48 +0x86 panic(0x2828220, 0x45abd80) /usr/lib/golang/src/runtime/panic.go:965 +0x1b9 github.com/konveyor/forklift-controller/pkg/controller/plan/adapter/vsphere.(*Builder).mapDisks(0xc0014a1300, 0xc000d6c400, 0xc00115f600, 0x1, 0x1, 0xc003e88118) /remote-source/app/pkg/controller/plan/adapter/vsphere/builder.go:485 +0x3fd github.com/konveyor/forklift-controller/pkg/controller/plan/adapter/vsphere.(*Builder).VirtualMachine(0xc0014a1300, 0xc003c9d8b0, 0x7, 0xc004d4b650, 0x2d, 0x0, 0x0, 0xc003e88118, 0xc00115f600, 0x1, ...) /remote-source/app/pkg/controller/plan/adapter/vsphere/builder.go:343 +0x485 github.com/konveyor/forklift-controller/pkg/controller/plan.(*KubeVirt).virtualMachine(0xc000e417d0, 0xc0007ba700, 0xc00053e480, 0x200000003, 0xc00053e480) /remote-source/app/pkg/controller/plan/kubevirt.go:729 +0x4d3 github.com/konveyor/forklift-controller/pkg/controller/plan.(*KubeVirt).EnsureVM(0xc000e417d0, 0xc0007ba700, 0xa, 0xc000db2b40) /remote-source/app/pkg/controller/plan/kubevirt.go:214 +0x50 github.com/konveyor/forklift-controller/pkg/controller/plan.(*Migration).execute(0xc000e417b8, 0xc0007ba700, 0x2, 0x2) /remote-source/app/pkg/controller/plan/migration.go:592 +0x2851 github.com/konveyor/forklift-controller/pkg/controller/plan.(*Migration).Run(0xc000e417b8, 0xb2d05e00, 0x0, 0x0) /remote-source/app/pkg/controller/plan/migration.go:168 +0x138 github.com/konveyor/forklift-controller/pkg/controller/plan.(*Reconciler).execute(0xc000fbba40, 0xc000595800, 0x0, 0x0, 0x0) /remote-source/app/pkg/controller/plan/controller.go:405 +0x854 github.com/konveyor/forklift-controller/pkg/controller/plan.Reconciler.Reconcile(0x30c1320, 0xc000e4bb00, 0x30e2478, 0xc0003a4ed0, 0xc00084d3e0, 0xc0008fc8b0, 0xd, 0xc000b311a0, 0x28, 0x0, ...) /remote-source/app/pkg/controller/plan/controller.go:263 +0x6e5 sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler(0xc00051e120, 0x295c020, 0xc0008f0340, 0x0) /remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:244 +0x2a9 sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem(0xc00051e120, 0x203000) /remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:218 +0xb0 sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker(...) /remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:197 k8s.io/apimachinery/pkg/util/wait.BackoffUntil.func1(0xc0008efd10) /remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:155 +0x5f k8s.io/apimachinery/pkg/util/wait.BackoffUntil(0xc0008efd10, 0x307fca0, 0xc000fbb9b0, 0x1, 0xc000a50c60) /remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:156 +0x9b k8s.io/apimachinery/pkg/util/wait.JitterUntil(0xc0008efd10, 0x3b9aca00, 0x0, 0x1, 0xc000a50c60) /remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:133 +0x98 k8s.io/apimachinery/pkg/util/wait.Until(0xc0008efd10, 0x3b9aca00, 0xc000a50c60) /remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:90 +0x4d created by sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).Start.func1 /remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:179 +0x3d6 panic: runtime error: invalid memory address or nil pointer dereference [recovered] panic: runtime error: invalid memory address or nil pointer dereference [signal SIGSEGV: segmentation violation code=0x1 addr=0x20 pc=0x1a0461d] Please verify with mtv-operator-bundle-container-2.2.0-103 / iib:140554, or later. since this bug was opened on panic on the inventory container on idle state without any load or migration in the background . I have used cloud10 with MTV version 2.2.0-104 for several migration cycles. ( for other BZ > https://bugzilla.redhat.com/show_bug.cgi?id=2024138 ) and left the cluster idle for the weekend for around 76 hours in total for tracking. during this time no PANIC or any other memory errors was spotted on the inventory or the main containers ( under the controller pod ) closing this BZ ( and we will keep tracking the controller pod, for any panics or errors ) [root@f01-h14-000-r640 ~]# oc get pods NAME READY STATUS RESTARTS AGE forklift-controller-6b86f597bc-fnzvx 2/2 Running 0 3d4h forklift-must-gather-api-646c6cfdc-bqd5p 1/1 Running 0 3d4h forklift-operator-d7968fc75-j2cp6 1/1 Running 0 3d4h forklift-ui-7b7c864d4c-7cj6p 1/1 Running 0 3d4h forklift-validation-594c88894-sfvj4 1/1 Running 0 3d4h verified on cloud10 MTV 2.2.0-104 Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (MTV 2.2.0 Images), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHEA-2021:5066 |
Description of problem: inventory container exit on PANIC on idle state exit on: "panic: runtime error: invalid memory address or nil pointer dereference" around 07:00 this morning Israel time without any usage of the system this leads to controller restart oc get pods : forklift-controller-795458684c-5gg8z 2/2 Running 1 (7h8m ago) 43h the PANIC message from the inventory container : panic: runtime error: invalid memory address or nil pointer dereference [signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x1971025] goroutine 964 [running]: github.com/vmware/govmomi.(*Client).RoundTrip(0x0, 0x30c9630, 0xc0026acfc0, 0x307d8c0, 0xc0017bd908, 0x307d8c0, 0xc0017bd920, 0x1313, 0x1) <autogenerated>:1 +0x5 github.com/vmware/govmomi/vim25/methods.WaitForUpdatesEx(0x30c9630, 0xc0026acfc0, 0x307bfc0, 0x0, 0xc0049e2fc0, 0xc00323e840, 0x2, 0x2) /remote-source/deps/gomod/pkg/mod/github.com/vmware/govmomi.1/vim25/methods/methods.go:17919 +0xb8 github.com/konveyor/forklift-controller/pkg/controller/provider/container/vsphere.(*Collector).getUpdates(0xc0026de120, 0x30c9630, 0xc0026acfc0, 0x0, 0x0) /remote-source/app/pkg/controller/provider/container/vsphere/collector.go:394 +0x575 github.com/konveyor/forklift-controller/pkg/controller/provider/container/vsphere.(*Collector).Start.func1() /remote-source/app/pkg/controller/provider/container/vsphere/collector.go:315 +0x145 created by github.com/konveyor/forklift-controller/pkg/controller/provider/container/vsphere.(*Collector).Start /remote-source/app/pkg/controller/provider/container/vsphere/collector.go:330 +0xc5 Version-Release number of selected component (if applicable): Cloud10 MTV version : 2.2.0-87 OCP - 4.9.7 Additional info: the full inventory container log can be found here : https://drive.google.com/drive/folders/1Y98K37g7asShI1PCC1-Gl1Gxor3TeH2-?usp=sharing