Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 2024554

Summary: "Inventory" container stops with "panic: runtime error: invalid memory address or nil pointer dereference" message
Product: Migration Toolkit for Virtualization Reporter: Tzahi Ashkenazi <tashkena>
Component: InventoryAssignee: Fabien Dupont <fdupont>
Status: CLOSED ERRATA QA Contact: Tzahi Ashkenazi <tashkena>
Severity: medium Docs Contact: Avital Pinnick <apinnick>
Priority: medium    
Version: 2.2.0CC: dvaanunu, fbladilo, jortel, mguetta
Target Milestone: ---   
Target Release: 2.2.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-12-09 19:21:13 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Tzahi Ashkenazi 2021-11-18 10:36:22 UTC
Description of problem:
inventory container exit on PANIC on idle state exit on:
"panic: runtime error: invalid memory address or nil pointer dereference"
around 07:00 this morning Israel time without any usage of the system 
this leads to controller restart 
oc get pods : 
forklift-controller-795458684c-5gg8z                               2/2     Running                      1 (7h8m ago)   43h

the PANIC message from the inventory container :

panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x1971025]

goroutine 964 [running]:
github.com/vmware/govmomi.(*Client).RoundTrip(0x0, 0x30c9630, 0xc0026acfc0, 0x307d8c0, 0xc0017bd908, 0x307d8c0, 0xc0017bd920, 0x1313, 0x1)
        <autogenerated>:1 +0x5
github.com/vmware/govmomi/vim25/methods.WaitForUpdatesEx(0x30c9630, 0xc0026acfc0, 0x307bfc0, 0x0, 0xc0049e2fc0, 0xc00323e840, 0x2, 0x2)
        /remote-source/deps/gomod/pkg/mod/github.com/vmware/govmomi.1/vim25/methods/methods.go:17919 +0xb8
github.com/konveyor/forklift-controller/pkg/controller/provider/container/vsphere.(*Collector).getUpdates(0xc0026de120, 0x30c9630, 0xc0026acfc0, 0x0, 0x0)
        /remote-source/app/pkg/controller/provider/container/vsphere/collector.go:394 +0x575
github.com/konveyor/forklift-controller/pkg/controller/provider/container/vsphere.(*Collector).Start.func1()
        /remote-source/app/pkg/controller/provider/container/vsphere/collector.go:315 +0x145
created by github.com/konveyor/forklift-controller/pkg/controller/provider/container/vsphere.(*Collector).Start
        /remote-source/app/pkg/controller/provider/container/vsphere/collector.go:330 +0xc5



Version-Release number of selected component (if applicable):

Cloud10 
MTV version : 2.2.0-87
OCP - 4.9.7

Additional info:

the full inventory container log can be found here :
https://drive.google.com/drive/folders/1Y98K37g7asShI1PCC1-Gl1Gxor3TeH2-?usp=sharing

Comment 1 Fabien Dupont 2021-11-22 14:07:34 UTC
I haven't reproduced this behavior in my lab and it was up for 4 days. I suspect that VMware's unavailability. We need to fix it, but I don't think it's a blocker. I've targeted the BZ to MTV 2.3.0.

Comment 2 Maayan Hadasi 2021-11-24 12:23:26 UTC
The issue was reproduced on a PSI-based cluster (mgn03) using MTV 2.2.0-96 while running a Warm migration.
forklift-controller pod was in a crash state till the migration plan was deleted.
Attachments: forklift-controller logs

Comment 5 Maayan Hadasi 2021-11-24 15:01:55 UTC
Looks like the panic bug is reproduced more often with MTV 2.2.0-96
It was reproduced today on mgn03 and mig01 - both PSI clusters

on mig01 the issue was reproduced while running a cold migration. The migration plan cannot end and MTV is not usable till the plan is deleted (via CLI)

Comment 6 Fabien Dupont 2021-11-29 15:19:29 UTC
This panic happens frequently enough to be painful and a fix would be beneficial in 2.2. Changing target release to 2.2.0.

Comment 7 Ilanit Stein 2021-11-29 16:20:58 UTC
Apparently the panic error mentioned in comment #2 is different than the one reported in the bug description.
The one mentioned in comment #2 is more frequent, and has another cause.

Comment 8 David Vaanunu 2021-11-30 10:43:51 UTC
Testing warm migration of 2 VMs using MTV-100 (BM) and MTV-102 (PSI)
forklift-controller main container crashed with the same error: 

E1130 08:51:29.000023       1 runtime.go:78] Observed a panic: "invalid memory address or nil pointer dereference" (runtime error: invalid memory address or nil pointer dereference)
goroutine 557 [running]:
k8s.io/apimachinery/pkg/util/runtime.logPanic(0x2828220, 0x45abd80)
	/remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/runtime/runtime.go:74 +0xa6
k8s.io/apimachinery/pkg/util/runtime.HandleCrash(0x0, 0x0, 0x0)
	/remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/runtime/runtime.go:48 +0x86
panic(0x2828220, 0x45abd80)
	/usr/lib/golang/src/runtime/panic.go:965 +0x1b9
github.com/konveyor/forklift-controller/pkg/controller/plan/adapter/vsphere.(*Builder).mapDisks(0xc0014a1300, 0xc000d6c400, 0xc00115f600, 0x1, 0x1, 0xc003e88118)
	/remote-source/app/pkg/controller/plan/adapter/vsphere/builder.go:485 +0x3fd
github.com/konveyor/forklift-controller/pkg/controller/plan/adapter/vsphere.(*Builder).VirtualMachine(0xc0014a1300, 0xc003c9d8b0, 0x7, 0xc004d4b650, 0x2d, 0x0, 0x0, 0xc003e88118, 0xc00115f600, 0x1, ...)
	/remote-source/app/pkg/controller/plan/adapter/vsphere/builder.go:343 +0x485
github.com/konveyor/forklift-controller/pkg/controller/plan.(*KubeVirt).virtualMachine(0xc000e417d0, 0xc0007ba700, 0xc00053e480, 0x200000003, 0xc00053e480)
	/remote-source/app/pkg/controller/plan/kubevirt.go:729 +0x4d3
github.com/konveyor/forklift-controller/pkg/controller/plan.(*KubeVirt).EnsureVM(0xc000e417d0, 0xc0007ba700, 0xa, 0xc000db2b40)
	/remote-source/app/pkg/controller/plan/kubevirt.go:214 +0x50
github.com/konveyor/forklift-controller/pkg/controller/plan.(*Migration).execute(0xc000e417b8, 0xc0007ba700, 0x2, 0x2)
	/remote-source/app/pkg/controller/plan/migration.go:592 +0x2851
github.com/konveyor/forklift-controller/pkg/controller/plan.(*Migration).Run(0xc000e417b8, 0xb2d05e00, 0x0, 0x0)
	/remote-source/app/pkg/controller/plan/migration.go:168 +0x138
github.com/konveyor/forklift-controller/pkg/controller/plan.(*Reconciler).execute(0xc000fbba40, 0xc000595800, 0x0, 0x0, 0x0)
	/remote-source/app/pkg/controller/plan/controller.go:405 +0x854
github.com/konveyor/forklift-controller/pkg/controller/plan.Reconciler.Reconcile(0x30c1320, 0xc000e4bb00, 0x30e2478, 0xc0003a4ed0, 0xc00084d3e0, 0xc0008fc8b0, 0xd, 0xc000b311a0, 0x28, 0x0, ...)
	/remote-source/app/pkg/controller/plan/controller.go:263 +0x6e5
sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler(0xc00051e120, 0x295c020, 0xc0008f0340, 0x0)
	/remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:244 +0x2a9
sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem(0xc00051e120, 0x203000)
	/remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:218 +0xb0
sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).worker(...)
	/remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:197
k8s.io/apimachinery/pkg/util/wait.BackoffUntil.func1(0xc0008efd10)
	/remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:155 +0x5f
k8s.io/apimachinery/pkg/util/wait.BackoffUntil(0xc0008efd10, 0x307fca0, 0xc000fbb9b0, 0x1, 0xc000a50c60)
	/remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:156 +0x9b
k8s.io/apimachinery/pkg/util/wait.JitterUntil(0xc0008efd10, 0x3b9aca00, 0x0, 0x1, 0xc000a50c60)
	/remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:133 +0x98
k8s.io/apimachinery/pkg/util/wait.Until(0xc0008efd10, 0x3b9aca00, 0xc000a50c60)
	/remote-source/deps/gomod/pkg/mod/k8s.io/apimachinery.3/pkg/util/wait/wait.go:90 +0x4d
created by sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).Start.func1
	/remote-source/deps/gomod/pkg/mod/sigs.k8s.io/controller-runtime.4/pkg/internal/controller/controller.go:179 +0x3d6
panic: runtime error: invalid memory address or nil pointer dereference [recovered]
	panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x20 pc=0x1a0461d]

Comment 9 Fabien Dupont 2021-12-01 10:54:50 UTC
Please verify with mtv-operator-bundle-container-2.2.0-103 / iib:140554, or later.

Comment 10 Tzahi Ashkenazi 2021-12-05 12:54:24 UTC
since this bug was opened on panic on the inventory container on idle state without any load or migration in the background .
I have used cloud10 with MTV version 2.2.0-104  for several migration cycles. ( for other BZ > https://bugzilla.redhat.com/show_bug.cgi?id=2024138 ) 
and left the cluster idle for the weekend for around 76 hours in total for tracking.
during this time no PANIC or any other memory errors was spotted on the inventory or the main containers  ( under the controller pod ) 
closing this BZ  ( and we will keep tracking the controller pod, for any panics or errors ) 


[root@f01-h14-000-r640 ~]# oc get pods
NAME                                       READY   STATUS      RESTARTS   AGE
forklift-controller-6b86f597bc-fnzvx       2/2     Running     0          3d4h
forklift-must-gather-api-646c6cfdc-bqd5p   1/1     Running     0          3d4h
forklift-operator-d7968fc75-j2cp6          1/1     Running     0          3d4h
forklift-ui-7b7c864d4c-7cj6p               1/1     Running     0          3d4h
forklift-validation-594c88894-sfvj4        1/1     Running     0          3d4h


verified on cloud10 
MTV 2.2.0-104

Comment 13 errata-xmlrpc 2021-12-09 19:21:13 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory (MTV 2.2.0 Images), and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHEA-2021:5066