Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: When heketi device remove is run on a device which has 150 volumes, it errored out after waiting for more than an hour. Looks like the gluster pod has gone down after a while. Any command to grep docker or oc logs gets hung on the node hosting this gluster pod. [root@dhcp47-175 ~]# heketi-cli device remove a0ebaa7ee2092b04324ee02e6cefc83d Error: Failed to remove device, error: Unable to execute command on glusterfs-mm42d: [root@dhcp47-175 ~]# [root@dhcp47-175 ~]# oc get pods | grep 'gluster' glusterfs-hnt3w 1/1 Running 1 2d glusterfs-m9q9h 1/1 Running 1 2d glusterfs-mm42d 0/1 Running 1 2d [root@dhcp47-175 ~]# oc get pods -o wide | grep 'gluster' glusterfs-hnt3w 1/1 Running 1 2d 10.70.47.171 dhcp47-171.lab.eng.blr.redhat.com glusterfs-m9q9h 1/1 Running 1 2d 10.70.47.168 dhcp47-168.lab.eng.blr.redhat.com glusterfs-mm42d 0/1 Running 1 2d 10.70.47.176 dhcp47-176.lab.eng.blr.redhat.com [root@dhcp47-175 ~]# heketi-cli node info a4a3353715414fa78778865fd873f554 Node Id: a4a3353715414fa78778865fd873f554 State: online Cluster Id: f19f0be52fa5147aad0071491b0f8da7 Zone: 1 Management Hostname: dhcp47-176.lab.eng.blr.redhat.com Storage Hostname: 10.70.47.176 Devices: Id:29f67ead9c4daf3dea14e8cf2010ab9a Name:/dev/sde State:online Size (GiB):99 Used (GiB):4 Free (GiB):95 Id:724d4c878d4f406cfeb4bca3bcc15bb0 Name:/dev/sdd State:online Size (GiB):99 Used (GiB):10 Free (GiB):89 Id:a0ebaa7ee2092b04324ee02e6cefc83d Name:/dev/sdf State:offline Size (GiB):299 Used (GiB):139 Free (GiB):160 heketi topology info, heketi logs shall be attached. Version-Release number of selected component (if applicable): heketi-client-4.0.0-6.el7rhgs.x86_64 How reproducible: 1/1 Steps to Reproduce: 1. create 150 volumes from the same device 2. Run remove device operation on this device which has 150 volumes 3. wait for remove operation to complete Actual results: remove operation fails after waiting for a very long time Expected results: remove operation should succeed Additional info:
Some logs from atomic-openshift-node service corresponding to the time when the node started timing out for heketi commands: Apr 11 00:32:11 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: I0411 00:32:11.086799 6217 operation_executor.go:917] MountVolume.SetUp succeeded for volume "kubernetes.io/secret/c80d740b-1e0a-11e7-8bb8-005056b3a26f-default-token-l8k05" (spec.Name: "default-token-l8k05") pod "c80d740b-1e0a-11e7-8bb8-005056b3a26f" (UID: "c80d740b-1e0a-11e7-8bb8-005056b3a26f"). Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190309 6217 glusterfs.go:133] glusterfs: failed to get endpoints pvc-fe80cde3-1e08-11e7-8bb8-005056b3a26f[an empty namespace may not be set when a resource name is provided] Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190369 6217 reconciler.go:437] Could not construct volume information: MountVolume.NewMounter failed for volume "kubernetes.io/glusterfs/01d50cb3-1e09-11e7-8bb8-005056b3a26f-pvc-fe80cde3-1e08-11e7-8bb8-005056b3a26f" (spec.Name: "pvc-fe80cde3-1e08-11e7-8bb8-005056b3a26f") pod "01d50cb3-1e09-11e7-8bb8-005056b3a26f" (UID: "01d50cb3-1e09-11e7-8bb8-005056b3a26f") with: an empty namespace may not be set when a resource name is provided Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190428 6217 glusterfs.go:133] glusterfs: failed to get endpoints pvc-09729254-1e0a-11e7-8bb8-005056b3a26f[an empty namespace may not be set when a resource name is provided] Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190442 6217 reconciler.go:437] Could not construct volume information: MountVolume.NewMounter failed for volume "kubernetes.io/glusterfs/0d87676e-1e0a-11e7-8bb8-005056b3a26f-pvc-09729254-1e0a-11e7-8bb8-005056b3a26f" (spec.Name: "pvc-09729254-1e0a-11e7-8bb8-005056b3a26f") pod "0d87676e-1e0a-11e7-8bb8-005056b3a26f" (UID: "0d87676e-1e0a-11e7-8bb8-005056b3a26f") with: an empty namespace may not be set when a resource name is provided
Created attachment 1270744 [details] more logs from atomic-openshift-node
Karthick, Reading logs would become easy if I knew how the operations were done. Did you create all 150 volumes first, then create mongodb pods and then perform remove device? Or was it all intermixed and parallel?
(In reply to Raghavendra Talur from comment #5) > Karthick, > > Reading logs would become easy if I knew how the operations were done. Did > you create all 150 volumes first, then create mongodb pods and then perform > remove device? Or was it all intermixed and parallel? I have a template to create mongodb pod, which creates pvc dynamically, creates service, route, dc and finally creates mongodb pod. [will attach the template] so creation of one pod would mean creation of all the above at once. This way, 150 pods were created one by one with a gap of 10 seconds. # for i in {1..100}; do oc new-app mongo.json --param=DATABASE_SERVICE_NAME=mongodb-$i --param=VOLUME_CAPACITY=4Gi; sleep 10; done Only after the creation of pods, device remove was run on a device. There was no parallel operation in this test. Hope this information helps.
Created attachment 1271034 [details] mongodb-template
This issue is not seen in cns 3.6. Verified in build - cns-deploy-5.0.0-34.el7rhgs.x86_64