Bug 1441044 - heketi device remove command on a device with 150 volumes fails
Summary: heketi device remove command on a device with 150 volumes fails
Keywords:
Status: CLOSED WORKSFORME
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: heketi
Version: cns-3.5
Hardware: Unspecified
OS: Unspecified
unspecified
high
Target Milestone: ---
: ---
Assignee: Ramakrishna Reddy Yekulla
QA Contact: krishnaram Karthick
URL:
Whiteboard:
Depends On: 1427156 1451471 1451474
Blocks:
TreeView+ depends on / blocked
 
Reported: 2017-04-11 04:23 UTC by krishnaram Karthick
Modified: 2018-09-18 08:00 UTC (History)
10 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2017-09-12 14:04:07 UTC
Embargoed:


Attachments (Terms of Use)
more logs from atomic-openshift-node (122.87 KB, text/plain)
2017-04-11 11:34 UTC, Raghavendra Talur
no flags Details
mongodb-template (7.77 KB, text/plain)
2017-04-12 05:16 UTC, krishnaram Karthick
no flags Details


Links
System ID Private Priority Status Summary Last Updated
Red Hat Bugzilla 1443103 0 unspecified CLOSED [Scale Testing] Scaling of gluster volumes beyond 300 failed as the PVC's got stuck in Pending state indefinitely 2021-02-22 00:41:40 UTC

Internal Links: 1443103

Description krishnaram Karthick 2017-04-11 04:23:25 UTC
Description of problem:

When heketi device remove is run on a device which has 150 volumes,  it errored out after waiting for more than an hour. Looks like the gluster pod has gone down after a while.

Any command to grep docker or oc logs gets hung on the node hosting this gluster pod.

[root@dhcp47-175 ~]# heketi-cli device remove a0ebaa7ee2092b04324ee02e6cefc83d

Error: Failed to remove device, error: Unable to execute command on glusterfs-mm42d:
[root@dhcp47-175 ~]#
[root@dhcp47-175 ~]# oc get pods | grep 'gluster'
glusterfs-hnt3w                  1/1       Running   1          2d
glusterfs-m9q9h                  1/1       Running   1          2d
glusterfs-mm42d                  0/1       Running   1          2d
[root@dhcp47-175 ~]# oc get pods -o wide | grep 'gluster'
glusterfs-hnt3w                  1/1       Running   1          2d        10.70.47.171   dhcp47-171.lab.eng.blr.redhat.com
glusterfs-m9q9h                  1/1       Running   1          2d        10.70.47.168   dhcp47-168.lab.eng.blr.redhat.com
glusterfs-mm42d                  0/1       Running   1          2d        10.70.47.176   dhcp47-176.lab.eng.blr.redhat.com


[root@dhcp47-175 ~]# heketi-cli  node info a4a3353715414fa78778865fd873f554
Node Id: a4a3353715414fa78778865fd873f554
State: online
Cluster Id: f19f0be52fa5147aad0071491b0f8da7
Zone: 1
Management Hostname: dhcp47-176.lab.eng.blr.redhat.com
Storage Hostname: 10.70.47.176
Devices:
Id:29f67ead9c4daf3dea14e8cf2010ab9a   Name:/dev/sde            State:online    Size (GiB):99      Used (GiB):4       Free (GiB):95      
Id:724d4c878d4f406cfeb4bca3bcc15bb0   Name:/dev/sdd            State:online    Size (GiB):99      Used (GiB):10      Free (GiB):89      
Id:a0ebaa7ee2092b04324ee02e6cefc83d   Name:/dev/sdf            State:offline   Size (GiB):299     Used (GiB):139     Free (GiB):160     

heketi topology info, heketi logs shall be attached. 

Version-Release number of selected component (if applicable):
heketi-client-4.0.0-6.el7rhgs.x86_64

How reproducible:
1/1

Steps to Reproduce:
1. create 150 volumes from the same device
2. Run remove device operation on this device which has 150 volumes
3. wait for remove operation to complete

Actual results:
remove operation fails after waiting for a very long time

Expected results:
remove operation should succeed

Additional info:

Comment 3 Raghavendra Talur 2017-04-11 11:20:09 UTC
Some logs from atomic-openshift-node service corresponding to the time when the node started timing out for heketi commands:

Apr 11 00:32:11 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: I0411 00:32:11.086799    6217 operation_executor.go:917] MountVolume.SetUp succeeded for volume "kubernetes.io/secret/c80d740b-1e0a-11e7-8bb8-005056b3a26f-default-token-l8k05" (spec.Name: "default-token-l8k05") pod "c80d740b-1e0a-11e7-8bb8-005056b3a26f" (UID: "c80d740b-1e0a-11e7-8bb8-005056b3a26f").
Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190309    6217 glusterfs.go:133] glusterfs: failed to get endpoints pvc-fe80cde3-1e08-11e7-8bb8-005056b3a26f[an empty namespace may not be set when a resource name is provided]
Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190369    6217 reconciler.go:437] Could not construct volume information: MountVolume.NewMounter failed for volume "kubernetes.io/glusterfs/01d50cb3-1e09-11e7-8bb8-005056b3a26f-pvc-fe80cde3-1e08-11e7-8bb8-005056b3a26f" (spec.Name: "pvc-fe80cde3-1e08-11e7-8bb8-005056b3a26f") pod "01d50cb3-1e09-11e7-8bb8-005056b3a26f" (UID: "01d50cb3-1e09-11e7-8bb8-005056b3a26f") with: an empty namespace may not be set when a resource name is provided
Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190428    6217 glusterfs.go:133] glusterfs: failed to get endpoints pvc-09729254-1e0a-11e7-8bb8-005056b3a26f[an empty namespace may not be set when a resource name is provided]
Apr 11 00:32:12 dhcp47-176.lab.eng.blr.redhat.com atomic-openshift-node[6217]: E0411 00:32:12.190442    6217 reconciler.go:437] Could not construct volume information: MountVolume.NewMounter failed for volume "kubernetes.io/glusterfs/0d87676e-1e0a-11e7-8bb8-005056b3a26f-pvc-09729254-1e0a-11e7-8bb8-005056b3a26f" (spec.Name: "pvc-09729254-1e0a-11e7-8bb8-005056b3a26f") pod "0d87676e-1e0a-11e7-8bb8-005056b3a26f" (UID: "0d87676e-1e0a-11e7-8bb8-005056b3a26f") with: an empty namespace may not be set when a resource name is provided

Comment 4 Raghavendra Talur 2017-04-11 11:34:16 UTC
Created attachment 1270744 [details]
more logs from atomic-openshift-node

Comment 5 Raghavendra Talur 2017-04-11 13:43:16 UTC
Karthick,

Reading logs would become easy if I knew how the operations were done. Did you create all 150 volumes first, then create mongodb pods and then perform remove device? Or was it all intermixed and parallel?

Comment 6 krishnaram Karthick 2017-04-12 05:15:14 UTC
(In reply to Raghavendra Talur from comment #5)
> Karthick,
> 
> Reading logs would become easy if I knew how the operations were done. Did
> you create all 150 volumes first, then create mongodb pods and then perform
> remove device? Or was it all intermixed and parallel?

I have a template to create mongodb pod, which creates pvc dynamically, creates service, route, dc and finally creates mongodb pod. [will attach the template]

so creation of one pod would mean creation of all the above at once. This way, 150 pods were created one by one with a gap of 10 seconds. 

# for i in {1..100}; do oc new-app mongo.json --param=DATABASE_SERVICE_NAME=mongodb-$i --param=VOLUME_CAPACITY=4Gi; sleep 10; done

Only after the creation of pods, device remove was run on a device. There was no parallel operation in this test.

Hope this information helps.

Comment 7 krishnaram Karthick 2017-04-12 05:16:30 UTC
Created attachment 1271034 [details]
mongodb-template

Comment 14 krishnaram Karthick 2017-09-12 14:04:07 UTC
This issue is not seen in cns 3.6. 

Verified in build - cns-deploy-5.0.0-34.el7rhgs.x86_64


Note You need to log in before you can comment on or make changes to this bug.