Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: In a 6 node setup: While creating 30 10GB mongodb pods in loop the glusterd service stops on 1 node Version-Release number of selected component (if applicable): oc version oc v3.10.0-0.58.0 kubernetes v1.10.0+b81c8f8 features: Basic-Auth GSSAPI Kerberos SPNEGO Server https://dhcp47-15.lab.eng.blr.redhat.com:8443 openshift v3.10.0-0.58.0 kubernetes v1.10.0+b81c8f8 oc rsh glusterfs-storage-fqk4h sh-4.2# rpm -qa | grep gluster glusterfs-client-xlators-3.8.4-54.8.el7rhgs.x86_64 glusterfs-fuse-3.8.4-54.8.el7rhgs.x86_64 glusterfs-geo-replication-3.8.4-54.8.el7rhgs.x86_64 glusterfs-libs-3.8.4-54.8.el7rhgs.x86_64 glusterfs-3.8.4-54.8.el7rhgs.x86_64 glusterfs-api-3.8.4-54.8.el7rhgs.x86_64 glusterfs-cli-3.8.4-54.8.el7rhgs.x86_64 glusterfs-server-3.8.4-54.8.el7rhgs.x86_64 gluster-block-0.2.1-18.el7rhgs.x86_64 oc rsh heketi-storage-1-fddnp sh-4.2# rpm -qa | grep heketi python-heketi-6.0.0-14.el7rhgs.x86_64 heketi-client-6.0.0-14.el7rhgs.x86_64 heketi-6.0.0-14.el7rhgs.x86_64 sh-4.2# How reproducible: 1/1 Steps to Reproduce: 1. Create a 6 node 3.10 OCP with 3.10 CNS 2. Creating 30 10GB mongodb app pods 3. 28 pods are up and running. 4. Observed 2 of the pods are in Error state. Actual results: Observed that 2 of the pods are in error state with error message 'Failed to provision volume with StorageClass "block-storage": failed to create volume: [heketi] failed to create volume: Unable to execute command on glusterfs-storage-fqk4h:' Expected results: All 30 pods need to created successfuly without any discrepancies Additional info: Adding oc get events output
Created attachment 1447760 [details] oc get events for the failed pvc
Gluster logs, heketi logs and sosreports are places at http://rhsqe-repo.lab.eng.blr.redhat.com/cns/sosreports/BZ-1585985/
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHEA-2018:2691