Bug 1593612
| Summary: | docker-registry pvc is trying to bind some provision volume with StorageClass "gp2" when already specify a nfs/glusterfs storage for docker-regsitry. | ||||||
|---|---|---|---|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | Johnny Liu <jialiu> | ||||
| Component: | Installer | Assignee: | Vadim Rutkovsky <vrutkovs> | ||||
| Status: | CLOSED ERRATA | QA Contact: | Johnny Liu <jialiu> | ||||
| Severity: | high | Docs Contact: | |||||
| Priority: | high | ||||||
| Version: | 3.10.0 | CC: | aclewett, aos-bugs, bleanhar, jarrpa, jialiu, jokerman, lxia, mawong, mmccomas, sdodson, vrutkovs, weshi, xtian | ||||
| Target Milestone: | --- | Keywords: | Regression, TestBlocker | ||||
| Target Release: | 3.10.z | ||||||
| Hardware: | Unspecified | ||||||
| OS: | Unspecified | ||||||
| Whiteboard: | |||||||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |||||
| Doc Text: | Story Points: | --- | |||||
| Clone Of: | Environment: | ||||||
| Last Closed: | 2018-07-30 20:22:32 UTC | Type: | Bug | ||||
| Regression: | --- | Mount Type: | --- | ||||
| Documentation: | --- | CRM: | |||||
| Verified Versions: | Category: | --- | |||||
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |||||
| Cloudforms Team: | --- | Target Upstream Version: | |||||
| Embargoed: | |||||||
| Attachments: |
|
||||||
As stated in #comment 0, registry pvc definition's storageclass is set to "" via a field named storageclass, but the correct field name should be storageClassName. It is ignored then. As there is a default storage class, so we set that name into the pvc. Might be caused by https://github.com/openshift/openshift-ansible/blob/master/roles/lib_utils/action_plugins/generate_pv_pvcs_list.py#L139 This is likely happening because we're configuring a default storage class on AWS. We should be using S3 storage for any registry running in AWS. Moving to 3.10.z There are two issue here: 1) openshift_hosted_registry_storage_createpvc is not defined, so its assumed True. openshift_hosted_registry_storage_storageclass is not defined to its set to "", which resolves into gp2 on AWS. In this case openshift-ansible installer creates a PVC and uses default storageclass set for this cloudprovider - its gp2. GP2 however doesn't support "ReadWriteMany" so it never gets bound. kind=nfs creates a volume, but it doesn't match default storageclass. There are two options here: 1) Create an NFS storageclass when NFS server is being deployed, set storageclass to the pvc if kind=nfs. This however would require NFS provisioner 2) Avoid using PVCs with openshift_hosted_registry_storage_create_pvc=false The latter option is however broken Description of problem: Glusterfs as docker registry back end storage deployment failed due to docker-registry pvc is trying to bind provision volume with IaaS default StorageClass. It's similarly issue as BZ #1593612, for CNS I believe it should fix in 3.10.0 since this is an basic scenarios. Version-Release number of the following components: openshift-ansible-3.10.5-1.git.204.3a9cb75.el7 How reproducible: 100% Steps to Reproduce: 1. Deploy OCP with CNS as docker registry back end storage. 2. 3. Actual results: TASK [openshift_hosted : Wait for registry pods] ******************************* Friday 22 June 2018 04:01:53 -0400 (0:00:00.750) 0:25:23.419 *********** FAILED - RETRYING: Wait for registry pods (60 retries left). ... FAILED - RETRYING: Wait for registry pods (1 retries left). fatal: [master.example.com]: FAILED! => {"attempts": 60, "changed": false, "failed": true, "results": {"cmd": "/usr/bin/oc get pod --selector=docker-registry=default -o json -n default", "results": [{"apiVersion": "v1", "items": [], "kind": "List", "metadata": {"resourceVersion": "", "selfLink": ""}}], "returncode": 0}, "state": "list"} Expected results: Should pass here Additional info: # oc get po NAME READY STATUS RESTARTS AGE docker-registry-1-deploy 1/1 Running 0 8m docker-registry-1-v2s6p 0/1 Pending 0 8m glusterblock-registry-provisioner-dc-1-c8f4h 1/1 Running 0 10m glusterfs-registry-4rc7g 1/1 Running 0 14m glusterfs-registry-9rh89 1/1 Running 0 14m glusterfs-registry-n4lkc 1/1 Running 0 14m heketi-registry-1-bccp7 1/1 Running 0 10m router-1-bt8st 1/1 Running 0 8m # oc describe po docker-registry-1-v2s6p Name: docker-registry-1-v2s6p Namespace: default Node: <none> Labels: deployment=docker-registry-1 deploymentconfig=docker-registry docker-registry=default Annotations: openshift.io/deployment-config.latest-version=1 openshift.io/deployment-config.name=docker-registry openshift.io/deployment.name=docker-registry-1 openshift.io/scc=restricted Status: Pending IP: Controlled By: ReplicationController/docker-registry-1 Containers: registry: Image: registry.reg-aws.openshift.com:443/openshift3/ose-docker-registry:v3.10 Port: 5000/TCP Host Port: 0/TCP Requests: cpu: 100m memory: 256Mi Liveness: http-get https://:5000/healthz delay=10s timeout=5s period=10s #success=1 #failure=3 Readiness: http-get https://:5000/healthz delay=0s timeout=5s period=10s #success=1 #failure=3 Environment: REGISTRY_HTTP_ADDR: :5000 REGISTRY_HTTP_NET: tcp REGISTRY_HTTP_SECRET: LyaZ4wm1KIK+g3l8kh5HDOperi0UcJOCUFyK3nrNzEA= REGISTRY_MIDDLEWARE_REPOSITORY_OPENSHIFT_ENFORCEQUOTA: false REGISTRY_OPENSHIFT_SERVER_ADDR: docker-registry.default.svc:5000 REGISTRY_HTTP_TLS_CERTIFICATE: /etc/secrets/registry.crt REGISTRY_HTTP_TLS_KEY: /etc/secrets/registry.key Mounts: /etc/secrets from registry-certificates (rw) /registry from registry-storage (rw) /var/run/secrets/kubernetes.io/serviceaccount from registry-token-5tw92 (ro) Conditions: Type Status PodScheduled False Volumes: registry-storage: Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace) ClaimName: registry-claim ReadOnly: false registry-certificates: Type: Secret (a volume populated by a Secret) SecretName: registry-certificates Optional: false registry-token-5tw92: Type: Secret (a volume populated by a Secret) SecretName: registry-token-5tw92 Optional: false QoS Class: Burstable Node-Selectors: registry=enabled role=node Tolerations: node.kubernetes.io/memory-pressure:NoSchedule Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedScheduling 2m (x26 over 9m) default-scheduler pod has unbound PersistentVolumeClaims # oc get sc NAME PROVISIONER AGE standard (default) kubernetes.io/gce-pd 11m # oc get pv NAME CAPACITY ACCESS MODES RECLAIM POLICY STATUS CLAIM STORAGECLASS REASON AGE registry-volume 5Gi RWX Retain Available 9m # oc get pvc NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE registry-claim Pending standard 9m # oc get pv -o yaml apiVersion: v1 items: - apiVersion: v1 kind: PersistentVolume metadata: creationTimestamp: 2018-06-22T08:02:11Z finalizers: - kubernetes.io/pv-protection name: registry-volume namespace: "" resourceVersion: "2986" selfLink: /api/v1/persistentvolumes/registry-volume uid: 916c75e8-75f2-11e8-9d01-42010af0004b spec: accessModes: - ReadWriteMany capacity: storage: 5Gi glusterfs: endpoints: glusterfs-registry-endpoints path: glusterfs-registry-volume persistentVolumeReclaimPolicy: Retain status: phase: Available kind: List metadata: resourceVersion: "" selfLink: "" # oc get pvc -o yaml apiVersion: v1 items: - apiVersion: v1 kind: PersistentVolumeClaim metadata: annotations: volume.beta.kubernetes.io/storage-provisioner: kubernetes.io/gce-pd creationTimestamp: 2018-06-22T08:02:14Z finalizers: - kubernetes.io/pvc-protection name: registry-claim namespace: default resourceVersion: "2992" selfLink: /api/v1/namespaces/default/persistentvolumeclaims/registry-claim uid: 92d9cec6-75f2-11e8-9d01-42010af0004b spec: accessModes: - ReadWriteMany resources: requests: storage: 5Gi storageClassName: standard status: phase: Pending kind: List metadata: resourceVersion: "" selfLink: "" # oc describe pvc registry-claim Name: registry-claim Namespace: default StorageClass: standard Status: Pending Volume: Labels: <none> Annotations: volume.beta.kubernetes.io/storage-provisioner=kubernetes.io/gce-pd Finalizers: [kubernetes.io/pvc-protection] Capacity: Access Modes: Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning ProvisioningFailed 3m (x26 over 9m) persistentvolume-controller Failed to provision volume with StorageClass "standard": invalid AccessModes [ReadWriteMany]: only AccessModes [ReadWriteOnce ReadOnlyMany] are supported (In reply to Vadim Rutkovsky from comment #3) > There are two options here: > > 1) Create an NFS storageclass when NFS server is being deployed, set > storageclass to the pvc if kind=nfs. This however would require NFS > provisioner > > 2) Avoid using PVCs with openshift_hosted_registry_storage_create_pvc=false > > The latter option is however broken Using NFS storageclass for pvc, or using specified static nfs pvc is two different scenarios, in this bug, I would like docker-registry use specific static nfs pvc. If I remember right, QE hit such issue recently, should be 3.10 regression. Based on comment 4, glusterfs testing is also hit such issue, we think we should fix it in 3.10. @Weshi will run some more testing to get know which old version does not have such issue to prove it is a regression. I've verify that this issue doesn't appear in openshift-ansible-3.10.1-1.git.157.2bb6250.el7. It appear since openshift-ansible-3.10.2-1.git.190.5abfddb.el7. QE's many automated TC is broken by this issue, hope we could have a quick fix for it. (In reply to Wenkai Shi from comment #4) > Description of problem: > Glusterfs as docker registry back end storage deployment failed due to > docker-registry pvc is trying to bind provision volume with IaaS default > StorageClass. > It's similarly issue as BZ #1593612, for CNS I believe it should fix in > 3.10.0 since this is an basic scenarios. This is a different issue, it fails as Access Mode was set incorrectly: > Warning ProvisioningFailed 3m (x26 over 9m) persistentvolume-controller > Failed to provision volume with StorageClass "standard": invalid AccessModes > [ReadWriteMany]: only AccessModes [ReadWriteOnce ReadOnlyMany] are supported (In reply to Johnny Liu from comment #7) > QE's many automated TC is broken by this issue, hope we could have a quick > fix for it. So this happens due to the default storageclass being set to gp2 in AWS case. PVC gets this storageclass automatically, cannot be bound (as AWS doesn't support ReadWriteMany). Its possible to add a new storageclass for NFS, but that'd probably require NFS provisioner. Current workaround - add `openshift_storageclass_default=false` to the inventory to avoid gp2 to be used when no storageclass is specified Moved to 3.10.z based on the provided workaround in comment 9 and the preference to deploy S3 based registries in AWS for production use. (In reply to Scott Dodson from comment #10) > Moved to 3.10.z based on the provided workaround in comment 9 and the For CNS part, the workaround in comment 9 also works. > preference to deploy S3 based registries in AWS for production use. For production use, cns customer may always meet this issue. OCP installer always create storageclass and mark it as default storageclass. For cns customer, glusterfs is always their choice whatever which IaaS they using. So it can not be avoid until they using the work around. Base on this, I believe it should fix in 3.10. (In reply to Vadim Rutkovsky from comment #9) > (In reply to Johnny Liu from comment #7) > > QE's many automated TC is broken by this issue, hope we could have a quick > > fix for it. > > So this happens due to the default storageclass being set to gp2 in AWS > case. PVC gets this storageclass automatically, cannot be bound (as AWS > doesn't support ReadWriteMany). If I remember right, this issue is really introduced recently. I just tried openshift-ansible-3.10.1-1.git.157.2bb6250.el7.noarch, totally the same configuration, no such issue, I am not sure what changed was done after 3.10.1, and what is the purpose of the change. [root@ip-172-18-0-56 ~]# oc get pvc NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE regpv-claim Bound regpv-volume 17G RWX 16m [root@ip-172-18-0-56 ~]# oc get pv NAME CAPACITY ACCESS MODES RECLAIM POLICY STATUS CLAIM STORAGECLASS REASON AGE regpv-volume 17G RWX Retain Bound default/regpv-claim 16m [root@ip-172-18-0-56 ~]# oc get sc NAME PROVISIONER AGE gp2 (default) kubernetes.io/aws-ebs 17m > Its possible to add a new storageclass for NFS, but that'd probably require > NFS provisioner. Just as what I explained in comment 5, enable storageclass for NFS is a different scenarios. In this bug, user specify a NFS storage for docker-registry, this PV is *only* for docker-registry PV. > Current workaround - add `openshift_storageclass_default=false` to the > inventory to avoid gp2 to be used when no storageclass is specified In my scenario, I want docker-registry to use the specified NFS storage, and also want to enable gp2 for other pod's storage dynamic provision. I do not think this is an edge scenario, I guess customer would also hit the same issue once this is released out. (In reply to Johnny Liu from comment #12) > (In reply to Vadim Rutkovsky from comment #9) > > (In reply to Johnny Liu from comment #7) > > > QE's many automated TC is broken by this issue, hope we could have a quick > > > fix for it. > > > > So this happens due to the default storageclass being set to gp2 in AWS > > case. PVC gets this storageclass automatically, cannot be bound (as AWS > > doesn't support ReadWriteMany). > If I remember right, this issue is really introduced recently. > I just tried openshift-ansible-3.10.1-1.git.157.2bb6250.el7.noarch, totally > the same configuration, no such issue, I am not sure what changed was done > after 3.10.1, and what is the purpose of the change. There was a change to set default storageclass on AWS, which is causing it > > Current workaround - add `openshift_storageclass_default=false` to the > > inventory to avoid gp2 to be used when no storageclass is specified > In my scenario, I want docker-registry to use the specified NFS storage, and > also want to enable gp2 for other pod's storage dynamic provision. > > I do not think this is an edge scenario, I guess customer would also hit the > same issue once this is released out. This configuration is not supported, we shouldn't be blocking 3.10 due to this in any case (In reply to Vadim Rutkovsky from comment #13) > (In reply to Johnny Liu from comment #12) > > (In reply to Vadim Rutkovsky from comment #9) > > > (In reply to Johnny Liu from comment #7) > > > > QE's many automated TC is broken by this issue, hope we could have a quick > > > > fix for it. > > > > > > So this happens due to the default storageclass being set to gp2 in AWS > > > case. PVC gets this storageclass automatically, cannot be bound (as AWS > > > doesn't support ReadWriteMany). > > If I remember right, this issue is really introduced recently. > > I just tried openshift-ansible-3.10.1-1.git.157.2bb6250.el7.noarch, totally > > the same configuration, no such issue, I am not sure what changed was done > > after 3.10.1, and what is the purpose of the change. > > There was a change to set default storageclass on AWS, which is causing it The change in question is https://github.com/openshift/openshift-ansible/pull/8858, which ensures default storageclass is set correctly for PVCs Could you please think about comment #11, if you agree with me, shall we quick fix this? It's really a blocker for CNS related test. (In reply to Vadim Rutkovsky from comment #13) > This configuration is not supported, we shouldn't be blocking 3.10 due to > this in any case One more confirm, so the following scenarios are not supported: 1. "docker-registry with specified nfs static PV + default storageclass enabled" is not supported. 2. "docker-registry with specified glusterfs static PV + default storageclass enabled" is not supported. If we do not want the PVC to be given the default StorageClass its storageClass should be set to "". We do not need to disable the default StorageClass Alternatively we may want to give the registry PV an arbitrary StorageClass name and set the PVC's storageClass to the same string to ensure they bind. Note that StorageClass on its own does not imply anything about dynamic provisioning, used in this way it would be like a label. (I have not read every detail of this bug thoroughly, I just want to clarify the binding logic) (In reply to Matthew Wong from comment #17) > If we do not want the PVC to be given the default StorageClass its > storageClass should be set to "". We do not need to disable the default > StorageClass I think this is what we're setting - and default storageclass is being used. > Alternatively we may want to give the registry PV an arbitrary StorageClass > name and set the PVC's storageClass to the same string to ensure they bind. > Note that StorageClass on its own does not imply anything about dynamic > provisioning, used in this way it would be like a label. So the recommended way is to create NFS storageclass and set registry's pv and pvc storageclass? Would it work automagically, no provisioners required? If yes then we can implement that, that would be a cleaner solution than relying on empty storageclass behaviour (In reply to Vadim Rutkovsky from comment #18) > (In reply to Matthew Wong from comment #17) > > If we do not want the PVC to be given the default StorageClass its > > storageClass should be set to "". We do not need to disable the default > > StorageClass > > I think this is what we're setting - and default storageclass is being used. > > > Alternatively we may want to give the registry PV an arbitrary StorageClass > > name and set the PVC's storageClass to the same string to ensure they bind. > > Note that StorageClass on its own does not imply anything about dynamic > > provisioning, used in this way it would be like a label. > > So the recommended way is to create NFS storageclass and set registry's pv > and pvc storageclass? Would it work automagically, no provisioners required? > If yes then we can implement that, that would be a cleaner solution than > relying on empty storageclass behaviour Matt, could you help me with understand what's needed to fix this? Yes, you can create a storageclass called 'x-registry' and set it on both the pv and pvc, there is no need for a provisioner to exist, the provisioner field can have basically anything in it as long as it fits a format like 'openshift.com/nfs-registry'. It would indeed be cleaner than relying on empty storageclass behaviour, because from what I can tell, changing our playbooks to rely on that behaviour will be difficult as our scripts & the templates do not make a distinction between '' and nil. https://github.com/openshift/openshift-ansible/blob/master/roles/openshift_persistent_volumes/templates/persistent-volume-claim.yml.j2 . I am no openshift-ansible expert though. There is another possible solution that will require a bit more work but will ensure pvc A always binds to pv B, documented here https://docs.openshift.com/container-platform/3.9/dev_guide/persistent_volumes.html#persistent-volumes-volumes-and-claim-prebinding . https://github.com/openshift/openshift-ansible/blob/192b8477faa018d31893b322f43f356ce3ac1c80/roles/openshift_provisioners/templates/pv.j2#L28 . This will be cleaner in the sense that the user won't see the single-use 'x-registry' storageclass every time they do 'oc get sc'. sorry for the late reply Created PR for master to fix it - https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would have empty storageClass and could be bound by particular PVC only (In reply to Vadim Rutkovsky from comment #27) > Created PR for master to fix it - > https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would > have empty storageClass and could be bound by particular PVC only Hi Vadim, Is this can solve the CNS issue at Comment 4 as well? (In reply to Wenkai Shi from comment #28) > (In reply to Vadim Rutkovsky from comment #27) > > Created PR for master to fix it - > > https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would > > have empty storageClass and could be bound by particular PVC only > > Hi Vadim, > Is this can solve the CNS issue at Comment 4 as well? Not yet, I'll prepare the fix for this as well (In reply to Vadim Rutkovsky from comment #29) > (In reply to Wenkai Shi from comment #28) > > (In reply to Vadim Rutkovsky from comment #27) > > > Created PR for master to fix it - > > > https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would > > > have empty storageClass and could be bound by particular PVC only > > > > Hi Vadim, > > Is this can solve the CNS issue at Comment 4 as well? > > Not yet, I'll prepare the fix for this as well Done in https://github.com/openshift/openshift-ansible/pull/9207 Merged, should be in next 3.10 build. *** Bug 1601083 has been marked as a duplicate of this bug. *** Check with version openshift-ansible-3.10.18-1.git.314.cfe4f91.el7, it's failed. Will check in next build. 3.10 PR for glusterfs - https://github.com/openshift/openshift-ansible/pull/9250, not yet in a build Fix is available in openshift-ansible-3.10.21-1 Now the latest puddle has openshift-ansible-3.10.18-1.git.314.cfe4f91.el7.noarch, NFS PR is already merged.
Verify NFS issue with openshift-ansible-3.10.18-1.git.314.cfe4f91.el7.noarch, and PASS.
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get pvc
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE
regpv-claim Bound regpv-volume 17G RWX 19m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get sc
NAME PROVISIONER AGE
standard (default) kubernetes.io/gce-pd 19m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get pv
NAME CAPACITY ACCESS MODES RECLAIM POLICY STATUS CLAIM STORAGECLASS REASON AGE
regpv-volume 17G RWX Retain Bound default/regpv-claim 19m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get po
NAME READY STATUS RESTARTS AGE
docker-registry-1-m9qn6 1/1 Running 0 18m
registry-console-1-x6qt8 1/1 Running 0 18m
router-1-m4pwg 1/1 Running 0 18m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc describe po docker-registry-1-m9qn6
Name: docker-registry-1-m9qn6
<--snip-->
Volumes:
registry-storage:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
ClaimName: regpv-claim
ReadOnly: false
<--snip-->
Keep ON_QA status, waiting newer puddle to run glusterfs issue verification with openshift-ansible-3.10.21-1.
Verified with version openshift-ansible-3.10.21-1.git.0.6446011.el7, installation can be done without error.
# oc get pvc
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE
registry-claim Bound registry-volume 5Gi RWX standard 14m
# oc get pv
NAME CAPACITY ACCESS MODES RECLAIM POLICY STATUS CLAIM STORAGECLASS REASON AGE
registry-volume 5Gi RWX Retain Bound default/registry-claim 14m
# oc get po
NAME READY STATUS RESTARTS AGE
docker-registry-1-sdn5k 1/1 Running 0 14m
glusterblock-registry-provisioner-dc-1-wq2sd 1/1 Running 0 16m
glusterfs-registry-bxwt2 1/1 Running 0 21m
glusterfs-registry-c4j6b 1/1 Running 0 21m
glusterfs-registry-dm7vx 1/1 Running 0 21m
heketi-registry-1-cqtzr 1/1 Running 0 17m
registry-console-1-7bjht 1/1 Running 0 13m
router-1-jv2bh 1/1 Running 0 15m
# oc describe po docker-registry-1-sdn5k
...
Volumes:
registry-storage:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
ClaimName: registry-claim
ReadOnly: false
...
According to this, customer can not use cns as docker-registry backend storage ever though glusterfs-registry group existing , which it works before. Is that's expect? @Jose
Wenkai, It seems the PVC is bound to PV correctly. Is PV being correctly provisioned as gluster? `oc describe` for the PV and PVC would be helpful to find out what's wrong here (In reply to Vadim Rutkovsky from comment #38) > Wenkai, > > It seems the PVC is bound to PV correctly. Is PV being correctly provisioned > as gluster? `oc describe` for the PV and PVC would be helpful to find out > what's wrong here Sorry, my bad. I've login to the pod and check the content. It's glusterfs. Just confused about the output of PVC's "StorageClass: standard". # oc describe pv Name: registry-volume Labels: <none> Annotations: <none> Finalizers: [kubernetes.io/pv-protection] StorageClass: Status: Bound Claim: default/registry-claim Reclaim Policy: Retain Access Modes: RWX Capacity: 5Gi Node Affinity: <none> Message: Source: Type: Glusterfs (a Glusterfs mount on the host that shares a pod's lifetime) EndpointsName: glusterfs-registry-endpoints Path: glusterfs-registry-volume ReadOnly: false Events: <none> # oc describe pvc Name: registry-claim Namespace: default StorageClass: standard Status: Bound Volume: registry-volume Labels: <none> Annotations: pv.kubernetes.io/bind-completed=yes pv.kubernetes.io/bound-by-controller=yes Finalizers: [kubernetes.io/pvc-protection] Capacity: 5Gi Access Modes: RWX Events: <none> # oc describe sc Name: standard IsDefaultClass: Yes Annotations: storageclass.beta.kubernetes.io/is-default-class=true Provisioner: kubernetes.io/gce-pd Parameters: type=pd-standard AllowVolumeExpansion: <unset> MountOptions: <none> ReclaimPolicy: Delete VolumeBindingMode: Immediate Events: <none> (In reply to Wenkai Shi from comment #39) > (In reply to Vadim Rutkovsky from comment #38) > > Wenkai, > > > > It seems the PVC is bound to PV correctly. Is PV being correctly provisioned > > as gluster? `oc describe` for the PV and PVC would be helpful to find out > > what's wrong here > > Sorry, my bad. I've login to the pod and check the content. It's glusterfs. > Just confused about the output of PVC's "StorageClass: standard". Yeah, that's expected - we're specifying empty StorageClass for PVC (so it gets 'standard'), but then manually binding PV to this PVC. Since we're not provisioning a new PV the storageclass here doesn't have an effect here. This is required for setup where PV and PVC is provisioned manually (NFS and GlusterFS). So this seems to work as expected, moving to ON_QA again. Move to VERIFIED per comment #40. This have been verified with version openshift-ansible-3.10.21-1.git.0.6446011.el7, it work as expected. I have also verified with openshift v3.10.24 and not happy with how this fix was done.
The action is correct, the registry PV does get created on the gluster_registry cluster. But the default SC is shown as where the PVC is claimed from. This deployment is on AWS and openshift_cloudprovider_kind=aws.
$ oc get sc
NAME PROVISIONER AGE
glusterfs-registry-block gluster.org/glusterblock 17h
glusterfs-storage kubernetes.io/glusterfs 17h
gp2 (default) kubernetes.io/aws-ebs 17h
STORAGECLASS shows 'gp2' for PVC, there is NO 10Gi EBS volume created
$ oc get pvc
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE
registry-claim Bound registry-volume 10Gi RWX gp2 17h
Again, the PVC looks like it was claimed from SC = gp2, it was not.
$ oc describe pvc registry-claim
Name: registry-claim
Namespace: default
StorageClass: gp2
Status: Bound
Volume: registry-volume
Labels: <none>
Annotations: pv.kubernetes.io/bind-completed=yes
pv.kubernetes.io/bound-by-controller=yes
Finalizers: [kubernetes.io/pvc-protection]
Capacity: 10Gi
Access Modes: RWX
Events: <none>
PV is Type: Glusterfs
$ oc describe pv registry-volume
Name: registry-volume
Labels: <none>
Annotations: <none>
Finalizers: [kubernetes.io/pv-protection]
StorageClass:
Status: Bound
Claim: default/registry-claim
Reclaim Policy: Retain
Access Modes: RWX
Capacity: 10Gi
Node Affinity: <none>
Message:
Source:
Type: Glusterfs (a Glusterfs mount on the host that shares a pod's lifetime)
EndpointsName: glusterfs-registry-endpoints
Path: glusterfs-registry-volume
ReadOnly: false
Events: <none>
I'm pretty sure this is just OpenShift/Kubernetes behavior due to the PVC's StorageClass being specified as "". Not sure this is something that can be remedied any time soon, and it would have to be fixed in Kubernetes first. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2018:2263 |
Created attachment 1453369 [details] installation log with inventory file embedded Description of problem: See the following details. Version-Release number of the following components: openshift-ansible-3.10.2-1.git.190.5abfddb.el7.noarch.rpm How reproducible: Always Steps to Reproduce: 1. Trigger an install on AWS with couldprovider enabled, and specify a nfs storage for docker-registry. openshift_hosted_registry_storage_kind=nfs openshift_enable_unsupported_configurations=true openshift_hosted_registry_storage_nfs_options="*(rw,root_squash,sync,no_wdelay)" openshift_hosted_registry_storage_nfs_directory=/var/lib/exports openshift_hosted_registry_storage_volume_name=regpv openshift_hosted_registry_storage_access_modes=["ReadWriteMany"] openshift_hosted_registry_storage_volume_size=17G 2. 3. Actual results: After installation, docker-registry get into "Pending" status due to its pvc is unbound. # oc get po NAME READY STATUS RESTARTS AGE docker-registry-1-bxmwz 0/1 Pending 0 8m docker-registry-1-deploy 1/1 Running 0 9m gp2 9m # oc describe po docker-registry-1-bxmwz <--snip--> Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedScheduling 2m (x26 over 9m) default-scheduler pod has unbound PersistentVolumeClaims # oc get pvc NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE regpv-claim Pending # oc get pv NAME CAPACITY ACCESS MODES RECLAIM POLICY STATUS CLAIM STORAGECLASS REASON AGE regpv-volume 17G RWX Retain Available 19m # oc get pvc regpv-claim -o yaml apiVersion: v1 kind: PersistentVolumeClaim metadata: annotations: volume.beta.kubernetes.io/storage-provisioner: kubernetes.io/aws-ebs creationTimestamp: 2018-06-21T07:35:45Z finalizers: - kubernetes.io/pvc-protection name: regpv-claim namespace: default resourceVersion: "1807" selfLink: /api/v1/namespaces/default/persistentvolumeclaims/regpv-claim uid: b5998ea2-7525-11e8-83c5-0e13c65a9cd6 spec: accessModes: - ReadWriteMany resources: requests: storage: 17G storageClassName: gp2 status: phase: Pending The registry pvc is trying to binding some provision volume with StorageClass "gp2". Expected results: default storage storage is created for any other app. docker-registry should use specified nfs storage pvc, but not some provision volume with StorageClass "gp2". Additional info: In the install log, registry pvc definition's storageclass is set to "", but after installation, the pvc is using the default "gp2" storageclass, I do not have idea what happened there. TASK [openshift_persistent_volumes : create standard pv and pvc lists] ********* Thursday 21 June 2018 03:35:39 -0400 (0:00:00.042) 0:11:18.092 ********* ok: [ec2-52-90-202-63.compute-1.amazonaws.com] => {"changed": false, "failed": false, "msg": "persistent_volumes list and persistent_volume_claims list created", "persistent_volume_claims": [{"access_modes": ["ReadWriteMany"], "annotations": [], "capacity": "17G", "name": "regpv-claim", "storageclass": ""}], "persistent_volumes": [{"access_modes": ["ReadWriteMany"], "capacity": "17G", "labels": {}, "name": "regpv-volume", "storage": {"nfs": {"path": "/var/lib/exports/regpv", "server": "ec2-52-90-202-63.compute-1.amazonaws.com"}}}]}