Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1593612

Summary: docker-registry pvc is trying to bind some provision volume with StorageClass "gp2" when already specify a nfs/glusterfs storage for docker-regsitry.
Product: OpenShift Container Platform Reporter: Johnny Liu <jialiu>
Component: InstallerAssignee: Vadim Rutkovsky <vrutkovs>
Status: CLOSED ERRATA QA Contact: Johnny Liu <jialiu>
Severity: high Docs Contact:
Priority: high    
Version: 3.10.0CC: aclewett, aos-bugs, bleanhar, jarrpa, jialiu, jokerman, lxia, mawong, mmccomas, sdodson, vrutkovs, weshi, xtian
Target Milestone: ---Keywords: Regression, TestBlocker
Target Release: 3.10.z   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2018-07-30 20:22:32 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Attachments:
Description Flags
installation log with inventory file embedded none

Description Johnny Liu 2018-06-21 08:06:02 UTC
Created attachment 1453369 [details]
installation log with inventory file embedded

Description of problem:
See the following details.

Version-Release number of the following components:
openshift-ansible-3.10.2-1.git.190.5abfddb.el7.noarch.rpm

How reproducible:
Always

Steps to Reproduce:
1. Trigger an install on AWS with couldprovider enabled, and specify a nfs storage for docker-registry.
openshift_hosted_registry_storage_kind=nfs
openshift_enable_unsupported_configurations=true
openshift_hosted_registry_storage_nfs_options="*(rw,root_squash,sync,no_wdelay)"
openshift_hosted_registry_storage_nfs_directory=/var/lib/exports
openshift_hosted_registry_storage_volume_name=regpv
openshift_hosted_registry_storage_access_modes=["ReadWriteMany"]
openshift_hosted_registry_storage_volume_size=17G
2.
3.

Actual results:
After installation, docker-registry get into "Pending" status due to its pvc is unbound.
# oc get po
NAME                       READY     STATUS    RESTARTS   AGE
docker-registry-1-bxmwz    0/1       Pending   0          8m
docker-registry-1-deploy   1/1       Running   0          9m                                      gp2            9m

# oc describe po docker-registry-1-bxmwz
<--snip-->
Events:
  Type     Reason            Age               From               Message
  ----     ------            ----              ----               -------
  Warning  FailedScheduling  2m (x26 over 9m)  default-scheduler  pod has unbound PersistentVolumeClaims


# oc get pvc
NAME          STATUS    VOLUME    CAPACITY   ACCESS MODES   STORAGECLASS   AGE
regpv-claim   Pending 

# oc get pv
NAME           CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS      CLAIM     STORAGECLASS   REASON    AGE
regpv-volume   17G        RWX            Retain           Available                                      19m

# oc get pvc regpv-claim -o yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  annotations:
    volume.beta.kubernetes.io/storage-provisioner: kubernetes.io/aws-ebs
  creationTimestamp: 2018-06-21T07:35:45Z
  finalizers:
  - kubernetes.io/pvc-protection
  name: regpv-claim
  namespace: default
  resourceVersion: "1807"
  selfLink: /api/v1/namespaces/default/persistentvolumeclaims/regpv-claim
  uid: b5998ea2-7525-11e8-83c5-0e13c65a9cd6
spec:
  accessModes:
  - ReadWriteMany
  resources:
    requests:
      storage: 17G
  storageClassName: gp2
status:
  phase: Pending

The registry pvc is trying to binding some provision volume with StorageClass "gp2".

Expected results:
default storage storage is created for any other app.
docker-registry should use specified nfs storage pvc, but not some provision volume with StorageClass "gp2".

Additional info:
In the install log, registry pvc definition's storageclass is set to "", but after installation, the pvc is using the default "gp2" storageclass, I do not have idea what happened there.

TASK [openshift_persistent_volumes : create standard pv and pvc lists] *********
Thursday 21 June 2018  03:35:39 -0400 (0:00:00.042)       0:11:18.092 ********* 
ok: [ec2-52-90-202-63.compute-1.amazonaws.com] => {"changed": false, "failed": false, "msg": "persistent_volumes list and persistent_volume_claims list created", "persistent_volume_claims": [{"access_modes": ["ReadWriteMany"], "annotations": [], "capacity": "17G", "name": "regpv-claim", "storageclass": ""}], "persistent_volumes": [{"access_modes": ["ReadWriteMany"], "capacity": "17G", "labels": {}, "name": "regpv-volume", "storage": {"nfs": {"path": "/var/lib/exports/regpv", "server": "ec2-52-90-202-63.compute-1.amazonaws.com"}}}]}

Comment 1 Liang Xia 2018-06-21 09:33:17 UTC
As stated in #comment 0, registry pvc definition's storageclass is set to "" via a field named storageclass, but the correct field name should be storageClassName. It is ignored then. As there is a default storage class, so we set that name into the pvc.

Might be caused by https://github.com/openshift/openshift-ansible/blob/master/roles/lib_utils/action_plugins/generate_pv_pvcs_list.py#L139

Comment 2 Scott Dodson 2018-06-21 13:13:40 UTC
This is likely happening because we're configuring a default storage class on AWS. We should be using S3 storage for any registry running in AWS. 

Moving to 3.10.z

Comment 3 Vadim Rutkovsky 2018-06-21 15:39:43 UTC
There are two issue here:
1) openshift_hosted_registry_storage_createpvc is not defined, so its assumed True. openshift_hosted_registry_storage_storageclass is not defined to its set to "", which resolves into gp2 on AWS.

In this case openshift-ansible installer creates a PVC and uses default storageclass set for this cloudprovider - its gp2. GP2 however doesn't support "ReadWriteMany" so it never gets bound.

kind=nfs creates a volume, but it doesn't match default storageclass.

There are two options here:

1) Create an NFS storageclass when NFS server is being deployed, set storageclass to the pvc if kind=nfs. This however would require NFS provisioner

2) Avoid using PVCs with openshift_hosted_registry_storage_create_pvc=false

The latter option is however broken

Comment 4 Wenkai Shi 2018-06-22 08:48:52 UTC
Description of problem:
Glusterfs as docker registry back end storage deployment failed due to docker-registry pvc is trying to bind provision volume with IaaS default StorageClass.
It's similarly issue as BZ #1593612, for CNS I believe it should fix in 3.10.0 since this is an basic scenarios.

Version-Release number of the following components:
openshift-ansible-3.10.5-1.git.204.3a9cb75.el7

How reproducible:
100%

Steps to Reproduce:
1. Deploy OCP with CNS as docker registry back end storage.
2.
3.

Actual results:

TASK [openshift_hosted : Wait for registry pods] *******************************
Friday 22 June 2018  04:01:53 -0400 (0:00:00.750)       0:25:23.419 *********** 
FAILED - RETRYING: Wait for registry pods (60 retries left).
...
FAILED - RETRYING: Wait for registry pods (1 retries left).
fatal: [master.example.com]: FAILED! => {"attempts": 60, "changed": false, "failed": true, "results": {"cmd": "/usr/bin/oc get pod --selector=docker-registry=default -o json -n default", "results": [{"apiVersion": "v1", "items": [], "kind": "List", "metadata": {"resourceVersion": "", "selfLink": ""}}], "returncode": 0}, "state": "list"}

Expected results:
Should pass here

Additional info:
# oc get po
NAME                                           READY     STATUS    RESTARTS   AGE
docker-registry-1-deploy                       1/1       Running   0          8m
docker-registry-1-v2s6p                        0/1       Pending   0          8m
glusterblock-registry-provisioner-dc-1-c8f4h   1/1       Running   0          10m
glusterfs-registry-4rc7g                       1/1       Running   0          14m
glusterfs-registry-9rh89                       1/1       Running   0          14m
glusterfs-registry-n4lkc                       1/1       Running   0          14m
heketi-registry-1-bccp7                        1/1       Running   0          10m
router-1-bt8st                                 1/1       Running   0          8m

# oc describe po docker-registry-1-v2s6p
Name:           docker-registry-1-v2s6p
Namespace:      default
Node:           <none>
Labels:         deployment=docker-registry-1
                deploymentconfig=docker-registry
                docker-registry=default
Annotations:    openshift.io/deployment-config.latest-version=1
                openshift.io/deployment-config.name=docker-registry
                openshift.io/deployment.name=docker-registry-1
                openshift.io/scc=restricted
Status:         Pending
IP:             
Controlled By:  ReplicationController/docker-registry-1
Containers:
  registry:
    Image:      registry.reg-aws.openshift.com:443/openshift3/ose-docker-registry:v3.10
    Port:       5000/TCP
    Host Port:  0/TCP
    Requests:
      cpu:      100m
      memory:   256Mi
    Liveness:   http-get https://:5000/healthz delay=10s timeout=5s period=10s #success=1 #failure=3
    Readiness:  http-get https://:5000/healthz delay=0s timeout=5s period=10s #success=1 #failure=3
    Environment:
      REGISTRY_HTTP_ADDR:                                     :5000
      REGISTRY_HTTP_NET:                                      tcp
      REGISTRY_HTTP_SECRET:                                   LyaZ4wm1KIK+g3l8kh5HDOperi0UcJOCUFyK3nrNzEA=
      REGISTRY_MIDDLEWARE_REPOSITORY_OPENSHIFT_ENFORCEQUOTA:  false
      REGISTRY_OPENSHIFT_SERVER_ADDR:                         docker-registry.default.svc:5000
      REGISTRY_HTTP_TLS_CERTIFICATE:                          /etc/secrets/registry.crt
      REGISTRY_HTTP_TLS_KEY:                                  /etc/secrets/registry.key
    Mounts:
      /etc/secrets from registry-certificates (rw)
      /registry from registry-storage (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from registry-token-5tw92 (ro)
Conditions:
  Type           Status
  PodScheduled   False 
Volumes:
  registry-storage:
    Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
    ClaimName:  registry-claim
    ReadOnly:   false
  registry-certificates:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  registry-certificates
    Optional:    false
  registry-token-5tw92:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  registry-token-5tw92
    Optional:    false
QoS Class:       Burstable
Node-Selectors:  registry=enabled
                 role=node
Tolerations:     node.kubernetes.io/memory-pressure:NoSchedule
Events:
  Type     Reason            Age               From               Message
  ----     ------            ----              ----               -------
  Warning  FailedScheduling  2m (x26 over 9m)  default-scheduler  pod has unbound PersistentVolumeClaims

# oc get sc
NAME                 PROVISIONER            AGE
standard (default)   kubernetes.io/gce-pd   11m

# oc get pv 
NAME              CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS      CLAIM     STORAGECLASS   REASON    AGE
registry-volume   5Gi        RWX            Retain           Available                                      9m

# oc get pvc 
NAME             STATUS    VOLUME    CAPACITY   ACCESS MODES   STORAGECLASS   AGE
registry-claim   Pending                                       standard       9m

# oc get pv -o yaml
apiVersion: v1
items:
- apiVersion: v1
  kind: PersistentVolume
  metadata:
    creationTimestamp: 2018-06-22T08:02:11Z
    finalizers:
    - kubernetes.io/pv-protection
    name: registry-volume
    namespace: ""
    resourceVersion: "2986"
    selfLink: /api/v1/persistentvolumes/registry-volume
    uid: 916c75e8-75f2-11e8-9d01-42010af0004b
  spec:
    accessModes:
    - ReadWriteMany
    capacity:
      storage: 5Gi
    glusterfs:
      endpoints: glusterfs-registry-endpoints
      path: glusterfs-registry-volume
    persistentVolumeReclaimPolicy: Retain
  status:
    phase: Available
kind: List
metadata:
  resourceVersion: ""
  selfLink: ""

# oc get pvc -o yaml
apiVersion: v1
items:
- apiVersion: v1
  kind: PersistentVolumeClaim
  metadata:
    annotations:
      volume.beta.kubernetes.io/storage-provisioner: kubernetes.io/gce-pd
    creationTimestamp: 2018-06-22T08:02:14Z
    finalizers:
    - kubernetes.io/pvc-protection
    name: registry-claim
    namespace: default
    resourceVersion: "2992"
    selfLink: /api/v1/namespaces/default/persistentvolumeclaims/registry-claim
    uid: 92d9cec6-75f2-11e8-9d01-42010af0004b
  spec:
    accessModes:
    - ReadWriteMany
    resources:
      requests:
        storage: 5Gi
    storageClassName: standard
  status:
    phase: Pending
kind: List
metadata:
  resourceVersion: ""
  selfLink: ""

# oc describe pvc registry-claim
Name:          registry-claim
Namespace:     default
StorageClass:  standard
Status:        Pending
Volume:        
Labels:        <none>
Annotations:   volume.beta.kubernetes.io/storage-provisioner=kubernetes.io/gce-pd
Finalizers:    [kubernetes.io/pvc-protection]
Capacity:      
Access Modes:  
Events:
  Type     Reason              Age               From                         Message
  ----     ------              ----              ----                         -------
  Warning  ProvisioningFailed  3m (x26 over 9m)  persistentvolume-controller  Failed to provision volume with StorageClass "standard": invalid AccessModes [ReadWriteMany]: only AccessModes [ReadWriteOnce ReadOnlyMany] are supported

Comment 5 Johnny Liu 2018-06-22 10:03:09 UTC
(In reply to Vadim Rutkovsky from comment #3)
> There are two options here:
> 
> 1) Create an NFS storageclass when NFS server is being deployed, set
> storageclass to the pvc if kind=nfs. This however would require NFS
> provisioner
> 
> 2) Avoid using PVCs with openshift_hosted_registry_storage_create_pvc=false
> 
> The latter option is however broken

Using NFS storageclass for pvc, or using specified static nfs pvc is two different scenarios, in this bug, I would like docker-registry use specific static nfs pvc.

If I remember right, QE hit such issue recently, should be 3.10 regression.

Based on comment 4, glusterfs testing is also hit such issue, we think we should fix it in 3.10.

@Weshi will run some more testing to get know which old version does not have such issue to prove it is a regression.

Comment 6 Wenkai Shi 2018-06-22 16:14:20 UTC
I've verify that this issue doesn't appear in openshift-ansible-3.10.1-1.git.157.2bb6250.el7. 
It appear since openshift-ansible-3.10.2-1.git.190.5abfddb.el7.

Comment 7 Johnny Liu 2018-06-25 07:37:54 UTC
QE's many automated TC is broken by this issue, hope we could have a quick fix for it.

Comment 8 Vadim Rutkovsky 2018-06-25 08:59:40 UTC
(In reply to Wenkai Shi from comment #4)
> Description of problem:
> Glusterfs as docker registry back end storage deployment failed due to
> docker-registry pvc is trying to bind provision volume with IaaS default
> StorageClass.
> It's similarly issue as BZ #1593612, for CNS I believe it should fix in
> 3.10.0 since this is an basic scenarios.

This is a different issue, it fails as Access Mode was set incorrectly:

>   Warning  ProvisioningFailed  3m (x26 over 9m)  persistentvolume-controller
> Failed to provision volume with StorageClass "standard": invalid AccessModes
> [ReadWriteMany]: only AccessModes [ReadWriteOnce ReadOnlyMany] are supported

Comment 9 Vadim Rutkovsky 2018-06-25 12:13:04 UTC
(In reply to Johnny Liu from comment #7)
> QE's many automated TC is broken by this issue, hope we could have a quick
> fix for it.

So this happens due to the default storageclass being set to gp2 in AWS case. PVC gets this storageclass automatically, cannot be bound (as AWS doesn't support ReadWriteMany).

Its possible to add a new storageclass for NFS, but that'd probably require NFS provisioner.

Current workaround - add `openshift_storageclass_default=false` to the inventory to avoid gp2 to be used when no storageclass is specified

Comment 10 Scott Dodson 2018-06-25 15:04:27 UTC
Moved to 3.10.z based on the provided workaround in comment 9 and the preference to deploy S3 based registries in AWS for production use.

Comment 11 Wenkai Shi 2018-06-26 05:41:21 UTC
(In reply to Scott Dodson from comment #10)
> Moved to 3.10.z based on the provided workaround in comment 9 and the

For CNS part, the workaround in comment 9 also works.

> preference to deploy S3 based registries in AWS for production use.

For production use, cns customer may always meet this issue. OCP installer always create storageclass and mark it as default storageclass. For cns customer,  glusterfs is always their choice whatever which IaaS they using. So it can not be avoid until they using the work around. 
Base on this, I believe it should fix in 3.10.

Comment 12 Johnny Liu 2018-06-26 11:13:02 UTC
(In reply to Vadim Rutkovsky from comment #9)
> (In reply to Johnny Liu from comment #7)
> > QE's many automated TC is broken by this issue, hope we could have a quick
> > fix for it.
> 
> So this happens due to the default storageclass being set to gp2 in AWS
> case. PVC gets this storageclass automatically, cannot be bound (as AWS
> doesn't support ReadWriteMany).
If I remember right, this issue is really introduced recently.
I just tried openshift-ansible-3.10.1-1.git.157.2bb6250.el7.noarch, totally the same configuration, no such issue, I am not sure what changed was done after 3.10.1, and what is the purpose of the change.
[root@ip-172-18-0-56 ~]# oc get pvc
NAME          STATUS    VOLUME         CAPACITY   ACCESS MODES   STORAGECLASS   AGE
regpv-claim   Bound     regpv-volume   17G        RWX                           16m
[root@ip-172-18-0-56 ~]# oc get pv
NAME           CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS    CLAIM                 STORAGECLASS   REASON    AGE
regpv-volume   17G        RWX            Retain           Bound     default/regpv-claim                            16m
[root@ip-172-18-0-56 ~]# oc get sc
NAME            PROVISIONER             AGE
gp2 (default)   kubernetes.io/aws-ebs   17m


> Its possible to add a new storageclass for NFS, but that'd probably require
> NFS provisioner.
Just as what I explained in comment 5, enable storageclass for NFS is a different scenarios.
In this bug, user specify a NFS storage for docker-registry, this PV is *only* for docker-registry PV.
 
> Current workaround - add `openshift_storageclass_default=false` to the
> inventory to avoid gp2 to be used when no storageclass is specified
In my scenario, I want docker-registry to use the specified NFS storage, and also want to enable gp2 for other pod's storage dynamic provision.

I do not think this is an edge scenario, I guess customer would also hit the same issue once this is released out.

Comment 13 Vadim Rutkovsky 2018-06-26 12:07:28 UTC
(In reply to Johnny Liu from comment #12)
> (In reply to Vadim Rutkovsky from comment #9)
> > (In reply to Johnny Liu from comment #7)
> > > QE's many automated TC is broken by this issue, hope we could have a quick
> > > fix for it.
> > 
> > So this happens due to the default storageclass being set to gp2 in AWS
> > case. PVC gets this storageclass automatically, cannot be bound (as AWS
> > doesn't support ReadWriteMany).
> If I remember right, this issue is really introduced recently.
> I just tried openshift-ansible-3.10.1-1.git.157.2bb6250.el7.noarch, totally
> the same configuration, no such issue, I am not sure what changed was done
> after 3.10.1, and what is the purpose of the change.

There was a change to set default storageclass on AWS, which is causing it

> > Current workaround - add `openshift_storageclass_default=false` to the
> > inventory to avoid gp2 to be used when no storageclass is specified
> In my scenario, I want docker-registry to use the specified NFS storage, and
> also want to enable gp2 for other pod's storage dynamic provision.
> 
> I do not think this is an edge scenario, I guess customer would also hit the
> same issue once this is released out.

This configuration is not supported, we shouldn't be blocking 3.10 due to this in any case

Comment 14 Vadim Rutkovsky 2018-06-26 12:34:24 UTC
(In reply to Vadim Rutkovsky from comment #13)
> (In reply to Johnny Liu from comment #12)
> > (In reply to Vadim Rutkovsky from comment #9)
> > > (In reply to Johnny Liu from comment #7)
> > > > QE's many automated TC is broken by this issue, hope we could have a quick
> > > > fix for it.
> > > 
> > > So this happens due to the default storageclass being set to gp2 in AWS
> > > case. PVC gets this storageclass automatically, cannot be bound (as AWS
> > > doesn't support ReadWriteMany).
> > If I remember right, this issue is really introduced recently.
> > I just tried openshift-ansible-3.10.1-1.git.157.2bb6250.el7.noarch, totally
> > the same configuration, no such issue, I am not sure what changed was done
> > after 3.10.1, and what is the purpose of the change.
> 
> There was a change to set default storageclass on AWS, which is causing it

The change in question is https://github.com/openshift/openshift-ansible/pull/8858, which ensures default storageclass is set correctly for PVCs

Comment 15 Wenkai Shi 2018-06-27 02:42:12 UTC
Could you please think about comment #11, if you agree with me, shall we quick fix this? It's really a blocker for CNS related test.

Comment 16 Johnny Liu 2018-06-27 05:59:05 UTC
(In reply to Vadim Rutkovsky from comment #13) 
> This configuration is not supported, we shouldn't be blocking 3.10 due to
> this in any case

One more confirm, so the following scenarios are not supported:
1. "docker-registry with specified nfs static PV + default storageclass enabled" is not supported. 
2. "docker-registry with specified glusterfs static PV + default storageclass enabled" is not supported.

Comment 17 Matthew Wong 2018-06-27 14:34:28 UTC
If we do not want the PVC to be given the default StorageClass its storageClass should be set to "". We do not need to disable the default StorageClass

Alternatively we may want to give the registry PV an arbitrary StorageClass name and set the PVC's storageClass to the same string to ensure they bind. Note that StorageClass on its own does not imply anything about dynamic provisioning, used in this way it would be like a label.

(I have not read every detail of this bug thoroughly, I just want to clarify the binding logic)

Comment 18 Vadim Rutkovsky 2018-06-27 15:45:30 UTC
(In reply to Matthew Wong from comment #17)
> If we do not want the PVC to be given the default StorageClass its
> storageClass should be set to "". We do not need to disable the default
> StorageClass

I think this is what we're setting - and default storageclass is being used.

> Alternatively we may want to give the registry PV an arbitrary StorageClass
> name and set the PVC's storageClass to the same string to ensure they bind.
> Note that StorageClass on its own does not imply anything about dynamic
> provisioning, used in this way it would be like a label.

So the recommended way is to create NFS storageclass and set registry's pv and pvc storageclass? Would it work automagically, no provisioners required? If yes then we can implement that, that would be a cleaner solution than relying on empty storageclass behaviour

Comment 24 Vadim Rutkovsky 2018-07-11 16:44:11 UTC
(In reply to Vadim Rutkovsky from comment #18)
> (In reply to Matthew Wong from comment #17)
> > If we do not want the PVC to be given the default StorageClass its
> > storageClass should be set to "". We do not need to disable the default
> > StorageClass
> 
> I think this is what we're setting - and default storageclass is being used.
> 
> > Alternatively we may want to give the registry PV an arbitrary StorageClass
> > name and set the PVC's storageClass to the same string to ensure they bind.
> > Note that StorageClass on its own does not imply anything about dynamic
> > provisioning, used in this way it would be like a label.
> 
> So the recommended way is to create NFS storageclass and set registry's pv
> and pvc storageclass? Would it work automagically, no provisioners required?
> If yes then we can implement that, that would be a cleaner solution than
> relying on empty storageclass behaviour

Matt, could you help me with understand what's needed to fix this?

Comment 25 Matthew Wong 2018-07-11 17:15:26 UTC
Yes, you can create a storageclass called 'x-registry' and set it on both the pv and pvc, there is no need for a provisioner to exist, the provisioner field can have basically anything in it as long as it fits a format like 'openshift.com/nfs-registry'.

It would indeed be cleaner than relying on empty storageclass behaviour, because from what I can tell, changing our playbooks to rely on that behaviour will be difficult as our scripts & the templates do not make a distinction between '' and nil. https://github.com/openshift/openshift-ansible/blob/master/roles/openshift_persistent_volumes/templates/persistent-volume-claim.yml.j2 . I am no openshift-ansible expert though.

There is another possible solution that will require a bit more work but will ensure pvc A always binds to pv B, documented here https://docs.openshift.com/container-platform/3.9/dev_guide/persistent_volumes.html#persistent-volumes-volumes-and-claim-prebinding . https://github.com/openshift/openshift-ansible/blob/192b8477faa018d31893b322f43f356ce3ac1c80/roles/openshift_provisioners/templates/pv.j2#L28 . This will be cleaner in the sense that the user won't see the single-use 'x-registry' storageclass every time they do 'oc get sc'.

sorry for the late reply

Comment 27 Vadim Rutkovsky 2018-07-12 12:07:28 UTC
Created PR for master to fix it - https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would have empty storageClass and could be bound by particular PVC only

Comment 28 Wenkai Shi 2018-07-12 14:27:08 UTC
(In reply to Vadim Rutkovsky from comment #27)
> Created PR for master to fix it -
> https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would
> have empty storageClass and could be bound by particular PVC only

Hi Vadim,
Is this can solve the CNS issue at Comment 4 as well?

Comment 29 Vadim Rutkovsky 2018-07-13 07:53:26 UTC
(In reply to Wenkai Shi from comment #28)
> (In reply to Vadim Rutkovsky from comment #27)
> > Created PR for master to fix it -
> > https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would
> > have empty storageClass and could be bound by particular PVC only
> 
> Hi Vadim,
> Is this can solve the CNS issue at Comment 4 as well?

Not yet, I'll prepare the fix for this as well

Comment 30 Vadim Rutkovsky 2018-07-16 12:40:10 UTC
(In reply to Vadim Rutkovsky from comment #29)
> (In reply to Wenkai Shi from comment #28)
> > (In reply to Vadim Rutkovsky from comment #27)
> > > Created PR for master to fix it -
> > > https://github.com/openshift/openshift-ansible/pull/9174. Now NFS PVCs would
> > > have empty storageClass and could be bound by particular PVC only
> > 
> > Hi Vadim,
> > Is this can solve the CNS issue at Comment 4 as well?
> 
> Not yet, I'll prepare the fix for this as well

Done in https://github.com/openshift/openshift-ansible/pull/9207

Comment 31 Scott Dodson 2018-07-18 20:22:17 UTC
Merged, should be in next 3.10 build.

Comment 32 Scott Dodson 2018-07-18 20:25:51 UTC
*** Bug 1601083 has been marked as a duplicate of this bug. ***

Comment 33 Wenkai Shi 2018-07-19 03:34:28 UTC
Check with version openshift-ansible-3.10.18-1.git.314.cfe4f91.el7, it's failed. Will check in next build.

Comment 34 Vadim Rutkovsky 2018-07-19 08:37:58 UTC
3.10 PR for glusterfs - https://github.com/openshift/openshift-ansible/pull/9250, not yet in a build

Comment 35 Vadim Rutkovsky 2018-07-20 07:56:02 UTC
Fix is available in openshift-ansible-3.10.21-1

Comment 36 Johnny Liu 2018-07-20 09:22:52 UTC
Now the latest puddle has openshift-ansible-3.10.18-1.git.314.cfe4f91.el7.noarch, NFS PR is already merged.

Verify NFS issue with openshift-ansible-3.10.18-1.git.314.cfe4f91.el7.noarch, and PASS.

[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get pvc
NAME          STATUS    VOLUME         CAPACITY   ACCESS MODES   STORAGECLASS   AGE
regpv-claim   Bound     regpv-volume   17G        RWX                           19m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get sc
NAME                 PROVISIONER            AGE
standard (default)   kubernetes.io/gce-pd   19m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get pv
NAME           CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS    CLAIM                 STORAGECLASS   REASON    AGE
regpv-volume   17G        RWX            Retain           Bound     default/regpv-claim                            19m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc get po 
NAME                       READY     STATUS    RESTARTS   AGE
docker-registry-1-m9qn6    1/1       Running   0          18m
registry-console-1-x6qt8   1/1       Running   0          18m
router-1-m4pwg             1/1       Running   0          18m
[root@qe-jialiu3100-master-etcd-nfs-1 ~]# oc describe po docker-registry-1-m9qn6
Name:           docker-registry-1-m9qn6
<--snip-->
Volumes:
  registry-storage:
    Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
    ClaimName:  regpv-claim
    ReadOnly:   false
<--snip-->


Keep ON_QA status, waiting newer puddle to run glusterfs issue verification with openshift-ansible-3.10.21-1.

Comment 37 Wenkai Shi 2018-07-23 10:15:10 UTC
Verified with version openshift-ansible-3.10.21-1.git.0.6446011.el7, installation can be done without error.

# oc get pvc 
NAME             STATUS    VOLUME            CAPACITY   ACCESS MODES   STORAGECLASS   AGE
registry-claim   Bound     registry-volume   5Gi        RWX            standard       14m

# oc get pv
NAME              CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS    CLAIM                    STORAGECLASS   REASON    AGE
registry-volume   5Gi        RWX            Retain           Bound     default/registry-claim                            14m

# oc get po
NAME                                           READY     STATUS    RESTARTS   AGE
docker-registry-1-sdn5k                        1/1       Running   0          14m
glusterblock-registry-provisioner-dc-1-wq2sd   1/1       Running   0          16m
glusterfs-registry-bxwt2                       1/1       Running   0          21m
glusterfs-registry-c4j6b                       1/1       Running   0          21m
glusterfs-registry-dm7vx                       1/1       Running   0          21m
heketi-registry-1-cqtzr                        1/1       Running   0          17m
registry-console-1-7bjht                       1/1       Running   0          13m
router-1-jv2bh                                 1/1       Running   0          15m

# oc describe po docker-registry-1-sdn5k
...
Volumes:
  registry-storage:
    Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
    ClaimName:  registry-claim
    ReadOnly:   false
...

According to this, customer can not use cns as docker-registry backend storage ever though glusterfs-registry group existing , which it works before. Is that's expect? @Jose

Comment 38 Vadim Rutkovsky 2018-07-23 13:24:46 UTC
Wenkai,

It seems the PVC is bound to PV correctly. Is PV being correctly provisioned as gluster? `oc describe` for the PV and PVC would be helpful to find out what's wrong here

Comment 39 Wenkai Shi 2018-07-23 13:33:36 UTC
(In reply to Vadim Rutkovsky from comment #38)
> Wenkai,
> 
> It seems the PVC is bound to PV correctly. Is PV being correctly provisioned
> as gluster? `oc describe` for the PV and PVC would be helpful to find out
> what's wrong here

Sorry, my bad. I've login to the pod and check the content. It's glusterfs.
Just confused about the output of PVC's "StorageClass:  standard".

# oc describe pv
Name:            registry-volume
Labels:          <none>
Annotations:     <none>
Finalizers:      [kubernetes.io/pv-protection]
StorageClass:    
Status:          Bound
Claim:           default/registry-claim
Reclaim Policy:  Retain
Access Modes:    RWX
Capacity:        5Gi
Node Affinity:   <none>
Message:         
Source:
    Type:           Glusterfs (a Glusterfs mount on the host that shares a pod's lifetime)
    EndpointsName:  glusterfs-registry-endpoints
    Path:           glusterfs-registry-volume
    ReadOnly:       false
Events:             <none>

# oc describe pvc 
Name:          registry-claim
Namespace:     default
StorageClass:  standard
Status:        Bound
Volume:        registry-volume
Labels:        <none>
Annotations:   pv.kubernetes.io/bind-completed=yes
               pv.kubernetes.io/bound-by-controller=yes
Finalizers:    [kubernetes.io/pvc-protection]
Capacity:      5Gi
Access Modes:  RWX
Events:        <none>

# oc describe sc 
Name:                  standard
IsDefaultClass:        Yes
Annotations:           storageclass.beta.kubernetes.io/is-default-class=true
Provisioner:           kubernetes.io/gce-pd
Parameters:            type=pd-standard
AllowVolumeExpansion:  <unset>
MountOptions:          <none>
ReclaimPolicy:         Delete
VolumeBindingMode:     Immediate
Events:                <none>

Comment 40 Vadim Rutkovsky 2018-07-23 13:48:34 UTC
(In reply to Wenkai Shi from comment #39)
> (In reply to Vadim Rutkovsky from comment #38)
> > Wenkai,
> > 
> > It seems the PVC is bound to PV correctly. Is PV being correctly provisioned
> > as gluster? `oc describe` for the PV and PVC would be helpful to find out
> > what's wrong here
> 
> Sorry, my bad. I've login to the pod and check the content. It's glusterfs.
> Just confused about the output of PVC's "StorageClass:  standard".

Yeah, that's expected - we're specifying empty StorageClass for PVC (so it gets 'standard'), but then manually binding PV to this PVC. Since we're not provisioning a new PV the storageclass here doesn't have an effect here. This is required for setup where PV and PVC is provisioned manually (NFS and GlusterFS).

So this seems to work as expected, moving to ON_QA again.

Comment 41 Wenkai Shi 2018-07-23 14:16:10 UTC
Move to VERIFIED per comment #40. This have been verified with version openshift-ansible-3.10.21-1.git.0.6446011.el7, it work as expected.

Comment 42 Annette Clewett 2018-07-26 15:10:35 UTC
I have also verified with openshift v3.10.24 and not happy with how this fix was done. 

The action is correct, the registry PV does get created on the gluster_registry cluster. But the default SC is shown as where the PVC is claimed from. This deployment is on AWS and openshift_cloudprovider_kind=aws.

$ oc get sc
NAME                       PROVISIONER                AGE
glusterfs-registry-block   gluster.org/glusterblock   17h
glusterfs-storage          kubernetes.io/glusterfs    17h
gp2 (default)              kubernetes.io/aws-ebs      17h


STORAGECLASS shows 'gp2' for PVC, there is NO 10Gi EBS volume created

$ oc get pvc
NAME             STATUS    VOLUME            CAPACITY   ACCESS MODES   STORAGECLASS   AGE
registry-claim   Bound     registry-volume   10Gi       RWX            gp2            17h

Again, the PVC looks like it was claimed from SC = gp2, it was not.

$ oc describe pvc registry-claim 
Name:          registry-claim
Namespace:     default
StorageClass:  gp2
Status:        Bound
Volume:        registry-volume
Labels:        <none>
Annotations:   pv.kubernetes.io/bind-completed=yes
               pv.kubernetes.io/bound-by-controller=yes
Finalizers:    [kubernetes.io/pvc-protection]
Capacity:      10Gi
Access Modes:  RWX
Events:        <none>

PV is Type: Glusterfs

$ oc describe pv registry-volume 
Name:            registry-volume
Labels:          <none>
Annotations:     <none>
Finalizers:      [kubernetes.io/pv-protection]
StorageClass:    
Status:          Bound
Claim:           default/registry-claim
Reclaim Policy:  Retain
Access Modes:    RWX
Capacity:        10Gi
Node Affinity:   <none>
Message:         
Source:
    Type:           Glusterfs (a Glusterfs mount on the host that shares a pod's lifetime)
    EndpointsName:  glusterfs-registry-endpoints
    Path:           glusterfs-registry-volume
    ReadOnly:       false
Events:             <none>

Comment 43 Jose A. Rivera 2018-07-26 15:20:52 UTC
I'm pretty sure this is just OpenShift/Kubernetes behavior due to the PVC's StorageClass being specified as "". Not sure this is something that can be remedied any time soon, and it would have to be fixed in Kubernetes first.

Comment 45 errata-xmlrpc 2018-07-30 20:22:32 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2018:2263