Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1548352

Summary: Fail to upgrade containerzied ocp due to master was not ready for webconsole during 3.8-3.9 phase
Product: OpenShift Container Platform Reporter: liujia <jiajliu>
Component: Cluster Version OperatorAssignee: Michael Gugino <mgugino>
Status: CLOSED DUPLICATE QA Contact: liujia <jiajliu>
Severity: urgent Docs Contact:
Priority: urgent    
Version: 3.9.0CC: aos-bugs, jokerman, mmccomas, wmeng, wsun
Target Milestone: ---   
Target Release: 3.9.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2018-03-05 15:03:50 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1548358    
Bug Blocks:    

Description liujia 2018-02-23 09:46:33 UTC
Description of problem:
Upgrade against non-ha containerized ocp, upgrade failed at task [openshift_web_console : Verify that the web console is running], but the root cause is that master node can not be ready for webconsole pod to be scheduled.

# oc get pod -n openshift-web-console
NAME                          READY     STATUS    RESTARTS   AGE
webconsole-54877f6577-fprlr   0/1       Unknown   0          19m
webconsole-54877f6577-qwhbp   0/1       Pending   0          14m

# oc get node
NAME                                STATUS     ROLES     AGE       VERSION
qe-jliu-c3-master-etcd-1            NotReady   master    1h        v1.8.5+440f8d36da
qe-jliu-c3-node-registry-router-1   Ready      <none>    1h        v1.7.6+a08f5eeb62

# docker ps
CONTAINER ID        IMAGE                                   COMMAND                  CREATED             STATUS              PORTS               NAMES
64a1f0ee3220        openshift3/ose:v3.9.0                   "/usr/bin/openshif..."   24 minutes ago      Up 24 minutes                           atomic-openshift-master-controllers
3c737407ec8c        openshift3/node:v3.9.0                  "/usr/local/bin/or..."   25 minutes ago      Up 25 minutes                           atomic-openshift-node
85526e79bcc5        openshift3/ose:v3.9.0                   "/usr/bin/openshif..."   26 minutes ago      Up 25 minutes                           atomic-openshift-master-api
6391f51257bb        openshift3/openvswitch:v3.9.0           "/usr/local/bin/ov..."   26 minutes ago      Up 26 minutes                           openvswitch
aad396c15c86        registry.access.redhat.com/rhel7/etcd   "/usr/bin/etcd"          26 minutes ago      Up 26 minutes                           etcd_container

Checked the log that task [openshift_node : Wait for node to be ready] just before task [openshift_web_console : Verify that the web console is running] is ok.

Version-Release number of the following components:
openshift-ansible-3.9.0-0.50.0.git.0.bb78b91.el7.noarch

How reproducible:
always

Steps to Reproduce:
1.Containerized install ocp v3.7 without service catalog deployed and docker version was 1.12.6.
2.Enable new repos included ocp 3.8 and 3.9 and docker-1.13.1
3.Upgrade above ocp to v3.9
# ansible-playbook -i hosts /usr/share/ansible/openshift-ansible/playbooks/byo/openshift-cluster/upgrades/v3_9/upgrade.yml
3.

Actual results:
Upgrade failed.

Expected results:
Upgrade succeed.

Additional info:
# oc describe pod webconsole-54877f6577-qwhbp -n openshift-web-console
Name:           webconsole-54877f6577-qwhbp
Namespace:      openshift-web-console
Node:           <none>
Labels:         pod-template-hash=1043392133
                webconsole=true
Annotations:    openshift.io/scc=restricted
Status:         Pending
IP:             
Controlled By:  ReplicaSet/webconsole-54877f6577
Containers:
  webconsole:
    Image:  registry.reg-aws.openshift.com:443/openshift3/ose-web-console:v3.9.0
    Port:   8443/TCP
    Command:
      /usr/bin/origin-web-console
      --audit-log-path=-
      -v=0
      --config=/var/webconsole-config/webconsole-config.yaml
    Requests:
      cpu:     100m
      memory:  100Mi
    Liveness:  exec [/bin/sh -i -c if [[ ! -f /tmp/webconsole-config.hash ]]; then \
  md5sum /var/webconsole-config/webconsole-config.yaml > /tmp/webconsole-config.hash; \
elif [[ $(md5sum /var/webconsole-config/webconsole-config.yaml) != $(cat /tmp/webconsole-config.hash) ]]; then \
  exit 1; \
fi && curl -k -f https://0.0.0.0:8443/console/] delay=0s timeout=1s period=10s #success=1 #failure=3
    Readiness:    http-get https://:8443/healthz delay=0s timeout=1s period=10s #success=1 #failure=3
    Environment:  <none>
    Mounts:
      /var/run/secrets/kubernetes.io/serviceaccount from webconsole-token-hcbw4 (ro)
      /var/serving-cert from serving-cert (rw)
      /var/webconsole-config from webconsole-config (rw)
Conditions:
  Type           Status
  PodScheduled   False 
Volumes:
  serving-cert:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  webconsole-serving-cert
    Optional:    false
  webconsole-config:
    Type:      ConfigMap (a volume populated by a ConfigMap)
    Name:      webconsole-config
    Optional:  false
  webconsole-token-hcbw4:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  webconsole-token-hcbw4
    Optional:    false
QoS Class:       Burstable
Node-Selectors:  node-role.kubernetes.io/master=true
Tolerations:     node.kubernetes.io/memory-pressure:NoSchedule
Events:
  Type     Reason            Age                From               Message
  ----     ------            ----               ----               -------
  Warning  FailedScheduling  4m (x37 over 14m)  default-scheduler  0/2 nodes are available: 1 MatchNodeSelector, 1 NodeNotReady, 1 NodeOutOfDisk.


# oc describe node qe-jliu-c3-master-etcd-1
Name:               qe-jliu-c3-master-etcd-1
Roles:              master
Labels:             beta.kubernetes.io/arch=amd64
                    beta.kubernetes.io/instance-type=n1-standard-1
                    beta.kubernetes.io/os=linux
                    failure-domain.beta.kubernetes.io/region=us-central1
                    failure-domain.beta.kubernetes.io/zone=us-central1-a
                    kubernetes.io/hostname=qe-jliu-c3-master-etcd-1
                    node-role.kubernetes.io/master=true
                    role=node
Annotations:        volumes.kubernetes.io/controller-managed-attach-detach=true
Taints:             <none>
CreationTimestamp:  Fri, 23 Feb 2018 02:17:52 -0500
Conditions:
  Type                 Status    LastHeartbeatTime                 LastTransitionTime                Reason              Message
  ----                 ------    -----------------                 ------------------                ------              -------
  NetworkUnavailable   False     Mon, 01 Jan 0001 00:00:00 +0000   Fri, 23 Feb 2018 02:17:52 -0500   RouteCreated        openshift-sdn cleared kubelet-set NoRouteCreated
  OutOfDisk            Unknown   Fri, 23 Feb 2018 03:07:32 -0500   Fri, 23 Feb 2018 03:10:51 -0500   NodeStatusUnknown   Kubelet stopped posting node status.
  MemoryPressure       Unknown   Fri, 23 Feb 2018 03:07:32 -0500   Fri, 23 Feb 2018 03:10:51 -0500   NodeStatusUnknown   Kubelet stopped posting node status.
  DiskPressure         Unknown   Fri, 23 Feb 2018 03:07:32 -0500   Fri, 23 Feb 2018 03:10:51 -0500   NodeStatusUnknown   Kubelet stopped posting node status.
  Ready                Unknown   Fri, 23 Feb 2018 03:07:32 -0500   Fri, 23 Feb 2018 03:10:51 -0500   NodeStatusUnknown   Kubelet stopped posting node status.
Addresses:
  InternalIP:  10.240.0.55
  ExternalIP:  35.192.157.52
  Hostname:    qe-jliu-c3-master-etcd-1
Capacity:
 cpu:     1
 memory:  3623852Ki
 pods:    250
Allocatable:
 cpu:     1
 memory:  3521452Ki
 pods:    250
System Info:
 Machine ID:                 6578704f71144944bcf05068370a5315
 System UUID:                4684624D-A1F7-E4FD-6C79-4590221982AB
 Boot ID:                    7a299c1d-02ec-41a1-8f24-efef3eef8b95
 Kernel Version:             3.10.0-693.11.1.el7.x86_64
 OS Image:                   Red Hat Enterprise Linux Server 7.4 (Maipo)
 Operating System:           linux
 Architecture:               amd64
 Container Runtime Version:  docker://1.13.1
 Kubelet Version:            v1.8.5+440f8d36da
 Kube-Proxy Version:         v1.8.5+440f8d36da
ExternalID:                  2774988619954389608
Non-terminated Pods:         (1 in total)
  Namespace                  Name                           CPU Requests  CPU Limits  Memory Requests  Memory Limits
  ---------                  ----                           ------------  ----------  ---------------  -------------
  openshift-web-console      webconsole-54877f6577-fprlr    100m (10%)    0 (0%)      100Mi (2%)       0 (0%)
Allocated resources:
  (Total limits may be over 100 percent, i.e., overcommitted.)
  CPU Requests  CPU Limits  Memory Requests  Memory Limits
  ------------  ----------  ---------------  -------------
  100m (10%)    0 (0%)      100Mi (2%)       0 (0%)
Events:
  Type     Reason                   Age                From                               Message
  ----     ------                   ----               ----                               -------
  Normal   NodeAllocatableEnforced  1h                 kubelet, qe-jliu-c3-master-etcd-1  Updated Node Allocatable limit across pods
  Normal   Starting                 1h                 kubelet, qe-jliu-c3-master-etcd-1  Starting kubelet.
  Normal   NodeHasSufficientDisk    1h (x2 over 1h)    kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientDisk
  Normal   NodeHasSufficientMemory  1h (x2 over 1h)    kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientMemory
  Normal   NodeHasNoDiskPressure    1h (x2 over 1h)    kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasNoDiskPressure
  Normal   NodeReady                1h                 kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeReady
  Normal   NodeNotSchedulable       1h                 kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeNotSchedulable
  Normal   Starting                 31m                kubelet, qe-jliu-c3-master-etcd-1  Starting kubelet.
  Normal   NodeAllocatableEnforced  31m                kubelet, qe-jliu-c3-master-etcd-1  Updated Node Allocatable limit across pods
  Normal   NodeNotReady             31m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeNotReady
  Normal   NodeHasSufficientDisk    31m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientDisk
  Normal   NodeHasNoDiskPressure    31m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasNoDiskPressure
  Normal   NodeHasSufficientMemory  31m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientMemory
  Normal   NodeReady                31m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeReady
  Normal   NodeSchedulable          31m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeSchedulable
  Normal   NodeNotSchedulable       24m (x2 over 31m)  kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeNotSchedulable
  Normal   Starting                 23m                kubelet, qe-jliu-c3-master-etcd-1  Starting kubelet.
  Normal   NodeHasSufficientDisk    23m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientDisk
  Normal   NodeHasSufficientMemory  23m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientMemory
  Normal   NodeHasNoDiskPressure    23m                kubelet, qe-jliu-c3-master-etcd-1  Node qe-jliu-c3-master-etcd-1 status is now: NodeHasNoDiskPressure
  Warning  ImageGCFailed            3m (x4 over 18m)   kubelet, qe-jliu-c3-master-etcd-1  failed to get imageFs info: unable to find data for container /

Comment 2 liujia 2018-02-23 09:47:47 UTC
Block container ocp upgrade.

Comment 3 liujia 2018-02-24 03:31:27 UTC
Still hit the issue on openshift-ansible-3.9.0-0.51.0.git.0.e26400f.el7.noarch

Comment 4 liujia 2018-02-28 05:57:34 UTC
Still hit the issue on openshift-ansible-3.9.1-1.git.0.9862628.el7.noarch.

Comment 5 liujia 2018-02-28 07:28:47 UTC
(In reply to liujia from comment #4)
> Still hit the issue on openshift-ansible-3.9.1-1.git.0.9862628.el7.noarch.

Steps:
Container install an old version ocp v3.9.0-0.48.0 without service-catalog deyployed.
Enable 3.8 and 3.9 repos
Run upgrade against above ocp to latest v3.9

Comment 6 Scott Dodson 2018-03-05 15:03:50 UTC
Believe that this is just another side effect of https://bugzilla.redhat.com/show_bug.cgi?id=1548358

*** This bug has been marked as a duplicate of bug 1548358 ***