Bug 1548352
| Summary: | Fail to upgrade containerzied ocp due to master was not ready for webconsole during 3.8-3.9 phase | ||
|---|---|---|---|
| Product: | OpenShift Container Platform | Reporter: | liujia <jiajliu> |
| Component: | Cluster Version Operator | Assignee: | Michael Gugino <mgugino> |
| Status: | CLOSED DUPLICATE | QA Contact: | liujia <jiajliu> |
| Severity: | urgent | Docs Contact: | |
| Priority: | urgent | ||
| Version: | 3.9.0 | CC: | aos-bugs, jokerman, mmccomas, wmeng, wsun |
| Target Milestone: | --- | ||
| Target Release: | 3.9.0 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2018-03-05 15:03:50 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | 1548358 | ||
| Bug Blocks: | |||
Block container ocp upgrade. Still hit the issue on openshift-ansible-3.9.0-0.51.0.git.0.e26400f.el7.noarch Still hit the issue on openshift-ansible-3.9.1-1.git.0.9862628.el7.noarch. (In reply to liujia from comment #4) > Still hit the issue on openshift-ansible-3.9.1-1.git.0.9862628.el7.noarch. Steps: Container install an old version ocp v3.9.0-0.48.0 without service-catalog deyployed. Enable 3.8 and 3.9 repos Run upgrade against above ocp to latest v3.9 Believe that this is just another side effect of https://bugzilla.redhat.com/show_bug.cgi?id=1548358 *** This bug has been marked as a duplicate of bug 1548358 *** |
Description of problem: Upgrade against non-ha containerized ocp, upgrade failed at task [openshift_web_console : Verify that the web console is running], but the root cause is that master node can not be ready for webconsole pod to be scheduled. # oc get pod -n openshift-web-console NAME READY STATUS RESTARTS AGE webconsole-54877f6577-fprlr 0/1 Unknown 0 19m webconsole-54877f6577-qwhbp 0/1 Pending 0 14m # oc get node NAME STATUS ROLES AGE VERSION qe-jliu-c3-master-etcd-1 NotReady master 1h v1.8.5+440f8d36da qe-jliu-c3-node-registry-router-1 Ready <none> 1h v1.7.6+a08f5eeb62 # docker ps CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES 64a1f0ee3220 openshift3/ose:v3.9.0 "/usr/bin/openshif..." 24 minutes ago Up 24 minutes atomic-openshift-master-controllers 3c737407ec8c openshift3/node:v3.9.0 "/usr/local/bin/or..." 25 minutes ago Up 25 minutes atomic-openshift-node 85526e79bcc5 openshift3/ose:v3.9.0 "/usr/bin/openshif..." 26 minutes ago Up 25 minutes atomic-openshift-master-api 6391f51257bb openshift3/openvswitch:v3.9.0 "/usr/local/bin/ov..." 26 minutes ago Up 26 minutes openvswitch aad396c15c86 registry.access.redhat.com/rhel7/etcd "/usr/bin/etcd" 26 minutes ago Up 26 minutes etcd_container Checked the log that task [openshift_node : Wait for node to be ready] just before task [openshift_web_console : Verify that the web console is running] is ok. Version-Release number of the following components: openshift-ansible-3.9.0-0.50.0.git.0.bb78b91.el7.noarch How reproducible: always Steps to Reproduce: 1.Containerized install ocp v3.7 without service catalog deployed and docker version was 1.12.6. 2.Enable new repos included ocp 3.8 and 3.9 and docker-1.13.1 3.Upgrade above ocp to v3.9 # ansible-playbook -i hosts /usr/share/ansible/openshift-ansible/playbooks/byo/openshift-cluster/upgrades/v3_9/upgrade.yml 3. Actual results: Upgrade failed. Expected results: Upgrade succeed. Additional info: # oc describe pod webconsole-54877f6577-qwhbp -n openshift-web-console Name: webconsole-54877f6577-qwhbp Namespace: openshift-web-console Node: <none> Labels: pod-template-hash=1043392133 webconsole=true Annotations: openshift.io/scc=restricted Status: Pending IP: Controlled By: ReplicaSet/webconsole-54877f6577 Containers: webconsole: Image: registry.reg-aws.openshift.com:443/openshift3/ose-web-console:v3.9.0 Port: 8443/TCP Command: /usr/bin/origin-web-console --audit-log-path=- -v=0 --config=/var/webconsole-config/webconsole-config.yaml Requests: cpu: 100m memory: 100Mi Liveness: exec [/bin/sh -i -c if [[ ! -f /tmp/webconsole-config.hash ]]; then \ md5sum /var/webconsole-config/webconsole-config.yaml > /tmp/webconsole-config.hash; \ elif [[ $(md5sum /var/webconsole-config/webconsole-config.yaml) != $(cat /tmp/webconsole-config.hash) ]]; then \ exit 1; \ fi && curl -k -f https://0.0.0.0:8443/console/] delay=0s timeout=1s period=10s #success=1 #failure=3 Readiness: http-get https://:8443/healthz delay=0s timeout=1s period=10s #success=1 #failure=3 Environment: <none> Mounts: /var/run/secrets/kubernetes.io/serviceaccount from webconsole-token-hcbw4 (ro) /var/serving-cert from serving-cert (rw) /var/webconsole-config from webconsole-config (rw) Conditions: Type Status PodScheduled False Volumes: serving-cert: Type: Secret (a volume populated by a Secret) SecretName: webconsole-serving-cert Optional: false webconsole-config: Type: ConfigMap (a volume populated by a ConfigMap) Name: webconsole-config Optional: false webconsole-token-hcbw4: Type: Secret (a volume populated by a Secret) SecretName: webconsole-token-hcbw4 Optional: false QoS Class: Burstable Node-Selectors: node-role.kubernetes.io/master=true Tolerations: node.kubernetes.io/memory-pressure:NoSchedule Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedScheduling 4m (x37 over 14m) default-scheduler 0/2 nodes are available: 1 MatchNodeSelector, 1 NodeNotReady, 1 NodeOutOfDisk. # oc describe node qe-jliu-c3-master-etcd-1 Name: qe-jliu-c3-master-etcd-1 Roles: master Labels: beta.kubernetes.io/arch=amd64 beta.kubernetes.io/instance-type=n1-standard-1 beta.kubernetes.io/os=linux failure-domain.beta.kubernetes.io/region=us-central1 failure-domain.beta.kubernetes.io/zone=us-central1-a kubernetes.io/hostname=qe-jliu-c3-master-etcd-1 node-role.kubernetes.io/master=true role=node Annotations: volumes.kubernetes.io/controller-managed-attach-detach=true Taints: <none> CreationTimestamp: Fri, 23 Feb 2018 02:17:52 -0500 Conditions: Type Status LastHeartbeatTime LastTransitionTime Reason Message ---- ------ ----------------- ------------------ ------ ------- NetworkUnavailable False Mon, 01 Jan 0001 00:00:00 +0000 Fri, 23 Feb 2018 02:17:52 -0500 RouteCreated openshift-sdn cleared kubelet-set NoRouteCreated OutOfDisk Unknown Fri, 23 Feb 2018 03:07:32 -0500 Fri, 23 Feb 2018 03:10:51 -0500 NodeStatusUnknown Kubelet stopped posting node status. MemoryPressure Unknown Fri, 23 Feb 2018 03:07:32 -0500 Fri, 23 Feb 2018 03:10:51 -0500 NodeStatusUnknown Kubelet stopped posting node status. DiskPressure Unknown Fri, 23 Feb 2018 03:07:32 -0500 Fri, 23 Feb 2018 03:10:51 -0500 NodeStatusUnknown Kubelet stopped posting node status. Ready Unknown Fri, 23 Feb 2018 03:07:32 -0500 Fri, 23 Feb 2018 03:10:51 -0500 NodeStatusUnknown Kubelet stopped posting node status. Addresses: InternalIP: 10.240.0.55 ExternalIP: 35.192.157.52 Hostname: qe-jliu-c3-master-etcd-1 Capacity: cpu: 1 memory: 3623852Ki pods: 250 Allocatable: cpu: 1 memory: 3521452Ki pods: 250 System Info: Machine ID: 6578704f71144944bcf05068370a5315 System UUID: 4684624D-A1F7-E4FD-6C79-4590221982AB Boot ID: 7a299c1d-02ec-41a1-8f24-efef3eef8b95 Kernel Version: 3.10.0-693.11.1.el7.x86_64 OS Image: Red Hat Enterprise Linux Server 7.4 (Maipo) Operating System: linux Architecture: amd64 Container Runtime Version: docker://1.13.1 Kubelet Version: v1.8.5+440f8d36da Kube-Proxy Version: v1.8.5+440f8d36da ExternalID: 2774988619954389608 Non-terminated Pods: (1 in total) Namespace Name CPU Requests CPU Limits Memory Requests Memory Limits --------- ---- ------------ ---------- --------------- ------------- openshift-web-console webconsole-54877f6577-fprlr 100m (10%) 0 (0%) 100Mi (2%) 0 (0%) Allocated resources: (Total limits may be over 100 percent, i.e., overcommitted.) CPU Requests CPU Limits Memory Requests Memory Limits ------------ ---------- --------------- ------------- 100m (10%) 0 (0%) 100Mi (2%) 0 (0%) Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal NodeAllocatableEnforced 1h kubelet, qe-jliu-c3-master-etcd-1 Updated Node Allocatable limit across pods Normal Starting 1h kubelet, qe-jliu-c3-master-etcd-1 Starting kubelet. Normal NodeHasSufficientDisk 1h (x2 over 1h) kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientDisk Normal NodeHasSufficientMemory 1h (x2 over 1h) kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientMemory Normal NodeHasNoDiskPressure 1h (x2 over 1h) kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasNoDiskPressure Normal NodeReady 1h kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeReady Normal NodeNotSchedulable 1h kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeNotSchedulable Normal Starting 31m kubelet, qe-jliu-c3-master-etcd-1 Starting kubelet. Normal NodeAllocatableEnforced 31m kubelet, qe-jliu-c3-master-etcd-1 Updated Node Allocatable limit across pods Normal NodeNotReady 31m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeNotReady Normal NodeHasSufficientDisk 31m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientDisk Normal NodeHasNoDiskPressure 31m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasNoDiskPressure Normal NodeHasSufficientMemory 31m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientMemory Normal NodeReady 31m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeReady Normal NodeSchedulable 31m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeSchedulable Normal NodeNotSchedulable 24m (x2 over 31m) kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeNotSchedulable Normal Starting 23m kubelet, qe-jliu-c3-master-etcd-1 Starting kubelet. Normal NodeHasSufficientDisk 23m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientDisk Normal NodeHasSufficientMemory 23m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasSufficientMemory Normal NodeHasNoDiskPressure 23m kubelet, qe-jliu-c3-master-etcd-1 Node qe-jliu-c3-master-etcd-1 status is now: NodeHasNoDiskPressure Warning ImageGCFailed 3m (x4 over 18m) kubelet, qe-jliu-c3-master-etcd-1 failed to get imageFs info: unable to find data for container /