Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1872398

Summary: Upgrade from 4.4.15 to 4.5 nightly failed, kube-controller-manager is degraded
Product: OpenShift Container Platform Reporter: Paige Rubendall <prubenda>
Component: kube-controller-managerAssignee: Maciej Szulik <maszulik>
Status: CLOSED DUPLICATE QA Contact: zhou ying <yinzhou>
Severity: medium Docs Contact:
Priority: medium    
Version: 4.6CC: aos-bugs, knarra, mfojtik, mifiedle, prubenda
Target Milestone: ---   
Target Release: 4.6.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-09-01 16:24:14 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Paige Rubendall 2020-08-25 16:14:00 UTC
Description of problem:
Upgrade from 4.4.15 to 4.5.0-0.nightly-2020-08-23-191713 fails due to kube-controller-manager being degraded

Version-Release number of selected component (if applicable): 4.4.15


How reproducible: 100%


Steps to Reproduce:
1. Create 4.4.15 cluster (Disconnected UPI on OSP13 with RHCOS & RHEL7.8 FIPS On & OVN with https_proxy)
2. Upgrade to 4.5.0-0.nightly-2020-08-23-191713 (oc adm upgrade --to-image registry.svc.ci.openshift.org/ocp/release:4.5.0-0.nightly-2020-08-23-191713 --force --allow-explicit-upgrade)
3.

Actual results:
Operator becomes degraded and upgrade stops


% oc get clusterversion
NAME      VERSION   AVAILABLE   PROGRESSING   SINCE   STATUS
version   4.4.15    True        True          116m    Unable to apply 4.5.0-0.nightly-2020-08-23-191713: the cluster operator kube-controller-manager is degraded


Expected results:
Upgrade succeeds and cluster becomes 4.5.0-0.nightly-2020-08-23-191713 version properly

Additional info:


% oc get co
NAME                                       VERSION                             AVAILABLE   PROGRESSING   DEGRADED   SINCE
authentication                             4.4.15                              True        False         False      174m
cloud-credential                           4.4.15                              True        False         False      3h18m
cluster-autoscaler                         4.4.15                              True        False         False      3h4m
config-operator                            4.5.0-0.nightly-2020-08-23-191713   True        False         False      118m
console                                    4.4.15                              True        False         False      176m
csi-snapshot-controller                    4.4.15                              True        True          False      116m
dns                                        4.4.15                              True        False         False      3h11m
etcd                                       4.5.0-0.nightly-2020-08-23-191713   True        False         False      3h10m
image-registry                             4.4.15                              True        False         False      3h4m
ingress                                    4.4.15                              True        False         False      3h4m
insights                                   4.4.15                              True        False         False      3h7m
kube-apiserver                             4.5.0-0.nightly-2020-08-23-191713   True        False         False      3h10m
kube-controller-manager                    4.5.0-0.nightly-2020-08-23-191713   True        False         True       3h9m
kube-scheduler                             4.5.0-0.nightly-2020-08-23-191713   True        False         False      3h10m
kube-storage-version-migrator              4.4.15                              True        False         False      3h4m
machine-api                                4.5.0-0.nightly-2020-08-23-191713   True        False         False      3h6m
machine-approver                                                                                                    
machine-config                             4.4.15                              True        False         False      3h12m
marketplace                                4.4.15                              True        False         False      3h6m
monitoring                                 4.4.15                              False       True          True       105m
network                                    4.4.15                              True        False         False      3h12m
node-tuning                                4.4.15                              True        False         False      3h12m
openshift-apiserver                        4.5.0-0.nightly-2020-08-23-191713   True        False         False      50m
openshift-controller-manager               4.4.15                              True        False         False      3h4m
openshift-samples                          4.4.15                              True        False         False      3h3m
operator-lifecycle-manager                 4.4.15                              True        False         False      3h12m
operator-lifecycle-manager-catalog         4.4.15                              True        False         False      3h12m
operator-lifecycle-manager-packageserver   4.4.15                              True        False         False      171m
service-ca                                 4.4.15                              True        False         False      3h12m
service-catalog-apiserver                  4.4.15                              True        False         False      169m
service-catalog-controller-manager         4.4.15                              True        False         False      169m
storage                                    4.4.15                              True        False         False      3h7m


% oc describe co kube-controller-manager 
Name:         kube-controller-manager
Namespace:    
Labels:       <none>
Annotations:  <none>
API Version:  config.openshift.io/v1
Kind:         ClusterOperator
Metadata:
  Creation Timestamp:  2020-08-25T12:53:35Z
  Generation:          1
  Resource Version:    91903
  Self Link:           /apis/config.openshift.io/v1/clusteroperators/kube-controller-manager
  UID:                 32820ac0-c70e-4c10-807a-9f8bce0d2a55
Spec:
Status:
  Conditions:
    Last Transition Time:  2020-08-25T14:17:49Z
    Message:               StaticPodsDegraded: pod/kube-controller-manager-qe-pr-osp44159-tdr4t-master-1 container "kube-controller-manager" is not ready: CrashLoopBackOff: back-off 5m0s restarting failed container=kube-controller-manager pod=kube-controller-manager-qe-pr-osp44159-tdr4t-master-1_openshift-kube-controller-manager(f16cdad7583b408b91d85aaf0384b707)
StaticPodsDegraded: pod/kube-controller-manager-qe-pr-osp44159-tdr4t-master-1 container "kube-controller-manager" is waiting: CrashLoopBackOff: back-off 5m0s restarting failed container=kube-controller-manager pod=kube-controller-manager-qe-pr-osp44159-tdr4t-master-1_openshift-kube-controller-manager(f16cdad7583b408b91d85aaf0384b707)
StaticPodsDegraded: pod/kube-controller-manager-qe-pr-osp44159-tdr4t-master-0 container "kube-controller-manager" is not ready: CrashLoopBackOff: back-off 5m0s restarting failed container=kube-controller-manager pod=kube-controller-manager-qe-pr-osp44159-tdr4t-master-0_openshift-kube-controller-manager(f16cdad7583b408b91d85aaf0384b707)
StaticPodsDegraded: pod/kube-controller-manager-qe-pr-osp44159-tdr4t-master-0 container "kube-controller-manager" is waiting: CrashLoopBackOff: back-off 5m0s restarting failed container=kube-controller-manager pod=kube-controller-manager-qe-pr-osp44159-tdr4t-master-0_openshift-kube-controller-manager(f16cdad7583b408b91d85aaf0384b707)
StaticPodsDegraded: pod/kube-controller-manager-qe-pr-osp44159-tdr4t-master-2 container "kube-controller-manager" is not ready: CrashLoopBackOff: back-off 5m0s restarting failed container=kube-controller-manager pod=kube-controller-manager-qe-pr-osp44159-tdr4t-master-2_openshift-kube-controller-manager(f16cdad7583b408b91d85aaf0384b707)
StaticPodsDegraded: pod/kube-controller-manager-qe-pr-osp44159-tdr4t-master-2 container "kube-controller-manager" is waiting: CrashLoopBackOff: back-off 5m0s restarting failed container=kube-controller-manager pod=kube-controller-manager-qe-pr-osp44159-tdr4t-master-2_openshift-kube-controller-manager(f16cdad7583b408b91d85aaf0384b707)
    Reason:                StaticPods_Error
    Status:                True
    Type:                  Degraded
    Last Transition Time:  2020-08-25T14:13:50Z
    Message:               NodeInstallerProgressing: 3 nodes are at revision 11
    Reason:                AsExpected
    Status:                False
    Type:                  Progressing
    Last Transition Time:  2020-08-25T12:56:17Z
    Message:               StaticPodsAvailable: 3 nodes are active; 3 nodes are at revision 11
    Reason:                AsExpected
    Status:                True
    Type:                  Available
    Last Transition Time:  2020-08-25T12:53:36Z
    Reason:                AsExpected
    Status:                True
    Type:                  Upgradeable
  Extension:               <nil>
  Related Objects:
    Group:     operator.openshift.io
    Name:      cluster
    Resource:  kubecontrollermanagers
    Group:     
    Name:      openshift-config
    Resource:  namespaces
    Group:     
    Name:      openshift-config-managed
    Resource:  namespaces
    Group:     
    Name:      openshift-kube-controller-manager
    Resource:  namespaces
    Group:     
    Name:      openshift-kube-controller-manager-operator
    Resource:  namespaces
  Versions:
    Name:     raw-internal
    Version:  4.5.0-0.nightly-2020-08-23-191713
    Name:     kube-controller-manager
    Version:  1.18.3
    Name:     operator
    Version:  4.5.0-0.nightly-2020-08-23-191713
Events:       <none>


I was not able to get a must gather for this cluster

Comment 2 RamaKasturi 2020-08-26 05:36:13 UTC
Hi paige,

   Could you please help check the kcm logs to see if the issue is same as [1] which is closed duplicate of [2]? If yes, there is a fix available for [2] and the bug is ON_QA. Could you please help try with the latest nightly and see if the issue still happens ? I am trying to test the upgrade today if no issues found will update the bug here, thanks !! 
# oc logs -f <kcm_pod_name> -n openshift-kube-controller-manager

[1] https://bugzilla.redhat.com/show_bug.cgi?id=1870553
[2] https://bugzilla.redhat.com/show_bug.cgi?id=1869962

Thanks
kasturi

Comment 3 Paige Rubendall 2020-09-01 16:24:14 UTC
Looks like the same issue as [1] you mentioned. Got the same error. I think this issue can be marked as a duplicate


% oc logs kube-controller-manager-qe-pr-osp4415o-pl8rv-master-0 -n openshift-kube-controller-manager 
…
 1 disruption.go:331] Starting disruption controller
I0826 15:23:58.151303       1 shared_informer.go:223] Waiting for caches to sync for disruption
F0826 15:23:58.151326       1 plugins.go:123] Could not create hostpath recycler pod from file /etc/kubernetes/manifests/recycler-pod.yaml: failed to read file path /etc/kubernetes/manifests/recycler-pod.yaml: open /etc/kubernetes/manifests/recycler-pod.yaml: no such file or directory

% oc get pods -n openshift-kube-controller-manager
kube-controller-manager-qe-pr-osp4415o-pl8rv-master-0   3/4     Running            32         82m
kube-controller-manager-qe-pr-osp4415o-pl8rv-master-1   3/4     CrashLoopBackOff   31         81m
kube-controller-manager-qe-pr-osp4415o-pl8rv-master-2   4/4     Running            33         82m


 % oc get nodes
NAME                            STATUS   ROLES    AGE    VERSION
qe-pr-osp4415o-pl8rv-master-0   Ready    master   170m   v1.17.1+3288478
qe-pr-osp4415o-pl8rv-master-1   Ready    master   166m   v1.17.1+3288478
qe-pr-osp4415o-pl8rv-master-2   Ready    master   161m   v1.17.1+3288478
qe-pr-osp4415o-pl8rv-worker-0   Ready    worker   155m   v1.17.1+3288478
qe-pr-osp4415o-pl8rv-worker-1   Ready    worker   155m   v1.17.1+3288478
qe-pr-osp4415o-pl8rv-worker-2   Ready    worker   153m   v1.17.1+3288478


% oc debug node/qe-pr-osp4415o-pl8rv-master-0 
Starting pod/qe-pr-osp4415o-pl8rv-master-0-debug ...
To use host binaries, run `chroot /host`
Pod IP: 192.168.0.167
sh-4.4# chroot /host
sh-4.4# cd /etc/kubernetes/manifests
sh-4.4# ls
coredns.yaml   haproxy.yaml     kube-apiserver-pod.yaml           kube-scheduler-pod.yaml
etcd-pod.yaml  keepalived.yaml  kube-controller-manager-pod.yaml  mdns-publisher.yaml

*** This bug has been marked as a duplicate of bug 1870553 ***