Bug 2314210 - Storagecluster remains stuck, MDS couldn't recover with CrashLoopBackOff during installation
Summary: Storagecluster remains stuck, MDS couldn't recover with CrashLoopBackOff duri...
Keywords:
Status: CLOSED DUPLICATE of bug 2315624
Alias: None
Product: Red Hat OpenShift Data Foundation
Classification: Red Hat Storage
Component: rook
Version: 4.17
Hardware: Unspecified
OS: Unspecified
unspecified
urgent
Target Milestone: ---
: ---
Assignee: Subham Rai
QA Contact: Neha Berry
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2024-09-23 12:51 UTC by Aman Agrawal
Modified: 2024-10-01 16:22 UTC (History)
4 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2024-10-01 16:22:18 UTC
Embargoed:


Attachments (Terms of Use)

Description Aman Agrawal 2024-09-23 12:51:11 UTC
Description of problem (please be detailed as possible and provide log
snippests):


Version of all relevant components (if applicable):
ODF 4.17.0-107
OCP 4.17.0-0.nightly-2024-09-22-211134
ceph version 18.2.1-229.el9cp (ef652b206f2487adfc86613646a4cac946f6b4e0) reef (stable)


Does this issue impact your ability to continue to work with the product
(please explain in detail what is the user impact)?


Is there any workaround available to the best of your knowledge?


Rate from 1 - 5 the complexity of the scenario you performed that caused this
bug (1 - very simple, 5 - very complex)?


Can this issue reproducible?


Can this issue reproduce from the UI?


If this is a regression, please provide more details to justify this:


Steps to Reproduce:
1. During ODF installation on a RDR setup, storagecluser remains stuck on one of the managed clusters C1.
2.
3.


Actual results: 

NAME                                                              READY   STATUS             RESTARTS         AGE     IP             NODE        NOMINATED NODE   READINESS GATES
rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5c5546b6kdvz9   1/2     CrashLoopBackOff   33 (55s ago)     3h8m    10.135.0.78    compute-0   <none>           <none>



oc describe pod rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5c5546b6kdvz9
Name:                 rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5c5546b6kdvz9
Namespace:            openshift-storage
Priority:             1000000000
Priority Class Name:  openshift-user-critical
Service Account:      rook-ceph-default
Node:                 compute-0/10.1.114.110
Start Time:           Mon, 23 Sep 2024 15:01:41 +0530
Labels:               app=rook-ceph-mds
                      app.kubernetes.io/component=cephfilesystems.ceph.rook.io
                      app.kubernetes.io/created-by=rook-ceph-operator
                      app.kubernetes.io/instance=ocs-storagecluster-cephfilesystem-b
                      app.kubernetes.io/managed-by=rook-ceph-operator
                      app.kubernetes.io/name=ceph-mds
                      app.kubernetes.io/part-of=ocs-storagecluster-cephfilesystem
                      ceph_daemon_id=ocs-storagecluster-cephfilesystem-b
                      ceph_daemon_type=mds
                      mds=ocs-storagecluster-cephfilesystem-b
                      odf-resource-profile=
                      pod-template-hash=5c5546b656
                      rook.io/operator-namespace=openshift-storage
                      rook_cluster=openshift-storage
                      rook_file_system=ocs-storagecluster-cephfilesystem
Annotations:          k8s.ovn.org/pod-networks:
                        {"default":{"ip_addresses":["10.135.0.78/23"],"mac_address":"0a:58:0a:87:00:4e","gateway_ips":["10.135.0.1"],"routes":[{"dest":"10.132.0.0...
                      k8s.v1.cni.cncf.io/network-status:
                        [{
                            "name": "ovn-kubernetes",
                            "interface": "eth0",
                            "ips": [
                                "10.135.0.78"
                            ],
                            "mac": "0a:58:0a:87:00:4e",
                            "default": true,
                            "dns": {}
                        }]
                      openshift.io/scc: rook-ceph
Status:               Running
IP:                   10.135.0.78
IPs:
  IP:           10.135.0.78
Controlled By:  ReplicaSet/rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5c5546b656
Init Containers:
  chown-container-data-dir:
    Container ID:  cri-o://3b2a0403f87ed97f78a0f26adc9aa32ad2094d02fe9929598e5a4d396aed141e
    Image:         registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:75bd8969ab3f86f2203a1ceb187876f44e54c9ee3b917518c4d696cf6cd88ce3
    Image ID:      registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:4f598dcdef399669e615b5624fd2ff3c4d152e44da2614e5aa5e286d628158ad
    Port:          <none>
    Host Port:     <none>
    Command:
      chown
    Args:
      --verbose
      --recursive
      ceph:ceph
      /var/log/ceph
      /var/lib/ceph/crash
      /run/ceph
      /var/lib/ceph/mds/ceph-ocs-storagecluster-cephfilesystem-b
    State:          Terminated
      Reason:       Completed
      Exit Code:    0
      Started:      Mon, 23 Sep 2024 15:01:42 +0530
      Finished:     Mon, 23 Sep 2024 15:01:42 +0530
    Ready:          True
    Restart Count:  0
    Limits:
      cpu:     2
      memory:  6Gi
    Requests:
      cpu:        2
      memory:     6Gi
    Environment:  <none>
    Mounts:
      /etc/ceph from rook-config-override (ro)
      /etc/ceph/keyring-store/ from rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-keyring (ro)
      /run/ceph from ceph-daemons-sock-dir (rw)
      /var/lib/ceph/crash from rook-ceph-crash (rw)
      /var/lib/ceph/mds/ceph-ocs-storagecluster-cephfilesystem-b from ceph-daemon-data (rw)
      /var/log/ceph from rook-ceph-log (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-lznvd (ro)
Containers:
  mds:
    Container ID:  cri-o://b8b7c482998386fba2b9cd7169dd1dc6859765acf4ab5866a18c6409bd14462a
    Image:         registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:75bd8969ab3f86f2203a1ceb187876f44e54c9ee3b917518c4d696cf6cd88ce3
    Image ID:      registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:4f598dcdef399669e615b5624fd2ff3c4d152e44da2614e5aa5e286d628158ad
    Port:          <none>
    Host Port:     <none>
    Command:
      ceph-mds
    Args:
      --fsid=36edbe41-be25-4f3b-87b9-26685471f6c4
      --keyring=/etc/ceph/keyring-store/keyring
      --default-log-to-stderr=true
      --default-err-to-stderr=true
      --default-mon-cluster-log-to-stderr=true
      --default-log-stderr-prefix=debug
      --default-log-to-file=false
      --default-mon-cluster-log-to-file=false
      --mon-host=$(ROOK_CEPH_MON_HOST)
      --mon-initial-members=$(ROOK_CEPH_MON_INITIAL_MEMBERS)
      --id=ocs-storagecluster-cephfilesystem-b
      --setuser=ceph
      --setgroup=ceph
      --foreground
      --public-addr=$(ROOK_POD_IP)
    State:          Waiting
      Reason:       CrashLoopBackOff
    Last State:     Terminated
      Reason:       Error
      Exit Code:    143
      Started:      Mon, 23 Sep 2024 18:17:02 +0530
      Finished:     Mon, 23 Sep 2024 18:20:01 +0530
    Ready:          False
    Restart Count:  35
    Limits:
      cpu:     2
      memory:  6Gi
    Requests:
      cpu:     2
      memory:  6Gi
    Liveness:  exec [timeout 20 sh -c #!/usr/bin/env bash
# do not use 'set -e -u' etc. because it is important to only fail this probe when failure is certain
# spurious failures risk destabilizing ceph or the filesystem

MDS_ID="ocs-storagecluster-cephfilesystem-b"
FILESYSTEM_NAME="ocs-storagecluster-cephfilesystem"
KEYRING="/etc/ceph/keyring-store/keyring"

outp="$(ceph fs dump --mon-host="$ROOK_CEPH_MON_HOST" --mon-initial-members="$ROOK_CEPH_MON_INITIAL_MEMBERS" --keyring "$KEYRING" --format json)"
rc=$?
if [ $rc -ne 0 ]; then
    echo "ceph MDS dump check failed with the following output:"
    echo "$outp"
    echo "passing probe to avoid restarting MDS. cannot determine if MDS is unhealthy. restarting MDS risks destabilizing ceph/filesystem, which is likely unreachable or in error state"
    exit 0
fi

# get the active and standby MDS in the fs map
standbyMds=$(echo "$outp" | jq ".standbys | map(.name) | any(.[]; . == \"$MDS_ID\")")
activeMds=$(echo "$outp" | jq ".filesystems[] | select(.mdsmap.fs_name == \"$FILESYSTEM_NAME\") | .mdsmap.info | map(.name) | any(.[]; . == \"$MDS_ID\")")

if [[ $standbyMds == true || $activeMds == true ]]; then
    echo "MDS ID present in MDS map, no need to re-start the container"
    exit 0
fi

echo "Error: MDS ID not present in MDS map"
exit 1
] delay=30s timeout=25s period=30s #success=1 #failure=5
    Environment:
      CONTAINER_IMAGE:                registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:75bd8969ab3f86f2203a1ceb187876f44e54c9ee3b917518c4d696cf6cd88ce3
      POD_NAME:                       rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5c5546b6kdvz9 (v1:metadata.name)
      POD_NAMESPACE:                  openshift-storage (v1:metadata.namespace)
      NODE_NAME:                       (v1:spec.nodeName)
      POD_MEMORY_LIMIT:               6442450944 (limits.memory)
      POD_MEMORY_REQUEST:             6442450944 (requests.memory)
      POD_CPU_LIMIT:                  2 (limits.cpu)
      POD_CPU_REQUEST:                2 (requests.cpu)
      CEPH_USE_RANDOM_NONCE:          true
      ROOK_MSGR2:                     msgr2_true_encryption_false_compression_false
      ROOK_CEPH_MON_HOST:             <set to the key 'mon_host' in secret 'rook-ceph-config'>             Optional: false
      ROOK_CEPH_MON_INITIAL_MEMBERS:  <set to the key 'mon_initial_members' in secret 'rook-ceph-config'>  Optional: false
      ROOK_POD_IP:                     (v1:status.podIP)
    Mounts:
      /etc/ceph from rook-config-override (ro)
      /etc/ceph/keyring-store/ from rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-keyring (ro)
      /run/ceph from ceph-daemons-sock-dir (rw)
      /var/lib/ceph/crash from rook-ceph-crash (rw)
      /var/lib/ceph/mds/ceph-ocs-storagecluster-cephfilesystem-b from ceph-daemon-data (rw)
      /var/log/ceph from rook-ceph-log (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-lznvd (ro)
  log-collector:
    Container ID:  cri-o://aecd3ca261e480ea352d2ab6a689c05ac5865ca694350ef6972d4206e8dade26
    Image:         registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:75bd8969ab3f86f2203a1ceb187876f44e54c9ee3b917518c4d696cf6cd88ce3
    Image ID:      registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:4f598dcdef399669e615b5624fd2ff3c4d152e44da2614e5aa5e286d628158ad
    Port:          <none>
    Host Port:     <none>
    Command:
      /bin/bash
      -x
      -e
      -m
      -c

      CEPH_CLIENT_ID=ceph-mds.ocs-storagecluster-cephfilesystem-b
      PERIODICITY=daily
      LOG_ROTATE_CEPH_FILE=/etc/logrotate.d/ceph
      LOG_MAX_SIZE=524M
      ROTATE=7

      # edit the logrotate file to only rotate a specific daemon log
      # otherwise we will logrotate log files without reloading certain daemons
      # this might happen when multiple daemons run on the same machine
      sed -i "s|*.log|$CEPH_CLIENT_ID.log|" "$LOG_ROTATE_CEPH_FILE"

      # replace default daily with given user input
      sed --in-place "s/daily/$PERIODICITY/g" "$LOG_ROTATE_CEPH_FILE"

      # replace rotate count, default 7 for all ceph daemons other than rbd-mirror
      sed --in-place "s/rotate 7/rotate $ROTATE/g" "$LOG_ROTATE_CEPH_FILE"

      if [ "$LOG_MAX_SIZE" != "0" ]; then
        # adding maxsize $LOG_MAX_SIZE at the 4th line of the logrotate config file with 4 spaces to maintain indentation
        sed --in-place "4i \ \ \ \ maxsize $LOG_MAX_SIZE" "$LOG_ROTATE_CEPH_FILE"
      fi

      while true; do
        # we don't force the logrorate but we let the logrotate binary handle the rotation based on user's input for periodicity and size
        logrotate --verbose "$LOG_ROTATE_CEPH_FILE"
        sleep 15m
      done

    State:          Running
      Started:      Mon, 23 Sep 2024 15:01:43 +0530
    Ready:          True
    Restart Count:  0
    Environment:    <none>
    Mounts:
      /etc/ceph from rook-config-override (ro)
      /run/ceph from ceph-daemons-sock-dir (rw)
      /var/lib/ceph/crash from rook-ceph-crash (rw)
      /var/log/ceph from rook-ceph-log (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-lznvd (ro)
Conditions:
  Type                        Status
  PodReadyToStartContainers   True
  Initialized                 True
  Ready                       False
  ContainersReady             False
  PodScheduled                True
Volumes:
  rook-config-override:
    Type:               Projected (a volume that contains injected data from multiple sources)
    ConfigMapName:      rook-config-override
    ConfigMapOptional:  <nil>
  rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-keyring:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-keyring
    Optional:    false
  ceph-daemons-sock-dir:
    Type:          HostPath (bare host directory volume)
    Path:          /var/lib/rook/exporter
    HostPathType:  DirectoryOrCreate
  rook-ceph-log:
    Type:          HostPath (bare host directory volume)
    Path:          /var/lib/rook/openshift-storage/log
    HostPathType:
  rook-ceph-crash:
    Type:          HostPath (bare host directory volume)
    Path:          /var/lib/rook/openshift-storage/crash
    HostPathType:
  ceph-daemon-data:
    Type:       EmptyDir (a temporary directory that shares a pod's lifetime)
    Medium:
    SizeLimit:  <unset>
  kube-api-access-lznvd:
    Type:                    Projected (a volume that contains injected data from multiple sources)
    TokenExpirationSeconds:  3607
    ConfigMapName:           kube-root-ca.crt
    ConfigMapOptional:       <nil>
    DownwardAPI:             true
    ConfigMapName:           openshift-service-ca.crt
    ConfigMapOptional:       <nil>
QoS Class:                   Burstable
Node-Selectors:              <none>
Tolerations:                 node.kubernetes.io/memory-pressure:NoSchedule op=Exists
                             node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 5s
                             node.ocs.openshift.io/storage=true:NoSchedule
Events:
  Type     Reason     Age   From     Message
  ----     ------     ----  ----     -------
  Warning  Unhealthy  178m  kubelet  Liveness probe failed: 2024-09-23T09:51:11.977+0000 7f85cb7fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:11.977+0000 7f85c3fff640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:14.976+0000 7f85c3fff640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:14.976+0000 7f85cb7fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:17.977+0000 7f85caffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:17.977+0000 7f85cb7fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:20.978+0000 7f85cb7fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:20.979+0000 7f85caffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:23.979+0000 7f85caffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:23.979+0000 7f85cb7fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
  Warning  Unhealthy  178m  kubelet  Liveness probe failed: 2024-09-23T09:51:43.379+0000 7f4d427fc640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:43.380+0000 7f4d42ffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:44.972+0000 7f4d437fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:44.974+0000 7f4d427fc640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:47.973+0000 7f4d42ffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:47.973+0000 7f4d437fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:50.973+0000 7f4d427fc640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:50.974+0000 7f4d437fe640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:53.974+0000 7f4d42ffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:53.974+0000 7f4d427fc640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:56.974+0000 7f4d427fc640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:56.974+0000 7f4d42ffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:59.974+0000 7f4d427fc640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:51:59.975+0000 7f4d42ffd640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
  Warning  Unhealthy  177m  kubelet  Liveness probe failed: 2024-09-23T09:52:11.970+0000 7f178dd9b640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:11.971+0000 7f178e59c640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:14.970+0000 7f178ed9d640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:14.970+0000 7f178dd9b640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:17.970+0000 7f178e59c640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:17.970+0000 7f178ed9d640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:20.970+0000 7f178ed9d640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:20.971+0000 7f178dd9b640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:23.972+0000 7f178dd9b640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:23.972+0000 7f178ed9d640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:26.972+0000 7f178e59c640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:26.972+0000 7f178dd9b640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:29.973+0000 7f178dd9b640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:29.973+0000 7f178ed9d640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
  Warning  Unhealthy  177m  kubelet  Liveness probe failed: 2024-09-23T09:52:41.985+0000 7f109d092640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:44.985+0000 7f109d092640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:47.985+0000 7f109d092640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:56.987+0000 7f109c891640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
2024-09-23T09:52:59.987+0000 7f109d092640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
  Warning  Unhealthy  176m                    kubelet  Liveness probe failed: 2024-09-23T09:53:21.171+0000 7fc2124eb640 -1 monclient(hunting): handle_auth_bad_method server allowed_methods [2] but i only support [2,1]
  Normal   Killing    173m (x2 over 176m)     kubelet  Container mds failed liveness probe, will be restarted
  Normal   Pulled     173m (x3 over 3h18m)    kubelet  Container image "registry.redhat.io/rhceph/rhceph-7-rhel9@sha256:75bd8969ab3f86f2203a1ceb187876f44e54c9ee3b917518c4d696cf6cd88ce3" already present on machine
  Normal   Created    173m (x3 over 3h18m)    kubelet  Created container mds
  Normal   Started    173m (x3 over 3h18m)    kubelet  Started container mds
  Warning  BackOff    8m26s (x336 over 155m)  kubelet  Back-off restarting failed container mds in pod rook-ceph-mds-ocs-storagecluster-cephfilesystem-b-5c5546b6kdvz9_openshift-storage(e453b242-7f69-4a32-95cf-aa08b99c31d4)
  Warning  Unhealthy  3m19s (x170 over 175m)  kubelet  Liveness probe failed:



------------------------------------------------------------------------------
Mon logs-
------------------------------------------------------------------------------
debug 2024-09-23T11:09:08.185+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.231+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.232+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.387+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.433+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.434+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.789+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.834+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.837+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:08.987+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:09.189+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:09.591+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:09.592+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:09.635+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:09.642+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:10.263+0000 7fe8a374f640  1 paxos.2).electionLogic(1843) init, last seen epoch 1843, mid-election, bumping
debug 2024-09-23T11:09:10.266+0000 7fe8a374f640  1 mon.d@2(electing) e6 collect_metadata sdb:  no unique device id for sdb: fallback method has model 'Virtual disk    ' but no serial'
debug 2024-09-23T11:09:10.393+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.185+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.232+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.233+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.388+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.433+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.436+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.611+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "mon remove", "name": "c", "format": "json"} v 0)
debug 2024-09-23T11:09:11.611+0000 7fe8a0f4a640  0 log_channel(audit) log [INF] : from='client.44914 10.132.2.40:0/4183111775' entity='client.admin' cmd={"prefix": "mon remove", "name": "c", "format": "json"} : dispatch
debug 2024-09-23T11:09:11.790+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.834+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.840+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.958+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:11.988+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.159+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.190+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.271+0000 7fe8a374f640 -1 mon.d@2(electing) e6 get_health_metrics reporting 5627 slow ops, oldest is log(1 entries from seq 42 at 2024-09-23T09:53:22.259800+0000)
debug 2024-09-23T11:09:12.274+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "mon remove", "name": "c", "format": "json"} v 0)
debug 2024-09-23T11:09:12.274+0000 7fe8a0f4a640  0 log_channel(audit) log [INF] : from='client.44914 10.132.2.40:0/4183111775' entity='client.admin' cmd={"prefix": "mon remove", "name": "c", "format": "json"} : dispatch
debug 2024-09-23T11:09:12.560+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.592+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.592+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.636+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:12.643+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:13.361+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:13.395+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:13.886+0000 7fe8a7113640  0 log_channel(audit) log [DBG] : from='admin socket' entity='admin socket' cmd='mon_status' args=[]: dispatch
debug 2024-09-23T11:09:13.887+0000 7fe8a7113640  0 log_channel(audit) log [DBG] : from='admin socket' entity='admin socket' cmd=mon_status args=[]: finished
debug 2024-09-23T11:09:14.185+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.232+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.233+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.387+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.433+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.435+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.789+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.835+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.837+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.958+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:14.988+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.159+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.190+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.266+0000 7fe8a374f640  1 paxos.2).electionLogic(1845) init, last seen epoch 1845, mid-election, bumping
debug 2024-09-23T11:09:15.269+0000 7fe8a374f640  1 mon.d@2(electing) e6 collect_metadata sdb:  no unique device id for sdb: fallback method has model 'Virtual disk    ' but no serial'
debug 2024-09-23T11:09:15.560+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.592+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.592+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.637+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:15.640+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:16.362+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:16.395+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.186+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.233+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.233+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.272+0000 7fe8a374f640 -1 mon.d@2(electing) e6 get_health_metrics reporting 5631 slow ops, oldest is log(1 entries from seq 42 at 2024-09-23T09:53:22.259800+0000)
debug 2024-09-23T11:09:17.388+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.417+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.434+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.435+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.619+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.790+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.835+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.838+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.959+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:17.988+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.022+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.160+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.192+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.562+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.593+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.597+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.637+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.641+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:18.825+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:19.364+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:19.401+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.186+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.233+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.269+0000 7fe8a374f640  1 paxos.2).electionLogic(1847) init, last seen epoch 1847, mid-election, bumping
debug 2024-09-23T11:09:20.272+0000 7fe8a374f640  1 mon.d@2(electing) e6 collect_metadata sdb:  no unique device id for sdb: fallback method has model 'Virtual disk    ' but no serial'
debug 2024-09-23T11:09:20.388+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.413+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.434+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.615+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.790+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.836+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.959+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:20.988+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.018+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.160+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.191+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.562+0000 7fe89ef46640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.580+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "osd blocklist ls", "format": "json"} v 0)
debug 2024-09-23T11:09:21.580+0000 7fe8a0f4a640  0 log_channel(audit) log [DBG] : from='mgr.20622 10.133.2.50:0/3476664645' entity='mgr.a' cmd={"prefix": "osd blocklist ls", "format": "json"} : dispatch
debug 2024-09-23T11:09:21.584+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "mon metadata", "id": "b"} v 0)
debug 2024-09-23T11:09:21.584+0000 7fe8a0f4a640  0 log_channel(audit) log [DBG] : from='mgr.20622 10.133.2.50:0/3476664645' entity='mgr.a' cmd={"prefix": "mon metadata", "id": "b"} : dispatch
debug 2024-09-23T11:09:21.584+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "mon metadata", "id": "c"} v 0)
debug 2024-09-23T11:09:21.584+0000 7fe8a0f4a640  0 log_channel(audit) log [DBG] : from='mgr.20622 10.133.2.50:0/3476664645' entity='mgr.a' cmd={"prefix": "mon metadata", "id": "c"} : dispatch
debug 2024-09-23T11:09:21.585+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "mon metadata", "id": "d"} v 0)
debug 2024-09-23T11:09:21.585+0000 7fe8a0f4a640  0 log_channel(audit) log [DBG] : from='mgr.20622 10.133.2.50:0/3476664645' entity='mgr.a' cmd={"prefix": "mon metadata", "id": "d"} : dispatch
debug 2024-09-23T11:09:21.585+0000 7fe8a0f4a640  0 mon.d@2(electing) e6 handle_command mon_command({"prefix": "mon metadata", "id": "e"} v 0)
debug 2024-09-23T11:09:21.585+0000 7fe8a0f4a640  0 log_channel(audit) log [DBG] : from='mgr.20622 10.133.2.50:0/3476664645' entity='mgr.a' cmd={"prefix": "mon metadata", "id": "e"} : dispatch
debug 2024-09-23T11:09:21.599+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.603+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.638+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:21.820+0000 7fe8a4f52640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:22.273+0000 7fe8a374f640 -1 mon.d@2(electing) e6 get_health_metrics reporting 5640 slow ops, oldest is log(1 entries from seq 42 at 2024-09-23T09:53:22.259800+0000)
debug 2024-09-23T11:09:22.363+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id
debug 2024-09-23T11:09:22.407+0000 7fe89f747640  1 mon.d@2(electing) e6 handle_auth_request failed to assign global_id


Expected results: MDS shouldn't crash and storagecluster should reach in Ready state


Additional info:


Note You need to log in before you can comment on or make changes to this bug.