Bug 2263256
| Summary: | [GSS] rook ceph operator crash loop backoff | ||
|---|---|---|---|
| Product: | [Red Hat Storage] Red Hat OpenShift Data Foundation | Reporter: | khover |
| Component: | ocs-operator | Assignee: | Malay Kumar parida <mparida> |
| Status: | CLOSED DEFERRED | QA Contact: | Mahesh Shetty <mashetty> |
| Severity: | medium | Docs Contact: | |
| Priority: | high | ||
| Version: | 4.12 | CC: | agogala, etamir, hnallurv, kramdoss, mashetty, mduasope, mparida, muagarwa, odf-bz-bot, sapillai, sheggodu, srai, tnielsen |
| Target Milestone: | --- | Keywords: | Reopened |
| Target Release: | ODF 4.12.12 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | 4.12.12-1 | Doc Type: | No Doc Update |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2024-04-01 13:42:32 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
khover
2024-02-07 21:54:17 UTC
Question: what was the state of the cluster before upgrading the cluster? Were mon's in quorum? Hello,
We dont have logs from before upgrade but customer stated no issues prior.
current state:
cluster:
id: e6d56c78-3697-4095-ba29-658a2745359e
health: HEALTH_OK
services:
mon: 5 daemons, quorum a,c,d,e,f (age 5m)
mgr: a(active, since 4d), standbys: b
mds: 1/1 daemons up, 1 hot standby
osd: 16 osds: 16 up (since 15h), 16 in (since 11w)
rgw: 2 daemons active (2 hosts, 1 zones)
data:
volumes: 1/1 healthy
pools: 12 pools, 449 pgs
objects: 376.25k objects, 1.2 TiB
usage: 4.9 TiB used, 18 TiB / 23 TiB avail
pgs: 449 active+clean
io:
client: 257 KiB/s rd, 22 MiB/s wr, 54 op/s rd, 2.86k op/s wr
Malay - could you pls share the steps to verify this bug? Steps to verify would be, Stretch cluster setup on ODF 4.12.10 Upgrade to ODF 4.12.11 & observer that rook-ceph-operator pod has gone to CLBO. Upgrade to ODF 4.12.12, now the rook-ceph-operator pod should be back to up and running. @tnielsen please add if anything else needs to be checked. Another verification step could be to look at the rook operator log and see many messages such as described in https://bugzilla.redhat.com/show_bug.cgi?id=2187952#c28 Closing the bug as we don't intend to test the fix in 4.12.12 for the reason above. |