Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
This project is now read‑only. Starting Monday, February 2, please use https://ibm-ceph.atlassian.net/ for all bug tracking management.

Bug 2149606

Summary: ceph orch daemon restart not working for monitors which went down 1+ days ago
Product: [Red Hat Storage] Red Hat Ceph Storage Reporter: Vasishta <vashastr>
Component: CephadmAssignee: Adam King <adking>
Status: CLOSED DUPLICATE QA Contact: Manisha Saini <msaini>
Severity: high Docs Contact: Anjana Suparna Sriram <asriram>
Priority: unspecified    
Version: 5.3CC: adking, cephqe-warriors
Target Milestone: ---   
Target Release: 6.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2022-12-08 15:39:47 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Vasishta 2022-11-30 11:11:49 UTC
Description of problem:
Trying ceph orch daemon on monitors which were down for 1+ days is not working.
Though orchestrator says that it Scheduled to restart, daemon doesn't get restarted on actual node.

Version-Release number of selected component (if applicable):
16.2.10-78.el8cp

How reproducible:
Tried on two different monitors on same node, reproducable on both nodes.

Steps to Reproduce:
1. Configured two clusters with different public and cluster network.
2. Enabled debug logs and log_to_file.
3. One monitor stopped and another monitor is in failed state as disk space got filled up.
4. Waited for a day and on mnitor node with stopped monitor, freed up space by deleting log files.
5. Tried ceph orch daemon restart 3-4 times with considerable gaps.
6. On the mon node, service was manually restarted, monitor was up and running.
7. On second monitor, ceph orch daemon restart was tried once, but it did not take place.
8. Without clearing space on second monitor node, manual restart was tried and service stopped immediately saying disk space is full.
9. Cleared space in second monitor 

Actual results:
ceph orch daemon restart doesn't restart mons

Expected results:
ceph orch daemon restart commands nedds to be able to restart mon once it schedules

Additional info: