Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
This project is now read‑only. Starting Monday, February 2, please use https://ibm-ceph.atlassian.net/ for all bug tracking management.

Bug 1594176

Summary: Error response from daemon: No such container: ceph-mon-controller-0
Product: [Red Hat Storage] Red Hat Ceph Storage Reporter: Filip Hubík <fhubik>
Component: Ceph-AnsibleAssignee: Guillaume Abrioux <gabrioux>
Status: CLOSED NOTABUG QA Contact: ceph-qe-bugs <ceph-qe-bugs>
Severity: medium Docs Contact:
Priority: unspecified    
Version: 3.1CC: aschoen, ceph-eng-bugs, fhubik, gfidente, gmeno, johfulto, nthomas, sankarshan
Target Milestone: rc   
Target Release: 3.*   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2018-06-28 20:15:07 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 1553640    
Attachments:
Description Flags
/var/lib/mistral/xyz/ansible.log
none
RC9 /var/lib/mistral/xyz/ansible.log
none
ansible.log with ceph-ansible RC9 package installed before OC deploy none

Description Filip Hubík 2018-06-22 10:22:53 UTC
Description of problem:
OSPd14 can not execute post-deployment ceph-ansible playbook. The error in /var/lib/mistral/xyz/ansible.log is:

...
"Thursday 21 June 2018  07:26:40 -0400 (0:00:00.062)       0:00:13.917 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : remove ceph nfs ganesha socket if exists and not used by a process] ***", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/check_socket_non_container.yml:194", "Thursday 21 June 2018  07:26:40 -0400 (0:00:00.048)       0:00:13.966 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : check if it is atomic host] ******************************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:2", "Thursday 21 June 2018  07:26:40 -0400 (0:00:00.047)       0:00:14.013 ********* ", "ok: [controller-0] => {\"changed\": false, \"stat\": {\"exists\": false}}", "", "TASK [ceph-defaults : set_fact is_atomic] **************************************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:7", "Thursday 21 June 2018  07:26:40 -0400 (0:00:00.501)       0:00:14.514 ********* ", "ok: [controller-0] => {\"ansible_facts\": {\"is_atomic\": false}, \"changed\": false}", "", "TASK [ceph-defaults : set_fact monitor_name ansible_hostname] ******************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:11", "Thursday 21 June 2018  07:26:40 -0400 (0:00:00.073)       0:00:14.588 ********* ", "ok: [controller-0] => {\"ansible_facts\": {\"monitor_name\": \"controller-0\"}, \"changed\": false}", "", "TASK [ceph-defaults : set_fact monitor_name ansible_fqdn] **********************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:17", "Thursday 21 June 2018  07:26:40 -0400 (0:00:00.077)       0:00:14.665 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : set_fact docker_exec_cmd] ********************************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:23", "Thursday 21 June 2018  07:26:40 -0400 (0:00:00.067)       0:00:14.732 ********* ", "ok: [controller-0 -> 192.168.24.8] => {\"ansible_facts\": {\"docker_exec_cmd\": \"docker exec ceph-mon-controller-0\"}, \"changed\": false}", "", "TASK [ceph-defaults : is ceph running already?] ********************************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:34", "Thursday 21 June 2018  07:26:41 -0400 (0:00:00.130)       0:00:14.863 ********* ", "ok: [controller-0 -> 192.168.24.8] => {\"changed\": false, \"cmd\": [\"timeout\", \"5\", \"docker\", \"exec\", \"ceph-mon-controller-0\", \"ceph\", \"--cluster\", \"ceph\", \"fsid\"], \"delta\": \"0:00:00.032773\", \"end\": \"2018-06-21 11:26:42.150915\", \"failed_when_result\": false, \"msg\": \"non-zero return code\", \"rc\": 1, \"start\": \"2018-06-21 11:26:42.118142\", \"stderr\": \"Error response from daemon: No such container: ceph-mon-controller-0\", \"stderr_lines\":

[\"Error response from daemon: No such container: ceph-mon-controller-0\"],

\"stdout\": \"\", \"stdout_lines\": []}", "", "TASK [ceph-defaults : check if /var/lib/mistral/2df872c0-03f8-4f54-928a-3e66a6fe858b/ceph-ansible/fetch_dir directory exists] ***", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:47", "Thursday 21 June 2018  07:26:41 -0400 (0:00:00.646)       0:00:15.509 ********* ", "ok: [controller-0 -> localhost] => {\"changed\": false, \"stat\": {\"exists\": false}}", "", "TASK [ceph-defaults : set_fact ceph_current_fsid rc 1] *************************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:57", "Thursday 21 June 2018  07:26:41 -0400 (0:00:00.186)       0:00:15.696 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : create a local fetch directory if it does not exist] *****", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:64", "Thursday 21 June 2018  07:26:41 -0400 (0:00:00.051)       0:00:15.747 ********* ", "ok: [controller-0 -> localhost] => {\"changed\": false, \"gid\": 985, \"group\": \"mistral\", \"mode\": \"0755\", \"owner\": \"mistral\", \"path\": \"/var/lib/mistral/2df872c0-03f8-4f54-928a-3e66a6fe858b/ceph-ansible/fetch_dir\", \"secontext\": \"system_u:object_r:var_lib_t:s0\", \"size\": 6, \"state\": \"directory\", \"uid\": 988}", "", "TASK [ceph-defaults : set_fact fsid ceph_current_fsid.stdout] ******************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:74", "Thursday 21 June 2018  07:26:42 -0400 (0:00:00.411)       0:00:16.159 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : set_fact ceph_release ceph_stable_release] ***************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:81", "Thursday 21 June 2018  07:26:42 -0400 (0:00:00.167)       0:00:16.327 ********* ", "ok: [controller-0] => {\"ansible_facts\": {\"ceph_release\": \"dummy\"}, \"changed\": false}", "", "TASK [ceph-defaults : generate cluster fsid] ***********************************", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:85", "Thursday 21 June 2018  07:26:42 -0400 (0:00:00.070)       0:00:16.398 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : reuse cluster fsid when cluster is already running] ******", "task path: /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:96", "Thursday 21 June 2018  07:26:42 -0400 (0:00:00.046)       0:00:16.444 ********* ", "skipping: [controller-0] => {\"changed\": false, \"skip_reason\": \"Conditional result was False\"}", "", "TASK [ceph-defaults : read cluster fsid if it already exists] ****

Version-Release number of selected component (if applicable):
OSPd14

How reproducible:
always

Steps to Reproduce:
1. Deploy OSPd14 using InfraRed topology 1:1.1:1
2. Workaround bug https://bugzilla.redhat.com/show_bug.cgi?id=1594169
using https://github.com/ceph/ceph-ansible/issues/2474
3. Redeploy overcloud using overcloud_deploy.sh again

Actual results:
Mistral execution fails with error:
[\"Error response from daemon: No such container: ceph-mon-controller-0\"],

Additional info:
ceph-ansible-3.1.0-0.1.beta7.el7cp.noarch
puppet-ceph-2.5.1-0.20180612095118.e6c31c4.el7ost.noarch
ansible-2.5.4-1.el7ae.noarch
python-tripleoclient-10.2.1-0.20180614123359.bac7317.el7ost.noarch
puppet-tripleo-9.1.1-0.20180614115802.5ad483e.el7ost.noarch
openstack-tripleo-common-9.1.1-0.20180614064914.33e8979.el7ost.noarch
openstack-tripleo-heat-templates-9.0.0-0.20180614125120.76c8fe9.el7ost.noarch
openstack-tripleo-ui-9.1.1-0.20180614162701.17501fb.el7ost.noarch
python-tripleoclient-heat-installer-10.2.1-0.20180614123359.bac7317.el7ost.noarch
openstack-tripleo-common-containers-9.1.1-0.20180614064914.33e8979.el7ost.noarch
openstack-tripleo-image-elements-9.0.0-0.20180601015717.2ac38dd.el7ost.noarch
openstack-tripleo-validations-9.1.1-0.20180613083358.1069c38.el7ost.noarch
ansible-role-tripleo-modify-image-0.0.1-0.20180531094252.f9e8e93.el7ost.noarch
openstack-tripleo-puppet-elements-9.0.0-0.20180602004307.939b586.el7ost.noarch
ansible-tripleo-ipsec-8.1.1-0.20180405121919.325d233.el7ost.noarch

Comment 1 Filip Hubík 2018-06-22 10:27:11 UTC
Created attachment 1453682 [details]
/var/lib/mistral/xyz/ansible.log

Comment 2 Filip Hubík 2018-06-22 10:31:50 UTC
First I considered this issue to be clone of https://bugzilla.redhat.com/show_bug.cgi?id=1555002 , but it looks like the mentioned upstream patch https://review.openstack.org/#/c/553364/1/workbooks/ceph-ansible.yaml is present in code (/usr/share/openstack-tripleo-common/workbooks/ceph-ansible.yaml undercloud) so this could be different issue?

Comment 5 John Fulton 2018-06-22 11:59:31 UTC
Filip, Would you please verify if you can reproduce this issue with ceph-ansible rc9 (you filed this with rc7) and then update the bug? A possibly related "no such container" bug was filed/fixed in rc9 as per 1590560.

Comment 7 Filip Hubík 2018-06-22 13:49:49 UTC
Created attachment 1453734 [details]
RC9 /var/lib/mistral/xyz/ansible.log

ansible.log with latest 3.1.0-0.1.rc9.el7cp.noarch.rpm installed in system.

Comment 8 Filip Hubík 2018-06-22 13:55:39 UTC
I retried OC deployment with ceph-ansible-3.1.0-0.1.rc9.el7cp - I still get the error, but OC deployment passed despite that. I attached ansible.log from run with rc9 package in place (https://bugzilla.redhat.com/attachment.cgi?id=1453734).

(In comparison with previous run I can see also additional "Error ENOENT: unrecognized pool 'xyz'" errors in log.)

Comment 9 Filip Hubík 2018-06-22 15:41:30 UTC
Note: Last log is from "unclean" deployment where just overcloud_deploy.sh was rerun without overcloud stack being deleted, so only post-deployment mistral steps were retried after rc9 rpm was installed. I have suspicion that it might have affected provided output here and some issues might be overlapping - so use such info for triage with caution.

Once I have free resources I'll provide clean ansible.log where rc9 package is applied in prior of first OC deploy run.

Comment 10 Filip Hubík 2018-06-25 10:29:54 UTC
Created attachment 1454318 [details]
ansible.log with ceph-ansible RC9 package installed before OC deploy

It seems like mentioned errors are still there, but they doesn't seem to be mandatory for OC deployment itself.

Looks like they are produced by tasks /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:34 and /usr/share/ceph-ansible/roles/ceph-osd/tasks/openstack_config.yml:12.


Are these expected to be bugs or can we consider them as Notabug?

Comment 11 John Fulton 2018-06-28 20:15:07 UTC
(In reply to Filip Hubík from comment #10)
> Created attachment 1454318 [details]
> ansible.log with ceph-ansible RC9 package installed before OC deploy
> 
> It seems like mentioned errors are still there, but they doesn't seem to be
> mandatory for OC deployment itself.
> 
> Looks like they are produced by tasks
> /usr/share/ceph-ansible/roles/ceph-defaults/tasks/facts.yml:34 and
> /usr/share/ceph-ansible/roles/ceph-osd/tasks/openstack_config.yml:12.
> 
> Are these expected to be bugs or can we consider them as Notabug?

There will be times when ceph-ansible checks if a container is available and then waits for it to be available which is normal. The ceph-ansible run looks good; no failures. I'm going to close this as not a bug, but if it comes up and if 'ceph-s' indicates that there's a problem with the deployed ceph cluster, then feel free to re-open.