Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1525580

Summary: OSP 12 with ceph deployment fails at step 4
Product: Red Hat OpenStack Reporter: Anil Dhingra <adhingra>
Component: openstack-tripleo-heat-templatesAssignee: Giulio Fidente <gfidente>
Status: CLOSED ERRATA QA Contact: Yogev Rabl <yrabl>
Severity: high Docs Contact:
Priority: high    
Version: 12.0 (Pike)CC: gfidente, jjoyce, jomurphy, m.andre, mburns, pkilambi, rhallise, rhel-osp-director-maint, yrabl
Target Milestone: z1Keywords: Triaged, ZStream
Target Release: 12.0 (Pike)   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: openstack-tripleo-heat-templates-7.0.3-18.el7ost Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2018-01-30 21:25:26 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Attachments:
Description Flags
step-4-failur- show none

Description Anil Dhingra 2017-12-13 15:48:55 UTC
Description of problem:

Not sure where exactly it fails & reason for failure , attached logs from deployment-show commands

017-12-13 17:28:06Z [overcloud-AllNodesDeploySteps-dahj4rhfhd6d-ControllerDeployment_Step4-ya4ku325efj3.0]: CREATE_FAILED  Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2
2017-12-13 17:28:06Z [overcloud-AllNodesDeploySteps-dahj4rhfhd6d-ControllerDeployment_Step4-ya4ku325efj3]: UPDATE_FAILED  Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2
2017-12-13 17:28:07Z [overcloud-AllNodesDeploySteps-dahj4rhfhd6d.ControllerDeployment_Step4]: UPDATE_FAILED  resources.ControllerDeployment_Step4: Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2
2017-12-13 17:28:08Z [overcloud-AllNodesDeploySteps-dahj4rhfhd6d]: UPDATE_FAILED  resources.ControllerDeployment_Step4: Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2
2017-12-13 17:28:09Z [AllNodesDeploySteps]: UPDATE_FAILED  resources.AllNodesDeploySteps: resources.ControllerDeployment_Step4: Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2
2017-12-13 17:28:10Z [overcloud]: UPDATE_FAILED  resources.AllNodesDeploySteps: resources.ControllerDeployment_Step4: Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2

 Stack overcloud UPDATE_FAILED 

overcloud.AllNodesDeploySteps.ControllerDeployment_Step4.0:
  resource_type: OS::Heat::StructuredDeployment
  physical_resource_id: 6da6defe-33d7-4c43-933c-4262a9457ac3
  status: CREATE_FAILED
  status_reason: |
    Error: resources[0]: Deployment to server failed: deploy_status_code : Deployment exited with non-zero status code: 2
  deploy_stdout: |
    ...
            "stdout: 74a21e89c5865403c8715204417259059fe6c0b2972e8066dff128e11f47b526", 
            "stdout: 9feca427b1142040bacf3c6bde07b50f5383193bbc4c72220e94bc4c3a466550"
        ], 
        "failed_when_result": true
    }
        to retry, use: --limit @/var/lib/heat-config/heat-config-ansible/6c57ef68-c38f-4520-a48d-116aa18ef9c4_playbook.retry
    
    PLAY RECAP *********************************************************************
    localhost                  : ok=7    changed=2    unreachable=0    failed=1   
    
    (truncated, view all with --long)
  deploy_stderr: |

Heat Stack update failed.
Heat Stack update failed.

it may be for cinder or ceph but no clue below failed file has nothing usefull

[heat-admin@overcloud-controller-0 ~]$ sudo cat /var/lib/heat-config/heat-config-ansible/6c57ef68-c38f-4520-a48d-116aa18ef9c4_playbook.retry
localhost


Version-Release number of selected component (if applicable):


How reproducible:


Steps to Reproduce:
1.
2.
3.

Actual results:


Expected results:


Additional info:

Comment 1 Anil Dhingra 2017-12-13 15:51:44 UTC
Created attachment 1367445 [details]
step-4-failur- show

Comment 3 Martin André 2017-12-13 16:19:09 UTC
AFAIK status code: 2 usually indicates an external deployment tool error, in this case ceph-ansible. Moving this to the ceph DFG.

Comment 4 Yogev Rabl 2018-01-10 14:40:52 UTC
Please notice this error from the logs:


Error running ['docker', 'run', '--name', 'gnocchi_db_sync', '--label', 'config_id=tripleo_step4', '--label', 'container_name=gnocchi_db_sync', '--label', 'managed_by=paunch', '--label', 'config_data={\\\"command\\\": \\\"/usr/bin/bootstrap_h
ost_exec gnocchi_api su gnocchi -s /bin/bash -c \\\\'/usr/bin/gnocchi-upgrade --sacks-number=128\\\\'\\\", \\\"user\\\": \\\"root\\\", \\\"volumes\\\": [\\\"/etc/hosts:/etc/hosts:ro\\\", \\\"/etc/localtime:/etc/localtime:ro\\\", \\\"/etc/puppet:/etc/puppet:ro\\\", \\\"/et
c/pki/ca-trust/extracted:/etc/pki/ca-trust/extracted:ro\\\", \\\"/etc/pki/tls/certs/ca-bundle.crt:/etc/pki/tls/certs/ca-bundle.crt:ro\\\", \\\"/etc/pki/tls/certs/ca-bundle.trust.crt:/etc/pki/tls/certs/ca-bundle.trust.crt:ro\\\", \\\"/etc/pki/tls/cert.pem:/etc/pki/tls/cert
.pem:ro\\\", \\\"/dev/log:/dev/log\\\", \\\"/etc/ssh/ssh_known_hosts:/etc/ssh/ssh_known_hosts:ro\\\", \\\"/var/lib/config-data/gnocchi/etc/my.cnf.d/tripleo.cnf:/etc/my.cnf.d/tripleo.cnf:ro\\\", \\\"/var/lib/config-data/gnocchi/etc/gnocchi/:/etc/gnocchi/:ro\\\", \\\"/var/l
og/containers/gnocchi:/var/log/gnocchi\\\", \\\"/var/log/containers/httpd/gnocchi-api:/var/log/httpd\\\", \\\"/etc/ceph:/etc/ceph:ro\\\"], \\\"image\\\": \\\"172.16.0.1:8787/rhosp12-beta/openstack-gnocchi-api:12.0-20171129.1\\\", \\\"detach\\\": false, \\\"net\\\": \\\"ho
st\\\", \\\"privileged\\\": false}', '--net=host', '--privileged=false', '--user=root', '--volume=/etc/hosts:/etc/hosts:ro', '--volume=/etc/localtime:/etc/localtime:ro', '--volume=/etc/puppet:/etc/puppet:ro', '--volume=/etc/pki/ca-trust/extracted:/etc/pki/ca-trust/extract
ed:ro', '--volume=/etc/pki/tls/certs/ca-bundle.crt:/etc/pki/tls/certs/ca-bundle.crt:ro', '--volume=/etc/pki/tls/certs/ca-bundle.trust.crt:/etc/pki/tls/certs/ca-bundle.trust.crt:ro', '--volume=/etc/pki/tls/cert.pem:/etc/pki/tls/cert.pem:ro', '--volume=/dev/log:/dev/log', '
--volume=/etc/ssh/ssh_known_hosts:/etc/ssh/ssh_known_hosts:ro', '--volume=/var/lib/config-data/gnocchi/etc/my.cnf.d/tripleo.cnf:/etc/my.cnf.d/tripleo.cnf:ro', '--volume=/var/lib/config-data/gnocchi/etc/gnocchi/:/etc/gnocchi/:ro', '--volume=/var/log/containers/gnocchi:/var
/log/gnocchi', '--volume=/var/log/containers/httpd/gnocchi-api:/var/log/httpd', '--volume=/etc/ceph:/etc/ceph:ro', '172.16.0.1:8787/rhosp12-beta/openstack-gnocchi-api:12.0-20171129.1', '/usr/bin/bootstrap_host_exec', 'gnocchi_api', 'su', 'gnocchi', '-s', '/bin/bash', '-c'
, \\\"'/usr/bin/gnocchi-upgrade\\\", \\\"--sacks-number=128'\\\"]. [1]\"


I'm not sure that it is related to Ceph, but to gnocchi deployment.

Comment 5 Pradeep Kilambi 2018-01-10 15:16:14 UTC
please add all the relevant logs so we can narrow down where the issue is. Its not really clear from the error message. It could be that gnocchi upgrade is failing as ceph isnt ready yet, or perhaps, the issue could be with keystone as well. We need logs.

Comment 6 Anil Dhingra 2018-01-11 03:08:33 UTC
for original problem as mentioned in  description logs were attached as per C#1

as this issue happened during hackfest & original env was deleted after multiple retries , but failure at same stage so we finally decided to test without ceph.

any finding as per attachment in C#1

Comment 7 Giulio Fidente 2018-01-17 13:45:18 UTC
I think this was fixed upstream, see [1]

Could

1. https://bugs.launchpad.net/tripleo/+bug/1734134

Comment 12 Yogev Rabl 2018-01-25 19:15:11 UTC
verified in openstack-tripleo-heat-templates-7.0.3-21.el7ost.noarch

Comment 18 errata-xmlrpc 2018-01-30 21:25:26 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2018:0253