Bug 1645554
| Summary: | Default deployment instructions for preprovisioned nodes deployment lead to errors with neutron-openvswitch-agent container on the compute node therefore making the environment unusable | ||
|---|---|---|---|
| Product: | Red Hat OpenStack | Reporter: | Punit Kundal <pkundal> |
| Component: | rhosp-director | Assignee: | RHOS Maint <rhos-maint> |
| Status: | CLOSED DUPLICATE | QA Contact: | Gurenko Alex <agurenko> |
| Severity: | high | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 13.0 (Queens) | CC: | bhaley, dbecker, jslagle, lars, mburns, morazi, pkundal |
| Target Milestone: | --- | ||
| Target Release: | --- | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2018-12-13 16:34:45 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
For openvswitch not starting, that is either of one of these 2 issues: https://bugzilla.redhat.com/show_bug.cgi?id=1642591 https://bugzilla.redhat.com/show_bug.cgi?id=1642588 As for the use of net-config-static.j2.yaml, that is just because DHCP from the director node is almost always never used for overcloud nodes when using pre-provisioned nodes. So, we default to using a template that sets the static IP that you have defined in DeployedServerPortMap. This is equivalent to the behavior when *not* using pre-provisioned nodes. In that case the default nic config template is actually net-config-noop.yaml, which doesn't create any bridges via the nic config either. The bridges you are used to seeing on compute nodes are probably created by Neutron itself, but in this case this did not happen due to the bugs I linked above around openvswitch not starting. In any case, our default nic config template mappings in environments/deployed-server-environment.j2.yaml is just a default. In almost all cases, you will want or need to override these mappings based on your environment and how you need to configure networking on the overcloud nodes. Hello, Installing openvswitch manually and starting the service prior to deploy helped in resolving the issue. Thanks for the help. Regards, Punit |
Description of problem: We have instructions in the documentation here at [1]; which dictate the procedure that one can use to create a basic overcloud with pre-provisioned nodes when one has no network isolation in place. Upon following these instructions it is noticed that the deployment completes but the openvswitch agent container is stuck in a restarting state: [root@compute ~]# docker ps -a CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES c39ae38974cd 192.168.24.1:8787/rhosp13/openstack-neutron-openvswitch-agent:13.0-59 "kolla_start" About an hour ago Restarting (1) 6 minutes ago neutron_ovs_agent 479cbd731f1d 192.168.24.1:8787/rhosp13/openstack-cron:13.0-60 "kolla_start" About an hour ago Up About an hour logrotate_crond d252a20faf5f 192.168.24.1:8787/rhosp13/openstack-nova-compute:13.0-62.1 "kolla_start" About an hour ago Up About an hour (healthy) nova_migration_target 287b0a526c46 192.168.24.1:8787/rhosp13/openstack-ceilometer-compute:13.0-55 "kolla_start" About an hour ago Up About an hour ceilometer_agent_compute 82b4fb472316 192.168.24.1:8787/rhosp13/openstack-nova-compute:13.0-62.1 "kolla_start" About an hour ago Up About an hour (healthy) nova_compute be9facace90f 192.168.24.1:8787/rhosp13/openstack-iscsid:13.0-54 "kolla_start" About an hour ago Up About an hour (healthy) iscsid eadde2634cfc 192.168.24.1:8787/rhosp13/openstack-nova-libvirt:13.0-64.1 "kolla_start" About an hour ago Up About an hour nova_libvirt 2fee704e0d10 192.168.24.1:8787/rhosp13/openstack-nova-libvirt:13.0-64.1 "kolla_start" About an hour ago Up About an hour nova_virtlogd 697e359d40cd 192.168.24.1:8787/rhosp13/openstack-nova-compute:13.0-62.1 "/docker-config-sc..." About an hour ago Exited (0) About an hour ago nova_statedir_owner 2bb50436b073 192.168.24.1:8787/rhosp13/openstack-neutron-server:13.0-58 "puppet apply --mo..." About an hour ago Exited (0) About an hour ago neutron_ovs_bridge The docker logs for the container show that the ovs agent is trying to connect to the ovsdb-server on the compute node: [root@compute ~]# docker logs c39ae38974cd .... ++ cat /run_command + CMD=/neutron_ovs_agent_launcher.sh + ARGS= + [[ ! -n '' ]] + . kolla_extend_start ++ [[ ! -d /var/log/kolla/neutron ]] +++ stat -c %a /var/log/kolla/neutron ++ [[ 2755 != \7\5\5 ]] ++ chmod 755 /var/log/kolla/neutron ++ . /usr/local/bin/kolla_neutron_extend_start + echo 'Running command: '\''/neutron_ovs_agent_launcher.sh'\''' Running command: '/neutron_ovs_agent_launcher.sh' + exec /neutron_ovs_agent_launcher.sh + /usr/bin/python -m neutron.cmd.destroy_patch_ports --config-file /usr/share/neutron/neutron-dist.conf --config-file /etc/neutron/neutron.conf --config-file /etc/neutron/plugins/ml2/openvswitch_agent.ini --config-dir /etc/neutron/conf.d/common --config-dir /etc/neutron/conf.d/neutron-openvswitch-agent Traceback (most recent call last): File "/usr/lib64/python2.7/runpy.py", line 162, in _run_module_as_main "__main__", fname, loader, pkg_name) File "/usr/lib64/python2.7/runpy.py", line 72, in _run_code exec code in run_globals File "/usr/lib/python2.7/site-packages/neutron/cmd/destroy_patch_ports.py", line 83, in <module> main() File "/usr/lib/python2.7/site-packages/neutron/cmd/destroy_patch_ports.py", line 78, in main port_cleaner = PatchPortCleaner(cfg.CONF) File "/usr/lib/python2.7/site-packages/neutron/cmd/destroy_patch_ports.py", line 44, in __init__ for bridge in mappings.values()] File "/usr/lib/python2.7/site-packages/neutron/agent/common/ovs_lib.py", line 217, in __init__ super(OVSBridge, self).__init__() File "/usr/lib/python2.7/site-packages/neutron/agent/common/ovs_lib.py", line 117, in __init__ self.ovsdb = ovsdb_api.from_config(self) File "/usr/lib/python2.7/site-packages/neutron/agent/ovsdb/api.py", line 31, in from_config return iface.api_factory(context) File "/usr/lib/python2.7/site-packages/neutron/agent/ovsdb/impl_idl.py", line 49, in api_factory idl=n_connection.idl_factory(), File "/usr/lib/python2.7/site-packages/neutron/agent/ovsdb/native/connection.py", line 69, in idl_factory helper = do_get_schema_helper() File "/usr/lib/python2.7/site-packages/tenacity/__init__.py", line 214, in wrapped_f return self.call(f, *args, **kw) File "/usr/lib/python2.7/site-packages/tenacity/__init__.py", line 295, in call start_time=start_time) File "/usr/lib/python2.7/site-packages/tenacity/__init__.py", line 265, in iter raise RetryError(fut).reraise() File "/usr/lib/python2.7/site-packages/tenacity/__init__.py", line 344, in reraise raise self.last_attempt.result() File "/usr/lib/python2.7/site-packages/concurrent/futures/_base.py", line 422, in result return self.__get_result() File "/usr/lib/python2.7/site-packages/tenacity/__init__.py", line 298, in call result = fn(*args, **kwargs) File "/usr/lib/python2.7/site-packages/neutron/agent/ovsdb/native/connection.py", line 67, in do_get_schema_helper return idlutils.get_schema_helper(conn, schema_name) File "/usr/lib/python2.7/site-packages/ovsdbapp/backend/ovs_idl/idlutils.py", line 128, in get_schema_helper 'err': os.strerror(err)}) Exception: Could not retrieve schema from tcp:127.0.0.1:6640: Connection refused but there is no process listening on port 6640 on the compute node: [root@compute ~]# netstat -tunpl | grep -i 6640 [root@compute ~]# The openvswitch service is dead on the compute node: [root@compute ~]# systemctl status openvswitch ● openvswitch.service - Open vSwitch Loaded: loaded (/usr/lib/systemd/system/openvswitch.service; disabled; vendor preset: disabled) Active: inactive (dead) Nov 02 13:36:18 compute.testlab.com systemd[1]: Dependency failed for Open vSwitch. Nov 02 13:36:18 compute.testlab.com systemd[1]: Job openvswitch.service/start failed with result 'dependency'. There is no bridge created on the compute node: [root@compute ~]# ip a 1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000 link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00 inet 127.0.0.1/8 scope host lo valid_lft forever preferred_lft forever inet6 ::1/128 scope host valid_lft forever preferred_lft forever 2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state UP group default qlen 1000 link/ether 52:54:00:d8:e2:5f brd ff:ff:ff:ff:ff:ff inet 192.168.24.215/24 brd 192.168.24.255 scope global eth0 valid_lft forever preferred_lft forever inet6 fe80::5054:ff:fed8:e25f/64 scope link valid_lft forever preferred_lft forever 3: eth1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state UP group default qlen 1000 link/ether 52:54:00:78:58:75 brd ff:ff:ff:ff:ff:ff inet 192.168.122.254/24 brd 192.168.122.255 scope global noprefixroute eth1 valid_lft forever preferred_lft forever inet6 fe80::5054:ff:fe78:5875/64 scope link valid_lft forever preferred_lft forever 4: docker0: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc noqueue state DOWN group default link/ether 02:42:3f:80:fe:15 brd ff:ff:ff:ff:ff:ff inet 172.31.0.1/24 scope global docker0 valid_lft forever preferred_lft forever inet6 fe80::42:3fff:fe80:fe15/64 scope link valid_lft forever preferred_lft forever I believe in a default configuration when no network isolation is in place, there should still be one br-ex created on the compute nodes along with the other bridges like br-int and br-tun for a basic vxlan network function but there is no such bridge created on the compute node. Whereas looking at the controller node, I see the basic networking layout being set: [root@controller ~]# ip a 1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000 link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00 inet 127.0.0.1/8 scope host lo valid_lft forever preferred_lft forever inet6 ::1/128 scope host valid_lft forever preferred_lft forever 2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast master ovs-system state UP group default qlen 1000 link/ether 52:54:00:cd:65:05 brd ff:ff:ff:ff:ff:ff inet6 fe80::5054:ff:fecd:6505/64 scope link valid_lft forever preferred_lft forever 3: eth1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state UP group default qlen 1000 link/ether 52:54:00:70:e3:7f brd ff:ff:ff:ff:ff:ff inet 192.168.122.61/24 brd 192.168.122.255 scope global noprefixroute eth1 valid_lft forever preferred_lft forever inet6 fe80::5054:ff:fe70:e37f/64 scope link valid_lft forever preferred_lft forever 4: ovs-system: <BROADCAST,MULTICAST> mtu 1500 qdisc noop state DOWN group default qlen 1000 link/ether 5a:cf:bc:7c:77:32 brd ff:ff:ff:ff:ff:ff 5: br-ex: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000 link/ether 52:54:00:cd:65:05 brd ff:ff:ff:ff:ff:ff inet 192.168.24.210/24 brd 192.168.24.255 scope global br-ex valid_lft forever preferred_lft forever inet 192.168.24.7/32 brd 192.168.24.255 scope global br-ex valid_lft forever preferred_lft forever inet 192.168.24.11/32 brd 192.168.24.255 scope global br-ex valid_lft forever preferred_lft forever inet6 fe80::5054:ff:fecd:6505/64 scope link valid_lft forever preferred_lft forever 6: docker0: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc noqueue state DOWN group default link/ether 02:42:b9:58:ca:1a brd ff:ff:ff:ff:ff:ff inet 172.31.0.1/24 scope global docker0 valid_lft forever preferred_lft forever inet6 fe80::42:b9ff:fe58:ca1a/64 scope link valid_lft forever preferred_lft forever 15: br-int: <BROADCAST,MULTICAST> mtu 1500 qdisc noop state DOWN group default qlen 1000 link/ether 2a:87:d0:31:12:43 brd ff:ff:ff:ff:ff:ff 16: br-tun: <BROADCAST,MULTICAST> mtu 1500 qdisc noop state DOWN group default qlen 1000 link/ether 3a:2c:cf:5b:3f:4c brd ff:ff:ff:ff:ff:ff [root@controller ~]# ovs-vsctl show dac7b242-818a-42a1-b133-8f0d4d99e843 Manager "ptcp:6640:127.0.0.1" is_connected: true Bridge br-int Controller "tcp:127.0.0.1:6633" is_connected: true fail_mode: secure Port patch-tun Interface patch-tun type: patch options: {peer=patch-int} Port int-br-ex Interface int-br-ex type: patch options: {peer=phy-br-ex} Port br-int Interface br-int type: internal Bridge br-tun Controller "tcp:127.0.0.1:6633" is_connected: true fail_mode: secure Port patch-int Interface patch-int type: patch options: {peer=patch-tun} Port br-tun Interface br-tun type: internal Bridge br-ex Controller "tcp:127.0.0.1:6633" is_connected: true fail_mode: secure Port phy-br-ex Interface phy-br-ex type: patch options: {peer=int-br-ex} Port br-ex Interface br-ex type: internal Port "eth0" Interface "eth0" ovs_version: "2.9.0" The reason behind no bridges being created on the compute node appears to be result of the below default template: +++ (undercloud) [stack@undercloud-13 ~]$ cat /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-environment.j2.yaml resource_registry: OS::TripleO::Server: ../deployed-server/deployed-server.yaml OS::TripleO::DeployedServer::ControlPlanePort: OS::Neutron::Port OS::TripleO::DeployedServer::Bootstrap: OS::Heat::None {% for role in roles %} # Default nic config mappings OS::TripleO::{{role.name}}::Net::SoftwareConfig: ../net-config-static.yaml {% endfor %} OS::TripleO::ControllerDeployedServer::Net::SoftwareConfig: ../net-config-static-bridge.yaml +++ net-config-static.yaml does not place any bridge configuration on the nodes on which it is applied: (undercloud) [stack@undercloud-13 ~]$ cat /usr/share/openstack-tripleo-heat-templates/net-config-static.j2.yaml heat_template_version: queens description: > Software Config to drive os-net-config for a simple interface with DHCP. parameters: ControlPlaneIp: default: '' description: IP address/subnet on the ctlplane network type: string {%- for network in networks %} {{network.name}}IpSubnet: default: '' description: IP address/subnet on the {{network.name_lower}} network type: string {%- endfor %} ControlPlaneSubnetCidr: # Override this via parameter_defaults default: '24' description: The subnet CIDR of the control plane network. type: string ControlPlaneDefaultRoute: # Override this via parameter_defaults description: The default route of the control plane network. type: string DnsServers: # Override this via parameter_defaults default: [] description: A list of DNS servers (2 max for some implementations) that will be added to resolv.conf. type: comma_delimited_list EC2MetadataIp: # Override this via parameter_defaults description: The IP address of the EC2 metadata server. type: string resources: OsNetConfigImpl: type: OS::Heat::SoftwareConfig properties: group: script config: str_replace: template: get_file: network/scripts/run-os-net-config.sh params: $network_config: network_config: - type: interface name: interface_name use_dhcp: false dns_servers: get_param: DnsServers addresses: - ip_netmask: list_join: - / - - get_param: ControlPlaneIp - get_param: ControlPlaneSubnetCidr routes: - ip_netmask: 169.254.169.254/32 next_hop: get_param: EC2MetadataIp - default: true next_hop: get_param: ControlPlaneDefaultRoute outputs: OS::stack_id: description: The OsNetConfigImpl resource. value: get_resource: OsNetConfigImpl it only configures the provisioning interface which is not enough for a normal vxlan use case Whereas this is what is being applied on the controller nodes: +++ (undercloud) [stack@undercloud-13 ~]$ cat /usr/share/openstack-tripleo-heat-templates/net-config-static-bridge.j2.yaml heat_template_version: queens description: > Software Config to drive os-net-config for a simple bridge configured with a static IP address for the ctlplane network. parameters: ControlPlaneIp: default: '' description: IP address/subnet on the ctlplane network type: string {%- for network in networks %} {{network.name}}IpSubnet: default: '' description: IP address/subnet on the {{network.name_lower}} network type: string {%- endfor %} ControlPlaneSubnetCidr: # Override this via parameter_defaults default: '24' description: The subnet CIDR of the control plane network. type: string ControlPlaneDefaultRoute: # Override this via parameter_defaults description: The default route of the control plane network. type: string DnsServers: # Override this via parameter_defaults default: [] description: A list of DNS servers (2 max for some implementations) that will be added to resolv.conf. type: comma_delimited_list EC2MetadataIp: # Override this via parameter_defaults description: The IP address of the EC2 metadata server. type: string resources: OsNetConfigImpl: type: OS::Heat::SoftwareConfig properties: group: script config: str_replace: template: get_file: network/scripts/run-os-net-config.sh params: $network_config: network_config: - type: ovs_bridge name: bridge_name use_dhcp: false dns_servers: get_param: DnsServers addresses: - ip_netmask: list_join: - / - - get_param: ControlPlaneIp - get_param: ControlPlaneSubnetCidr routes: - ip_netmask: 169.254.169.254/32 next_hop: get_param: EC2MetadataIp - default: true next_hop: get_param: ControlPlaneDefaultRoute members: - type: interface name: interface_name # force the MAC address of the bridge to this interface primary: true outputs: OS::stack_id: description: The OsNetConfigImpl resource. value: get_resource: OsNetConfigImpl +++ which can work when one has no network isolation in place. The documentation suggests to include: /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-environment.yaml Here is the deployment command that I used: +++ (undercloud) [stack@undercloud-13 ~]$ openstack overcloud deploy --disable-validations --templates -e /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-environment.yaml -e /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-bootstrap-environment-rhel.yaml -e /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-pacemaker-environment.yaml -r /usr/share/openstack-tripleo-heat-templates/deployed-server/deployed-server-roles-data.yaml -e /home/stack/templates/overcloud_images.yaml -e /home/stack/templates/ctlplane-assignments.yaml +++ The deployment completed: +++ (undercloud) [stack@undercloud-13 ~]$ heat stack-list WARNING (shell) "heat stack-list" is deprecated, please use "openstack stack list" instead +--------------------------------------+------------+-----------------+----------------------+----------------------+----------------------------------+ | id | stack_name | stack_status | creation_time | updated_time | project | +--------------------------------------+------------+-----------------+----------------------+----------------------+----------------------------------+ | 6bb3d4e3-fa04-421c-83e1-f0c957da2bb6 | overcloud | UPDATE_COMPLETE | 2018-11-01T19:41:45Z | 2018-11-02T13:58:59Z | 1375d2a7289c4e90be0a94c7315133c8 | +--------------------------------------+------------+-----------------+----------------------+----------------------+----------------------------------+ +++ There are other limitations in the documentation as well, for example, the documented example of deploy command doesn't include ctlplane-assignments.yaml: +++ $ source ~/stackrc (undercloud) $ openstack overcloud deploy \ [other arguments] \ --disable-validations \ -e /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-environment.yaml \ -e /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-bootstrap-environment-rhel.yaml \ -e /usr/share/openstack-tripleo-heat-templates/environments/deployed-server-pacemaker-environment.yaml \ -r /usr/share/openstack-tripleo-heat-templates/deployed-server/deployed-server-roles-data.yaml +++ Secondly the order in which the templates are supposed to be passed in incorrect, the documented example is confusing since it does not clarify that one should put their own custom templates at the end of the deployment command so that the director doesn't override them with the defaults. This should be made clear. Even though the deployment completed, the neutron-openvswitch agent on the compute node was not listed in the list of neutron agents being run on this environment: +++ (overcloud) [stack@undercloud-13 ~]$ neutron agent-list neutron CLI is deprecated and will be removed in the future. Use openstack CLI instead. +--------------------------------------+--------------------+------------------------+-------------------+-------+----------------+---------------------------+ | id | agent_type | host | availability_zone | alive | admin_state_up | binary | +--------------------------------------+--------------------+------------------------+-------------------+-------+----------------+---------------------------+ | 0625dbe3-6fde-4ba5-b98e-92317c3f43e0 | Metadata agent | controller.localdomain | | :-) | True | neutron-metadata-agent | | 3bc7bd5c-c91c-4065-9694-e15d03a7d0b8 | DHCP agent | controller.localdomain | nova | :-) | True | neutron-dhcp-agent | | 60ec3741-b219-483d-bc9f-248c526cae2d | L3 agent | controller.localdomain | nova | :-) | True | neutron-l3-agent | | d45ca831-6e5b-49e5-909f-35d8517c41e6 | Open vSwitch agent | controller.localdomain | | :-) | True | neutron-openvswitch-agent | +--------------------------------------+--------------------+------------------------+-------------------+-------+----------------+---------------------------+ +++ [1] https://access.redhat.com/documentation/en-us/red_hat_openstack_platform/13/html/director_installation_and_usage/chap-configuring_basic_overcloud_requirements_on_pre_provisioned_nodes Version-Release number of selected component (if applicable): +++ (undercloud) [stack@undercloud-13 ~]$ rpm -qa | grep -i tripleo python-tripleoclient-9.2.3-4.el7ost.noarch openstack-tripleo-image-elements-8.0.1-1.el7ost.noarch openstack-tripleo-common-containers-8.6.3-13.el7ost.noarch ansible-tripleo-ipsec-8.1.1-0.20180308133440.8f5369a.el7ost.noarch openstack-tripleo-puppet-elements-8.0.1-1.el7ost.noarch openstack-tripleo-common-8.6.3-13.el7ost.noarch puppet-tripleo-8.3.4-5.el7ost.noarch openstack-tripleo-heat-templates-8.0.4-20.el7ost.noarch +++ Actual results: Expected results: Additional info: sosreport from the compute node sosreport from the controller node templates used