Bug 1883741
| Summary: | [OSP13] Unable to start service systemd-modules-load.service during fast forward upgrade on controller nodes | ||
|---|---|---|---|
| Product: | Red Hat OpenStack | Reporter: | Vadim Khitrin <vkhitrin> |
| Component: | tripleo-ansible | Assignee: | RHOS Maint <rhos-maint> |
| Status: | CLOSED DUPLICATE | QA Contact: | Sasha Smolyak <ssmolyak> |
| Severity: | high | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 13.0 (Queens) | CC: | jpretori |
| Target Milestone: | --- | ||
| Target Release: | --- | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2020-09-30 09:53:05 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
Sorry, was not sure to which component to file this BZ for, so I opted for 'tripleo-ansible'. Please let me know if this is the wrong component. *** This bug has been marked as a duplicate of bug 1880979 *** |
Description of problem: When attempting to perform fast forward upgrade from OSP13z13 (2020-09-16.1) to OSP16.1.2 (RHOS-16.1-RHEL-8-20200925.n.1), there appears to be an issue when attempting to upgrade overcloud controller nodes. Undercloud upgrade and overcloud preparation to upgrade have been executed successfully. The failing Ansible task for controller node: 2020-09-29 17:40:33 | TASK [tripleo-kernel : Modules reload] ***************************************** 2020-09-29 17:40:33 | Tuesday 29 September 2020 17:40:30 -0400 (0:00:01.830) 0:03:54.764 ***** 2020-09-29 17:40:33 | fatal: [controller-0]: FAILED! => {"changed": false, "msg": "Unable to start service systemd-modules-load.service: Job for systemd-modules-load.service failed because the control process exited with error code.\nSee \"systemctl status systemd-modules-load.service\" and \"journalctl -xe\" for details.\n"} While looking on the controller node, it was successfully updated to 8.2 from 7.9 using leapp, and looking at the failing daemon: journalctl -u systemd-modules-load.service -- Logs begin at Tue 2020-09-29 12:55:08 UTC, end at Wed 2020-09-30 05:33:39 UTC. -- Sep 29 21:07:03 controller-0 systemd[1]: Starting Load Kernel Modules... Sep 29 21:07:03 controller-0 systemd-modules-load[693]: Failed to find module 'nf_conntrack_proto_sctp' Sep 29 21:07:03 controller-0 systemd[1]: systemd-modules-load.service: Main process exited, code=exited, status=1/FAILURE Sep 29 21:07:03 controller-0 systemd[1]: systemd-modules-load.service: Failed with result 'exit-code'. Sep 29 21:07:03 controller-0 systemd[1]: Failed to start Load Kernel Modules. -- Reboot -- Sep 29 21:12:19 controller-0 systemd[1]: Starting Load Kernel Modules... Sep 29 21:12:19 controller-0 systemd-modules-load[694]: Failed to find module 'nf_conntrack_proto_sctp' Sep 29 21:12:19 controller-0 systemd[1]: systemd-modules-load.service: Main process exited, code=exited, status=1/FAILURE Sep 29 21:12:19 controller-0 systemd[1]: systemd-modules-load.service: Failed with result 'exit-code'. Sep 29 21:12:19 controller-0 systemd[1]: Failed to start Load Kernel Modules. Sep 29 21:40:31 controller-0 systemd[1]: Starting Load Kernel Modules... Sep 29 21:40:31 controller-0 systemd[1]: systemd-modules-load.service: Main process exited, code=exited, status=1/FAILURE Sep 29 21:40:31 controller-0 systemd[1]: systemd-modules-load.service: Failed with result 'exit-code'. Sep 29 21:40:31 controller-0 systemd[1]: Failed to start Load Kernel Modules. It apepars that there is a file '/etc/modules-load.d/nf_conntrack_proto_sctp.conf' which attempts to load 'nf_conntrack_proto_sctp' kernel module which is not present in RHEL 8.2 (and also not present in 7.9), this causes the service fail to start. Unsure how '/etc/modules-load.d/' is populated since it is empty in overcloucd images. During OSP13 fresh install service 'systemd-modules-load.service' is not started and is skipped due to a condition in systemd unit file: ● systemd-modules-load.service - Load Kernel Modules Loaded: loaded (/usr/lib/systemd/system/systemd-modules-load.service; static; vendor preset: disabled) Active: inactive (dead) Condition: start condition failed at Sun 2020-09-27 18:51:29 UTC; 1 day 1h ago Docs: man:systemd-modules-load.service(8) man:modules-load.d(5) In previous attempt, when re-running the overcloud upgrade command for that node, on next execution the playbook was executed successfully. I did not encounter this when performing FFU from OSP13z12 to OSP16.1.1 How reproducible: Tried twice on same puddles, encountered this issue every time. Steps to Reproduce: 1. Deploy fresh OSP13z13 2. Attempt to perform fast forward upgrade Actual results: Fast forward upgrade fails Expected results: Fast forward upgrade succeeds Additional info: Will upload sosreport and logs in comment.