Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1883741

Summary: [OSP13] Unable to start service systemd-modules-load.service during fast forward upgrade on controller nodes
Product: Red Hat OpenStack Reporter: Vadim Khitrin <vkhitrin>
Component: tripleo-ansibleAssignee: RHOS Maint <rhos-maint>
Status: CLOSED DUPLICATE QA Contact: Sasha Smolyak <ssmolyak>
Severity: high Docs Contact:
Priority: unspecified    
Version: 13.0 (Queens)CC: jpretori
Target Milestone: ---   
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-09-30 09:53:05 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Vadim Khitrin 2020-09-30 05:38:52 UTC
Description of problem:
When attempting to perform fast forward upgrade from OSP13z13 (2020-09-16.1) to OSP16.1.2 (RHOS-16.1-RHEL-8-20200925.n.1), there appears to be an issue when attempting to upgrade overcloud controller nodes.

Undercloud upgrade and overcloud preparation to upgrade have been executed successfully.

The failing Ansible task for controller node:
2020-09-29 17:40:33 | TASK [tripleo-kernel : Modules reload] *****************************************
2020-09-29 17:40:33 | Tuesday 29 September 2020  17:40:30 -0400 (0:00:01.830)       0:03:54.764 *****
2020-09-29 17:40:33 | fatal: [controller-0]: FAILED! => {"changed": false, "msg": "Unable to start service systemd-modules-load.service: Job for systemd-modules-load.service failed because the control process exited with error code.\nSee \"systemctl status systemd-modules-load.service\" and \"journalctl -xe\" for details.\n"}

While looking on the controller node, it was successfully updated to 8.2 from 7.9 using leapp, and looking at the failing daemon:
journalctl -u systemd-modules-load.service
-- Logs begin at Tue 2020-09-29 12:55:08 UTC, end at Wed 2020-09-30 05:33:39 UTC. --
Sep 29 21:07:03 controller-0 systemd[1]: Starting Load Kernel Modules...
Sep 29 21:07:03 controller-0 systemd-modules-load[693]: Failed to find module 'nf_conntrack_proto_sctp'
Sep 29 21:07:03 controller-0 systemd[1]: systemd-modules-load.service: Main process exited, code=exited, status=1/FAILURE
Sep 29 21:07:03 controller-0 systemd[1]: systemd-modules-load.service: Failed with result 'exit-code'.
Sep 29 21:07:03 controller-0 systemd[1]: Failed to start Load Kernel Modules.
-- Reboot --
Sep 29 21:12:19 controller-0 systemd[1]: Starting Load Kernel Modules...
Sep 29 21:12:19 controller-0 systemd-modules-load[694]: Failed to find module 'nf_conntrack_proto_sctp'
Sep 29 21:12:19 controller-0 systemd[1]: systemd-modules-load.service: Main process exited, code=exited, status=1/FAILURE
Sep 29 21:12:19 controller-0 systemd[1]: systemd-modules-load.service: Failed with result 'exit-code'.
Sep 29 21:12:19 controller-0 systemd[1]: Failed to start Load Kernel Modules.
Sep 29 21:40:31 controller-0 systemd[1]: Starting Load Kernel Modules...
Sep 29 21:40:31 controller-0 systemd[1]: systemd-modules-load.service: Main process exited, code=exited, status=1/FAILURE
Sep 29 21:40:31 controller-0 systemd[1]: systemd-modules-load.service: Failed with result 'exit-code'.
Sep 29 21:40:31 controller-0 systemd[1]: Failed to start Load Kernel Modules.

It apepars that there is a file '/etc/modules-load.d/nf_conntrack_proto_sctp.conf' which attempts to load 'nf_conntrack_proto_sctp' kernel module which is not present in RHEL 8.2 (and also not present in 7.9), this causes the service fail to start.

Unsure how '/etc/modules-load.d/' is populated since it is empty in overcloucd images.

During OSP13 fresh install service 'systemd-modules-load.service' is not started and is skipped due to a condition in systemd unit file:
● systemd-modules-load.service - Load Kernel Modules
   Loaded: loaded (/usr/lib/systemd/system/systemd-modules-load.service; static; vendor preset: disabled)
   Active: inactive (dead)
Condition: start condition failed at Sun 2020-09-27 18:51:29 UTC; 1 day 1h ago
     Docs: man:systemd-modules-load.service(8)
           man:modules-load.d(5)

In previous attempt, when re-running the overcloud upgrade command for that node, on next execution the playbook was executed successfully.

I did not encounter this when performing FFU from OSP13z12 to OSP16.1.1


How reproducible:
Tried twice on same puddles, encountered this issue every time.


Steps to Reproduce:
1. Deploy fresh OSP13z13
2. Attempt to perform fast forward upgrade

Actual results:
Fast forward upgrade fails


Expected results:
Fast forward upgrade succeeds

Additional info:
Will upload sosreport and logs in comment.

Comment 1 Vadim Khitrin 2020-09-30 05:39:40 UTC
Sorry, was not sure to which component to file this BZ for, so I opted for 'tripleo-ansible'. Please let me know if this is the wrong component.

Comment 2 Jesse Pretorius 2020-09-30 09:53:05 UTC

*** This bug has been marked as a duplicate of bug 1880979 ***