Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 2054766

Summary: [OSP16.2] SAR doesn't work post FFU
Product: Red Hat OpenStack Reporter: ggrimaux
Component: openstack-tripleo-heat-templatesAssignee: Lukas Bezdicka <lbezdick>
Status: CLOSED DUPLICATE QA Contact: Joe H. Rahme <jhakimra>
Severity: medium Docs Contact:
Priority: medium    
Version: 16.2 (Train)CC: bshephar, bwelterl, cbesson, jpretori, lbezdick, mburns, ramishra
Target Milestone: z4Keywords: Triaged
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2024-01-22 16:23:55 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description ggrimaux 2022-02-15 16:37:48 UTC
Description of problem:
Client did FFU from 16.1.5 then minor to 16.2.1

While debugging an issue on the controller nodes I noticed something in the /var/log/sa/ dir:
[root@controller1 sa]# pwd
/var/log/sa
[root@controller1 sa]# ls -ltR
.:
total 725804
-rw-r--r--. 1 root root     1044 Jan 25 20:01 sa25
-rw-r--r--. 1 root root     1044 Jan 19 10:57 sa19
-rw-r--r--. 1 root root     1044 Jan 13 07:37 sa13
-rw-r--r--. 1 root root     1044 Dec 27 12:58 sa27
-rw-r--r--. 1 root root     1072 Dec 23 20:47 sa23
-rw-r--r--. 1 root root     1044 Dec  8 08:35 sa08
-rw-r--r--. 1 root root     1044 Oct 14 18:07 sa14
-rw-r--r--. 1 root root     1044 Oct 11 11:06 sa11
-rw-r--r--. 1 root root     1044 Sep 21 14:55 sa21
-rw-r--r--. 1 root root     1072 Jun 29  2021 sa29
-rw-r--r--. 1 root root     1044 Jun 12  2021 sa12
-rw-r--r--. 1 root root 13204293 Jun 11  2021 sar11
-rw-r--r--. 1 root root 18012128 Jun 10  2021 sa10
-rw-r--r--. 1 root root 13204293 Jun 10  2021 sar10

FFU from OSP13 was started on June 12.

The issue is that with RHEL8, sar is not managed through cron job anymore but from systemd timer.
But here the timers are disabled:
sos_commands/systemd/systemctl_list-unit-files:sysstat.service                                                  enabled        
sos_commands/systemd/systemctl_list-unit-files:sysstat-collect.timer                                            disabled       
sos_commands/systemd/systemctl_list-unit-files:sysstat-summary.timer                                            disabled 

Why the files above have different dates after June is that each time you restart a server it runs for one time.

I had the client delete all files and restart sysstart.service and this is what is seen after a few days:
[root@controller1 ~]# ls -lhZ /var/log/sa/
total 4.0K
-rw-r--r--. 1 root root system_u:object_r:sysstat_log_t:s0 1.1K Feb 11 10:42 sa11
[root@controller1 ~]#

I spoke to someone from sbr-services and it is a bug in the post deployment of the package.

I'll let him explain below the things he found and what is needed.

So not sure if this will be inside THT.

Thank you.



Version-Release number of selected component (if applicable):
OSP16.2.1

How reproducible:
Probably 100% (noticed only now)

Steps to Reproduce:
1.RHEL7 + sysstat installed (SAR)
2.FFU RHEL8
3.

Actual results:
Files are not being written anymore in /var/log/sa/

Expected results:
Files being written in /var/log/sa/

Additional info:
I have sosreport from a node.

Comment 1 Welterlen Benoit 2022-02-15 16:44:33 UTC
Hello,

The issue is that post install script are not executed (expected for a normal upgrade):
----
if [ $1 -eq 1 ] ; then 
        # Initial installation 
        systemctl --no-reload preset sysstat.service sysstat-collect.timer sysstat-summary.timer &>/dev/null || : 
fi
----

But it is expected that this has been executed one time, from a previous rpm install. But the previous rpm was a RHEL7 one, with this:

---
if [ $1 -eq 1 ] ; then 
        # Initial installation 
        systemctl preset sysstat.service >/dev/null 2>&1 || : 
fi
---

=> thus only sysstat service was managed in RHEL7 (of course, no timer, because based on crond).

Then we have only sysstat enabled, without required timer services to collect data...

Don't really know if this is really a sysstat rpm issue or should be managed by leapp.

Thank you !

Comment 2 Brendan Shephard 2022-02-21 02:10:54 UTC
Hmm, in my opinion and since this is a package provided by RHEL. I think this should be handled during the Leapp upgrade. If packages need to be re-installed to re-run the install scripts, I think that should happen during Leapp, rather than something that we implement a hack for in tripleo.

I don't think this problem would be unique to OpenStack deployments, we should probably be seeing it in all environments that have been upgraded from 7 > 8 using Leapp.

Comment 3 Christophe Besson 2022-02-21 08:20:53 UTC
Note that a dedicated BZ for leapp has been created:
https://bugzilla.redhat.com/show_bug.cgi?id=2055117

Comment 4 ggrimaux 2022-02-22 15:00:56 UTC
Quick workaround is to enable the two systemd timer:
systemctl enable --now sysstat-collect.timer
systemctl enable --now sysstat-summary.timer

Comment 5 Lukas Bezdicka 2024-01-22 16:23:55 UTC

*** This bug has been marked as a duplicate of bug 2055117 ***