Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.

Bug 1962690

Summary: stalld service crash on openshift 4.8
Product: Red Hat Enterprise Linux 8 Reporter: Sebastian Scheinkman <sscheink>
Component: stalldAssignee: Fernando Pacheco <fpacheco>
Status: CLOSED CURRENTRELEASE QA Contact: Mark Simmons <msimmons>
Severity: unspecified Docs Contact:
Priority: high    
Version: 8.4CC: bhu, fpacheco, jlelli, kcarcia, keyoung, williams
Target Milestone: betaKeywords: Triaged
Target Release: ---Flags: pm-rhel: mirror+
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-07-30 00:43:25 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Sebastian Scheinkman 2021-05-20 13:41:10 UTC
Description of problem:
stalld is not able to run on ocp 4.8, the systemctl service failed to start


Version-Release number of selected component (if applicable):
oc version
Client Version: 4.8.0-0.nightly-2021-05-18-205323
Server Version: 4.8.0-0.nightly-2021-05-18-205323
Kubernetes Version: v1.21.0-rc.0+9d99e1c

cnfdc8.clus2.t5g.lab.eng.bos.redhat.com          Ready    worker,worker-cnf   19h   v1.21.0-rc.0+9d99e1c   10.19.16.57    <none>        Red Hat Enterprise Linux CoreOS 48.84.202105180118-0 (Ootpa)   4.18.0-293.rt7.59.el8.x86_64   cri-o://1.21.0-93.rhaos4.8.git8f8bcd9.el8

Steps to Reproduce:
install performance addon operator and apply a policy

apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
  annotations:
    kubectl.kubernetes.io/last-applied-configuration: |
      {"apiVersion":"performance.openshift.io/v2","kind":"PerformanceProfile","metadata":{"annotations":{},"name":"performance"},"spec":{"additionalKernelArgs":["nmi_watchdog=0","audit=0","mce=off","processor.max_cstate=1","idle=poll","intel_idle.max_cstate=0","nosmt","tsc=reliable"],"cpu":{"isolated":"1,3,5,7,9,11-51","reserved":"0,2,4,6,8,10"},"hugepages":{"defaultHugepagesSize":"1G","pages":[{"count":10,"size":"1G"}]},"nodeSelector":{"node-role.kubernetes.io/worker-cnf":""},"numa":{"topologyPolicy":"best-effort"},"realTimeKernel":{"enabled":true}}}
  creationTimestamp: "2021-05-19T16:55:51Z"
  finalizers:
  - foreground-deletion
  generation: 1
  name: performance
  resourceVersion: "131670"
  uid: eb4086c3-7b67-4d13-b84a-5fd71cc60821
spec:
  additionalKernelArgs:
  - nmi_watchdog=0
  - audit=0
  - mce=off
  - processor.max_cstate=1
  - idle=poll
  - intel_idle.max_cstate=0
  - nosmt
  - tsc=reliable
  cpu:
    isolated: 1,3,5,7,9,11-51
    reserved: 0,2,4,6,8,10
  hugepages:
    defaultHugepagesSize: 1G
    pages:
    - count: 10
      size: 1G
  nodeSelector:
    node-role.kubernetes.io/worker-cnf: ""
  numa:
    topologyPolicy: best-effort
  realTimeKernel:
    enabled: true


Logs:
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778940]: Disabled RT throttling
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778942]: boosted pid 0 using SCHED_DEADLINE
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778942]: using SCHED_DEADLINE for boosting
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778942]: sched_debug is getting larger, increasing the buffer to 204800
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778942]:   detect_task_format: sched_debug size greater than 102400 bytes didn't get full read
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778942]: unable to find 'runnable tasks' in buffer, invalid input
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com systemd[1]: stalld.service: Control process exited, code=exited status=255
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com stalld[778948]: Restored RT throttling
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com systemd[1]: stalld.service: Failed with result 'exit-code'.
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com systemd[1]: Failed to start Stall Monitor.
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com systemd[1]: stalld.service: Consumed 19ms CPU time
May 19 16:56:00 cnfdc8.clus2.t5g.lab.eng.bos.redhat.com systemd[1]: Reloading.

Comment 1 Fernando Pacheco 2021-05-20 16:52:13 UTC
Looks like the buffer used to hold sched_debug output isn't being re-sized.
The error happens when we look for a specific marker in the output and well
we don't have all of the output..

This should be fixed in newer versions of stalld where the buffer is dynamically resized
to accommodate the larger output.

commit dc2dc32e: "stalld.c: rework detect_task_format and buffer_size logic"

Comment 2 Sebastian Scheinkman 2021-05-20 16:55:13 UTC
Hi Fernando,

thanks for the replay!

Do you know if this fix is already in rhel 8.2 and next in rhcos 4.6/8.2?

Comment 3 Clark Williams 2021-05-20 18:52:56 UTC
(In reply to Sebastian Scheinkman from comment #2)
> Hi Fernando,
> 
> thanks for the replay!
> 
> Do you know if this fix is already in rhel 8.2 and next in rhcos 4.6/8.2?

We didn't release stalld until 8.4. I thought that OCP did their own container of stalld for the 4.6 releases?

Comment 4 Sebastian Scheinkman 2021-05-24 12:28:19 UTC
Hi Clark,

you are right it's part of the tuned operator and the version there for stalld is v1.9.0 (https://github.com/openshift/cluster-node-tuning-operator/tree/release-4.8/assets/tuned/stalld) is that version contains the fix or we should move this bug to the tuned operator team?

Thanks!
Sebastian

Comment 5 Fernando Pacheco 2021-05-25 18:24:17 UTC
(In reply to Sebastian Scheinkman from comment #4)
> Hi Clark,
> 
> you are right it's part of the tuned operator and the version there for
> stalld is v1.9.0
> (https://github.com/openshift/cluster-node-tuning-operator/tree/release-4.8/
> assets/tuned/stalld) is that version contains the fix or we should move this
> bug to the tuned operator team?
> 
> Thanks!
> Sebastian

Hi Sebastian,

The commit mentioned is there for stalld v1.9.0.

What version of stalld was used in the run that generated the logs in the description?

AFAICT, this log line was removed in v1.7.0:
"detect_task_format: sched_debug size greater than 102400 bytes didn't get full read"

Thanks,
Fernando

Comment 6 Sebastian Scheinkman 2021-05-26 09:27:02 UTC
Hi Fernando,

I don't know what was the version it was on our CI system that deploy every night.

I will wait for another CI running and check again maybe the tuned was not updated on that nightly.

Thanks for the help!
Sebastian

Comment 17 Red Hat Bugzilla 2023-09-15 01:06:57 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days