Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.

Bug 1790469

Summary: There is latency spike(exceeds 2us~6us) in kvm-rt with all mitigations off
Product: Red Hat Enterprise Linux 7 Reporter: Pei Zhang <pezhang>
Component: kernel-rtAssignee: Nitesh Narayan Lal <nilal>
kernel-rt sub component: KVM QA Contact: Pei Zhang <pezhang>
Status: CLOSED WORKSFORME Docs Contact:
Severity: medium    
Priority: medium CC: bhu, chayang, daolivei, jinzhao, juzhang, lcapitulino, lgoncalv, nilal, peterx, trix, virt-maint, williams
Version: 7.8   
Target Milestone: rc   
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-05-19 13:31:04 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 1655690, 1672377    

Description Pei Zhang 2020-01-13 12:45:37 UTC
Description of problem:

After disabling all mitigations, the 24h cyclictest max latency exceeds 20us a bit, about 2us~6us.

Version-Release number of selected component (if applicable):
kernel-rt-3.10.0-1121.rt56.1085.el7.x86_64
rt-tests-1.5-9.el7.x86_64
qemu-kvm-rhev-2.12.0-41.el7.x86_64
libvirt-4.5.0-31.el7.x86_64
microcode_ctl-2.1-61.el7.x86_64
tuned-2.11.0-8.el7.noarch

How reproducible:
1/1

Steps to Reproduce:
1. Setup rt host

2. Boot rt guest

3. Disable all mitigations

# systool -vm kvm | grep nx_huge_pages/sys/kernel/debug/x86/
    nx_huge_pages_recovery_ratio= "0"
    nx_huge_pages       = "N"
 
# cat /sys/kernel/debug/x86/ibpb_enabled
1
# cat /sys/kernel/debug/x86/ibrs_enabled
0
# cat /sys/kernel/debug/x86/pti_enabled
0
# cat /sys/kernel/debug/x86/retp_enabled
0

4. Start 24h cyclictest in guest. Meanwhile compiling kernel in both host and guest housekeeping cores. 

# taskset -c 2,3,4,5,6,7,8,9 cyclictest -m -q -p95 -D 24h -h60 -t 8 -a 2,3,4,5,6,7,8,9 -i 200

5. Check max latency result, exceeds 20us a bit, it's 22us.

- Single VM with 8 rt vCPUs:
# Min Latencies: 00006 00008 00008 00008 00008 00008 00008 00008
# Avg Latencies: 00008 00008 00008 00008 00008 00008 00008 00008
# Max Latencies: 00019 00019 00019 00022 00019 00020 00019 00020


Actual results:
The max latency exceeds 20us.

Expected results:
We expect the max latency should not exceed 20us.

Additional info:
1. We also hit this little spike in rhel7.6.z and rhel7.7z. 

7.6.z kernel-rt-3.10.0-957.43.1.rt56.957.el7.x86_64(max latency is 24us):

- Single VM with 8 rt vCPUs:
# Min Latencies: 00005 00006 00006 00006 00005 00006 00005 00005
# Avg Latencies: 00007 00006 00007 00006 00006 00006 00006 00006
# Max Latencies: 00021 00018 00024 00018 00022 00018 00018 00018

7.7.z kernel-rt-3.10.0-1062.12.1.rt56.1039.el7.x86_64(max latency is 26us):
- Single VM with 8 rt vCPUs:
# Min Latencies: 00006 00009 00009 00008 00008 00008 00009 00009
# Avg Latencies: 00008 00009 00009 00009 00009 00009 00009 00009
# Max Latencies: 00025 00024 00023 00023 00023 00023 00022 00023

- Multiple VMs each with 1 rt vCPU:
- VM1
# Min Latencies: 00006
# Avg Latencies: 00008
# Max Latencies: 00023

- VM2
# Min Latencies: 00006
# Avg Latencies: 00008
# Max Latencies: 00023

- VM3
# Min Latencies: 00006
# Avg Latencies: 00008
# Max Latencies: 00023

- VM4
# Min Latencies: 00007
# Avg Latencies: 00008
# Max Latencies: 00026

Comment 3 Luiz Capitulino 2020-01-13 21:52:54 UTC
Pei,

Since this is not a big spike, I'm not considering it urgent for now.

You mentioned that you were planning to test older kernels to see if
this is a regression. However, Marcelo thinks this could a dupe of bug 1772082.

So, instead of debugging this issue, what about running the measurements
against RHEL7.8. If you get the same results, then we could ask Nitesh
to build a 7.8 kernel with the fix for bug 1772082. If it is indeed a dupe,
we may want to fix bug 1772082 for 7.8.

Comment 5 Nitesh Narayan Lal 2020-05-06 12:38:16 UTC
Recently, Jean from the QE team who was running PvP testing has reported a max latency of >22us on a 7.8 rt kernel with all mitigations disabled.
One such +3us spike that I was able to reproduce on his setup is provided below:

Test started at Thu Apr 30 13:16:02 EDT 2020

Test duration:    1h
Run rteval:       n
Run stress:       y
Isolated CPUs:    1,2,3,4
Kernel:           3.10.0-1127.rt56.1093.el7.x86_64
Kernel cmd-line:  BOOT_IMAGE=/vmlinuz-3.10.0-1127.rt56.1093.el7.x86_64 root=UUID=ee7ba82b-2889-424d-aad5-446464c910ce ro rhgb quiet crashkernel=auto spectre_v2=retpoline console=ttyS0,115200 skew_tick=1 isolcpus=1-4 intel_pstate=disable nosoftlockup nohz=on nohz_full=1-4 rcu_nocbs=1-4
x86 debug opts:   retp_enabled=0 pti_enabled=0 ibrs_enabled=0 ibpb_enabled=1
Machine:          localhost.localdomain
CPU:              Intel(R) Xeon(R) Gold 6132 CPU @ 2.60GHz
Results dir:      /root/results/cyclictest-results.whvPXJ

running stress
   taskset -c 1 stress --cpu 1
   taskset -c 2 stress --cpu 1
   taskset -c 3 stress --cpu 1
   taskset -c 4 stress --cpu 1

starting Thu Apr 30 13:16:03 EDT 2020
   taskset -c 1,2,3,4 cyclictest -m -q -p95 -D 1h -h60 -i 200 -t 4 -a 1,2,3,4
ended Thu Apr 30 14:16:03 EDT 2020

output dir is /root/results/cyclictest-results.whvPXJ


# Min Latencies: 00006 00008 00008 00008
# Avg Latencies: 00010 00010 00010 00010
# Max Latencies: 00019 00014 00024 00016

Hence I am keeping this bug open for now, for further debugging (Although a few us spike is minor and relatively harder to debug).

Comment 6 Nitesh Narayan Lal 2020-05-12 13:37:01 UTC
Pei,

Can you please confirm if your housekeeping and RT vCPUs were pinned to the same NUMA node when you were triggering this spike?

Comment 7 Pei Zhang 2020-05-13 04:06:40 UTC
(In reply to Nitesh Narayan Lal from comment #6)
> Pei,
> 
> Can you please confirm if your housekeeping and RT vCPUs were pinned to the
> same NUMA node when you were triggering this spike?

Hi Nitesh,


The housekeeping vCPUs and RT vCPUs are pinned to different NUMA nodes. Please see below.

<vcpu placement='static'>10</vcpu>
  <cputune>
    <vcpupin vcpu='0' cpuset='16'/>
    <vcpupin vcpu='1' cpuset='18'/>
    <vcpupin vcpu='2' cpuset='1'/>
    <vcpupin vcpu='3' cpuset='3'/>
    <vcpupin vcpu='4' cpuset='5'/>
    <vcpupin vcpu='5' cpuset='7'/>
    <vcpupin vcpu='6' cpuset='9'/>
    <vcpupin vcpu='7' cpuset='11'/>
    <vcpupin vcpu='8' cpuset='13'/>
    <vcpupin vcpu='9' cpuset='15'/>
    <emulatorpin cpuset='2,4,6,8,10'/>
    <vcpusched vcpus='0' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='1' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='2' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='3' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='4' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='5' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='6' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='7' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='8' scheduler='fifo' priority='1'/>
    <vcpusched vcpus='9' scheduler='fifo' priority='1'/>
  </cputune>


Host NUMA topology:
# lscpu | grep NUMA
NUMA node(s):        2
NUMA node0 CPU(s):   0,2,4,6,8,10,12,14,16,18
NUMA node1 CPU(s):   1,3,5,7,9,11,13,15,17,19

Host cores isolation:
# cat /proc/cmdline 
BOOT_IMAGE=(hd0,msdos1)/vmlinuz... console=ttyS0,115200n81 skew_tick=1 isolcpus=managed_irq,domain,1,3,5,7,9,11,13,15,17,19,12,14,16,18 intel_pstate=disable nosoftlockup nohz=on nohz_full=1,3,5,7,9,11,13,15,17,19,12,14,16,18 rcu_nocbs=1,3,5,7,9,11,13,15,17,19,12,14,16,18 default_hugepagesz=1G ...



Best reagards,

Pei

Comment 8 Pei Zhang 2020-05-13 04:16:58 UTC
(In reply to Pei Zhang from comment #7)
> (In reply to Nitesh Narayan Lal from comment #6)
> > Pei,
> > 
> > Can you please confirm if your housekeeping and RT vCPUs were pinned to the
> > same NUMA node when you were triggering this spike?
> 
> Hi Nitesh,
> 
> 
> The housekeeping vCPUs and RT vCPUs are pinned to different NUMA nodes.
> Please see below.
> 
> <vcpu placement='static'>10</vcpu>
>   <cputune>
>     <vcpupin vcpu='0' cpuset='16'/>
>     <vcpupin vcpu='1' cpuset='18'/>
>     <vcpupin vcpu='2' cpuset='1'/>
>     <vcpupin vcpu='3' cpuset='3'/>
>     <vcpupin vcpu='4' cpuset='5'/>
>     <vcpupin vcpu='5' cpuset='7'/>
>     <vcpupin vcpu='6' cpuset='9'/>
>     <vcpupin vcpu='7' cpuset='11'/>
>     <vcpupin vcpu='8' cpuset='13'/>
>     <vcpupin vcpu='9' cpuset='15'/>
>     <emulatorpin cpuset='2,4,6,8,10'/>
>     <vcpusched vcpus='0' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='1' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='2' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='3' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='4' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='5' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='6' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='7' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='8' scheduler='fifo' priority='1'/>
>     <vcpusched vcpus='9' scheduler='fifo' priority='1'/>
>   </cputune>

vCPUs 0,1 are guest housekeeping vCPUs.
vCPUs 2~9 are guest RT vCPUs.

Guest cores isolation:
# cat /proc/cmdline 
BOOT_IMAGE=/vmlinuz... default_hugepagesz=1G iommu=pt intel_iommu=on skew_tick=1 isolcpus=2,3,4,5,6,7,8,9 intel_pstate=disable nosoftlockup nohz=on nohz_full=2,3,4,5,6,7,8,9 rcu_nocbs=2,3,4,5,6,7,8,9 spectre_v2=off nopti kvm-intel.vmentry_l1d_flush=never

Comment 9 Nitesh Narayan Lal 2020-05-13 13:45:53 UTC
(In reply to Pei Zhang from comment #8)
> (In reply to Pei Zhang from comment #7)
> > (In reply to Nitesh Narayan Lal from comment #6)
> > > Pei,
> > > 
> > > Can you please confirm if your housekeeping and RT vCPUs were pinned to the
> > > same NUMA node when you were triggering this spike?
> > 
> > Hi Nitesh,
> > 
> > 
> > The housekeeping vCPUs and RT vCPUs are pinned to different NUMA nodes.
> > Please see below.
> > 
> > <vcpu placement='static'>10</vcpu>
> >   <cputune>
> >     <vcpupin vcpu='0' cpuset='16'/>
> >     <vcpupin vcpu='1' cpuset='18'/>
> >     <vcpupin vcpu='2' cpuset='1'/>
> >     <vcpupin vcpu='3' cpuset='3'/>
> >     <vcpupin vcpu='4' cpuset='5'/>
> >     <vcpupin vcpu='5' cpuset='7'/>
> >     <vcpupin vcpu='6' cpuset='9'/>
> >     <vcpupin vcpu='7' cpuset='11'/>
> >     <vcpupin vcpu='8' cpuset='13'/>
> >     <vcpupin vcpu='9' cpuset='15'/>
> >     <emulatorpin cpuset='2,4,6,8,10'/>
> >     <vcpusched vcpus='0' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='1' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='2' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='3' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='4' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='5' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='6' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='7' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='8' scheduler='fifo' priority='1'/>
> >     <vcpusched vcpus='9' scheduler='fifo' priority='1'/>
> >   </cputune>
> 
> vCPUs 0,1 are guest housekeeping vCPUs.
> vCPUs 2~9 are guest RT vCPUs.
> 
> Guest cores isolation:
> # cat /proc/cmdline 
> BOOT_IMAGE=/vmlinuz... default_hugepagesz=1G iommu=pt intel_iommu=on
> skew_tick=1 isolcpus=2,3,4,5,6,7,8,9 intel_pstate=disable nosoftlockup
> nohz=on nohz_full=2,3,4,5,6,7,8,9 rcu_nocbs=2,3,4,5,6,7,8,9 spectre_v2=off
> nopti kvm-intel.vmentry_l1d_flush=never


Thank you for sharing the information.
Also, if I am not mistaken then you are not able to reproduce the reported spike with the latest 7.8.z rt-kernel, is that correct?

Comment 11 Nitesh Narayan Lal 2020-05-19 13:31:04 UTC
I have not been able to reproduce the latency spike in several of my last test runs with the latest 7.8.z RT kernel and Pei has also reported that she is not able to reproduce the spike anymore, hence I am closing this Bug.

Latency spike in comment5:
This spike is probably because of pinning housekeeping and RT vCPUs on the same NUMA node. Unfortunately, I don't have access to that system anymore for verification. However, I did manage to create a similar spike in one of my testbeds by following this non-recommended pinning and the latency spike was gone as soon as I switched to the recommended pinning. As per KVM-RT guidelines, it is always recommended to pin housekeeping and RT vCPU on different NUMA nodes, this is to avoid cache thrashing issues. Although, the impact of pinning RT: Housekeeping vCPUs on the same NUMA is hugely dependent upon the workload that is running on the housekeeping vCPUs.

Latency spike originally reported:
The latency spike was not consistently reproduced to gather enough tracing data to make a definitive conclusion. However, the occasional spikes could be due to the missing timer advancement support in RHEL7 codebase, which sometimes leads to timer interrupts being injected before the actual guest timer expiration and hence resulting in the generation of another timer reprogramming request. 
Note: The latency spike due to this is hard to trigger.


Following is a test run on 4VM running cyclictest simultatensouly with stress for 8h:
Environment:
Test duration:          8h
Run rteval:             n
Run stress:             y
Isolated CPUs:          1
Nr Housekeeping CPUs:   1
Kernel:                 3.10.0-1127.8.2.rt56.1103.el7.x86_64
Kernel cmd-line:        BOOT_IMAGE=/vmlinuz-3.10.0-1127.8.2.rt56.1103.el7.x86_64 root=/dev/mapper/rhel-root ro crashkernel=auto rd.lvm.lv=rhel/root rd.lvm.lv=rhel/swap rhgb quiet console=ttyS0 LANG=en_US.UTF-8 skew_tick=1 isolcpus=1 intel_pstate=disable nosoftlockup nohz=on nohz_full=1 rcu_nocbs=1
x86 debug opts:         retp_enabled=0 pti_enabled=0 ibrs_enabled=0 ibpb_enabled=1
Machine:                test-vm1
CPU:                    Intel(R) Xeon(R) CPU E5-2640 v3 @ 2.60GHz


Results:
VM1:
# Min Latencies: 00011
# Avg Latencies: 00014
# Max Latencies: 00022
VM2:
# Min Latencies: 00011
# Avg Latencies: 00014
# Max Latencies: 00022
VM3:
# Min Latencies: 00011
# Avg Latencies: 00014
# Max Latencies: 00022
VM4:
# Min Latencies: 00011
# Avg Latencies: 00014
# Max Latencies: 00022