Bug 1790469
| Summary: | There is latency spike(exceeds 2us~6us) in kvm-rt with all mitigations off | ||
|---|---|---|---|
| Product: | Red Hat Enterprise Linux 7 | Reporter: | Pei Zhang <pezhang> |
| Component: | kernel-rt | Assignee: | Nitesh Narayan Lal <nilal> |
| kernel-rt sub component: | KVM | QA Contact: | Pei Zhang <pezhang> |
| Status: | CLOSED WORKSFORME | Docs Contact: | |
| Severity: | medium | ||
| Priority: | medium | CC: | bhu, chayang, daolivei, jinzhao, juzhang, lcapitulino, lgoncalv, nilal, peterx, trix, virt-maint, williams |
| Version: | 7.8 | ||
| Target Milestone: | rc | ||
| Target Release: | --- | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2020-05-19 13:31:04 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | |||
| Bug Blocks: | 1655690, 1672377 | ||
Pei, Since this is not a big spike, I'm not considering it urgent for now. You mentioned that you were planning to test older kernels to see if this is a regression. However, Marcelo thinks this could a dupe of bug 1772082. So, instead of debugging this issue, what about running the measurements against RHEL7.8. If you get the same results, then we could ask Nitesh to build a 7.8 kernel with the fix for bug 1772082. If it is indeed a dupe, we may want to fix bug 1772082 for 7.8. Recently, Jean from the QE team who was running PvP testing has reported a max latency of >22us on a 7.8 rt kernel with all mitigations disabled. One such +3us spike that I was able to reproduce on his setup is provided below: Test started at Thu Apr 30 13:16:02 EDT 2020 Test duration: 1h Run rteval: n Run stress: y Isolated CPUs: 1,2,3,4 Kernel: 3.10.0-1127.rt56.1093.el7.x86_64 Kernel cmd-line: BOOT_IMAGE=/vmlinuz-3.10.0-1127.rt56.1093.el7.x86_64 root=UUID=ee7ba82b-2889-424d-aad5-446464c910ce ro rhgb quiet crashkernel=auto spectre_v2=retpoline console=ttyS0,115200 skew_tick=1 isolcpus=1-4 intel_pstate=disable nosoftlockup nohz=on nohz_full=1-4 rcu_nocbs=1-4 x86 debug opts: retp_enabled=0 pti_enabled=0 ibrs_enabled=0 ibpb_enabled=1 Machine: localhost.localdomain CPU: Intel(R) Xeon(R) Gold 6132 CPU @ 2.60GHz Results dir: /root/results/cyclictest-results.whvPXJ running stress taskset -c 1 stress --cpu 1 taskset -c 2 stress --cpu 1 taskset -c 3 stress --cpu 1 taskset -c 4 stress --cpu 1 starting Thu Apr 30 13:16:03 EDT 2020 taskset -c 1,2,3,4 cyclictest -m -q -p95 -D 1h -h60 -i 200 -t 4 -a 1,2,3,4 ended Thu Apr 30 14:16:03 EDT 2020 output dir is /root/results/cyclictest-results.whvPXJ # Min Latencies: 00006 00008 00008 00008 # Avg Latencies: 00010 00010 00010 00010 # Max Latencies: 00019 00014 00024 00016 Hence I am keeping this bug open for now, for further debugging (Although a few us spike is minor and relatively harder to debug). Pei, Can you please confirm if your housekeeping and RT vCPUs were pinned to the same NUMA node when you were triggering this spike? (In reply to Nitesh Narayan Lal from comment #6) > Pei, > > Can you please confirm if your housekeeping and RT vCPUs were pinned to the > same NUMA node when you were triggering this spike? Hi Nitesh, The housekeeping vCPUs and RT vCPUs are pinned to different NUMA nodes. Please see below. <vcpu placement='static'>10</vcpu> <cputune> <vcpupin vcpu='0' cpuset='16'/> <vcpupin vcpu='1' cpuset='18'/> <vcpupin vcpu='2' cpuset='1'/> <vcpupin vcpu='3' cpuset='3'/> <vcpupin vcpu='4' cpuset='5'/> <vcpupin vcpu='5' cpuset='7'/> <vcpupin vcpu='6' cpuset='9'/> <vcpupin vcpu='7' cpuset='11'/> <vcpupin vcpu='8' cpuset='13'/> <vcpupin vcpu='9' cpuset='15'/> <emulatorpin cpuset='2,4,6,8,10'/> <vcpusched vcpus='0' scheduler='fifo' priority='1'/> <vcpusched vcpus='1' scheduler='fifo' priority='1'/> <vcpusched vcpus='2' scheduler='fifo' priority='1'/> <vcpusched vcpus='3' scheduler='fifo' priority='1'/> <vcpusched vcpus='4' scheduler='fifo' priority='1'/> <vcpusched vcpus='5' scheduler='fifo' priority='1'/> <vcpusched vcpus='6' scheduler='fifo' priority='1'/> <vcpusched vcpus='7' scheduler='fifo' priority='1'/> <vcpusched vcpus='8' scheduler='fifo' priority='1'/> <vcpusched vcpus='9' scheduler='fifo' priority='1'/> </cputune> Host NUMA topology: # lscpu | grep NUMA NUMA node(s): 2 NUMA node0 CPU(s): 0,2,4,6,8,10,12,14,16,18 NUMA node1 CPU(s): 1,3,5,7,9,11,13,15,17,19 Host cores isolation: # cat /proc/cmdline BOOT_IMAGE=(hd0,msdos1)/vmlinuz... console=ttyS0,115200n81 skew_tick=1 isolcpus=managed_irq,domain,1,3,5,7,9,11,13,15,17,19,12,14,16,18 intel_pstate=disable nosoftlockup nohz=on nohz_full=1,3,5,7,9,11,13,15,17,19,12,14,16,18 rcu_nocbs=1,3,5,7,9,11,13,15,17,19,12,14,16,18 default_hugepagesz=1G ... Best reagards, Pei (In reply to Pei Zhang from comment #7) > (In reply to Nitesh Narayan Lal from comment #6) > > Pei, > > > > Can you please confirm if your housekeeping and RT vCPUs were pinned to the > > same NUMA node when you were triggering this spike? > > Hi Nitesh, > > > The housekeeping vCPUs and RT vCPUs are pinned to different NUMA nodes. > Please see below. > > <vcpu placement='static'>10</vcpu> > <cputune> > <vcpupin vcpu='0' cpuset='16'/> > <vcpupin vcpu='1' cpuset='18'/> > <vcpupin vcpu='2' cpuset='1'/> > <vcpupin vcpu='3' cpuset='3'/> > <vcpupin vcpu='4' cpuset='5'/> > <vcpupin vcpu='5' cpuset='7'/> > <vcpupin vcpu='6' cpuset='9'/> > <vcpupin vcpu='7' cpuset='11'/> > <vcpupin vcpu='8' cpuset='13'/> > <vcpupin vcpu='9' cpuset='15'/> > <emulatorpin cpuset='2,4,6,8,10'/> > <vcpusched vcpus='0' scheduler='fifo' priority='1'/> > <vcpusched vcpus='1' scheduler='fifo' priority='1'/> > <vcpusched vcpus='2' scheduler='fifo' priority='1'/> > <vcpusched vcpus='3' scheduler='fifo' priority='1'/> > <vcpusched vcpus='4' scheduler='fifo' priority='1'/> > <vcpusched vcpus='5' scheduler='fifo' priority='1'/> > <vcpusched vcpus='6' scheduler='fifo' priority='1'/> > <vcpusched vcpus='7' scheduler='fifo' priority='1'/> > <vcpusched vcpus='8' scheduler='fifo' priority='1'/> > <vcpusched vcpus='9' scheduler='fifo' priority='1'/> > </cputune> vCPUs 0,1 are guest housekeeping vCPUs. vCPUs 2~9 are guest RT vCPUs. Guest cores isolation: # cat /proc/cmdline BOOT_IMAGE=/vmlinuz... default_hugepagesz=1G iommu=pt intel_iommu=on skew_tick=1 isolcpus=2,3,4,5,6,7,8,9 intel_pstate=disable nosoftlockup nohz=on nohz_full=2,3,4,5,6,7,8,9 rcu_nocbs=2,3,4,5,6,7,8,9 spectre_v2=off nopti kvm-intel.vmentry_l1d_flush=never (In reply to Pei Zhang from comment #8) > (In reply to Pei Zhang from comment #7) > > (In reply to Nitesh Narayan Lal from comment #6) > > > Pei, > > > > > > Can you please confirm if your housekeeping and RT vCPUs were pinned to the > > > same NUMA node when you were triggering this spike? > > > > Hi Nitesh, > > > > > > The housekeeping vCPUs and RT vCPUs are pinned to different NUMA nodes. > > Please see below. > > > > <vcpu placement='static'>10</vcpu> > > <cputune> > > <vcpupin vcpu='0' cpuset='16'/> > > <vcpupin vcpu='1' cpuset='18'/> > > <vcpupin vcpu='2' cpuset='1'/> > > <vcpupin vcpu='3' cpuset='3'/> > > <vcpupin vcpu='4' cpuset='5'/> > > <vcpupin vcpu='5' cpuset='7'/> > > <vcpupin vcpu='6' cpuset='9'/> > > <vcpupin vcpu='7' cpuset='11'/> > > <vcpupin vcpu='8' cpuset='13'/> > > <vcpupin vcpu='9' cpuset='15'/> > > <emulatorpin cpuset='2,4,6,8,10'/> > > <vcpusched vcpus='0' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='1' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='2' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='3' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='4' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='5' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='6' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='7' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='8' scheduler='fifo' priority='1'/> > > <vcpusched vcpus='9' scheduler='fifo' priority='1'/> > > </cputune> > > vCPUs 0,1 are guest housekeeping vCPUs. > vCPUs 2~9 are guest RT vCPUs. > > Guest cores isolation: > # cat /proc/cmdline > BOOT_IMAGE=/vmlinuz... default_hugepagesz=1G iommu=pt intel_iommu=on > skew_tick=1 isolcpus=2,3,4,5,6,7,8,9 intel_pstate=disable nosoftlockup > nohz=on nohz_full=2,3,4,5,6,7,8,9 rcu_nocbs=2,3,4,5,6,7,8,9 spectre_v2=off > nopti kvm-intel.vmentry_l1d_flush=never Thank you for sharing the information. Also, if I am not mistaken then you are not able to reproduce the reported spike with the latest 7.8.z rt-kernel, is that correct? I have not been able to reproduce the latency spike in several of my last test runs with the latest 7.8.z RT kernel and Pei has also reported that she is not able to reproduce the spike anymore, hence I am closing this Bug. Latency spike in comment5: This spike is probably because of pinning housekeeping and RT vCPUs on the same NUMA node. Unfortunately, I don't have access to that system anymore for verification. However, I did manage to create a similar spike in one of my testbeds by following this non-recommended pinning and the latency spike was gone as soon as I switched to the recommended pinning. As per KVM-RT guidelines, it is always recommended to pin housekeeping and RT vCPU on different NUMA nodes, this is to avoid cache thrashing issues. Although, the impact of pinning RT: Housekeeping vCPUs on the same NUMA is hugely dependent upon the workload that is running on the housekeeping vCPUs. Latency spike originally reported: The latency spike was not consistently reproduced to gather enough tracing data to make a definitive conclusion. However, the occasional spikes could be due to the missing timer advancement support in RHEL7 codebase, which sometimes leads to timer interrupts being injected before the actual guest timer expiration and hence resulting in the generation of another timer reprogramming request. Note: The latency spike due to this is hard to trigger. Following is a test run on 4VM running cyclictest simultatensouly with stress for 8h: Environment: Test duration: 8h Run rteval: n Run stress: y Isolated CPUs: 1 Nr Housekeeping CPUs: 1 Kernel: 3.10.0-1127.8.2.rt56.1103.el7.x86_64 Kernel cmd-line: BOOT_IMAGE=/vmlinuz-3.10.0-1127.8.2.rt56.1103.el7.x86_64 root=/dev/mapper/rhel-root ro crashkernel=auto rd.lvm.lv=rhel/root rd.lvm.lv=rhel/swap rhgb quiet console=ttyS0 LANG=en_US.UTF-8 skew_tick=1 isolcpus=1 intel_pstate=disable nosoftlockup nohz=on nohz_full=1 rcu_nocbs=1 x86 debug opts: retp_enabled=0 pti_enabled=0 ibrs_enabled=0 ibpb_enabled=1 Machine: test-vm1 CPU: Intel(R) Xeon(R) CPU E5-2640 v3 @ 2.60GHz Results: VM1: # Min Latencies: 00011 # Avg Latencies: 00014 # Max Latencies: 00022 VM2: # Min Latencies: 00011 # Avg Latencies: 00014 # Max Latencies: 00022 VM3: # Min Latencies: 00011 # Avg Latencies: 00014 # Max Latencies: 00022 VM4: # Min Latencies: 00011 # Avg Latencies: 00014 # Max Latencies: 00022 |
Description of problem: After disabling all mitigations, the 24h cyclictest max latency exceeds 20us a bit, about 2us~6us. Version-Release number of selected component (if applicable): kernel-rt-3.10.0-1121.rt56.1085.el7.x86_64 rt-tests-1.5-9.el7.x86_64 qemu-kvm-rhev-2.12.0-41.el7.x86_64 libvirt-4.5.0-31.el7.x86_64 microcode_ctl-2.1-61.el7.x86_64 tuned-2.11.0-8.el7.noarch How reproducible: 1/1 Steps to Reproduce: 1. Setup rt host 2. Boot rt guest 3. Disable all mitigations # systool -vm kvm | grep nx_huge_pages/sys/kernel/debug/x86/ nx_huge_pages_recovery_ratio= "0" nx_huge_pages = "N" # cat /sys/kernel/debug/x86/ibpb_enabled 1 # cat /sys/kernel/debug/x86/ibrs_enabled 0 # cat /sys/kernel/debug/x86/pti_enabled 0 # cat /sys/kernel/debug/x86/retp_enabled 0 4. Start 24h cyclictest in guest. Meanwhile compiling kernel in both host and guest housekeeping cores. # taskset -c 2,3,4,5,6,7,8,9 cyclictest -m -q -p95 -D 24h -h60 -t 8 -a 2,3,4,5,6,7,8,9 -i 200 5. Check max latency result, exceeds 20us a bit, it's 22us. - Single VM with 8 rt vCPUs: # Min Latencies: 00006 00008 00008 00008 00008 00008 00008 00008 # Avg Latencies: 00008 00008 00008 00008 00008 00008 00008 00008 # Max Latencies: 00019 00019 00019 00022 00019 00020 00019 00020 Actual results: The max latency exceeds 20us. Expected results: We expect the max latency should not exceed 20us. Additional info: 1. We also hit this little spike in rhel7.6.z and rhel7.7z. 7.6.z kernel-rt-3.10.0-957.43.1.rt56.957.el7.x86_64(max latency is 24us): - Single VM with 8 rt vCPUs: # Min Latencies: 00005 00006 00006 00006 00005 00006 00005 00005 # Avg Latencies: 00007 00006 00007 00006 00006 00006 00006 00006 # Max Latencies: 00021 00018 00024 00018 00022 00018 00018 00018 7.7.z kernel-rt-3.10.0-1062.12.1.rt56.1039.el7.x86_64(max latency is 26us): - Single VM with 8 rt vCPUs: # Min Latencies: 00006 00009 00009 00008 00008 00008 00009 00009 # Avg Latencies: 00008 00009 00009 00009 00009 00009 00009 00009 # Max Latencies: 00025 00024 00023 00023 00023 00023 00022 00023 - Multiple VMs each with 1 rt vCPU: - VM1 # Min Latencies: 00006 # Avg Latencies: 00008 # Max Latencies: 00023 - VM2 # Min Latencies: 00006 # Avg Latencies: 00008 # Max Latencies: 00023 - VM3 # Min Latencies: 00006 # Avg Latencies: 00008 # Max Latencies: 00023 - VM4 # Min Latencies: 00007 # Avg Latencies: 00008 # Max Latencies: 00026