Bug 1848665
| Summary: | [RHEL8.3] all MVAPICH2 benchmarks with "mpirun" fail with "src/mpid/ch3/channels/mrail/src/gen2/ibv_channel_manager.c:548:" on MLX5 IB0 devices | ||||||
|---|---|---|---|---|---|---|---|
| Product: | Red Hat Enterprise Linux 8 | Reporter: | Brian Chae <bchae> | ||||
| Component: | mvapich2 | Assignee: | Honggang LI <honli> | ||||
| Status: | CLOSED WONTFIX | QA Contact: | Afom T. Michael <tmichael> | ||||
| Severity: | unspecified | Docs Contact: | |||||
| Priority: | unspecified | ||||||
| Version: | 8.3 | CC: | hwkernel-mgr, linville, rdma-dev-team, tmichael | ||||
| Target Milestone: | rc | Keywords: | Triaged | ||||
| Target Release: | 8.4 | Flags: | pm-rhel:
mirror+
|
||||
| Hardware: | Unspecified | ||||||
| OS: | Unspecified | ||||||
| Whiteboard: | |||||||
| Fixed In Version: | Doc Type: | If docs needed, set a value | |||||
| Doc Text: | Story Points: | --- | |||||
| Clone Of: | Environment: | ||||||
| Last Closed: | 2021-12-18 07:26:57 UTC | Type: | Bug | ||||
| Regression: | --- | Mount Type: | --- | ||||
| Documentation: | --- | CRM: | |||||
| Verified Versions: | Category: | --- | |||||
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |||||
| Cloudforms Team: | --- | Target Upstream Version: | |||||
| Embargoed: | |||||||
| Attachments: |
|
||||||
|
Description
Brian Chae
2020-06-18 17:23:26 UTC
Description of problem:
This is a regression on MVAPICH2 on MLX5 IB0 devices, where all benchmarks with "mpirun" fail from RHEL8.2.
All benchmarks failed with the following error message.
[rdma-virt-02.lab.bos.redhat.com:mpi_rank_0][handle_cqe] Send desc error in msg to 1, wc_opcode=0
[rdma-virt-02.lab.bos.redhat.com:mpi_rank_0][handle_cqe] Msg from 1: wc.status=12, wc.wr_id=0x559ffa82c040, wc.opcode=0, vbuf->phead->type=0 = MPIDI_CH3_PKT_EAGER_SEND
[rdma-virt-02.lab.bos.redhat.com:mpi_rank_0][handle_cqe] src/mpid/ch3/channels/mrail/src/gen2/ibv_channel_manager.c:548: [] Got completion with error 12, vendor code=0x81, dest rank=1
: Protocol not supported (93)
[mpiexec.bos.redhat.com] HYDU_sock_write (utils/sock/sock.c:294): write error (Bad file descriptor)
[mpiexec.bos.redhat.com] HYD_pmcd_pmiserv_send_signal (pm/pmiserv/pmiserv_cb.c:177): unable to write data to proxy
[mpiexec.bos.redhat.com] ui_cmd_cb (pm/pmiserv/pmiserv_pmci.c:79): unable to send signal downstream
[mpiexec.bos.redhat.com] HYDT_dmxu_poll_wait_for_event (tools/demux/demux_poll.c:76): callback returned error status
[mpiexec.bos.redhat.com] HYD_pmci_wait_for_completion (pm/pmiserv/pmiserv_pmci.c:198): error waiting for event
[mpiexec.bos.redhat.com] main (ui/mpich/mpiexec.c:340): process manager error waiting for completion
+ [20-06-17 13:02:25] __MPI_check_result 255 mpitests-mvapich2 IMB-MPI1 PingPong mpirun /root/hfile_one_core
Version-Release number of selected component (if applicable):
DISTRO=RHEL-8.3.0-20200616.0
+ [20-06-17 12:59:16] cat /etc/redhat-release
Red Hat Enterprise Linux release 8.3 Beta (Ootpa)
+ [20-06-17 12:59:16] uname -a
Linux rdma-dev-21.lab.bos.redhat.com 4.18.0-214.el8.x86_64 #1 SMP Fri Jun 12 08:55:08 UTC 2020 x86_64 x86_64 x86_64 GNU/Linux
+ [20-06-17 12:59:16] cat /proc/cmdline
BOOT_IMAGE=(hd0,msdos1)/vmlinuz-4.18.0-214.el8.x86_64 root=UUID=3cfc7774-3ae5-434c-8fc0-65f4e7497ca2 ro intel_idle.max_cstate=0 processor.max_cstate=0 intel_iommu=on iommu=on console=tty0 rd_NO_PLYMOUTH crashkernel=auto resume=UUID=d2129f6f-6378-42c4-a1b2-b73e2d143d8f console=ttyS1,115200n81
+ [20-06-17 12:59:16] rpm -q rdma-core linux-firmware
rdma-core-29.0-3.el8.x86_64
linux-firmware-20200512-98.gitb2cad6a2.el8.noarch
+ [20-06-17 12:59:16] tail /sys/class/infiniband/mlx5_0/fw_ver /sys/class/infiniband/mlx5_1/fw_ver /sys/class/infiniband/mlx5_2/fw_ver
==> /sys/class/infiniband/mlx5_0/fw_ver <==
12.23.1020
==> /sys/class/infiniband/mlx5_1/fw_ver <==
12.23.1020
==> /sys/class/infiniband/mlx5_2/fw_ver <==
12.23.1020
+ [20-06-17 12:59:16] lspci
+ [20-06-17 12:59:16] grep -i -e ethernet -e infiniband -e omni -e ConnectX
01:00.0 Ethernet controller: Broadcom Inc. and subsidiaries NetXtreme BCM5720 2-port Gigabit Ethernet PCIe
01:00.1 Ethernet controller: Broadcom Inc. and subsidiaries NetXtreme BCM5720 2-port Gigabit Ethernet PCIe
02:00.0 Ethernet controller: Broadcom Inc. and subsidiaries NetXtreme BCM5720 2-port Gigabit Ethernet PCIe
02:00.1 Ethernet controller: Broadcom Inc. and subsidiaries NetXtreme BCM5720 2-port Gigabit Ethernet PCIe
04:00.0 Ethernet controller: Mellanox Technologies MT27700 Family [ConnectX-4]
82:00.0 Infiniband controller: Mellanox Technologies MT27700 Family [ConnectX-4]
82:00.1 Infiniband controller: Mellanox Technologies MT27700 Family [ConnectX-4]
How reproducible:
100% on MLX5 ConnectX-4 devices
Steps to Reproduce:
1. mpirun -hostfile <hostfile> -n 2 <benchmark>
Actual results:
One of the failed benchmark shows:
+ [20-06-17 12:59:25] timeout --preserve-status --kill-after=5m 3m mpirun -hostfile /root/hfile_one_core -np 2 mpitests-IMB-MPI1 PingPong -time 1.5
[src/mpid/ch3/channels/mrail/src/gen2/rdma_iba_priv.c:1672] Could not modify qpto RTR
#------------------------------------------------------------
# Intel(R) MPI Benchmarks 2019 Update 6, MPI-1 part
#------------------------------------------------------------
# Date : Wed Jun 17 12:59:25 2020
# Machine : x86_64
# System : Linux
# Release : 4.18.0-214.el8.x86_64
# Version : #1 SMP Fri Jun 12 08:55:08 UTC 2020
# MPI Version : 3.1
# MPI Thread Environment:
# Calling sequence was:
# mpitests-IMB-MPI1 PingPong -time 1.5
# Minimum message length in bytes: 0
# Maximum message length in bytes: 4194304
#
# MPI_Datatype : MPI_BYTE
# MPI_Datatype for reductions : MPI_FLOAT
# MPI_Op : MPI_SUM
#
#
# List of Benchmarks to run:
# PingPong
[rdma-virt-02.lab.bos.redhat.com:mpi_rank_0][handle_cqe] Send desc error in msg to 1, wc_opcode=0
[rdma-virt-02.lab.bos.redhat.com:mpi_rank_0][handle_cqe] Msg from 1: wc.status=12, wc.wr_id=0x559ffa82c040, wc.opcode=0, vbuf->phead->type=0 = MPIDI_CH3_PKT_EAGER_SEND
[rdma-virt-02.lab.bos.redhat.com:mpi_rank_0][handle_cqe] src/mpid/ch3/channels/mrail/src/gen2/ibv_channel_manager.c:548: [] Got completion with error 12, vendor code=0x81, dest rank=1
: Protocol not supported (93)
[mpiexec.bos.redhat.com] HYDU_sock_write (utils/sock/sock.c:294): write error (Bad file descriptor)
[mpiexec.bos.redhat.com] HYD_pmcd_pmiserv_send_signal (pm/pmiserv/pmiserv_cb.c:177): unable to write data to proxy
[mpiexec.bos.redhat.com] ui_cmd_cb (pm/pmiserv/pmiserv_pmci.c:79): unable to send signal downstream
[mpiexec.bos.redhat.com] HYDT_dmxu_poll_wait_for_event (tools/demux/demux_poll.c:76): callback returned error status
[mpiexec.bos.redhat.com] HYD_pmci_wait_for_completion (pm/pmiserv/pmiserv_pmci.c:198): error waiting for event
[mpiexec.bos.redhat.com] main (ui/mpich/mpiexec.c:340): process manager error waiting for completion
+ [20-06-17 13:02:25] __MPI_check_result 255 mpitests-mvapich2 IMB-MPI1 PingPong mpirun /root/hfile_one_core
Expected results:
+ [20-06-18 06:06:46] timeout --preserve-status --kill-after=5m 3m mpirun -hostfile /root/hfile_one_core -np 2 mpitests-IMB-MPI1 PingPong -time 1.5
#------------------------------------------------------------
# Intel(R) MPI Benchmarks 2019 Update 6, MPI-1 part
#------------------------------------------------------------
# Date : Thu Jun 18 06:06:46 2020
# Machine : x86_64
# System : Linux
# Release : 4.18.0-214.el8.x86_64
# Version : #1 SMP Fri Jun 12 08:55:08 UTC 2020
# MPI Version : 3.1
# MPI Thread Environment:
# Calling sequence was:
# mpitests-IMB-MPI1 PingPong -time 1.5
# Minimum message length in bytes: 0
# Maximum message length in bytes: 4194304
#
# MPI_Datatype : MPI_BYTE
# MPI_Datatype for reductions : MPI_FLOAT
# MPI_Op : MPI_SUM
#
#
# List of Benchmarks to run:
# PingPong
#---------------------------------------------------
# Benchmarking PingPong
# #processes = 2
#---------------------------------------------------
#bytes #repetitions t[usec] Mbytes/sec
0 1000 1.03 0.00
1 1000 1.07 0.93
2 1000 1.07 1.87
4 1000 1.07 3.75
8 1000 1.07 7.47
16 1000 1.09 14.62
32 1000 1.11 28.76
64 1000 1.16 55.28
128 1000 1.25 102.68
256 1000 1.89 135.23
512 1000 2.02 253.47
1024 1000 2.29 446.39
2048 1000 2.84 720.75
4096 1000 3.94 1039.19
8192 1000 5.20 1574.55
16384 1000 7.36 2224.58
32768 1000 9.56 3426.12
65536 640 14.61 4486.54
131072 320 24.94 5256.26
262144 160 45.61 5747.86
524288 80 86.89 6034.12
1048576 40 169.70 6179.09
2097152 20 335.09 6258.45
4194304 10 667.62 6282.47
# All processes entering MPI_Finalize
Additional info:
Created attachment 1697999 [details]
client log for mvapich2 on mlx5 ib0 where all benchmarks with mpirun failed with same reason
https://bugzilla.redhat.com/show_bug.cgi?id=1672767#c3 The test was run over machines with multiple HCAs. Please setup environment variable 'MV2_IBA_HCA'. (In reply to Honggang LI from comment #3) > https://bugzilla.redhat.com/show_bug.cgi?id=1672767#c3 > > The test was run over machines with multiple HCAs. Please setup environment > variable 'MV2_IBA_HCA'. Honggang, with "MV2_IBA_HCA=mlx5_1" for mlx5_0 device under test, the "mpirun" test for mpapich2 benchmarks successfully run. However, this CANNOT be the fix as we need to test on different IB HCAs. (In reply to Brian Chae from comment #4) > (In reply to Honggang LI from comment #3) > > https://bugzilla.redhat.com/show_bug.cgi?id=1672767#c3 > > > > The test was run over machines with multiple HCAs. Please setup environment > > variable 'MV2_IBA_HCA'. > > Honggang, with "MV2_IBA_HCA=mlx5_1" for mlx5_0 device under test, the > "mpirun" test for mpapich2 benchmarks successfully run. Well, it is a known issue in upstream mvapich2. > However, this CANNOT be the fix as we need to test on different IB HCAs. Maybe we need to update our beaker mvapich2 case? (In reply to Honggang LI from comment #5) > (In reply to Brian Chae from comment #4) > > (In reply to Honggang LI from comment #3) > > > https://bugzilla.redhat.com/show_bug.cgi?id=1672767#c3 > > > > > > The test was run over machines with multiple HCAs. Please setup environment > > > variable 'MV2_IBA_HCA'. > > > > Honggang, with "MV2_IBA_HCA=mlx5_1" for mlx5_0 device under test, the > > "mpirun" test for mpapich2 benchmarks successfully run. > > Well, it is a known issue in upstream mvapich2. > > > However, this CANNOT be the fix as we need to test on different IB HCAs. > > Maybe we need to update our beaker mvapich2 case? Sure, we can do that, as long as we keep this bug open. (In reply to Brian Chae from comment #6) > (In reply to Honggang LI from comment #5) > > (In reply to Brian Chae from comment #4) > > > (In reply to Honggang LI from comment #3) > > > > https://bugzilla.redhat.com/show_bug.cgi?id=1672767#c3 > > > > > > > > The test was run over machines with multiple HCAs. Please setup environment > > > > variable 'MV2_IBA_HCA'. > > > > > > Honggang, with "MV2_IBA_HCA=mlx5_1" for mlx5_0 device under test, the > > > "mpirun" test for mpapich2 benchmarks successfully run. > > > > Well, it is a known issue in upstream mvapich2. > > > > > However, this CANNOT be the fix as we need to test on different IB HCAs. > > > > Maybe we need to update our beaker mvapich2 case? > > Sure, we can do that, as long as we keep this bug open. FYI, both "mpirun" and "mpirun_rsh" for mvapich2 benchmarks on MLX5 multi-HCA IB device works with this workaround. Additional info: We have MLX5 multi-HCAs on rdma-dev-21/22, and rdma-dev-19/20 hosts, but this bug is not observed for "mpirun" mvapich2 benchmark runs. So, this bug applies only to rmda-perf-02/03 (single HCA MLX5) and rmda-virt-02/03 (multi HCA MLX5). After evaluating this issue, there are no plans to address it further or fix it in an upcoming release. Therefore, it is being closed. If plans change such that this issue will be fixed in an upcoming release, then the bug can be reopened. |