Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
The FDP team is no longer accepting new bugs in Bugzilla. Please report your issues under FDP project in Jira. Thanks.

Bug 2059758

Summary: [DPDK][Mellanox][NIC Partitioning] Port creation fails on CX4 VF with higher order of PCI address.
Product: Red Hat Enterprise Linux Fast Datapath Reporter: Timothy Redaelli <tredaelli>
Component: openvswitch2.16Assignee: Open vSwitch development team <ovs-team>
Status: CLOSED ERRATA QA Contact: liting <tli>
Severity: medium Docs Contact:
Priority: medium    
Version: RHEL 8.0CC: apevec, cfields, cfontain, chrisw, ctrautma, dmarchan, ekuris, hakhande, jhsiao, ksundara, ktraynor, ralongi, tli
Target Milestone: ---Keywords: Triaged
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: openvswitch2.16-2.16.0-58.el8fdp Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: 2048601
: 2077451 (view as bug list) Environment:
Last Closed: 2022-03-30 16:28:58 UTC Type: ---
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 2048601    
Bug Blocks: 2077451    

Description Timothy Redaelli 2022-03-01 22:27:38 UTC
+++ This bug was initially created as a clone of Bug #2048601 +++

+++ This bug was initially created as a clone of Bug #2040292 +++

Description of problem:

The OSP deployment involves creation of 64 VFs on 4 CX4 ports (256 VFs), spread across 2 NUMA. A dpdk bond is created between VF0 of port 0 (NUMA0) and VF0 of port 3 (NUMA1). The port creation fails with below error.

    Bridge br-link0
        fail_mode: standalone
        datapath_type: netdev
        Port dpdkbond0
            Interface dpdk1
                type: dpdk
                options: {dpdk-devargs="0000:d8:00.2", n_rxq="1"}
                error: "Error attaching device '0000:d8:00.2' to DPDK"
            Interface dpdk0
                type: dpdk
                options: {dpdk-devargs="0000:12:00.2", n_rxq="1"}
        Port br-link0
            Interface br-link0
                type: internal
    ovs_version: "2.13.4"

How reproducible:

As a quick reproducer with limited VFs, we could build Openvswitch with CONFIG_RTE_MAX_ETHPORTS set to 32 and create a user bridge with dpdk bonds involving a VF with higher order PCI address.


Actual results:
Port creation fails

Expected results:
Port creation shall be successful.

SOS reports are attached in the case.
PS: The port creations were successful when the customer went ahead with 32 VFs per port.

--- Additional comment from Christophe Fontaine on 2022-01-13 12:34:39 UTC ---

When OVS-DPDK starts, it will try to probe all available devices.
With MLX NICs, because of the bifurcated driver, ovs will probe all the VFs, and hits the limit (RTE_MAX_ETHPORTS).

RTE_MAX_ETHPORTS was set to 32 in the log below, but is set to 128 in what we ship.


2022-01-06T08:11:48.252Z|00088|dpdk|INFO|EAL: Probe PCI driver: mlx5_pci (15b3:1018) device: 0000:c4:04.0 (socket 0)
2022-01-06T08:11:48.281Z|00089|dpdk|ERR|mlx5_net: unable to allocate switch domain: Operation not supported
2022-01-06T08:11:48.281Z|00090|dpdk|ERR|mlx5_net: probe of PCI device 0000:c4:04.0 aborted after encountering an error: Operation not supported
2022-01-06T08:11:48.281Z|00091|dpdk|ERR|mlx5_common: Failed to load driver mlx5_eth
2022-01-06T08:11:48.282Z|00092|dpdk|ERR|EAL: Requested device 0000:c4:04.0 cannot be used
[...]
2022-01-06T08:12:00.314Z|00748|dpdk|ERR|EAL: Driver cannot attach the device (0000:c4:04.0)
2022-01-06T08:12:00.314Z|00749|dpdk|ERR|EAL: Failed to attach device on primary process
2022-01-06T08:12:00.314Z|00750|netdev_dpdk|WARN|Error attaching device '0000:c4:04.0' to DPDK
2022-01-06T08:12:00.314Z|00751|netdev|WARN|dpdk0: could not set configuration (Invalid argument)
2022-01-06T08:12:00.314Z|00752|dpdk|ERR|Invalid port_id=32

As a workaround, we can give the allowlist/whitelist (-a or -w) parameter to the EAL, so we don't probe any device:
ovs-vsctl set open_vswitch . other_config:dpdk-extra="-a 0000:00:00.0"

Then, we can hot plug (aka  probe & attach ) the required VFs as needed.

--- Additional comment from Chris Fields on 2022-01-25 20:12:10 UTC ---

In order to write a KCS on this we need template examples.  I'll suggest a 'post-install.yaml' like this one [1] except with this scripted action (instead of the mcast options shown currently):

sudo ovs-vsctl set open_vswitch . other_config:dpdk-extra="-a 0000:00:00.0"

Thoughts?  

For the subsequent probe and attachment of VF's; what are the actions?  

CFields

[1] https://code.engineering.redhat.com/gerrit/gitweb?p=nfv-qe.git;a=blob;f=ospd-16.2-vxlan-dpdk-sriov-ctlplane-dataplane-bonding-lacp-igmp-hybrid/post-install.yaml;hb=bc775078f303ad6948a0acbe67f0c088d63370bc

--- Additional comment from Christophe Fontaine on 2022-01-26 14:57:27 UTC ---

Instead of a post-install.yaml, can't we use the variable "tripleo_ovs_dpdk_extra" ?

https://github.com/openstack/tripleo-ansible/blob/8165e3e94209bbe9ab52cae06de25a56a8b4c812/tripleo_ansible/roles/tripleo_ovs_dpdk/defaults/main.yml#L25

--- Additional comment from Haresh Khandelwal on 2022-01-31 15:18:08 UTC ---

We have discussed this bug internally and below is what we think.
- post-install comes very late in the deployment cycle and this allowlisting is needed much earlier.
- we have tripleo-ansible role with tripleo_ovs_dpdk_extra But dont have corresponding tripleo heat parameter.
- NFV team will own this bz and provide a fix for tripleo-heat-templates. 
- Clone this Bz against FDP product so ovs team can fix actual issue

Thanks

--- Additional comment from Anita Tragler on 2022-01-31 17:15:56 CET ---

This is needed for OSP 16.2 customers requiring to use NIC partitioning
This issue is generic and affects OVS_DPDK NICs COnnectX-4 conenctx-5 that support more than 32 VFs , 64 VFs or higher (28 VFs)
We need DPDK support for higher order PCI IDs.
Is it possible to fix it in DPDK 21.11 and  backport to 20.11?

--- Additional comment from David Marchand on 2022-02-03 14:55:00 CET ---

Strictly speaking, DPDK does not limit "higher order" PCI IDs and there is nothing to "fix" in DPDK.

The issue is how OVS lets DPDK scan all PCI devices, regardless of those devices being used in OVS later.
Because of this, a limit on the number of ports in DPDK (RTE_ETH_MAXPORTS) is reached, and no other mlx device can be initialised in DPDK.

- One option is to make RTE_ETH_MAXPORTS dynamic, but this breaks API/ABI with a huge rework in DPDK.

- One other option is to avoid the initial DPDK scan (this is what the workaround I provided to NFV team).
  DPDK then only probes PCI devices which OVS passes on port addition.
  One note though, this partially breaks support of devices like mlx4 which have multiple ports for a single PCI device because it breaks use of syntax 'class=eth,mac=$MAC' for them.
  This does not affect other devices we currently support, and we dropped mlx4 support not so long ago, so this is not a big problem.

- One last option, probably the best compromise, is to raise the number of accepted ports (RTE_ETH_MAXPORTS) in the embedded DPDK in OVS downstream packages.
  We must decide on a best value (from report, 1024 seems high enough, the max value possible being UINT16_MAX, i.e. 64k).
  This requires QE to do sanity checks with new config.

--- Additional comment from Chris Fields on 2022-02-15 15:16:22 CET ---

Haresh, do you have feedback on the options David presented?

--- Additional comment from Haresh Khandelwal on 2022-02-21 08:04:27 CET ---

(In reply to David Marchand from comment #2)
> Strictly speaking, DPDK does not limit "higher order" PCI IDs and there is
> nothing to "fix" in DPDK.
> 
> The issue is how OVS lets DPDK scan all PCI devices, regardless of those
> devices being used in OVS later.
> Because of this, a limit on the number of ports in DPDK (RTE_ETH_MAXPORTS)
> is reached, and no other mlx device can be initialised in DPDK.
> 
> - One option is to make RTE_ETH_MAXPORTS dynamic, but this breaks API/ABI
> with a huge rework in DPDK.
> 
> - One other option is to avoid the initial DPDK scan (this is what the
> workaround I provided to NFV team).
>   DPDK then only probes PCI devices which OVS passes on port addition.
>   One note though, this partially breaks support of devices like mlx4 which
> have multiple ports for a single PCI device because it breaks use of syntax
> 'class=eth,mac=$MAC' for them.
>   This does not affect other devices we currently support, and we dropped
> mlx4 support not so long ago, so this is not a big problem.
> 
> - One last option, probably the best compromise, is to raise the number of
> accepted ports (RTE_ETH_MAXPORTS) in the embedded DPDK in OVS downstream
> packages.
>   We must decide on a best value (from report, 1024 seems high enough, the
> max value possible being UINT16_MAX, i.e. 64k).
>   This requires QE to do sanity checks with new config.

Ack on this 3rd option David, we can have sanity/perf jobs runs and confirm any regression.

Thanks

--- Additional comment from Red Hat Bugzilla on 2022-02-22 06:43:34 CET ---

remove performed by PnT Account Manager <pnt-expunge>

Comment 1 OvS team 2022-03-02 16:02:56 UTC
* Tue Mar 01 2022 Timothy Redaelli <tredaelli> - 2.16.0-58
- Change RTE_ETH_MAXPORTS to 1024 [RH git: 81ff7c5a60] (#2059758)
    Resolves: #2059758

Comment 5 liting 2022-03-11 07:59:50 UTC
I verified on CX5 card and passed. Do I still need to verify on CX4 card?
For CX5 card:
Reproduce on openvswitch2.15-2.15.0-57.el8fdp.x86_64.

echo 64 > /sys/devices/pci0000:00/0000:00:02.0/0000:04:00.1/sriov_numvfs
echo 64 > /sys/devices/pci0000:00/0000:00:02.0/0000:04:00.0/sriov_numvfs

ovs-vsctl --no-wait set Open_vSwitch . other_config:dpdk-init=true
 ovs-vsctl --no-wait set Open_vSwitch . other_config:dpdk-socket-mem=1024,1024
 ovs-vsctl set Open_vSwitch . other_config:pmd-cpu-mask=0x400000400000
 ovs-vsctl add-br ovsbr0 -- set bridge ovsbr0 datapath_type=netdev
ovs-vsctl add-bond ovsbr0 dpdkbond dpdk0 dpdk1 "bond_mode=active-backup" -- set Interface dpdk0 type=dpdk options:dpdk-devargs=0000:04:09.7 -- set Interface dpdk1 type=dpdk options:dpdk-devargs=0000:04:10.0

[root@dell-per730-56 ~]# ovs-vsctl add-bond ovsbr0 dpdkbond dpdk0 dpdk1 "bond_mode=active-backup" -- set Interface dpdk0 type=dpdk options:dpdk-devargs=0000:04:09.7 -- set Interface dpdk1 type=dpdk options:dpdk-devargs=0000:04:10.0
ovs-vsctl: Error detected while setting up 'dpdk1': Error attaching device '0000:04:10.0' to DPDK.  See ovs-vswitchd log for details.
ovs-vsctl: The default log directory is "/var/log/openvswitch".

[root@dell-per730-56 ~]# ovs-vsctl show
59055c3d-2595-478f-b6ca-26ed5778404e
    Bridge ovsbr0
        datapath_type: netdev
        Port ovsbr0
            Interface ovsbr0
                type: internal
        Port dpdkbond
            Interface dpdk1
                type: dpdk
                options: {dpdk-devargs="0000:04:10.0"}
                error: "Error attaching device '0000:04:10.0' to DPDK"
            Interface dpdk0
                type: dpdk
                options: {dpdk-devargs="0000:04:09.7"}
    ovs_version: "2.15.4"

[root@dell-per730-56 ~]# tail -f /var/log/openvswitch/ovs-vswitchd.log 
2022-03-11T07:47:51.862Z|01203|dpdk|INFO|EAL: Probe PCI driver: mlx5_pci (15b3:1018) device: 0000:04:10.0 (socket 0)
2022-03-11T07:47:51.876Z|01204|dpdk|ERR|mlx5_pci: unable to allocate switch domain: File exists
2022-03-11T07:47:51.876Z|01205|netdev_dpdk|WARN|Error attaching device '0000:04:10.0' to DPDK
2022-03-11T07:47:51.876Z|01206|netdev|WARN|dpdk1: could not set configuration (Invalid argument)
2022-03-11T07:47:51.876Z|01207|dpdk|ERR|Invalid port_id=128

Verified on openvswitch2.16-2.16.0-58.el8fdp.x86_64
steps as same as above
[root@dell-per730-56 ~]# ovs-vsctl show
59055c3d-2595-478f-b6ca-26ed5778404e
    Bridge ovsbr0
        datapath_type: netdev
        Port dpdkbond
            Interface dpdk1
                type: dpdk
                options: {dpdk-devargs="0000:04:10.0"}
            Interface dpdk0
                type: dpdk
                options: {dpdk-devargs="0000:04:09.7"}
        Port ovsbr0
            Interface ovsbr0
                type: internal
    ovs_version: "2.16.3"

Comment 7 liting 2022-03-16 03:09:49 UTC
Verify pass on CX4 card.
[root@netqe30 ~]# lspci|grep af
af:00.0 Ethernet controller: Mellanox Technologies MT27700 Family [ConnectX-4]
af:00.1 Ethernet controller: Mellanox Technologies MT27700 Family [ConnectX-4]

Reproduced on openvswitch2.13-2.13.0-72.el8fdp.x86_64.
[root@netqe30 ~]# echo 64 > /sys/devices/pci0000:ae/0000:ae:00.0/0000:af:00.0/sriov_numvfs
[root@netqe30 ~]# echo 64 > /sys/devices/pci0000:ae/0000:ae:00.0/0000:af:00.1/sriov_numvfs
[root@netqe30 ~]# ovs-vsctl --no-wait set Open_vSwitch . other_config:dpdk-init=true
[root@netqe30 ~]#  ovs-vsctl --no-wait set Open_vSwitch . other_config:dpdk-socket-mem=1024,1024
[root@netqe30 ~]#  ovs-vsctl set Open_vSwitch . other_config:pmd-cpu-mask=0x40000004000000
[root@netqe30 ~]#  ovs-vsctl add-br ovsbr0 -- set bridge ovsbr0 datapath_type=netdev
[root@netqe30 ~]# ovs-vsctl add-bond ovsbr0 dpdkbond dpdk0 dpdk1 "bond_mode=active-backup" -- set Interface dpdk0 type=dpdk options:dpdk-devargs=0000:af:07.7 -- set Interface dpdk1 type=dpdk options:dpdk-devargs=0000:af:18.0

[root@netqe30 ~]# tail -f /var/log/openvswitch/ovs-vswitchd.log 
2022-03-16T03:06:20.827Z|00490|dpif_netdev|INFO|PMD thread on numa_id: 0, core id: 26 created.
2022-03-16T03:06:20.830Z|00491|dpif_netdev|INFO|PMD thread on numa_id: 0, core id: 54 created.
2022-03-16T03:06:20.830Z|00492|dpif_netdev|INFO|There are 2 pmd threads on numa node 0
2022-03-16T03:06:20.830Z|00493|dpdk|INFO|Device with port_id=65 already stopped
2022-03-16T03:06:22.258Z|00494|netdev_dpdk|INFO|Port 65: b6:cb:41:18:e0:d3
2022-03-16T03:06:22.259Z|00495|dpif_netdev|WARN|There's no available (non-isolated) pmd thread on numa node 1. Queue 0 on port 'dpdk0' will be assigned to the pmd on core 26 (numa node 0). Expect reduced performance.
2022-03-16T03:06:22.260Z|00496|bridge|INFO|bridge ovsbr0: added interface dpdk0 on port 1
2022-03-16T03:06:22.297Z|00497|dpdk|INFO|EAL: PCI device 0000:af:18.0 on NUMA socket 1
2022-03-16T03:06:22.297Z|00498|dpdk|INFO|EAL:   probe driver: 15b3:1014 net_mlx5
2022-03-16T03:06:22.323Z|00499|dpdk|ERR|net_mlx5: unable to allocate switch domain: File exists
2022-03-16T03:06:22.325Z|00500|netdev_dpdk|WARN|Error attaching device '0000:af:18.0' to DPDK
2022-03-16T03:06:22.325Z|00501|netdev|WARN|dpdk1: could not set configuration (Invalid argument)
2022-03-16T03:06:22.325Z|00502|dpdk|ERR|Invalid port_id=128

verify pass on openvswitch2.16-2.16.0-58.el8fdp
[root@netqe30 ~]# ovs-vsctl add-bond ovsbr0 dpdkbond dpdk0 dpdk1 "bond_mode=active-backup" -- set Interface dpdk0 type=dpdk options:dpdk-devargs=0000:af:07.7 -- set Interface dpdk1 type=dpdk options:dpdk-devargs=0000:af:18.0

Comment 9 errata-xmlrpc 2022-03-30 16:28:58 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory (openvswitch2.16 bug fix and enhancement update), and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2022:1146