Bug 2030197 - Nmstate fails to configure trunk bridge
Summary: Nmstate fails to configure trunk bridge
Keywords:
Status: CLOSED DUPLICATE of bug 2026621
Alias: None
Product: Container Native Virtualization (CNV)
Classification: Red Hat
Component: Networking
Version: 4.9.10
Hardware: x86_64
OS: Linux
high
high
Target Milestone: ---
: ---
Assignee: Petr Horáček
QA Contact: Meni Yakove
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2021-12-08 07:50 UTC by Karim Latouche
Modified: 2022-07-01 06:03 UTC (History)
8 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2022-01-20 11:34:26 UTC
Target Upstream Version:
Embargoed:


Attachments (Terms of Use)
oc get nnce -o jsonpath='{.status.conditions[?(@.type=="Failing")].message}' > nnce.log (246.16 KB, text/plain)
2021-12-08 07:50 UTC, Karim Latouche
no flags Details

Description Karim Latouche 2021-12-08 07:50:08 UTC
Created attachment 1845194 [details]
oc get nnce -o jsonpath='{.status.conditions[?(@.type=="Failing")].message}' > nnce.log

Description of problem:

Customer wants to run CNV VMs in specific vlans.


Creating a trunk bridge with nmstate fails 


Version-Release number of selected component (if applicable):

4.9.0
How reproducible:

When deploying a bridge on workers using nmstate 




Steps to Reproduce:
1.
cat << EOF | oc create -f -
apiVersion: nmstate.io/v1beta1
kind: NodeNetworkConfigurationPolicy
metadata:
  name: br0-ens2f0np0-policy-workers
spec:
  nodeSelector:
    node-role.kubernetes.io/worker: ""
  desiredState:
    interfaces:
      - name: linux-br0
        description: Linux bridge with ens2f0np0 as a port
        type: linux-bridge
        state: up
        bridge:
          options:
            stp:
              enabled: false
          port:
          - name: ens2f0np0

EOF



  


Actual results:

[root@CentOS-85-64-minimal ~]# oc get nncp
NAME                      STATUS
br0-ens2f0np0-policy-workers      Failing

Expected results:
[root@CentOS-85-64-minimal ~]# oc get nncp
NAME                      STATUS
br0-ens2f0np0-policy-workers      Available

Additional info:

Related bug:
https://bugzilla.redhat.com/show_bug.cgi?id=2026621

Workaround:

This needs to be done for each vlans 

cat << EOF | oc create -f -
apiVersion: nmstate.io/v1beta1
kind: NodeNetworkConfigurationPolicy
metadata:
  name: br0-ens2f0np0-113-policy-workers
spec:
  nodeSelector:
    node-role.kubernetes.io/worker: ""
  desiredState:
    interfaces:
      - name: br0-113
        type: linux-bridge
        state: up
        ipv4:
          enabled: false
        ipv6:
          enabled: false
        bridge:
          options:
            stp:
              enabled: false
          port:
          - name: ens2f0np0.113
      - name: ens2f0np0.113
        type: vlan
        state: up
        vlan:
          base-iface: ens2f0np0
          id: 113
EOF

cat << EOF | oc create -f -
apiVersion: "k8s.cni.cncf.io/v1"
kind: NetworkAttachmentDefinition
metadata:
  name: br-net
  annotations:
    k8s.v1.cni.cncf.io/resourceName: bridge.network.kubevirt.io/br0-113
spec:
  config: '{
    "cniVersion": "0.3.1",
    "name": "br-net", 
    "type": "cnv-bridge", 
    "bridge": "br0-113"
  }'
EOF

Comment 1 Karim Latouche 2021-12-08 14:08:35 UTC
Adding more details:

OCP4 on Baremetal deployment

Workers NICs

[core@ocp4-worker1 ~]$ lspci | egrep -i --color 'network|ethernet'

37:00.0 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller
37:00.1 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller (rev 01)
5d:00.0 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller
5d:00.1 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller (rev 01)

Comment 4 Ben Nemec 2021-12-17 23:10:18 UTC
Since CNV is mentioned in the description I'm going to send it to the CNV component. The standalone operator shouldn't be installed if CNV is.

(I have no idea if the version is correct - I just picked the latest 4.9, but I have no idea if the CNV versions correspond to OCP)

Comment 5 Karim Latouche 2021-12-18 00:55:01 UTC
CNV version is 4.9.10 
Standalone NMstate operator was not installed

Comment 6 Petr Horáček 2022-01-06 08:53:15 UTC
Thanks for the detailed bug description.

Would you be able to check the driver of the used NIC?

sudo ethtool -i ens2f0np0

That may help us confirm whether it is related to the similar https://bugzilla.redhat.com/show_bug.cgi?id=2026621 you linked.

Comment 7 Karim Latouche 2022-01-12 18:21:00 UTC
Sorry I just saw your request:
Here are the needed infos

[root@ocp4-worker0 ~]# sudo ethtool -i ens2f0np0
driver: bnxt_en
version: 4.18.0-305.28.1.el8_4.x86_64
firmware-version: 20.6.135.0/pkg 20.06.0601
expansion-rom-version: 
bus-info: 0000:37:00.0
supports-statistics: yes
supports-test: yes
supports-eeprom-access: yes
supports-register-dump: yes
supports-priv-flags: no

Comment 8 Petr Horáček 2022-01-13 10:01:52 UTC
Gris, I may need your help here. Is it possible that the same issue as we saw with X710 can happen also here with NetXtreme, despite using a different driver? The error seems to be the same:

 state: up
 bridge:
   options:
+    group-addr: 01:80:C2:00:00:00
+    group-forward-mask: 0
...
     stp:
       enabled: false
+      forward-delay: 15
...
   port:
   - name: ens2f0np0
-    vlan:
-      enable-native: false
-      mode: trunk
-      trunk-tags:
-      - id: 2
-      - id: 3
-      - id: 4
...

Comment 9 Gris Ge 2022-01-20 11:34:26 UTC
I have reproduced the problem locally. 

The build in https://bugzilla.redhat.com/show_bug.cgi?id=2026621 would fix this problem.

The fix will ship to RHEL 8.4.0 and 8.5.0 through next batch update of zstream/EUS.

Please reopen this bug if rpm build there cannot fix the problem.

Thank you!

*** This bug has been marked as a duplicate of bug 2026621 ***

Comment 10 Karim Latouche 2022-01-20 12:55:25 UTC
Is this fix gonna be pushed to COreOS as well ?
How do we push it to our existing OCP 4 deployment?
Thx

Comment 11 Petr Horáček 2022-01-20 13:00:28 UTC
Once the fix on nmstate gets released in RHEL 8.4, we would automatically rebuild OpenShift Virtualization container and consume the bug fix.


Note You need to log in before you can comment on or make changes to this bug.