Fedora Account System
Red Hat Associate
Red Hat Customer
Created attachment 1845194 [details] oc get nnce -o jsonpath='{.status.conditions[?(@.type=="Failing")].message}' > nnce.log Description of problem: Customer wants to run CNV VMs in specific vlans. Creating a trunk bridge with nmstate fails Version-Release number of selected component (if applicable): 4.9.0 How reproducible: When deploying a bridge on workers using nmstate Steps to Reproduce: 1. cat << EOF | oc create -f - apiVersion: nmstate.io/v1beta1 kind: NodeNetworkConfigurationPolicy metadata: name: br0-ens2f0np0-policy-workers spec: nodeSelector: node-role.kubernetes.io/worker: "" desiredState: interfaces: - name: linux-br0 description: Linux bridge with ens2f0np0 as a port type: linux-bridge state: up bridge: options: stp: enabled: false port: - name: ens2f0np0 EOF Actual results: [root@CentOS-85-64-minimal ~]# oc get nncp NAME STATUS br0-ens2f0np0-policy-workers Failing Expected results: [root@CentOS-85-64-minimal ~]# oc get nncp NAME STATUS br0-ens2f0np0-policy-workers Available Additional info: Related bug: https://bugzilla.redhat.com/show_bug.cgi?id=2026621 Workaround: This needs to be done for each vlans cat << EOF | oc create -f - apiVersion: nmstate.io/v1beta1 kind: NodeNetworkConfigurationPolicy metadata: name: br0-ens2f0np0-113-policy-workers spec: nodeSelector: node-role.kubernetes.io/worker: "" desiredState: interfaces: - name: br0-113 type: linux-bridge state: up ipv4: enabled: false ipv6: enabled: false bridge: options: stp: enabled: false port: - name: ens2f0np0.113 - name: ens2f0np0.113 type: vlan state: up vlan: base-iface: ens2f0np0 id: 113 EOF cat << EOF | oc create -f - apiVersion: "k8s.cni.cncf.io/v1" kind: NetworkAttachmentDefinition metadata: name: br-net annotations: k8s.v1.cni.cncf.io/resourceName: bridge.network.kubevirt.io/br0-113 spec: config: '{ "cniVersion": "0.3.1", "name": "br-net", "type": "cnv-bridge", "bridge": "br0-113" }' EOF
Adding more details: OCP4 on Baremetal deployment Workers NICs [core@ocp4-worker1 ~]$ lspci | egrep -i --color 'network|ethernet' 37:00.0 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller 37:00.1 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller (rev 01) 5d:00.0 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller 5d:00.1 Ethernet controller: Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller (rev 01)
Since CNV is mentioned in the description I'm going to send it to the CNV component. The standalone operator shouldn't be installed if CNV is. (I have no idea if the version is correct - I just picked the latest 4.9, but I have no idea if the CNV versions correspond to OCP)
CNV version is 4.9.10 Standalone NMstate operator was not installed
Thanks for the detailed bug description. Would you be able to check the driver of the used NIC? sudo ethtool -i ens2f0np0 That may help us confirm whether it is related to the similar https://bugzilla.redhat.com/show_bug.cgi?id=2026621 you linked.
Sorry I just saw your request: Here are the needed infos [root@ocp4-worker0 ~]# sudo ethtool -i ens2f0np0 driver: bnxt_en version: 4.18.0-305.28.1.el8_4.x86_64 firmware-version: 20.6.135.0/pkg 20.06.0601 expansion-rom-version: bus-info: 0000:37:00.0 supports-statistics: yes supports-test: yes supports-eeprom-access: yes supports-register-dump: yes supports-priv-flags: no
Gris, I may need your help here. Is it possible that the same issue as we saw with X710 can happen also here with NetXtreme, despite using a different driver? The error seems to be the same: state: up bridge: options: + group-addr: 01:80:C2:00:00:00 + group-forward-mask: 0 ... stp: enabled: false + forward-delay: 15 ... port: - name: ens2f0np0 - vlan: - enable-native: false - mode: trunk - trunk-tags: - - id: 2 - - id: 3 - - id: 4 ...
I have reproduced the problem locally. The build in https://bugzilla.redhat.com/show_bug.cgi?id=2026621 would fix this problem. The fix will ship to RHEL 8.4.0 and 8.5.0 through next batch update of zstream/EUS. Please reopen this bug if rpm build there cannot fix the problem. Thank you! *** This bug has been marked as a duplicate of bug 2026621 ***
Is this fix gonna be pushed to COreOS as well ? How do we push it to our existing OCP 4 deployment? Thx
Once the fix on nmstate gets released in RHEL 8.4, we would automatically rebuild OpenShift Virtualization container and consume the bug fix.