Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
The FDP team is no longer accepting new bugs in Bugzilla. Please report your issues under FDP project in Jira. Thanks.

Bug 1795295

Summary: [ovn] With large SB databases, large reconnection times are observed and ovsdb-server consumes a lot of CPU and memory
Product: Red Hat Enterprise Linux Fast Datapath Reporter: Daniel Alvarez Sanchez <dalvarez>
Component: ovn2.11Assignee: OVN Team <ovnteam>
Status: CLOSED NOTABUG QA Contact: Jianlin Shi <jishi>
Severity: high Docs Contact:
Priority: high    
Version: FDP 20.ACC: ctrautma, fhallal, mmichels, rkhan
Target Milestone: ---   
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-03-22 13:20:15 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Daniel Alvarez Sanchez 2020-01-27 16:17:35 UTC
With large SB databases with many Logical Flows (see below), ovsdb-server consumes a lot of CPU and memory making reconnections from ovn-controller taking a lot of time (30-40 minutes sometimes until everything settles):

[root@ovn-central /]# ovn-nbctl list logical_switch | grep _uuid | wc -l
51
[root@ovn-central /]# ovn-nbctl list logical_switch_port | grep _uuid | wc -l
1247
[root@ovn-central /]# ovn-nbctl list logical_router_port | grep _uuid | wc -l
53
[root@ovn-central /]# ovn-nbctl list logical_router | grep _uuid | wc -l
8
[root@ovn-central /]# ovn-nbctl list ACL | grep _uuid | wc -l
1166
[root@ovn-central /]# ovn-nbctl list port_Group | grep _uuid | wc -l
152
 
 
[root@ovn-central /]# time ovn-sbctl dump-flows | wc -l
680386


A simple script that runs "ovn-sbctl dump-flows" concurrently 20 times leads to an OOM (over 7GB of memory consumption) to SB ovsdb-server after 5 minutes.

As a side-effect, when running ovsdb-server in active-passive mode with pacemaker, monitoring commands using ovs-appctl may timeout causing failovers that makes things worse as it'll force the reconnection of all the clients.

Please, feel free to change the component to openvswitch if you think that ovn is not the right component.

Comment 1 Mark Michelson 2020-05-27 17:44:01 UTC
It seems like we have two issues:

1) High CPU usage by SBDB. This results in ovn-controller taking a long time to converge. It also results in monitoring systems failing over.
2) High memory usage by SBDB. An example of this is that running simultaneous commands to dump flows results in an eventual OOM.

Let's break these down

Issue 1)

How many clients are connecting to the southbound database (this includes the number of ovn-controllers)? For clients that are not ovn-controller or ovn-northd, what are they doing? Are they issuing ovn-sbctl commands, and if so which ones and how frequently? Is the logical network being updated when you see the high CPU usage, or is the network configuration "idle" at this point?

In performing tests with OpenShift-like configurations, we found that the conditional monitoring logic in the southbound database was a point of contention. During your testing, do you have ovn-monitor-all enabled in the ovn-controllers? This can ease some of the load from the southbound database but has the tradeoff of putting more traffic on the network between the SBDB and ovn-controller and putting some more load on ovn-controller. It would be good to see what sort of effect this option has on this scenario.

Issue 2)

This again can be broken into a couple of problems. First is the memory usage of the SBDB during "regular" operation, and the other is the memory usage you described from running `ovn-sbctl dump-flows` in a script.

When you run `ovn-sbctl dump-flows` in a script, each invocation of ovn-sbctl is going to download a sizable portion of the SBDB. ovsdb-server has to package the DB contents in a series of jsonrpc messages for each ovn-sbctl. Given the CPU usage of the sbdb, it likely means that many requests from ovn-sbctl are being serviced at once, meaning lots of memory is being consumed creating those jsonrpc requests. It makes sense that we'd see huge memory consumption and eventually an OOM under this sort of stress. This should be improved, but given how ovsdb-server currently works, I'm not surprised by this.

When it comes to "regular" operation, it's hard to say that the same problem is occurring. If the setup you're using has frequent ovn-sbctl calls occurring, then we may be seeing the same memory problems here as well, just not at quite the same accelerated pace. Otherwise, the memory consumption on ovsdb-server is likely from different causes. However, the memory usage could be a side effect of the CPU usage. If an iteration though ovsdb-server takes a long time to complete, then the following iteration could be tasked with doing many operations, each requiring memory allocations, thus increasing the memory usage.

You mentioned that in your `ovn-sbctl dump-flows` test, you saw 7GB of memory used before the OOM. What sort of memory usage are you seeing during typical usage? Does the memory continue to grow over a large time period, or does it eventually plateau at some level? Part of the issue is that as the database grows, it results in more memory being used in ovsdb-server to represent it.


Overall, I think this requires some performance and memory analysis, and doing so with as close a setup to the one you're using is best here. If you're able to share your SBDB, then we could attempt to set up something similar and see what we can find.

Comment 2 Mark Michelson 2021-03-01 14:41:37 UTC
Hey Daniel, this issue has been open for a long time. It was opened against OVN2.11, but since then we've released OVN2.13, and OVN2.13 has lots of performance-improving changes.

Is this issue still relevant? Can we close?

Comment 3 Mark Michelson 2021-03-22 13:20:15 UTC
Closing. This issue is mitigated by lots of fixes that have been made to OVN since it was initially opened. When questioned if this was still an issue, got no response.

If there are still performance issues with regards to the southbound database, then please feel free to open a new issue.

Comment 4 Red Hat Bugzilla 2023-09-15 00:21:01 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days