Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
The FDP team is no longer accepting new bugs in Bugzilla. Please report your issues under FDP project in Jira. Thanks.

Bug 1827769

Summary: ovn-northd does not release memory after cleanup
Product: Red Hat Enterprise Linux Fast Datapath Reporter: Joe Talerico <jtaleric>
Component: ovn2.13Assignee: Ilya Maximets <i.maximets>
Status: CLOSED ERRATA QA Contact: Jianlin Shi <jishi>
Severity: urgent Docs Contact:
Priority: urgent    
Version: RHEL 7.6CC: avishnoi, ctrautma, dcbw, i.maximets, jishi, mmichels, ralongi, rkhan, trozet
Target Milestone: ---Keywords: TestBlocker
Target Release: ---   
Hardware: All   
OS: All   
Whiteboard: aos-scalability-45
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of:
: 1834838 (view as bug list) Environment:
Last Closed: 2020-05-26 14:07:18 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 1834838    
Attachments:
Description Flags
perf record of northd
none
northd-debug-log none

Description Joe Talerico 2020-04-24 18:26:38 UTC
Description of problem:

sbdb seems to free the memory after ~ 5 hours[1]. However, nbdb doesn't ever freed the memory.

[1] https://snapshot.raintank.io/dashboard/snapshot/e3lrRBySJ47DBI0heOYYYMxmtVHzGJhd?orgId=2&fullscreen&panelId=156

Version-Release number of selected component (if applicable):

Not sure if Version is set correctly to RHEL7.6 , more info below. 

OCP 4.5
ovn2.13-2.13.0-16.el7fdp.x86_64

How reproducible:
100%

Steps to Reproduce:
1. Deploy 4.5 on AWS
2. Run Mastervert 500
3. Delete all newly created projects after mastervert completes 
4. Observe sbdb / nbdb memory growth

Possibly need to update the RAFT timer to avoid many leader elections. 

Actual results:
https://snapshot.raintank.io/dashboard/snapshot/e3lrRBySJ47DBI0heOYYYMxmtVHzGJhd?orgId=2&fullscreen&panelId=156


Expected results:
sbdb eventually did free memory, but it was over a 5 hour period. 

nbdb never cleaned up

Additional info:

Comment 1 Dan Williams 2020-04-25 02:16:02 UTC
Joe, does the database itself actually shrink, just that memory is not freed?  eg, let's make sure ovn-kubernetes is doing the right thing first by deleting all the relevant objects from the NBDB.

Next, can you grab and attach the nbdb and sbdb files (should be in /var/lib/ovn I believe) from a master node after giving ovn-kubernetes enough time to process all the kube API deletion events?

Comment 2 Joe Talerico 2020-04-27 14:19:32 UTC
(In reply to Dan Williams from comment #1)
> Joe, does the database itself actually shrink, just that memory is not
> freed?  eg, let's make sure ovn-kubernetes is doing the right thing first by
> deleting all the relevant objects from the NBDB.

I did not take note the db size. I will capture that this next iteration. The cluster became unstable over the weekend[1].

> 
> Next, can you grab and attach the nbdb and sbdb files (should be in
> /var/lib/ovn I believe) from a master node after giving ovn-kubernetes
> enough time to process all the kube API deletion events?

The link I shared in Comment #1 shows that sbdb frees the memory after ~5 hours... nbdb never freed the memory.

I will rerun the test and capture the results.

[1] https://gist.github.com/jtaleric/5d7719d86b913177cc257bdf8c45cae1

Comment 3 Joe Talerico 2020-04-27 19:09:25 UTC
(In reply to Joe Talerico from comment #2)
> (In reply to Dan Williams from comment #1)
> > Joe, does the database itself actually shrink, just that memory is not
> > freed?  eg, let's make sure ovn-kubernetes is doing the right thing first by
> > deleting all the relevant objects from the NBDB.
> 
> I did not take note the db size. I will capture that this next iteration.
> The cluster became unstable over the weekend[1].
> 
> > 
> > Next, can you grab and attach the nbdb and sbdb files (should be in
> > /var/lib/ovn I believe) from a master node after giving ovn-kubernetes
> > enough time to process all the kube API deletion events?
> 
> The link I shared in Comment #1 shows that sbdb frees the memory after ~5
> hours... nbdb never freed the memory.
> 
> I will rerun the test and capture the results.
> 
> [1] https://gist.github.com/jtaleric/5d7719d86b913177cc257bdf8c45cae1

The database size at 25 nodes, 25 nodes w/ 50 projects 545 Running pods, after cleanup, and ~60 minutes after cleanup

root@ip-172-31-68-73: ~/pprof # oc exec -n openshift-ovn-kubernetes ovnkube-master-5nhvq -c sbdb --  ls -talr /etc/ovn/ | tee 25_nodes_dbsize                                                                                                                 
total 18944
-rw-------. 1 root root        0 Apr 27 15:44 .ovnnb_db.db.~lock~
drwxr-xr-x. 1 root root       17 Apr 27 15:44 ..
-rw-------. 1 root root        0 Apr 27 15:44 .ovnsb_db.db.~lock~
drwxr-xr-x. 2 root root       98 Apr 27 16:37 .
-rw-r--r--. 1 root root  3667363 Apr 27 16:46 ovnnb_db.db
-rw-r--r--. 1 root root 13693365 Apr 27 16:46 ovnsb_db.db
root@ip-172-31-68-73: ~/pprof # oc exec -n openshift-ovn-kubernetes ovnkube-master-5nhvq -c sbdb --  ls -talr /etc/ovn/ | tee 25_nodes-50_projects-545Pods_dbsize                                                                                             
total 54720
-rw-------. 1 root root        0 Apr 27 15:44 .ovnnb_db.db.~lock~
drwxr-xr-x. 1 root root       17 Apr 27 15:44 ..
-rw-------. 1 root root        0 Apr 27 15:44 .ovnsb_db.db.~lock~
drwxr-xr-x. 2 root root       98 Apr 27 17:12 .
-rw-r--r--. 1 root root   551612 Apr 27 17:14 ovnnb_db.db
-rw-r--r--. 1 root root 23979842 Apr 27 17:14 ovnsb_db.db
root@ip-172-31-68-73: ~/pprof # oc exec -n openshift-ovn-kubernetes ovnkube-master-5nhvq -c sbdb --  ls -talr /etc/ovn/ | tee 25_nodes-cleanup-rightafter-50_projects-545Pods_dbsize
total 136768
-rw-------. 1 root root        0 Apr 27 15:44 .ovnnb_db.db.~lock~
drwxr-xr-x. 1 root root       17 Apr 27 15:44 ..
-rw-------. 1 root root        0 Apr 27 15:44 .ovnsb_db.db.~lock~
drwxr-xr-x. 2 root root       98 Apr 27 17:12 .
-rw-r--r--. 1 root root 23395173 Apr 27 17:18 ovnnb_db.db
-rw-r--r--. 1 root root 47149068 Apr 27 17:18 ovnsb_db.db
root@ip-172-31-68-73: ~/pprof # oc exec -n openshift-ovn-kubernetes ovnkube-master-5nhvq -c sbdb --  ls -talr /etc/ovn/ | tee 25_nodes-cleanup-60min-50_projects-545Pods_dbsize                                                                               
total 17216
-rw-------. 1 root root        0 Apr 27 15:44 .ovnnb_db.db.~lock~
drwxr-xr-x. 1 root root       17 Apr 27 15:44 ..
-rw-------. 1 root root        0 Apr 27 15:44 .ovnsb_db.db.~lock~
drwxr-xr-x. 2 root root       98 Apr 27 17:57 .
-rw-r--r--. 1 root root   347775 Apr 27 17:58 ovnnb_db.db
-rw-r--r--. 1 root root 11739159 Apr 27 17:58 ovnsb_db.db

We can see here[1], sbdb does eventually free the memory, however nbdb grows during the cleanup (~13:16 timeframe). We also witnessed this in the initial BZ comment[2] (@1700). 

I ran the workload twice, each time cleaning up after the run. OVN nbdb didn't grow the second iteration. It seems to allocate memory, and possibly reuse the memory it already allocated. However, when reviewing other components of OVN, I noticed that northd was growing with each iteration[3]


[1] https://snapshot.raintank.io/dashboard/snapshot/3SZSO97byd26A39871j2semDHxHhVoPT?orgId=2
[2] https://snapshot.raintank.io/dashboard/snapshot/e3lrRBySJ47DBI0heOYYYMxmtVHzGJhd?orgId=2&fullscreen&panelId=156
[3] https://snapshot.raintank.io/dashboard/snapshot/L4vkD6a4SFk9fNTv0DlOrWeQyu3Nar8W?orgId=2&from=1588008663771&to=1588012400097

Comment 4 Joe Talerico 2020-04-29 10:44:23 UTC
northd continues to consume memory, without any workload running in the cluster[1] 

root@ip-172-31-68-73: ~ # diff northd_1day_later northd_mem_before
89,90c89,90
< 55f383db8000-55f52e25c000 rw-p 00000000 00:00 0                          [heap]
< Size:            6984336 kB
---
> 55f383db8000-55f3d1a09000 rw-p 00000000 00:00 0                          [heap]
> Size:            1274180 kB

Capture with a 5 minute pause and re-capture.

root@ip-172-31-68-73: ~ # diff rollup_1 rollup_2 
2,3c2,3
< Rss:             7038436 kB
< Pss:             7033454 kB
---
> Rss:             7051548 kB
> Pss:             7046566 kB
7,9c7,9
< Private_Dirty:   7031200 kB
< Referenced:      7038436 kB
< Anonymous:       7031200 kB
---
> Private_Dirty:   7044312 kB
> Referenced:      7051548 kB
> Anonymous:       7044312 kB


[1] https://snapshot.raintank.io/dashboard/snapshot/1p2wGXtulkohVJm27C4HOeB7UYkgkQ0A?orgId=2

Comment 5 Joe Talerico 2020-04-29 19:24:55 UTC
Created attachment 1683051 [details]
perf record of northd

Comment 7 Joe Talerico 2020-04-29 19:36:04 UTC
Created attachment 1683064 [details]
northd-debug-log

Comment 10 Dan Williams 2020-05-01 22:04:25 UTC
@joe does memory go down if you rsh to each of the master 'nbdb' or 'sbdb' containers and run: 

ovs-appctl -t /var/run/ovn/ovnnb_db.ctl ovsdb-server/compact
ovs-appctl -t /var/run/ovn/ovnsb_db.ctl ovsdb-server/compact

Comment 11 Joe Talerico 2020-05-04 10:15:04 UTC
(In reply to Dan Williams from comment #10)
> @joe does memory go down if you rsh to each of the master 'nbdb' or 'sbdb'
> containers and run: 
> 
> ovs-appctl -t /var/run/ovn/ovnnb_db.ctl ovsdb-server/compact
> ovs-appctl -t /var/run/ovn/ovnsb_db.ctl ovsdb-server/compact

Hey Dan - I ran the compact across the masters, waited for a couple of minutes to see if we saw any kind of dip, and it didn't.[1]

I wanted to also note, that I redeployed a clean OCP4.5 cluster last week (3 masters, 2 infra, 3 worker nodes). I let the cluster sit over the weekend, here is the result of a seemingly small "idle" cluster[2]. 

I say "idle", since the networking bits are never really "idle" in cloud.

[1] https://snapshot.raintank.io/dashboard/snapshot/fGRivIr0KsJGI3dO96bUlX1Ink2irOfV
[2] https://snapshot.raintank.io/dashboard/snapshot/lveTNSxypW73uJg6omtBFGOVsScqqERR?orgId=2

Comment 12 Dan Williams 2020-05-04 15:21:07 UTC
Thanks Joe; a stab in the dark there but good to rule out.

Comment 13 Rashid Khan 2020-05-04 19:49:15 UTC
Hi Joe, 
Quick question, this bug is listed against RHEL7.6?
Is it because you were running that version in the VM? or some other reason?

Comment 14 Joe Talerico 2020-05-05 12:20:32 UTC
Possibly by mistake? This is OCP4.5 - which is : 

sh-4.2# cat /etc/redhat-release 
CentOS Linux release 7.7.1908 (Core)

Comment 15 Ilya Maximets 2020-05-07 14:25:14 UTC
I think, I managed to reproduce the issue with databases and northd not freeing the memory.

I've set up fake cluster with 100 nodes and 500 ports, than I removed all 500 VIFs from OVS
instances and removed all the corresponding logical switch ports from NB db.  After removing
VIFs memory consumption of ovn-northd almost doubled, memory of  ovsdb-servers for NB and SB
DBs slightly increased.  After removing logical switch ports for NB db memory consumption
didn't go down and still almost same values.

ovsdb-server/compact didn't change the amount of consumed virtual memory, but it significantly
reduced the size of NB database on the disk (10 times in my case).

I'll allow this setup to run for some time more to see if something will change eventually.
After that I'll stop it to collect valgrind data.
It actually took so long for me to setup just because I wanted to run components under valgrind
and hit lots of setup/runtime issues.

Will report here in more details and numbers along with valgrind reports, if any.

For now I didn't see the continuous growing of memory consumption by ovn-northd reported here
that is still very suspicious.  After this test I'm going to run again with exact version of
OVN packages. This test was with upstream OVN.


Joe, I see your reports that NB DB goes up to 13GB of ram.  How many logical switch ports this
setup had while pods still was there?

Comment 16 Joe Talerico 2020-05-08 16:36:15 UTC
(In reply to Ilya Maximets from comment #15)
> I think, I managed to reproduce the issue with databases and northd not
> freeing the memory.
> 
> I've set up fake cluster with 100 nodes and 500 ports, than I removed all
> 500 VIFs from OVS
> instances and removed all the corresponding logical switch ports from NB db.
> After removing
> VIFs memory consumption of ovn-northd almost doubled, memory of 
> ovsdb-servers for NB and SB
> DBs slightly increased.  After removing logical switch ports for NB db
> memory consumption
> didn't go down and still almost same values.
> 
> ovsdb-server/compact didn't change the amount of consumed virtual memory,
> but it significantly
> reduced the size of NB database on the disk (10 times in my case).
> 
> I'll allow this setup to run for some time more to see if something will
> change eventually.
> After that I'll stop it to collect valgrind data.
> It actually took so long for me to setup just because I wanted to run
> components under valgrind
> and hit lots of setup/runtime issues.
> 
> Will report here in more details and numbers along with valgrind reports, if
> any.
> 
> For now I didn't see the continuous growing of memory consumption by
> ovn-northd reported here

Interesting. I ran a small test again on the latest 4.5 nightly I could get my hands on, and easily reproduced the memory growth issue. https://snapshot.raintank.io/dashboard/snapshot/kiQweW8RpsM26yDZdGOTniK3D59zPFW4?orgId=2

This env currently has : 

root@ip-172-31-68-73: ~/dittybopper # oc get nodes | grep Read | wc -l
31
root@ip-172-31-68-73: ~/dittybopper # oc get pods -A | wc -l
445


> that is still very suspicious.  After this test I'm going to run again with
> exact version of
> OVN packages. This test was with upstream OVN.
> 
> 
> Joe, I see your reports that NB DB goes up to 13GB of ram.  How many logical
> switch ports this
> setup had while pods still was there?


I am not sure how many I had at that time. However from the current environment, where I have 31 nodes, 445 pods :

sh-4.2# export OVN_NB_DB=ssl:10.0.143.150:9641,ssl:10.0.149.197:9641,ssl:10.0.175.79:9641
sh-4.2# ovn-nbctl -p /ovn-cert/tls.key -c /ovn-cert/tls.crt -C /ovn-ca/ca-bundle.crt ls-list | wc -l
93

Comment 17 Dan Williams 2020-05-11 17:05:43 UTC
(In reply to Joe Talerico from comment #16)
> (In reply to Ilya Maximets from comment #15)
> > I think, I managed to reproduce the issue with databases and northd not
> > freeing the memory.
> > 
> > I've set up fake cluster with 100 nodes and 500 ports, than I removed all
> > 500 VIFs from OVS
> > instances and removed all the corresponding logical switch ports from NB db.
> > After removing
> > VIFs memory consumption of ovn-northd almost doubled, memory of 
> > ovsdb-servers for NB and SB
> > DBs slightly increased.  After removing logical switch ports for NB db
> > memory consumption
> > didn't go down and still almost same values.
> > 
> > ovsdb-server/compact didn't change the amount of consumed virtual memory,
> > but it significantly
> > reduced the size of NB database on the disk (10 times in my case).
> > 
> > I'll allow this setup to run for some time more to see if something will
> > change eventually.
> > After that I'll stop it to collect valgrind data.
> > It actually took so long for me to setup just because I wanted to run
> > components under valgrind
> > and hit lots of setup/runtime issues.
> > 
> > Will report here in more details and numbers along with valgrind reports, if
> > any.
> > 
> > For now I didn't see the continuous growing of memory consumption by
> > ovn-northd reported here
> 
> Interesting. I ran a small test again on the latest 4.5 nightly I could get
> my hands on, and easily reproduced the memory growth issue.
> https://snapshot.raintank.io/dashboard/snapshot/
> kiQweW8RpsM26yDZdGOTniK3D59zPFW4?orgId=2
> 
> This env currently has : 
> 
> root@ip-172-31-68-73: ~/dittybopper # oc get nodes | grep Read | wc -l
> 31
> root@ip-172-31-68-73: ~/dittybopper # oc get pods -A | wc -l
> 445
> 
> 
> > that is still very suspicious.  After this test I'm going to run again with
> > exact version of
> > OVN packages. This test was with upstream OVN.
> > 
> > 
> > Joe, I see your reports that NB DB goes up to 13GB of ram.  How many logical
> > switch ports this
> > setup had while pods still was there?
> 
> 
> I am not sure how many I had at that time. However from the current
> environment, where I have 31 nodes, 445 pods :
> 
> sh-4.2# export
> OVN_NB_DB=ssl:10.0.143.150:9641,ssl:10.0.149.197:9641,ssl:10.0.175.79:9641
> sh-4.2# ovn-nbctl -p /ovn-cert/tls.key -c /ovn-cert/tls.crt -C
> /ovn-ca/ca-bundle.crt ls-list | wc -l
> 93

That looks about right for logical switch numbers. Could you get "lsp-list" to get logical switch ports instead?

Comment 18 Dan Williams 2020-05-12 14:06:07 UTC
Making this bug *only* about the northd issue since they will have separate fixes, and we already have one for northd.  See https://bugzilla.redhat.com/show_bug.cgi?id=1834838 for the NBDB issue.

Comment 19 Dan Williams 2020-05-12 14:07:35 UTC
Upstream fix for main ovn-northd issue is http://patchwork.ozlabs.org/project/openvswitch/patch/20200512104406.839463-1-i.maximets@ovn.org/

Comment 22 Jianlin Shi 2020-05-13 06:20:21 UTC
tried with following steps:

setup basic env:
systemctl start openvswitch
systemctl start ovn-northd   
ovn-nbctl set-connection ptcp:6641   
ovn-sbctl set-connection ptcp:6642                   
ovs-vsctl set open . external_ids:system-id=hv1 external_ids:ovn-remote=tcp:10.16.56.76:6642 external_ids:ovn-encap-type=geneve external_ids:ovn-encap-ip=10.16.56.76
systemctl restart ovn-controller
ovn-nbctl ls-add ls1
                             
ovn-nbctl lr-add lr1                 
ovn-nbctl lrp-add lr1 lr1-ls1 00:00:00:00:00:01 172.16.1.1/24
ovn-nbctl lsp-add ls1 ls1-lr1             
ovn-nbctl lsp-set-type ls1-lr1 router
ovn-nbctl lsp-set-options ls1-lr1 router-port=lr1-ls1
ovn-nbctl lsp-set-addresses ls1-lr1 router
                                                                      
ovn-nbctl lsp-add ls1 lnls1
ovn-nbctl lsp-set-options lnls1 network_name=provider
ovn-nbctl lsp-set-type lnls1 localnet                            
ovn-nbctl lsp-set-addresses lnls1 unknown    
                                                                   
ovn-nbctl set logical_router lr1 options:chassis=hv1
                         
ovn-nbctl lrp-add lr1 lr1-ls2 00:00:00:00:00:02 172.16.1.2/24
ovn-nbctl lrp-add lr1 lr1-ls3 00:00:00:00:00:03 172.16.1.3/24
ovn-nbctl ls-add ls2
ovn-nbctl lsp-add ls2 ls2-lr1                                           
ovn-nbctl lsp-set-type ls2-lr1 router                        
ovn-nbctl lsp-set-options ls2-lr1 router-port=lr1-ls2        
ovn-nbctl lsp-set-addresses ls2-lr1 router                   

ovn-nbctl ls-add ls3    
ovn-nbctl lsp-add ls3 ls3-lr1 
ovn-nbctl lsp-set-type ls3-lr1 router
ovn-nbctl lsp-set-options ls3-lr1 router-port=lr1-ls3
ovn-nbctl lsp-set-addresses ls3-lr1 router

ovs-vsctl add-br br-test                                              
ip link set br-test up                                        
ovs-vsctl set open . external-ids:ovn-bridge-mappings=provider:br-test

ip netns add server0
ip link add veth0_s0 netns server0 type veth peer name veth0_s0_p
ip netns exec server0 ip link set veth0_s0 up
ip netns exec server0 ip addr add 2001:1db8:3333::2/64 dev veth0_s0
ovs-vsctl add-port br-test veth0_s0_p
ip link set veth0_s0_p up

ip addr add 2001:1db8:3333::1/64 dev br-test

ovn-nbctl set logical_router_port lr1-ls1 options:prefix_delegation=true
ovn-nbctl set logical_router_port lr1-ls1 options:prefix=true
ovn-nbctl set logical_router_port lr1-ls2 options:prefix=true
ovn-nbctl set logical_router_port lr1-ls3 options:prefix=true

cat > dhcpd6.conf << EOF
option dhcp-rebinding-time 15;
option dhcp-renewal-time 10;
option dhcp6.unicast fe80::f455:8ff:fe20:6d66;
subnet6 2001:1db8:3333::/64 {

        # Some /64 prefixes available for Prefix Delegation (RFC 3633)
        prefix6 2001:1db8:3333:100:: 2001:1db8:3333:111:: /80;
}
EOF

ip netns exec server0 dhcpd -6 -cf ./dhcpd6.conf veth0_s0

script to add much ports:

[root@kvm-04-guest09 bz1826623]# cat add_port.sh
for num in `seq 4 250`
do
        ovn-nbctl lrp-add lr1 lr1-ls$num 00:00:00:00:00:$(printf %x $num) 172.16.1.$num/24
ovn-nbctl ls-add ls$num
ovn-nbctl lsp-add ls$num ls${num}-lr1                   
ovn-nbctl lsp-set-type ls${num}-lr1 router           
ovn-nbctl lsp-set-options ls${num}-lr1 router-port=lr1-ls$num
ovn-nbctl lsp-set-addresses ls${num}-lr1 router               
ovn-nbctl set logical_router_port lr1-ls${num} options:prefix=true
                                   
done

script to del ports:

[root@kvm-04-guest09 bz1826623]# cat del_port.sh
for num in `seq 4 250`
do
        ovn-nbctl lrp-del lr1-ls$num
ovn-nbctl ls-del ls$num
                                   
done

test on ovn2.13.0-27:

[root@kvm-04-guest09 ovn2.13.0-30]# rpm -qa | grep -E "openvswitch|ovn"
ovn2.13-2.13.0-27.el8fdp.x86_64
openvswitch-selinux-extra-policy-1.0-23.el8fdp.noarch
ovn2.13-central-2.13.0-27.el8fdp.x86_64
openvswitch2.13-2.13.0-18.el8fdp.x86_64
ovn2.13-host-2.13.0-27.el8fdp.x86_64

show VmRSS for ovn-northd after setup:

[root@kvm-04-guest09 bz1826623]# cat /var/run/ovn/ovn-northd.pid   
7174 
[root@kvm-04-guest09 bz1826623]# grep VmRSS /proc/7174/status
VmRSS:      4728 kB 

add port with add_port.sh, then check VmRSS:

[root@kvm-04-guest09 bz1826623]# ./add_port.sh
[root@kvm-04-guest09 bz1826623]# ovn-nbctl list logical_router_port lr1-ls250                                                                                                                 
_uuid               : 7306fe2d-7e95-4027-a79f-9aa7dde0d475                                                       
enabled             : []                                     
external_ids        : {}                                     
gateway_chassis     : []                                                                                                                                             
ha_chassis_group    : []                                          
ipv6_prefix         : ["2001:1db8:3333:110:afb0::/80"]       
ipv6_ra_configs     : {}           
mac                 : "00:00:00:00:00:fa"            
name                : lr1-ls250                              
networks            : ["172.16.1.250/24"]            
options             : {prefix="true"}         
peer                : []

[root@kvm-04-guest09 bz1826623]# grep VmRSS /proc/7174/status         
VmRSS:    207440 kB

<==== memory used increased much

del port with del_port.sh, then check VmRSS:

[root@kvm-04-guest09 bz1826623]# ./del_port.sh                
[root@kvm-04-guest09 bz1826623]# grep VmRSS /proc/7174/status         
VmRSS:    228872 kB


tested on ovn2.13.0-30:

[root@kvm-04-guest09 ovn2.13.0-30]# rpm -qa | grep -E "openvswitch|ovn"
ovn2.13-2.13.0-30.el8fdp.x86_64
openvswitch-selinux-extra-policy-1.0-23.el8fdp.noarch
ovn2.13-central-2.13.0-30.el8fdp.x86_64
openvswitch2.13-2.13.0-18.el8fdp.x86_64
ovn2.13-host-2.13.0-30.el8fdp.x86_64

VmRSS after setup:

[root@kvm-04-guest09 bz1826623]# grep VmRSS /proc/10941/status 
VmRSS:      4236 kB

VmRSS after add port:

[root@kvm-04-guest09 bz1826623]# ./add_port.sh 
[root@kvm-04-guest09 bz1826623]# ovn-nbctl list logical_router_port lr1-ls250
_uuid               : d9dd0f5f-b569-4d4d-aa52-e5d54ed89390
enabled             : []
external_ids        : {}
gateway_chassis     : []
ha_chassis_group    : []
ipv6_prefix         : ["2001:1db8:3333:110:59be::/80"]
ipv6_ra_configs     : {}
mac                 : "00:00:00:00:00:fa"
name                : lr1-ls250
networks            : ["172.16.1.250/24"]
options             : {prefix="true"}
peer                : []
[root@kvm-04-guest09 bz1826623]# grep VmRSS /proc/10941/status 
VmRSS:     45268 kB

<=== less memory increased compared to 207440 kB

VmRSS after delete ports:

[root@kvm-04-guest09 bz1826623]# ./del_port.sh 
[root@kvm-04-guest09 bz1826623]# grep VmRSS /proc/10941/status 
VmRSS:     45268 kB


from above result, less memory is increased after add much ports on ovn2.13.0-30

Comment 23 Jianlin Shi 2020-05-13 06:38:58 UTC
result on rhel7 version.

on ovn2.13.0-27.e7:

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# rpm -qa | grep -E "openvswitch|ovn"
openvswitch2.13-2.13.0-17.el7fdp.x86_64
ovn2.13-central-2.13.0-27.el7fdp.x86_64
openvswitch-selinux-extra-policy-1.0-15.el7fdp.noarch
ovn2.13-2.13.0-27.el7fdp.x86_64
ovn2.13-host-2.13.0-27.el7fdp.x86_64

after setup:
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ovn-nbctl list logical_router_port lr1-ls3
_uuid               : 3539eff0-6add-47bf-9d71-9fbe25534409
enabled             : []
external_ids        : {}
gateway_chassis     : []
ha_chassis_group    : []
ipv6_prefix         : ["2001:1db8:3333:110:bebf::/80"]
ipv6_ra_configs     : {}
mac                 : "00:00:00:00:00:03"
name                : lr1-ls3
networks            : ["172.16.1.3/24"]
options             : {prefix="true"}
peer                : []
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# grep VmRSS /proc/26730/status 
VmRSS:      2904 kB

after add ports:

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ./add_port.sh 
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ovn-nbctl list logical_router_port lr1-ls250
_uuid               : b4e7fa81-f880-4888-b795-d7e81828b2a5
enabled             : []
external_ids        : {}
gateway_chassis     : []
ha_chassis_group    : []
ipv6_prefix         : ["2001:1db8:3333:110:9c0c::/80"]
ipv6_ra_configs     : {}
mac                 : "00:00:00:00:00:fa"
name                : lr1-ls250
networks            : ["172.16.1.250/24"]
options             : {prefix="true"}
peer                : []
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# grep VmRSS /proc/26730/status 
VmRSS:    140520 kB


after del ports:

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ./del_port.sh 
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# grep VmRSS /proc/26730/status 
VmRSS:    164444 kB


result on ovn2.13.0-30.el7:

VmRSS after setup:

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ovn-nbctl list logical_router_port lr1-ls3
_uuid               : 17718adf-98a4-425b-a263-5336e50209ba
enabled             : []
external_ids        : {}
gateway_chassis     : []
ha_chassis_group    : []
ipv6_prefix         : ["2001:1db8:3333:110:bebf::/80"]
ipv6_ra_configs     : {}
mac                 : "00:00:00:00:00:03"
name                : lr1-ls3
networks            : ["172.16.1.3/24"]
options             : {prefix="true"}
peer                : []

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# cat /var/run/ovn/ovn-northd.pid 
29691
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# grep VmRSS /proc/29691/status
VmRSS:      2768 kB

VmRSS after add ports:

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ./add_port.sh 
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# grep VmRSS /proc/29691/status
VmRSS:     36872 kB

VmRSS after delete ports:

[root@hpe-dl380pgen8-02-vm-13 bz1826623]# ./del_port.sh 
[root@hpe-dl380pgen8-02-vm-13 bz1826623]# grep VmRSS /proc/29691/status
VmRSS:     36876 kB


after add ports, VmRSS 36872 on ovn2.13.0-30 is much less than 140520 on ovn2.13.0-27.

Comment 24 Jianlin Shi 2020-05-13 06:40:51 UTC
Hi Ilya,

can comment 22 and comment 23 prove that the issue is fixed?

Comment 25 Ilya Maximets 2020-05-13 10:14:57 UTC
(In reply to Jianlin Shi from comment #24)
> Hi Ilya,
> 
> can comment 22 and comment 23 prove that the issue is fixed?

Yes.  Since memory usage decreased with the new build and also it doesn't grow after removing ports.

Comment 26 Jianlin Shi 2020-05-14 00:18:54 UTC
set VERIFIED per comment 25

Comment 30 errata-xmlrpc 2020-05-26 14:07:18 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2020:2317