Bug 1720564 - [Ganesha] Ganesha occupies increased memory when IO's and lookups are running in parallel and consumption is same post test completion
Summary: [Ganesha] Ganesha occupies increased memory when IO's and lookups are running...
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat Gluster Storage
Classification: Red Hat Storage
Component: nfs-ganesha
Version: rhgs-3.5
Hardware: Unspecified
OS: Unspecified
unspecified
high
Target Milestone: ---
: RHGS 3.5.0
Assignee: Daniel Gryniewicz
QA Contact: Manisha Saini
URL:
Whiteboard:
Depends On:
Blocks: 1696809
TreeView+ depends on / blocked
 
Reported: 2019-06-14 08:43 UTC by Manisha Saini
Modified: 2019-10-30 12:15 UTC (History)
9 users (show)

Fixed In Version: nfs-ganesha-2.7.3-5
Doc Type: No Doc Update
Doc Text:
Clone Of:
Environment:
Last Closed: 2019-10-30 12:15:39 UTC
Embargoed:


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Red Hat Product Errata RHEA-2019:3252 0 None None None 2019-10-30 12:15:56 UTC

Description Manisha Saini 2019-06-14 08:43:39 UTC
Description of problem:

High memory consumption for Ganesha process is observed (RES ~9G) when large dirs untars and lookups are running from multiple clients. The memory is not getting released and is same even when test is completed.

-----------
Tasks:   1 total,   0 running,   1 sleeping,   0 stopped,   0 zombie
%Cpu(s):  0.0 us,  0.0 sy,  0.0 ni,100.0 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
KiB Mem : 19646235+total, 16030555+free, 15752392 used, 20404408 buff/cache
KiB Swap:  4194300 total,  4194300 free,        0 used. 17988723+avail Mem 

   PID USER      PR  NI    VIRT    RES    SHR S  %CPU %MEM     TIME+ COMMAND                                                                    
  4376 root      20   0   14.6g   9.0g   7116 S   0.0  4.8 995:23.00 ganesha.nfsd   

-----------

Version-Release number of selected component (if applicable):

# rpm -qa | grep ganesha
nfs-ganesha-gluster-2.7.3-4.el7rhgs.x86_64
nfs-ganesha-debuginfo-2.7.3-4.el7rhgs.x86_64
glusterfs-ganesha-6.0-5.el7rhgs.x86_64
nfs-ganesha-2.7.3-4.el7rhgs.x86_64

# cat /etc/redhat-release 
Red Hat Enterprise Linux Server release 7.7 Beta (Maipo)


How reproducible:
2/2


Steps to Reproduce:
1.Create 4 node Ganesha cluster
2.Create 12*3 Distributed-Replicate Volume
3.Export the volume via Ganesha
4.Mount the volume on 4 nfs clients via v4.1
5.Run the following workload
Client 1: Large dirs untars
Client 2: du -sh in loop
Client 3: ls -lRt in loop
Client 4: find . -mindepth 1 -type f -name _04_* in loop 

Actual results:

Even when the test is completed and lookups are stopped from other clients,Memory consumption still observed high i.e 9G on ideal setup


Expected results:
The memory consumption for Ganesha process should not be high

Additional info:

Post test is completed,total of 102 directories, 1145134 files were created on mount point as part of linux untars.
Total storage consumption was 18GB.

Comment 4 Manisha Saini 2019-06-14 19:40:07 UTC
Deleted all files and dirs from mount point.Still the memory consumption on ideal setup is 8.4G and is same from past 6 hrs...

----------
Tasks:   1 total,   0 running,   1 sleeping,   0 stopped,   0 zombie
%Cpu(s):  0.0 us,  0.0 sy,  0.0 ni,100.0 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
KiB Mem : 19646235+total, 17871353+free, 14047236 used,  3701576 buff/cache
KiB Swap:  4194300 total,  4194300 free,        0 used. 18156003+avail Mem 

   PID USER      PR  NI    VIRT    RES    SHR S  %CPU %MEM     TIME+ COMMAND                                                                   
  4376 root      20   0   13.9g   8.4g   7152 S   0.0  4.5   1026:51 ganesha.nfsd    
-------------

Comment 5 Manisha Saini 2019-06-17 17:57:52 UTC
Attempt 2:
========= 

Ganesha process took over 30.5G memory while performing following workload-

1. Create 4 node Ganesha cluster
2. Create 12*3 Distributed-Replicate Volume
3. Mount the volume on 4 clients
4. Run the following worload-

Client 1:Run Crefy which does following IO operation

for i in {create,chmod,hardlink,chgrp,symlink,hardlink,truncate,hardlink}; do crefi --multi -n 1000 -b 10 -d 10 --max=10K --min=500 --random -T 10 -t text --fop=$i /mnt/ganesha/dir1 ; sleep 5 ; done

Client 2: Run bonnie
Client 3: Run dbench
Client 4: Run ls -lRt in loop

Observation: Triggered this test over weekend.Stopped the IO's on sunday Night and left the setup idle in same state.On having look into system on Monday,observed Ganesha memory consumption was still same i.e 30.5G. This memory keep on increasing if I again trigger some IO's and lookups.

There were 5910 directories, 5000000 files created on mount point as part of this test

I will be attaching statedumps for ganesha process (taken at 3 different intervals) and sosreport for this run shortly.

Further:
==========

I will be deleting all files and dirs from mount point today overnight and monitor memory consumption post deleting content.Will update this bug by tomorrow with further observation

Comment 8 Manisha Saini 2019-06-18 06:21:35 UTC
(In reply to Manisha Saini from comment #5)
> Attempt 2:
> ========= 
> 
> Ganesha process took over 30.5G memory while performing following workload-
> 
> 1. Create 4 node Ganesha cluster
> 2. Create 12*3 Distributed-Replicate Volume
> 3. Mount the volume on 4 clients
> 4. Run the following worload-
> 
> Client 1:Run Crefy which does following IO operation
> 
> for i in {create,chmod,hardlink,chgrp,symlink,hardlink,truncate,hardlink};
> do crefi --multi -n 1000 -b 10 -d 10 --max=10K --min=500 --random -T 10 -t
> text --fop=$i /mnt/ganesha/dir1 ; sleep 5 ; done
> 
> Client 2: Run bonnie
> Client 3: Run dbench
> Client 4: Run ls -lRt in loop
> 
> Observation: Triggered this test over weekend.Stopped the IO's on sunday
> Night and left the setup idle in same state.On having look into system on
> Monday,observed Ganesha memory consumption was still same i.e 30.5G. This
> memory keep on increasing if I again trigger some IO's and lookups.
> 
> There were 5910 directories, 5000000 files created on mount point as part of
> this test
> 
> I will be attaching statedumps for ganesha process (taken at 3 different
> intervals) and sosreport for this run shortly.
> 
> Further:
> ==========
> 
> I will be deleting all files and dirs from mount point today overnight and
> monitor memory consumption post deleting content.Will update this bug by
> tomorrow with further observation


Deleted all the contents from mount point overnight and kept the setup in idle state to monitor memory utilization.

I see memory consumption after 6 hours is- 25.0G (which is way too high)

----
Tasks:   1 total,   0 running,   1 sleeping,   0 stopped,   0 zombie
%Cpu(s):  0.0 us,  0.0 sy,  0.0 ni, 99.9 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
KiB Mem : 19646235+total, 14792275+free, 39134500 used,  9405096 buff/cache
KiB Swap:  4194300 total,  4194300 free,        0 used. 15632041+avail Mem 

   PID USER      PR  NI    VIRT    RES    SHR S  %CPU %MEM     TIME+ COMMAND                                                                                                   
 37983 root      20   0   31.8g  25.0g   7276 S   0.3 13.3   1044:03 ganesha.nfsd   
----

Comment 9 Frank Filz 2019-06-18 15:16:05 UTC
It would be interesting to see valgrind massif reports. In the past those have shown that Ganesha consumes just a small portion of the memory footprint. Also, as we discussed last year when we all came to BLR, the memory allocators do not necessarily return memory to the system, so the memory reports will continue to show a large footprint.

Comment 10 Jiffin 2019-06-19 13:05:29 UTC
(In reply to Frank Filz from comment #9)
> It would be interesting to see valgrind massif reports. In the past those
> have shown that Ganesha consumes just a small portion of the memory
> footprint. Also, as we discussed last year when we all came to BLR, the
> memory allocators do not necessarily return memory to the system, so the
> memory reports will continue to show a large footprint.

I was trying to run ganesha with valgrind massif for ganesha, but when I try to do some operations on the mount point valgrind exits with the following error : 
valgrind: m_execontext.c:411 (record_ExeContext_wrk2): Assertion 'n_ips >= 1 && n_ips <= VG_(clo_backtrace_size)' failed. 
Following is the command I have used for running valgrind
valgrind  --log-file=/tmp/vg.out --tool=massif --max-stackframe=16777216 `which ganesha.nfsd` -C -F  -f /etc/ganesha/ganesha.conf -L /var/log/ganesha.log -N NIV_INFO

Don't know whether it is a configuration issue for my command or genuine crash in ganesha. All the threads in the output file look fine for me.

So I was running the test with Linux untar in a loop (without valgrind), and memory usage reached upto 7G over 3 days.

p lru_state
$16 = {entries_hiwat = 125000, entries_used = 1259305, chunks_hiwat = 100000, chunks_used = 70726, 
  fds_system_imposed = 4096, fds_hard_limit = 4055, fds_hiwat = 3686, fds_lowat = 2048, futility = 0, 
  per_lane_work = 50, biggest_window = 1638, prev_fd_count = 0, prev_time = 1560943711, fd_state = 0}

in the live ganesha process, the entries_used is crossed entries_hiwat(more than 10 times) and is still increasing. I saw(Soumya mentioned) upstream user reported similar and Dan suggested the same fix mentioned in this bug. So I triggered a new run including Dan's fix today.(I still didn't understand how this fix solves the issue which we are seeing "creating file increasing memory consumption)

And one more thing I checked the last statedump shared by manisha post deleting the files. The mem usage decreased to MBs(below 25) from GB's. But still ganesha overall consumption was the same. So I am suspecting the leak is from ganesha layer.

--

Jiffin

Comment 11 Daniel Gryniewicz 2019-06-19 14:04:54 UTC
That commit fixed a bug where readdir would leave entries refcounted causing them to never be reaped.

I've only personally run massif on ganesha a few times.  Your massif command line looks fine to me, and I've never seen that valgrind error, so I have no idea what it means.

1259305 entries is 1.4 GB.  So it's a lot, but it's a very long way from 7G.

Comment 12 Frank Filz 2019-06-19 17:00:36 UTC
The massif invocation I've used is:

valgrind  --log-file=/tmp/vg.out --tool=massif `which ganesha.nfsd` -F  -f /etc/ganesha/top21.conf -L /var/log/ganesha -N NIV_INFO

I've not see an error like you have.

When I did do massif analysis before, Ganesha used a fraction of the memory compared to Gluster.

Comment 15 Jiffin 2019-06-24 09:17:31 UTC
Changing the state to POST since fix is available in upstream

Comment 18 Manisha Saini 2019-08-26 01:48:42 UTC
# rpm -qa | grep ganesha
nfs-ganesha-2.7.3-7.el7rhgs.x86_64
glusterfs-ganesha-6.0-11.el7rhgs.x86_64
nfs-ganesha-gluster-2.7.3-7.el7rhgs.x86_64


Steps to Reproduce:
==============
1.Create 8 node ganesha cluster
2.Create 8*3 Distributed-Replicate Volume
3.Export the volume via ganesha
4.Mount the volume on 5 clients via v4.1
5.Run the following workload
Client 1: Linux untars for large dirs
Client 2: du -sh in loop
Client 3: ls -lRt in loop
Client 4: find . -mindepth 1 -type f -name _04_* in loop
Client 5:  find . -mindepth 1 -type f in loop

=============
Tasks:   1 total,   0 running,   1 sleeping,   0 stopped,   0 zombie
%Cpu(s):  5.2 us,  4.6 sy,  0.0 ni, 90.2 id,  0.0 wa,  0.0 hi,  0.0 si,  0.0 st
KiB Mem : 65758176 total, 49721856 free,  5234248 used, 10802072 buff/cache
KiB Swap: 32964604 total, 32964604 free,        0 used. 59847572 avail Mem 

  PID USER      PR  NI    VIRT    RES    SHR S  %CPU %MEM     TIME+ COMMAND                                                                                  
 6127 root      20   0 6578724   1.2g   6632 S 120.0  1.9  12812:55 ganesha.nfsd    

=============

No increased memory consumption was observed during and post test completion.Moving tis BZ to verified state.

Comment 20 errata-xmlrpc 2019-10-30 12:15:39 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHEA-2019:3252


Note You need to log in before you can comment on or make changes to this bug.