Bug 1720564
| Summary: | [Ganesha] Ganesha occupies increased memory when IO's and lookups are running in parallel and consumption is same post test completion | ||
|---|---|---|---|
| Product: | [Red Hat Storage] Red Hat Gluster Storage | Reporter: | Manisha Saini <msaini> |
| Component: | nfs-ganesha | Assignee: | Daniel Gryniewicz <dang> |
| Status: | CLOSED ERRATA | QA Contact: | Manisha Saini <msaini> |
| Severity: | high | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | rhgs-3.5 | CC: | dang, ffilz, grajoria, jthottan, mbenjamin, rhs-bugs, skoduri, storage-qa-internal, vdas |
| Target Milestone: | --- | ||
| Target Release: | RHGS 3.5.0 | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | nfs-ganesha-2.7.3-5 | Doc Type: | No Doc Update |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2019-10-30 12:15:39 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | |||
| Bug Blocks: | 1696809 | ||
|
Description
Manisha Saini
2019-06-14 08:43:39 UTC
Maybe this: https://github.com/nfs-ganesha/nfs-ganesha/commit/136df4f262c3f9bc29df456eac50921321826da3 Deleted all files and dirs from mount point.Still the memory consumption on ideal setup is 8.4G and is same from past 6 hrs... ---------- Tasks: 1 total, 0 running, 1 sleeping, 0 stopped, 0 zombie %Cpu(s): 0.0 us, 0.0 sy, 0.0 ni,100.0 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st KiB Mem : 19646235+total, 17871353+free, 14047236 used, 3701576 buff/cache KiB Swap: 4194300 total, 4194300 free, 0 used. 18156003+avail Mem PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND 4376 root 20 0 13.9g 8.4g 7152 S 0.0 4.5 1026:51 ganesha.nfsd ------------- Attempt 2:
=========
Ganesha process took over 30.5G memory while performing following workload-
1. Create 4 node Ganesha cluster
2. Create 12*3 Distributed-Replicate Volume
3. Mount the volume on 4 clients
4. Run the following worload-
Client 1:Run Crefy which does following IO operation
for i in {create,chmod,hardlink,chgrp,symlink,hardlink,truncate,hardlink}; do crefi --multi -n 1000 -b 10 -d 10 --max=10K --min=500 --random -T 10 -t text --fop=$i /mnt/ganesha/dir1 ; sleep 5 ; done
Client 2: Run bonnie
Client 3: Run dbench
Client 4: Run ls -lRt in loop
Observation: Triggered this test over weekend.Stopped the IO's on sunday Night and left the setup idle in same state.On having look into system on Monday,observed Ganesha memory consumption was still same i.e 30.5G. This memory keep on increasing if I again trigger some IO's and lookups.
There were 5910 directories, 5000000 files created on mount point as part of this test
I will be attaching statedumps for ganesha process (taken at 3 different intervals) and sosreport for this run shortly.
Further:
==========
I will be deleting all files and dirs from mount point today overnight and monitor memory consumption post deleting content.Will update this bug by tomorrow with further observation
(In reply to Manisha Saini from comment #5) > Attempt 2: > ========= > > Ganesha process took over 30.5G memory while performing following workload- > > 1. Create 4 node Ganesha cluster > 2. Create 12*3 Distributed-Replicate Volume > 3. Mount the volume on 4 clients > 4. Run the following worload- > > Client 1:Run Crefy which does following IO operation > > for i in {create,chmod,hardlink,chgrp,symlink,hardlink,truncate,hardlink}; > do crefi --multi -n 1000 -b 10 -d 10 --max=10K --min=500 --random -T 10 -t > text --fop=$i /mnt/ganesha/dir1 ; sleep 5 ; done > > Client 2: Run bonnie > Client 3: Run dbench > Client 4: Run ls -lRt in loop > > Observation: Triggered this test over weekend.Stopped the IO's on sunday > Night and left the setup idle in same state.On having look into system on > Monday,observed Ganesha memory consumption was still same i.e 30.5G. This > memory keep on increasing if I again trigger some IO's and lookups. > > There were 5910 directories, 5000000 files created on mount point as part of > this test > > I will be attaching statedumps for ganesha process (taken at 3 different > intervals) and sosreport for this run shortly. > > Further: > ========== > > I will be deleting all files and dirs from mount point today overnight and > monitor memory consumption post deleting content.Will update this bug by > tomorrow with further observation Deleted all the contents from mount point overnight and kept the setup in idle state to monitor memory utilization. I see memory consumption after 6 hours is- 25.0G (which is way too high) ---- Tasks: 1 total, 0 running, 1 sleeping, 0 stopped, 0 zombie %Cpu(s): 0.0 us, 0.0 sy, 0.0 ni, 99.9 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st KiB Mem : 19646235+total, 14792275+free, 39134500 used, 9405096 buff/cache KiB Swap: 4194300 total, 4194300 free, 0 used. 15632041+avail Mem PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND 37983 root 20 0 31.8g 25.0g 7276 S 0.3 13.3 1044:03 ganesha.nfsd ---- It would be interesting to see valgrind massif reports. In the past those have shown that Ganesha consumes just a small portion of the memory footprint. Also, as we discussed last year when we all came to BLR, the memory allocators do not necessarily return memory to the system, so the memory reports will continue to show a large footprint. (In reply to Frank Filz from comment #9) > It would be interesting to see valgrind massif reports. In the past those > have shown that Ganesha consumes just a small portion of the memory > footprint. Also, as we discussed last year when we all came to BLR, the > memory allocators do not necessarily return memory to the system, so the > memory reports will continue to show a large footprint. I was trying to run ganesha with valgrind massif for ganesha, but when I try to do some operations on the mount point valgrind exits with the following error : valgrind: m_execontext.c:411 (record_ExeContext_wrk2): Assertion 'n_ips >= 1 && n_ips <= VG_(clo_backtrace_size)' failed. Following is the command I have used for running valgrind valgrind --log-file=/tmp/vg.out --tool=massif --max-stackframe=16777216 `which ganesha.nfsd` -C -F -f /etc/ganesha/ganesha.conf -L /var/log/ganesha.log -N NIV_INFO Don't know whether it is a configuration issue for my command or genuine crash in ganesha. All the threads in the output file look fine for me. So I was running the test with Linux untar in a loop (without valgrind), and memory usage reached upto 7G over 3 days. p lru_state $16 = {entries_hiwat = 125000, entries_used = 1259305, chunks_hiwat = 100000, chunks_used = 70726, fds_system_imposed = 4096, fds_hard_limit = 4055, fds_hiwat = 3686, fds_lowat = 2048, futility = 0, per_lane_work = 50, biggest_window = 1638, prev_fd_count = 0, prev_time = 1560943711, fd_state = 0} in the live ganesha process, the entries_used is crossed entries_hiwat(more than 10 times) and is still increasing. I saw(Soumya mentioned) upstream user reported similar and Dan suggested the same fix mentioned in this bug. So I triggered a new run including Dan's fix today.(I still didn't understand how this fix solves the issue which we are seeing "creating file increasing memory consumption) And one more thing I checked the last statedump shared by manisha post deleting the files. The mem usage decreased to MBs(below 25) from GB's. But still ganesha overall consumption was the same. So I am suspecting the leak is from ganesha layer. -- Jiffin That commit fixed a bug where readdir would leave entries refcounted causing them to never be reaped. I've only personally run massif on ganesha a few times. Your massif command line looks fine to me, and I've never seen that valgrind error, so I have no idea what it means. 1259305 entries is 1.4 GB. So it's a lot, but it's a very long way from 7G. The massif invocation I've used is: valgrind --log-file=/tmp/vg.out --tool=massif `which ganesha.nfsd` -F -f /etc/ganesha/top21.conf -L /var/log/ganesha -N NIV_INFO I've not see an error like you have. When I did do massif analysis before, Ganesha used a fraction of the memory compared to Gluster. Changing the state to POST since fix is available in upstream # rpm -qa | grep ganesha nfs-ganesha-2.7.3-7.el7rhgs.x86_64 glusterfs-ganesha-6.0-11.el7rhgs.x86_64 nfs-ganesha-gluster-2.7.3-7.el7rhgs.x86_64 Steps to Reproduce: ============== 1.Create 8 node ganesha cluster 2.Create 8*3 Distributed-Replicate Volume 3.Export the volume via ganesha 4.Mount the volume on 5 clients via v4.1 5.Run the following workload Client 1: Linux untars for large dirs Client 2: du -sh in loop Client 3: ls -lRt in loop Client 4: find . -mindepth 1 -type f -name _04_* in loop Client 5: find . -mindepth 1 -type f in loop ============= Tasks: 1 total, 0 running, 1 sleeping, 0 stopped, 0 zombie %Cpu(s): 5.2 us, 4.6 sy, 0.0 ni, 90.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st KiB Mem : 65758176 total, 49721856 free, 5234248 used, 10802072 buff/cache KiB Swap: 32964604 total, 32964604 free, 0 used. 59847572 avail Mem PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND 6127 root 20 0 6578724 1.2g 6632 S 120.0 1.9 12812:55 ganesha.nfsd ============= No increased memory consumption was observed during and post test completion.Moving tis BZ to verified state. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHEA-2019:3252 |