Bug 2042210 - Redis pod crashing with OOMKilled in ACM 2.4 running on OCP 4.8
Summary: Redis pod crashing with OOMKilled in ACM 2.4 running on OCP 4.8
Keywords:
Status: CLOSED NEXTRELEASE
Alias: None
Product: Red Hat Advanced Cluster Management for Kubernetes
Classification: Red Hat
Component: Search / Analytics
Version: rhacm-2.4
Hardware: Unspecified
OS: Unspecified
urgent
high
Target Milestone: ---
: rhacm-2.4.5
Assignee: Jorge Padilla
QA Contact: Xiang Yin
Mikela Dockery
URL:
Whiteboard:
: 2047273 2052910 2138423 (view as bug list)
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2022-01-19 02:30 UTC by Nikhil Gupta
Modified: 2025-10-03 11:31 UTC (History)
12 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2022-09-02 15:03:19 UTC
Target Upstream Version:
Embargoed:
ashafi: qe_test_coverage-
bot-tracker-sync: rhacm-2.4.z+


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Github stolostron backlog issues 19228 0 None None None 2022-01-19 05:16:21 UTC
Red Hat Knowledge Base (Solution) 6671301 0 None None None 2022-02-03 11:04:56 UTC

Description Nikhil Gupta 2022-01-19 02:30:22 UTC
Description of the problem:

Redis pod crashing with OOMKilled in ACM 2.4 running on OCP 4.8 and getting following error on the ACM Overview page:
~~~
An unexpected error occurred. Search service is unavailable
~~~

Following are the findings from the pod logs:
- search-prod-c0b92-search-aggregator-6c9c9fc4f7-m7dm7
~~~
2022-01-17T20:40:41.553464795Z E0117 20:40:41.553010       1 redisinterfacev2.go:47] Error fetching results from RedisGraph V2 : LOADING Redis is loading the dataset in memory
2022-01-17T20:40:41.553464795Z W0117 20:40:41.553037       1 clusterWatch.go:150] Error on UpdateByName() LOADING Redis is loading the dataset in memory
~~~

- search-prod-c0b92-search-api-864784d684-gpz7t
~~~
2022-01-17T20:34:21.196934589Z [2022-01-17T20:34:21.195] [ERROR] [search-api] [server] Unable to resolve search request because RedisGraph is unavailable.
2022-01-17T20:34:56.468684530Z [2022-01-17T20:34:56.468] [INFO] [search-api] [server] Role configuration has changed. User RBAC cache has been deleted

2022-01-17T20:38:23.659061144Z [2022-01-17T20:38:23.658] [INFO] [search-api] [server] The Redis connection has ended.
2022-01-17T20:38:23.659338130Z [2022-01-17T20:38:23.659] [INFO] [search-api] [server] Error with Redis connection.
2022-01-17T20:38:23.864374597Z [2022-01-17T20:38:23.864] [INFO] [search-api] [server] Error with Redis connection.
2022-01-17T20:38:24.212032608Z [2022-01-17T20:38:24.211] [INFO] [search-api] [server] Error with Redis connection.
2022-01-17T20:38:31.913318498Z [2022-01-17T20:38:31.912] [INFO] [search-api] [server] Redis Client connected.
~~~

- search-prod-c0b92-search-collector-869545f6bd-cxht6
~~~
2022-01-17T20:30:46.888225546Z W0117 20:30:46.888117       1 sender.go:169] Received error response [POST to: https://search-aggregator.open-cluster-management.svc:3010/aggregator/clusters/local-cluster/sync responded with error. StatusCode: 400  Message: 400 Bad Request] from Aggregator. Resending in 600000 ms after resetting config.
~~~

- search-redisgraph-0
~~~
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG6[16]: Peer certificate not required
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[16]: TLS state (accept): before SSL initialization
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG3[16]: SSL_accept: Peer suddenly disconnected
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG5[16]: Connection reset: 0 byte(s) sent to TLS, 0 byte(s) sent to socket
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[16]: Local descriptor (FD=14) closed
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[16]: Service [redis] finished (3 left)
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[ui]: FD=8 events=0x2001 revents=0x1
2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[ui]: Service [redis] accepted (FD=14) from ::ffff:10.59.2.1:55500
~~~

- search-ui-85f85bcbd7-5d59n
~~~
2022-01-17T18:30:53.691183884Z {"level":"info","time":1642444253688,"instance":1,"msg":"Not Modified","status":304,"method":"GE
T","url":"/search/static/js/0.9aea736a.chunk.js","ms":8.56}
2022-01-17T18:30:53.691514710Z {"level":"info","time":1642444253689,"instance":1,"msg":"Not Modified","status":304,"method":"GE
T","url":"/search/static/js/7.479fc9f4.chunk.js","ms":8.3}
2022-01-17T18:30:53.709557785Z {"level":"info","time":1642444253708,"instance":1,"msg":"Not Modified","status":304,"method":"GE
T","url":"/search/static/media/RHACM-Logo.8a59e4b5.svg","ms":0.58}
2022-01-17T18:30:53.950240013Z {"level":"info","time":1642444253950,"instance":1,"msg":"Not Modified","status":304,"method":"GE
T","url":"/search/static/media/RedHatText-Regular.4a43a00f.woff2","ms":1.14}
2022-01-17T18:30:53.951709016Z {"level":"info","time":1642444253951,"instance":1,"msg":"Not Modified","status":304,"method":"GET","url":"/search/static/media/RedHatText-Medium.0227c8bb.woff2","ms":0.44}
2022-01-17T18:37:03.861858408Z {"level":"warn","time":1642444623861,"instance":1,"msg":"Not Found","status":404,"method":"GET","url":"/search/locales/en/translation.json","ms":1.31}
~~~

Release version:
ACM 2.4.1

OCP version:
OCP 4.8.25

Steps to reproduce:
1. Install ACM 2.4.1 on OCP 4.8.25
2. Enable Search operator

Actual results: 
The `search-redisgraph-0` pod is running out of memory and getting terminated

~~~
$ omg get pod search-redisgraph-0 -o yaml
apiVersion: v1
kind: Pod
metadata:
  annotations:
....
    resources:
      limits:
	memory: 4Gi
      requests:
	cpu: 25m
	memory: 128Mi
....
    lastState:
      terminated:
	containerID: cri-o://13188df8f22de6e1302ee3ac80aba22035827909c3d9c71c72390efef909ad36
	exitCode: 137
	finishedAt: '2022-01-17T20:38:23Z'
	reason: OOMKilled
	startedAt: '2022-01-17T20:27:22Z'
    name: redisgraph
    ready: true
    restartCount: 361
~~~

Expected results:
The `search-redisgraph-0` pod should be running.

Additional info:
The error is still there even after the pod restart.

Comment 3 Jorge Padilla 2022-01-20 21:45:51 UTC
Thanks for the detailed description. The root cause is the Redis pod running out of memory. This topic explains how to increase the memory limits.
https://access.redhat.com/documentation/en-us/red_hat_advanced_cluster_management_for_kubernetes/2.4/html/web_console/web-console#options-increase-memory

The connection errors on the other search-* pods should recover once the Redis pod stabilizes, but you could also manually restart the pods to reestablish the connection.

Comment 4 Felix Dewaleyne 2022-01-24 10:35:36 UTC
expending the memory limit wasn't enough, the connection errors continued as reported by the customer (raising to 10gb and then 16gb).
could there be other changes to make?

no logs were provided, so I did ask for a new set.

Comment 5 Felix Dewaleyne 2022-01-26 10:20:24 UTC
it turns out that the PV what the Redis is using was full and there were no space left. By default, ACM installed with 10GB PV, and that was full after few months of running it. A new volume was created (the customer didn't see how to extend it dynamically), used different storage class and increased it to 50GB. no more crashes.

do we have a better process I could use to create a kbase article, and how much storage would be require for a setup where ACM manags between 20 to 40 clusters?

Comment 6 Jorge Padilla 2022-01-27 20:01:13 UTC
To prevent this problem from happening, the search operator should check that the PVC is larger than the memory limit.
We should also update this topic. https://access.redhat.com/documentation/en-us/red_hat_advanced_cluster_management_for_kubernetes/2.4/html/web_console/web-console#options-increase-memory


The guidance in this topic needs to be reviewed. We are seeing a larger number of resources per cluster, which increases the required memory.
https://access.redhat.com/documentation/en-us/red_hat_advanced_cluster_management_for_kubernetes/2.4/html/install/installing#search-scalability

Comment 7 Ryan Spagnola 2022-01-28 15:55:56 UTC
From customer:

> Compared to the secret creation date, could you provide a timeline of when ACM was first installed, when ACM was upgraded, and when this problem was noticed?

Looking at the creation timestamps of a few resources, I think we originally installed RHACM 2.3 on 10/28/2021 (creation of namespaces open-cluster-management and open-cluster-management-hub), and we upgraded it to RHACM 2.4 on 1/5/2022 (creation of advanced-cluster-management.v2.4.1-* cluster roles). That would mean that the redisgraph-user-secret is the original one from when we installed RHACM.

> Is there a reason why the search operator pod was created yesterday? 

We are working on automating post-install configuration items that include changes to machine config resources. Every time we deploy a new version of those, all worker nodes are rebooted, which causes most of the pods to be recreated.

Comment 8 Jorge Padilla 2022-01-28 21:50:43 UTC
It sounds like the restart of the nodes is triggering this problem. I'd like to confirm that this.  Also it would help to collect must-gather before and after the restart event.

I'll try to recreate this scenario in another cluster.

Comment 20 Jorge Padilla 2022-02-09 16:47:21 UTC
*** Bug 2047273 has been marked as a duplicate of this bug. ***

Comment 22 Jorge Padilla 2022-02-09 17:47:05 UTC
@fdewaley 
Please ask the customer to run the debug script in comment 18. It would help us if they can capture 2 snapshots, first with the system stable and another one while observing high memory consumption. Also, it would help to capture must gather data around the same time.

It's possible that some data is not getting deleted, the debug script should give us some indication if there's a leak.

Comment 27 Jorge Padilla 2022-02-23 04:14:09 UTC
We have identified a couple factors contributing to the increased memory usage.

1. In some scenarios, saving and loading data from a PVC is too slow and requires too much memory. This completely negates the benefit from persisting the Redis state. In this case, turning off persistence can make the Redis instance more stable.  More information in this bugzilla: https://bugzilla.redhat.com/show_bug.cgi?id=2036197#c29

2. Sometimes when deleting a cluster from ACM, the data is not correctly removed from Redisgraph. This can happen if either search-redisgraph or the search-aggregator pods experience problems around the time that a cluster is removed from ACM. We have seen this problem in ACM environments where managed clusters are created and deleted often. We are working on a fix for this scenario.

Comment 28 Felix Dewaleyne 2022-02-23 11:44:28 UTC
(In reply to Jorge Padilla from comment #27)
> We have identified a couple factors contributing to the increased memory
> usage.
> 
> 1. In some scenarios, saving and loading data from a PVC is too slow and
> requires too much memory. This completely negates the benefit from
> persisting the Redis state. In this case, turning off persistence can make
> the Redis instance more stable.  More information in this bugzilla:
> https://bugzilla.redhat.com/show_bug.cgi?id=2036197#c29
> 
> 2. Sometimes when deleting a cluster from ACM, the data is not correctly
> removed from Redisgraph. This can happen if either search-redisgraph or the
> search-aggregator pods experience problems around the time that a cluster is
> removed from ACM. We have seen this problem in ACM environments where
> managed clusters are created and deleted often. We are working on a fix for
> this scenario.

regarding point 2 is there a way to clear up the data.

Comment 29 Jorge Padilla 2022-02-23 21:25:32 UTC
Aa a workaround to clear the Redis data for deleted clusters you could execute these queries directly in the search-redisgraph-0 pod.

This command lists all clusters for which Redisgraph has data.
---
oc exec -it search-redisgraph-0 -n open-cluster-management -- /bin/bash <<EOF
    export REDISCLI_AUTH=\$REDIS_PASSWORD
    redis-cli graph.query search-db "MATCH (n) RETURN DISTINCT n.cluster AS clusterName"
EOF
---


This command will delete data for clusterA and clusterB (replace with your actual cluster names)
---
oc exec -it search-redisgraph-0 -n open-cluster-management -- /bin/bash <<EOF
    export REDISCLI_AUTH=\$REDIS_PASSWORD
    redis-cli graph.query search-db "MATCH (n) WHERE n.cluster IN ['clusterA', 'clusterB'] DELETE n"
EOF
---

Comment 37 Jorge Padilla 2022-04-27 15:27:22 UTC
*** Bug 2052910 has been marked as a duplicate of this bug. ***

Comment 39 bot-tracker-sync 2022-08-16 22:25:11 UTC
G2Bsync 1124031852 comment 
 crizzo71 Wed, 11 May 2022 17:11:55 UTC 
 G2Bsync No code changes in 2.4.5 The solution is Search V2

Comment 40 Xavier 2022-09-02 15:03:19 UTC
Solution to this problem is addressed in next version of search feature. You can try out the DEV-PREVIEW operator here https://github.com/stolostron/search-v2-operator . Give it a shot and let us know your feedback. Plan is to include this in next major version of ACM.

Closing this BZ as FUTURE target.

Comment 41 Tuan 2022-11-04 14:33:04 UTC
*** Bug 2138423 has been marked as a duplicate of this bug. ***


Note You need to log in before you can comment on or make changes to this bug.