Fedora Account System
Red Hat Associate
Red Hat Customer
Description of the problem: Redis pod crashing with OOMKilled in ACM 2.4 running on OCP 4.8 and getting following error on the ACM Overview page: ~~~ An unexpected error occurred. Search service is unavailable ~~~ Following are the findings from the pod logs: - search-prod-c0b92-search-aggregator-6c9c9fc4f7-m7dm7 ~~~ 2022-01-17T20:40:41.553464795Z E0117 20:40:41.553010 1 redisinterfacev2.go:47] Error fetching results from RedisGraph V2 : LOADING Redis is loading the dataset in memory 2022-01-17T20:40:41.553464795Z W0117 20:40:41.553037 1 clusterWatch.go:150] Error on UpdateByName() LOADING Redis is loading the dataset in memory ~~~ - search-prod-c0b92-search-api-864784d684-gpz7t ~~~ 2022-01-17T20:34:21.196934589Z [2022-01-17T20:34:21.195] [ERROR] [search-api] [server] Unable to resolve search request because RedisGraph is unavailable. 2022-01-17T20:34:56.468684530Z [2022-01-17T20:34:56.468] [INFO] [search-api] [server] Role configuration has changed. User RBAC cache has been deleted 2022-01-17T20:38:23.659061144Z [2022-01-17T20:38:23.658] [INFO] [search-api] [server] The Redis connection has ended. 2022-01-17T20:38:23.659338130Z [2022-01-17T20:38:23.659] [INFO] [search-api] [server] Error with Redis connection. 2022-01-17T20:38:23.864374597Z [2022-01-17T20:38:23.864] [INFO] [search-api] [server] Error with Redis connection. 2022-01-17T20:38:24.212032608Z [2022-01-17T20:38:24.211] [INFO] [search-api] [server] Error with Redis connection. 2022-01-17T20:38:31.913318498Z [2022-01-17T20:38:31.912] [INFO] [search-api] [server] Redis Client connected. ~~~ - search-prod-c0b92-search-collector-869545f6bd-cxht6 ~~~ 2022-01-17T20:30:46.888225546Z W0117 20:30:46.888117 1 sender.go:169] Received error response [POST to: https://search-aggregator.open-cluster-management.svc:3010/aggregator/clusters/local-cluster/sync responded with error. StatusCode: 400 Message: 400 Bad Request] from Aggregator. Resending in 600000 ms after resetting config. ~~~ - search-redisgraph-0 ~~~ 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG6[16]: Peer certificate not required 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[16]: TLS state (accept): before SSL initialization 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG3[16]: SSL_accept: Peer suddenly disconnected 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG5[16]: Connection reset: 0 byte(s) sent to TLS, 0 byte(s) sent to socket 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[16]: Local descriptor (FD=14) closed 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[16]: Service [redis] finished (3 left) 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[ui]: FD=8 events=0x2001 revents=0x1 2022-01-17T20:40:16.590566358Z 20:40:16 stunnel.1 | 2022.01.17 20:40:16 LOG7[ui]: Service [redis] accepted (FD=14) from ::ffff:10.59.2.1:55500 ~~~ - search-ui-85f85bcbd7-5d59n ~~~ 2022-01-17T18:30:53.691183884Z {"level":"info","time":1642444253688,"instance":1,"msg":"Not Modified","status":304,"method":"GE T","url":"/search/static/js/0.9aea736a.chunk.js","ms":8.56} 2022-01-17T18:30:53.691514710Z {"level":"info","time":1642444253689,"instance":1,"msg":"Not Modified","status":304,"method":"GE T","url":"/search/static/js/7.479fc9f4.chunk.js","ms":8.3} 2022-01-17T18:30:53.709557785Z {"level":"info","time":1642444253708,"instance":1,"msg":"Not Modified","status":304,"method":"GE T","url":"/search/static/media/RHACM-Logo.8a59e4b5.svg","ms":0.58} 2022-01-17T18:30:53.950240013Z {"level":"info","time":1642444253950,"instance":1,"msg":"Not Modified","status":304,"method":"GE T","url":"/search/static/media/RedHatText-Regular.4a43a00f.woff2","ms":1.14} 2022-01-17T18:30:53.951709016Z {"level":"info","time":1642444253951,"instance":1,"msg":"Not Modified","status":304,"method":"GET","url":"/search/static/media/RedHatText-Medium.0227c8bb.woff2","ms":0.44} 2022-01-17T18:37:03.861858408Z {"level":"warn","time":1642444623861,"instance":1,"msg":"Not Found","status":404,"method":"GET","url":"/search/locales/en/translation.json","ms":1.31} ~~~ Release version: ACM 2.4.1 OCP version: OCP 4.8.25 Steps to reproduce: 1. Install ACM 2.4.1 on OCP 4.8.25 2. Enable Search operator Actual results: The `search-redisgraph-0` pod is running out of memory and getting terminated ~~~ $ omg get pod search-redisgraph-0 -o yaml apiVersion: v1 kind: Pod metadata: annotations: .... resources: limits: memory: 4Gi requests: cpu: 25m memory: 128Mi .... lastState: terminated: containerID: cri-o://13188df8f22de6e1302ee3ac80aba22035827909c3d9c71c72390efef909ad36 exitCode: 137 finishedAt: '2022-01-17T20:38:23Z' reason: OOMKilled startedAt: '2022-01-17T20:27:22Z' name: redisgraph ready: true restartCount: 361 ~~~ Expected results: The `search-redisgraph-0` pod should be running. Additional info: The error is still there even after the pod restart.
Thanks for the detailed description. The root cause is the Redis pod running out of memory. This topic explains how to increase the memory limits. https://access.redhat.com/documentation/en-us/red_hat_advanced_cluster_management_for_kubernetes/2.4/html/web_console/web-console#options-increase-memory The connection errors on the other search-* pods should recover once the Redis pod stabilizes, but you could also manually restart the pods to reestablish the connection.
expending the memory limit wasn't enough, the connection errors continued as reported by the customer (raising to 10gb and then 16gb). could there be other changes to make? no logs were provided, so I did ask for a new set.
it turns out that the PV what the Redis is using was full and there were no space left. By default, ACM installed with 10GB PV, and that was full after few months of running it. A new volume was created (the customer didn't see how to extend it dynamically), used different storage class and increased it to 50GB. no more crashes. do we have a better process I could use to create a kbase article, and how much storage would be require for a setup where ACM manags between 20 to 40 clusters?
To prevent this problem from happening, the search operator should check that the PVC is larger than the memory limit. We should also update this topic. https://access.redhat.com/documentation/en-us/red_hat_advanced_cluster_management_for_kubernetes/2.4/html/web_console/web-console#options-increase-memory The guidance in this topic needs to be reviewed. We are seeing a larger number of resources per cluster, which increases the required memory. https://access.redhat.com/documentation/en-us/red_hat_advanced_cluster_management_for_kubernetes/2.4/html/install/installing#search-scalability
From customer: > Compared to the secret creation date, could you provide a timeline of when ACM was first installed, when ACM was upgraded, and when this problem was noticed? Looking at the creation timestamps of a few resources, I think we originally installed RHACM 2.3 on 10/28/2021 (creation of namespaces open-cluster-management and open-cluster-management-hub), and we upgraded it to RHACM 2.4 on 1/5/2022 (creation of advanced-cluster-management.v2.4.1-* cluster roles). That would mean that the redisgraph-user-secret is the original one from when we installed RHACM. > Is there a reason why the search operator pod was created yesterday? We are working on automating post-install configuration items that include changes to machine config resources. Every time we deploy a new version of those, all worker nodes are rebooted, which causes most of the pods to be recreated.
It sounds like the restart of the nodes is triggering this problem. I'd like to confirm that this. Also it would help to collect must-gather before and after the restart event. I'll try to recreate this scenario in another cluster.
*** Bug 2047273 has been marked as a duplicate of this bug. ***
@fdewaley Please ask the customer to run the debug script in comment 18. It would help us if they can capture 2 snapshots, first with the system stable and another one while observing high memory consumption. Also, it would help to capture must gather data around the same time. It's possible that some data is not getting deleted, the debug script should give us some indication if there's a leak.
We have identified a couple factors contributing to the increased memory usage. 1. In some scenarios, saving and loading data from a PVC is too slow and requires too much memory. This completely negates the benefit from persisting the Redis state. In this case, turning off persistence can make the Redis instance more stable. More information in this bugzilla: https://bugzilla.redhat.com/show_bug.cgi?id=2036197#c29 2. Sometimes when deleting a cluster from ACM, the data is not correctly removed from Redisgraph. This can happen if either search-redisgraph or the search-aggregator pods experience problems around the time that a cluster is removed from ACM. We have seen this problem in ACM environments where managed clusters are created and deleted often. We are working on a fix for this scenario.
(In reply to Jorge Padilla from comment #27) > We have identified a couple factors contributing to the increased memory > usage. > > 1. In some scenarios, saving and loading data from a PVC is too slow and > requires too much memory. This completely negates the benefit from > persisting the Redis state. In this case, turning off persistence can make > the Redis instance more stable. More information in this bugzilla: > https://bugzilla.redhat.com/show_bug.cgi?id=2036197#c29 > > 2. Sometimes when deleting a cluster from ACM, the data is not correctly > removed from Redisgraph. This can happen if either search-redisgraph or the > search-aggregator pods experience problems around the time that a cluster is > removed from ACM. We have seen this problem in ACM environments where > managed clusters are created and deleted often. We are working on a fix for > this scenario. regarding point 2 is there a way to clear up the data.
Aa a workaround to clear the Redis data for deleted clusters you could execute these queries directly in the search-redisgraph-0 pod. This command lists all clusters for which Redisgraph has data. --- oc exec -it search-redisgraph-0 -n open-cluster-management -- /bin/bash <<EOF export REDISCLI_AUTH=\$REDIS_PASSWORD redis-cli graph.query search-db "MATCH (n) RETURN DISTINCT n.cluster AS clusterName" EOF --- This command will delete data for clusterA and clusterB (replace with your actual cluster names) --- oc exec -it search-redisgraph-0 -n open-cluster-management -- /bin/bash <<EOF export REDISCLI_AUTH=\$REDIS_PASSWORD redis-cli graph.query search-db "MATCH (n) WHERE n.cluster IN ['clusterA', 'clusterB'] DELETE n" EOF ---
*** Bug 2052910 has been marked as a duplicate of this bug. ***
G2Bsync 1124031852 comment crizzo71 Wed, 11 May 2022 17:11:55 UTC G2Bsync No code changes in 2.4.5 The solution is Search V2
Solution to this problem is addressed in next version of search feature. You can try out the DEV-PREVIEW operator here https://github.com/stolostron/search-v2-operator . Give it a shot and let us know your feedback. Plan is to include this in next major version of ACM. Closing this BZ as FUTURE target.
*** Bug 2138423 has been marked as a duplicate of this bug. ***