Bug 1964161 - 504 error when trying to pull up a pod in search view that resides on a remote cluster
Summary: 504 error when trying to pull up a pod in search view that resides on a remot...
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat Advanced Cluster Management for Kubernetes
Classification: Red Hat
Component: Server Foundation
Version: rhacm-2.2
Hardware: All
OS: All
unspecified
high
Target Milestone: ---
: rhacm-2.2.5
Assignee: Jian Qiu
QA Contact: Song Lai
Christopher Dawson
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2021-05-24 20:37 UTC by Tuan
Modified: 2025-10-03 11:04 UTC (History)
4 users (show)

Fixed In Version:
Doc Type: If docs needed, set a value
Doc Text:
Clone Of:
Environment:
Last Closed: 2021-06-28 23:24:27 UTC
Target Upstream Version:
Embargoed:
slai: qe_test_coverage+
ming: rhacm-2.2.z+


Attachments (Terms of Use)
must-gather (6.94 MB, application/zip)
2021-05-26 19:45 UTC, Jared Deubel
no flags Details


Links
System ID Private Priority Status Summary Last Updated
Github open-cluster-management backlog issues 12776 0 None None None 2021-05-25 15:46:52 UTC
Red Hat Product Errata RHBA-2021:2558 0 None None None 2021-06-28 23:24:30 UTC

Comment 8 Jared Deubel 2021-05-26 19:45:11 UTC
Created attachment 1787359 [details]
must-gather

Comment 13 Jorge Padilla 2021-06-03 18:53:07 UTC
Each managed cluster maps to a namespace with the same name in the hub.

- Is it possible that the namespace `gp-2-poc` exists but the logged user isn't authorized. It seems unlikely given the output of the `oc auth` command, but let's confirm that the user has the correct permissions for this namespace.
- Is it possible that the namespace on the hub doesn't match the cluster name?  I'm not aware of any setting to overwrite this name, but if that's the case then the UI code could be making the wrong assumption.

Comment 17 zyin@redhat.com 2021-06-04 03:02:52 UTC
Tuan, 

please check the status of pod open-cluster-management-agent-addon/klusterlet-addon-workmgr-xx and the events in the ns open-cluster-management-agent-addon on managed cluster gp-1-poc. 

and help collect the logs of pod open-cluster-management-agent-addon/klusterlet-addon-workmgr-xx on managed cluster gp-1-poc after sample-managed-cluster-view clusterview is created.

thanks.

Comment 19 zyin@redhat.com 2021-06-07 12:38:10 UTC
the pods klusterlet-addon-workmgr-645cf6c946-jzpl9 crashed since Liveness/Readiness porbe failed from the info above.

I suspect this issue is the same with the one https://bugzilla.redhat.com/show_bug.cgi?id=1960793 you reported.


which platform do the managed clsuters host? as the same with the env of  https://bugzilla.redhat.com/show_bug.cgi?id=1960793?


please help collect the logs of this pod klusterlet-addon-workmgr-xxx (kubectl logs -n open-cluster-management-agent-addon klusterlet-addon-workmgr-xxx).
 and the cpu/mem usages of nodes on the managed clusters.(kubectl describe nodes)

Comment 23 zyin@redhat.com 2021-06-10 00:21:14 UTC
this is a known issue that takes a long time to request APIs in some cases (like insufficient cpu/mem or bad network etc.) and we are planning to fix it in the next release.

Comment 25 zyin@redhat.com 2021-06-16 03:40:16 UTC
Hi Tuan, 

The Search uses ManagedClusterView as backend to query APIs of managedCluster.

Get 504 error from Search since it failed to query APIs using ManagedClusterView on the managedCluster.

From the logs we can see that the resources failed to be cached since there are many throttling requests that took a long time and failed as timeouts.

This problem occurs when the load of concurrent requests is large because the default for client to request APIs are low (5 and 10).

That is the reason why it failed to request APIs using ManagedClusterView.

logs:
```
E0604 17:04:09.112302       1 memcache.go:196] couldn't get resource list for quota.openshift.io/v1: the server is currently unable to handle the request
I0604 17:04:15.294933       1 request.go:645] Throttling request took 11.190788102s, request: GET:https://172.29.0.1:443/apis/networking.istio.io/v1beta1?timeout=32s
2021-06-04T17:04:19.156Z	ERROR	controllers.ManagedClusterView	failed to query resource	{"ManagedClusterView": "gp-1-poc/0c44bfa690c9476e57fdbbfae2155d18427291f2", "error": "pods \"archer-v1-executor-5464b7475f-nn7bm\" not found"}
```


We will have a patch to improve the caches mechanism and the requests throttle QPS and Burst in 2.2.5 to address this issue.


as you mentioned it is ok to search resources of local-cluster, could help to compare the hub and managed clusters, like platform, ocp version, network etc? let me try to figure out why the load of requests is large.

Comment 33 errata-xmlrpc 2021-06-28 23:24:27 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory (Red Hat Advanced Cluster Management 2.2.5 bug fix update), and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2021:2558

Comment 34 Red Hat Bugzilla 2023-09-15 01:07:05 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days


Note You need to log in before you can comment on or make changes to this bug.