Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1994491

Summary: horizontalpodautoscaler/rc-light - reason/FailedGetResourceMetric failed to get cpu utilization
Product: OpenShift Container Platform Reporter: Maciej Szulik <maszulik>
Component: kube-controller-managerAssignee: Filip Krepinsky <fkrepins>
Status: CLOSED DUPLICATE QA Contact: zhou ying <yinzhou>
Severity: medium Docs Contact:
Priority: medium    
Version: 4.9CC: aos-bugs, ercohen, jchaloup, mfojtik
Target Milestone: ---   
Target Release: 4.9.0   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard: tag-ci
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2021-08-17 14:19:36 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Maciej Szulik 2021-08-17 11:52:35 UTC
There's a great uptake in:

[sig-arch] events should not repeat pathologically

synthetic test failure. Pretty good chunk of these are specifically failing:

horizontalpodautoscaler/rc-light - reason/FailedGetResourceMetric failed to get cpu utilization

For now I've added an exception in https://github.com/openshift/origin/blob/master/pkg/synthetictests/duplicated_events.go#L86

But from the initial look I had it looks like the problem might be metric.

A bunch of useful links:

https://search.ci.openshift.org/?search=reason%2FFailedGetResourceMetric+failed+to+get+&maxAge=24h&context=1&type=junit&name=&excludeName=&maxMatches=5&maxBytes=20971520&groupBy=job

Particular failure I looked at:

https://prow.ci.openshift.org/view/gs/origin-ci-test/pr-logs/pull/26377/pull-ci-openshift-origin-master-e2e-gcp/1427556641120718848

where 

[sig-autoscaling] [Feature:HPA] Horizontal pod autoscaling (scale resource: CPU) ReplicationController light Should scale from 1 pod to 2 pods 

passed and failed once, so it was marked as flake. But the amount of events (79) in that failure instance over 15m (it timed out) caused the synthetic test to fail.

Checking kcm logs I do see a lot of:

E0817 09:57:57.833193       1 horizontal.go:227] failed to compute desired number of replicas based on listed metrics for ReplicationController/e2e-horizontal-pod-autoscaling-8508/rc-light: invalid metrics (1 invalid out of 1), first error is: failed to get cpu utilization: unable to get metrics for resource cpu: no metrics returned from resource metrics API

during that failed instance.

Comment 2 Jan Chaloupka 2021-08-17 14:19:36 UTC

*** This bug has been marked as a duplicate of bug 1993985 ***