Created attachment 1700998 [details] telemetry-config configmap file Description of problem: It's "code:apiserver_request_count:rate:sum" in telemetry-config configmap, but there is not such metrics, should be "code:apiserver_request_total:rate:sum" # oc -n openshift-monitoring get cm telemetry-config -oyaml ... - '{__name__="code:apiserver_request_count:rate:sum"}' ... # token=`oc sa get-token prometheus-k8s -n openshift-monitoring` # oc -n openshift-monitoring exec -c prometheus prometheus-k8s-0 -- curl -k -H "Authorization: Bearer $token" 'https://prometheus-k8s.openshift-monitoring.svc:9091/api/v1/label/__name__/values' | jq | grep "code:apiserver_request" "code:apiserver_request_total:increase30d", "code:apiserver_request_total:rate:sum", Version-Release number of selected component (if applicable): 4.5.0-0.nightly-2020-07-14-022827 How reproducible: always Steps to Reproduce: 1. see the description 2. 3. Actual results: Expected results: Additional info:
Thanks Junqi, I was going to open a Bugzilla for this already, it should be fixed in 4.6 already. But we need another one for 4.4 as well, correct?
Fix for 4.6 was included in https://github.com/openshift/cluster-monitoring-operator/pull/821
(In reply to Lili Cosic from comment #1) > Thanks Junqi, I was going to open a Bugzilla for this already, it should be > fixed in 4.6 already. But we need another one for 4.4 as well, correct? yes, 4.4 bug: https://bugzilla.redhat.com/show_bug.cgi?id=1856767 # oc get clusterversion NAME VERSION AVAILABLE PROGRESSING SINCE STATUS version 4.4.0-0.nightly-2020-07-12-055624 True False 5h58m Cluster version is 4.4.0-0.nightly-2020-07-12-055624 # oc -n openshift-monitoring get cm telemetry-config -oyaml | grep "code:apiserver_request" # code:apiserver_request_count:rate:sum identifies average of occurances - '{__name__="code:apiserver_request_count:rate:sum"}' # token=`oc sa get-token prometheus-k8s -n openshift-monitoring` # oc -n openshift-monitoring exec -c prometheus prometheus-k8s-0 -- curl -k -H "Authorization: Bearer $token" 'https://prometheus-k8s.openshift-monitoring.svc:9091/api/v1/label/__name__/values' | jq | grep "code:apiserver_request" "code:apiserver_request_total:rate:sum",
Reassigned to Frederic.
since 4.6 bug 1859164 is fixed, set the Target Release to 4.5.z
Increasing to high severity as we are unable to observe apiserver requests metrics which is one of our most important signals.
Tested with 4.5.0-0.nightly-2020-07-23-201307, issue is fixed # oc get clusterversion NAME VERSION AVAILABLE PROGRESSING SINCE STATUS version 4.5.0-0.nightly-2020-07-23-201307 True False 83m Cluster version is 4.5.0-0.nightly-2020-07-23-201307 # oc -n openshift-monitoring get cm telemetry-config -oyaml | grep "code:apiserver_request_total:rate:sum" # (@openshift/openshift-team-olm) code:apiserver_request_total:rate:sum identifies average of occurences - '{__name__="code:apiserver_request_total:rate:sum"}' # token=`oc sa get-token prometheus-k8s -n openshift-monitoring` # oc -n openshift-monitoring exec -c prometheus prometheus-k8s-0 -- curl -k -H "Authorization: Bearer $token" 'https://prometheus-k8s.openshift-monitoring.svc:9091/api/v1/label/__name__/values' | jq | grep "code:apiserver_request_total:rate:sum" "code:apiserver_request_total:rate:sum",
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2020:3028