Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1986119

Summary: Compact jobs alerting on etcdGRPCRequestsSlow
Product: OpenShift Container Platform Reporter: Stephen Benjamin <stbenjam>
Component: NetworkingAssignee: Dan Winship <danw>
Networking sub component: openshift-sdn QA Contact: zhaozhanqi <zzhao>
Status: CLOSED DUPLICATE Docs Contact:
Severity: medium    
Priority: medium CC: astoycos, dmistry, sippy, vpickard, wking
Version: 4.9   
Target Milestone: ---   
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard: tag-ci
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
job=periodic-ci-openshift-release-master-ci-4.9-e2e-aws-compact=all job=periodic-ci-openshift-release-master-ci-4.9-e2e-aws-compact-upgrade=all job=periodic-ci-openshift-release-master-ci-4.9-e2e-azure-compact=all job=periodic-ci-openshift-release-master-ci-4.9-e2e-gcp-compact-upgrade=all job=periodic-ci-openshift-release-master-ci-4.9-e2e-azure-compact-upgrade=all job=periodic-ci-openshift-release-master-ci-4.9-e2e-gcp-compact=all job=periodic-ci-openshift-release-master-ci-4.9-upgrade-from-stable-4.8-e2e-aws-compact-upgrade=all job=release-openshift-ocp-installer-e2e-metal-compact-4.9=all
Last Closed: 2021-08-03 13:59:15 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Stephen Benjamin 2021-07-26 17:32:52 UTC
job:
periodic-ci-openshift-release-master-ci-4.9-e2e-aws-compact
periodic-ci-openshift-release-master-ci-4.9-e2e-azure-compact
periodic-ci-openshift-release-master-ci-4.9-e2e-gcp-compact

is failing frequently in CI, see testgrid results:
https://testgrid.k8s.io/redhat-openshift-ocp-release-4.9-informing#periodic-ci-openshift-release-master-ci-4.9-e2e-aws-compact

Example job:
https://prow.ci.openshift.org/view/gs/origin-ci-test/logs/periodic-ci-openshift-release-master-ci-4.9-e2e-gcp-compact/1416473242087985152

This alert is firing:
alert etcdGRPCRequestsSlow fired for 150 seconds with labels: {endpoint="etcd-metrics", grpc_method="Txn", grpc_service="etcdserverpb.KV", instance="10.0.0.5:9979", job="etcd", namespace="openshift-etcd", pod="etcd-ci-op-0ptgfjiz-d7295-c8njd-master-1", service="etcd", severity="critical"}

From what I can tell this started around 7/17 for the compact jobs, probably due to this being merged: https://github.com/openshift/cluster-etcd-operator/pull/566

Comment 4 Stephen Benjamin 2021-07-29 16:42:19 UTC
Looking at some other tests, we see the CPU increase pretty clearly on 6/22:

https://search.ci.openshift.org/graph/metrics?metric=cluster%3Ausage%3Acpu%3Atest%3Aseconds&job=periodic-ci-openshift-release-master-ci-4.9-e2e-aws&job=periodic-ci-openshift-release-master-ci-4.9-e2e-aws-compact&job=periodic-ci-openshift-release-master-ci-4.9-e2e-azure-compact&job=periodic-ci-openshift-release-master-ci-4.9-e2e-gcp&job=periodic-ci-openshift-release-master-ci-4.9-e2e-gcp-compact

We enabled the network policy tests on 6/22, so my guess is it's that -- this PR merged: https://github.com/openshift/origin/pull/26231


I took a look at PromeCleus for this run --  https://prow.ci.openshift.org/view/gs/origin-ci-test/logs/periodic-ci-openshift-release-master-ci-4.9-e2e-gcp-compact/1420097739735175168

The alerts start happening at 19:34:51, and I see a NetworkPolicy test kick off right at 19:34:38.  I think there's a pretty good chance it's these tests.

Some additional evidence: e2e-metal-ipi which runs the compact instances as VM's on extremely overpowered hardware doesn't suffer from this problem.  https://prow.ci.openshift.org/view/gs/origin-ci-test/logs/periodic-ci-openshift-release-master-nightly-4.9-e2e-metal-ipi-compact/1419447547176423424


I'm not sure if the tests can be optimized to reduce CPU usage, or if we just need to use bigger instances across the board for compact jobs. Sending this to the SDN team for their opinion.

Comment 6 Victor Pickard 2021-08-02 15:20:11 UTC
Dan, can you please take a look to see if network policy tests are causing this issue? Thanks

Comment 7 Dan Winship 2021-08-03 13:09:25 UTC
(In reply to Victor Pickard from comment #6)
> Dan, can you please take a look to see if network policy tests are causing
> this issue? Thanks

I mean, the timing is pretty suspicious, and we've seen other problems with the new NetPol tests overloading the cluster. They are also pretty flaky. Maybe we should just disable them in all e2e jobs for now?

Comment 8 Dan Winship 2021-08-03 13:59:15 UTC

*** This bug has been marked as a duplicate of bug 1980141 ***