Bug 2006970
| Summary: | Percent CPU frequently goes above 100 | ||
|---|---|---|---|
| Product: | Red Hat OpenStack | Reporter: | Paul Leimer <pleimer> |
| Component: | collectd-libpod-stats | Assignee: | Nobody <nobody> |
| Status: | CLOSED ERRATA | QA Contact: | Leonid Natapov <lnatapov> |
| Severity: | medium | Docs Contact: | |
| Priority: | medium | ||
| Version: | 16.2 (Train) | CC: | csibbitt, jamsmith, joflynn, lmadsen, spower |
| Target Milestone: | z2 | Keywords: | Triaged, ZStream |
| Target Release: | 16.2 (Train on RHEL 8.4) | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | collectd-libpod-stats-1.0.4-1.el8ost | Doc Type: | Bug Fix |
| Doc Text: |
In cases where high CPU use was monitored in a multi-core system, the calculation for CPU use was inaccurate.
+
With this update, the calculation of CPU use in a multi-core scenario is now accurate. The latest STF dashboards have been adjusted to incorporate this update.
|
Story Points: | --- |
| Clone Of: | Environment: | ||
| Last Closed: | 2022-03-23 22:11:40 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Paul Leimer
2021-09-22 18:21:14 UTC
I consider the dashboard adjustment in GH pull #40 workaround to mitigate the effects of this bug in dashboards. A patch to libpod-stats must still be completed. Reproduce this bug locally: 1. Launch collectd with the libpod-stats plugin loaded, be sure that collectd is writing to a data store like Prometheus so that metrics can be graphed 2. Start another container that was not previously running 3. CPU percentage calculations for the container in step 2 will spike After further evaluation, the above process does not necessarily reproduce the bug. Rather, it demonstrates a > 100% usage when multiple cores are working hard at once, in which case >100% is expected behavior. The real bug is suspected to be because of the usage of unsigned integers to calculate difference between cpu utilization at different points in time. If a counter resets, the numerator might result in a unsigned int where the high order bits have been flipped (as a result of 2's-complement). This has been fixed upstream in the attached PR. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory (Release of components for Red Hat OpenStack Platform 16.2.2), and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2022:1001 |