Fedora Account System
Red Hat Associate
Red Hat Customer
Hi Gellert, The preliminary results seems to show that the api is being requested for report results, reaching the 1 GB memory threshold and being killed, with a replacement then started. Because there is only 1 configured web service worker, any subsequent request would fail to reach an available worker until the replacement is started and available. The preliminary suggestion is to increase the memory threshold for the web service worker and increase the worker count from 1. It's unclear what is the correct memory threshold so it's recommended to grow it until the web service worker can fulfill the request without getting too close to the threshold. We should also evaluate what is causing the large memory growth to ensure there's not already a fix for this. The "can't modify frozen hash" is a side effect of workers continually restarting and we'll try to evaluate and see if we can fix it. That fix may not prevent the reported downtime accessing the api though so resolving the underlying problem is what we should do first, so please increase the count and memory threshold for the web service worker.
Based on the logging provided by the customer, we have been able to finally recreate this error and have a solution we're working on to fix it. The solution will not fix the underlying problems causing this error to occur: ui/web service worker exceeding memory/not responding but will no longer skip one iteration of the worker monitor loop with an ugly error.
https://github.com/ManageIQ/manageiq/pull/19638
New commit detected on ManageIQ/manageiq/ivanchuk: https://github.com/ManageIQ/manageiq/commit/6ee115953abeb0c1f5f2da40de8b59d5cd5a7d40 commit 6ee115953abeb0c1f5f2da40de8b59d5cd5a7d40 Author: Joe Rafaniello <jrafanie> AuthorDate: Tue Dec 10 16:41:29 2019 -0500 Commit: Joe Rafaniello <jrafanie> CommitDate: Tue Dec 10 16:41:29 2019 -0500 Query current/starting miq_workers, bypassing cache/stale objects Fixes https://bugzilla.redhat.com/show_bug.cgi?id=1781202 If the server deletes the row in the prior loop's call to this method because the worker exceeded memory, the subsequent call to post_message_for_workers could fail with "can't modify frozen Hash" when trying to update the worker's heartbeat. Incidentally fixed on master in commit: https://github.com/ManageIQ/manageiq/commit/e6cbfa8b4625147741919c723eacac540f056099 app/models/miq_server/worker_management/heartbeat.rb | 17 +- app/models/mixins/miq_web_server_worker_mixin.rb | 1 + spec/lib/workers/heartbeat_spec.rb | 22 + 3 files changed, 32 insertions(+), 8 deletions(-)
FIXED. Verified on 5.11.2.0.20200113212029_18edbd8.
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2020:0452
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 120 days