Bug 1781202 - MIQ(MiqServer#monitor) can't modify frozen Hash after UI worker restarts
Summary: MIQ(MiqServer#monitor) can't modify frozen Hash after UI worker restarts
Keywords:
Status: CLOSED ERRATA
Alias: None
Product: Red Hat CloudForms Management Engine
Classification: Red Hat
Component: Appliance
Version: 5.10.9
Hardware: Unspecified
OS: Unspecified
urgent
urgent
Target Milestone: GA
: 5.11.2
Assignee: Joe Rafaniello
QA Contact: Parthvi Vala
Red Hat CloudForms Documentation
URL:
Whiteboard:
Depends On: 1507131
Blocks: 1788289
TreeView+ depends on / blocked
 
Reported: 2019-12-09 14:23 UTC by Gellert Kis
Modified: 2023-09-18 00:19 UTC (History)
13 users (show)

Fixed In Version: 5.11.2.0
Doc Type: If docs needed, set a value
Doc Text:
Clone Of: 1507131
: 1788289 (view as bug list)
Environment:
5.10.9.1
Last Closed: 2020-02-12 05:02:24 UTC
Category: Bug
Cloudforms Team: CFME Core
Target Upstream Version:
Embargoed:
pm-rhel: cfme-5.11.z+
dmetzger: mirror+


Attachments (Terms of Use)


Links
System ID Private Priority Status Summary Last Updated
Red Hat Product Errata RHBA-2020:0452 0 None None None 2020-02-12 05:02:36 UTC

Comment 4 Joe Rafaniello 2019-12-09 16:18:00 UTC
Hi Gellert,

The preliminary results seems to show that the api is being requested for report results, reaching the 1 GB memory threshold and being killed, with a replacement then started.  Because there is only 1 configured web service worker, any subsequent request would fail to reach an available worker until the replacement is started and available.

The preliminary suggestion is to increase the memory threshold for the web service worker and increase the worker count from 1.  It's unclear what is the correct memory threshold so it's recommended to grow it until the web service worker can fulfill the request without getting too close to the threshold.  We should also evaluate what is causing the large memory growth to ensure there's not already a fix for this.

The "can't modify frozen hash" is a side effect of workers continually restarting and we'll try to evaluate and see if we can fix it.  That fix may not prevent the reported downtime accessing the api though so resolving the underlying problem is what we should do first, so please increase the count and memory threshold for the web service worker.

Comment 5 Joe Rafaniello 2019-12-11 19:31:02 UTC
Based on the logging provided by the customer, we have been able to finally recreate this error and have a solution we're working on to fix it.   The solution will not fix the underlying problems causing this error to occur:  ui/web service worker exceeding memory/not responding but will no longer skip one iteration of the worker monitor loop with an ugly error.

Comment 7 CFME Bot 2020-01-06 21:01:39 UTC
New commit detected on ManageIQ/manageiq/ivanchuk:

https://github.com/ManageIQ/manageiq/commit/6ee115953abeb0c1f5f2da40de8b59d5cd5a7d40
commit 6ee115953abeb0c1f5f2da40de8b59d5cd5a7d40
Author:     Joe Rafaniello <jrafanie>
AuthorDate: Tue Dec 10 16:41:29 2019 -0500
Commit:     Joe Rafaniello <jrafanie>
CommitDate: Tue Dec 10 16:41:29 2019 -0500

    Query current/starting miq_workers, bypassing cache/stale objects

    Fixes https://bugzilla.redhat.com/show_bug.cgi?id=1781202

    If the server deletes the row in the prior loop's call to this method because the
    worker exceeded memory, the subsequent call to post_message_for_workers could fail
    with "can't modify frozen Hash" when trying to update the worker's heartbeat.

    Incidentally fixed on master in commit:
    https://github.com/ManageIQ/manageiq/commit/e6cbfa8b4625147741919c723eacac540f056099

 app/models/miq_server/worker_management/heartbeat.rb | 17 +-
 app/models/mixins/miq_web_server_worker_mixin.rb | 1 +
 spec/lib/workers/heartbeat_spec.rb | 22 +
 3 files changed, 32 insertions(+), 8 deletions(-)

Comment 9 Parthvi Vala 2020-01-22 07:31:23 UTC
FIXED. Verified on 5.11.2.0.20200113212029_18edbd8.

Comment 11 errata-xmlrpc 2020-02-12 05:02:24 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2020:0452

Comment 12 Red Hat Bugzilla 2023-09-18 00:19:05 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 120 days


Note You need to log in before you can comment on or make changes to this bug.