Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1426835

Summary: heat-engine thread consumes large memory
Product: Red Hat OpenStack Reporter: Jaison Raju <jraju>
Component: openstack-heatAssignee: Zane Bitter <zbitter>
Status: CLOSED CURRENTRELEASE QA Contact: Amit Ugol <augol>
Severity: high Docs Contact:
Priority: high    
Version: 8.0 (Liberty)CC: ipetrova, jraju, mburns, pablo.iranzo, rhel-osp-director-maint, sbaker, shardy, srevivo, therve, zbitter
Target Milestone: ---Keywords: Triaged, ZStream
Target Release: 11.0 (Ocata)   
Hardware: All   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2017-07-28 18:07:09 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1430433    
Bug Blocks:    

Description Jaison Raju 2017-02-25 13:22:30 UTC
Description of problem:
For an environment with 40 computes , heat-engine thread consumes large memory
while deployment .
For a system with 160G , the 8 heat-engine threads starts with 1% of system memory & rises upto 10% at the end of installation .
The threads cannot be reduced as it will cause the deployment time to increase .

Version-Release number of selected component (if applicable):
RHOS8
openstack-heat-engine-5.0.1-9.el7ost.noarch

How reproducible:
Always on customer environment

Steps to Reproduce:
1.
2.
3.

Actual results:


Expected results:


Additional info:
https://bugs.launchpad.net/heat/+bug/1626675 seems related & can be backported to improve this . one of this patch (c392e9138211f757f02a02b29712cd9431c55616) is present in osp10 .

Comment 4 Zane Bitter 2017-02-27 20:30:29 UTC
TripleO is a very aggressive use case for Heat, because it contains hundreds of deeply-nested stacks even in a small deployment, complex environment files with hundreds of resource type mappings, and enormous files inputs, and it must all run on a single server. All of these are extremely unusual and had not been optimised for until TripleO demonstrated the problems they caused.

The single largest improvement here occurred early in the Newton cycle, when we changed Heat to keep only a single copy per worker of the input files in memory, rather than one per (nested) stack: https://bugs.launchpad.net/heat/+bug/1570983

Unfortunately, however, this change is wholly unsuitable for backporting, even to OSP 9. It is large, risky, depends on other large and risky changes, and modifies the database schema in a major way.

Over the course of Newton development, the TripleO templates increased even further in complexity and further effort was put into improving the memory usage, tracked in the bug https://bugs.launchpad.net/heat/+bug/1626675 that you identified. The results of this effort are fairly well documented upstream: http://lists.openstack.org/pipermail/openstack-dev/2017-January/109748.html

We can ignore the improvement in memory usage by YAQL functions, since these were introduced only in Newton.

https://review.openstack.org/382068/ (addressing the very large product of resource type mappings times number of nested stacks) is a very simple change that could certainly be backported. Its merging is temporally correlated with a very large reduction in memory usage... much greater than could have been theoretically expected. So it may or may not have a significant effect. (And it would certainly be reduced, because much of the savings was from the increased number of nested stacks in Newton TripleO.) https://review.openstack.org/377061/ may or may not have contributed as well, though there is no point backporting it unless the many, many other patches for circular references are also backported (since the point is get rid of _all_ of them).

https://review.openstack.org/#/c/386247/ and https://review.openstack.org/#/c/383839/ reduced the loading of child stacks in-memory along with their parent stacks - a legacy of a change in Kilo to do many, but unfortunately not all, nested stack operations over RPC instead of in a single worker. Both would be very challenging to backport.


Backporting major architectural changes to a version of the project that is EOL upstream (so we have e.g. no access to upstream CI testing) while maintaining the stability our customers expect is not feasible.

The best solution here would be to upgrade to OSP10. Failing that, I know of no reason in principle that you couldn't use the OSP10 version of Heat in the undercloud while continuing to use the OSP8 templates, though it hasn't been tested. Otherwise, you probably need to accept that Heat in an OSP8 undercloud will use a lot of memory and size the undercloud machine accordingly.

Comment 5 Zane Bitter 2017-02-27 20:48:10 UTC
This patch to python-heatclient might conceivably help: https://review.openstack.org/#/c/402081/