Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1632313

Summary: Ignore spurious nested stack locks in convergence
Product: Red Hat OpenStack Reporter: Zane Bitter <zbitter>
Component: openstack-heatAssignee: Zane Bitter <zbitter>
Status: CLOSED ERRATA QA Contact: Ronnie Rasouli <rrasouli>
Severity: medium Docs Contact:
Priority: medium    
Version: 13.0 (Queens)CC: bshephar, cylopez, jschluet, mburns, pamadio, rrasouli, sbaker, shardy, srevivo, therve, zbitter
Target Milestone: z3Keywords: Triaged, ZStream
Target Release: 13.0 (Queens)   
Hardware: All   
OS: Linux   
Whiteboard:
Fixed In Version: openstack-heat-10.0.2-2.el7ost Doc Type: Bug Fix
Doc Text:
Cause: When convergence is enabled for a stack, the "stack check" operation uses a different type of locking (stack-level instead of resource-level) from regular operations such as "stack update" or "stack delete". However, when processing nested stacks they are not considered complete until any stack-level lock is released. Consequence: Any stale locks left behind in a nested stack by a stack check will cause the parent stack to time out when updating or deleting the nested stack. Fix: Ignore the stack-level locks when checking the completion of any nested stack operation (like update or delete in convergence stacks) that uses only resource-level locks. Result: Any stray locks left behind in the database will not prevent further operations on the nested stack.
Story Points: ---
Clone Of: 1632295 Environment:
Last Closed: 2018-11-13 22:14:23 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On: 1626367, 1632295    
Bug Blocks:    

Description Zane Bitter 2018-09-24 15:31:45 UTC
+++ This bug was initially created as a clone of Bug #1632295 +++

+++ This bug was initially created as a clone of Bug #1626367 +++

This is a bug to request a backport of this change on RHOSP11:

https://review.openstack.org/#/c/598354/

Several operations (e.g. stack check) are yet to be converted to convergence-style workflows, and still create locks. While we try to always remove dead locks, it's possible that if one of these gets left behind we won't notice, since convergence doesn't actually use stack locks for most regular operations. When doing _check_status_complete() on a nested stack resource, we wait for the operation to release the nested stack's lock before reporting the resource completed, to ensure that for legacy operations the nested stack is ready to perform another action on as soon as the resource is complete. For convergence stacks this is an unnecessary DB call, and it can lead to resources never completing if a stray lock happens to be left in the database. Only check the nested stack's stack lock for operations where we are not taking a resource lock. This corresponds exactly to legacy-style operations.

Comment 2 Zane Bitter 2018-10-08 16:04:51 UTC
Merged upstream in stable/queens.

Comment 3 Cyril Lopez 2018-10-18 16:02:41 UTC
This could affect osp 10 too ?

Comment 4 Zane Bitter 2018-10-18 16:30:58 UTC
(In reply to Cyril Lopez from comment #3)
> This could affect osp 10 too ?

Yes, in the overcloud.

Comment 16 errata-xmlrpc 2018-11-13 22:14:23 UTC
Since the problem described in this bug report should be
resolved in a recent advisory, it has been closed with a
resolution of ERRATA.

For information on the advisory, and where to find the updated
files, follow the link below.

If the solution does not work for you, open a new bug report.

https://access.redhat.com/errata/RHBA-2018:3603