Bug 1245639
| Summary: | KeyError after mariadb restart | ||
|---|---|---|---|
| Product: | Red Hat OpenStack | Reporter: | Attila Fazekas <afazekas> |
| Component: | python-sqlalchemy | Assignee: | Michael Bayer <mbayer> |
| Status: | CLOSED ERRATA | QA Contact: | Leonid Natapov <lnatapov> |
| Severity: | high | Docs Contact: | |
| Priority: | high | ||
| Version: | 7.0 (Kilo) | CC: | apevec, dnavale, lhh, yeylon |
| Target Milestone: | z2 | Keywords: | Rebase, Triaged, ZStream |
| Target Release: | 7.0 (Kilo) | ||
| Hardware: | Unspecified | ||
| OS: | Unspecified | ||
| Whiteboard: | |||
| Fixed In Version: | python-sqlalchemy-1.0.8-1.el7ost | Doc Type: | Rebase: Bug Fixes Only |
| Doc Text: |
Previously, a connection-level event handler in SQLAlchemy was invoked with the wrong state when a database reconnection attempt had failed. A separate change introduced in SQLAlchemy 1.0.2 made a slight change to the scope of the connection-level datastructure (Connection.info) during the reconnection process. The oslo.db OpenStack library depends on the state, which it stores in the datastructure within the event handler, and due to these two changes together, the datastructure unexpectedly blanked out if multiple reconnection attempts failed before eventually succeeding. As a result, when OpenStack applications attempted to recover after an unexpected database disconnect, they would in some cases encounter this situation, raise a stack trace from within oslo.db and fail to continue.
With this update, the rebase of SQLAlchemy to version 1.0.8 fixes the issue with the event handler so that the correct state is passed to the connection event, allowing oslo.db to correctly maintain the information in Connection.info it expects, resulting in the multiple database reconnection attempts in OpenStack applications no longer causing oslo.db to fail.
|
Story Points: | --- |
| Clone Of: | Environment: | ||
| Last Closed: | 2015-10-08 12:21:47 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | --- | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
|
Description
Attila Fazekas
2015-07-22 12:48:08 UTC
this issue is local to oslo.db. The cause is that we have a protection mechanism which ensures that a connection which is used in a certain process is only used in that same process subsequently. This is the 'pid' key we place in .info. What's not clear is why the key would be missing. I thought perhaps that if the connection were recycled, maybe the 'connect' event which populates the 'pid' is not getting called, but this event is always called in the two areas where the connection is recycled and .info is cleared. I would need to see if there's a way to reproduce this condition with oslo.db. I'm not at the moment seeing any codepath that could produce this unless something else were affecting that .info dictionary. this issue is occurring in several OS projects and I've identified the probable upstream cause here: https://bitbucket.org/zzzeek/sqlalchemy/issues/3497/connectionrec-recycle-can-create-situation this is an oslo.db-level reproduction case:
from oslo_db.sqlalchemy.session import create_engine
engine = create_engine("mysql://scott:tiger@localhost/test")
c1 = engine.connect()
c2 = engine.connect()
c1.scalar("select 1")
c2.scalar("select 1")
c1.close()
c2.close()
raw_input("shutdown")
try:
conn = engine.connect()
except Exception as e:
print "expected error: %s" % e
assert engine.pool._invalidate_time
try:
c2 = engine.connect()
except Exception as e:
print "expected error: %s" % e
c3 = engine.connect()
when the script pauses on "shutdown", shut off the MySQL database, then press enter to watch the failure. The failure will occur in SQLAlchemy 1.0.3 or greater, and should not occur in 1.0.2 or earlier. The underlying issue is still present in those versions as well as the 0.9 series however the specific case of the 'info' dictionary being involved occurs in 1.0.3, due to changes in the connection lifecycle to suit the HAAlchemy project.
Looking here to rebase python-sqlalchemy-1.0.5 to python-sqlalchemy-1.0.8. changelog is at http://docs.sqlalchemy.org/en/rel_1_0/changelog/changelog_10.html. The SQLAlchemy 1.0 series is only doing bugfixes and only a tiny amount of maximally conservative feature adds, all new feature dev is in the 1.1 series now; I've reviewed all changes between 1.0.5 and 1.0.8 and none have any backwards-incompatible implications. python-sqlalchemy-1.0.8-1.el7ost.x86_64 No errors after restarting mariadb on undercloud and deploying new system. Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHBA-2015:1875 |