Note: This bug is displayed in read-only format because
the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.
OSP 10 (newton) puddle:
- 2016-10-04.2
The test that is failing is our basic "pingtest."
We have observed this in 2 successive runs with the same puddle, and reproduced
it on a machine in islocation to facilitate debugging. Please contact weshay
or myoung for access. The (failing) test does the following:
1. deploy HA overcloud (undercloud, 3 controller, 1 compute) in virt (using tripleo-quickstart)
2. Launch basic stack template that launches a single VM/domain
3. Attempt to ping VM and get a response.
Observed behavior:
- the console log is empty for the VM created and launched by the heat stack
- unable to connect to console via virsh
We confirmed today the following (with debug help from dansmith):
1. the domain is booting --> bios, but the kernel is not booting (at least far enough to init serial port). This explains why where is not a console log present. Here's a screenshot: http://imgur.com/a/bX79E
2. the VM launched by the heat stack is configured to boot from a cinder volume. The block device is created and is readable, the entire cirros image can be dd'd successfully. This resolved a working hypothesis: even though the volume is created and block device present, the VM (after load bios) was attempting to read initial blocks from the volume and hanging on a read().
3. (later) reproduces without a cinder volume at all, booting from an ephemeral disk. this confirms #2, this is not storage/cinder related.
4. reset on domain (power cycle) seems to not be responsive, or it's rebooting so quickly it's not registering on the VGA console. We did not determine which. However destroying the domain and restarting it yields the domain wedged in a similar fashion.
5. Have reproduced on my own (myoung) hardware (again virt, HA).
We've got initial confirmation that switching CI --> KVM resolves this issue, and are working to land patches and fully validate.
Per discussion, changed subject / focus of this particular issue to be QEMU specific. We still clearly have an bug here, but it's not blocking CI/automation, and this (nested virt + qemu) clearly isn't a recommended customer configuration. Dropping severity to medium to reflect this.
Well...no...there's still a bug here, but it's not something blocking CI any more. This configuration isn't a recommended config for nova, but it seems like there's still a bug here (perhaps deeper, qemu/kvm)...yes?
(hit save too soon)
Also, this reliably reproduces in HA (3 controller), and reliably works with 1 controller. Is this from your perspective not a nova issue because qemu was being used?
Sorry, I am new to this, can you describe the setup in a little more detail ?
Also can you please attach the host dmesg ? And if possible, also the qemu command line on the host when this happens.