Note: This bug is displayed in read-only format because
the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.
Created attachment 1638570[details]
rdsosreport
Description of problem:
Installed a new hypervisor 2-3 weeks ago
NFS storage backing the Hypervisors
SHE install for RHEV-M
Power outage affected hypervisor
When it reboots it failed to starte and dropped to draccut could not boot prompt.
This is the second time this has happened.
Version-Release number of selected component (if applicable):
4.3.6EL7.7
How reproducible:
install the Async 4.3.6 EL7.7 relese and unplugg the server randomly. Reboot the server should drop into darcut.
Steps to Reproduce:
1.
2.
3.
Actual results:
Expected results:
Server should reboot
Additional info:
rmsosreport attach. It appears that during reboot the LVM manager didn't start the LVM disks. Turing off the thin_check_executable repaird the drives and activated them. Need to see why LVM was not starting the disks.
total 16
lrwxrwxrwx 1 root 0 7 Oct 23 22:49 bin -> usr/bin
drwxr-xr-x 16 root 0 2800 Nov 19 12:43 dev
-rw-r--r-- 1 root 0 409 Nov 19 12:28 dracut-state.sh
-rw-r--r-- 1 root 0 2 Oct 23 22:49 early_cpio
drwxr-xr-x 13 root 0 0 Nov 19 12:32 etc
lrwxrwxrwx 1 root 0 23 Oct 23 22:49 init -> usr/lib/systemd/systemd
drwxr-xr-x 3 root 0 0 Oct 23 22:49 kernel
lrwxrwxrwx 1 root 0 7 Oct 23 22:49 lib -> usr/lib
lrwxrwxrwx 1 root 0 9 Oct 23 22:49 lib64 -> usr/lib64
-rw-r--r-- 1 root 0 0 Nov 19 13:06 output.txt
dr-xr-xr-x 183 root 0 0 Nov 19 12:28 proc
drwxr-xr-x 2 root 0 0 Oct 23 22:49 root
drwxr-xr-x 12 root 0 340 Nov 19 12:32 run
lrwxrwxrwx 1 root 0 8 Oct 23 22:49 sbin -> usr/sbin
-rwxr-xr-x 1 root 0 3117 Jun 19 13:04 shutdown
dr-xr-xr-x 13 root 0 0 Nov 19 12:36 sys
drwxr-xr-x 2 root 0 0 Oct 23 22:49 sysroot
drwxr-xr-x 2 root 0 0 Nov 19 12:28 tmp
drwxr-xr-x 4 root 0 4096 Jan 1 1970 usb-drive
drwxr-xr-x 8 root 0 0 Oct 23 22:49 usr
drwxr-xr-x 3 root 0 0 Nov 19 12:28 var
PV /dev/sda2 VG rhvh lvm2 [<222.57 GiB / 42.68 GiB free]
Total: 1 [<222.57 GiB] / in use: 1 [<222.57 GiB] / in no VG: 0 [0 ]
Reading all physical volumes. This may take a while...
Found volume group "rhvh" using metadata type lvm2
ACTIVE '/dev/rhvh/swap' [<15.69 GiB] inherit
inactive '/dev/rhvh/pool00' [<162.20 GiB] inherit
inactive '/dev/rhvh/var_log_audit' [2.00 GiB] inherit
inactive '/dev/rhvh/var_log' [8.00 GiB] inherit
inactive '/dev/rhvh/var' [15.00 GiB] inherit
inactive '/dev/rhvh/tmp' [1.00 GiB] inherit
inactive '/dev/rhvh/home' [1.00 GiB] inherit
inactive '/dev/rhvh/root' [<135.20 GiB] inherit
inactive '/dev/rhvh/rhvh-4.3.6.2-0.20190924.0' [<135.20 GiB] inherit
inactive '/dev/rhvh/rhvh-4.3.6.2-0.20190924.0+1' [<135.20 GiB] inherit
inactive '/dev/rhvh/var_crash' [10.00 GiB] inherit
command to Repair: lvm change --config 'global{ think_check_executable=""} -v -a y
Is this bug report about:
1. No traces of failing thin_check in the dmesg, or
2. Need to repair a pool after power failure?
In case it is missing error message, yes, it perhaps should be reported to dmesg.
In case of thin pool needing repair after a power failure NOTABUG I think. The pool needs checking, and in case of inconsistencies repairing and user is expected to check the results of repair.
BTW, the "command to Repair: lvm change --config 'global{ think_check_executable=""} -v -a y" repairs nothing, it just masks problems, skips repair, and the pool is probably corrupted.
Comment 6Jonathan Earl Brassow
2020-05-05 14:28:32 UTC
We don't have much time to work on this for rhel7.9. Can we get an answer for comment5?
It does seem like this is not a bug, but a need to repair the thin-pool(?).
This happened a number of times after power failures, such as power being unplugged. I have not tested this since the initial bug report. My thought would be that it should check to fix the thin-pool at boot, but I am way out of my depth for what should happen, and the process to get it back up and running was not straightforward or easy.
When thin-pool is incorrectly stopped (i.e. unexpected power-off) - it usually DOES REQUIRE thin_check and validate all the mappings are valid/correct.
It may need to run 'lvconvert --repair' to fix problems in metadata.
There is known one case where thin-pool metadata are marked as invalid (mismatch in use-count) when this case will be fixed with plain 'thin_check' execution in future.
Since 'disabling' thin_check leads to usable a thin-pool (with likely hidden issues) - I tend to believe this looks like duplicate of bug 1834944.
To confirm the theory - we would need to get uploaded 'metadata' of such unbootable thin-pool - which can be easily obtained just by activate _tmeta LV standalon and grabbing and packing its content to a file with 'dd' & 'gzip'.
Without getting this evindence - the bug likely will be closed as duplicate.
Comment 9Jonathan Earl Brassow
2020-09-09 13:47:53 UTC
(In reply to Zdenek Kabelac from comment #8)
> Without getting this evindence - the bug likely will be closed as duplicate.
Closing this bug as a duplicate.
*** This bug has been marked as a duplicate of bug 1834944 ***
Comment 11Red Hat Bugzilla
2023-09-15 00:19:48 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days
Created attachment 1638570 [details] rdsosreport Description of problem: Installed a new hypervisor 2-3 weeks ago NFS storage backing the Hypervisors SHE install for RHEV-M Power outage affected hypervisor When it reboots it failed to starte and dropped to draccut could not boot prompt. This is the second time this has happened. Version-Release number of selected component (if applicable): 4.3.6EL7.7 How reproducible: install the Async 4.3.6 EL7.7 relese and unplugg the server randomly. Reboot the server should drop into darcut. Steps to Reproduce: 1. 2. 3. Actual results: Expected results: Server should reboot Additional info: rmsosreport attach. It appears that during reboot the LVM manager didn't start the LVM disks. Turing off the thin_check_executable repaird the drives and activated them. Need to see why LVM was not starting the disks. total 16 lrwxrwxrwx 1 root 0 7 Oct 23 22:49 bin -> usr/bin drwxr-xr-x 16 root 0 2800 Nov 19 12:43 dev -rw-r--r-- 1 root 0 409 Nov 19 12:28 dracut-state.sh -rw-r--r-- 1 root 0 2 Oct 23 22:49 early_cpio drwxr-xr-x 13 root 0 0 Nov 19 12:32 etc lrwxrwxrwx 1 root 0 23 Oct 23 22:49 init -> usr/lib/systemd/systemd drwxr-xr-x 3 root 0 0 Oct 23 22:49 kernel lrwxrwxrwx 1 root 0 7 Oct 23 22:49 lib -> usr/lib lrwxrwxrwx 1 root 0 9 Oct 23 22:49 lib64 -> usr/lib64 -rw-r--r-- 1 root 0 0 Nov 19 13:06 output.txt dr-xr-xr-x 183 root 0 0 Nov 19 12:28 proc drwxr-xr-x 2 root 0 0 Oct 23 22:49 root drwxr-xr-x 12 root 0 340 Nov 19 12:32 run lrwxrwxrwx 1 root 0 8 Oct 23 22:49 sbin -> usr/sbin -rwxr-xr-x 1 root 0 3117 Jun 19 13:04 shutdown dr-xr-xr-x 13 root 0 0 Nov 19 12:36 sys drwxr-xr-x 2 root 0 0 Oct 23 22:49 sysroot drwxr-xr-x 2 root 0 0 Nov 19 12:28 tmp drwxr-xr-x 4 root 0 4096 Jan 1 1970 usb-drive drwxr-xr-x 8 root 0 0 Oct 23 22:49 usr drwxr-xr-x 3 root 0 0 Nov 19 12:28 var PV /dev/sda2 VG rhvh lvm2 [<222.57 GiB / 42.68 GiB free] Total: 1 [<222.57 GiB] / in use: 1 [<222.57 GiB] / in no VG: 0 [0 ] Reading all physical volumes. This may take a while... Found volume group "rhvh" using metadata type lvm2 ACTIVE '/dev/rhvh/swap' [<15.69 GiB] inherit inactive '/dev/rhvh/pool00' [<162.20 GiB] inherit inactive '/dev/rhvh/var_log_audit' [2.00 GiB] inherit inactive '/dev/rhvh/var_log' [8.00 GiB] inherit inactive '/dev/rhvh/var' [15.00 GiB] inherit inactive '/dev/rhvh/tmp' [1.00 GiB] inherit inactive '/dev/rhvh/home' [1.00 GiB] inherit inactive '/dev/rhvh/root' [<135.20 GiB] inherit inactive '/dev/rhvh/rhvh-4.3.6.2-0.20190924.0' [<135.20 GiB] inherit inactive '/dev/rhvh/rhvh-4.3.6.2-0.20190924.0+1' [<135.20 GiB] inherit inactive '/dev/rhvh/var_crash' [10.00 GiB] inherit command to Repair: lvm change --config 'global{ think_check_executable=""} -v -a y