Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
RHEL Engineering is moving the tracking of its product development work on RHEL 6 through RHEL 9 to Red Hat Jira (issues.redhat.com). If you're a Red Hat customer, please continue to file support cases via the Red Hat customer portal. If you're not, please head to the "RHEL project" in Red Hat Jira and file new tickets here. Individual Bugzilla bugs in the statuses "NEW", "ASSIGNED", and "POST" are being migrated throughout September 2023. Bugs of Red Hat partners with an assigned Engineering Partner Manager (EPM) are migrated in late September as per pre-agreed dates. Bugs against components "kernel", "kernel-rt", and "kpatch" are only migrated if still in "NEW" or "ASSIGNED". If you cannot log in to RH Jira, please consult article #7032570. That failing, please send an e-mail to the RH Jira admins at rh-issues@redhat.com to troubleshoot your issue as a user management inquiry. The email creates a ServiceNow ticket with Red Hat. Individual Bugzilla bugs that are migrated will be moved to status "CLOSED", resolution "MIGRATED", and set with "MigratedToJIRA" in "Keywords". The link to the successor Jira issue will be found under "Links", have a little "two-footprint" icon next to it, and direct you to the "RHEL project" in Red Hat Jira (issue links are of type "https://issues.redhat.com/browse/RHEL-XXXX", where "X" is a digit). This same link will be available in a blue banner at the top of the page informing you that that bug has been migrated.

Bug 1775371

Summary: RHV-H Power outage lvm manager doesn't restart disk
Product: Red Hat Enterprise Linux 7 Reporter: Jason Holt <jholt>
Component: lvm2Assignee: LVM and device-mapper development team <lvm-team>
lvm2 sub component: Activating existing Logical Volumes QA Contact: cluster-qe <cluster-qe>
Status: CLOSED DUPLICATE Docs Contact:
Severity: low    
Priority: unspecified CC: agk, cshao, dougsland, heinzm, jbrassow, jholt, lsvaty, mavital, mcsontos, msnitzer, nlevy, peyu, prajnoha, qiyuan, sbonazzo, shlei, weiwang, yaniwang, yturgema, zkabelac
Version: 7.7   
Target Milestone: rc   
Target Release: ---   
Hardware: x86_64   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: If docs needed, set a value
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2020-09-09 13:47:53 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Attachments:
Description Flags
rdsosreport none

Description Jason Holt 2019-11-21 19:56:22 UTC
Created attachment 1638570 [details]
rdsosreport

Description of problem:
Installed a new hypervisor 2-3 weeks ago
NFS storage backing the Hypervisors
SHE install for RHEV-M
Power outage affected hypervisor
When it reboots it failed to starte and dropped to draccut could not boot prompt.
This is the second time this has happened.

Version-Release number of selected component (if applicable):
4.3.6EL7.7

How reproducible:
install the Async 4.3.6 EL7.7 relese and unplugg the server randomly.  Reboot the server should drop into darcut. 

Steps to Reproduce:
1.
2.
3.

Actual results:


Expected results:
Server should reboot

Additional info:
rmsosreport attach.  It appears that during reboot the LVM manager didn't start the LVM disks.  Turing off the thin_check_executable repaird the drives and activated them.  Need to see why LVM was not starting the disks. 

  total 16
lrwxrwxrwx   1 root 0    7 Oct 23 22:49 bin -> usr/bin
drwxr-xr-x  16 root 0 2800 Nov 19 12:43 dev
-rw-r--r--   1 root 0  409 Nov 19 12:28 dracut-state.sh
-rw-r--r--   1 root 0    2 Oct 23 22:49 early_cpio
drwxr-xr-x  13 root 0    0 Nov 19 12:32 etc
lrwxrwxrwx   1 root 0   23 Oct 23 22:49 init -> usr/lib/systemd/systemd
drwxr-xr-x   3 root 0    0 Oct 23 22:49 kernel
lrwxrwxrwx   1 root 0    7 Oct 23 22:49 lib -> usr/lib
lrwxrwxrwx   1 root 0    9 Oct 23 22:49 lib64 -> usr/lib64
-rw-r--r--   1 root 0    0 Nov 19 13:06 output.txt
dr-xr-xr-x 183 root 0    0 Nov 19 12:28 proc
drwxr-xr-x   2 root 0    0 Oct 23 22:49 root
drwxr-xr-x  12 root 0  340 Nov 19 12:32 run
lrwxrwxrwx   1 root 0    8 Oct 23 22:49 sbin -> usr/sbin
-rwxr-xr-x   1 root 0 3117 Jun 19 13:04 shutdown
dr-xr-xr-x  13 root 0    0 Nov 19 12:36 sys
drwxr-xr-x   2 root 0    0 Oct 23 22:49 sysroot
drwxr-xr-x   2 root 0    0 Nov 19 12:28 tmp
drwxr-xr-x   4 root 0 4096 Jan  1  1970 usb-drive
drwxr-xr-x   8 root 0    0 Oct 23 22:49 usr
drwxr-xr-x   3 root 0    0 Nov 19 12:28 var


 PV /dev/sda2   VG rhvh            lvm2 [<222.57 GiB / 42.68 GiB free]
  Total: 1 [<222.57 GiB] / in use: 1 [<222.57 GiB] / in no VG: 0 [0   ]

 Reading all physical volumes.  This may take a while...
  Found volume group "rhvh" using metadata type lvm2

  ACTIVE            '/dev/rhvh/swap' [<15.69 GiB] inherit
  inactive          '/dev/rhvh/pool00' [<162.20 GiB] inherit
  inactive          '/dev/rhvh/var_log_audit' [2.00 GiB] inherit
  inactive          '/dev/rhvh/var_log' [8.00 GiB] inherit
  inactive          '/dev/rhvh/var' [15.00 GiB] inherit
  inactive          '/dev/rhvh/tmp' [1.00 GiB] inherit
  inactive          '/dev/rhvh/home' [1.00 GiB] inherit
  inactive          '/dev/rhvh/root' [<135.20 GiB] inherit
  inactive          '/dev/rhvh/rhvh-4.3.6.2-0.20190924.0' [<135.20 GiB] inherit
  inactive          '/dev/rhvh/rhvh-4.3.6.2-0.20190924.0+1' [<135.20 GiB] inherit
  inactive          '/dev/rhvh/var_crash' [10.00 GiB] inherit


command to Repair:  lvm change --config 'global{ think_check_executable=""} -v -a y

Comment 4 Sandro Bonazzola 2020-02-04 08:47:56 UTC
Moving to LVM

Comment 5 Marian Csontos 2020-02-18 10:25:05 UTC
Is this bug report about:

1. No traces of failing thin_check in the dmesg, or
2. Need to repair a pool after power failure?

In case it is missing error message, yes, it perhaps should be reported to dmesg.

In case of thin pool needing repair after a power failure NOTABUG I think. The pool needs checking, and in case of inconsistencies repairing and user is expected to check the results of repair.

BTW, the "command to Repair:  lvm change --config 'global{ think_check_executable=""} -v -a y" repairs nothing, it just masks problems, skips repair, and the pool is probably corrupted.

Comment 6 Jonathan Earl Brassow 2020-05-05 14:28:32 UTC
We don't have much time to work on this for rhel7.9.  Can we get an answer for comment5?

It does seem like this is not a bug, but a need to repair the thin-pool(?).

Comment 7 Jason Holt 2020-05-06 11:09:42 UTC
This happened a number of times after power failures, such as power being unplugged. I have not tested this since the initial bug report.  My thought would be that it should check to fix the thin-pool at boot, but I am way out of my depth for what should happen, and the process to get it back up and running was not straightforward or easy.

Comment 8 Zdenek Kabelac 2020-07-16 13:08:46 UTC
When thin-pool is incorrectly stopped  (i.e. unexpected power-off) - it usually DOES REQUIRE thin_check and validate all the mappings are valid/correct. 

It may need to run 'lvconvert --repair'  to fix problems in metadata.

There is known one case where thin-pool metadata are marked as invalid (mismatch in use-count) when this case will be fixed with plain 'thin_check' execution in future.

Since 'disabling' thin_check leads to usable a thin-pool (with likely hidden issues) - I tend to believe this looks like duplicate of bug 1834944.

To confirm the theory - we would need to get uploaded  'metadata' of such unbootable thin-pool - which can be easily obtained just by activate _tmeta LV standalon and grabbing and packing its content to a file with   'dd' & 'gzip'.

Without getting this evindence - the bug likely will be closed as duplicate.

Comment 9 Jonathan Earl Brassow 2020-09-09 13:47:53 UTC
(In reply to Zdenek Kabelac from comment #8)

> Without getting this evindence - the bug likely will be closed as duplicate.

Closing this bug as a duplicate.

*** This bug has been marked as a duplicate of bug 1834944 ***

Comment 11 Red Hat Bugzilla 2023-09-15 00:19:48 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 500 days