Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 1102530

Summary: Losing Gluster storage links under /rhev/data
Product: [Retired] oVirt Reporter: Peter <doilooksensible>
Component: vdsmAssignee: Nir Soffer <nsoffer>
Status: CLOSED INSUFFICIENT_DATA QA Contact: Gil Klein <gklein>
Severity: high Docs Contact:
Priority: unspecified    
Version: 3.4CC: amureini, bazulay, bugs, doilooksensible, fsimonce, gklein, iheim, mgoldboi, nsoffer, rbalakri, sabose, Sustugriel, yeylon
Target Milestone: ---   
Target Release: 3.5.1   
Hardware: x86_64   
OS: Linux   
Whiteboard: storage
Fixed In Version: Doc Type: Bug Fix
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2014-11-02 19:36:32 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: Storage RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 1193195    

Description Peter 2014-05-29 07:06:29 UTC
My first bug submit, so please excuse where I have missed importatnt stuff.

Description of problem:
I have recently tried to migrate my ovirt 3.3.4 + NFS environment to ovirt 3.4.1 and Gluster.

During that process, I have had issues during import whereby the gluster storage appears to be active in the ovirt gui, but the link to the gluster storage in /rhev/data-center is missing, and the import therefore fails. 

The error always seems to boild down to the following error on engine.log (the following specifically for the create of a new disk image on a gluster storage domain):

2014-05-20 08:51:21,136 INFO  [org.ovirt.engine.core.dal.dbbroker.auditloghandling.AuditLogDirector] (ajp--127.0.0.1-8702-9) [4637af09] Correlation ID: 2b0b55ab, Job ID: 1a583643-e28a-4f09-a39d-46e4fc6d20b8, Call Stack: null, Custom Event ID: -1, Message: Add-Disk operation of rhel-7_Disk1 was initiated on VM rhel-7 by peter.harris.
2014-05-20 08:51:21,137 INFO  [org.ovirt.engine.core.bll.SPMAsyncTask] (ajp--127.0.0.1-8702-9) [4637af09] BaseAsyncTask::startPollingTask: Starting to poll task 720b4d92-1425-478c-8351-4ff827b8f728.
2014-05-20 08:51:28,077 INFO  [org.ovirt.engine.core.bll.AsyncTaskManager] (DefaultQuartzScheduler_Worker-19) Polling and updating Async Tasks: 1 tasks, 1 tasks to poll now
2014-05-20 08:51:28,084 ERROR [org.ovirt.engine.core.vdsbroker.vdsbroker.HSMGetAllTasksStatusesVDSCommand] (DefaultQuartzScheduler_Worker-19) Failed in HSMGetAllTasksStatusesVDS method
2014-05-20 08:51:28,085 INFO  [org.ovirt.engine.core.bll.SPMAsyncTask] (DefaultQuartzScheduler_Worker-19) SPMAsyncTask::PollTask: Polling task 720b4d92-1425-478c-8351-4ff827b8f728 (Parent Command AddDisk, Parameters Type org.ovirt.engine.core.common.asynctasks.AsyncTaskParameters) returned status finished, result 'cleanSuccess'.
2014-05-20 08:51:28,104 ERROR [org.ovirt.engine.core.bll.SPMAsyncTask] (DefaultQuartzScheduler_Worker-19) BaseAsyncTask::LogEndTaskFailure: Task 720b4d92-1425-478c-8351-4ff827b8f728 (Parent Command AddDisk, Parameters Type org.ovirt.engine.core.common.asynctasks.AsyncTaskParameters) ended with failure:^M
-- Result: cleanSuccess^M
-- Message: VDSGenericException: VDSErrorException: Failed to HSMGetAllTasksStatusesVDS, error = [Errno 2] No such file or directory: '/rhev/data-center/06930787-a091-49a3-8217-1418c5a9881e/967aec77-46d5-418b-8979-d0a86389a77b/images/7726b997-7e58-45f8-a5a6-9cb9a689a45a', code = 100,^M
-- Exception: VDSGenericException: VDSErrorException: Failed to HSMGetAllTasksStatusesVDS, error = [Errno 2] No such file or directory: '/rhev/data-center/06930787-a091-49a3-8217-1418c5a9881e/967aec77-46d5-418b-8979-d0a86389a77b/images/7726b997-7e58-45f8-a5a6-9cb9a689a45a', code = 100

Certainly, if I check /rhev/data-center/06930787-a091-49a3-8217-1418c5a9881e/ on the SPM server, there is no 967aec77-46d5-418b-8979-d0a86389a77b subdirectory. The only elements I have are NFS mounts.

There appear to be no errors in the SPM vdsm.log for this disk


Last nightm I also had events whereby vms on one node paused due to unknown storage error:
2014-05-28 20:09:22,788 WARN  [org.ovirt.engine.core.vdsbroker.irsbroker.IrsBrokerCommand] (org.ovirt.thread.pool-6-thread-40) domain 615647e2-1f60-47e1-8e55-be9f7ead6f15:strg-vm_infrastructure in problem. vds: vmhost2
2014-05-28 20:10:38,646 INFO  [org.ovirt.engine.core.vdsbroker.VdsUpdateRunTimeInfo] (DefaultQuartzScheduler_Worker-70) VM inf-ipa02 9175e0c0-6ebd-4882-b6b1-0a50c265c683 moved from Up --> Paused
2014-05-28 20:10:38,706 INFO  [org.ovirt.engine.core.dal.dbbroker.auditloghandling.AuditLogDirector] (DefaultQuartzScheduler_Worker-70) Correlation ID: null, Call Stack: null, Custom Event ID: -1, Message: VM inf-ipa02 has paused due to unknown storage error.
2014-05-28 20:10:38,711 INFO  [org.ovirt.engine.core.vdsbroker.VdsUpdateRunTimeInfo] (DefaultQuartzScheduler_Worker-70) VM inf-zenoss ceb362eb-22f4-488b-bd7b-5db4aa423411 moved from Up --> Paused
2014-05-28 20:10:38,723 INFO  [org.ovirt.engine.core.dal.dbbroker.auditloghandling.AuditLogDirector] (DefaultQuartzScheduler_Worker-70) Correlation ID: null, Call Stack: null, Custom Event ID: -1, Message: VM inf-zenoss has paused due to unknown storage error.
2014-05-28 20:10:38,727 INFO  [org.ovirt.engine.core.vdsbroker.VdsUpdateRunTimeInfo] (DefaultQuartzScheduler_Worker-70) VM sup-CMS 3cbbbe18-6eaf-47d2-b085-0eda878642cd moved from Up --> Paused
2014-05-28 20:10:38,748 INFO  [org.ovirt.engine.core.dal.dbbroker.auditloghandling.AuditLogDirector] (DefaultQuartzScheduler_Worker-70) Correlation ID: null, Call Stack: null, Custom Event ID: -1, Message: VM sup-CMS has paused due to storage I/O problem.
2014-05-28 20:10:38,810 INFO  [org.ovirt.engine.core.vdsbroker.irsbroker.IrsBrokerCommand] (org.ovirt.thread.pool-6-thread-46) Domain 615647e2-1f60-47e1-8e55-be9f7ead6f15:strg-vm_infrastructure recovered from problem. vds: vmhost2
2014-05-28 20:10:38,811 INFO  [org.ovirt.engine.core.vdsbroker.irsbroker.IrsBrokerCommand] (org.ovirt.thread.pool-6-thread-46) Domain 615647e2-1f60-47e1-8e55-be9f7ead6f15:strg-vm_infrastructure has recovered from problem. No active host in the DC is reporting it as problematic, so clearing the domain recovery timer.

Checking this morning, and /rhev/data-center link has disappeared over night. on most of the nodes.

Is there a way of reactivating these links without going through the process of:
- Stop All VMs
- set ALL gluster nodes to maintenance in the data center (NOTE: NFS links are fine)
- Activate All gluster nodes
- restart VMs.

This migration is dead in the water if I can't get this to work, or at least recover with out VM stop/start

Version-Release number of selected component (if applicable):
- gluster 3.5.0 Release 2.el6
- vdsm 4.14.8.1 Release 0.el6
- ovirt-engine 3.4.1 Release 1.el6


How reproducible:
I am not sure how to replicate this I am afraid. It sometimes occurs when I am doing an import, and that could possibly be when I try and import multiple VMs. Importing VMs one by one seems to fail less often. So may be a network timeout?


Steps to Reproduce:
1.
2.
3.

Actual results:


Expected results:


Additional info:

This appears to be a reoccurence of 
https://bugzilla.redhat.com/show_bug.cgi?id=1069772

My environment is as follows:
OLD
vmhost1 - ovirt 3.3.4 - NFS storages

NEW
ovirtmgr - ovirt 3.4.1 (virt only setup) - gluster storage domains
vmhost2 - cluster1
vmhost3 - cluster2
vmhost4 - cluster2
vmhost5 - cluster3
vmhost6 - cluster3

My gluster volumes are created via gluster command line and I have

All hosts are running scientific Linux 6.5, and the intention is to migrate vmhost1 to new environment cluster1.

I have an NFS export storage domain which I am using to migrate VMs from vmhost1.
Volume Name: vol-vminf
Type: Distributed-Replicate
Volume ID: b0b456bb-76e9-42e7-bb95-3415db79d631
Status: Started
Number of Bricks: 2 x 2 = 4
Transport-type: tcp
Bricks:
Brick1: vmhost3:/storage/inf/br-inf
Brick2: vmhost4:/storage/inf/br-inf
Brick3: vmhost5:/storage/inf/br-inf
Brick4: vmhost6:/storage/inf/br-inf
Options Reconfigured:
storage.owner-gid: 36
storage.owner-uid: 36
server.allow-insecure: on

Volume Name: vol-vmimages
Type: Distribute
Volume ID: 91e2cf8b-2662-4c26-b937-84b8f5b62e2b
Status: Started
Number of Bricks: 4
Transport-type: tcp
Bricks:
Brick1: vmhost3:/storage/vmimages/br-vmimages
Brick2: vmhost3:/storage/vmimages/br-vmimages
Brick3: vmhost3:/storage/vmimages/br-vmimages
Brick4: vmhost3:/storage/vmimages/br-vmimages
Options Reconfigured:
storage.owner-gid: 36
storage.owner-uid: 36
server.allow-insecure: on

Comment 1 Peter 2014-05-29 09:23:54 UTC
Sorry, copy and paste finger trouble. Last volume should read:
Volume Name: vol-vmimages
Type: Distribute
Volume ID: 91e2cf8b-2662-4c26-b937-84b8f5b62e2b
Status: Started
Number of Bricks: 4
Transport-type: tcp
Bricks:
Brick1: vmhost3:/storage/vmimages/br-vmimages
Brick2: vmhost4:/storage/vmimages/br-vmimages
Brick3: vmhost5:/storage/vmimages/br-vmimages
Brick4: vmhost6:/storage/vmimages/br-vmimages
Options Reconfigured:
storage.owner-gid: 36
storage.owner-uid: 36
server.allow-insecure: on

Comment 3 Nathan Hill 2014-09-18 16:54:04 UTC
Having a similar issue on Ovirt-Engine 3.4.3-1.el6:

[root@fileserver 00000002-0002-0002-0002-00000000001f]# ll
total 16
lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 196c83f6-7d10-4385-8fd4-54bb0bfd918c -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c
lrwxrwxrwx. 1 vdsm kvm 113 Sep 18 10:07 3bba1f89-7768-4033-a3dc-3522c9e6e3c6 -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_iso/3bba1f89-7768-4033-a3dc-3522c9e6e3c6
lrwxrwxrwx. 1 vdsm kvm 107 Sep 18 10:07 654f4315-2a71-41e8-a763-80e48bdfa1d6 -> /rhev/data-center/mnt/nas.styx.local:_mnt_Storage-RaidZ_virtual_export/654f4315-2a71-41e8-a763-80e48bdfa1d6
lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 mastersd -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c
[root@fileserver 00000002-0002-0002-0002-00000000001f]# pwd
/rhev/data-center/00000002-0002-0002-0002-00000000001f
[root@fileserver 00000002-0002-0002-0002-00000000001f]#

In /rhev/data-center/mnt it correctly lists a glusterSD storage domain and the proper brick. The symbolic link has disappeared from the above directory.

Other bug reports claim this was fixed in 3.4.1. Bug 1069772. This doesn't seem to be the case. I am also unclear as to what causes this.

Is there a course of action without detaching and reattaching the storage domain? Environment is production and global shutdown is not available.

Comment 4 Nathan Hill 2014-09-22 13:35:47 UTC
An extremely hackey workaround for this issue:

1. Shutdown all virtual machines
2. Put Gluster storage domain in maintenance mode on attached datacenter.
3. Activate Gluster Storage Domain
4. Start virtual machines.

This is definitely not a feasible workaround due to the downtime. 

Below displays the proper links for my setup:

[root@fileserver 00000002-0002-0002-0002-00000000001f]# ll
total 20
lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 196c83f6-7d10-4385-8fd4-54bb0bfd918c -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c
lrwxrwxrwx. 1 vdsm kvm 113 Sep 18 10:07 3bba1f89-7768-4033-a3dc-3522c9e6e3c6 -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_iso/3bba1f89-7768-4033-a3dc-3522c9e6e3c6
lrwxrwxrwx. 1 vdsm kvm 107 Sep 18 10:07 654f4315-2a71-41e8-a763-80e48bdfa1d6 -> /rhev/data-center/mnt/nas.styx.local:_mnt_Storage-RaidZ_virtual_export/654f4315-2a71-41e8-a763-80e48bdfa1d6
lrwxrwxrwx. 1 vdsm kvm  94 Sep 21 19:10 a96b9a1a-4dce-4de5-b70b-57111027ee84 -> /rhev/data-center/mnt/glusterSD/virt.styx.local:glusterFS/a96b9a1a-4dce-4de5-b70b-57111027ee84
lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 mastersd -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c
[root@fileserver 00000002-0002-0002-0002-00000000001f]# pwd
/rhev/data-center/00000002-0002-0002-0002-00000000001f
[root@fileserver 00000002-0002-0002-0002-00000000001f]#


I'm curious if, I make note of the storage domain location and the ID number, if I could simply:

<ln -s /rhev/data-center/mnt/glusterSD/virt.styx.local:glusterFS/a96b9a1a-4dce-4de5-b70b-57111027ee84 a96b9a1a-4dce-4de5-b70b-57111027ee84>

If this reoccurs, I will try the above.

Also, I think this might have something to do with SPM switching when using multiple hosts on GlusterFS. If the above works it could be a feasible coding switch if it's missing.

Any thoughts?

Comment 6 Nathan Hill 2014-09-29 16:40:54 UTC
A followup,

Running the following:

<ln -s /rhev/data-center/mnt/glusterSD/virt.styx.local:glusterFS/a96b9a1a-4dce-4de5-b70b-57111027ee84 a96b9a1a-4dce-4de5-b70b-57111027ee84>

Re-added the storage links after a recurrence. No daemons required restarts, no reattaching storage. I would consider this a very feasible workaround for this problem.

Simply make note of your storage UUID's.

Comment 7 Allon Mureinik 2014-09-29 21:12:54 UTC
Nir, is this a dup of bug 1116585?

Comment 8 Allon Mureinik 2014-10-01 07:21:38 UTC
Pushing out to oVirt 3.5.1 as to not block the GA.
However, as noted in comment 7, this is most probably a duplicate report of an issue that's already solved in 3.5.0 (pending Nir's or Federico's confirmation).

Comment 9 Nir Soffer 2014-10-01 12:48:50 UTC
Looks like duplicate of bug 1146401 - not sure which one should be the duplicate.

Comment 10 Nathan Hill 2014-10-03 13:11:05 UTC
The later one? This one dates back to May.

Whether it's fixed in 3.5 or not isn't exactly relevant as we are running 3.4 stable. 

I found that a suitable workaround is to re-add the link manually to the SPM. Can anyone else confirm who has this problem?

Comment 11 Nir Soffer 2014-10-03 15:52:02 UTC
Nathan, we need more info to continue with this bug:

1. Complete logs describing the operation you are trying to do:
vdsm.log
engine.log
sanlock.log

The log should start when you create your gluster storage domain until the links are "lost" and recovered.

2. Please specify vdsm pacakge version on the hosts:
rpm -q vdsm

The version you mention (vdsm 4.14.8.1 Release 0.el6) is very old. Did you try with latest (4.14.17)?

Comment 12 Nir Soffer 2014-10-03 15:52:53 UTC
Adding back need info for Federico, removed by mistake.

Comment 13 Nathan Hill 2014-10-03 16:58:56 UTC
My testing was with latest:

[root@fileserver /]# rpm -qa | grep -i vdsm
vdsm-xmlrpc-4.14.17-0.el6.noarch
vdsm-python-4.14.17-0.el6.x86_64
vdsm-4.14.17-0.el6.x86_64
vdsm-gluster-4.14.17-0.el6.noarch
vdsm-cli-4.14.17-0.el6.noarch
vdsm-python-zombiereaper-4.14.17-0.el6.noarch
vdsm-bootstrap-4.14.17-0.el6.noarch

Stand by while I browse logs.

Comment 14 Federico Simoncelli 2014-10-23 12:40:45 UTC
Fixing needinfo flags. Waiting on further investigation from Nathan.

Comment 15 Nir Soffer 2014-11-02 19:36:32 UTC
We are waiting for info since 2014-10-03 (30 days). Please reopen if you have new information.

Comment 16 Red Hat Bugzilla 2023-09-14 02:08:45 UTC
The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days