Bug 1102530
| Summary: | Losing Gluster storage links under /rhev/data | ||
|---|---|---|---|
| Product: | [Retired] oVirt | Reporter: | Peter <doilooksensible> |
| Component: | vdsm | Assignee: | Nir Soffer <nsoffer> |
| Status: | CLOSED INSUFFICIENT_DATA | QA Contact: | Gil Klein <gklein> |
| Severity: | high | Docs Contact: | |
| Priority: | unspecified | ||
| Version: | 3.4 | CC: | amureini, bazulay, bugs, doilooksensible, fsimonce, gklein, iheim, mgoldboi, nsoffer, rbalakri, sabose, Sustugriel, yeylon |
| Target Milestone: | --- | ||
| Target Release: | 3.5.1 | ||
| Hardware: | x86_64 | ||
| OS: | Linux | ||
| Whiteboard: | storage | ||
| Fixed In Version: | Doc Type: | Bug Fix | |
| Doc Text: | Story Points: | --- | |
| Clone Of: | Environment: | ||
| Last Closed: | 2014-11-02 19:36:32 UTC | Type: | Bug |
| Regression: | --- | Mount Type: | --- |
| Documentation: | --- | CRM: | |
| Verified Versions: | Category: | --- | |
| oVirt Team: | Storage | RHEL 7.3 requirements from Atomic Host: | |
| Cloudforms Team: | --- | Target Upstream Version: | |
| Embargoed: | |||
| Bug Depends On: | |||
| Bug Blocks: | 1193195 | ||
|
Description
Peter
2014-05-29 07:06:29 UTC
Sorry, copy and paste finger trouble. Last volume should read: Volume Name: vol-vmimages Type: Distribute Volume ID: 91e2cf8b-2662-4c26-b937-84b8f5b62e2b Status: Started Number of Bricks: 4 Transport-type: tcp Bricks: Brick1: vmhost3:/storage/vmimages/br-vmimages Brick2: vmhost4:/storage/vmimages/br-vmimages Brick3: vmhost5:/storage/vmimages/br-vmimages Brick4: vmhost6:/storage/vmimages/br-vmimages Options Reconfigured: storage.owner-gid: 36 storage.owner-uid: 36 server.allow-insecure: on Having a similar issue on Ovirt-Engine 3.4.3-1.el6: [root@fileserver 00000002-0002-0002-0002-00000000001f]# ll total 16 lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 196c83f6-7d10-4385-8fd4-54bb0bfd918c -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c lrwxrwxrwx. 1 vdsm kvm 113 Sep 18 10:07 3bba1f89-7768-4033-a3dc-3522c9e6e3c6 -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_iso/3bba1f89-7768-4033-a3dc-3522c9e6e3c6 lrwxrwxrwx. 1 vdsm kvm 107 Sep 18 10:07 654f4315-2a71-41e8-a763-80e48bdfa1d6 -> /rhev/data-center/mnt/nas.styx.local:_mnt_Storage-RaidZ_virtual_export/654f4315-2a71-41e8-a763-80e48bdfa1d6 lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 mastersd -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c [root@fileserver 00000002-0002-0002-0002-00000000001f]# pwd /rhev/data-center/00000002-0002-0002-0002-00000000001f [root@fileserver 00000002-0002-0002-0002-00000000001f]# In /rhev/data-center/mnt it correctly lists a glusterSD storage domain and the proper brick. The symbolic link has disappeared from the above directory. Other bug reports claim this was fixed in 3.4.1. Bug 1069772. This doesn't seem to be the case. I am also unclear as to what causes this. Is there a course of action without detaching and reattaching the storage domain? Environment is production and global shutdown is not available. An extremely hackey workaround for this issue: 1. Shutdown all virtual machines 2. Put Gluster storage domain in maintenance mode on attached datacenter. 3. Activate Gluster Storage Domain 4. Start virtual machines. This is definitely not a feasible workaround due to the downtime. Below displays the proper links for my setup: [root@fileserver 00000002-0002-0002-0002-00000000001f]# ll total 20 lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 196c83f6-7d10-4385-8fd4-54bb0bfd918c -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c lrwxrwxrwx. 1 vdsm kvm 113 Sep 18 10:07 3bba1f89-7768-4033-a3dc-3522c9e6e3c6 -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_iso/3bba1f89-7768-4033-a3dc-3522c9e6e3c6 lrwxrwxrwx. 1 vdsm kvm 107 Sep 18 10:07 654f4315-2a71-41e8-a763-80e48bdfa1d6 -> /rhev/data-center/mnt/nas.styx.local:_mnt_Storage-RaidZ_virtual_export/654f4315-2a71-41e8-a763-80e48bdfa1d6 lrwxrwxrwx. 1 vdsm kvm 94 Sep 21 19:10 a96b9a1a-4dce-4de5-b70b-57111027ee84 -> /rhev/data-center/mnt/glusterSD/virt.styx.local:glusterFS/a96b9a1a-4dce-4de5-b70b-57111027ee84 lrwxrwxrwx. 1 vdsm kvm 116 Sep 18 10:07 mastersd -> /rhev/data-center/mnt/fileserver.styx.local:_media_DataPartRaid5_virtual_images/196c83f6-7d10-4385-8fd4-54bb0bfd918c [root@fileserver 00000002-0002-0002-0002-00000000001f]# pwd /rhev/data-center/00000002-0002-0002-0002-00000000001f [root@fileserver 00000002-0002-0002-0002-00000000001f]# I'm curious if, I make note of the storage domain location and the ID number, if I could simply: <ln -s /rhev/data-center/mnt/glusterSD/virt.styx.local:glusterFS/a96b9a1a-4dce-4de5-b70b-57111027ee84 a96b9a1a-4dce-4de5-b70b-57111027ee84> If this reoccurs, I will try the above. Also, I think this might have something to do with SPM switching when using multiple hosts on GlusterFS. If the above works it could be a feasible coding switch if it's missing. Any thoughts? A followup, Running the following: <ln -s /rhev/data-center/mnt/glusterSD/virt.styx.local:glusterFS/a96b9a1a-4dce-4de5-b70b-57111027ee84 a96b9a1a-4dce-4de5-b70b-57111027ee84> Re-added the storage links after a recurrence. No daemons required restarts, no reattaching storage. I would consider this a very feasible workaround for this problem. Simply make note of your storage UUID's. Nir, is this a dup of bug 1116585? Pushing out to oVirt 3.5.1 as to not block the GA. However, as noted in comment 7, this is most probably a duplicate report of an issue that's already solved in 3.5.0 (pending Nir's or Federico's confirmation). Looks like duplicate of bug 1146401 - not sure which one should be the duplicate. The later one? This one dates back to May. Whether it's fixed in 3.5 or not isn't exactly relevant as we are running 3.4 stable. I found that a suitable workaround is to re-add the link manually to the SPM. Can anyone else confirm who has this problem? Nathan, we need more info to continue with this bug: 1. Complete logs describing the operation you are trying to do: vdsm.log engine.log sanlock.log The log should start when you create your gluster storage domain until the links are "lost" and recovered. 2. Please specify vdsm pacakge version on the hosts: rpm -q vdsm The version you mention (vdsm 4.14.8.1 Release 0.el6) is very old. Did you try with latest (4.14.17)? Adding back need info for Federico, removed by mistake. My testing was with latest: [root@fileserver /]# rpm -qa | grep -i vdsm vdsm-xmlrpc-4.14.17-0.el6.noarch vdsm-python-4.14.17-0.el6.x86_64 vdsm-4.14.17-0.el6.x86_64 vdsm-gluster-4.14.17-0.el6.noarch vdsm-cli-4.14.17-0.el6.noarch vdsm-python-zombiereaper-4.14.17-0.el6.noarch vdsm-bootstrap-4.14.17-0.el6.noarch Stand by while I browse logs. Fixing needinfo flags. Waiting on further investigation from Nathan. We are waiting for info since 2014-10-03 (30 days). Please reopen if you have new information. The needinfo request[s] on this closed bug have been removed as they have been unresolved for 1000 days |