Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.
For bugs related to Red Hat Enterprise Linux 5 product line. The current stable release is 5.10. For Red Hat Enterprise Linux 6 and above, please visit Red Hat JIRA https://issues.redhat.com/secure/CreateIssue!default.jspa?pid=12332745 to report new issues.

Bug 514627

Summary: qdisk not working with 16 nodes
Product: Red Hat Enterprise Linux 5 Reporter: Nate Straz <nstraz>
Component: cmanAssignee: Lon Hohberger <lhh>
Status: CLOSED DUPLICATE QA Contact: Cluster QE <mspqa-list>
Severity: medium Docs Contact:
Priority: low    
Version: 5.4CC: cluster-maint, edamato
Target Milestone: rc   
Target Release: ---   
Hardware: All   
OS: Linux   
Whiteboard:
Fixed In Version: Doc Type: Bug Fix
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2009-09-22 21:03:56 UTC Type: ---
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:

Description Nate Straz 2009-07-29 22:03:04 UTC
Description of problem:

While trying to configure a 16 node cluster with qdisk, I was not able to obtain qdisk membership on all nodes.  I was able to get at most 7 out of 16 nodes acknowledging membership in qdisk.

I reconfigured the storage array to allow 32 concurrent host-to-lun connections and switched to the deadline I/O scheduler for the device.  None of it worked.


Version-Release number of selected component (if applicable):
kernel-2.6.18-157.el5
cman-2.0.110-1.el5


How reproducible:
Every time with 16 nodes

Steps to Reproduce:
1. Create 16 node cluster
2. mkqdisk
3. service qdiskd start
4. check qdisk membership with cman_tool nodes
  
Actual results:

[nstraz@nsew ~]$ doall west.xml cman_tool nodes -F id,type| grep -w 0 | sort
[west-01] 0 M
[west-02] 0 X
[west-03] 0 X
[west-04] 0 X
[west-05] 0 X
[west-06] 0 M
[west-07] 0 M
[west-08] 0 X
[west-09] 0 M
[west-10] 0 X
[west-11] 0 X
[west-12] 0 M
[west-13] 0 X
[west-14] 0 M
[west-15] 0 M
[west-16] 0 X


Expected results:
[nstraz@nsew ~]$ doall west.xml cman_tool nodes -F id,type| grep -w 0 | sort
[west-01] 0 M
[west-02] 0 M
[west-03] 0 M
[west-04] 0 M
[west-05] 0 M
[west-06] 0 M
[west-07] 0 M
[west-08] 0 M
[west-09] 0 M
[west-10] 0 M
[west-11] 0 M
[west-12] 0 M
[west-13] 0 M
[west-14] 0 M
[west-15] 0 M
[west-16] 0 M

Additional info:

Comment 1 Lon Hohberger 2009-07-30 15:14:48 UTC
Jul 29 09:57:53 west-01 qdiskd[10132]: <warning> qdiskd: read (system call) has hung for 5 seconds
Jul 29 09:57:53 west-01 qdiskd[10132]: <warning> In 5 more seconds, we will be evicted
Jul 29 09:58:00 west-01 openais[9729]: [CMAN ] lost contact with quorum device
Jul 29 10:00:26 west-01 qdiskd[10132]: <warning> qdisk cycle took more than 1 second to complete (157.490000)


There's something causing problems here.

Comment 2 Lon Hohberger 2009-09-22 21:03:56 UTC

*** This bug has been marked as a duplicate of bug 496130 ***