Bug 1663785 - opensm's custom .launch script uses option no longer valid from upstream
Summary: opensm's custom .launch script uses option no longer valid from upstream
Keywords:
Status: CLOSED RAWHIDE
Alias: None
Product: Fedora
Classification: Fedora
Component: opensm
Version: rawhide
Hardware: x86_64
OS: Linux
unspecified
unspecified
Target Milestone: ---
Assignee: Honggang LI
QA Contact: Fedora Extras Quality Assurance
URL:
Whiteboard:
Depends On:
Blocks:
TreeView+ depends on / blocked
 
Reported: 2019-01-07 04:17 UTC by jamespharvey20
Modified: 2019-01-21 03:38 UTC (History)
8 users (show)

Fixed In Version: opensm-3.3.21-2.fc30
Clone Of:
: 1664575 (view as bug list)
Environment:
Last Closed: 2019-01-21 03:38:15 UTC
Type: Bug
Embargoed:


Attachments (Terms of Use)

Description jamespharvey20 2019-01-07 04:17:26 UTC
Description of problem: Upstream opensm does not provide a systemd service file, and at least in the past did not provide a way to run multiple instances for multiple ports.  Fedora adds additional files for these capabilities.  Fedora's own files work fine for running with a single port, but fail when using Fedora's multiple ports capability addition, because upstream's `opensm` no longer has `--subnet_prefix` as a valid option.  I'm not sure when upstream removed it.


Version-Release number of selected component (if applicable): 3.3.21-1.fc30


How reproducible: 100%


Steps to Reproduce:
1. Install opensm
2. Edit Fedora's opensm.sysconfig, adding in a LINE of GUIDs separated by spaces.
3. `systemctl start opensm`
4. `journalctl -xe | tail -n 1000`


Actual results: `journalctl` will show `opensm` is complaining about the invalid option.  No instances of `opensm` are actually started, just Fedora's `opensm.launch` script.


Expected results: 2 instances of opensm running, one for each port


Additional info:
* Simply editing Fedora's `opensm.launch` and removing the ` --subnet_prefix $SUBNET_PREFIX` appears to work at first glance.  `opensm` doesn't complain about the multiple instances.  I'm assuming Fedora did this, because maybe in the past `opensm` would complain if there were multiple instances on the same subnet prefix, but I'm of course not sure about that.
* Fedora's `opensm.launch` script runs `opensm` in a while/sleep loop like it does, because `opensm` used to have a timing bug that would intermittently cause signal 15 failures on start.  I read and noted this years ago from Fedora, but am not sure from where.  I'm not sure if `opensm` still has this intermittent bug, but it certainly doesn't hurt anything to have it started this way in case.
* Someone started a request for a systemd service file upstream here: https://github.com/linux-rdma/opensm/issues/9
** I noted there the systemd service file Fedora distributes.  No idea if they'll start providing one, or comment on the timing startup bug I mentioned.

Comment 1 jamespharvey20 2019-01-07 04:19:29 UTC
I'll also add that I have no idea if upstream now provides a way to run on multiple ports.

Comment 4 Honggang LI 2019-01-08 12:10:49 UTC
Sorry for late reply. I managed to get a machine with dual IB ports to reproduce this issue.

Comment 5 Honggang LI 2019-01-08 12:49:04 UTC
(In reply to jamespharvey20 from comment #0)
> but fail when using Fedora's multiple ports capability addition, because
> upstream's `opensm` no longer has `--subnet_prefix` as a valid option.  I'm
> not sure when upstream removed it.

The '--subnet_prefix' is a rhel specific option as it was introduced in a rhel
specific opensm patch. That patch was hang on for many years but never got merged
into upstream. When I rebased opensm to latest upstream release, I deleted the
outdated patch as it was broken. Now it introduced this regression issue.

I will create a patch to add the option '--subnet_prefix' into upstream opensm
repo.

[root@rdma05 ~]# opensm -c default.conf 
-------------------------------------------------
OpenSM 3.3.21
 Reading Cached Option File: /etc/rdma/opensm.conf
Command Line Arguments:
 Creating config file template 'default.conf'.
 Log File: /var/log/opensm.log
-------------------------------------------------
[root@rdma05 ~]# grep -n subnet_prefix default.conf
34:subnet_prefix 0xfe80000000000000

As you see, the default configuration file generated by opensm includes "subnet_prefix".
That means opensm can parse such option. We should create an option for it.


> Steps to Reproduce:
> 1. Install opensm
> 2. Edit Fedora's opensm.sysconfig, adding in a LINE of GUIDs separated by
                   ^^^^^^^^^^^^^^^^

In fact, it is /etc/rdma/opensm.conf


> * Simply editing Fedora's `opensm.launch` and removing the ` --subnet_prefix
> $SUBNET_PREFIX` appears to work at first glance.  

It is workaround.

> `opensm` doesn't complain
> about the multiple instances.  I'm assuming Fedora did this, because maybe
> in the past `opensm` would complain if there were multiple instances on the
> same subnet prefix, but I'm of course not sure about that.

I don't see this. Why do you think opensm should complain if there are multiple
instances on the same subnet prefix.


> * Fedora's `opensm.launch` script runs `opensm` in a while/sleep loop like
> it does, because `opensm` used to have a timing bug that would
> intermittently cause signal 15 failures on start.  I read and noted this
> years ago from Fedora, but am not sure from where.  I'm not sure if `opensm`
> still has this intermittent bug, but it certainly doesn't hurt anything to
> have it started this way in case.

I checked internal both fedora and rhel repo of opensm. Failed to find the bug
you mentioned. The sleep command introduced before rhel-7.0. So, I don't understand
"cause signal 15 failure on start".

I also searched bugzilla database, failed to find the bug you mentioned.

If you can provide the bug link, it will be very helpful to narrow the
signal 15 failure issue.

> * Someone started a request for a systemd service file upstream here:
> https://github.com/linux-rdma/opensm/issues/9
> ** I noted there the systemd service file Fedora distributes.  No idea if
> they'll start providing one, or comment on the timing startup bug I
> mentioned.

Will backport it for Fedora once upstream fixed this issue.

Comment 6 Honggang LI 2019-01-08 13:30:04 UTC
https://koji.fedoraproject.org/koji/taskinfo?taskID=31892595

Please test this scratch build. It was built with latest upstream code and the specific patch deleted by me which introduced this regression issue.

 opensm (master)]$ cat 0001-opensm-main.c-Add-subnet_prefix-option.patch
From cf2d6c23f9989e8d150d022ab6fb4f0f530b81de Mon Sep 17 00:00:00 2001
From: Honggang Li <honli>
Date: Tue, 8 Jan 2019 21:08:03 +0800
Subject: [PATCH] opensm/main.c: Add '--subnet_prefix' option

The original patch was wrote by Doug Ledford. We need this patch to
fix a regression for Fedora opensm package.

Signed-off-by: Honggang Li <honli>
---
 man/opensm.8.in | 6 ++++++
 opensm/main.c   | 9 +++++++++
 2 files changed, 15 insertions(+)

diff --git a/man/opensm.8.in b/man/opensm.8.in
index 1ee2b16d17a5..0be24c2f70fd 100644
--- a/man/opensm.8.in
+++ b/man/opensm.8.in
@@ -11,6 +11,7 @@ opensm \- InfiniBand subnet manager and administration (SM/SA)
 [\-g(uid) <GUID in hex>]
 [\-l(mc) <LMC>]
 [\-p(riority) <PRIORITY>]
+[\-\-subnet_prefix <PREFIX in hex>]
 [\-\-smkey <SM_Key>]
 [\-\-sm_sl <SL number>]
 [\-r(eassign_lids)]
@@ -136,6 +137,11 @@ This will effect the handover cases, where master
 is chosen by priority and GUID.  Range goes from 0
 (default and lowest priority) to 15 (highest).
 .TP
+\fB\-\-subnet_prefix\fR <PREFIX in hex>
+This option specifies the subnet prefix to use in
+on the fabric.  The default prefix is
+0xfe80000000000000.
+.TP
 \fB\-\-smkey\fR <SM_Key value>
 This option specifies the SM\'s SM_Key (64 bits).
 This will effect SM authentication.
diff --git a/opensm/main.c b/opensm/main.c
index 0b50b43d75a0..af20a16591c1 100644
--- a/opensm/main.c
+++ b/opensm/main.c
@@ -161,6 +161,9 @@ static void show_usage(void)
 	       "          This will effect the handover cases, where master\n"
 	       "          is chosen by priority and GUID.  Range goes\n"
 	       "          from 0 (lowest priority) to 15 (highest).\n\n");
+	printf("--subnet_prefix <prefix>\n"
+	       "          Set the subnet prefix to something other than the\n"
+	       "          default value of 0xfe80000000000000\n\n");
 	printf("--smkey, -k <SM_Key>\n"
 	       "          This option specifies the SM's SM_Key (64 bits).\n"
 	       "          This will effect SM authentication.\n"
@@ -665,6 +668,7 @@ int main(int argc, char *argv[])
 		{"once", 0, NULL, 'o'},
 		{"reassign_lids", 0, NULL, 'r'},
 		{"priority", 1, NULL, 'p'},
+		{"subnet_prefix", 1, NULL, 16},
 		{"smkey", 1, NULL, 'k'},
 		{"routing_engine", 1, NULL, 'R'},
 		{"ucast_cache", 0, NULL, 'A'},
@@ -1008,6 +1012,11 @@ int main(int argc, char *argv[])
 			printf(" Priority = %d\n", temp);
 			break;
 
+		case 16:
+			opt.subnet_prefix = cl_hton64(strtoull(optarg, NULL, 16));
+			printf(" Subnet_Prefix = <0x%" PRIx64 ">\n", cl_hton64(opt.subnet_prefix));
+			break;
+
 		case 'k':
 			sm_key = cl_hton64(strtoull(optarg, NULL, 16));
 			printf(" SM Key <0x%" PRIx64 ">\n", cl_hton64(sm_key));
-- 
2.15.0-rc1

Comment 7 Doug Ledford 2019-01-08 17:21:02 UTC
(In reply to Honggang LI from comment #5)
> (In reply to jamespharvey20 from comment #0)
> > but fail when using Fedora's multiple ports capability addition, because
> > upstream's `opensm` no longer has `--subnet_prefix` as a valid option.  I'm
> > not sure when upstream removed it.
> 
> The '--subnet_prefix' is a rhel specific option as it was introduced in a
> rhel
> specific opensm patch. That patch was hang on for many years but never got
> merged
> into upstream. When I rebased opensm to latest upstream release, I deleted
> the
> outdated patch as it was broken. Now it introduced this regression issue.
> 
> I will create a patch to add the option '--subnet_prefix' into upstream
> opensm
> repo.

This wasn't added to upstream before because there are a couple other items that also ought to be command line options to make things truly work.  The most important one is the log directory.  In our own cluster we use separate config files because of this.  I would suggest logging into rdma-master and comparing the two opensm config file to see all of the differences and pick out the ones that need to be available on the command line in order for two instances to share a single config file.


> > Steps to Reproduce:
> > 1. Install opensm
> > 2. Edit Fedora's opensm.sysconfig, adding in a LINE of GUIDs separated by
>                    ^^^^^^^^^^^^^^^^
> 
> In fact, it is /etc/rdma/opensm.conf

No, /etc/rdma/opensm.conf is the full opensm config file, /etc/sysconfig/opensm is the file that the launch script uses to determine if it needs to launch multiple opensm instances.  Putting multiple GUIDs into this file causes the launch script to attempt to use the same opensm.conf file for multiple launches and uses command line options to try and differentiate the instances.  Since that's what this bug is about, the /etc/sysconfig/opensm file is the right place to replicate this bug.

> 
> > * Simply editing Fedora's `opensm.launch` and removing the ` --subnet_prefix
> > $SUBNET_PREFIX` appears to work at first glance.  
> 
> It is workaround.

It will allow multiple instances and run multiple fabrics, but they will all have the same subnet prefix and at least openmpi will choke on this.

> > `opensm` doesn't complain
> > about the multiple instances.  I'm assuming Fedora did this, because maybe
> > in the past `opensm` would complain if there were multiple instances on the
> > same subnet prefix, but I'm of course not sure about that.
> 
> I don't see this. Why do you think opensm should complain if there are
> multiple
> instances on the same subnet prefix.

No, it's because other software (openmpi being a primary one) will refuse to run if you have two physically separate subnets with the same subnet prefix.  The subnet prefix is supposed to be unique on each subnet.  It's how software can tell for certain whether or not to end points can communicate with each other.  The subnet prefix is also embedded in things like the 20byte IPoIB MAC address for instance.  If you have two physically separate subnets with the same subnet prefix, it makes it look like the two IPoIB instances should be able to talk to each other when they really can't (the broadcast multicast group will be the same on both).

> 
> > * Fedora's `opensm.launch` script runs `opensm` in a while/sleep loop like
> > it does, because `opensm` used to have a timing bug that would
> > intermittently cause signal 15 failures on start.

No.  We run opensm in a while loop because certain conditions can cause opensm to exit.  But, opensm is supposed to be a long running daemon service.  If I recall correctly, unplugging the cable on the machine running opensm can cause opensm to exit.  Or an interface that can't be brought up can cause opensm to exit.  In case the issue is a transient hardware failure, we run opensm in a while loop.  That way, whenever the issue is resolved, a new instance of opensm will be started and the fabric will be managed.  The alternative makes it such that the wrong transient failure on the host can stop opensm and eventually render the entire fabric offline.  Because opensm really needs to behave as a mission critical daemon where failure/exit is not an option, and where infinite retries until it can manage the fabric again are the norm, and because it *doesn't* behave that way, the while loop was added to resolve the issue.

> >  I read and noted this
> > years ago from Fedora, but am not sure from where.  I'm not sure if `opensm`
> > still has this intermittent bug, but it certainly doesn't hurt anything to
> > have it started this way in case.
> 
> I checked internal both fedora and rhel repo of opensm. Failed to find the
> bug
> you mentioned. The sleep command introduced before rhel-7.0. So, I don't
> understand
> "cause signal 15 failure on start".

The sleep command is because a transient hardware failure, such as card down with a cable unplugged, can result in that while loop pegging a cpu at 100% just restarting opensm infinitely.  This makes log files explode and all sorts of bad things happen.  The sleep makes the retries infinite without being an out of control loop.

> I also searched bugzilla database, failed to find the bug you mentioned.
> 
> If you can provide the bug link, it will be very helpful to narrow the
> signal 15 failure issue.
> 
> > * Someone started a request for a systemd service file upstream here:
> > https://github.com/linux-rdma/opensm/issues/9
> > ** I noted there the systemd service file Fedora distributes.  No idea if
> > they'll start providing one, or comment on the timing startup bug I
> > mentioned.
> 
> Will backport it for Fedora once upstream fixed this issue.

Comment 8 Doug Ledford 2019-01-08 17:27:57 UTC
(In reply to Doug Ledford from comment #7)

> This wasn't added to upstream before because there are a couple other items
> that also ought to be command line options to make things truly work.  The
> most important one is the log directory.  In our own cluster we use separate
> config files because of this.  I would suggest logging into rdma-master and
> comparing the two opensm config file to see all of the differences and pick
> out the ones that need to be available on the command line in order for two
> instances to share a single config file.

Sorry, I forgot we changed the opensm setup in the cluster so rdma-master only runs a single instance now.  Here's the diff between the old config files:

[dledford@haswell-e rdma-testing ((28324a0bcc69...))]$ diff -u opensm.conf.[12]-rdma-master
--- opensm.conf.1-rdma-master	2019-01-08 12:25:43.319507660 -0500
+++ opensm.conf.2-rdma-master	2019-01-08 12:25:43.320507708 -0500
@@ -2,7 +2,7 @@
 # DEVICE ATTRIBUTES OPTIONS
 #
 # The port GUID on which the OpenSM is running
-guid 0xf4521403007be131
+guid 0x001175000078b68e
 
 # M_Key value sent to all ports qualifying all Set(PortInfo)
 m_key 0x0000000000000000
@@ -31,7 +31,7 @@
 # on a little endian machine.
 
 # Subnet prefix used on this subnet
-subnet_prefix 0xfe80000000000000
+subnet_prefix 0xfe80000000000001
 
 # The LMC value used on this subnet
 lmc 1
@@ -120,7 +120,7 @@
 # PARTITIONING OPTIONS
 #
 # Partition configuration file to be used
-partition_config_file /etc/rdma/partitions-ib0.conf
+partition_config_file /etc/rdma/partitions-ib1.conf
 
 # Disable partition enforcement by switches (DEPRECATED)
 # This option is DEPRECATED. Please use part_enforce instead
@@ -396,7 +396,7 @@
 force_log_flush FALSE
 
 # Log file to be used
-log_file /var/log/opensm/ib0/opensm.log
+log_file /var/log/opensm/ib1/opensm.log
 
 # Limit the size of the log file in MB. If overrun, log is restarted
 log_max_size 0
@@ -412,7 +412,7 @@
 per_module_logging_file /etc/rdma/per-module-logging.conf
 
 # The directory to hold the file OpenSM dumps
-dump_files_dir /var/log/opensm/ib0
+dump_files_dir /var/log/opensm/ib1
 
 # If TRUE enables new high risk options and hardware specific quirks
 enable_quirks FALSE


Pretty much everything in this diff needs to be settable from the command line for multiple instances via a single config file and command line option overrides to work.

Comment 9 Honggang LI 2019-01-09 03:56:38 UTC
(In reply to Doug Ledford from comment #7)

> > > Steps to Reproduce:
> > > 1. Install opensm
> > > 2. Edit Fedora's opensm.sysconfig, adding in a LINE of GUIDs separated by
> >                    ^^^^^^^^^^^^^^^^
> > 
> > In fact, it is /etc/rdma/opensm.conf

opps, it is "copy and paste" error. I mean /etc/sysconfig/opensm.
> 

> 
> No, it's because other software (openmpi being a primary one) will refuse to
> run if you have two physically separate subnets with the same subnet prefix.
> The subnet prefix is supposed to be unique on each subnet.  It's how
> software can tell for certain whether or not to end points can communicate
> with each other.  The subnet prefix is also embedded in things like the
> 20byte IPoIB MAC address for instance.  If you have two physically separate
> subnets with the same subnet prefix, it makes it look like the two IPoIB
> instances should be able to talk to each other when they really can't (the
> broadcast multicast group will be the same on both).


> No.  We run opensm in a while loop because certain conditions can cause
> opensm to exit.  But, opensm is supposed to be a long running daemon
> service.  If I recall correctly, unplugging the cable on the machine running
> opensm can cause opensm to exit.  Or an interface that can't be brought up
> can cause opensm to exit.  In case the issue is a transient hardware
> failure, we run opensm in a while loop.  That way, whenever the issue is
> resolved, a new instance of opensm will be started and the fabric will be
> managed.  The alternative makes it such that the wrong transient failure on
> the host can stop opensm and eventually render the entire fabric offline. 
> Because opensm really needs to behave as a mission critical daemon where
> failure/exit is not an option, and where infinite retries until it can
> manage the fabric again are the norm, and because it *doesn't* behave that
> way, the while loop was added to resolve the issue.

Thanks for those comments. Now I see why need the loop.

> The sleep command is because a transient hardware failure, such as card down
> with a cable unplugged, can result in that while loop pegging a cpu at 100%
> just restarting opensm infinitely.  This makes log files explode and all
> sorts of bad things happen.  The sleep makes the retries infinite without
> being an out of control loop.

Comment 10 Honggang LI 2019-01-09 04:05:39 UTC
(In reply to Doug Ledford from comment #8)

Except the "dump_files_dir" option, other options are available in upstream or the patch.

>  # The port GUID on which the OpenSM is running
> -guid 0xf4521403007be131
> +guid 0x001175000078b68e

This option is available in upstream.

--guid, -g <GUID in hex>
          This option specifies the local port GUID value
          with which OpenSM should bind.  OpenSM may be
          bound to 1 port at a time.
          If GUID given is 0, OpenSM displays a list
          of possible port GUIDs and waits for user input.
          Without -g, OpenSM tries to use the default port. 

 
>  # Subnet prefix used on this subnet
> -subnet_prefix 0xfe80000000000000
> +subnet_prefix 0xfe80000000000001

The patch adds this option.
  

>  # Partition configuration file to be used
> -partition_config_file /etc/rdma/partitions-ib0.conf
> +partition_config_file /etc/rdma/partitions-ib1.conf

This option is available in upstream.

--Pconfig, -P <partition-config-file>
          This option defines the optional partition configuration file.
          The default name is '/etc/rdma/partitions.conf'.

>  # Log file to be used
> -log_file /var/log/opensm/ib0/opensm.log
> +log_file /var/log/opensm/ib1/opensm.log

This option is available in upstream.
--log_file, -f <log-file-name>
          This option defines the log to be the given file.
          By default, the log goes to /var/log/opensm.log.
          For the log to go to standard output use -f stdout.


>  # The directory to hold the file OpenSM dumps
> -dump_files_dir /var/log/opensm/ib0
> +dump_files_dir /var/log/opensm/ib1

Well, this one is missing in upstream. I will create a patch for this.

Comment 11 jamespharvey20 2019-01-09 08:34:25 UTC
(In reply to Honggang LI from comment #6)
> https://koji.fedoraproject.org/koji/taskinfo?taskID=31892595
> 
> Please test this scratch build. It was built with latest upstream code and
> the specific patch deleted by me which introduced this regression issue.
> 
>  opensm (master)]$ cat 0001-opensm-main.c-Add-subnet_prefix-option.patch

Your patch works great, thanks!  I'm actually on Arch Linux, so I couldn't use the scratch build, but I applied it on top of upstream's most recent release, 3.3.21.  I'm one of the maintainers for Arch Linux's AUR opensm (and other InfiniBand) packages.  (To be clear, AUR packages are maintained by any user who adopts the packages - I'm not part of the official Arch team.)  I'm using your custom .launch script and systemd .service file, since upstream doesn't provide alternatives.

Using your patch, I see there are two instances of opensm running, each with the proper gid, and each with a distinct subnet_prefix.  On both machines, "ibstatus" shows ACTIVE/LinkUp, and the second port shows the distinct subnet_prefix that the custom .launch script generated.

(In reply to Honggang LI from comment #5)
> (In reply to jamespharvey20 from comment #0)
> > `opensm` doesn't complain
> > about the multiple instances.  I'm assuming Fedora did this, because maybe
> > in the past `opensm` would complain if there were multiple instances on the
> > same subnet prefix, but I'm of course not sure about that.
> 
> I don't see this. Why do you think opensm should complain if there are
> multiple
> instances on the same subnet prefix.

At the time, I wasn't aware that "--subnet_prefix" had been used by the custom .launch script so certain things (like openmpi) wouldn't choke.  Not knowing that, it was just my best guess on why it was being used.

> > * Fedora's `opensm.launch` script runs `opensm` in a while/sleep loop like
> > it does, because `opensm` used to have a timing bug that would
> > intermittently cause signal 15 failures on start.  I read and noted this
> > years ago from Fedora, but am not sure from where.  I'm not sure if `opensm`
> > still has this intermittent bug, but it certainly doesn't hurt anything to
> > have it started this way in case.
> 
> I checked internal both fedora and rhel repo of opensm. Failed to find the
> bug
> you mentioned. The sleep command introduced before rhel-7.0. So, I don't
> understand
> "cause signal 15 failure on start".
> 
> I also searched bugzilla database, failed to find the bug you mentioned.
> 
> If you can provide the bug link, it will be very helpful to narrow the
> signal 15 failure issue.

I was able to find where I got that idea from.  An Arch user (known to use multiple ports) was having a problem where half the time they were starting opensm, it was immediately exiting with signal 15.  They found Fedora's custom .launch script, and said it was done that way because of the same timing bug.  Hearing Doug's explanation, the Arch user who said this must have been mistaken.  The loop acts as a workaround for the bug he was experiencing, as well.  I don't know how feasible it would be, but it would be nice for an upstream patch to change its default behavior, so it wouldn't close in many of the situations that it currently does.


Note You need to log in before you can comment on or make changes to this bug.