Note: This bug is displayed in read-only format because the product is no longer active in Red Hat Bugzilla.

Bug 884629

Summary: sortkey is not reset back to default for starvation ordering
Product: Red Hat Enterprise MRG Reporter: Lubos Trilety <ltrilety>
Component: condorAssignee: Erik Erlandson <eerlands>
Status: CLOSED NOTABUG QA Contact: MRG Quality Engineering <mrgqe-bugs>
Severity: unspecified Docs Contact:
Priority: unspecified    
Version: DevelopmentCC: matt, tstclair
Target Milestone: ---   
Target Release: ---   
Hardware: Unspecified   
OS: Unspecified   
Whiteboard:
Fixed In Version: Doc Type: Bug Fix
Doc Text:
Story Points: ---
Clone Of: Environment:
Last Closed: 2013-01-09 15:54:35 UTC Type: Bug
Regression: --- Mount Type: ---
Documentation: --- CRM:
Verified Versions: Category: ---
oVirt Team: --- RHEL 7.3 requirements from Atomic Host:
Cloudforms Team: --- Target Upstream Version:
Embargoed:
Bug Depends On:    
Bug Blocks: 785283    

Description Lubos Trilety 2012-12-06 13:00:14 UTC
Description of problem:
By default, when the GROUP_SORT_EXPR is not set in configuration, accounting group negotiation-ordering should be in "starvation-order". It means that group's 'sortkey' is first set to default value (previously FLT_MAX), then it has zero starvation-ratio, as its allocation is nonzero but no jobs are yet running. Then starvation ratio will become 1, as all allocation is filled. And finally, after all jobs are completed, it should goes back to default value. But it never happens instead it is zero again.

Version-Release number of selected component (if applicable):
condor-7.8.7-0.5

How reproducible:
100%

Steps to Reproduce:
1. configuration:
NEGOTIATOR_DEBUG = D_FULLDEBUG
NEGOTIATOR_USE_SLOT_WEIGHTS = FALSE
NEGOTIATOR_INTERVAL = 30

SCHEDD_INTERVAL	= 15

CLAIM_WORKLIFE = 0

NUM_CPUS = 10

# turn off round robin and multiple allocation rounds
HFS_ROUND_ROBIN_RATE = 100000000
HFS_MAX_ALLOCATION_ROUNDS = 1

GROUP_NAMES = a, b

GROUP_QUOTA_a = 5
GROUP_QUOTA_b = 5

GROUP_AUTOREGROUP = FALSE
GROUP_ACCEPT_SURPLUS = FALSE

2. Submit the following file:
cmd = /bin/sleep
args = 60
should_transfer_files = if_needed
when_to_transfer_output = on_exit
+AccountingGroup="a.user"
queue 5
+AccountingGroup="b.user"
queue 10

3. see Negotiator log file
$ cat NegotiatorLog | grep -e WARNING -e sortkey
12/06/12 13:50:18 Group a - sortkey= 3.3e+38
12/06/12 13:50:18 Group b - sortkey= 3.3e+38
12/06/12 13:50:18 Group <none> - sortkey= 3.4e+38
12/06/12 13:50:48 Group a - sortkey= 0
12/06/12 13:50:48 Group b - sortkey= 0
12/06/12 13:50:48 Group <none> - sortkey= 3.4e+38
12/06/12 13:51:18 Group a - sortkey= 0
12/06/12 13:51:18 Group b - sortkey= 0
12/06/12 13:51:18 Group <none> - sortkey= 3.4e+38
12/06/12 13:51:38 Group a - sortkey= 0
12/06/12 13:51:38 Group b - sortkey= 0
12/06/12 13:51:38 Group <none> - sortkey= 3.4e+38
12/06/12 13:52:08 Group a - sortkey= 1
12/06/12 13:52:08 Group b - sortkey= 1
12/06/12 13:52:08 Group <none> - sortkey= 3.4e+38
12/06/12 13:52:38 Group a - sortkey= 1
12/06/12 13:52:38 Group b - sortkey= 1
12/06/12 13:52:38 Group <none> - sortkey= 3.4e+38
12/06/12 13:53:08 Group a - sortkey= 0
12/06/12 13:53:08 Group b - sortkey= 0
12/06/12 13:53:08 Group <none> - sortkey= 3.4e+38
12/06/12 13:53:38 Group a - sortkey= 0
12/06/12 13:53:38 Group b - sortkey= 1
12/06/12 13:53:38 Group <none> - sortkey= 3.4e+38
12/06/12 13:54:09 Group a - sortkey= 0
12/06/12 13:54:09 Group b - sortkey= 1
12/06/12 13:54:09 Group <none> - sortkey= 3.4e+38
12/06/12 13:54:39 Group a - sortkey= 0
12/06/12 13:54:39 Group b - sortkey= 0
12/06/12 13:54:39 Group <none> - sortkey= 3.4e+38
12/06/12 13:55:09 Group a - sortkey= 0
12/06/12 13:55:09 Group b - sortkey= 0
...
  
Actual results:
starvation ratios are zeroed

Expected results:
those values should be set to default again after related jobs complete

Additional info:

Comment 2 Erik Erlandson 2012-12-06 23:21:49 UTC
In the original implementation the default value of GROUP_SORT_EXPR was:

ifThenElse(GroupResourcesAllocated>0.0,GroupResourcesInUse/GroupResourcesAllocated,3.40282e+38)

The denominator was GroupResourcesAllocated, which is resources allocated to the group: that is, the lesser of quota or the number of jobs submitted against the group.   So in the case of my original test writeup on Bug 785283, when there were no jobs, 'allocated' dropped to zero, and it was then set to 'max-float'

In the 7.8 series, the default expression is now:

ifThenElse(AccountingGroup=?="<none>",3.4e+38,ifThenElse(GroupQuota>0,GroupResourcesInUse/GroupQuota,3.3e+38))

See that the denominator is now GroupQuota.  A group's quota does not drop to zero when there are no jobs submitted, and so the expression will evaluate to zero.

My guess about the 1st cycle where it is evaluating to 'max-float' is that the slot ads haven't yet reported to the collector, and so all quotas will be zero.

Comment 3 Erik Erlandson 2013-01-09 15:54:35 UTC
I'm closing out for reasons discussed in Comment 2 and:

https://bugzilla.redhat.com/show_bug.cgi?id=785283#c8