feat(gun_harness): scale-aware selector thresholds; default = path + relative

The selection thresholds were calibrated for a rate scale that does not exist.
MEASURED on an exact offline replay of a fogged live WorldState vs DrussGT
(1397 selection ticks), the 0.10 absolute floor fires on 53.0% of point-metric
ticks and forces HeadOn, which has a REAL hit rate of 2.0-4.4% - worst or
near-worst of 13 guns. HeadOn's selection share: 69.1% (abs+point) -> 43.5%
(rel+point). My earlier claim that the floor fires ALWAYS is REFUTED - it is
53%, because bestRate is a max over gun x bin and a >=50-sample bin
occasionally clears 10%. The mechanism is confirmed; the literal statement was
not.

Scale-aware mode (GUN_SELECTOR_MODE, absolute|relative, default relative):
  RelTieMargin   = 0.20  dimensionless FRACTION of bestRate, replacing the
                         fixed 2pp band so the band scales with the metric
  FloorPeakFrac  = 0.25  the floor fires iff bestRate < 0.25 * peakRateRef,
  SelectorWindow = 256   where peakRateRef is the field-best rate over the last
                         256 selection ticks - keeping the original 'don't trust
                         a collapsed field' purpose but only when the field is
                         bad RELATIVE TO ITS OWN RECENT BEST, and counting only
                         guns with >= MinObsBeforeCompete samples so cold-start
                         100% spikes cannot pin HeadOn
  also pools the rate over power bins instead of taking the max over bins, so
  one lucky bin no longer wins
absolute mode is preserved byte-for-byte for rollback.

A/B vs DrussGT, real server hit rate, 3 runs x 10 rounds per config, one frozen
binary:
  absolute+point  3.66 / 2.45 / 5.01   pooled 3.76%
  absolute+path   7.55 / 8.21 / 6.83   pooled 7.57%
  relative+point  7.66 / 6.18 / 5.79   pooled 6.59%
  relative+path   7.15 / 7.55 / 6.90   pooled 7.21%
absolute+point is SEPARATED from all three (p < 0.0001); the other three
OVERLAP each other (p = 0.18-0.64). So the METRIC is the dominant lever and
under path the two threshold models are statistically tied.

DEFAULT SET: metric = path, thresholds = relative. absolute+path was nominally
0.35pp higher but indistinguishable (p = 0.64); relative is the principled
scale-aware fix, is the only model that works under BOTH metrics, and prevents
the point-metric catastrophe if anyone switches back. Shipping absolute would
ship the accidental side-effect this work exists to remove.

STILL NOT SOLVED: the selector remains only a moderate ranker.
Spearman(virtual rank, real rank) is 0.52 for the winning config, 0.36 pooled
for path and 0.04 for point - and it is INCONSISTENT across run sets. The
metric switch won by de-selecting HeadOn, not by ranking guns better. That is
the next problem.

TASK B, report only: do NOT drive selection from raw real hit rates yet.
Only the selected gun fires, so unselected guns get near-zero real shots
(GuessFactor 20, Linear 24 vs HeadOn 733); noise is fatal (n=470 at p=10% gives
+/-2.8pp, most guns n<200 gives +/-5pp+ across a 3-15% spread); and real rate is
conditional on when the gun was selected. A blended signal with forced
exploration and shrinkage is defensible in principle but needs thousands of
shots per gun across many battles. Real rate is best used OFFLINE as the
evaluation metric - which is exactly what this A/B did.

RELATED BUG FLAGGED, not fixed: MinHitRate = 0.40 in bestPower is on the same
wrong scale - no bin ever clears 40%, so once every bin has data, power
selection falls back to bin 0 (power 1.0) late in a round.

Verified: 33/33 guard checks, 11/11 metric checks, tsetlin green, 12/12
offline==online acceptance under the shipped default, run_range rc=0 over 20
fixtures. Adds analyze_selector.nim to measure floor/tie/bestRate/HeadOn-share
per config on any fixture.
This commit is contained in:
2026-09-21 04:33:52 +02:00
parent 3b5d70b7c3
commit dea4dcb574
6 changed files with 320 additions and 58 deletions
+7 -2
View File
@@ -330,10 +330,15 @@ proc testRangeGroundTruthStationary() =
r[0].shots > 300 and r[0].hits == r[0].shots
proc testRangeConstantVelocityLinearWins() =
# Model-specific property: under the point model HeadOn aims at the current
# position and misses a moving target, while Linear leads it. (Under the
# shipped path model every gun scores 100% here, so the check pins the
# metric it was calibrated against.)
let fx = synthesizeConstantVelocity(ticks = 120)
let r = replayFixture(fx, @[makeDriver("HeadOn", HeadOnGun()),
makeDriver("Linear", LinearGun())])
check "range: constant velocity -> Linear beats HeadOn",
makeDriver("Linear", LinearGun())],
metric = bmPoint)
check "range: constant velocity -> Linear beats HeadOn (point model)",
r[1].hitRate() > r[0].hitRate()
proc testEnergyThresholdFixtureRule() =