Files
SirRoboGarage/docs/selector_negative_value.md
T
SirStone a54ae6a162 SETTLED: no small rack beats Pattern alone; the selector is negative value on a GOOD rack
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.

6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.

  arm          guns (selector active?)                    real %  dmg/run  p vs onlyPattern
  onlyPattern  Pattern, NO selection                       10.36    264     --
  lean8        HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce  6.31  146  0.0169
  lean6        lean8 - HeadOn                               8.83    212     0.0262
  pairPC       Pattern + Circular                           8.23    185     0.0460
  pairPK       Pattern + KNN                                9.80    264     0.3998
  pairPL       Pattern + Linear                             8.23    200     0.0035

The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).

THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
  lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
  lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
  pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
  pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
  pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
          most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.

VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.

REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.

Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
2026-09-22 01:35:51 +02:00

7.6 KiB
Raw Blame History

The gun selector is currently NEGATIVE value

Measured. The best single gun beats the full rack, and not by a little.

arm runs shots hits real % dmg/run per-run range exact two-sided p vs full
full (shipped rack) 7 3898 270 6.93 159 2.54–10.37 —
onlyPattern 7 4582 494 10.78 287 9.14–11.85 0.0012
onlyKNN 7 4033 207 5.13 119 4.15–6.47 0.1340
onlyLinear 7 3215 105 3.27 65 1.85–4.69 0.0082
onlyGF 7 3193 72 2.25 45 1.52–3.47 0.0012

Firing Pattern alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012. It also fires MORE shots (4582 vs 3898), so it dominates on rate and volume.

This is not "any single gun wins" — full beats onlyLinear, onlyGF and onlyKNN. It is specifically "Pattern alone beats the rack".

Follow-up: a small rack of GOOD guns STILL loses to Pattern alone

The obvious objection to the result above — "the rack is bloated, so of course it loses; prune it" — was tested directly. Same protocol (one frozen binary, rack knobs only, real DrussGT, server-side real hit rate), 6 arms × 7 runs × 7 rounds. lean8/lean6 are the offline audit's recommended racks, with the selector ACTIVE; onlyPattern is the no-selection control.

arm guns (selector active?) runs shots hits real % dmg/run per-run range exact two-sided p vs onlyPattern
onlyPattern (control) Pattern, NO selection 7 4325 448 10.36 264 8.96–11.47 —
lean8 HeadOn, Linear, Circular, Accel, Pattern, GF, KNN, WallBounce 7 3711 234 6.31 146 2.64–12.99 0.0169
lean6 lean8 − HeadOn 7 4113 363 8.83 212 7.35–10.87 0.0262
pairPC Pattern + Circular 7 3862 318 8.23 185 3.36–11.64 0.0460
pairPK Pattern + KNN 7 4571 448 9.80 264 7.90–12.76 0.3998
pairPL Pattern + Linear 7 4109 338 8.23 200 6.14–9.90 0.0035

No small rack beats Pattern alone. Every selector-active arm is worse or (in one case) statistically indistinguishable:

  • lean8 is far worse — 6.31% vs 10.36%, p = 0.0169. Pruning to the audit's recommended good guns did NOT rescue the selector.
  • lean6 (HeadOn removed) is still worse — 8.83% vs 10.36%, p = 0.0262. Removing the single most-over-selected gun helps (+2.5 pp over lean8) but still does not beat no-selection.
  • pairPK is the only near-tie — 9.80% vs 10.36%, p = 0.3998, identical dmg/run. It ties only because the selector happens to pick Pattern 86.4% of the time on a 2-gun rack; it does not beat Pattern.

Verdict: the selector is negative value on a good rack too — DISABLE it

lean6 and lean8 (selector active) are both significantly WORSE than Pattern alone. This contradicts the standing directive to keep the virtual-fitness selection mechanism, so it is stated plainly: the selection apparatus should be disabled (TR_RACK_<every gun but PATTERN>=off) pending a better fitness signal. Keeping it costs ~1.5–4 pp of real hit rate on a good rack.

Why the pruned racks still lose: the same mis-ranking, visible in the mix

The selected-gun mix (liveness proof) shows the virtual signal still over-selects the wrong guns:

  • lean8: HeadOn 46.2% of ticks → only 2.0% real (34/1670); WallBounce 21.0% → 6.5% real (35/540); Pattern only 14.4% → 12.6% real (65/514). HeadOn and WallBounce crowd out the best gun.
  • lean6 (HeadOn gone): Pattern 29.2% → 10.4% real (127/1216) — the best real gun — while KNN 25.3% → 8.1%, Linear 15.2% → 8.0%, Accel 13.8% → 8.7%. The rack is still diluted.
  • pairPL: the selector gives Linear 57.7% of ticks → 6.0% real (103/1712), and Pattern only 42.3% → 11.1% real. Over-selecting Linear is why pairPL is the worst 2-gun arm (p = 0.0035).

The failure is not "the rack is too big". It is that the virtual fitness signal ranks the wrong guns first, on a big rack and on a small one alike.

Ship recommendation

onlyPattern — Pattern alone, selector bypassed. Measured: 10.36% real hit rate, 264 damage/run (vs lean8 6.31%/146 and full 6.93%/159). This replicates the original finding (10.78%/287).

Method note: one frozen binary built from clean e0666a5 via git archive (other agents had common_libs/guns/* dirty and have since committed further changes; this build predates them), sha256 df8d4f2e…. 6 arms × 7 runs × 7 rounds, 8 concurrent bridge battles. Liveness verified for every arm from gun_stats.jsonl; the [rack] active=… line confirmed each arm's membership at process start.

Method

One frozen binary built from clean HEAD via git archive (other agents had common_libs/guns/* dirty), sha256 f02481d8…. Arms selected with the rack knobs only (TR_RACK_<GUN>=off), no source edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates (C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.

Harness preserved at tools/ab/which_gun_*.sh and tools/ab/which_gun_analyze.py.

Why: the virtual fitness signal mis-ranks guns vs real outcomes

From the full arm's own selection mix and per-gun real rates:

  • HeadOn is massively over-selected — 31.4% of ticks, the most real shots (1070), but only 4.5% real. It alone drags the rack down.
  • Pattern has the best virtual rank and near-best real rate (11.9%, real rank 2), yet is selected only 22.6% of the time.
  • Linear's apparent strength was SELECTION BIAS. Conditional on being selected it looked like 15.2% (n=33); its unconditional rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded — this experiment is the clean measurement.

What this does NOT yet settle

  • One adversary. Everything here is vs DrussGT. Pattern should be re-checked against other bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline audit found Pattern is the only gun competitive in every distance/speed bucket.)
  • Whether a SMALL good rack beats Pattern alone — SETTLED (follow-up above): NO. A 6-arm follow-up (lean8, lean6, pairPC, pairPK, pairPL) found every selector-active rack worse or tied; lean6 8.83% and lean8 6.31% both lose to Pattern alone 10.36% (p = 0.026 / 0.017). The selector is negative value on a good rack too.
  • The user's directive was to KEEP the virtual-fitness selection mechanism — now directly tested. The follow-up tested the selector on small, good racks instead of assuming either answer; it lost there too, so the directive conflicts with the evidence and disabling is recommended pending a better signal.

Prior context: three failed selection-side attempts

attempt result
hysteresis (commit to incumbent) 7.02% → 5.10%, p=0.002
commitment (remove the random draw) 7.17% → 4.44%, p=0.0012
arrival-accuracy tie-break (rank by path, narrow by point) 7.08%, p=0.88 — null

So the per-tick random draw is load-bearing on three independent measurements, and no attempt to "smarten" the tied band has helped. This experiment shows the problem is one level up: which guns are in the rack, and the fact that the virtual signal ranks them wrongly.