Files
SirRoboGarage/docs/selector_negative_value.md
T
SirStone e0666a562d The gun selector is NEGATIVE value: Pattern alone beats the full rack (p=0.0012)
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.

  arm             runs  shots  hits  real %  dmg/run  p vs full
  full (shipped)     7   3898   270    6.93     159      --
  onlyPattern        7   4582   494   10.78     287      0.0012  <- BETTER
  onlyKNN            7   4033   207    5.13     119      0.1340
  onlyLinear         7   3215   105    3.27      65      0.0082
  onlyGF             7   3193    72    2.25      45      0.0012

Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".

WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
  but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
  is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
  looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
  Every earlier per-gun "real rate" in this repo is conditional on selection and
  is therefore confounded. This experiment is the clean measurement.

NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
  other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
  value on the CURRENT bloated rack; that does not prove it is negative value on
  a rack of only good guns. That is the next experiment and it decides whether
  the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
  measurement conflicts with it, so the next step tests the selector on a small
  good rack rather than assuming either answer.

Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.

Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).
2026-09-22 01:21:14 +02:00

3.6 KiB
Raw Blame History

The gun selector is currently NEGATIVE value

Measured. The best single gun beats the full rack, and not by a little.

arm runs shots hits real % dmg/run per-run range exact two-sided p vs full
full (shipped rack) 7 3898 270 6.93 159 2.54–10.37 —
onlyPattern 7 4582 494 10.78 287 9.14–11.85 0.0012
onlyKNN 7 4033 207 5.13 119 4.15–6.47 0.1340
onlyLinear 7 3215 105 3.27 65 1.85–4.69 0.0082
onlyGF 7 3193 72 2.25 45 1.52–3.47 0.0012

Firing Pattern alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012. It also fires MORE shots (4582 vs 3898), so it dominates on rate and volume.

This is not "any single gun wins" — full beats onlyLinear, onlyGF and onlyKNN. It is specifically "Pattern alone beats the rack".

Method

One frozen binary built from clean HEAD via git archive (other agents had common_libs/guns/* dirty), sha256 f02481d8…. Arms selected with the rack knobs only (TR_RACK_<GUN>=off), no source edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates (C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.

Harness preserved at tools/ab/which_gun_*.sh and tools/ab/which_gun_analyze.py.

Why: the virtual fitness signal mis-ranks guns vs real outcomes

From the full arm's own selection mix and per-gun real rates:

  • HeadOn is massively over-selected — 31.4% of ticks, the most real shots (1070), but only 4.5% real. It alone drags the rack down.
  • Pattern has the best virtual rank and near-best real rate (11.9%, real rank 2), yet is selected only 22.6% of the time.
  • Linear's apparent strength was SELECTION BIAS. Conditional on being selected it looked like 15.2% (n=33); its unconditional rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded — this experiment is the clean measurement.

What this does NOT yet settle

  • One adversary. Everything here is vs DrussGT. Pattern should be re-checked against other bots before it becomes the default on the strength of this alone. (Supporting evidence: an offline audit found Pattern is the only gun competitive in every distance/speed bucket.)
  • Whether a SMALL good rack beats Pattern alone. The selector is negative value on the current bloated rack; that does not prove it is negative value on a rack of only good guns. That is the next experiment, and it decides whether the selection apparatus gets fixed or disabled.
  • The user's directive was to KEEP the virtual-fitness selection mechanism. This measurement conflicts with that directive, so the next step is to test the selector on a small, good rack rather than to assume either answer.

Prior context: three failed selection-side attempts

attempt result
hysteresis (commit to incumbent) 7.02% → 5.10%, p=0.002
commitment (remove the random draw) 7.17% → 4.44%, p=0.0012
arrival-accuracy tie-break (rank by path, narrow by point) 7.08%, p=0.88 — null

So the per-tick random draw is load-bearing on three independent measurements, and no attempt to "smarten" the tied band has helped. This experiment shows the problem is one level up: which guns are in the rack, and the fact that the virtual signal ranks them wrongly.