Files
SirRoboGarage/docs/selector_negative_value.md
T
SirStone 394b3deeed SETTLED: Pattern alone is the best single gun IN GENERAL, not just vs DrussGT
Closes the one-adversary caveat that blocked shipping `onlyPattern`. Two arms
(`full` vs `onlyPattern`, rack knobs only), one frozen binary from clean HEAD
(a54ae6a, sha256 04d63cd8...), 10 adversaries, 120 battles, real server-side hit
rate, exact two-sided permutation test per adversary.

  adversary      full %   onlyPattern %   diff    exact p   winner
  drussgt         6.88        9.99        +3.11   0.0023    Pattern
  corners        85.10       88.48        +3.38   0.0012    Pattern
  crazy          39.16       48.42        +9.26   0.0006    Pattern
  patternmover   59.97       67.36        +7.39   0.0159    Pattern
  spinbot        83.37       81.20        -2.17   0.4720    full (n.s.)
  ramfire        97.06       95.58        -1.48   0.2214    full (n.s.)
  randommover    46.06       43.15        -2.91   0.3968    full (n.s.)
  sittingduck    96.37       96.42        +0.04   0.9048    tie
  oscillator     65.13       67.08        +1.94   0.5238    tie
  wavesurfer     41.09       42.68        +1.59   0.5159    tie

Pattern significantly WINS on 4 adversaries, ties on 3, and the full rack's three
nominal "wins" are all NON-SIGNIFICANT and only on saturated bots (83-97% hit
rate, where any gun works). **The full rack never significantly beats Pattern
alone on any adversary.** DrussGT replicates across a different frozen binary
(full 6.88 vs 6.93 prior; onlyPattern 9.99 vs 10.36/10.78 prior).

HONEST LIMITATION on the opponents: the real classic SpinBot/Corners/Crazy/RamFire
JARS are NOT runnable through the shim - `tools/robocode_shim/BotHost.java`
hardcodes `loadClass("jk.mega.DrussGT")`, and generalising it needs source edits.
So those four are the Tank Royale SAMPLE-BOT PORTS (the same opponents the shim
README validated against classic captures), not the original jars. Only DrussGT is
a real classic jar via the shim. Stated explicitly rather than implied.

Also noted: `full` in the harness passes TR_RACK_<GUN>=off for every gun, which
yields an empty admitted set and hits the documented FULL fallback ("an empty
membership admits every gun") - i.e. the shipped all-`both` rack. Verified against
the source, not assumed. Liveness proven per arm ([rack] active=..., gun_stats
100% Pattern for onlyPattern) and owner identification validated in 120/120 battles.

Follow-up if a stricter generalisation is wanted: add a generic classic-bot host to
the shim (a source change) and re-run those four as real jars.
2026-09-22 01:49:57 +02:00

17 KiB
Raw Blame History

The gun selector is currently NEGATIVE value

Measured. The best single gun beats the full rack, and not by a little.

GENERALISED (this update): the "vs DrussGT only" caveat is closed. Against ten adversaries, Pattern alone significantly beats the full rack on DrussGT, Corners, Crazy and PatternMover, ties on the other six, and never significantly loses. Ship onlyPattern — see the Generalisation section.

arm runs shots hits real % dmg/run per-run range exact two-sided p vs full
full (shipped rack) 7 3898 270 6.93 159 2.54–10.37 —
onlyPattern 7 4582 494 10.78 287 9.14–11.85 0.0012
onlyKNN 7 4033 207 5.13 119 4.15–6.47 0.1340
onlyLinear 7 3215 105 3.27 65 1.85–4.69 0.0082
onlyGF 7 3193 72 2.25 45 1.52–3.47 0.0012

Firing Pattern alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012. It also fires MORE shots (4582 vs 3898), so it dominates on rate and volume.

This is not "any single gun wins" — full beats onlyLinear, onlyGF and onlyKNN. It is specifically "Pattern alone beats the rack".


Generalisation (SETTLED): Pattern alone wins or ties on ALL ten adversaries

Every number in the section above is against DrussGT. This section closes that caveat. Same protocol (one frozen binary, rack knobs only, server-side events sidecar, exact two-sided permutation test on per-run rates), two arms — full (shipped rack, control) and onlyPattern — against ten adversaries.

MEASURED verdict. onlyPattern significantly beats the full rack on DrussGT (+3.11 pp, p=0.0023), Corners (+3.38 pp, p=0.0012), Crazy (+9.26 pp, p=0.0006) and PatternMover (+7.39 pp, p=0.0159), and is statistically indistinguishable on the other six (p ≥ 0.22). It never significantly loses. The full rack never significantly beats Pattern alone on any adversary. The full rack's only nominal edges are on saturated bots where nearly every shot already hits (SpinBot −2.17 pp p=0.47, RamFire −1.48 pp p=0.22, RandomMover −2.91 pp p=0.40) — none significant.

Direct answer: Pattern is the best single gun in general, not only against DrussGT. The selector is not merely redundant vs this one boss — on the discriminating opponents (DrussGT 7→10 %, Crazy 39→48 %, Corners 85→88 %) removing it is a real, significant gain. There is no adversary in this set that justifies keeping the selector.

MEASURED: every hit rate, damage figure, run count and exact p-value above; the [rack] active= lines; the Pattern-100 % selection mix; the per-battle owner identification. INFERRED: that the result generalises beyond this finite ten-bot sample to "best gun in general"; that the ~83–97 % hit-rate bots (SpinBot, RamFire, SittingDuck) are saturated and therefore weak discriminators rather than genuine ties; and that the Tank Royale sample ports are a fair stand-in for the classic bots they are ported from.

Adversaries actually available (what is a real classic opponent, and what is not)

  • drussgt — the real classic jar through the robocode_shim bridge. This is the only classic jar the shim can host today: tools/robocode_shim/src/robocode_shim/BotHost.java hardcodes loader.loadClass("jk.mega.DrussGT"). Driving the real classic SpinBot/Corners/Crazy/RamFire jars through the shim would require shim source edits, which this measurement job forbids — so they are not available as shim opponents.
  • spinbot, corners, crazy, ramfire — the Tank Royale sample-bot ports (/home/davide/Projects/tank-royale/sample-bots/java/build/archive/…), i.e. the exact opponents §5.8 of tools/robocode_shim/README.md already validated against the classic captures. They are ports, not the original classic jars — stated plainly rather than passed off as the real thing.
  • sittingduck, oscillator, randommover, patternmover, wavesurfer — the in-repo adversaries (common_libs/test_framework/adversaries/). Deliberately weak; included as cheap extra data points, but they are not the evidence base.

Result table

7 runs × 7 rounds for DrussGT and the four sample ports; 5 runs × 7 rounds for the five weak in-repo bots. Real hit rate = server-side events sidecar (fire/hit), not the bot's own attribution.

adversary adversary class arm runs shots hits real % dmg/run per-run range exact p vs full
drussgt shim (real classic jar) full 7 4087 281 6.88 168 5.11-8.83 —
drussgt shim (real classic jar) onlyPattern 7 4285 428 9.99 253 7.84-11.42 0.0023
spinbot TR sample port full 7 848 707 83.37 656 74.66-90.23 —
spinbot TR sample port onlyPattern 7 883 717 81.20 666 77.37-90.48 0.4720
corners TR sample port full 7 1134 965 85.10 662 81.68-87.12 —
corners TR sample port onlyPattern 7 1146 1014 88.48 662 86.71-90.07 0.0012
crazy TR sample port full 7 1213 475 39.16 616 33.79-43.97 —
crazy TR sample port onlyPattern 7 1105 535 48.42 644 46.15-51.27 0.0006
ramfire TR sample port full 7 442 429 97.06 637 93.94-100.00 —
ramfire TR sample port onlyPattern 7 430 411 95.58 627 92.98-98.44 0.2214
sittingduck in-repo (weak) full 5 662 638 96.37 676 93.60-98.06 —
sittingduck in-repo (weak) onlyPattern 5 726 700 96.42 686 95.59-98.26 0.9048
oscillator in-repo (weak) full 5 674 439 65.13 678 61.08-69.30 —
oscillator in-repo (weak) onlyPattern 5 814 546 67.08 644 53.24-80.92 0.5238
randommover in-repo (weak) full 5 1016 468 46.06 654 38.35-51.85 —
randommover in-repo (weak) onlyPattern 5 1008 435 43.15 643 39.37-50.98 0.3968
patternmover in-repo (weak) full 5 697 418 59.97 647 56.25-63.71 —
patternmover in-repo (weak) onlyPattern 5 622 419 67.36 659 62.20-77.39 0.0159
wavesurfer in-repo (weak) full 5 859 353 41.09 576 36.14-48.23 —
wavesurfer in-repo (weak) onlyPattern 5 888 379 42.68 616 37.21-49.24 0.5159

Per-run real hit rates (%)

  drussgt       full         r1=8.37 r2=5.70 r3=7.21 r4=5.11 r5=6.08 r6=6.38 r7=8.83
  drussgt       onlyPattern  r1=10.79 r2=10.79 r3=7.84 r4=8.90 r5=11.11 r6=9.51 r7=11.42
  spinbot       full         r1=88.32 r2=81.73 r3=90.23 r4=74.66 r5=88.31 r6=82.03 r7=80.49
  spinbot       onlyPattern  r1=90.48 r2=85.29 r3=78.17 r4=79.39 r5=80.77 r6=80.47 r7=77.37
  corners       full         r1=86.01 r2=84.09 r3=87.12 r4=86.42 r5=86.33 r6=85.00 r7=81.68
  corners       onlyPattern  r1=89.09 r2=88.32 r3=90.07 r4=88.16 r5=89.50 r6=87.70 r7=86.71
  crazy         full         r1=38.46 r2=43.97 r3=41.42 r4=39.39 r5=39.20 r6=40.26 r7=33.79
  crazy         onlyPattern  r1=46.15 r2=51.27 r3=46.30 r4=50.96 r5=46.90 r6=47.37 r7=49.70
  ramfire       full         r1=98.28 r2=93.94 r3=95.24 r4=100.00 r5=98.48 r6=94.12 r7=100.00
  ramfire       onlyPattern  r1=93.75 r2=98.18 r3=95.38 r4=95.31 r5=95.08 r6=98.44 r7=92.98
  sittingduck   full         r1=97.37 r2=98.06 r3=97.14 r4=96.92 r5=93.60
  sittingduck   onlyPattern  r1=96.64 r2=96.27 r3=95.76 r4=98.26 r5=95.59
  oscillator    full         r1=68.97 r2=61.08 r3=65.00 r4=63.25 r5=69.30
  oscillator    onlyPattern  r1=64.33 r2=53.24 r3=80.92 r4=73.08 r5=72.14
  randommover   full         r1=51.85 r2=48.00 r3=38.35 r4=49.17 r5=47.34
  randommover   onlyPattern  r1=46.94 r2=42.20 r3=50.98 r4=39.37 r5=40.68
  patternmover  full         r1=60.61 r2=56.25 r3=58.50 r4=63.71 r5=61.33
  patternmover  onlyPattern  r1=77.39 r2=67.96 r3=64.34 r4=62.20 r5=66.42
  wavesurfer    full         r1=36.14 r2=40.24 r3=41.76 r4=41.21 r5=48.23
  wavesurfer    onlyPattern  r1=49.24 r2=42.86 r3=40.38 r4=37.21 r5=47.09

Liveness (each arm was live)

The [rack] line printed at process start confirms membership for every run:

  full         [rack] mode=1v1 active=FULL     [rack] mode=melee active=FULL
  onlyPattern  [rack] mode=1v1 active=PATTERN  [rack] mode=melee active=PATTERN

(full emits TR_RACK_<GUN>=off for every gun; the empty admitted set triggers the documented rackActive == "FULL" fallback, which is behaviourally the shipped all-both rack — verified against selectShotPolicy: "An empty membership admits every gun".) The selected-gun mix in gun_stats.jsonl confirms onlyPattern selects Pattern 100 % of ticks on every adversary, while full distributes across the rack (e.g. vs DrussGT: HeadOn 25.5 %, Pattern 18.4 %, KNN 10.0 %, Accel 8.7 %, …).

Protocol (reproduction)

  • One frozen binary for every arm and every adversary, built from clean a54ae6a via git archive (the working tree was dirty from other agents) → /tmp/ModularBot_generalise, sha256 04d63cd87ea834fc6f9c55efb037c511fb79883389259809d57228f6d7e25de1, bot dir name ModularBot.
  • Rack knobs only (TR_RACK_<GUN>=off); no source edits. 8 concurrent battles, distinct output files. ModularBot owner id is not stable across runs (it flipped 1↔2 randomly), so the analysis identifies it independently per battle (power-bin subset for the shim; subject bulletsFired count otherwise). All 120 battles passed the identification check.
  • DrussGT rows replicate the earlier runs across a different frozen binary (full 6.88 % vs the earlier 6.93 %; onlyPattern 9.99 % vs 10.36–10.78 %) — MEASURED cross-build stability.
  • Harness adapted from tools/ab/which_gun_*.sh + which_gun_analyze.py (exact permutation test reused verbatim); generalised copies at /tmp/whichgun_gen/.

Follow-up: a small rack of GOOD guns STILL loses to Pattern alone

The obvious objection to the result above — "the rack is bloated, so of course it loses; prune it" — was tested directly. Same protocol (one frozen binary, rack knobs only, real DrussGT, server-side real hit rate), 6 arms × 7 runs × 7 rounds. lean8/lean6 are the offline audit's recommended racks, with the selector ACTIVE; onlyPattern is the no-selection control.

arm guns (selector active?) runs shots hits real % dmg/run per-run range exact two-sided p vs onlyPattern
onlyPattern (control) Pattern, NO selection 7 4325 448 10.36 264 8.96–11.47 —
lean8 HeadOn, Linear, Circular, Accel, Pattern, GF, KNN, WallBounce 7 3711 234 6.31 146 2.64–12.99 0.0169
lean6 lean8 − HeadOn 7 4113 363 8.83 212 7.35–10.87 0.0262
pairPC Pattern + Circular 7 3862 318 8.23 185 3.36–11.64 0.0460
pairPK Pattern + KNN 7 4571 448 9.80 264 7.90–12.76 0.3998
pairPL Pattern + Linear 7 4109 338 8.23 200 6.14–9.90 0.0035

No small rack beats Pattern alone. Every selector-active arm is worse or (in one case) statistically indistinguishable:

  • lean8 is far worse — 6.31% vs 10.36%, p = 0.0169. Pruning to the audit's recommended good guns did NOT rescue the selector.
  • lean6 (HeadOn removed) is still worse — 8.83% vs 10.36%, p = 0.0262. Removing the single most-over-selected gun helps (+2.5 pp over lean8) but still does not beat no-selection.
  • pairPK is the only near-tie — 9.80% vs 10.36%, p = 0.3998, identical dmg/run. It ties only because the selector happens to pick Pattern 86.4% of the time on a 2-gun rack; it does not beat Pattern.

Verdict: the selector is negative value on a good rack too — DISABLE it

lean6 and lean8 (selector active) are both significantly WORSE than Pattern alone. This contradicts the standing directive to keep the virtual-fitness selection mechanism, so it is stated plainly: the selection apparatus should be disabled (TR_RACK_<every gun but PATTERN>=off) pending a better fitness signal. Keeping it costs ~1.5–4 pp of real hit rate on a good rack.

Why the pruned racks still lose: the same mis-ranking, visible in the mix

The selected-gun mix (liveness proof) shows the virtual signal still over-selects the wrong guns:

  • lean8: HeadOn 46.2% of ticks → only 2.0% real (34/1670); WallBounce 21.0% → 6.5% real (35/540); Pattern only 14.4% → 12.6% real (65/514). HeadOn and WallBounce crowd out the best gun.
  • lean6 (HeadOn gone): Pattern 29.2% → 10.4% real (127/1216) — the best real gun — while KNN 25.3% → 8.1%, Linear 15.2% → 8.0%, Accel 13.8% → 8.7%. The rack is still diluted.
  • pairPL: the selector gives Linear 57.7% of ticks → 6.0% real (103/1712), and Pattern only 42.3% → 11.1% real. Over-selecting Linear is why pairPL is the worst 2-gun arm (p = 0.0035).

The failure is not "the rack is too big". It is that the virtual fitness signal ranks the wrong guns first, on a big rack and on a small one alike.

Ship recommendation

onlyPattern — Pattern alone, selector bypassed. Measured: 10.36% real hit rate, 264 damage/run (vs lean8 6.31%/146 and full 6.93%/159). This replicates the original finding (10.78%/287).

Method note: one frozen binary built from clean e0666a5 via git archive (other agents had common_libs/guns/* dirty and have since committed further changes; this build predates them), sha256 df8d4f2e…. 6 arms × 7 runs × 7 rounds, 8 concurrent bridge battles. Liveness verified for every arm from gun_stats.jsonl; the [rack] active=… line confirmed each arm's membership at process start.

Method

One frozen binary built from clean HEAD via git archive (other agents had common_libs/guns/* dirty), sha256 f02481d8…. Arms selected with the rack knobs only (TR_RACK_<GUN>=off), no source edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates (C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.

Harness preserved at tools/ab/which_gun_*.sh and tools/ab/which_gun_analyze.py.

Why: the virtual fitness signal mis-ranks guns vs real outcomes

From the full arm's own selection mix and per-gun real rates:

  • HeadOn is massively over-selected — 31.4% of ticks, the most real shots (1070), but only 4.5% real. It alone drags the rack down.
  • Pattern has the best virtual rank and near-best real rate (11.9%, real rank 2), yet is selected only 22.6% of the time.
  • Linear's apparent strength was SELECTION BIAS. Conditional on being selected it looked like 15.2% (n=33); its unconditional rate (onlyLinear) is 3.27%. Every earlier per-gun "real rate" in this repo is conditional on selection and is therefore confounded — this experiment is the clean measurement.

What this does NOT yet settle

  • One adversary — SETTLED (generalisation section above): NO LONGER A CAVEAT. The ten-adversary run shows onlyPattern significantly beats the full rack on DrussGT, Corners, Crazy and PatternMover, ties on the other six, and never significantly loses. Pattern is the best single gun in general, not just vs DrussGT. (Supporting evidence: an offline audit found Pattern is the only gun competitive in every distance/speed bucket.)
  • Whether a SMALL good rack beats Pattern alone — SETTLED (follow-up above): NO. A 6-arm follow-up (lean8, lean6, pairPC, pairPK, pairPL) found every selector-active rack worse or tied; lean6 8.83% and lean8 6.31% both lose to Pattern alone 10.36% (p = 0.026 / 0.017). The selector is negative value on a good rack too.
  • The user's directive was to KEEP the virtual-fitness selection mechanism — now directly tested. The follow-up tested the selector on small, good racks instead of assuming either answer; it lost there too, so the directive conflicts with the evidence and disabling is recommended pending a better signal.

Prior context: three failed selection-side attempts

attempt result
hysteresis (commit to incumbent) 7.02% → 5.10%, p=0.002
commitment (remove the random draw) 7.17% → 4.44%, p=0.0012
arrival-accuracy tie-break (rank by path, narrow by point) 7.08%, p=0.88 — null

So the per-tick random draw is load-bearing on three independent measurements, and no attempt to "smarten" the tied band has helped. This experiment shows the problem is one level up: which guns are in the rack, and the fact that the virtual signal ranks them wrongly.