Files
SirRoboGarage/docs/selector_negative_value.md
T
SirStone 394b3deeed SETTLED: Pattern alone is the best single gun IN GENERAL, not just vs DrussGT
Closes the one-adversary caveat that blocked shipping `onlyPattern`. Two arms
(`full` vs `onlyPattern`, rack knobs only), one frozen binary from clean HEAD
(a54ae6a, sha256 04d63cd8...), 10 adversaries, 120 battles, real server-side hit
rate, exact two-sided permutation test per adversary.

  adversary      full %   onlyPattern %   diff    exact p   winner
  drussgt         6.88        9.99        +3.11   0.0023    Pattern
  corners        85.10       88.48        +3.38   0.0012    Pattern
  crazy          39.16       48.42        +9.26   0.0006    Pattern
  patternmover   59.97       67.36        +7.39   0.0159    Pattern
  spinbot        83.37       81.20        -2.17   0.4720    full (n.s.)
  ramfire        97.06       95.58        -1.48   0.2214    full (n.s.)
  randommover    46.06       43.15        -2.91   0.3968    full (n.s.)
  sittingduck    96.37       96.42        +0.04   0.9048    tie
  oscillator     65.13       67.08        +1.94   0.5238    tie
  wavesurfer     41.09       42.68        +1.59   0.5159    tie

Pattern significantly WINS on 4 adversaries, ties on 3, and the full rack's three
nominal "wins" are all NON-SIGNIFICANT and only on saturated bots (83-97% hit
rate, where any gun works). **The full rack never significantly beats Pattern
alone on any adversary.** DrussGT replicates across a different frozen binary
(full 6.88 vs 6.93 prior; onlyPattern 9.99 vs 10.36/10.78 prior).

HONEST LIMITATION on the opponents: the real classic SpinBot/Corners/Crazy/RamFire
JARS are NOT runnable through the shim - `tools/robocode_shim/BotHost.java`
hardcodes `loadClass("jk.mega.DrussGT")`, and generalising it needs source edits.
So those four are the Tank Royale SAMPLE-BOT PORTS (the same opponents the shim
README validated against classic captures), not the original jars. Only DrussGT is
a real classic jar via the shim. Stated explicitly rather than implied.

Also noted: `full` in the harness passes TR_RACK_<GUN>=off for every gun, which
yields an empty admitted set and hits the documented FULL fallback ("an empty
membership admits every gun") - i.e. the shipped all-`both` rack. Verified against
the source, not assumed. Liveness proven per arm ([rack] active=..., gun_stats
100% Pattern for onlyPattern) and owner identification validated in 120/120 battles.

Follow-up if a stricter generalisation is wanted: add a generic classic-bot host to
the shim (a source change) and re-run those four as real jars.
2026-09-22 01:49:57 +02:00

255 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The gun selector is currently NEGATIVE value
**Measured. The best single gun beats the full rack, and not by a little.**
> **GENERALISED (this update):** the "vs DrussGT only" caveat is closed. Against **ten** adversaries,
> `Pattern` alone **significantly beats** the full rack on DrussGT, Corners, Crazy and PatternMover,
> **ties** on the other six, and **never significantly loses**. Ship `onlyPattern` — see the
> [Generalisation](#generalisation-settled-pattern-alone-wins-or-ties-on-all-ten-adversaries) section.
| arm | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `full` |
|---|---:|---:|---:|---:|---:|---|---:|
| `full` (shipped rack) | 7 | 3898 | 270 | **6.93** | 159 | 2.54–10.37 | — |
| **`onlyPattern`** | 7 | 4582 | 494 | **10.78** | **287** | 9.14–11.85 | **0.0012** |
| `onlyKNN` | 7 | 4033 | 207 | 5.13 | 119 | 4.15–6.47 | 0.1340 |
| `onlyLinear` | 7 | 3215 | 105 | 3.27 | 65 | 1.85–4.69 | 0.0082 |
| `onlyGF` | 7 | 3193 | 72 | 2.25 | 45 | 1.52–3.47 | 0.0012 |
**Firing `Pattern` alone: +3.85 pp pooled hit rate, +80% damage per run, p = 0.0012.**
It also fires MORE shots (4582 vs 3898), so it dominates on rate *and* volume.
This is **not** "any single gun wins" — `full` beats `onlyLinear`, `onlyGF` and `onlyKNN`.
It is specifically **"Pattern alone beats the rack"**.
---
## Generalisation (SETTLED): `Pattern` alone wins or ties on ALL ten adversaries
Every number in the section above is against DrussGT. This section closes that caveat. Same protocol
(one frozen binary, rack knobs only, server-side events sidecar, exact two-sided permutation test on
per-run rates), two arms — `full` (shipped rack, control) and `onlyPattern` — against **ten** adversaries.
**MEASURED verdict.** `onlyPattern` **significantly beats** the full rack on **DrussGT** (+3.11 pp,
p=0.0023), **Corners** (+3.38 pp, p=0.0012), **Crazy** (+9.26 pp, p=0.0006) and **PatternMover**
(+7.39 pp, p=0.0159), and is **statistically indistinguishable** on the other six (p ≥ 0.22). It
**never significantly loses**. The full rack **never significantly beats** `Pattern` alone on any
adversary. The full rack's only nominal edges are on **saturated** bots where nearly every shot already
hits (SpinBot −2.17 pp p=0.47, RamFire −1.48 pp p=0.22, RandomMover −2.91 pp p=0.40) — none significant.
**Direct answer: `Pattern` is the best single gun in general, not only against DrussGT.** The selector
is not merely redundant vs this one boss — on the discriminating opponents (DrussGT 7→10 %, Crazy
39→48 %, Corners 85→88 %) removing it is a real, significant gain. There is no adversary in this set
that justifies keeping the selector.
**MEASURED:** every hit rate, damage figure, run count and exact p-value above; the `[rack] active=`
lines; the `Pattern`-100 % selection mix; the per-battle owner identification. **INFERRED:** that the
result generalises beyond this finite ten-bot sample to "best gun in general"; that the ~83–97 %
hit-rate bots (SpinBot, RamFire, SittingDuck) are *saturated* and therefore weak discriminators rather
than genuine ties; and that the Tank Royale sample ports are a fair stand-in for the classic bots they
are ported from.
### Adversaries actually available (what is a *real classic* opponent, and what is not)
- **`drussgt`** — the real classic jar through the `robocode_shim` bridge. This is the **only** classic
jar the shim can host today: `tools/robocode_shim/src/robocode_shim/BotHost.java` hardcodes
`loader.loadClass("jk.mega.DrussGT")`. Driving the real classic SpinBot/Corners/Crazy/RamFire jars
through the shim would require shim source edits, which this measurement job forbids — so they are
**not available as shim opponents**.
- **`spinbot`, `corners`, `crazy`, `ramfire`** — the Tank Royale **sample-bot ports**
(`/home/davide/Projects/tank-royale/sample-bots/java/build/archive/…`), i.e. the exact opponents §5.8
of `tools/robocode_shim/README.md` already validated against the classic captures. They are **ports,
not the original classic jars** — stated plainly rather than passed off as the real thing.
- **`sittingduck`, `oscillator`, `randommover`, `patternmover`, `wavesurfer`** — the in-repo adversaries
(`common_libs/test_framework/adversaries/`). Deliberately **weak**; included as cheap extra data
points, but they are not the evidence base.
### Result table
7 runs × 7 rounds for DrussGT and the four sample ports; 5 runs × 7 rounds for the five weak in-repo
bots. Real hit rate = server-side events sidecar (`fire`/`hit`), not the bot's own attribution.
| adversary | adversary class | arm | runs | shots | hits | real % | dmg/run | per-run range | exact p vs `full` |
|---|---|---:|---:|---:|---:|---:|---:|---|---:|
| `drussgt` | shim (real classic jar) | `full` | 7 | 4087 | 281 | 6.88 | 168 | 5.11-8.83 | — |
| `drussgt` | shim (real classic jar) | **`onlyPattern`** | 7 | 4285 | 428 | **9.99** | **253** | 7.84-11.42 | **0.0023** |
| `spinbot` | TR sample port | `full` | 7 | 848 | 707 | 83.37 | 656 | 74.66-90.23 | — |
| `spinbot` | TR sample port | **`onlyPattern`** | 7 | 883 | 717 | **81.20** | **666** | 77.37-90.48 | **0.4720** |
| `corners` | TR sample port | `full` | 7 | 1134 | 965 | 85.10 | 662 | 81.68-87.12 | — |
| `corners` | TR sample port | **`onlyPattern`** | 7 | 1146 | 1014 | **88.48** | **662** | 86.71-90.07 | **0.0012** |
| `crazy` | TR sample port | `full` | 7 | 1213 | 475 | 39.16 | 616 | 33.79-43.97 | — |
| `crazy` | TR sample port | **`onlyPattern`** | 7 | 1105 | 535 | **48.42** | **644** | 46.15-51.27 | **0.0006** |
| `ramfire` | TR sample port | `full` | 7 | 442 | 429 | 97.06 | 637 | 93.94-100.00 | — |
| `ramfire` | TR sample port | **`onlyPattern`** | 7 | 430 | 411 | **95.58** | **627** | 92.98-98.44 | **0.2214** |
| `sittingduck` | in-repo (weak) | `full` | 5 | 662 | 638 | 96.37 | 676 | 93.60-98.06 | — |
| `sittingduck` | in-repo (weak) | **`onlyPattern`** | 5 | 726 | 700 | **96.42** | **686** | 95.59-98.26 | **0.9048** |
| `oscillator` | in-repo (weak) | `full` | 5 | 674 | 439 | 65.13 | 678 | 61.08-69.30 | — |
| `oscillator` | in-repo (weak) | **`onlyPattern`** | 5 | 814 | 546 | **67.08** | **644** | 53.24-80.92 | **0.5238** |
| `randommover` | in-repo (weak) | `full` | 5 | 1016 | 468 | 46.06 | 654 | 38.35-51.85 | — |
| `randommover` | in-repo (weak) | **`onlyPattern`** | 5 | 1008 | 435 | **43.15** | **643** | 39.37-50.98 | **0.3968** |
| `patternmover` | in-repo (weak) | `full` | 5 | 697 | 418 | 59.97 | 647 | 56.25-63.71 | — |
| `patternmover` | in-repo (weak) | **`onlyPattern`** | 5 | 622 | 419 | **67.36** | **659** | 62.20-77.39 | **0.0159** |
| `wavesurfer` | in-repo (weak) | `full` | 5 | 859 | 353 | 41.09 | 576 | 36.14-48.23 | — |
| `wavesurfer` | in-repo (weak) | **`onlyPattern`** | 5 | 888 | 379 | **42.68** | **616** | 37.21-49.24 | **0.5159** |
### Per-run real hit rates (%)
```
drussgt full r1=8.37 r2=5.70 r3=7.21 r4=5.11 r5=6.08 r6=6.38 r7=8.83
drussgt onlyPattern r1=10.79 r2=10.79 r3=7.84 r4=8.90 r5=11.11 r6=9.51 r7=11.42
spinbot full r1=88.32 r2=81.73 r3=90.23 r4=74.66 r5=88.31 r6=82.03 r7=80.49
spinbot onlyPattern r1=90.48 r2=85.29 r3=78.17 r4=79.39 r5=80.77 r6=80.47 r7=77.37
corners full r1=86.01 r2=84.09 r3=87.12 r4=86.42 r5=86.33 r6=85.00 r7=81.68
corners onlyPattern r1=89.09 r2=88.32 r3=90.07 r4=88.16 r5=89.50 r6=87.70 r7=86.71
crazy full r1=38.46 r2=43.97 r3=41.42 r4=39.39 r5=39.20 r6=40.26 r7=33.79
crazy onlyPattern r1=46.15 r2=51.27 r3=46.30 r4=50.96 r5=46.90 r6=47.37 r7=49.70
ramfire full r1=98.28 r2=93.94 r3=95.24 r4=100.00 r5=98.48 r6=94.12 r7=100.00
ramfire onlyPattern r1=93.75 r2=98.18 r3=95.38 r4=95.31 r5=95.08 r6=98.44 r7=92.98
sittingduck full r1=97.37 r2=98.06 r3=97.14 r4=96.92 r5=93.60
sittingduck onlyPattern r1=96.64 r2=96.27 r3=95.76 r4=98.26 r5=95.59
oscillator full r1=68.97 r2=61.08 r3=65.00 r4=63.25 r5=69.30
oscillator onlyPattern r1=64.33 r2=53.24 r3=80.92 r4=73.08 r5=72.14
randommover full r1=51.85 r2=48.00 r3=38.35 r4=49.17 r5=47.34
randommover onlyPattern r1=46.94 r2=42.20 r3=50.98 r4=39.37 r5=40.68
patternmover full r1=60.61 r2=56.25 r3=58.50 r4=63.71 r5=61.33
patternmover onlyPattern r1=77.39 r2=67.96 r3=64.34 r4=62.20 r5=66.42
wavesurfer full r1=36.14 r2=40.24 r3=41.76 r4=41.21 r5=48.23
wavesurfer onlyPattern r1=49.24 r2=42.86 r3=40.38 r4=37.21 r5=47.09
```
### Liveness (each arm was live)
The `[rack]` line printed at process start confirms membership for every run:
```
full [rack] mode=1v1 active=FULL [rack] mode=melee active=FULL
onlyPattern [rack] mode=1v1 active=PATTERN [rack] mode=melee active=PATTERN
```
(`full` emits `TR_RACK_<GUN>=off` for every gun; the empty admitted set triggers the documented
`rackActive == "FULL"` fallback, which is behaviourally the shipped all-`both` rack — verified against
`selectShotPolicy`: "An empty membership admits every gun".) The selected-gun mix in `gun_stats.jsonl`
confirms `onlyPattern` selects `Pattern` **100 %** of ticks on every adversary, while `full` distributes
across the rack (e.g. vs DrussGT: HeadOn 25.5 %, Pattern 18.4 %, KNN 10.0 %, Accel 8.7 %, …).
### Protocol (reproduction)
- **One frozen binary for every arm and every adversary**, built from **clean `a54ae6a`** via
`git archive` (the working tree was dirty from other agents) → `/tmp/ModularBot_generalise`,
sha256 `04d63cd87ea834fc6f9c55efb037c511fb79883389259809d57228f6d7e25de1`, bot dir name `ModularBot`.
- Rack knobs only (`TR_RACK_<GUN>=off`); **no source edits**. 8 concurrent battles, distinct output
files. ModularBot owner id is **not** stable across runs (it flipped 1↔2 randomly), so the analysis
identifies it independently per battle (power-bin subset for the shim; subject `bulletsFired` count
otherwise). All 120 battles passed the identification check.
- DrussGT rows replicate the earlier runs across a different frozen binary (`full` 6.88 % vs the earlier
6.93 %; `onlyPattern` 9.99 % vs 10.36–10.78 %) — **MEASURED** cross-build stability.
- Harness adapted from `tools/ab/which_gun_*.sh` + `which_gun_analyze.py` (exact permutation test
reused verbatim); generalised copies at `/tmp/whichgun_gen/`.
## Follow-up: a small rack of GOOD guns STILL loses to `Pattern` alone
The obvious objection to the result above — "the rack is bloated, so of course it loses; prune it" — was
tested directly. Same protocol (one frozen binary, rack knobs only, real DrussGT, server-side real hit
rate), **6 arms × 7 runs × 7 rounds**. `lean8`/`lean6` are the offline audit's recommended racks, with the
selector ACTIVE; `onlyPattern` is the no-selection control.
| arm | guns (selector active?) | runs | shots | hits | real % | dmg/run | per-run range | exact two-sided p vs `onlyPattern` |
|---|---|---:|---:|---:|---:|---:|---|---:|
| **`onlyPattern`** (control) | Pattern, NO selection | 7 | 4325 | 448 | **10.36** | **264** | 8.96–11.47 | — |
| `lean8` | HeadOn, Linear, Circular, Accel, Pattern, GF, KNN, WallBounce | 7 | 3711 | 234 | 6.31 | 146 | 2.64–12.99 | 0.0169 |
| `lean6` | lean8 − HeadOn | 7 | 4113 | 363 | 8.83 | 212 | 7.35–10.87 | 0.0262 |
| `pairPC` | Pattern + Circular | 7 | 3862 | 318 | 8.23 | 185 | 3.36–11.64 | 0.0460 |
| `pairPK` | Pattern + KNN | 7 | 4571 | 448 | 9.80 | 264 | 7.90–12.76 | 0.3998 |
| `pairPL` | Pattern + Linear | 7 | 4109 | 338 | 8.23 | 200 | 6.14–9.90 | 0.0035 |
**No small rack beats `Pattern` alone.** Every selector-active arm is worse or (in one case) statistically
indistinguishable:
- **`lean8` is far worse** — 6.31% vs 10.36%, **p = 0.0169**. Pruning to the audit's recommended good guns
did NOT rescue the selector.
- **`lean6` (HeadOn removed) is still worse** — 8.83% vs 10.36%, **p = 0.0262**. Removing the single
most-over-selected gun helps (+2.5 pp over lean8) but still does not beat no-selection.
- **`pairPK` is the only near-tie** — 9.80% vs 10.36%, p = 0.3998, identical dmg/run. It ties only because
the selector happens to pick `Pattern` **86.4%** of the time on a 2-gun rack; it does not *beat* Pattern.
### Verdict: the selector is negative value on a good rack too — DISABLE it
`lean6` and `lean8` (selector active) are both significantly WORSE than `Pattern` alone. This contradicts
the standing directive to keep the virtual-fitness selection mechanism, so it is stated plainly: **the
selection apparatus should be disabled (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal.** Keeping it costs ~1.5–4 pp of real hit rate on a good rack.
### Why the pruned racks still lose: the same mis-ranking, visible in the mix
The selected-gun mix (liveness proof) shows the virtual signal still over-selects the wrong guns:
- **`lean8`**: `HeadOn` **46.2%** of ticks → only **2.0% real** (34/1670); `WallBounce` **21.0%** → 6.5%
real (35/540); `Pattern` only **14.4%** → **12.6% real** (65/514). HeadOn and WallBounce crowd out the
best gun.
- **`lean6`** (HeadOn gone): `Pattern` **29.2%** → **10.4% real** (127/1216) — the best real gun — while
`KNN` 25.3% → 8.1%, `Linear` 15.2% → 8.0%, `Accel` 13.8% → 8.7%. The rack is still diluted.
- **`pairPL`**: the selector gives `Linear` **57.7%** of ticks → **6.0% real** (103/1712), and `Pattern`
only 42.3% → **11.1% real**. Over-selecting Linear is why `pairPL` is the *worst* 2-gun arm (p = 0.0035).
The failure is not "the rack is too big". It is that the **virtual fitness signal ranks the wrong guns
first**, on a big rack and on a small one alike.
### Ship recommendation
**`onlyPattern`** — `Pattern` alone, selector bypassed. Measured: **10.36% real hit rate, 264 damage/run**
(vs `lean8` 6.31%/146 and `full` 6.93%/159). This replicates the original finding (10.78%/287).
Method note: one frozen binary built from clean `e0666a5` via `git archive` (other agents had
`common_libs/guns/*` dirty and have since committed further changes; this build predates them), sha256
`df8d4f2e…`. 6 arms × 7 runs × 7 rounds, 8 concurrent bridge battles. Liveness verified for every arm from
`gun_stats.jsonl`; the `[rack] active=…` line confirmed each arm's membership at process start.
## Method
One frozen binary built from clean `HEAD` via `git archive` (other agents had `common_libs/guns/*`
dirty), sha256 `f02481d8…`. Arms selected with the rack knobs only (`TR_RACK_<GUN>=off`), no source
edits. 5 arms × 7 runs × 7 rounds, 8 concurrent bridge battles, real DrussGT, judged ONLY on
server-side real hit rate from the events sidecar. Exact two-sided permutation test on per-run rates
(C(14,7)=3432 splits). Each arm's liveness verified from the selected-gun mix.
Harness preserved at `tools/ab/which_gun_*.sh` and `tools/ab/which_gun_analyze.py`.
## Why: the virtual fitness signal mis-ranks guns vs real outcomes
From the `full` arm's own selection mix and per-gun real rates:
- **`HeadOn` is massively over-selected** — **31.4% of ticks**, the most real shots (1070), but only
**4.5% real**. It alone drags the rack down.
- **`Pattern`** has the best *virtual* rank and near-best *real* rate (11.9%, real rank 2), yet is
selected only **22.6%** of the time.
- **`Linear`'s apparent strength was SELECTION BIAS.** Conditional on being selected it looked like
15.2% (n=33); its *unconditional* rate (`onlyLinear`) is **3.27%**. Every earlier per-gun "real
rate" in this repo is conditional on selection and is therefore confounded — this experiment is
the clean measurement.
## What this does NOT yet settle
- **One adversary — SETTLED (generalisation section above): NO LONGER A CAVEAT.** The ten-adversary run
shows `onlyPattern` significantly beats the full rack on DrussGT, Corners, Crazy and PatternMover, ties
on the other six, and never significantly loses. `Pattern` is the best single gun in general, not just
vs DrussGT. (Supporting evidence: an offline audit found `Pattern` is the only gun competitive in
*every* distance/speed bucket.)
- **Whether a SMALL good rack beats `Pattern` alone — SETTLED (follow-up above): NO.** A 6-arm
follow-up (`lean8`, `lean6`, `pairPC`, `pairPK`, `pairPL`) found every selector-active rack worse or tied;
`lean6` 8.83% and `lean8` 6.31% both lose to `Pattern` alone 10.36% (p = 0.026 / 0.017). The selector is
negative value on a good rack too.
- **The user's directive was to KEEP the virtual-fitness selection mechanism — now directly tested.** The
follow-up tested the selector on small, good racks instead of assuming either answer; it lost there too,
so the directive conflicts with the evidence and disabling is recommended pending a better signal.
## Prior context: three failed selection-side attempts
| attempt | result |
|---|---|
| hysteresis (commit to incumbent) | 7.02% → 5.10%, p=0.002 |
| commitment (remove the random draw) | 7.17% → 4.44%, p=0.0012 |
| arrival-accuracy tie-break (rank by path, narrow by point) | 7.08%, p=0.88 — null |
So the per-tick random draw is load-bearing on three independent measurements, and no attempt to
"smarten" the tied band has helped. This experiment shows the problem is one level up: **which guns
are in the rack, and the fact that the virtual signal ranks them wrongly.**