Closes the one-adversary caveat that blocked shipping `onlyPattern`. Two arms
(`full` vs `onlyPattern`, rack knobs only), one frozen binary from clean HEAD
(a54ae6a, sha256 04d63cd8...), 10 adversaries, 120 battles, real server-side hit
rate, exact two-sided permutation test per adversary.
adversary full % onlyPattern % diff exact p winner
drussgt 6.88 9.99 +3.11 0.0023 Pattern
corners 85.10 88.48 +3.38 0.0012 Pattern
crazy 39.16 48.42 +9.26 0.0006 Pattern
patternmover 59.97 67.36 +7.39 0.0159 Pattern
spinbot 83.37 81.20 -2.17 0.4720 full (n.s.)
ramfire 97.06 95.58 -1.48 0.2214 full (n.s.)
randommover 46.06 43.15 -2.91 0.3968 full (n.s.)
sittingduck 96.37 96.42 +0.04 0.9048 tie
oscillator 65.13 67.08 +1.94 0.5238 tie
wavesurfer 41.09 42.68 +1.59 0.5159 tie
Pattern significantly WINS on 4 adversaries, ties on 3, and the full rack's three
nominal "wins" are all NON-SIGNIFICANT and only on saturated bots (83-97% hit
rate, where any gun works). **The full rack never significantly beats Pattern
alone on any adversary.** DrussGT replicates across a different frozen binary
(full 6.88 vs 6.93 prior; onlyPattern 9.99 vs 10.36/10.78 prior).
HONEST LIMITATION on the opponents: the real classic SpinBot/Corners/Crazy/RamFire
JARS are NOT runnable through the shim - `tools/robocode_shim/BotHost.java`
hardcodes `loadClass("jk.mega.DrussGT")`, and generalising it needs source edits.
So those four are the Tank Royale SAMPLE-BOT PORTS (the same opponents the shim
README validated against classic captures), not the original jars. Only DrussGT is
a real classic jar via the shim. Stated explicitly rather than implied.
Also noted: `full` in the harness passes TR_RACK_<GUN>=off for every gun, which
yields an empty admitted set and hits the documented FULL fallback ("an empty
membership admits every gun") - i.e. the shipped all-`both` rack. Verified against
the source, not assumed. Liveness proven per arm ([rack] active=..., gun_stats
100% Pattern for onlyPattern) and owner identification validated in 120/120 battles.
Follow-up if a stricter generalisation is wanted: add a generic classic-bot host to
the shim (a source change) and re-run those four as real jars.
Follow-up to e0666a5, which showed Pattern alone (10.78%) beats the full rack
(6.93%). That left two open questions: is a SMALL rack of good guns better than
Pattern alone, and does the selector add value on a good rack (rather than only
on the bloated one)? Both are now answered: NO and NO.
6 arms x 7 runs x 7 rounds, one frozen binary from CLEAN HEAD e0666a5 (built via
`git archive`, source verified byte-identical to the clean tree), rack knobs
only, 8 concurrent battles, real server-side hit rate vs the real DrussGT, exact
two-sided permutation test on per-run rates.
arm guns (selector active?) real % dmg/run p vs onlyPattern
onlyPattern Pattern, NO selection 10.36 264 --
lean8 HeadOn,Linear,Circular,Accel,Pattern,GF,KNN,WallBounce 6.31 146 0.0169
lean6 lean8 - HeadOn 8.83 212 0.0262
pairPC Pattern + Circular 8.23 185 0.0460
pairPK Pattern + KNN 9.80 264 0.3998
pairPL Pattern + Linear 8.23 200 0.0035
The control replicates the prior run (10.36% vs 10.78% before; same binary tree,
different build path).
THE MECHANISM, from the per-arm selected-gun mix - the virtual signal keeps
ranking the WRONG guns first, even on a two-gun rack:
lean8: HeadOn 46.2% of ticks at 2.0% REAL; Pattern only 14.4% (12.6% real)
lean6: Pattern 29.2% (10.4% real) vs KNN 25.3% (8.1%) and Linear 15.2% (8.0%)
pairPC: Circular 66.8% (6.9% real) vs Pattern 33.2% (11.2% real) - over-picks Circular
pairPL: Linear 57.7% (6.0% real) vs Pattern 42.3% (11.1% real) - over-picks Linear
pairPK: Pattern 86.4% - ties ONLY because the selector happens to pick Pattern
most of the time; it is numerically lower with identical dmg/run
So the failure is NOT rack size. Pruning does not fix it; the ranking is wrong.
VERDICT: ship `onlyPattern` - Pattern alone with selection bypassed - at 10.36%
real and 264 dmg/run, vs lean8 6.31%/146 and the prior full rack 6.93%/159.
This DIRECTLY CONTRADICTS the standing user directive to keep virtual-fitness
selection, so it is recorded here plainly rather than quietly acted on: disable
the selector (`TR_RACK_<every gun but PATTERN>=off`) pending a better fitness
signal. The mechanism itself is left intact and functional so it can be re-enabled
with one env var, and so it can be fixed rather than discarded.
REMAINING CAVEAT: ONE ADVERSARY. All of this is vs DrussGT. Pattern as the default
must be re-checked against other bots first - that is the next job.
Extends the reusable harness (tools/ab/which_gun_arm_env.sh now has lean8/lean6/
pairPC/pairPK/pairPL; which_gun_analyze.py is parameterised by WHICHGUN_OUT and
compares against both `full` and `onlyPattern`).
MEASURED against the real DrussGT, one frozen binary built from clean HEAD, rack
knobs only (no source edits), 5 arms x 7 runs x 7 rounds, 8 concurrent battles,
judged ONLY on server-side real hit rate from the events sidecar, exact
two-sided permutation test on per-run rates.
arm runs shots hits real % dmg/run p vs full
full (shipped) 7 3898 270 6.93 159 --
onlyPattern 7 4582 494 10.78 287 0.0012 <- BETTER
onlyKNN 7 4033 207 5.13 119 0.1340
onlyLinear 7 3215 105 3.27 65 0.0082
onlyGF 7 3193 72 2.25 45 0.0012
Firing Pattern ALONE gives +3.85pp pooled hit rate and +80% damage per run, and
it fires MORE shots (4582 vs 3898) - it dominates on rate and volume. This is not
"any single gun wins" (full beats Linear, GF and KNN); it is specifically
"Pattern alone beats the rack".
WHY - the virtual fitness signal mis-ranks guns against real outcomes:
- HeadOn is massively over-selected: 31.4% of ticks, the most real shots (1070),
but only 4.5% REAL. It alone drags the rack down.
- Pattern has the best virtual rank and near-best real rate (11.9%, rank 2), yet
is selected only 22.6% of the time.
- Linear's apparent strength was SELECTION BIAS: conditional on being selected it
looked like 15.2% (n=33), but its UNCONDITIONAL rate (onlyLinear) is 3.27%.
Every earlier per-gun "real rate" in this repo is conditional on selection and
is therefore confounded. This experiment is the clean measurement.
NOT YET SETTLED (do not overclaim):
- ONE ADVERSARY. All of this is vs DrussGT. Pattern must be re-checked against
other bots before it becomes the default on this evidence alone.
- Whether a SMALL rack of good guns beats Pattern alone. The selector is negative
value on the CURRENT bloated rack; that does not prove it is negative value on
a rack of only good guns. That is the next experiment and it decides whether
the selection apparatus is fixed or disabled.
- The user's standing directive is to KEEP virtual-fitness selection. This
measurement conflicts with it, so the next step tests the selector on a small
good rack rather than assuming either answer.
Context - three prior selection-side attempts all failed: hysteresis (7.02% ->
5.10%, p=0.002), commitment (7.17% -> 4.44%, p=0.0012), arrival-accuracy
tie-break (7.08%, p=0.88 null). The per-tick random draw is load-bearing on
three independent measurements. This experiment locates the real problem one
level up: which guns are in the rack, and that the virtual signal ranks them
wrongly.
Preserves the reusable harness (tools/ab/which_gun_run_one.sh,
which_gun_arm_env.sh, which_gun_analyze.py) and the full writeup
(docs/selector_negative_value.md).