Files
SirRoboGarage/docs/gun_campaign.md
T

11 KiB
Raw Blame History

Gun campaign — ledger

Goal (owner's mandate, 2026-09-26 overnight): the movement campaign produced a replicated champion (TR_MOVEMENT=strafe, docs/movement_campaign.md) which a parallel job is shipping. "Continue until you found an amazing movement. When found do the same over for a gun." This file is the GUN campaign's single source of truth: later jobs append a ## Batch N section and never edit an earlier one (a wrong earlier number gets a correction line, not a rewrite). This file is the gun successor to docs/movement_campaign.md; read that first, then this.


0. The question

The shipped rack admits Pattern only (onlyPattern). That decision was taken because the virtual-fitness selector measured negative value at every rack size tested, and Pattern is the best single gun by the DrussGT hit-rate table (docs/selector_negative_value.md, docs/gun_rack_analysis.md, commit e0666a5). But every one of those measurements — like every pre-campaign movement claim — is DrussGT-heavy, and the movement campaign's standing lesson is that a one-opponent result is not a result.

So Batch 1 asks the direct question, on a frozen 15-opponent panel:

Does any rack gun, or any small gun configuration, beat the shipped Pattern on damage/run AND round wins across 15 opponents — or is onlyPattern CONFIRMED rather than merely assumed?

A null is a successful outcome here: it upgrades onlyPattern from an assumption to a measured verdict across the panel, which is itself valuable.


1. Protocol (identical to the movement campaign; the instrument is reused verbatim)

tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py + tools/ab/panel_movement.txt (frozen panel), reused unchanged. The unit of evidence is the number of opponents, not the number of runs.

Element Rule
Subject ONE frozen binary built from git archive HEAD (tournament_run.sh records commit + binary sha256 in session.json)
Arms env dicts only — no per-arm rebuild, ever; the arm file is a committed file (tools/ab/arms_gun_b1.txt)
Panel the frozen tools/ab/panel_movement.txt (15 opponents). Adding/removing an opponent starts a new batch number
Pairing per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate
Isolation per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir
Serialization one battle fleet at a time via --wait-arena; never a broad pkill robocode_shim
Liveness every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named
Never shipped this is a measurement campaign: git status clean, shipped defaults untouched, .gitignore untouched

Movement is PINNED in every arm: TR_MOVEMENT=strafe

Reason (hard rule). A movement-default flip is landing from job j120 during this run, so the default engine can change underneath us. Every arm therefore sets TR_MOVEMENT=strafe explicitly. This (a) removes movement as a confound, (b) makes all arms share exactly one movement, so every delta is a pure gun delta, and (c) satisfies the analyzer's liveness rule for every arm (each declared token must appear verbatim; an undeclared leaked TR_MOVEMENT is fatal, but here every arm declares it). The reference arm is therefore the shipped gun rack under the pinned movement, not under the (mutable) default.


2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)

  1. Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate, mean distance. Never hit rate alone — that trap has inverted six verdicts in this project.
  2. BETTER than the reference iff one primary metric is up with a cross-opponent sign test p < 0.05 while the other does not go down; the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (a real answer, not a failure). Both the strict reading (the other metric's mean delta >= 0) and the substantive reading (not detectably down: not significant and smaller than that metric's MDE) are printed.
  3. A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as not detectable — never as absent, never as a win.
  4. Somewhere to stop: if no arm beats the shipped pattern by rule 2 in Batch 1 and no arm shows a >= +MDE damage gain with p<0.10, then the gun stage's first phase is closed with "the shipped onlyPattern rack is the best gun configuration we have measured across the panel" — that is a successful outcome, and the campaign moves to a named next axis rather than inventing more arms. See What would make us stop.
  5. No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
  6. Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.

3. Stage 0 — what we already know (given, not re-derived)

fact value source
shipped rack onlyPattern (id 5), all other guns off common_libs/gun_harness/selector.nim
why selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries docs/selector_negative_value.md, docs/gun_rack_analysis.md
Pattern hit rate vs DrussGT ~10.8% given
KNN / Linear / Circular / WallBounce / GF vs DrussGT 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % given
BitBrain when idle == Pattern (30 runs/arm: 97/210 vs 97/210 wins) docs/bitbrain_vs_tmhorizon_ab.md
BitBrain across 32 legacy opponents neutral: -3.6 dmg/run, +0.02 wins/run, p=0.53/0.91 docs/gauntlet_bitbrain_vs_pattern.md
lead amplitude dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern docs/bitbrain_campaign.md
offline prediction checks veto-only — may kill a design, never select one docs/offline_harness_trust.md
spinner-specific claim UNTESTED: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) docs/gauntlet_bitbrain_vs_pattern.md

Movement context (why the pin): the movement champion is TR_MOVEMENT=strafe at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win over the shipped tfil in four independent sessions (Batch 4: 52.6% vs 38.5% round-win rate). Movement is held constant at that champion for every gun arm.


Batch 1 — does anything beat the shipped Pattern?

Design. One frozen binary, six env-only arms, one frozen panel (tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 1 wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive megas), 3 runs × 3 rounds per (opponent, arm) = 270 battles. Arm file: tools/ab/arms_gun_b1.txt. Reference: pattern.

# arm env what it isolates
1 pattern TR_MOVEMENT=strafe the arm to beat (shipped onlyPattern rack)
2 bitbrain + TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard)
3 tmhorizon + TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TMHorizon-only corrector (predicted to lose; never panel-tested)
4 knn + TR_RACK_PATTERN=off TR_RACK_KNN=both best non-Pattern single gun (never panel-tested)
5 rack_pk + TR_RACK_KNN=both Pattern + KNN, selector ON (different-family hedge)
6 rack_pt + TR_RACK_TMHORIZON=both Pattern + TMHorizon, selector ON (corrector-family hedge)

(+ = the pinned TR_MOVEMENT=strafe is present in every arm; see §1.)

Pre-registered prediction (written BEFORE the battles finished):

  1. No arm beats pattern on BOTH primaries, and the point estimates sit inside the MDE. onlyPattern is confirmed across the panel.
  2. bitbrain is a wash vs pattern (replicating the 32-opponent gauntlet: ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.
  3. tmhorizon is WORSE than pattern (its own offline work predicts it needs ~80% side accuracy and can reach ~60%); no damage win, no win win.
  4. knn is WORSE than pattern on damage (its DrussGT hit rate is roughly half Pattern's) and not better on wins.
  5. The two small racks do not beat pattern; rack_pk lands closer to Pattern than knn-alone does (the selector at least partially hedges back), but the selector's poor ranking keeps it below the reference.
  6. If any surprise exists, it is a rack arm on damage without wins — the ring-mover mirror-image trap — and it will not be read as a win.

What to try next

(Rewritten after Batch 1 — these are recommendations, not results. Filled in the results commit below.)

What would make us stop

  • Stop the first phase once a batch's best arm cannot beat pattern beyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point "onlyPattern is the measured optimum of this rack" is the conclusion, not a failure (§2 rule 4).
  • Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.
  • Do not invent more arms on the same axis once two consecutive batches fail to improve on pattern beyond the MDE; move to a named new axis instead (candidate list in "What to try next").

How to run a batch (exact commands)

# 0. wait for the arena (j120's movement run may still be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
    --arms    tools/ab/arms_gun_b1.txt \
    --panel   tools/ab/panel_movement.txt \
    --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
    --reference pattern \
    --outdir /tmp/ab/j121_g1

# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern

--reference may be ANY arm of the session: re-analyzing an old session with a different reference is a free pairwise comparison with no battles.

Session log (outdirs are in /tmp and are NOT committed)

session commit battles arms verdict
/tmp/ab/j121_g1 (recorded at run time) 270 pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt (Batch 1 outcome below)