# Gun campaign — ledger **Goal (owner's mandate, 2026-09-26 overnight):** the movement campaign produced a replicated champion (`TR_MOVEMENT=strafe`, `docs/movement_campaign.md`) which a parallel job is shipping. *"Continue until you found an amazing movement. When found do the same over for a gun."* This file is the GUN campaign's single source of truth: later jobs **append** a `## Batch N` section and never edit an earlier one (a wrong earlier number gets a correction line, not a rewrite). This file is the gun successor to `docs/movement_campaign.md`; read that first, then this. --- ## 0. The question The shipped rack admits **`Pattern` only** (`onlyPattern`). That decision was taken because the virtual-fitness selector measured **negative value** at every rack size tested, and Pattern is the best single gun by the DrussGT hit-rate table (`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit `e0666a5`). But every one of those measurements — like every pre-campaign movement claim — is **DrussGT-heavy**, and the movement campaign's standing lesson is that a one-opponent result is not a result. So Batch 1 asks the direct question, on a frozen 15-opponent panel: > **Does any rack gun, or any small gun configuration, beat the shipped > `Pattern` on damage/run AND round wins across 15 opponents — or is > `onlyPattern` CONFIRMED rather than merely assumed?** A **null is a successful outcome here**: it upgrades `onlyPattern` from an assumption to a measured verdict across the panel, which is itself valuable. --- ## 1. Protocol (identical to the movement campaign; the instrument is reused verbatim) `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` + `tools/ab/panel_movement.txt` (frozen panel), reused **unchanged**. The unit of evidence is the **number of opponents**, not the number of runs. | Element | Rule | |---|---| | Subject | ONE frozen binary built from `git archive HEAD` (`tournament_run.sh` records commit + binary sha256 in `session.json`) | | Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file (`tools/ab/arms_gun_b1.txt`) | | Panel | the **frozen** `tools/ab/panel_movement.txt` (15 opponents). Adding/removing an opponent starts a new batch number | | Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate | | Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir | | Serialization | **one battle fleet at a time** via `--wait-arena`; never a broad `pkill robocode_shim` | | Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named | | Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched | ### Movement is PINNED in every arm: `TR_MOVEMENT=strafe` **Reason (hard rule).** A movement-default flip is landing from job j120 during this run, so the default engine can change underneath us. Every arm therefore sets **`TR_MOVEMENT=strafe` explicitly**. This (a) removes movement as a confound, (b) makes all arms share exactly one movement, so every delta is a pure **gun** delta, and (c) satisfies the analyzer's liveness rule for every arm (each declared token must appear verbatim; an *undeclared* leaked `TR_MOVEMENT` is fatal, but here every arm declares it). The reference arm is therefore the shipped **gun** rack under the pinned movement, not under the (mutable) default. --- ## 2. Pre-registered decision rules (fixed BEFORE Batch 1 ran) 1. **Primary metrics:** **damage/run** and **ROUND WINS**. Secondary/explanation only: damage taken/run, incoming hit rate, mean distance. **Never hit rate alone** — that trap has inverted six verdicts in this project. 2. **BETTER than the reference** iff one primary metric is up with a cross-opponent **sign test p < 0.05** while the other does **not** go down; the mirror image for **WORSE**. Anything else is **NOT DISTINGUISHABLE** (a real answer, not a failure). Both the strict reading (the other metric's mean delta `>= 0`) and the substantive reading (not *detectably* down: not significant **and** smaller than that metric's MDE) are printed. 3. **A verdict must survive the between-opponent spread:** the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the **MDE** (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as *not detectable* — never as *absent*, never as a win. 4. **Somewhere to stop:** if no arm beats the shipped `pattern` by rule 2 in Batch 1 **and** no arm shows a `>= +MDE` damage gain with p<0.10, then the gun stage's first phase is closed with *"the shipped `onlyPattern` rack is the best gun configuration we have measured across the panel"* — that is a **successful** outcome, and the campaign moves to a named next axis rather than inventing more arms. See *What would make us stop*. 5. **No promotion off a single metric, a single opponent, or a single run.** A change that wins damage by losing wins (or vice-versa) is not a win. 6. Every batch is shot with a **pre-registered prediction** stated in its section *before* the battles finish; a prediction that turns out wrong is recorded as wrong. --- ## 3. Stage 0 — what we already know (given, not re-derived) | fact | value | source | |---|---|---| | shipped rack | `onlyPattern` (id 5), all other guns `off` | `common_libs/gun_harness/selector.nim` | | why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | `docs/selector_negative_value.md`, `docs/gun_rack_analysis.md` | | Pattern hit rate vs DrussGT | ~10.8% | given | | KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given | | BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | `docs/bitbrain_vs_tmhorizon_ab.md` | | BitBrain across 32 legacy opponents | neutral: **-3.6 dmg/run**, +0.02 wins/run, p=0.53/0.91 | `docs/gauntlet_bitbrain_vs_pattern.md` | | lead **amplitude** | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | `docs/bitbrain_campaign.md` | | offline prediction checks | **veto-only** — may kill a design, never select one | `docs/offline_harness_trust.md` | | spinner-specific claim | **UNTESTED**: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | `docs/gauntlet_bitbrain_vs_pattern.md` | **Movement context (why the pin):** the movement champion is `TR_MOVEMENT=strafe` at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win over the shipped `tfil` in four independent sessions (Batch 4: 52.6% vs 38.5% round-win rate). Movement is held constant at that champion for every gun arm. --- ## Batch 1 — does anything beat the shipped `Pattern`? **Design.** One frozen binary, six env-only arms, one frozen panel (`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 1 wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file: `tools/ab/arms_gun_b1.txt`. Reference: `pattern`. | # | arm | env | what it isolates | |---|---|---|---| | 1 | `pattern` | `TR_MOVEMENT=strafe` | the arm to beat (shipped `onlyPattern` rack) | | 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) | | 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only corrector (predicted to lose; never panel-tested) | | 4 | `knn` | `+ TR_RACK_PATTERN=off TR_RACK_KNN=both` | best non-Pattern single gun (never panel-tested) | | 5 | `rack_pk` | `+ TR_RACK_KNN=both` | Pattern + KNN, **selector ON** (different-family hedge) | | 6 | `rack_pt` | `+ TR_RACK_TMHORIZON=both` | Pattern + TMHorizon, **selector ON** (corrector-family hedge) | *(`+` = the pinned `TR_MOVEMENT=strafe` is present in every arm; see §1.)* **Pre-registered prediction (written BEFORE the battles finished):** 1. **No arm beats `pattern` on BOTH primaries**, and the point estimates sit inside the MDE. `onlyPattern` is confirmed across the panel. 2. `bitbrain` is a **wash** vs `pattern` (replicating the 32-opponent gauntlet: ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry. 3. `tmhorizon` is **WORSE** than `pattern` (its own offline work predicts it needs ~80% side accuracy and can reach ~60%); no damage win, no win win. 4. `knn` is **WORSE** than `pattern` on damage (its DrussGT hit rate is roughly half Pattern's) and not better on wins. 5. The two small racks **do not beat `pattern`**; `rack_pk` lands closer to Pattern than `knn`-alone does (the selector at least partially hedges back), but the selector's poor ranking keeps it below the reference. 6. If any surprise exists, it is a **rack arm on damage without wins** — the `ring`-mover mirror-image trap — and it will not be read as a win. ### Outcome — direct answer **MEASURED.** Session `/tmp/ab/j121_g1`, commit `1d8143a15e4a038c39dfc2153ff518fa003b8612` (the pre-registration commit), frozen binary sha256 `aa49a45fec20…`, 15 opponents × 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started, 0 liveness exclusions**. Every arm declared `TR_MOVEMENT=strafe`, verified in each bot's own raw-env report; the reference rack line reads `PATTERN` for `pattern` and the overridden rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the others. > **DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped > `Pattern` across the panel on damage/run AND round wins.** Every arm except > `knn` is **NOT DISTINGUISHABLE** from `pattern` under both the strict and the > substantive reading of the pre-registered rule. The best nominal challenger, > `tmhorizon`, is **+0.13 wins/run** (MDE 0.36) and **+8.7 dmg/run** (MDE 13.8) > — an effect ~3× smaller than the design can detect — and `bitbrain` is > `+0.11` wins / `+10.3` dmg, also inside the MDE. `knn` is **detectably WORSE** > on both primaries. The shipped `onlyPattern` rack is therefore **CONFIRMED** > rather than merely assumed, at the resolution of this panel. That is a real > result, not a failure. #### Pooled dashboard (all valid runs — explanation only, NOT the verdict) | arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance | |---|---:|---:|---:|---:|---:|---:|---:|---:| | `pattern` (REF) | 45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 | | `bitbrain` | 45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 | | `tmhorizon` | 45 | 108.3 | 153.7 | **1.64** | **74/135** | **54.8%** | **12.53%** | 428 | | `knn` | 45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 | | `rack_pk` | 45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 | | `rack_pt` | 45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 | #### Cross-opponent aggregation (the verdict layer, damage + wins) | arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE | |---|---|---:|---:|---|---:|---:|---:|---:|---:| | `bitbrain` | damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 | | `bitbrain` | wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 | | `tmhorizon` | damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 | | `tmhorizon` | wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 | | `knn` | damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 | | `knn` | wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 | | `rack_pk` | damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 | | `rack_pk` | wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 | | `rack_pt` | damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 | | `rack_pt` | wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 | #### The pre-registered verdict (verbatim) | rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) | |---:|---|---:|---:|---|---|---|---| | 1 | `tmhorizon` | +0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** | | 2 | `rack_pt` | +0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | **not distinguishable** | **not distinguishable** | | 3 | `bitbrain` | +0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** | | 4 | `rack_pk` | −0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | **not distinguishable** | **not distinguishable** | | 5 | `knn` | −0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | **WORSE** | **not distinguishable** | Reference `pattern`: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px. #### Reading (MEASURED, with the mechanism) * **`knn` is a clean, large loss and is dropped.** It deals **−39 dmg/run** and loses **−0.62 wins/run**, positive on **0/15** opponents on damage and 1/11 on wins (p=6e-5 / p=0.012). Its incoming hit rate is *no worse* (13.31% vs 12.84%) — it simply cannot aim: the same movement, a worse gun. * **The two small racks do not beat `pattern`.** `rack_pk` (Pattern+KNN, selector ON) is *not* the average of Pattern and KNN: it recovers most of KNN's damage loss (+7.7 vs KNN-alone's −39.3) but still lands **nominally below** Pattern on wins (−0.09). `rack_pt` (Pattern+TMHorizon) is statistically indistinguishable from `tmhorizon`-alone — adding Pattern to the rack changes nothing, i.e. the selector picks the corrector essentially always. **The selector's 16-gun negative verdict reproduces at a 2-gun rack**: it does not manufacture a win it did not have. * **The only signal is a damage-side hint on the two Pattern-lead correctors** (`bitbrain` +10.3, `tmhorizon` +8.7, both 11/15 opponents). It is *below the MDE* (12.0 / 13.8) and *not significant by the pre-registered sign test* (p=0.1185), so it is **not a win and not a null**. Note the sign-flip test on the *mean* is nominally significant for `bitbrain` (p=0.032) while the sign test is not — the positive deltas are larger than the negatives — but with 5 arms compared this is not compelling after multiplicity, and the effect is under the MDE. **This is exactly the "damage without wins" pattern this project keeps paying for** (the `ring` mover: +31 dmg/run and *fewer* wins); here the win deltas (+0.11/+0.13) are themselves sub-MDE. * **The reference is stable:** `pattern` under `TR_MOVEMENT=strafe` measures 99.6 dmg/run and 50.4% round wins here, matching the movement campaign's `strafe` reference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same panel, different job. #### Pre-registered predictions — scorecard (an honest count) | # | prediction | outcome | |---|---|---| | 1 | no arm beats `pattern` on both primaries, point estimates inside the MDE; `onlyPattern` confirmed | **CORRECT** | | 2 | `bitbrain` is a wash vs `pattern` (gauntlet: −3.6 dmg, ≈0 wins) | **CORRECT on the verdict, WRONG on the point estimate**: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds | | 3 | `tmhorizon` is WORSE than `pattern` (its own docs predict it loses) | **WRONG**: it is nominally *better* on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable | | 4 | `knn` is WORSE than `pattern` on damage, not better on wins | **CORRECT** (detectably on both) | | 5 | the small racks do not beat `pattern`; `rack_pk` lands closer to Pattern than `knn`-alone | **CORRECT** | | 6 | any surprise is a rack arm on damage without wins | **PARTLY CORRECT**: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat | --- ## Batch 2 — wider-panel calibration of the damage-side hint The Batch-1 answer is decisive on the primary question (nothing beats `pattern`) but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids calling either way. Resolving it needs **more opponents, not more runs**, so this batch re-runs exactly those two arms plus the reference on a wider panel. **Design.** One frozen binary, three env-only arms, panel `tools/ab/panel_gun_b2.txt` = the j117 32-opponent legacy roster **minus Aurora** (inert in j117: one fire in 6 battles — it cannot reveal a gun regression) **plus SpinBot** = **33 opponents** (a superset of Batch 1's 15). **3 runs × 3 rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file: `tools/ab/arms_gun_b2.txt`. Reference: `pattern`. Movement pinned `TR_MOVEMENT=strafe` in every arm (same rationale as Batch 1). | # | arm | env | role | |---|---|---|---| | 1 | `pattern` | `TR_MOVEMENT=strafe` | reference | | 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | largest Batch-1 damage edge (+10.3) | | 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | second Batch-1 damage edge (+8.7) | **Pre-registered prediction (written BEFORE the battles finished):** 1. The `+10` dmg/run hint **does not survive** the wider panel: `bitbrain` and `tmhorizon` damage deltas fall toward zero and stay inside the new (smaller) MDE; the per-opponent sign test remains non-significant. Rationale: both arms come from the same Pattern-lead-correction base, and the prior 32-opponent gauntlet measured `bitbrain` at **−3.6 dmg/run** on a largely overlapping roster; +10 on 15 opponents is plausibly a small-panel fluctuation. 2. Round wins remain **flat** for both arms — if anything they regress toward zero. 3. If instead the damage hint **survives** at a significant cross-opponent sign test with wins not down, it is recorded as the first gun to **beat** `pattern` by rule 2 — and the campaign immediately looks for what the corrector is exploiting. That would be a genuine up-set, and it would **not** be spun as a null. ### Outcome — Batch 2 *(filled in by the Batch-2 results commit)* --- ## What to try next *(Rewritten after Batch 1 — these are recommendations, not results. Filled in the results commit below.)* ## What would make us stop * **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point *"`onlyPattern` is the measured optimum of this rack"* is the conclusion, not a failure (§2 rule 4). * **Stop a single batch early** only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring. * **Do not invent more arms on the same axis** once two consecutive batches fail to improve on `pattern` beyond the MDE; move to a *named* new axis instead (candidate list in "What to try next"). --- ## How to run a batch (exact commands) ```sh # 0. wait for the arena (j120's movement run may still be fighting) TOURNAMENT_NIMCACHE=/tmp/nc_j121 \ tools/ab/tournament_run.sh \ --arms tools/ab/arms_gun_b1.txt \ --panel tools/ab/panel_movement.txt \ --runs 3 --rounds 3 --conc 6 --wait-arena 45 \ --reference pattern \ --outdir /tmp/ab/j121_g1 # 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern ``` `--reference` may be ANY arm of the session: re-analyzing an old session with a different reference is a free pairwise comparison with no battles. ## Session log (outdirs are in `/tmp` and are NOT committed) | session | commit | battles | arms | verdict | |---|---|---:|---|---| | `/tmp/ab/j121_g1` | *(recorded at run time)* | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | *(Batch 1 outcome below)* |