diff --git a/docs/gun_campaign.md b/docs/gun_campaign.md new file mode 100644 index 0000000..2355e7d --- /dev/null +++ b/docs/gun_campaign.md @@ -0,0 +1,194 @@ +# Gun campaign — ledger + +**Goal (owner's mandate, 2026-09-26 overnight):** the movement campaign produced +a replicated champion (`TR_MOVEMENT=strafe`, `docs/movement_campaign.md`) which a +parallel job is shipping. *"Continue until you found an amazing movement. When +found do the same over for a gun."* This file is the GUN campaign's single source +of truth: later jobs **append** a `## Batch N` section and never edit an earlier +one (a wrong earlier number gets a correction line, not a rewrite). This file is +the gun successor to `docs/movement_campaign.md`; read that first, then this. + +--- + +## 0. The question + +The shipped rack admits **`Pattern` only** (`onlyPattern`). That decision was +taken because the virtual-fitness selector measured **negative value** at every +rack size tested, and Pattern is the best single gun by the DrussGT hit-rate +table (`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit +`e0666a5`). But every one of those measurements — like every pre-campaign +movement claim — is **DrussGT-heavy**, and the movement campaign's standing +lesson is that a one-opponent result is not a result. + +So Batch 1 asks the direct question, on a frozen 15-opponent panel: + +> **Does any rack gun, or any small gun configuration, beat the shipped +> `Pattern` on damage/run AND round wins across 15 opponents — or is +> `onlyPattern` CONFIRMED rather than merely assumed?** + +A **null is a successful outcome here**: it upgrades `onlyPattern` from an +assumption to a measured verdict across the panel, which is itself valuable. + +--- + +## 1. Protocol (identical to the movement campaign; the instrument is reused verbatim) + +`tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` + +`tools/ab/panel_movement.txt` (frozen panel), reused **unchanged**. The unit of +evidence is the **number of opponents**, not the number of runs. + +| Element | Rule | +|---|---| +| Subject | ONE frozen binary built from `git archive HEAD` (`tournament_run.sh` records commit + binary sha256 in `session.json`) | +| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file (`tools/ab/arms_gun_b1.txt`) | +| Panel | the **frozen** `tools/ab/panel_movement.txt` (15 opponents). Adding/removing an opponent starts a new batch number | +| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate | +| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir | +| Serialization | **one battle fleet at a time** via `--wait-arena`; never a broad `pkill robocode_shim` | +| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named | +| Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched | + +### Movement is PINNED in every arm: `TR_MOVEMENT=strafe` + +**Reason (hard rule).** A movement-default flip is landing from job j120 during +this run, so the default engine can change underneath us. Every arm therefore +sets **`TR_MOVEMENT=strafe` explicitly**. This (a) removes movement as a +confound, (b) makes all arms share exactly one movement, so every delta is a pure +**gun** delta, and (c) satisfies the analyzer's liveness rule for every arm +(each declared token must appear verbatim; an *undeclared* leaked `TR_MOVEMENT` +is fatal, but here every arm declares it). The reference arm is therefore the +shipped **gun** rack under the pinned movement, not under the (mutable) default. + +--- + +## 2. Pre-registered decision rules (fixed BEFORE Batch 1 ran) + +1. **Primary metrics:** **damage/run** and **ROUND WINS**. Secondary/explanation + only: damage taken/run, incoming hit rate, mean distance. **Never hit rate + alone** — that trap has inverted six verdicts in this project. +2. **BETTER than the reference** iff one primary metric is up with a + cross-opponent **sign test p < 0.05** while the other does **not** go down; + the mirror image for **WORSE**. Anything else is **NOT DISTINGUISHABLE** (a + real answer, not a failure). Both the strict reading (the other metric's mean + delta `>= 0`) and the substantive reading (not *detectably* down: not + significant **and** smaller than that metric's MDE) are printed. +3. **A verdict must survive the between-opponent spread:** the pooled mean delta + is reported with the SD across opponents, its SE, a 95% CI, and the **MDE** + (α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as + *not detectable* — never as *absent*, never as a win. +4. **Somewhere to stop:** if no arm beats the shipped `pattern` by rule 2 in + Batch 1 **and** no arm shows a `>= +MDE` damage gain with p<0.10, then the + gun stage's first phase is closed with *"the shipped `onlyPattern` rack is the + best gun configuration we have measured across the panel"* — that is a + **successful** outcome, and the campaign moves to a named next axis rather + than inventing more arms. See *What would make us stop*. +5. **No promotion off a single metric, a single opponent, or a single run.** + A change that wins damage by losing wins (or vice-versa) is not a win. +6. Every batch is shot with a **pre-registered prediction** stated in its section + *before* the battles finish; a prediction that turns out wrong is recorded as + wrong. + +--- + +## 3. Stage 0 — what we already know (given, not re-derived) + +| fact | value | source | +|---|---|---| +| shipped rack | `onlyPattern` (id 5), all other guns `off` | `common_libs/gun_harness/selector.nim` | +| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | `docs/selector_negative_value.md`, `docs/gun_rack_analysis.md` | +| Pattern hit rate vs DrussGT | ~10.8% | given | +| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given | +| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | `docs/bitbrain_vs_tmhorizon_ab.md` | +| BitBrain across 32 legacy opponents | neutral: **-3.6 dmg/run**, +0.02 wins/run, p=0.53/0.91 | `docs/gauntlet_bitbrain_vs_pattern.md` | +| lead **amplitude** | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | `docs/bitbrain_campaign.md` | +| offline prediction checks | **veto-only** — may kill a design, never select one | `docs/offline_harness_trust.md` | +| spinner-specific claim | **UNTESTED**: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | `docs/gauntlet_bitbrain_vs_pattern.md` | + +**Movement context (why the pin):** the movement champion is `TR_MOVEMENT=strafe` +at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win +over the shipped `tfil` in four independent sessions (Batch 4: 52.6% vs 38.5% +round-win rate). Movement is held constant at that champion for every gun arm. + +--- + +## Batch 1 — does anything beat the shipped `Pattern`? + +**Design.** One frozen binary, six env-only arms, one frozen panel +(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 1 +wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive +megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file: +`tools/ab/arms_gun_b1.txt`. Reference: `pattern`. + +| # | arm | env | what it isolates | +|---|---|---|---| +| 1 | `pattern` | `TR_MOVEMENT=strafe` | the arm to beat (shipped `onlyPattern` rack) | +| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) | +| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only corrector (predicted to lose; never panel-tested) | +| 4 | `knn` | `+ TR_RACK_PATTERN=off TR_RACK_KNN=both` | best non-Pattern single gun (never panel-tested) | +| 5 | `rack_pk` | `+ TR_RACK_KNN=both` | Pattern + KNN, **selector ON** (different-family hedge) | +| 6 | `rack_pt` | `+ TR_RACK_TMHORIZON=both` | Pattern + TMHorizon, **selector ON** (corrector-family hedge) | + +*(`+` = the pinned `TR_MOVEMENT=strafe` is present in every arm; see §1.)* + +**Pre-registered prediction (written BEFORE the battles finished):** +1. **No arm beats `pattern` on BOTH primaries**, and the point estimates sit + inside the MDE. `onlyPattern` is confirmed across the panel. +2. `bitbrain` is a **wash** vs `pattern` (replicating the 32-opponent gauntlet: + ±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry. +3. `tmhorizon` is **WORSE** than `pattern` (its own offline work predicts it + needs ~80% side accuracy and can reach ~60%); no damage win, no win win. +4. `knn` is **WORSE** than `pattern` on damage (its DrussGT hit rate is roughly + half Pattern's) and not better on wins. +5. The two small racks **do not beat `pattern`**; `rack_pk` lands closer to + Pattern than `knn`-alone does (the selector at least partially hedges back), + but the selector's poor ranking keeps it below the reference. +6. If any surprise exists, it is a **rack arm on damage without wins** — the + `ring`-mover mirror-image trap — and it will not be read as a win. + + + +--- + +## What to try next + +*(Rewritten after Batch 1 — these are recommendations, not results. Filled in the +results commit below.)* + +## What would make us stop + +* **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond + the MDE, or when a gun arm's damage gain is bought with a detectable win loss. + At that point *"`onlyPattern` is the measured optimum of this rack"* is the + conclusion, not a failure (§2 rule 4). +* **Stop a single batch early** only for a contract violation (arena not free, + liveness FAIL, non-zero exit rate) — never because the numbers look boring. +* **Do not invent more arms on the same axis** once two consecutive batches fail + to improve on `pattern` beyond the MDE; move to a *named* new axis instead + (candidate list in "What to try next"). + +--- + +## How to run a batch (exact commands) + +```sh +# 0. wait for the arena (j120's movement run may still be fighting) +TOURNAMENT_NIMCACHE=/tmp/nc_j121 \ +tools/ab/tournament_run.sh \ + --arms tools/ab/arms_gun_b1.txt \ + --panel tools/ab/panel_movement.txt \ + --runs 3 --rounds 3 --conc 6 --wait-arena 45 \ + --reference pattern \ + --outdir /tmp/ab/j121_g1 + +# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict +python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern +``` + +`--reference` may be ANY arm of the session: re-analyzing an old session with a +different reference is a free pairwise comparison with no battles. + +## Session log (outdirs are in `/tmp` and are NOT committed) + +| session | commit | battles | arms | verdict | +|---|---|---:|---|---| +| `/tmp/ab/j121_g1` | *(recorded at run time)* | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | *(Batch 1 outcome below)* | diff --git a/tools/ab/arms_gun_b1.txt b/tools/ab/arms_gun_b1.txt new file mode 100644 index 0000000..56ff167 --- /dev/null +++ b/tools/ab/arms_gun_b1.txt @@ -0,0 +1,66 @@ +# ───────────────────────────────────────────────────────────────────────────── +# arms_gun_b1.txt — BATCH 1 of the GUN campaign: does ANY rack gun or gun +# configuration beat the shipped `Pattern` across the frozen panel? +# +# Format: name | ENV=value ENV=value | label +# +# ONE frozen binary (tournament_run.sh builds it from `git archive HEAD`); every +# arm below differs ONLY by its env dict. No per-arm rebuild. +# +# MOVEMENT IS PINNED IN EVERY ARM (`TR_MOVEMENT=strafe`). Reason (hard rule): +# a parallel job (j120) is shipping a movement-default flip during this run, so +# the default engine may change underneath us. Pinning the engine in EVERY arm +# removes movement as a confound, makes all arms share one movement, and keeps +# the comparison a pure GUN comparison. Because every arm declares TR_MOVEMENT, +# the analyzer's liveness rule (a token must appear verbatim in the bot's own +# raw-env report) is also satisfied for every arm. +# +# Prior data this batch is built on (taken as given; NOT re-derived): +# * shipped rack = `onlyPattern` (Pattern id 5, all others off); chosen +# because the selector measured NEGATIVE value at every rack size tested +# (docs/selector_negative_value.md, docs/gun_rack_analysis.md). +# * per-gun single-gun hit rate vs DrussGT: Pattern ~10.8%, KNN 5.6%, +# Linear 3.0%, Circular 2.9%, WallBounce 2.7%, GF 2.1%. +# * BitBrain == Pattern when idle (30 runs/arm 97/210 vs 97/210 wins) and is +# NEUTRAL across 32 legacy opponents (-3.6 dmg/run, docs/gauntlet_bitbrain_vs_pattern.md). +# Its learned-gain config was worse on DrussGT ONLY. +# * lead AMPLITUDE is a dead axis (full gain sweep run, nothing beats Pattern); +# lead INFORMATION is the open one. +# +# The question this batch answers: is `onlyPattern` CONFIRMED across 15 opponents +# (rather than merely assumed from a DrussGT-heavy history), or does some gun / +# small rack beat it on damage-per-run AND round wins? +# ───────────────────────────────────────────────────────────────────────────── + +# 1. REFERENCE. The shipped `onlyPattern` rack, movement pinned. No gun env at +# all, so this is what the shipped bot's gun does. +pattern | TR_MOVEMENT=strafe | shipped onlyPattern rack (REFERENCE) + +# 2. BitBrain-only rack: the learned-gain corrector on Pattern's lead. This is +# the exact config the owner most likely ran (gains 1.0..2.0, decay memory); +# the 32-opponent gauntlet had it as a wash, so this is a PANEL replication at +# the new (movement-pinned) standard. BitBrain computes Pattern's prediction +# internally, so Pattern does not need to be in the rack. +bitbrain | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay | BitBrain-only, learned gain (gains 1.0..2.0, decay) + +# 3. TMHorizon-only rack: the horizon Tsetlin corrector on Pattern's lead. Its +# own design docs PREDICT it loses (needs ~80% side accuracy to break even; +# achievable ~60%). Never measured across the panel. This arm closes that. +tmhorizon | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon-only rack + +# 4. KNN-only rack: the best non-Pattern single gun by the old DrussGT hit-rate +# table (5.6% vs Pattern 10.8%) and NEVER measured across the panel. Included +# to close "the second-best single gun was never panel-tested". +knn | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_KNN=both | KNN-only rack + +# 5. TWO-GUN RACK, SELECTOR ON: Pattern (default both) + KNN. The selector was +# measured NEGATIVE with 16 guns and a rolling hit-rate; this re-tests it at +# a SMALL rack under the new panel standard, because the movement campaign +# overturned two earlier single-opponent conclusions. If KNN ever wins a +# matchup, the selector can hedge into it. +rack_pk | TR_MOVEMENT=strafe TR_RACK_KNN=both | Pattern + KNN, selector ON + +# 6. TWO-GUN RACK, SELECTOR ON: Pattern (default both) + TMHorizon. Same +# selector question against a CORRECTOR family (TMHorizon = Pattern +/- a +# learned few-degree shift) instead of a different-family gun (KNN). +rack_pt | TR_MOVEMENT=strafe TR_RACK_TMHORIZON=both | Pattern + TMHorizon, selector ON