Files
SirRoboGarage/docs/gun_campaign.md
T

195 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gun campaign — ledger
**Goal (owner's mandate, 2026-09-26 overnight):** the movement campaign produced
a replicated champion (`TR_MOVEMENT=strafe`, `docs/movement_campaign.md`) which a
parallel job is shipping. *"Continue until you found an amazing movement. When
found do the same over for a gun."* This file is the GUN campaign's single source
of truth: later jobs **append** a `## Batch N` section and never edit an earlier
one (a wrong earlier number gets a correction line, not a rewrite). This file is
the gun successor to `docs/movement_campaign.md`; read that first, then this.
---
## 0. The question
The shipped rack admits **`Pattern` only** (`onlyPattern`). That decision was
taken because the virtual-fitness selector measured **negative value** at every
rack size tested, and Pattern is the best single gun by the DrussGT hit-rate
table (`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit
`e0666a5`). But every one of those measurements — like every pre-campaign
movement claim — is **DrussGT-heavy**, and the movement campaign's standing
lesson is that a one-opponent result is not a result.
So Batch 1 asks the direct question, on a frozen 15-opponent panel:
> **Does any rack gun, or any small gun configuration, beat the shipped
> `Pattern` on damage/run AND round wins across 15 opponents — or is
> `onlyPattern` CONFIRMED rather than merely assumed?**
A **null is a successful outcome here**: it upgrades `onlyPattern` from an
assumption to a measured verdict across the panel, which is itself valuable.
---
## 1. Protocol (identical to the movement campaign; the instrument is reused verbatim)
`tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` +
`tools/ab/panel_movement.txt` (frozen panel), reused **unchanged**. The unit of
evidence is the **number of opponents**, not the number of runs.
| Element | Rule |
|---|---|
| Subject | ONE frozen binary built from `git archive HEAD` (`tournament_run.sh` records commit + binary sha256 in `session.json`) |
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file (`tools/ab/arms_gun_b1.txt`) |
| Panel | the **frozen** `tools/ab/panel_movement.txt` (15 opponents). Adding/removing an opponent starts a new batch number |
| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | **one battle fleet at a time** via `--wait-arena`; never a broad `pkill robocode_shim` |
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
| Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched |
### Movement is PINNED in every arm: `TR_MOVEMENT=strafe`
**Reason (hard rule).** A movement-default flip is landing from job j120 during
this run, so the default engine can change underneath us. Every arm therefore
sets **`TR_MOVEMENT=strafe` explicitly**. This (a) removes movement as a
confound, (b) makes all arms share exactly one movement, so every delta is a pure
**gun** delta, and (c) satisfies the analyzer's liveness rule for every arm
(each declared token must appear verbatim; an *undeclared* leaked `TR_MOVEMENT`
is fatal, but here every arm declares it). The reference arm is therefore the
shipped **gun** rack under the pinned movement, not under the (mutable) default.
---
## 2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)
1. **Primary metrics:** **damage/run** and **ROUND WINS**. Secondary/explanation
only: damage taken/run, incoming hit rate, mean distance. **Never hit rate
alone** — that trap has inverted six verdicts in this project.
2. **BETTER than the reference** iff one primary metric is up with a
cross-opponent **sign test p < 0.05** while the other does **not** go down;
the mirror image for **WORSE**. Anything else is **NOT DISTINGUISHABLE** (a
real answer, not a failure). Both the strict reading (the other metric's mean
delta `>= 0`) and the substantive reading (not *detectably* down: not
significant **and** smaller than that metric's MDE) are printed.
3. **A verdict must survive the between-opponent spread:** the pooled mean delta
is reported with the SD across opponents, its SE, a 95% CI, and the **MDE**
(α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as
*not detectable* — never as *absent*, never as a win.
4. **Somewhere to stop:** if no arm beats the shipped `pattern` by rule 2 in
Batch 1 **and** no arm shows a `>= +MDE` damage gain with p<0.10, then the
gun stage's first phase is closed with *"the shipped `onlyPattern` rack is the
best gun configuration we have measured across the panel"* — that is a
**successful** outcome, and the campaign moves to a named next axis rather
than inventing more arms. See *What would make us stop*.
5. **No promotion off a single metric, a single opponent, or a single run.**
A change that wins damage by losing wins (or vice-versa) is not a win.
6. Every batch is shot with a **pre-registered prediction** stated in its section
*before* the battles finish; a prediction that turns out wrong is recorded as
wrong.
---
## 3. Stage 0 — what we already know (given, not re-derived)
| fact | value | source |
|---|---|---|
| shipped rack | `onlyPattern` (id 5), all other guns `off` | `common_libs/gun_harness/selector.nim` |
| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | `docs/selector_negative_value.md`, `docs/gun_rack_analysis.md` |
| Pattern hit rate vs DrussGT | ~10.8% | given |
| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given |
| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | `docs/bitbrain_vs_tmhorizon_ab.md` |
| BitBrain across 32 legacy opponents | neutral: **-3.6 dmg/run**, +0.02 wins/run, p=0.53/0.91 | `docs/gauntlet_bitbrain_vs_pattern.md` |
| lead **amplitude** | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | `docs/bitbrain_campaign.md` |
| offline prediction checks | **veto-only** — may kill a design, never select one | `docs/offline_harness_trust.md` |
| spinner-specific claim | **UNTESTED**: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | `docs/gauntlet_bitbrain_vs_pattern.md` |
**Movement context (why the pin):** the movement champion is `TR_MOVEMENT=strafe`
at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win
over the shipped `tfil` in four independent sessions (Batch 4: 52.6% vs 38.5%
round-win rate). Movement is held constant at that champion for every gun arm.
---
## Batch 1 — does anything beat the shipped `Pattern`?
**Design.** One frozen binary, six env-only arms, one frozen panel
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 1
wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive
megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file:
`tools/ab/arms_gun_b1.txt`. Reference: `pattern`.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | `pattern` | `TR_MOVEMENT=strafe` | the arm to beat (shipped `onlyPattern` rack) |
| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) |
| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only corrector (predicted to lose; never panel-tested) |
| 4 | `knn` | `+ TR_RACK_PATTERN=off TR_RACK_KNN=both` | best non-Pattern single gun (never panel-tested) |
| 5 | `rack_pk` | `+ TR_RACK_KNN=both` | Pattern + KNN, **selector ON** (different-family hedge) |
| 6 | `rack_pt` | `+ TR_RACK_TMHORIZON=both` | Pattern + TMHorizon, **selector ON** (corrector-family hedge) |
*(`+` = the pinned `TR_MOVEMENT=strafe` is present in every arm; see §1.)*
**Pre-registered prediction (written BEFORE the battles finished):**
1. **No arm beats `pattern` on BOTH primaries**, and the point estimates sit
inside the MDE. `onlyPattern` is confirmed across the panel.
2. `bitbrain` is a **wash** vs `pattern` (replicating the 32-opponent gauntlet:
±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.
3. `tmhorizon` is **WORSE** than `pattern` (its own offline work predicts it
needs ~80% side accuracy and can reach ~60%); no damage win, no win win.
4. `knn` is **WORSE** than `pattern` on damage (its DrussGT hit rate is roughly
half Pattern's) and not better on wins.
5. The two small racks **do not beat `pattern`**; `rack_pk` lands closer to
Pattern than `knn`-alone does (the selector at least partially hedges back),
but the selector's poor ranking keeps it below the reference.
6. If any surprise exists, it is a **rack arm on damage without wins** — the
`ring`-mover mirror-image trap — and it will not be read as a win.
<!-- OUTCOME-ANCHOR -->
---
## What to try next
*(Rewritten after Batch 1 — these are recommendations, not results. Filled in the
results commit below.)*
## What would make us stop
* **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond
the MDE, or when a gun arm's damage gain is bought with a detectable win loss.
At that point *"`onlyPattern` is the measured optimum of this rack"* is the
conclusion, not a failure (§2 rule 4).
* **Stop a single batch early** only for a contract violation (arena not free,
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
* **Do not invent more arms on the same axis** once two consecutive batches fail
to improve on `pattern` beyond the MDE; move to a *named* new axis instead
(candidate list in "What to try next").
---
## How to run a batch (exact commands)
```sh
# 0. wait for the arena (j120's movement run may still be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference pattern \
--outdir /tmp/ab/j121_g1
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
```
`--reference` may be ANY arm of the session: re-analyzing an old session with a
different reference is a free pairwise comparison with no battles.
## Session log (outdirs are in `/tmp` and are NOT committed)
| session | commit | battles | arms | verdict |
|---|---|---:|---|---|
| `/tmp/ab/j121_g1` | *(recorded at run time)* | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | *(Batch 1 outcome below)* |