Files
SirRoboGarage/docs/gun_campaign.md
T

334 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gun campaign — ledger
**Goal (owner's mandate, 2026-09-26 overnight):** the movement campaign produced
a replicated champion (`TR_MOVEMENT=strafe`, `docs/movement_campaign.md`) which a
parallel job is shipping. *"Continue until you found an amazing movement. When
found do the same over for a gun."* This file is the GUN campaign's single source
of truth: later jobs **append** a `## Batch N` section and never edit an earlier
one (a wrong earlier number gets a correction line, not a rewrite). This file is
the gun successor to `docs/movement_campaign.md`; read that first, then this.
---
## 0. The question
The shipped rack admits **`Pattern` only** (`onlyPattern`). That decision was
taken because the virtual-fitness selector measured **negative value** at every
rack size tested, and Pattern is the best single gun by the DrussGT hit-rate
table (`docs/selector_negative_value.md`, `docs/gun_rack_analysis.md`, commit
`e0666a5`). But every one of those measurements — like every pre-campaign
movement claim — is **DrussGT-heavy**, and the movement campaign's standing
lesson is that a one-opponent result is not a result.
So Batch 1 asks the direct question, on a frozen 15-opponent panel:
> **Does any rack gun, or any small gun configuration, beat the shipped
> `Pattern` on damage/run AND round wins across 15 opponents — or is
> `onlyPattern` CONFIRMED rather than merely assumed?**
A **null is a successful outcome here**: it upgrades `onlyPattern` from an
assumption to a measured verdict across the panel, which is itself valuable.
---
## 1. Protocol (identical to the movement campaign; the instrument is reused verbatim)
`tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` +
`tools/ab/panel_movement.txt` (frozen panel), reused **unchanged**. The unit of
evidence is the **number of opponents**, not the number of runs.
| Element | Rule |
|---|---|
| Subject | ONE frozen binary built from `git archive HEAD` (`tournament_run.sh` records commit + binary sha256 in `session.json`) |
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file (`tools/ab/arms_gun_b1.txt`) |
| Panel | the **frozen** `tools/ab/panel_movement.txt` (15 opponents). Adding/removing an opponent starts a new batch number |
| Pairing | per opponent: average the arm's runs, subtract the reference's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | **one battle fleet at a time** via `--wait-arena`; never a broad `pkill robocode_shim` |
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
| Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched |
### Movement is PINNED in every arm: `TR_MOVEMENT=strafe`
**Reason (hard rule).** A movement-default flip is landing from job j120 during
this run, so the default engine can change underneath us. Every arm therefore
sets **`TR_MOVEMENT=strafe` explicitly**. This (a) removes movement as a
confound, (b) makes all arms share exactly one movement, so every delta is a pure
**gun** delta, and (c) satisfies the analyzer's liveness rule for every arm
(each declared token must appear verbatim; an *undeclared* leaked `TR_MOVEMENT`
is fatal, but here every arm declares it). The reference arm is therefore the
shipped **gun** rack under the pinned movement, not under the (mutable) default.
---
## 2. Pre-registered decision rules (fixed BEFORE Batch 1 ran)
1. **Primary metrics:** **damage/run** and **ROUND WINS**. Secondary/explanation
only: damage taken/run, incoming hit rate, mean distance. **Never hit rate
alone** — that trap has inverted six verdicts in this project.
2. **BETTER than the reference** iff one primary metric is up with a
cross-opponent **sign test p < 0.05** while the other does **not** go down;
the mirror image for **WORSE**. Anything else is **NOT DISTINGUISHABLE** (a
real answer, not a failure). Both the strict reading (the other metric's mean
delta `>= 0`) and the substantive reading (not *detectably* down: not
significant **and** smaller than that metric's MDE) are printed.
3. **A verdict must survive the between-opponent spread:** the pooled mean delta
is reported with the SD across opponents, its SE, a 95% CI, and the **MDE**
(α=0.05 two-sided, 80% power). An effect smaller than the MDE is reported as
*not detectable* — never as *absent*, never as a win.
4. **Somewhere to stop:** if no arm beats the shipped `pattern` by rule 2 in
Batch 1 **and** no arm shows a `>= +MDE` damage gain with p<0.10, then the
gun stage's first phase is closed with *"the shipped `onlyPattern` rack is the
best gun configuration we have measured across the panel"* — that is a
**successful** outcome, and the campaign moves to a named next axis rather
than inventing more arms. See *What would make us stop*.
5. **No promotion off a single metric, a single opponent, or a single run.**
A change that wins damage by losing wins (or vice-versa) is not a win.
6. Every batch is shot with a **pre-registered prediction** stated in its section
*before* the battles finish; a prediction that turns out wrong is recorded as
wrong.
---
## 3. Stage 0 — what we already know (given, not re-derived)
| fact | value | source |
|---|---|---|
| shipped rack | `onlyPattern` (id 5), all other guns `off` | `common_libs/gun_harness/selector.nim` |
| why | selector measured negative value at every rack size (16, lean8, lean6, pairPC/PK/PL) and 10/10 adversaries | `docs/selector_negative_value.md`, `docs/gun_rack_analysis.md` |
| Pattern hit rate vs DrussGT | ~10.8% | given |
| KNN / Linear / Circular / WallBounce / GF vs DrussGT | 5.6 / 3.0 / 2.9 / 2.7 / 2.1 % | given |
| BitBrain when idle | == Pattern (30 runs/arm: 97/210 vs 97/210 wins) | `docs/bitbrain_vs_tmhorizon_ab.md` |
| BitBrain across 32 legacy opponents | neutral: **-3.6 dmg/run**, +0.02 wins/run, p=0.53/0.91 | `docs/gauntlet_bitbrain_vs_pattern.md` |
| lead **amplitude** | dead axis — full gain sweep {0, 0.25-1.0, 1.0, 1.25, 1.5, 2.0} run, nothing beats Pattern | `docs/bitbrain_campaign.md` |
| offline prediction checks | **veto-only** — may kill a design, never select one | `docs/offline_harness_trust.md` |
| spinner-specific claim | **UNTESTED**: no constant-turn spinner exists in the legacy roster (a note for a later job, not built here unless nearly free) | `docs/gauntlet_bitbrain_vs_pattern.md` |
**Movement context (why the pin):** the movement champion is `TR_MOVEMENT=strafe`
at its current defaults (range 325, tol 25, tilt 15/0.10); it replicated its win
over the shipped `tfil` in four independent sessions (Batch 4: 52.6% vs 38.5%
round-win rate). Movement is held constant at that champion for every gun arm.
---
## Batch 1 — does anything beat the shipped `Pattern`?
**Design.** One frozen binary, six env-only arms, one frozen panel
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 1
wall-follower, 1 corner-camper, 1 spinner, 1 rammer, 1 brawler, 2 aggressive
megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file:
`tools/ab/arms_gun_b1.txt`. Reference: `pattern`.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | `pattern` | `TR_MOVEMENT=strafe` | the arm to beat (shipped `onlyPattern` rack) |
| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, learned gain (panel replication of the 32-opp gauntlet at the pinned-movement standard) |
| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only corrector (predicted to lose; never panel-tested) |
| 4 | `knn` | `+ TR_RACK_PATTERN=off TR_RACK_KNN=both` | best non-Pattern single gun (never panel-tested) |
| 5 | `rack_pk` | `+ TR_RACK_KNN=both` | Pattern + KNN, **selector ON** (different-family hedge) |
| 6 | `rack_pt` | `+ TR_RACK_TMHORIZON=both` | Pattern + TMHorizon, **selector ON** (corrector-family hedge) |
*(`+` = the pinned `TR_MOVEMENT=strafe` is present in every arm; see §1.)*
**Pre-registered prediction (written BEFORE the battles finished):**
1. **No arm beats `pattern` on BOTH primaries**, and the point estimates sit
inside the MDE. `onlyPattern` is confirmed across the panel.
2. `bitbrain` is a **wash** vs `pattern` (replicating the 32-opponent gauntlet:
±3.6 dmg/run, ≈0 wins) — the DrussGT-only gain-config penalty does not carry.
3. `tmhorizon` is **WORSE** than `pattern` (its own offline work predicts it
needs ~80% side accuracy and can reach ~60%); no damage win, no win win.
4. `knn` is **WORSE** than `pattern` on damage (its DrussGT hit rate is roughly
half Pattern's) and not better on wins.
5. The two small racks **do not beat `pattern`**; `rack_pk` lands closer to
Pattern than `knn`-alone does (the selector at least partially hedges back),
but the selector's poor ranking keeps it below the reference.
6. If any surprise exists, it is a **rack arm on damage without wins** — the
`ring`-mover mirror-image trap — and it will not be read as a win.
### Outcome — direct answer
**MEASURED.** Session `/tmp/ab/j121_g1`, commit `1d8143a15e4a038c39dfc2153ff518fa003b8612`
(the pre-registration commit), frozen binary sha256 `aa49a45fec20…`, 15 opponents
× 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started, 0
liveness exclusions**. Every arm declared `TR_MOVEMENT=strafe`, verified in each
bot's own raw-env report; the reference rack line reads `PATTERN` for `pattern`
and the overridden rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the others.
> **DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped
> `Pattern` across the panel on damage/run AND round wins.** Every arm except
> `knn` is **NOT DISTINGUISHABLE** from `pattern` under both the strict and the
> substantive reading of the pre-registered rule. The best nominal challenger,
> `tmhorizon`, is **+0.13 wins/run** (MDE 0.36) and **+8.7 dmg/run** (MDE 13.8)
> — an effect ~3× smaller than the design can detect — and `bitbrain` is
> `+0.11` wins / `+10.3` dmg, also inside the MDE. `knn` is **detectably WORSE**
> on both primaries. The shipped `onlyPattern` rack is therefore **CONFIRMED**
> rather than merely assumed, at the resolution of this panel. That is a real
> result, not a failure.
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `pattern` (REF) | 45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 |
| `bitbrain` | 45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 |
| `tmhorizon` | 45 | 108.3 | 153.7 | **1.64** | **74/135** | **54.8%** | **12.53%** | 428 |
| `knn` | 45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 |
| `rack_pk` | 45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 |
| `rack_pt` | 45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 |
#### Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---:|---:|---|---:|---:|---:|---:|---:|
| `bitbrain` | damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 |
| `bitbrain` | wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 |
| `tmhorizon` | damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 |
| `tmhorizon` | wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 |
| `knn` | damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 |
| `knn` | wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 |
| `rack_pk` | damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 |
| `rack_pk` | wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 |
| `rack_pt` | damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 |
| `rack_pt` | wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 |
#### The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---:|---|---:|---:|---|---|---|---|
| 1 | `tmhorizon` | +0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** |
| 2 | `rack_pt` | +0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | **not distinguishable** | **not distinguishable** |
| 3 | `bitbrain` | +0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** |
| 4 | `rack_pk` | −0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | **not distinguishable** | **not distinguishable** |
| 5 | `knn` | −0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | **WORSE** | **not distinguishable** |
Reference `pattern`: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.
#### Reading (MEASURED, with the mechanism)
* **`knn` is a clean, large loss and is dropped.** It deals **−39 dmg/run** and
loses **−0.62 wins/run**, positive on **0/15** opponents on damage and 1/11 on
wins (p=6e-5 / p=0.012). Its incoming hit rate is *no worse* (13.31% vs
12.84%) — it simply cannot aim: the same movement, a worse gun.
* **The two small racks do not beat `pattern`.** `rack_pk` (Pattern+KNN,
selector ON) is *not* the average of Pattern and KNN: it recovers most of KNN's
damage loss (+7.7 vs KNN-alone's −39.3) but still lands **nominally below**
Pattern on wins (−0.09). `rack_pt` (Pattern+TMHorizon) is statistically
indistinguishable from `tmhorizon`-alone — adding Pattern to the rack changes
nothing, i.e. the selector picks the corrector essentially always. **The
selector's 16-gun negative verdict reproduces at a 2-gun rack**: it does not
manufacture a win it did not have.
* **The only signal is a damage-side hint on the two Pattern-lead correctors**
(`bitbrain` +10.3, `tmhorizon` +8.7, both 11/15 opponents). It is *below the
MDE* (12.0 / 13.8) and *not significant by the pre-registered sign test*
(p=0.1185), so it is **not a win and not a null**. Note the sign-flip test on
the *mean* is nominally significant for `bitbrain` (p=0.032) while the sign
test is not — the positive deltas are larger than the negatives — but with 5
arms compared this is not compelling after multiplicity, and the effect is
under the MDE. **This is exactly the "damage without wins" pattern this
project keeps paying for** (the `ring` mover: +31 dmg/run and *fewer* wins);
here the win deltas (+0.11/+0.13) are themselves sub-MDE.
* **The reference is stable:** `pattern` under `TR_MOVEMENT=strafe` measures
99.6 dmg/run and 50.4% round wins here, matching the movement campaign's
`strafe` reference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same
panel, different job.
#### Pre-registered predictions — scorecard (an honest count)
| # | prediction | outcome |
|---|---|---|
| 1 | no arm beats `pattern` on both primaries, point estimates inside the MDE; `onlyPattern` confirmed | **CORRECT** |
| 2 | `bitbrain` is a wash vs `pattern` (gauntlet: −3.6 dmg, ≈0 wins) | **CORRECT on the verdict, WRONG on the point estimate**: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds |
| 3 | `tmhorizon` is WORSE than `pattern` (its own docs predict it loses) | **WRONG**: it is nominally *better* on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable |
| 4 | `knn` is WORSE than `pattern` on damage, not better on wins | **CORRECT** (detectably on both) |
| 5 | the small racks do not beat `pattern`; `rack_pk` lands closer to Pattern than `knn`-alone | **CORRECT** |
| 6 | any surprise is a rack arm on damage without wins | **PARTLY CORRECT**: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat |
---
## Batch 2 — wider-panel calibration of the damage-side hint
The Batch-1 answer is decisive on the primary question (nothing beats `pattern`)
but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids
calling either way. Resolving it needs **more opponents, not more runs**, so this
batch re-runs exactly those two arms plus the reference on a wider panel.
**Design.** One frozen binary, three env-only arms, panel
`tools/ab/panel_gun_b2.txt` = the j117 32-opponent legacy roster **minus Aurora**
(inert in j117: one fire in 6 battles — it cannot reveal a gun regression)
**plus SpinBot** = **33 opponents** (a superset of Batch 1's 15). **3 runs × 3
rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file:
`tools/ab/arms_gun_b2.txt`. Reference: `pattern`. Movement pinned
`TR_MOVEMENT=strafe` in every arm (same rationale as Batch 1).
| # | arm | env | role |
|---|---|---|---|
| 1 | `pattern` | `TR_MOVEMENT=strafe` | reference |
| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | largest Batch-1 damage edge (+10.3) |
| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | second Batch-1 damage edge (+8.7) |
**Pre-registered prediction (written BEFORE the battles finished):**
1. The `+10` dmg/run hint **does not survive** the wider panel: `bitbrain` and
`tmhorizon` damage deltas fall toward zero and stay inside the new (smaller)
MDE; the per-opponent sign test remains non-significant. Rationale: both arms
come from the same Pattern-lead-correction base, and the prior 32-opponent
gauntlet measured `bitbrain` at **−3.6 dmg/run** on a largely overlapping
roster; +10 on 15 opponents is plausibly a small-panel fluctuation.
2. Round wins remain **flat** for both arms — if anything they regress toward
zero.
3. If instead the damage hint **survives** at a significant cross-opponent sign
test with wins not down, it is recorded as the first gun to **beat** `pattern`
by rule 2 — and the campaign immediately looks for what the corrector is
exploiting. That would be a genuine up-set, and it would **not** be spun as a
null.
### Outcome — Batch 2
*(filled in by the Batch-2 results commit)*
---
## What to try next
*(Rewritten after Batch 1 — these are recommendations, not results. Filled in the
results commit below.)*
## What would make us stop
* **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond
the MDE, or when a gun arm's damage gain is bought with a detectable win loss.
At that point *"`onlyPattern` is the measured optimum of this rack"* is the
conclusion, not a failure (§2 rule 4).
* **Stop a single batch early** only for a contract violation (arena not free,
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
* **Do not invent more arms on the same axis** once two consecutive batches fail
to improve on `pattern` beyond the MDE; move to a *named* new axis instead
(candidate list in "What to try next").
---
## How to run a batch (exact commands)
```sh
# 0. wait for the arena (j120's movement run may still be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference pattern \
--outdir /tmp/ab/j121_g1
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
```
`--reference` may be ANY arm of the session: re-analyzing an old session with a
different reference is a free pairwise comparison with no battles.
## Session log (outdirs are in `/tmp` and are NOT committed)
| session | commit | battles | arms | verdict |
|---|---|---:|---|---|
| `/tmp/ab/j121_g1` | *(recorded at run time)* | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | *(Batch 1 outcome below)* |