diff --git a/docs/gun_campaign.md b/docs/gun_campaign.md index 9b4b8a8..8d00a4a 100644 --- a/docs/gun_campaign.md +++ b/docs/gun_campaign.md @@ -48,6 +48,14 @@ evidence is the **number of opponents**, not the number of runs. | Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named | | Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched | +> **Movement-default flip note (MEASURED).** During this job a parallel job +> (j120, commit `3fd6db9`) flipped the shipped `TR_MOVEMENT` default from `tfil` +> to `strafe`. Batch 1 built before the flip, Batch 2 after. **Every arm in both +> batches declares `TR_MOVEMENT=strafe`**, so the effective movement is identical +> and the two batches are directly comparable — the pin did its job. The reference +> rack lines in the boot reports read `PATTERN` for `pattern`, and the active 1v1 +> rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the overridden arms. + ### Movement is PINNED in every arm: `TR_MOVEMENT=strafe` **Reason (hard rule).** A movement-default flip is landing from job j120 during @@ -283,22 +291,141 @@ rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file: ### Outcome — Batch 2 -*(filled in by the Batch-2 results commit)* +**MEASURED.** Session `/tmp/ab/j121_g2`, commit +`c343c00aaf800bc05ce9bf1b277054c11f5f577c`, frozen binary sha256 `d267ab78f6ce…`, +**33 opponents × 3 arms × 3 runs × 3 rounds = 297 battles, 0 failed, 0 never +started, 0 liveness exclusions**. + +**A movement-default flip landed between the two batches** (commit `3fd6db9`, +`TR_MOVEMENT` default `tfil` → `strafe`; the only source change). Because +**every** arm in **both** batches pins `TR_MOVEMENT=strafe`, the effective +movement is identical in Batch 1 and Batch 2 — the batches are directly +comparable and the gun comparison is unconfounded. This is exactly what the pin +was for. + +> **DIRECT ANSWER (MEASURED, n=33 opponents): the Batch-1 damage hint did NOT +> survive.** `bitbrain` collapses to **+1.8 dmg/run** (19/33 opponents, sign test +> p=0.49) with **+0.11 wins/run** — a wash. `tmhorizon` **reverses sign and is +> detectably WORSE on damage**: **−10.4 dmg/run** (10/33 opponents, sign test +> p=0.035, sign-flip p=0.0047, 95% CI [−17.2, −3.6]) at unchanged wins (−0.03). +> The only open signal in Batch 1 was a **small-panel fluctuation**, and the +> wider panel resolves it **against** the correctors. Combined with Batch 1: +> **nothing beats the shipped `Pattern` on damage/run AND round wins**; +> `onlyPattern` is CONFIRMED on 15 opponents and again on 33. + +#### Pooled dashboard (all valid runs — explanation only, NOT the verdict) + +| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance | +|---|---:|---:|---:|---:|---:|---:|---:|---:| +| `pattern` (REF) | 99 | 131.2 | 126.4 | 1.90 | 188/297 | 63.3% | 12.93% | 417 | +| `bitbrain` | 99 | 133.0 | 126.7 | **2.01** | **199/297** | **67.0%** | 12.91% | 417 | +| `tmhorizon` | 99 | 120.8 | 133.2 | 1.87 | 185/297 | 62.3% | 13.04% | 422 | + +#### Cross-opponent aggregation (the verdict layer, damage + wins) + +| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE | +|---|---|---:|---:|---|---:|---:|---:|---:|---:| +| `bitbrain` | damage | +1.80 | 18.95 | [−4.67, +8.26] | 19/33 | 0.4869 | 0.5914 | 0.5675 | 9.24 | +| `bitbrain` | wins | +0.11 | 0.48 | [−0.05, +0.28] | 12/19 | 0.3593 | 0.2436 | 0.5716 | 0.24 | +| `tmhorizon` | damage | **−10.40** | 20.01 | **[−17.22, −3.57]** | 10/33 | **0.03508** | **0.00472** | **0.006975** | 9.76 | +| `tmhorizon` | wins | −0.03 | 0.59 | [−0.23, +0.17] | 9/19 | 1.0 | 0.8459 | 0.4804 | 0.29 | + +#### The pre-registered verdict (verbatim) + +| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) | +|---:|---|---:|---:|---|---|---|---| +| 1 | `bitbrain` | +0.11 | +1.8 | 12/19 p=0.3593 | 19/33 p=0.4869 | **not distinguishable** | **not distinguishable** | +| 2 | `tmhorizon` | −0.03 | −10.4 | 9/19 p=1 | 10/33 p=0.03508 | **WORSE** | **WORSE** | + +Reference `pattern`: 131.2 dmg/run, 1.90 wins/run, 12.93% incoming, 417 px. + +#### Reading + +* **BitBrain is a wash — now measured three times.** vs the shipped Pattern: + −3.6 dmg/run on the 32-opponent legacy gauntlet (with `tfil` movement, + `docs/gauntlet_bitbrain_vs_pattern.md`), **+10.3** on the 15-opponent Batch-1 + panel, **+1.8** here on 33 opponents. All three sit inside their own MDEs; the + pooled evidence is a zero. The Batch-1 `+10` was a 15-opponent fluctuation. +* **TMHorizon is the one arm the wider panel separates — and it loses.** Its + Batch-1 sign (+8.7) does not merely fail to replicate, it **inverts** + (−10.4, 10/33, p=0.035). The damage loss is broad (24/33 opponents negative, + worst `Aristocles` −61, `Jen` −54, `WallAvoider` −41), so it is not one + outlier. Since TMHorizon was already predicted to lose by its own design docs, + this is the first *confirming* live panel measurement of that prediction. +* **Wins are flat for both arms** (+0.11 / −0.03, MDEs 0.24 / 0.29): the + correctors change damage at most, never survival. In this harness round wins + are survival wins, so neither is a win. +* **The corrector family is closed** as a source of a `pattern`-beating gun at + the resolutions tested. + +#### Pre-registered predictions — scorecard + +| # | prediction | outcome | +|---|---|---| +| 1 | the +10 dmg hint does not survive the wider panel; deltas fall inside the new MDE; sign test non-significant | **CORRECT** (`bitbrain` +1.8, p=0.49; `tmhorizon` even reversed to −10.4, p=0.035) | +| 2 | wins remain flat | **CORRECT** (+0.11 / −0.03) | +| 3 | if the hint survived at p<0.05 with wins not down, it is the first gun to beat `pattern` | **not invoked** | + +**Final direct answer for the phase (MEASURED).** Across two frozen panels +(15 and 33 opponents), a single frozen-binary-per-session design, 567 battles, +seven distinct gun configurations, and movement pinned identically in every arm: +**no rack gun and no small rack beats the shipped `Pattern` on damage/run AND +round wins.** Every candidate is either not distinguishable (`bitbrain`, +`tmhorizon`, `rack_pk`, `rack_pt` at n=15), a detectable loss +(`knn`; `tmhorizon` at n=33), or a sub-MDE damage-only wobble +(`bitbrain`). The shipped **`onlyPattern` rack is CONFIRMED**, not merely +assumed — the gun axis's first phase ends in a measured optimum of the tested +design space. --- -## What to try next +## What to try next (rewritten AFTER Batches 1–2 — recommendations, not results) -*(Rewritten after Batch 1 — these are recommendations, not results. Filled in the -results commit below.)* +**The primary question is answered:** nothing beats the shipped `pattern` across +the frozen panel, and `onlyPattern` is CONFIRMED. Ranked by value per battle: + +1. **The damage-side hint is now CLOSED.** Batch 2 (33 opponents) killed it: + `bitbrain` is a wash (+1.8 dmg/run, p=0.49) and `tmhorizon` is detectably + **worse** (−10.4 dmg/run, p=0.035). Do **not** spend another batch on the + Pattern-lead correctors; three independent measurements of `bitbrain` vs + Pattern now average ~0 (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs). +2. **Lead INFORMATION remains the only named gun axis** (given: amplitude is + dead). But the two live implementations of it (BitBrain's gain, TMHorizon's + shift) are both measured negative/neutral across panels. A future axis would + need a *different* information source, not another knob on these two — e.g. a + new base prediction, or a gun-agnostic ensemble. Do not re-open the correctors. +3. **The selector is confirmed negative at a 2-gun rack** (`rack_pt` == + `tmhorizon`; `rack_pk` below `pattern` on wins). Do not re-open it with a + larger rack: the 16-gun negative already stands and Batch 1 shows the + mechanism does not manufacture a win. +4. **Drop `knn`** (detectably worse on both primaries, 0/15 on damage). +5. **Never chase damage without wins.** This harness's round wins are survival + wins, so the gun's job is to *kill*; a damage gain that does not raise the win + rate is the `ring`-mover trap in a new costume (and `tmhorizon`'s n=33 result + is a *negative* damage signal with flat wins — also not a win). +6. **Spinner-specific claim (owner's) remains UNTESTED.** The legacy roster has + **no constant-turn spinner**; SpinBot is a periodic circle-mover and is in + both panels (Batch 1: `bitbrain` +25.3 dmg / +0 wins, `tmhorizon` +15.2 / +0 + wins; on the 33-panel SpinBot's deltas are in the per-opponent tables). + Building a true constant-turn spinner opponent is a nearly-free follow-up and + is the one part of the owner's claim this campaign has not touched. ## What would make us stop +**PHASE STATUS (updated after Batch 2): STOPPED — the condition fired.** Across +Batches 1–2 no arm beats `pattern` beyond the MDE, and no arm shows a +`>= +MDE` damage gain with p<0.10 (`bitbrain` +1.8 vs MDE 9.24; `tmhorizon` +actually negative). The gun stage's first phase is therefore **closed** with +*"the shipped `onlyPattern` rack is the best gun configuration we have measured +across the panels (15 and 33 opponents)"* — a successful outcome, not a failure. +Further gun work must open a **named new axis** (see *What to try next*), not +another arm on the same axis. + * **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond the MDE, or when a gun arm's damage gain is bought with a detectable win loss. At that point *"`onlyPattern` is the measured optimum of this rack"* is the - conclusion, not a failure (§2 rule 4). + conclusion, not a failure (§2 rule 4). **This fired after Batch 2.** * **Stop a single batch early** only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring. * **Do not invent more arms on the same axis** once two consecutive batches fail @@ -310,7 +437,7 @@ results commit below.)* ## How to run a batch (exact commands) ```sh -# 0. wait for the arena (j120's movement run may still be fighting) +# 0. wait for the arena (other campaign jobs may be fighting) TOURNAMENT_NIMCACHE=/tmp/nc_j121 \ tools/ab/tournament_run.sh \ --arms tools/ab/arms_gun_b1.txt \ @@ -319,6 +446,10 @@ tools/ab/tournament_run.sh \ --reference pattern \ --outdir /tmp/ab/j121_g1 +# Batch 2 (the wider-panel calibration) was the same command with +# --arms tools/ab/arms_gun_b2.txt --panel tools/ab/panel_gun_b2.txt \ +# --outdir /tmp/ab/j121_g2 --wait-arena 50 + # 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern ``` @@ -330,4 +461,5 @@ different reference is a free pairwise comparison with no battles. | session | commit | battles | arms | verdict | |---|---|---:|---|---| -| `/tmp/ab/j121_g1` | *(recorded at run time)* | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | *(Batch 1 outcome below)* | +| `/tmp/ab/j121_g1` | `1d8143a` | 270 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | **nothing beats the shipped `pattern`**; all arms not distinguishable except `knn` = WORSE; `onlyPattern` CONFIRMED | +| `/tmp/ab/j121_g2` | `c343c00` | 297 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon | **the Batch-1 damage hint does not survive**: `bitbrain` +1.8 dmg/run p=0.49 (wash); `tmhorizon` −10.4 dmg/run p=0.035 (WORSE); wins flat |