From 087e26955f2e363dc1a0383b4ab7d32352b5ded5 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sat, 26 Sep 2026 02:58:41 +0200 Subject: [PATCH] gun j121 Batch 1 outcome: onlyPattern CONFIRMED across the 15-opponent panel; pre-register Batch 2 wider-panel calibration --- docs/gun_campaign.md | 141 +++++++++++++++++++++++++++++++++++++- tools/ab/arms_gun_b2.txt | 30 ++++++++ tools/ab/panel_gun_b2.txt | 53 ++++++++++++++ 3 files changed, 223 insertions(+), 1 deletion(-) create mode 100644 tools/ab/arms_gun_b2.txt create mode 100644 tools/ab/panel_gun_b2.txt diff --git a/docs/gun_campaign.md b/docs/gun_campaign.md index 2355e7d..9b4b8a8 100644 --- a/docs/gun_campaign.md +++ b/docs/gun_campaign.md @@ -145,7 +145,146 @@ megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file: 6. If any surprise exists, it is a **rack arm on damage without wins** — the `ring`-mover mirror-image trap — and it will not be read as a win. - +### Outcome — direct answer + +**MEASURED.** Session `/tmp/ab/j121_g1`, commit `1d8143a15e4a038c39dfc2153ff518fa003b8612` +(the pre-registration commit), frozen binary sha256 `aa49a45fec20…`, 15 opponents +× 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started, 0 +liveness exclusions**. Every arm declared `TR_MOVEMENT=strafe`, verified in each +bot's own raw-env report; the reference rack line reads `PATTERN` for `pattern` +and the overridden rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the others. + +> **DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped +> `Pattern` across the panel on damage/run AND round wins.** Every arm except +> `knn` is **NOT DISTINGUISHABLE** from `pattern` under both the strict and the +> substantive reading of the pre-registered rule. The best nominal challenger, +> `tmhorizon`, is **+0.13 wins/run** (MDE 0.36) and **+8.7 dmg/run** (MDE 13.8) +> — an effect ~3× smaller than the design can detect — and `bitbrain` is +> `+0.11` wins / `+10.3` dmg, also inside the MDE. `knn` is **detectably WORSE** +> on both primaries. The shipped `onlyPattern` rack is therefore **CONFIRMED** +> rather than merely assumed, at the resolution of this panel. That is a real +> result, not a failure. + +#### Pooled dashboard (all valid runs — explanation only, NOT the verdict) + +| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance | +|---|---:|---:|---:|---:|---:|---:|---:|---:| +| `pattern` (REF) | 45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 | +| `bitbrain` | 45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 | +| `tmhorizon` | 45 | 108.3 | 153.7 | **1.64** | **74/135** | **54.8%** | **12.53%** | 428 | +| `knn` | 45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 | +| `rack_pk` | 45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 | +| `rack_pt` | 45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 | + +#### Cross-opponent aggregation (the verdict layer, damage + wins) + +| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE | +|---|---|---:|---:|---|---:|---:|---:|---:|---:| +| `bitbrain` | damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 | +| `bitbrain` | wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 | +| `tmhorizon` | damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 | +| `tmhorizon` | wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 | +| `knn` | damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 | +| `knn` | wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 | +| `rack_pk` | damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 | +| `rack_pk` | wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 | +| `rack_pt` | damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 | +| `rack_pt` | wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 | + +#### The pre-registered verdict (verbatim) + +| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) | +|---:|---|---:|---:|---|---|---|---| +| 1 | `tmhorizon` | +0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** | +| 2 | `rack_pt` | +0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | **not distinguishable** | **not distinguishable** | +| 3 | `bitbrain` | +0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** | +| 4 | `rack_pk` | −0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | **not distinguishable** | **not distinguishable** | +| 5 | `knn` | −0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | **WORSE** | **not distinguishable** | + +Reference `pattern`: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px. + +#### Reading (MEASURED, with the mechanism) + +* **`knn` is a clean, large loss and is dropped.** It deals **−39 dmg/run** and + loses **−0.62 wins/run**, positive on **0/15** opponents on damage and 1/11 on + wins (p=6e-5 / p=0.012). Its incoming hit rate is *no worse* (13.31% vs + 12.84%) — it simply cannot aim: the same movement, a worse gun. +* **The two small racks do not beat `pattern`.** `rack_pk` (Pattern+KNN, + selector ON) is *not* the average of Pattern and KNN: it recovers most of KNN's + damage loss (+7.7 vs KNN-alone's −39.3) but still lands **nominally below** + Pattern on wins (−0.09). `rack_pt` (Pattern+TMHorizon) is statistically + indistinguishable from `tmhorizon`-alone — adding Pattern to the rack changes + nothing, i.e. the selector picks the corrector essentially always. **The + selector's 16-gun negative verdict reproduces at a 2-gun rack**: it does not + manufacture a win it did not have. +* **The only signal is a damage-side hint on the two Pattern-lead correctors** + (`bitbrain` +10.3, `tmhorizon` +8.7, both 11/15 opponents). It is *below the + MDE* (12.0 / 13.8) and *not significant by the pre-registered sign test* + (p=0.1185), so it is **not a win and not a null**. Note the sign-flip test on + the *mean* is nominally significant for `bitbrain` (p=0.032) while the sign + test is not — the positive deltas are larger than the negatives — but with 5 + arms compared this is not compelling after multiplicity, and the effect is + under the MDE. **This is exactly the "damage without wins" pattern this + project keeps paying for** (the `ring` mover: +31 dmg/run and *fewer* wins); + here the win deltas (+0.11/+0.13) are themselves sub-MDE. +* **The reference is stable:** `pattern` under `TR_MOVEMENT=strafe` measures + 99.6 dmg/run and 50.4% round wins here, matching the movement campaign's + `strafe` reference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same + panel, different job. + +#### Pre-registered predictions — scorecard (an honest count) + +| # | prediction | outcome | +|---|---|---| +| 1 | no arm beats `pattern` on both primaries, point estimates inside the MDE; `onlyPattern` confirmed | **CORRECT** | +| 2 | `bitbrain` is a wash vs `pattern` (gauntlet: −3.6 dmg, ≈0 wins) | **CORRECT on the verdict, WRONG on the point estimate**: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds | +| 3 | `tmhorizon` is WORSE than `pattern` (its own docs predict it loses) | **WRONG**: it is nominally *better* on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable | +| 4 | `knn` is WORSE than `pattern` on damage, not better on wins | **CORRECT** (detectably on both) | +| 5 | the small racks do not beat `pattern`; `rack_pk` lands closer to Pattern than `knn`-alone | **CORRECT** | +| 6 | any surprise is a rack arm on damage without wins | **PARTLY CORRECT**: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat | + +--- + +## Batch 2 — wider-panel calibration of the damage-side hint + +The Batch-1 answer is decisive on the primary question (nothing beats `pattern`) +but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids +calling either way. Resolving it needs **more opponents, not more runs**, so this +batch re-runs exactly those two arms plus the reference on a wider panel. + +**Design.** One frozen binary, three env-only arms, panel +`tools/ab/panel_gun_b2.txt` = the j117 32-opponent legacy roster **minus Aurora** +(inert in j117: one fire in 6 battles — it cannot reveal a gun regression) +**plus SpinBot** = **33 opponents** (a superset of Batch 1's 15). **3 runs × 3 +rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file: +`tools/ab/arms_gun_b2.txt`. Reference: `pattern`. Movement pinned +`TR_MOVEMENT=strafe` in every arm (same rationale as Batch 1). + +| # | arm | env | role | +|---|---|---|---| +| 1 | `pattern` | `TR_MOVEMENT=strafe` | reference | +| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | largest Batch-1 damage edge (+10.3) | +| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | second Batch-1 damage edge (+8.7) | + +**Pre-registered prediction (written BEFORE the battles finished):** +1. The `+10` dmg/run hint **does not survive** the wider panel: `bitbrain` and + `tmhorizon` damage deltas fall toward zero and stay inside the new (smaller) + MDE; the per-opponent sign test remains non-significant. Rationale: both arms + come from the same Pattern-lead-correction base, and the prior 32-opponent + gauntlet measured `bitbrain` at **−3.6 dmg/run** on a largely overlapping + roster; +10 on 15 opponents is plausibly a small-panel fluctuation. +2. Round wins remain **flat** for both arms — if anything they regress toward + zero. +3. If instead the damage hint **survives** at a significant cross-opponent sign + test with wins not down, it is recorded as the first gun to **beat** `pattern` + by rule 2 — and the campaign immediately looks for what the corrector is + exploiting. That would be a genuine up-set, and it would **not** be spun as a + null. + +### Outcome — Batch 2 + +*(filled in by the Batch-2 results commit)* + --- diff --git a/tools/ab/arms_gun_b2.txt b/tools/ab/arms_gun_b2.txt new file mode 100644 index 0000000..19ff221 --- /dev/null +++ b/tools/ab/arms_gun_b2.txt @@ -0,0 +1,30 @@ +# ───────────────────────────────────────────────────────────────────────────── +# arms_gun_b2.txt — BATCH 2 of the GUN campaign: a WIDER-PANEL calibration of the +# two arms that carried the only nominal positive signal in Batch 1. +# +# Batch 1 (docs/gun_campaign.md, 15-opponent frozen panel, 270 battles) found NO +# arm beating the shipped `pattern` beyond the MDE, but two arms sat on the +# positive side of damage with a flat win count: +# BitBrain Δdmg +10.3 (11/15 opponents, sign test p=0.1185, sign-flip +# p=0.0325), Δwins +0.11 (MDE 0.35) +# TMHorizon Δdmg +8.7 (11/15, p=0.1185, sign-flip p=0.0992), Δwins +0.13 +# Both damage effects are BELOW the Batch-1 MDE (~12-14 dmg/run), so rule 3 says +# they are not detectable — they must not be called a win OR a null. Resolving +# them needs more OPPONENTS, not more runs, so this batch re-runs exactly those +# two arms plus the reference on tools/ab/panel_gun_b2.txt (33 opponents). +# +# Format: name | ENV=value ENV=value | label +# MOVEMENT IS PINNED (`TR_MOVEMENT=strafe`) in EVERY arm — same rationale as +# Batch 1: a default flip may be landing from j120, and the pin makes every +# delta a pure gun delta and satisfies the analyzer liveness rule. +# ───────────────────────────────────────────────────────────────────────────── + +# 1. REFERENCE — shipped onlyPattern rack, movement pinned. +pattern | TR_MOVEMENT=strafe | shipped onlyPattern rack (REFERENCE) + +# 2. BitBrain learned-gain corrector (the arm with the largest nominal damage +# edge in Batch 1: +10.3 dmg/run, 11/15). +bitbrain | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay | BitBrain-only, learned gain (gains 1.0..2.0, decay) + +# 3. TMHorizon corrector (Batch 1: +8.7 dmg/run, +0.13 wins/run, 11/15). +tmhorizon | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon-only rack diff --git a/tools/ab/panel_gun_b2.txt b/tools/ab/panel_gun_b2.txt new file mode 100644 index 0000000..743de3d --- /dev/null +++ b/tools/ab/panel_gun_b2.txt @@ -0,0 +1,53 @@ +# ───────────────────────────────────────────────────────────────────────────── +# panel_gun_b2.txt — the GUN campaign's WIDER PANEL (Batch 2, 2026-09-26). +# +# DO NOT ADD, REMOVE OR REORDER AN OPPONENT without starting a new batch number +# in docs/gun_campaign.md. Batch 1 used the frozen 15-opponent +# `panel_movement.txt`; Batch 2 widens to this 32-opponent superset because the +# Batch-1 arms with a nominal damage edge (BitBrain +10.3, TMHorizon +8.7 +# dmg/run, 11/15 opponents) sat INSIDE the Batch-1 MDE (~12-14 dmg/run), so a +# wider panel is required to resolve them (rule 3: an effect under the MDE is +# not detectable, so it must not be called a win OR a null). +# +# Composition: the j117 32-opponent legacy roster MINUS Aurora (inert in j117: +# a single fire in 6 battles — it cannot reveal a gun regression) PLUS SpinBot +# (the Batch-1 spinner control, absent from the legacy roster) = 31+1 = 32. +# Every batch-1 opponent except none is covered; 17 opponents are added. +# +# Format: |