gun j121 Batch 1 outcome: onlyPattern CONFIRMED across the 15-opponent panel; pre-register Batch 2 wider-panel calibration

This commit is contained in:
2026-09-26 02:58:41 +02:00
parent 514674886d
commit 087e26955f
3 changed files with 223 additions and 1 deletions
+140 -1
View File
@@ -145,7 +145,146 @@ megas), **3 runs × 3 rounds per (opponent, arm) = 270 battles**. Arm file:
6. If any surprise exists, it is a **rack arm on damage without wins** — the
`ring`-mover mirror-image trap — and it will not be read as a win.
<!-- OUTCOME-ANCHOR -->
### Outcome — direct answer
**MEASURED.** Session `/tmp/ab/j121_g1`, commit `1d8143a15e4a038c39dfc2153ff518fa003b8612`
(the pre-registration commit), frozen binary sha256 `aa49a45fec20…`, 15 opponents
× 6 arms × 3 runs × 3 rounds = **270 battles, 0 failed, 0 never started, 0
liveness exclusions**. Every arm declared `TR_MOVEMENT=strafe`, verified in each
bot's own raw-env report; the reference rack line reads `PATTERN` for `pattern`
and the overridden rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the others.
> **DIRECT ANSWER (MEASURED, n=15 opponents): NO — nothing beats the shipped
> `Pattern` across the panel on damage/run AND round wins.** Every arm except
> `knn` is **NOT DISTINGUISHABLE** from `pattern` under both the strict and the
> substantive reading of the pre-registered rule. The best nominal challenger,
> `tmhorizon`, is **+0.13 wins/run** (MDE 0.36) and **+8.7 dmg/run** (MDE 13.8)
> — an effect ~3× smaller than the design can detect — and `bitbrain` is
> `+0.11` wins / `+10.3` dmg, also inside the MDE. `knn` is **detectably WORSE**
> on both primaries. The shipped `onlyPattern` rack is therefore **CONFIRMED**
> rather than merely assumed, at the resolution of this panel. That is a real
> result, not a failure.
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `pattern` (REF) | 45 | 99.6 | 155.9 | 1.51 | 68/135 | 50.4% | 12.84% | 435 |
| `bitbrain` | 45 | 109.9 | 151.2 | 1.62 | 73/135 | 54.1% | 12.87% | 427 |
| `tmhorizon` | 45 | 108.3 | 153.7 | **1.64** | **74/135** | **54.8%** | **12.53%** | 428 |
| `knn` | 45 | 60.3 | 175.7 | 0.89 | 40/135 | 29.6% | 13.31% | 430 |
| `rack_pk` | 45 | 107.3 | 161.3 | 1.42 | 64/135 | 47.4% | 13.65% | 432 |
| `rack_pt` | 45 | 105.3 | 149.9 | 1.64 | 74/135 | 54.8% | 12.90% | 424 |
#### Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---:|---:|---|---:|---:|---:|---:|---:|
| `bitbrain` | damage | +10.34 | 16.65 | [+1.12, +19.57] | 11/15 | 0.1185 | 0.03247 | 0.05006 | 12.05 |
| `bitbrain` | wins | +0.11 | 0.48 | [-0.16, +0.38] | 6/9 | 0.5078 | 0.4922 | 0.4764 | 0.35 |
| `tmhorizon` | damage | +8.71 | 19.10 | [-1.86, +19.29] | 11/15 | 0.1185 | 0.09918 | 0.1183 | 13.82 |
| `tmhorizon` | wins | +0.13 | 0.50 | [-0.14, +0.41] | 5/9 | 1.0 | 0.4062 | 0.3118 | 0.36 |
| `knn` | damage | −39.33 | 23.15 | [−52.15, −26.51] | 0/15 | 6.1e-5 | 6.1e-5 | 0.0007 | 16.75 |
| `knn` | wins | −0.62 | 0.59 | [−0.95, −0.30] | 1/11 | 0.01172 | 0.001953 | 0.004948 | 0.43 |
| `rack_pk` | damage | +7.68 | 32.09 | [−10.10, +25.45] | 9/15 | 0.6072 | 0.423 | 0.5895 | 23.21 |
| `rack_pk` | wins | −0.09 | 0.60 | [−0.42, +0.24] | 6/11 | 1.0 | 0.6738 | 0.5932 | 0.43 |
| `rack_pt` | damage | +5.69 | 16.50 | [−3.45, +14.83] | 10/15 | 0.3018 | 0.2026 | 0.2681 | 11.93 |
| `rack_pt` | wins | +0.13 | 0.41 | [−0.10, +0.36] | 6/10 | 0.7539 | 0.3242 | 0.1997 | 0.30 |
#### The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---:|---|---:|---:|---|---|---|---|
| 1 | `tmhorizon` | +0.13 | +8.7 | 5/9 p=1 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** |
| 2 | `rack_pt` | +0.13 | +5.7 | 6/10 p=0.7539 | 10/15 p=0.3018 | **not distinguishable** | **not distinguishable** |
| 3 | `bitbrain` | +0.11 | +10.3 | 6/9 p=0.5078 | 11/15 p=0.1185 | **not distinguishable** | **not distinguishable** |
| 4 | `rack_pk` | −0.09 | +7.7 | 6/11 p=1 | 9/15 p=0.6072 | **not distinguishable** | **not distinguishable** |
| 5 | `knn` | −0.62 | −39.3 | 1/11 p=0.01172 | 0/15 p=6.104e-05 | **WORSE** | **not distinguishable** |
Reference `pattern`: 99.6 dmg/run, 1.51 wins/run, 12.84% incoming, 435 px.
#### Reading (MEASURED, with the mechanism)
* **`knn` is a clean, large loss and is dropped.** It deals **−39 dmg/run** and
loses **−0.62 wins/run**, positive on **0/15** opponents on damage and 1/11 on
wins (p=6e-5 / p=0.012). Its incoming hit rate is *no worse* (13.31% vs
12.84%) — it simply cannot aim: the same movement, a worse gun.
* **The two small racks do not beat `pattern`.** `rack_pk` (Pattern+KNN,
selector ON) is *not* the average of Pattern and KNN: it recovers most of KNN's
damage loss (+7.7 vs KNN-alone's −39.3) but still lands **nominally below**
Pattern on wins (−0.09). `rack_pt` (Pattern+TMHorizon) is statistically
indistinguishable from `tmhorizon`-alone — adding Pattern to the rack changes
nothing, i.e. the selector picks the corrector essentially always. **The
selector's 16-gun negative verdict reproduces at a 2-gun rack**: it does not
manufacture a win it did not have.
* **The only signal is a damage-side hint on the two Pattern-lead correctors**
(`bitbrain` +10.3, `tmhorizon` +8.7, both 11/15 opponents). It is *below the
MDE* (12.0 / 13.8) and *not significant by the pre-registered sign test*
(p=0.1185), so it is **not a win and not a null**. Note the sign-flip test on
the *mean* is nominally significant for `bitbrain` (p=0.032) while the sign
test is not — the positive deltas are larger than the negatives — but with 5
arms compared this is not compelling after multiplicity, and the effect is
under the MDE. **This is exactly the "damage without wins" pattern this
project keeps paying for** (the `ring` mover: +31 dmg/run and *fewer* wins);
here the win deltas (+0.11/+0.13) are themselves sub-MDE.
* **The reference is stable:** `pattern` under `TR_MOVEMENT=strafe` measures
99.6 dmg/run and 50.4% round wins here, matching the movement campaign's
`strafe` reference in Batch 4 (103.5 dmg/run, 52.6%) — same movement, same
panel, different job.
#### Pre-registered predictions — scorecard (an honest count)
| # | prediction | outcome |
|---|---|---|
| 1 | no arm beats `pattern` on both primaries, point estimates inside the MDE; `onlyPattern` confirmed | **CORRECT** |
| 2 | `bitbrain` is a wash vs `pattern` (gauntlet: −3.6 dmg, ≈0 wins) | **CORRECT on the verdict, WRONG on the point estimate**: +10.3 dmg here vs −3.6 in the 32-opponent gauntlet; both inside their MDEs. Direction differs, verdict (wash) holds |
| 3 | `tmhorizon` is WORSE than `pattern` (its own docs predict it loses) | **WRONG**: it is nominally *better* on both primaries (+8.7 dmg, +0.13 wins), though not distinguishable |
| 4 | `knn` is WORSE than `pattern` on damage, not better on wins | **CORRECT** (detectably on both) |
| 5 | the small racks do not beat `pattern`; `rack_pk` lands closer to Pattern than `knn`-alone | **CORRECT** |
| 6 | any surprise is a rack arm on damage without wins | **PARTLY CORRECT**: the nominal damage edge is on the single-gun correctors and the racks; every one of them is win-flat |
---
## Batch 2 — wider-panel calibration of the damage-side hint
The Batch-1 answer is decisive on the primary question (nothing beats `pattern`)
but the two corrector arms carry a sub-MDE damage hint that rule 3 forbids
calling either way. Resolving it needs **more opponents, not more runs**, so this
batch re-runs exactly those two arms plus the reference on a wider panel.
**Design.** One frozen binary, three env-only arms, panel
`tools/ab/panel_gun_b2.txt` = the j117 32-opponent legacy roster **minus Aurora**
(inert in j117: one fire in 6 battles — it cannot reveal a gun regression)
**plus SpinBot** = **33 opponents** (a superset of Batch 1's 15). **3 runs × 3
rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file:
`tools/ab/arms_gun_b2.txt`. Reference: `pattern`. Movement pinned
`TR_MOVEMENT=strafe` in every arm (same rationale as Batch 1).
| # | arm | env | role |
|---|---|---|---|
| 1 | `pattern` | `TR_MOVEMENT=strafe` | reference |
| 2 | `bitbrain` | `+ TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | largest Batch-1 damage edge (+10.3) |
| 3 | `tmhorizon` | `+ TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | second Batch-1 damage edge (+8.7) |
**Pre-registered prediction (written BEFORE the battles finished):**
1. The `+10` dmg/run hint **does not survive** the wider panel: `bitbrain` and
`tmhorizon` damage deltas fall toward zero and stay inside the new (smaller)
MDE; the per-opponent sign test remains non-significant. Rationale: both arms
come from the same Pattern-lead-correction base, and the prior 32-opponent
gauntlet measured `bitbrain` at **−3.6 dmg/run** on a largely overlapping
roster; +10 on 15 opponents is plausibly a small-panel fluctuation.
2. Round wins remain **flat** for both arms — if anything they regress toward
zero.
3. If instead the damage hint **survives** at a significant cross-opponent sign
test with wins not down, it is recorded as the first gun to **beat** `pattern`
by rule 2 — and the campaign immediately looks for what the corrector is
exploiting. That would be a genuine up-set, and it would **not** be spun as a
null.
### Outcome — Batch 2
*(filled in by the Batch-2 results commit)*
---
+30
View File
@@ -0,0 +1,30 @@
# ─────────────────────────────────────────────────────────────────────────────
# arms_gun_b2.txt — BATCH 2 of the GUN campaign: a WIDER-PANEL calibration of the
# two arms that carried the only nominal positive signal in Batch 1.
#
# Batch 1 (docs/gun_campaign.md, 15-opponent frozen panel, 270 battles) found NO
# arm beating the shipped `pattern` beyond the MDE, but two arms sat on the
# positive side of damage with a flat win count:
# BitBrain Δdmg +10.3 (11/15 opponents, sign test p=0.1185, sign-flip
# p=0.0325), Δwins +0.11 (MDE 0.35)
# TMHorizon Δdmg +8.7 (11/15, p=0.1185, sign-flip p=0.0992), Δwins +0.13
# Both damage effects are BELOW the Batch-1 MDE (~12-14 dmg/run), so rule 3 says
# they are not detectable — they must not be called a win OR a null. Resolving
# them needs more OPPONENTS, not more runs, so this batch re-runs exactly those
# two arms plus the reference on tools/ab/panel_gun_b2.txt (33 opponents).
#
# Format: name | ENV=value ENV=value | label
# MOVEMENT IS PINNED (`TR_MOVEMENT=strafe`) in EVERY arm — same rationale as
# Batch 1: a default flip may be landing from j120, and the pin makes every
# delta a pure gun delta and satisfies the analyzer liveness rule.
# ─────────────────────────────────────────────────────────────────────────────
# 1. REFERENCE — shipped onlyPattern rack, movement pinned.
pattern | TR_MOVEMENT=strafe | shipped onlyPattern rack (REFERENCE)
# 2. BitBrain learned-gain corrector (the arm with the largest nominal damage
# edge in Batch 1: +10.3 dmg/run, 11/15).
bitbrain | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay | BitBrain-only, learned gain (gains 1.0..2.0, decay)
# 3. TMHorizon corrector (Batch 1: +8.7 dmg/run, +0.13 wins/run, 11/15).
tmhorizon | TR_MOVEMENT=strafe TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon-only rack
+53
View File
@@ -0,0 +1,53 @@
# ─────────────────────────────────────────────────────────────────────────────
# panel_gun_b2.txt — the GUN campaign's WIDER PANEL (Batch 2, 2026-09-26).
#
# DO NOT ADD, REMOVE OR REORDER AN OPPONENT without starting a new batch number
# in docs/gun_campaign.md. Batch 1 used the frozen 15-opponent
# `panel_movement.txt`; Batch 2 widens to this 32-opponent superset because the
# Batch-1 arms with a nominal damage edge (BitBrain +10.3, TMHorizon +8.7
# dmg/run, 11/15 opponents) sat INSIDE the Batch-1 MDE (~12-14 dmg/run), so a
# wider panel is required to resolve them (rule 3: an effect under the MDE is
# not detectable, so it must not be called a win OR a null).
#
# Composition: the j117 32-opponent legacy roster MINUS Aurora (inert in j117:
# a single fire in 6 battles — it cannot reveal a gun regression) PLUS SpinBot
# (the Batch-1 spinner control, absent from the legacy roster) = 31+1 = 32.
# Every batch-1 opponent except none is covered; 17 opponents are added.
#
# Format: <bot dir or name> | <style group> | <note>. `style` is INFERRED from
# robots.json / the bot name, never decompiled, and is used ONLY to explain a
# result, never to decide one.
# ─────────────────────────────────────────────────────────────────────────────
/tmp/tr_bots/Aristocles | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/Ascendant | other | Batch-1 panel member
/tmp/tr_bots/BlitzBat | regular | Batch-1 panel member
/tmp/tr_bots/BrokenSword | other | added by Batch 2 (wider roster)
/tmp/tr_bots/CassiusClay | dodger | Batch-1 panel member
/tmp/tr_bots/Cigaret | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/CigaretBH | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/Coriantumr | other | Batch-1 panel member
/tmp/tr_bots/Diamond | dodger | Batch-1 panel member
/tmp/tr_bots/DiamondHawk | other | added by Batch 2 (wider roster)
/tmp/tr_bots/DiamondStealer | regular | Batch-1 panel member
/tmp/tr_bots/Dookious | dodger | Batch-1 panel member
/tmp/tr_bots/DrussGT | other | Batch-1 panel member
/tmp/tr_bots/FloodMini | regular | added by Batch 2 (wider roster)
/tmp/tr_bots/GresSuffurd | dodger | Batch-1 panel member
/tmp/tr_bots/HawkOnFire | regular | Batch-1 panel member
/tmp/tr_bots/Jen | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/Komarious | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/KurtWaveSurfer | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/LightningBug | other | added by Batch 2 (wider roster)
/tmp/tr_bots/LionWWSVMvoid | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/Lukious | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/PatternRobot | regular | added by Batch 2 (wider roster)
/tmp/tr_bots/Phoenix | other | added by Batch 2 (wider roster)
/tmp/tr_bots/RetroGirl | other | Batch-1 panel member
/tmp/tr_bots/RougeDC | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/Shadow | other | added by Batch 2 (wider roster)
/tmp/tr_bots/SpinBot | spinner | added by Batch 2 (wider roster)
/tmp/tr_bots/TripHammer | other | Batch-1 panel member
/tmp/tr_bots/WallAvoider | regular | Batch-1 panel member
/tmp/tr_bots/WaveSurferGF | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/WaveSurferPG | dodger | added by Batch 2 (wider roster)
/tmp/tr_bots/YersiniaPestis | other | Batch-1 panel member