gun j121 Batch 2 outcome: the wider panel kills the damage hint (bitbrain wash, tmhorizon WORSE); onlyPattern confirmed; phase closed

This commit is contained in:
2026-09-26 03:54:52 +02:00
parent c343c00aaf
commit 2a5ea5c17b
+139 -7
View File
@@ -48,6 +48,14 @@ evidence is the **number of opponents**, not the number of runs.
| Liveness | every declared env token must appear verbatim in OUR bot's own raw-env report, else the run is excluded and named |
| Never shipped | this is a measurement campaign: `git status` clean, shipped defaults untouched, `.gitignore` untouched |
> **Movement-default flip note (MEASURED).** During this job a parallel job
> (j120, commit `3fd6db9`) flipped the shipped `TR_MOVEMENT` default from `tfil`
> to `strafe`. Batch 1 built before the flip, Batch 2 after. **Every arm in both
> batches declares `TR_MOVEMENT=strafe`**, so the effective movement is identical
> and the two batches are directly comparable — the pin did its job. The reference
> rack lines in the boot reports read `PATTERN` for `pattern`, and the active 1v1
> rack reads `BITBRAIN` / `TMHORIZON` / `KNN` for the overridden arms.
### Movement is PINNED in every arm: `TR_MOVEMENT=strafe`
**Reason (hard rule).** A movement-default flip is landing from job j120 during
@@ -283,22 +291,141 @@ rounds = 297 battles**, `--conc 6`, `--wait-arena`. Arm file:
### Outcome — Batch 2
*(filled in by the Batch-2 results commit)*
**MEASURED.** Session `/tmp/ab/j121_g2`, commit
`c343c00aaf800bc05ce9bf1b277054c11f5f577c`, frozen binary sha256 `d267ab78f6ce…`,
**33 opponents × 3 arms × 3 runs × 3 rounds = 297 battles, 0 failed, 0 never
started, 0 liveness exclusions**.
**A movement-default flip landed between the two batches** (commit `3fd6db9`,
`TR_MOVEMENT` default `tfil` → `strafe`; the only source change). Because
**every** arm in **both** batches pins `TR_MOVEMENT=strafe`, the effective
movement is identical in Batch 1 and Batch 2 — the batches are directly
comparable and the gun comparison is unconfounded. This is exactly what the pin
was for.
> **DIRECT ANSWER (MEASURED, n=33 opponents): the Batch-1 damage hint did NOT
> survive.** `bitbrain` collapses to **+1.8 dmg/run** (19/33 opponents, sign test
> p=0.49) with **+0.11 wins/run** — a wash. `tmhorizon` **reverses sign and is
> detectably WORSE on damage**: **−10.4 dmg/run** (10/33 opponents, sign test
> p=0.035, sign-flip p=0.0047, 95% CI [−17.2, −3.6]) at unchanged wins (−0.03).
> The only open signal in Batch 1 was a **small-panel fluctuation**, and the
> wider panel resolves it **against** the correctors. Combined with Batch 1:
> **nothing beats the shipped `Pattern` on damage/run AND round wins**;
> `onlyPattern` is CONFIRMED on 15 opponents and again on 33.
#### Pooled dashboard (all valid runs — explanation only, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `pattern` (REF) | 99 | 131.2 | 126.4 | 1.90 | 188/297 | 63.3% | 12.93% | 417 |
| `bitbrain` | 99 | 133.0 | 126.7 | **2.01** | **199/297** | **67.0%** | 12.91% | 417 |
| `tmhorizon` | 99 | 120.8 | 133.2 | 1.87 | 185/297 | 62.3% | 13.04% | 422 |
#### Cross-opponent aggregation (the verdict layer, damage + wins)
| arm | metric | mean Δ | spread (SD) | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---:|---:|---|---:|---:|---:|---:|---:|
| `bitbrain` | damage | +1.80 | 18.95 | [−4.67, +8.26] | 19/33 | 0.4869 | 0.5914 | 0.5675 | 9.24 |
| `bitbrain` | wins | +0.11 | 0.48 | [−0.05, +0.28] | 12/19 | 0.3593 | 0.2436 | 0.5716 | 0.24 |
| `tmhorizon` | damage | **−10.40** | 20.01 | **[−17.22, −3.57]** | 10/33 | **0.03508** | **0.00472** | **0.006975** | 9.76 |
| `tmhorizon` | wins | −0.03 | 0.59 | [−0.23, +0.17] | 9/19 | 1.0 | 0.8459 | 0.4804 | 0.29 |
#### The pre-registered verdict (verbatim)
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---:|---|---:|---:|---|---|---|---|
| 1 | `bitbrain` | +0.11 | +1.8 | 12/19 p=0.3593 | 19/33 p=0.4869 | **not distinguishable** | **not distinguishable** |
| 2 | `tmhorizon` | −0.03 | −10.4 | 9/19 p=1 | 10/33 p=0.03508 | **WORSE** | **WORSE** |
Reference `pattern`: 131.2 dmg/run, 1.90 wins/run, 12.93% incoming, 417 px.
#### Reading
* **BitBrain is a wash — now measured three times.** vs the shipped Pattern:
−3.6 dmg/run on the 32-opponent legacy gauntlet (with `tfil` movement,
`docs/gauntlet_bitbrain_vs_pattern.md`), **+10.3** on the 15-opponent Batch-1
panel, **+1.8** here on 33 opponents. All three sit inside their own MDEs; the
pooled evidence is a zero. The Batch-1 `+10` was a 15-opponent fluctuation.
* **TMHorizon is the one arm the wider panel separates — and it loses.** Its
Batch-1 sign (+8.7) does not merely fail to replicate, it **inverts**
(−10.4, 10/33, p=0.035). The damage loss is broad (24/33 opponents negative,
worst `Aristocles` −61, `Jen` −54, `WallAvoider` −41), so it is not one
outlier. Since TMHorizon was already predicted to lose by its own design docs,
this is the first *confirming* live panel measurement of that prediction.
* **Wins are flat for both arms** (+0.11 / −0.03, MDEs 0.24 / 0.29): the
correctors change damage at most, never survival. In this harness round wins
are survival wins, so neither is a win.
* **The corrector family is closed** as a source of a `pattern`-beating gun at
the resolutions tested.
#### Pre-registered predictions — scorecard
| # | prediction | outcome |
|---|---|---|
| 1 | the +10 dmg hint does not survive the wider panel; deltas fall inside the new MDE; sign test non-significant | **CORRECT** (`bitbrain` +1.8, p=0.49; `tmhorizon` even reversed to −10.4, p=0.035) |
| 2 | wins remain flat | **CORRECT** (+0.11 / −0.03) |
| 3 | if the hint survived at p<0.05 with wins not down, it is the first gun to beat `pattern` | **not invoked** |
**Final direct answer for the phase (MEASURED).** Across two frozen panels
(15 and 33 opponents), a single frozen-binary-per-session design, 567 battles,
seven distinct gun configurations, and movement pinned identically in every arm:
**no rack gun and no small rack beats the shipped `Pattern` on damage/run AND
round wins.** Every candidate is either not distinguishable (`bitbrain`,
`tmhorizon`, `rack_pk`, `rack_pt` at n=15), a detectable loss
(`knn`; `tmhorizon` at n=33), or a sub-MDE damage-only wobble
(`bitbrain`). The shipped **`onlyPattern` rack is CONFIRMED**, not merely
assumed — the gun axis's first phase ends in a measured optimum of the tested
design space.
---
## What to try next
## What to try next (rewritten AFTER Batches 1–2 — recommendations, not results)
*(Rewritten after Batch 1 — these are recommendations, not results. Filled in the
results commit below.)*
**The primary question is answered:** nothing beats the shipped `pattern` across
the frozen panel, and `onlyPattern` is CONFIRMED. Ranked by value per battle:
1. **The damage-side hint is now CLOSED.** Batch 2 (33 opponents) killed it:
`bitbrain` is a wash (+1.8 dmg/run, p=0.49) and `tmhorizon` is detectably
**worse** (−10.4 dmg/run, p=0.035). Do **not** spend another batch on the
Pattern-lead correctors; three independent measurements of `bitbrain` vs
Pattern now average ~0 (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs).
2. **Lead INFORMATION remains the only named gun axis** (given: amplitude is
dead). But the two live implementations of it (BitBrain's gain, TMHorizon's
shift) are both measured negative/neutral across panels. A future axis would
need a *different* information source, not another knob on these two — e.g. a
new base prediction, or a gun-agnostic ensemble. Do not re-open the correctors.
3. **The selector is confirmed negative at a 2-gun rack** (`rack_pt` ==
`tmhorizon`; `rack_pk` below `pattern` on wins). Do not re-open it with a
larger rack: the 16-gun negative already stands and Batch 1 shows the
mechanism does not manufacture a win.
4. **Drop `knn`** (detectably worse on both primaries, 0/15 on damage).
5. **Never chase damage without wins.** This harness's round wins are survival
wins, so the gun's job is to *kill*; a damage gain that does not raise the win
rate is the `ring`-mover trap in a new costume (and `tmhorizon`'s n=33 result
is a *negative* damage signal with flat wins — also not a win).
6. **Spinner-specific claim (owner's) remains UNTESTED.** The legacy roster has
**no constant-turn spinner**; SpinBot is a periodic circle-mover and is in
both panels (Batch 1: `bitbrain` +25.3 dmg / +0 wins, `tmhorizon` +15.2 / +0
wins; on the 33-panel SpinBot's deltas are in the per-opponent tables).
Building a true constant-turn spinner opponent is a nearly-free follow-up and
is the one part of the owner's claim this campaign has not touched.
## What would make us stop
**PHASE STATUS (updated after Batch 2): STOPPED — the condition fired.** Across
Batches 1–2 no arm beats `pattern` beyond the MDE, and no arm shows a
`>= +MDE` damage gain with p<0.10 (`bitbrain` +1.8 vs MDE 9.24; `tmhorizon`
actually negative). The gun stage's first phase is therefore **closed** with
*"the shipped `onlyPattern` rack is the best gun configuration we have measured
across the panels (15 and 33 opponents)"* — a successful outcome, not a failure.
Further gun work must open a **named new axis** (see *What to try next*), not
another arm on the same axis.
* **Stop the first phase** once a batch's best arm cannot beat `pattern` beyond
the MDE, or when a gun arm's damage gain is bought with a detectable win loss.
At that point *"`onlyPattern` is the measured optimum of this rack"* is the
conclusion, not a failure (§2 rule 4).
conclusion, not a failure (§2 rule 4). **This fired after Batch 2.**
* **Stop a single batch early** only for a contract violation (arena not free,
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
* **Do not invent more arms on the same axis** once two consecutive batches fail
@@ -310,7 +437,7 @@ results commit below.)*
## How to run a batch (exact commands)
```sh
# 0. wait for the arena (j120's movement run may still be fighting)
# 0. wait for the arena (other campaign jobs may be fighting)
TOURNAMENT_NIMCACHE=/tmp/nc_j121 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b1.txt \
@@ -319,6 +446,10 @@ tools/ab/tournament_run.sh \
--reference pattern \
--outdir /tmp/ab/j121_g1
# Batch 2 (the wider-panel calibration) was the same command with
# --arms tools/ab/arms_gun_b2.txt --panel tools/ab/panel_gun_b2.txt \
# --outdir /tmp/ab/j121_g2 --wait-arena 50
# 1. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j121_g1 --reference pattern
```
@@ -330,4 +461,5 @@ different reference is a free pairwise comparison with no battles.
| session | commit | battles | arms | verdict |
|---|---|---:|---|---|
| `/tmp/ab/j121_g1` | *(recorded at run time)* | 270 | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | *(Batch 1 outcome below)* |
| `/tmp/ab/j121_g1` | `1d8143a` | 270 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon, knn, rack_pk, rack_pt | **nothing beats the shipped `pattern`**; all arms not distinguishable except `knn` = WORSE; `onlyPattern` CONFIRMED |
| `/tmp/ab/j121_g2` | `c343c00` | 297 (0 failed, 0 never started, 0 excluded) | pattern, bitbrain, tmhorizon | **the Batch-1 damage hint does not survive**: `bitbrain` +1.8 dmg/run p=0.49 (wash); `tmhorizon` −10.4 dmg/run p=0.035 (WORSE); wins flat |