closing j125: verify movement ship end-to-end (clean build SuccessX, full 16-test guard suite all counts match, real default+revert battles) and restructure gun/movement ledgers outcome-first

This commit is contained in:
2026-09-26 05:06:49 +02:00
parent cf77d0d647
commit 8e109bae0b
2 changed files with 125 additions and 0 deletions
+96
View File
@@ -10,6 +10,102 @@ the gun successor to `docs/movement_campaign.md`; read that first, then this.
---
## OUTCOME (read this first)
**The shipped gun is UNCHANGED: `onlyPattern`** — the rack admits `Pattern` only
(`common_libs/gun_harness/selector.nim`, commit `e0666a5`). **Nothing beat it.**
The gun design space explored here is **CLOSED on evidence**, not on belief.
**Headline (MEASURED).** Four frozen-binary tournament sessions on the frozen
**15- and 33-opponent panels**: **1,233 battles** (phase 1: 270 + 297; phase 2:
270 + 396), **0 failed, 0 never-started, 0 liveness exclusions**. Each session is
ONE binary built from `git archive HEAD`; arms differ only by env; movement is
pinned `TR_MOVEMENT=strafe` in every arm, so every delta is a pure gun delta.
With the earlier 16-gun selector campaign (≈660 battles) the gun ledger now
stands on **≈1,900 measured battles**. What was **detectably worse**: `knn`
(−39.3 dmg/run, 0/15 opponents; −0.62 wins/run, p=0.012) and, on the wider
33-opponent panel, `TMHorizon` (−10.4 dmg/run, sign p=0.035). What was
**not distinguishable** (inside the MDE, no verdict): `bitbrain`,
`tmhorizon`@15, `rack_pk`, `rack_pt`, `len6`, `len16`, `depth100`, the two radial
controls.
### CLOSED axes (one-line reason each)
| axis | reason it is closed |
|---|---|
| other single guns | `knn` detectably worse; `bitbrain` a wash measured 3× (−3.6 / +10.3 / +1.8 dmg/run, all inside their MDEs); `TMHorizon` worse on 33 (`docs/gauntlet_bitbrain_vs_pattern.md`, Batches 1–2) |
| 2-gun selector racks | `rack_pk` nominally below `pattern` on wins; `rack_pt` == `tmhorizon`-alone → the selector just picks the corrector (`## Batch 1`) |
| selector mechanics | the virtual-fitness selector is negative value at every rack size tested (16, lean8, lean6, pairs); it never manufactured a win (`docs/selector_negative_value.md`) |
| lead amplitude | a full gain sweep (1.0/1.5/2.0/3.0) makes Pattern strictly worse at every range band — the lever is lead *information*, not amplitude (`docs/bitbrain_campaign.md` §0.3.2) |
| Pattern's own match-length / history params | `len6`'s single-session +0.49 wins (p=0.039 on 15) reversed to a wash on 33 (−0.04, p=1); `len16`/`depth100` flat; two batches, no replication (`## Phase 2`) |
| radial knobs (`TR_PATTERN_RAD_*`) | bearing-invariant **by construction** (job j99 proved `bmPath` is a structural no-op and scale/offset cannot change the aim bearing); in Batch 2 `rad_offset` is nominally *negative* |
### THE ONE OPEN AXIS
**A genuinely new source of lead INFORMATION** — not another knob on
BitBrain/TMHorizon (both measured neutral→negative), not amplitude (dead), not
radial (bearing-invariant). The standing mechanism is `docs/bitbrain_campaign.md`
§0.3.3: at 450+ px Pattern's own lead correlation with the required lead is only
**0.165**, so it is adding variance to a nearly uninformative signal.
**Honest caveat (MEASURED, `docs/state_window_gate.md`).** The best surviving
idea from the state line — a *single* wave-relative state at Q=4 — **does carry
signal at coarse resolution**: it predicts the miss-offset *bin* on held-out
battles at 0.4094 (vs 0.2348 majority), but those bins are **36/42/60 px wide =
4.58/5.34/7.63° at 450 px**, and its implied hit probability is **still the base
rate (~0.09)** — i.e. it is informative about **which side** the miss falls on,
not about **whether** the shot hits. The live hit window is
`atan(18/450) = 2.29°` half-width; **the measured signal is at ~4.6–7.6°, not at
the ~2° the 450 px hit window needs.** A temporal window does not fix this (the
window gate is negative; long contexts never recur). So the open axis is a *new
information source that resolves the arrival offset to better than ~2°*, which
does not exist yet.
**Cost estimate for trying it (INFERRED).** Build and *offline-gate* the new
information source first — require it to beat Pattern's 0.165 lead correlation
and, ideally, predict the arrival offset to <2.29° on held-out shots (days of
work; the offline instruments exist, e.g. `common_libs/tests/state_window_gate.py`).
Only then spend a live batch: one frozen binary, 6 env-only arms × 15 opponents ×
3 runs × 3 rounds = **270 battles ≈ 1–2 h wall** at conc 6 with `--wait-arena`
(exact command in *How to run a batch*). Do **not** open with a live batch.
### What would change our mind
A named new information source that (1) offline predicts the arrival offset to
better than the **~2.29° hit half-window at 450 px** on held-out battles **and**
(2) in a live frozen-panel batch beats `pattern` on **wins/run** with the winning
metric's CI excluding 0 and sign-flip p<0.05 while **damage is not detectably
down**. A damage-only gain without wins is the `ring`-mover trap and does **not**
count.
---
## What changed tonight (2026-09-26)
* **SHIPPED:** nothing in the gun. The rack is still `onlyPattern`.
* **NOT shipped, and why:** every candidate (other gun, small rack, selector
mechanic, match-length/history parameter, radial knob) failed to beat `pattern`
beyond the MDE on the frozen panels. Phase 2's one signal (`len6`, +0.49
wins/run on 15 opponents) **did not replicate** on 33 opponents. The only
shipped change of the night is the **movement** default (`tfil -> strafe`, see
`docs/movement_campaign.md`); the gun was left untouched.
* **Revert:** no gun revert is needed (no gun default changed). Movement:
`TR_MOVEMENT=tfil`.
* **Reproduce the key gun evidence (one command + the analyzer)** — phase 2,
Batch 2, the session that killed `len6`:
```sh
TOURNAMENT_NIMCACHE=/tmp/nc_j123 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_gun_b4.txt \
--panel tools/ab/panel_gun_b2.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference pattern \
--outdir /tmp/ab/j123_b2
python3 tools/ab/tournament_analyze.py /tmp/ab/j123_b2 --reference pattern
```
---
## 0. The question
The shipped rack admits **`Pattern` only** (`onlyPattern`). That decision was
+29
View File
@@ -57,6 +57,35 @@ intact**; gate v2's pre-registration and results are at the bottom of this file.
---
## What changed tonight (2026-09-26)
* **SHIPPED:** the default 1v1 movement is now **`TR_MOVEMENT=strafe`** (flipped
from `tfil` at `ModularBot_garage/src/ModularBot.nim:118`, commit `3fd6db9`).
Gate v2 on **300 fresh battles**: Δwins/run **+0.30**, 95% CI **[+0.02, +0.58]**,
sign-flip permutation **p = 0.04517**; cost **−10.97 dmg/run** (CI [−19.87,
−2.06]) — *survive far more for slightly less output*, net-positive on the
server score (+0.30 wins × 50 survival − 11 damage ≈ +4/run).
* **NOT shipped, and why:** every other arm on the frozen panel failed to beat
`strafe` beyond the MDE (`field_strong` / `field_off` were detectably
**worse**); nothing was promoted without replication. Gate **v1** was NOT
reinterpreted — it failed its required sign-test leg and the default was *not*
flipped until gate v2 passed on genuinely fresh data.
* **Revert:** `TR_MOVEMENT=tfil` (env only, no rebuild; the bot reports
`TR_MOVEMENT = tfil (source: env)`).
* **Reproduce the key evidence (one command + the analyzer):**
```sh
TOURNAMENT_NIMCACHE=/tmp/nc_j122 \
tools/ab/tournament_run.sh \
--arms tools/ab/arms_movement_v2.txt \
--panel tools/ab/panel_movement.txt \
--runs 10 --rounds 3 --conc 6 --wait-arena 45 \
--reference tfil \
--outdir /tmp/ab/j122_v2
python3 tools/ab/tournament_analyze.py /tmp/ab/j122_v2 --reference tfil
```
---
## Final confirmation + SHIP
> **Provenance.** The ship criterion below was pre-registered and committed in