j159: live 2-arm A/B result — geometry draw weighting costs 8.8 damage/run, wins null; keep TR_TFIL_GEO_MODE=off

This commit is contained in:
2026-09-27 12:44:38 +02:00
parent 7c5bc9ccbd
commit 4a1f3e1c88
+84
View File
@@ -123,3 +123,87 @@ A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need
## MEASURED ## MEASURED
*(appended after the battles — everything above was committed first)* *(appended after the battles — everything above was committed first)*
### MEASURED — the live A/B, 420 battles (j159)
* **Provenance.** Session `/tmp/ab/j159_geo`, commit `7c5bc9c`, frozen binary
sha256 `9f116e7a9eb9…`, panel `tools/ab/panel_movement.txt` (FROZEN, 15
opponents), 15 x 2 x **14 runs** x 3 rounds = **420 battles, 0 failed, 0 never
started, 1110 s**. Per-arm env files `/tmp/j159_geo/env/{A_off,B_geo}.env`.
* **Env verification.** All **420** runs carry their intended arm: 210/210
`A_off` show `TR_TFIL_GEO_MODE=off` / `TR_TFIL_GEO_TAU=0` (parsed `off`/`0.0`),
210/210 `B_geo` show `both-rej`/`60` (parsed `both`/`60.0`), every run reports
`env file: /tmp/j159_geo/env/<arm>.env (source: TR_ENV_FILE)` and
`move.effective = tfil`. No `.env` exists anywhere in the session dir, and the
loader's fallbacks are `./.env` then `.env` next to the executable — it does
not walk up parents — so the owner's file is unreachable. **0 runs mis-set.**
* **Record correction (no battle re-run).** The arms file declared only
`TR_ENV_FILE`, which the analyzer's liveness guard reads from `session.json`
and treats as an undeclared `TR_MOVEMENT` (fatal contamination). `session.json`
and a corrected arms file were rewritten to declare the effective env — the
original is kept as `/tmp/j159_geo/{session.json.orig,arms_tfil_geo.txt.orig}`.
The guard then re-verified all 420 declared values verbatim in the boot
reports: `liveness: 0 run(s) excluded (420 total)`.
**Pooled dashboard (descriptive, NOT the verdict):**
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `A_off` | 210 | 112.5 | 195.2 | 1.21 | 255/630 | 40.5% | 17.64% | 393 |
| `B_geo` | 210 | 103.7 | 188.0 | 1.16 | 244/630 | 38.7% | 17.42% | 419 |
**Verdict layer** (per-opponent paired deltas, arm − reference; sign-flip is
the exact 2^15 permutation the pre-registration names as the decision test):
| arm | metric | mean Δ | 95% CI | sign test | p(sign) | **p(sign-flip)** | Wilcoxon p | MDE |
|---|---|---:|---|---:|---:|---:|---:|---:|
| `B_geo` | **damage/run** | **-8.83** | [-14.69, -2.97] | 5/15 | 0.3018 | **0.006104** | 0.0115 | 7.65 |
| `B_geo` | round wins | -0.05 | [-0.18, +0.08] | 4/11 | 0.5488 | 0.4619 | 0.3496 | 0.17 |
| `B_geo` | damage taken | -7.24 | [-20.84, +6.36] | 8/15 | 1 | 0.2786 | 0.4777 | 17.77 |
| `B_geo` | incoming hit rate | -1.19 pp | [-3.88, +1.51] | 8/15 | 1 | 0.3962 | 0.5895 | 3.51 |
| `B_geo` | mean distance | **+26.3 px** | [+16.3, +36.4] | **15/15** | 6.1e-05 | 6.1e-05 | 0.0007 | 13.15 |
**Per-opponent damage/run** (the pattern/ram rows carry the loss): Coriantumr
-25.1, CassiusClay -24.2, SpinBot -26.2, WallAvoider -13.3, BlitzBat -11.2,
HawkOnFire -11.7, Diamond -10.9, TripHammer -9.0, YersiniaPestis -8.9,
Dookious -7.3 vs GresSuffurd +2.3, DiamondStealer +3.2, Ascendant +3.7,
DrussGT +0.1. Round wins: HawkOnFire +0.50 and WallAvoider +0.14 (both closer
opponents) against Coriantumr -0.57, TripHammer -0.21, YersiniaPestis -0.21.
**Mechanism, offline (the committed ruler, re-run unchanged for this doc):**
`measure_tfil_pick_defects` reproduces the pre-registered prediction exactly —
REACH **4.5% -> 29.4%**, hot-on-arrival **31.0% -> 24.0%**, and the diversity cost
**top-1 tile share 7.0% -> 10.6%** (normalised entropy 0.87 -> 0.85). The live
mean distance **+26 px on 15/15 opponents** is the same mechanism seen end to
end: the geometry weight prefers tiles that are far better to *arrive* in, and
the bot sits further out and deals **less** damage.
### VERDICT — DO NOT ADOPT
1. **H1 is rejected.** Round wins are flat (-0.05/run, p(sign-flip) = 0.46, under
the 0.17 MDE) and damage/run is **down 8.83** (p(sign-flip) = 0.0061, above
the 7.65 MDE, Wilcoxon p = 0.011). The pre-registration required BOTH
primaries up; one is down and significant by the test it named.
2. **The MDE, restated.** 14 runs/arm on the frozen 15-opponent panel resolves
**0.17 wins/run** and **7.65 damage/run** (better than the 0.28 pre-registered
estimate). So this run excludes a large *benefit*; it also positively measures
a small *harm* in damage. It says nothing about effects below those numbers.
3. **The mechanism moved exactly as predicted, and that is what makes it bad.**
REACH/arrival 4.5% -> 29.4% and hot-on-arrival 31.0% -> 24.0% are real and
reproducible, but they bought **+26 px of distance** and fewer damage points,
not survival: incoming hit rate moved -1.19 pp, a fifth of its own 3.51 MDE.
"Arrive at a safe tile" turned out to mean "arrive further away".
4. **Diversity worsened, as the offline sweep warned.** Top-1 tile share
7.0% -> 10.6%, and the live losses concentrate against the opponents that
punish a long-range mover (SpinBot -26.2 damage with a -14.5 pp hit-rate
shift, the pattern guns -9 to -25). Outcomes are null-to-negative AND
diversity is worse: that is the argument against shipping, exactly the
pre-registered case.
5. **A null on wins lets us claim only "no large win".** It does not show the
knob is inert, and it does not retract the geometric measurement — but the
geometry measurement is not an argument for shipping when the live
consequence of it is measurably less damage from measurably further away.
**Ship state: `TR_TFIL_GEO_MODE` stays default `off`. Nothing changes.** Not
adopted, not adopted default-off. This is the **fifth** mechanism-positive /
outcome-not-positive result in the movement campaign (j144, j145, j146, j147,
j159) — and the first one where the mechanism is *anti*-correlated with damage.