# j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win? **Written and committed BEFORE a single battle of this experiment ran.** No result in the "MEASURED" section below existed when this section was written. ## The question `TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` (shipped default `off`, j152) shape the **draw** over the safe tiles the tfil picker chooses from. The offline sweep on recorded fixtures predicted a real geometric improvement — | offline metric (`measure_tfil_pick_defects`) | `off` | `both-rej`, tau 60 | |---|---:|---:| | REACH / arrival within feasible time | 4.5% | 29.4% | | hot on arrival (`hotAtTta`) | 31.0% | 24.0% | | **top-1 tile share (pick DIVERSITY)** | **7.0%** | **10.6%** | — and a **diversity cost**, because the geometry weight concentrates the draw. The offline ruler replays the real picker on real fixtures; it says nothing about whether a tile that is geometrically reachable and less hot actually wins a round. That is what this run measures. **Hypothesis H1.** Weighting the draw by tile geometry (`both-rej`, tau 60) raises damage/run and round-win rate over the shipped uniform draw, because the mover arrives at safe tiles instead of merely picking them. **Direction is pre-registered as two-sided.** A regression is as interesting as a win (the diversity cost makes one plausible) and re-deciding the direction after seeing the data is exactly what this document exists to prevent. ## Arms — identical except the geo knob Both arms: `TR_MOVEMENT=tfil`, everything else at the shipped defaults, one frozen binary built once from `git archive HEAD` (session 2223ca6). | arm | env | role | |---|---|---| | `A_off` | `TR_TFIL_GEO_MODE=off` (`TAU=0`) | **REFERENCE** — the shipped uniform draw | | `B_geo` | `TR_TFIL_GEO_MODE=both-rej` `TR_TFIL_GEO_TAU=60` | treatment | **Contamination control (the main risk).** The owner has `both-rej` in a personal `.env` (`ModularBot_garage/out/.env`, a copy at `tr_bots/ModularBot_geo/.env`), and this bot's dotenv loader gives the FILE priority over shell exports. If the tournament's bot instances resolved that file, arm A would silently become arm B and the whole run would be void. Three guarantees, all verifiable from the logs: 1. each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this job generated** (`/tmp/j159_geo/env/A_off.env`, `/tmp/j159_geo/env/B_geo.env`) in this job's own outdir, so the ONLY `.env` the loader can resolve is mine; 2. the tournament's per-run botdir (`$OUTDIR/.work///run/bots/ModularBot`) contains only `ModularBot.json`, `ModularBot.sh` and a symlink to the frozen binary — **no `.env`**, and the loader's fallback is `./.env` then `.env` next to the executable (it does not walk up parent directories), so the owner's file is not reachable; 3. every single run's `[env]` boot report is checked for its intended `TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` before any number is read. **Any run whose `[env]` disagrees with its arm invalidates the session** and the run is reported as void rather than analysed. ## Design * Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` (unmodified). * Panel: `tools/ab/panel_movement.txt` — the **FROZEN 15-opponent movement panel**, unchanged. Unit of evidence is the opponent, not the battle. * 15 opponents x 2 arms x **14 runs** x 3 rounds = **420 battles**. * Battles serialised: `--wait-arena 45`, one session at a time. ### Primary metrics (pre-registered, fixed) 1. **damage/run** (our damage dealt per run) 2. **round-win rate** (rounds won / rounds fought) **Hit rate is NOT a primary metric** — it hid a survival regression once already. ### Secondary / mechanism (reported, never a verdict) * incoming hit rate (the survival channel the mechanism actually runs through); * damage taken/run; * the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile share). The live battle logs do not contain the per-pick tile or the arrival state, so the live run **cannot** re-measure them; that is stated in the verdict rather than papered over. No new instrumentation is built for this. ### Statistical treatment Same as every previous movement gate: per-opponent paired deltas (arm − reference), mean delta, SD, SE, 95% CI, a sign test and a **sign-flip permutation test** (exact when `2^n <= 2^20`, else Monte-Carlo), Wilcoxon as a cross-check, plus the MDE the analyzer reports for the reference arm's n. ### MDE — stated up front, and it is LARGE At **14 runs/arm** over the frozen 15-opponent panel this design resolves about **0.28 wins/run** (and the corresponding damage/run MDE the analyzer prints). A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need ~2.2 h and ~2,900 battles — **which we are NOT doing.** **Consequences, recorded before any data:** * **A null is the likely outcome.** This would be the *fifth* consecutive mechanism-positive / outcome-null result in this campaign (after j144, j145, j146, j147). * A null here **excludes only a LARGE effect** (>= ~0.28 wins/run). It does not show the knob does nothing, and it does not retract the offline geometric measurement. * Because the offline sweep also measured a **diversity regression** (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity cost is an argument **against** shipping, not for it. ### Verdict rule (fixed now, not re-read later) * **Adopt** only if BOTH primaries move in B's favour with `p(sign-flip) < 0.05` and the effect is at or above the reported MDE. One primary at p<0.05 with the other not down is reported as a partial signal, not a win. * Otherwise **do not ship**; the knob stays default `off`. * The mechanism is reported as measured, with no vote in the verdict. * No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result. --- ## MEASURED *(appended after the battles — everything above was committed first)* ### MEASURED — the live A/B, 420 battles (j159) * **Provenance.** Session `/tmp/ab/j159_geo`, commit `7c5bc9c`, frozen binary sha256 `9f116e7a9eb9…`, panel `tools/ab/panel_movement.txt` (FROZEN, 15 opponents), 15 x 2 x **14 runs** x 3 rounds = **420 battles, 0 failed, 0 never started, 1110 s**. Per-arm env files `/tmp/j159_geo/env/{A_off,B_geo}.env`. * **Env verification.** All **420** runs carry their intended arm: 210/210 `A_off` show `TR_TFIL_GEO_MODE=off` / `TR_TFIL_GEO_TAU=0` (parsed `off`/`0.0`), 210/210 `B_geo` show `both-rej`/`60` (parsed `both`/`60.0`), every run reports `env file: /tmp/j159_geo/env/.env (source: TR_ENV_FILE)` and `move.effective = tfil`. No `.env` exists anywhere in the session dir, and the loader's fallbacks are `./.env` then `.env` next to the executable — it does not walk up parents — so the owner's file is unreachable. **0 runs mis-set.** * **Record correction (no battle re-run).** The arms file declared only `TR_ENV_FILE`, which the analyzer's liveness guard reads from `session.json` and treats as an undeclared `TR_MOVEMENT` (fatal contamination). `session.json` and a corrected arms file were rewritten to declare the effective env — the original is kept as `/tmp/j159_geo/{session.json.orig,arms_tfil_geo.txt.orig}`. The guard then re-verified all 420 declared values verbatim in the boot reports: `liveness: 0 run(s) excluded (420 total)`. **Pooled dashboard (descriptive, NOT the verdict):** | arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance | |---|---:|---:|---:|---:|---:|---:|---:|---:| | `A_off` | 210 | 112.5 | 195.2 | 1.21 | 255/630 | 40.5% | 17.64% | 393 | | `B_geo` | 210 | 103.7 | 188.0 | 1.16 | 244/630 | 38.7% | 17.42% | 419 | **Verdict layer** (per-opponent paired deltas, arm − reference; sign-flip is the exact 2^15 permutation the pre-registration names as the decision test): | arm | metric | mean Δ | 95% CI | sign test | p(sign) | **p(sign-flip)** | Wilcoxon p | MDE | |---|---|---:|---|---:|---:|---:|---:|---:| | `B_geo` | **damage/run** | **-8.83** | [-14.69, -2.97] | 5/15 | 0.3018 | **0.006104** | 0.0115 | 7.65 | | `B_geo` | round wins | -0.05 | [-0.18, +0.08] | 4/11 | 0.5488 | 0.4619 | 0.3496 | 0.17 | | `B_geo` | damage taken | -7.24 | [-20.84, +6.36] | 8/15 | 1 | 0.2786 | 0.4777 | 17.77 | | `B_geo` | incoming hit rate | -1.19 pp | [-3.88, +1.51] | 8/15 | 1 | 0.3962 | 0.5895 | 3.51 | | `B_geo` | mean distance | **+26.3 px** | [+16.3, +36.4] | **15/15** | 6.1e-05 | 6.1e-05 | 0.0007 | 13.15 | **Per-opponent damage/run** (the pattern/ram rows carry the loss): Coriantumr -25.1, CassiusClay -24.2, SpinBot -26.2, WallAvoider -13.3, BlitzBat -11.2, HawkOnFire -11.7, Diamond -10.9, TripHammer -9.0, YersiniaPestis -8.9, Dookious -7.3 vs GresSuffurd +2.3, DiamondStealer +3.2, Ascendant +3.7, DrussGT +0.1. Round wins: HawkOnFire +0.50 and WallAvoider +0.14 (both closer opponents) against Coriantumr -0.57, TripHammer -0.21, YersiniaPestis -0.21. **Mechanism, offline (the committed ruler, re-run unchanged for this doc):** `measure_tfil_pick_defects` reproduces the pre-registered prediction exactly — REACH **4.5% -> 29.4%**, hot-on-arrival **31.0% -> 24.0%**, and the diversity cost **top-1 tile share 7.0% -> 10.6%** (normalised entropy 0.87 -> 0.85). The live mean distance **+26 px on 15/15 opponents** is the same mechanism seen end to end: the geometry weight prefers tiles that are far better to *arrive* in, and the bot sits further out and deals **less** damage. ### VERDICT — DO NOT ADOPT 1. **H1 is rejected.** Round wins are flat (-0.05/run, p(sign-flip) = 0.46, under the 0.17 MDE) and damage/run is **down 8.83** (p(sign-flip) = 0.0061, above the 7.65 MDE, Wilcoxon p = 0.011). The pre-registration required BOTH primaries up; one is down and significant by the test it named. 2. **The MDE, restated.** 14 runs/arm on the frozen 15-opponent panel resolves **0.17 wins/run** and **7.65 damage/run** (better than the 0.28 pre-registered estimate). So this run excludes a large *benefit*; it also positively measures a small *harm* in damage. It says nothing about effects below those numbers. 3. **The mechanism moved exactly as predicted, and that is what makes it bad.** REACH/arrival 4.5% -> 29.4% and hot-on-arrival 31.0% -> 24.0% are real and reproducible, but they bought **+26 px of distance** and fewer damage points, not survival: incoming hit rate moved -1.19 pp, a fifth of its own 3.51 MDE. "Arrive at a safe tile" turned out to mean "arrive further away". 4. **Diversity worsened, as the offline sweep warned.** Top-1 tile share 7.0% -> 10.6%, and the live losses concentrate against the opponents that punish a long-range mover (SpinBot -26.2 damage with a -14.5 pp hit-rate shift, the pattern guns -9 to -25). Outcomes are null-to-negative AND diversity is worse: that is the argument against shipping, exactly the pre-registered case. 5. **A null on wins lets us claim only "no large win".** It does not show the knob is inert, and it does not retract the geometric measurement — but the geometry measurement is not an argument for shipping when the live consequence of it is measurably less damage from measurably further away. **Ship state: `TR_TFIL_GEO_MODE` stays default `off`. Nothing changes.** Not adopted, not adopted default-off. This is the **fifth** mechanism-positive / outcome-not-positive result in the movement campaign (j144, j145, j146, j147, j159) — and the first one where the mechanism is *anti*-correlated with damage.