# j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win? **Written and committed BEFORE a single battle of this experiment ran.** No result in the "MEASURED" section below existed when this section was written. ## The question `TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` (shipped default `off`, j152) shape the **draw** over the safe tiles the tfil picker chooses from. The offline sweep on recorded fixtures predicted a real geometric improvement — | offline metric (`measure_tfil_pick_defects`) | `off` | `both-rej`, tau 60 | |---|---:|---:| | REACH / arrival within feasible time | 4.5% | 29.4% | | hot on arrival (`hotAtTta`) | 31.0% | 24.0% | | **top-1 tile share (pick DIVERSITY)** | **7.0%** | **10.6%** | — and a **diversity cost**, because the geometry weight concentrates the draw. The offline ruler replays the real picker on real fixtures; it says nothing about whether a tile that is geometrically reachable and less hot actually wins a round. That is what this run measures. **Hypothesis H1.** Weighting the draw by tile geometry (`both-rej`, tau 60) raises damage/run and round-win rate over the shipped uniform draw, because the mover arrives at safe tiles instead of merely picking them. **Direction is pre-registered as two-sided.** A regression is as interesting as a win (the diversity cost makes one plausible) and re-deciding the direction after seeing the data is exactly what this document exists to prevent. ## Arms — identical except the geo knob Both arms: `TR_MOVEMENT=tfil`, everything else at the shipped defaults, one frozen binary built once from `git archive HEAD` (session 2223ca6). | arm | env | role | |---|---|---| | `A_off` | `TR_TFIL_GEO_MODE=off` (`TAU=0`) | **REFERENCE** — the shipped uniform draw | | `B_geo` | `TR_TFIL_GEO_MODE=both-rej` `TR_TFIL_GEO_TAU=60` | treatment | **Contamination control (the main risk).** The owner has `both-rej` in a personal `.env` (`ModularBot_garage/out/.env`, a copy at `tr_bots/ModularBot_geo/.env`), and this bot's dotenv loader gives the FILE priority over shell exports. If the tournament's bot instances resolved that file, arm A would silently become arm B and the whole run would be void. Three guarantees, all verifiable from the logs: 1. each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this job generated** (`/tmp/j159_geo/env/A_off.env`, `/tmp/j159_geo/env/B_geo.env`) in this job's own outdir, so the ONLY `.env` the loader can resolve is mine; 2. the tournament's per-run botdir (`$OUTDIR/.work///run/bots/ModularBot`) contains only `ModularBot.json`, `ModularBot.sh` and a symlink to the frozen binary — **no `.env`**, and the loader's fallback is `./.env` then `.env` next to the executable (it does not walk up parent directories), so the owner's file is not reachable; 3. every single run's `[env]` boot report is checked for its intended `TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` before any number is read. **Any run whose `[env]` disagrees with its arm invalidates the session** and the run is reported as void rather than analysed. ## Design * Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` (unmodified). * Panel: `tools/ab/panel_movement.txt` — the **FROZEN 15-opponent movement panel**, unchanged. Unit of evidence is the opponent, not the battle. * 15 opponents x 2 arms x **14 runs** x 3 rounds = **420 battles**. * Battles serialised: `--wait-arena 45`, one session at a time. ### Primary metrics (pre-registered, fixed) 1. **damage/run** (our damage dealt per run) 2. **round-win rate** (rounds won / rounds fought) **Hit rate is NOT a primary metric** — it hid a survival regression once already. ### Secondary / mechanism (reported, never a verdict) * incoming hit rate (the survival channel the mechanism actually runs through); * damage taken/run; * the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile share). The live battle logs do not contain the per-pick tile or the arrival state, so the live run **cannot** re-measure them; that is stated in the verdict rather than papered over. No new instrumentation is built for this. ### Statistical treatment Same as every previous movement gate: per-opponent paired deltas (arm − reference), mean delta, SD, SE, 95% CI, a sign test and a **sign-flip permutation test** (exact when `2^n <= 2^20`, else Monte-Carlo), Wilcoxon as a cross-check, plus the MDE the analyzer reports for the reference arm's n. ### MDE — stated up front, and it is LARGE At **14 runs/arm** over the frozen 15-opponent panel this design resolves about **0.28 wins/run** (and the corresponding damage/run MDE the analyzer prints). A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need ~2.2 h and ~2,900 battles — **which we are NOT doing.** **Consequences, recorded before any data:** * **A null is the likely outcome.** This would be the *fifth* consecutive mechanism-positive / outcome-null result in this campaign (after j144, j145, j146, j147). * A null here **excludes only a LARGE effect** (>= ~0.28 wins/run). It does not show the knob does nothing, and it does not retract the offline geometric measurement. * Because the offline sweep also measured a **diversity regression** (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity cost is an argument **against** shipping, not for it. ### Verdict rule (fixed now, not re-read later) * **Adopt** only if BOTH primaries move in B's favour with `p(sign-flip) < 0.05` and the effect is at or above the reported MDE. One primary at p<0.05 with the other not down is reported as a partial signal, not a win. * Otherwise **do not ship**; the knob stays default `off`. * The mechanism is reported as measured, with no vote in the verdict. * No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result. --- ## MEASURED *(appended after the battles — everything above was committed first)*