From 7c5bc9ccbd29021ff100890e5b049cfeb273831b Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sun, 27 Sep 2026 12:16:36 +0200 Subject: [PATCH] j159: PRE-REGISTER the tfil geo A/B (geo draw weighting vs uniform) before any battle --- docs/tfil_geo_ab.md | 125 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 125 insertions(+) create mode 100644 docs/tfil_geo_ab.md diff --git a/docs/tfil_geo_ab.md b/docs/tfil_geo_ab.md new file mode 100644 index 0000000..934cf1b --- /dev/null +++ b/docs/tfil_geo_ab.md @@ -0,0 +1,125 @@ +# j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win? + +**Written and committed BEFORE a single battle of this experiment ran.** No +result in the "MEASURED" section below existed when this section was written. + +## The question + +`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` (shipped default `off`, j152) shape the +**draw** over the safe tiles the tfil picker chooses from. The offline sweep on +recorded fixtures predicted a real geometric improvement — + +| offline metric (`measure_tfil_pick_defects`) | `off` | `both-rej`, tau 60 | +|---|---:|---:| +| REACH / arrival within feasible time | 4.5% | 29.4% | +| hot on arrival (`hotAtTta`) | 31.0% | 24.0% | +| **top-1 tile share (pick DIVERSITY)** | **7.0%** | **10.6%** | + +— and a **diversity cost**, because the geometry weight concentrates the draw. +The offline ruler replays the real picker on real fixtures; it says nothing +about whether a tile that is geometrically reachable and less hot actually wins +a round. That is what this run measures. + +**Hypothesis H1.** Weighting the draw by tile geometry (`both-rej`, tau 60) +raises damage/run and round-win rate over the shipped uniform draw, because the +mover arrives at safe tiles instead of merely picking them. + +**Direction is pre-registered as two-sided.** A regression is as interesting as +a win (the diversity cost makes one plausible) and re-deciding the direction +after seeing the data is exactly what this document exists to prevent. + +## Arms — identical except the geo knob + +Both arms: `TR_MOVEMENT=tfil`, everything else at the shipped defaults, one +frozen binary built once from `git archive HEAD` (session 2223ca6). + +| arm | env | role | +|---|---|---| +| `A_off` | `TR_TFIL_GEO_MODE=off` (`TAU=0`) | **REFERENCE** — the shipped uniform draw | +| `B_geo` | `TR_TFIL_GEO_MODE=both-rej` `TR_TFIL_GEO_TAU=60` | treatment | + +**Contamination control (the main risk).** The owner has `both-rej` in a personal +`.env` (`ModularBot_garage/out/.env`, a copy at `tr_bots/ModularBot_geo/.env`), +and this bot's dotenv loader gives the FILE priority over shell exports. If the +tournament's bot instances resolved that file, arm A would silently become arm B +and the whole run would be void. Three guarantees, all verifiable from the logs: + +1. each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this + job generated** (`/tmp/j159_geo/env/A_off.env`, `/tmp/j159_geo/env/B_geo.env`) + in this job's own outdir, so the ONLY `.env` the loader can resolve is mine; +2. the tournament's per-run botdir (`$OUTDIR/.work///run/bots/ModularBot`) + contains only `ModularBot.json`, `ModularBot.sh` and a symlink to the frozen + binary — **no `.env`**, and the loader's fallback is `./.env` then `.env` next + to the executable (it does not walk up parent directories), so the owner's file + is not reachable; +3. every single run's `[env]` boot report is checked for its intended + `TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` before any number is read. **Any run + whose `[env]` disagrees with its arm invalidates the session** and the run is + reported as void rather than analysed. + +## Design + +* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py` + (unmodified). +* Panel: `tools/ab/panel_movement.txt` — the **FROZEN 15-opponent movement + panel**, unchanged. Unit of evidence is the opponent, not the battle. +* 15 opponents x 2 arms x **14 runs** x 3 rounds = **420 battles**. +* Battles serialised: `--wait-arena 45`, one session at a time. + +### Primary metrics (pre-registered, fixed) + +1. **damage/run** (our damage dealt per run) +2. **round-win rate** (rounds won / rounds fought) + +**Hit rate is NOT a primary metric** — it hid a survival regression once already. + +### Secondary / mechanism (reported, never a verdict) + +* incoming hit rate (the survival channel the mechanism actually runs through); +* damage taken/run; +* the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile + share). The live battle logs do not contain the per-pick tile or the arrival + state, so the live run **cannot** re-measure them; that is stated in the + verdict rather than papered over. No new instrumentation is built for this. + +### Statistical treatment + +Same as every previous movement gate: per-opponent paired deltas (arm − +reference), mean delta, SD, SE, 95% CI, a sign test and a **sign-flip +permutation test** (exact when `2^n <= 2^20`, else Monte-Carlo), Wilcoxon as a +cross-check, plus the MDE the analyzer reports for the reference arm's n. + +### MDE — stated up front, and it is LARGE + +At **14 runs/arm** over the frozen 15-opponent panel this design resolves about +**0.28 wins/run** (and the corresponding damage/run MDE the analyzer prints). +A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need +~2.2 h and ~2,900 battles — **which we are NOT doing.** + +**Consequences, recorded before any data:** + +* **A null is the likely outcome.** This would be the *fifth* consecutive + mechanism-positive / outcome-null result in this campaign (after j144, j145, + j146, j147). +* A null here **excludes only a LARGE effect** (>= ~0.28 wins/run). It does not + show the knob does nothing, and it does not retract the offline geometric + measurement. +* Because the offline sweep also measured a **diversity regression** + (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity + cost is an argument **against** shipping, not for it. + +### Verdict rule (fixed now, not re-read later) + +* **Adopt** only if BOTH primaries move in B's favour with `p(sign-flip) < 0.05` + and the effect is at or above the reported MDE. One primary at p<0.05 with + the other not down is reported as a partial signal, not a win. +* Otherwise **do not ship**; the knob stays default `off`. +* The mechanism is reported as measured, with no vote in the verdict. +* No subsetting, no dropping opponents, no re-running to chase a p-value. A + clean null is a fully acceptable result. + +--- + +## MEASURED + +*(appended after the battles — everything above was committed first)*