126 lines
5.9 KiB
Markdown
126 lines
5.9 KiB
Markdown
# j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?
|
||
|
||
**Written and committed BEFORE a single battle of this experiment ran.** No
|
||
result in the "MEASURED" section below existed when this section was written.
|
||
|
||
## The question
|
||
|
||
`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` (shipped default `off`, j152) shape the
|
||
**draw** over the safe tiles the tfil picker chooses from. The offline sweep on
|
||
recorded fixtures predicted a real geometric improvement —
|
||
|
||
| offline metric (`measure_tfil_pick_defects`) | `off` | `both-rej`, tau 60 |
|
||
|---|---:|---:|
|
||
| REACH / arrival within feasible time | 4.5% | 29.4% |
|
||
| hot on arrival (`hotAtTta`) | 31.0% | 24.0% |
|
||
| **top-1 tile share (pick DIVERSITY)** | **7.0%** | **10.6%** |
|
||
|
||
— and a **diversity cost**, because the geometry weight concentrates the draw.
|
||
The offline ruler replays the real picker on real fixtures; it says nothing
|
||
about whether a tile that is geometrically reachable and less hot actually wins
|
||
a round. That is what this run measures.
|
||
|
||
**Hypothesis H1.** Weighting the draw by tile geometry (`both-rej`, tau 60)
|
||
raises damage/run and round-win rate over the shipped uniform draw, because the
|
||
mover arrives at safe tiles instead of merely picking them.
|
||
|
||
**Direction is pre-registered as two-sided.** A regression is as interesting as
|
||
a win (the diversity cost makes one plausible) and re-deciding the direction
|
||
after seeing the data is exactly what this document exists to prevent.
|
||
|
||
## Arms — identical except the geo knob
|
||
|
||
Both arms: `TR_MOVEMENT=tfil`, everything else at the shipped defaults, one
|
||
frozen binary built once from `git archive HEAD` (session 2223ca6).
|
||
|
||
| arm | env | role |
|
||
|---|---|---|
|
||
| `A_off` | `TR_TFIL_GEO_MODE=off` (`TAU=0`) | **REFERENCE** — the shipped uniform draw |
|
||
| `B_geo` | `TR_TFIL_GEO_MODE=both-rej` `TR_TFIL_GEO_TAU=60` | treatment |
|
||
|
||
**Contamination control (the main risk).** The owner has `both-rej` in a personal
|
||
`.env` (`ModularBot_garage/out/.env`, a copy at `tr_bots/ModularBot_geo/.env`),
|
||
and this bot's dotenv loader gives the FILE priority over shell exports. If the
|
||
tournament's bot instances resolved that file, arm A would silently become arm B
|
||
and the whole run would be void. Three guarantees, all verifiable from the logs:
|
||
|
||
1. each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this
|
||
job generated** (`/tmp/j159_geo/env/A_off.env`, `/tmp/j159_geo/env/B_geo.env`)
|
||
in this job's own outdir, so the ONLY `.env` the loader can resolve is mine;
|
||
2. the tournament's per-run botdir (`$OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot`)
|
||
contains only `ModularBot.json`, `ModularBot.sh` and a symlink to the frozen
|
||
binary — **no `.env`**, and the loader's fallback is `./.env` then `.env` next
|
||
to the executable (it does not walk up parent directories), so the owner's file
|
||
is not reachable;
|
||
3. every single run's `[env]` boot report is checked for its intended
|
||
`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` before any number is read. **Any run
|
||
whose `[env]` disagrees with its arm invalidates the session** and the run is
|
||
reported as void rather than analysed.
|
||
|
||
## Design
|
||
|
||
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`
|
||
(unmodified).
|
||
* Panel: `tools/ab/panel_movement.txt` — the **FROZEN 15-opponent movement
|
||
panel**, unchanged. Unit of evidence is the opponent, not the battle.
|
||
* 15 opponents x 2 arms x **14 runs** x 3 rounds = **420 battles**.
|
||
* Battles serialised: `--wait-arena 45`, one session at a time.
|
||
|
||
### Primary metrics (pre-registered, fixed)
|
||
|
||
1. **damage/run** (our damage dealt per run)
|
||
2. **round-win rate** (rounds won / rounds fought)
|
||
|
||
**Hit rate is NOT a primary metric** — it hid a survival regression once already.
|
||
|
||
### Secondary / mechanism (reported, never a verdict)
|
||
|
||
* incoming hit rate (the survival channel the mechanism actually runs through);
|
||
* damage taken/run;
|
||
* the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile
|
||
share). The live battle logs do not contain the per-pick tile or the arrival
|
||
state, so the live run **cannot** re-measure them; that is stated in the
|
||
verdict rather than papered over. No new instrumentation is built for this.
|
||
|
||
### Statistical treatment
|
||
|
||
Same as every previous movement gate: per-opponent paired deltas (arm −
|
||
reference), mean delta, SD, SE, 95% CI, a sign test and a **sign-flip
|
||
permutation test** (exact when `2^n <= 2^20`, else Monte-Carlo), Wilcoxon as a
|
||
cross-check, plus the MDE the analyzer reports for the reference arm's n.
|
||
|
||
### MDE — stated up front, and it is LARGE
|
||
|
||
At **14 runs/arm** over the frozen 15-opponent panel this design resolves about
|
||
**0.28 wins/run** (and the corresponding damage/run MDE the analyzer prints).
|
||
A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need
|
||
~2.2 h and ~2,900 battles — **which we are NOT doing.**
|
||
|
||
**Consequences, recorded before any data:**
|
||
|
||
* **A null is the likely outcome.** This would be the *fifth* consecutive
|
||
mechanism-positive / outcome-null result in this campaign (after j144, j145,
|
||
j146, j147).
|
||
* A null here **excludes only a LARGE effect** (>= ~0.28 wins/run). It does not
|
||
show the knob does nothing, and it does not retract the offline geometric
|
||
measurement.
|
||
* Because the offline sweep also measured a **diversity regression**
|
||
(top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity
|
||
cost is an argument **against** shipping, not for it.
|
||
|
||
### Verdict rule (fixed now, not re-read later)
|
||
|
||
* **Adopt** only if BOTH primaries move in B's favour with `p(sign-flip) < 0.05`
|
||
and the effect is at or above the reported MDE. One primary at p<0.05 with
|
||
the other not down is reported as a partial signal, not a win.
|
||
* Otherwise **do not ship**; the knob stays default `off`.
|
||
* The mechanism is reported as measured, with no vote in the verdict.
|
||
* No subsetting, no dropping opponents, no re-running to chase a p-value. A
|
||
clean null is a fully acceptable result.
|
||
|
||
---
|
||
|
||
## MEASURED
|
||
|
||
*(appended after the battles — everything above was committed first)*
|