5.9 KiB
j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?
Written and committed BEFORE a single battle of this experiment ran. No result in the "MEASURED" section below existed when this section was written.
The question
TR_TFIL_GEO_MODE / TR_TFIL_GEO_TAU (shipped default off, j152) shape the
draw over the safe tiles the tfil picker chooses from. The offline sweep on
recorded fixtures predicted a real geometric improvement —
offline metric (measure_tfil_pick_defects) |
off |
both-rej, tau 60 |
|---|---|---|
| REACH / arrival within feasible time | 4.5% | 29.4% |
hot on arrival (hotAtTta) |
31.0% | 24.0% |
| top-1 tile share (pick DIVERSITY) | 7.0% | 10.6% |
— and a diversity cost, because the geometry weight concentrates the draw. The offline ruler replays the real picker on real fixtures; it says nothing about whether a tile that is geometrically reachable and less hot actually wins a round. That is what this run measures.
Hypothesis H1. Weighting the draw by tile geometry (both-rej, tau 60)
raises damage/run and round-win rate over the shipped uniform draw, because the
mover arrives at safe tiles instead of merely picking them.
Direction is pre-registered as two-sided. A regression is as interesting as a win (the diversity cost makes one plausible) and re-deciding the direction after seeing the data is exactly what this document exists to prevent.
Arms — identical except the geo knob
Both arms: TR_MOVEMENT=tfil, everything else at the shipped defaults, one
frozen binary built once from git archive HEAD (session 2223ca6).
| arm | env | role |
|---|---|---|
A_off |
TR_TFIL_GEO_MODE=off (TAU=0) |
REFERENCE — the shipped uniform draw |
B_geo |
TR_TFIL_GEO_MODE=both-rej TR_TFIL_GEO_TAU=60 |
treatment |
Contamination control (the main risk). The owner has both-rej in a personal
.env (ModularBot_garage/out/.env, a copy at tr_bots/ModularBot_geo/.env),
and this bot's dotenv loader gives the FILE priority over shell exports. If the
tournament's bot instances resolved that file, arm A would silently become arm B
and the whole run would be void. Three guarantees, all verifiable from the logs:
- each arm is launched with
TR_ENV_FILEpointing at a per-arm file this job generated (/tmp/j159_geo/env/A_off.env,/tmp/j159_geo/env/B_geo.env) in this job's own outdir, so the ONLY.envthe loader can resolve is mine; - the tournament's per-run botdir (
$OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot) contains onlyModularBot.json,ModularBot.shand a symlink to the frozen binary — no.env, and the loader's fallback is./.envthen.envnext to the executable (it does not walk up parent directories), so the owner's file is not reachable; - every single run's
[env]boot report is checked for its intendedTR_TFIL_GEO_MODE/TR_TFIL_GEO_TAUbefore any number is read. Any run whose[env]disagrees with its arm invalidates the session and the run is reported as void rather than analysed.
Design
- Harness:
tools/ab/tournament_run.sh+tools/ab/tournament_analyze.py(unmodified). - Panel:
tools/ab/panel_movement.txt— the FROZEN 15-opponent movement panel, unchanged. Unit of evidence is the opponent, not the battle. - 15 opponents x 2 arms x 14 runs x 3 rounds = 420 battles.
- Battles serialised:
--wait-arena 45, one session at a time.
Primary metrics (pre-registered, fixed)
- damage/run (our damage dealt per run)
- round-win rate (rounds won / rounds fought)
Hit rate is NOT a primary metric — it hid a survival regression once already.
Secondary / mechanism (reported, never a verdict)
- incoming hit rate (the survival channel the mechanism actually runs through);
- damage taken/run;
- the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile share). The live battle logs do not contain the per-pick tile or the arrival state, so the live run cannot re-measure them; that is stated in the verdict rather than papered over. No new instrumentation is built for this.
Statistical treatment
Same as every previous movement gate: per-opponent paired deltas (arm −
reference), mean delta, SD, SE, 95% CI, a sign test and a sign-flip
permutation test (exact when 2^n <= 2^20, else Monte-Carlo), Wilcoxon as a
cross-check, plus the MDE the analyzer reports for the reference arm's n.
MDE — stated up front, and it is LARGE
At 14 runs/arm over the frozen 15-opponent panel this design resolves about 0.28 wins/run (and the corresponding damage/run MDE the analyzer prints). A two-arm run is 420 battles, ~1 hour. Resolving 0.10 wins/run would need ~2.2 h and ~2,900 battles — which we are NOT doing.
Consequences, recorded before any data:
- A null is the likely outcome. This would be the fifth consecutive mechanism-positive / outcome-null result in this campaign (after j144, j145, j146, j147).
- A null here excludes only a LARGE effect (>= ~0.28 wins/run). It does not show the knob does nothing, and it does not retract the offline geometric measurement.
- Because the offline sweep also measured a diversity regression (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity cost is an argument against shipping, not for it.
Verdict rule (fixed now, not re-read later)
- Adopt only if BOTH primaries move in B's favour with
p(sign-flip) < 0.05and the effect is at or above the reported MDE. One primary at p<0.05 with the other not down is reported as a partial signal, not a win. - Otherwise do not ship; the knob stays default
off. - The mechanism is reported as measured, with no vote in the verdict.
- No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result.
MEASURED
(appended after the battles — everything above was committed first)