Files
SirRoboGarage/docs/tfil_geo_ab.md
T

5.9 KiB
Raw Blame History

j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?

Written and committed BEFORE a single battle of this experiment ran. No result in the "MEASURED" section below existed when this section was written.

The question

TR_TFIL_GEO_MODE / TR_TFIL_GEO_TAU (shipped default off, j152) shape the draw over the safe tiles the tfil picker chooses from. The offline sweep on recorded fixtures predicted a real geometric improvement —

offline metric (measure_tfil_pick_defects) off both-rej, tau 60
REACH / arrival within feasible time 4.5% 29.4%
hot on arrival (hotAtTta) 31.0% 24.0%
top-1 tile share (pick DIVERSITY) 7.0% 10.6%

— and a diversity cost, because the geometry weight concentrates the draw. The offline ruler replays the real picker on real fixtures; it says nothing about whether a tile that is geometrically reachable and less hot actually wins a round. That is what this run measures.

Hypothesis H1. Weighting the draw by tile geometry (both-rej, tau 60) raises damage/run and round-win rate over the shipped uniform draw, because the mover arrives at safe tiles instead of merely picking them.

Direction is pre-registered as two-sided. A regression is as interesting as a win (the diversity cost makes one plausible) and re-deciding the direction after seeing the data is exactly what this document exists to prevent.

Arms — identical except the geo knob

Both arms: TR_MOVEMENT=tfil, everything else at the shipped defaults, one frozen binary built once from git archive HEAD (session 2223ca6).

arm env role
A_off TR_TFIL_GEO_MODE=off (TAU=0) REFERENCE — the shipped uniform draw
B_geo TR_TFIL_GEO_MODE=both-rej TR_TFIL_GEO_TAU=60 treatment

Contamination control (the main risk). The owner has both-rej in a personal .env (ModularBot_garage/out/.env, a copy at tr_bots/ModularBot_geo/.env), and this bot's dotenv loader gives the FILE priority over shell exports. If the tournament's bot instances resolved that file, arm A would silently become arm B and the whole run would be void. Three guarantees, all verifiable from the logs:

  1. each arm is launched with TR_ENV_FILE pointing at a per-arm file this job generated (/tmp/j159_geo/env/A_off.env, /tmp/j159_geo/env/B_geo.env) in this job's own outdir, so the ONLY .env the loader can resolve is mine;
  2. the tournament's per-run botdir ($OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot) contains only ModularBot.json, ModularBot.sh and a symlink to the frozen binary — no .env, and the loader's fallback is ./.env then .env next to the executable (it does not walk up parent directories), so the owner's file is not reachable;
  3. every single run's [env] boot report is checked for its intended TR_TFIL_GEO_MODE / TR_TFIL_GEO_TAU before any number is read. Any run whose [env] disagrees with its arm invalidates the session and the run is reported as void rather than analysed.

Design

  • Harness: tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py (unmodified).
  • Panel: tools/ab/panel_movement.txt — the FROZEN 15-opponent movement panel, unchanged. Unit of evidence is the opponent, not the battle.
  • 15 opponents x 2 arms x 14 runs x 3 rounds = 420 battles.
  • Battles serialised: --wait-arena 45, one session at a time.

Primary metrics (pre-registered, fixed)

  1. damage/run (our damage dealt per run)
  2. round-win rate (rounds won / rounds fought)

Hit rate is NOT a primary metric — it hid a survival regression once already.

Secondary / mechanism (reported, never a verdict)

  • incoming hit rate (the survival channel the mechanism actually runs through);
  • damage taken/run;
  • the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile share). The live battle logs do not contain the per-pick tile or the arrival state, so the live run cannot re-measure them; that is stated in the verdict rather than papered over. No new instrumentation is built for this.

Statistical treatment

Same as every previous movement gate: per-opponent paired deltas (arm − reference), mean delta, SD, SE, 95% CI, a sign test and a sign-flip permutation test (exact when 2^n <= 2^20, else Monte-Carlo), Wilcoxon as a cross-check, plus the MDE the analyzer reports for the reference arm's n.

MDE — stated up front, and it is LARGE

At 14 runs/arm over the frozen 15-opponent panel this design resolves about 0.28 wins/run (and the corresponding damage/run MDE the analyzer prints). A two-arm run is 420 battles, ~1 hour. Resolving 0.10 wins/run would need ~2.2 h and ~2,900 battles — which we are NOT doing.

Consequences, recorded before any data:

  • A null is the likely outcome. This would be the fifth consecutive mechanism-positive / outcome-null result in this campaign (after j144, j145, j146, j147).
  • A null here excludes only a LARGE effect (>= ~0.28 wins/run). It does not show the knob does nothing, and it does not retract the offline geometric measurement.
  • Because the offline sweep also measured a diversity regression (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity cost is an argument against shipping, not for it.

Verdict rule (fixed now, not re-read later)

  • Adopt only if BOTH primaries move in B's favour with p(sign-flip) < 0.05 and the effect is at or above the reported MDE. One primary at p<0.05 with the other not down is reported as a partial signal, not a win.
  • Otherwise do not ship; the knob stays default off.
  • The mechanism is reported as measured, with no vote in the verdict.
  • No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result.

MEASURED

(appended after the battles — everything above was committed first)