Files
SirRoboGarage/docs/tfil_geo_ab.md
T

11 KiB
Raw Blame History

j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?

Written and committed BEFORE a single battle of this experiment ran. No result in the "MEASURED" section below existed when this section was written.

The question

TR_TFIL_GEO_MODE / TR_TFIL_GEO_TAU (shipped default off, j152) shape the draw over the safe tiles the tfil picker chooses from. The offline sweep on recorded fixtures predicted a real geometric improvement —

offline metric (measure_tfil_pick_defects) off both-rej, tau 60
REACH / arrival within feasible time 4.5% 29.4%
hot on arrival (hotAtTta) 31.0% 24.0%
top-1 tile share (pick DIVERSITY) 7.0% 10.6%

— and a diversity cost, because the geometry weight concentrates the draw. The offline ruler replays the real picker on real fixtures; it says nothing about whether a tile that is geometrically reachable and less hot actually wins a round. That is what this run measures.

Hypothesis H1. Weighting the draw by tile geometry (both-rej, tau 60) raises damage/run and round-win rate over the shipped uniform draw, because the mover arrives at safe tiles instead of merely picking them.

Direction is pre-registered as two-sided. A regression is as interesting as a win (the diversity cost makes one plausible) and re-deciding the direction after seeing the data is exactly what this document exists to prevent.

Arms — identical except the geo knob

Both arms: TR_MOVEMENT=tfil, everything else at the shipped defaults, one frozen binary built once from git archive HEAD (session 2223ca6).

arm env role
A_off TR_TFIL_GEO_MODE=off (TAU=0) REFERENCE — the shipped uniform draw
B_geo TR_TFIL_GEO_MODE=both-rej TR_TFIL_GEO_TAU=60 treatment

Contamination control (the main risk). The owner has both-rej in a personal .env (ModularBot_garage/out/.env, a copy at tr_bots/ModularBot_geo/.env), and this bot's dotenv loader gives the FILE priority over shell exports. If the tournament's bot instances resolved that file, arm A would silently become arm B and the whole run would be void. Three guarantees, all verifiable from the logs:

  1. each arm is launched with TR_ENV_FILE pointing at a per-arm file this job generated (/tmp/j159_geo/env/A_off.env, /tmp/j159_geo/env/B_geo.env) in this job's own outdir, so the ONLY .env the loader can resolve is mine;
  2. the tournament's per-run botdir ($OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot) contains only ModularBot.json, ModularBot.sh and a symlink to the frozen binary — no .env, and the loader's fallback is ./.env then .env next to the executable (it does not walk up parent directories), so the owner's file is not reachable;
  3. every single run's [env] boot report is checked for its intended TR_TFIL_GEO_MODE / TR_TFIL_GEO_TAU before any number is read. Any run whose [env] disagrees with its arm invalidates the session and the run is reported as void rather than analysed.

Design

  • Harness: tools/ab/tournament_run.sh + tools/ab/tournament_analyze.py (unmodified).
  • Panel: tools/ab/panel_movement.txt — the FROZEN 15-opponent movement panel, unchanged. Unit of evidence is the opponent, not the battle.
  • 15 opponents x 2 arms x 14 runs x 3 rounds = 420 battles.
  • Battles serialised: --wait-arena 45, one session at a time.

Primary metrics (pre-registered, fixed)

  1. damage/run (our damage dealt per run)
  2. round-win rate (rounds won / rounds fought)

Hit rate is NOT a primary metric — it hid a survival regression once already.

Secondary / mechanism (reported, never a verdict)

  • incoming hit rate (the survival channel the mechanism actually runs through);
  • damage taken/run;
  • the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile share). The live battle logs do not contain the per-pick tile or the arrival state, so the live run cannot re-measure them; that is stated in the verdict rather than papered over. No new instrumentation is built for this.

Statistical treatment

Same as every previous movement gate: per-opponent paired deltas (arm − reference), mean delta, SD, SE, 95% CI, a sign test and a sign-flip permutation test (exact when 2^n <= 2^20, else Monte-Carlo), Wilcoxon as a cross-check, plus the MDE the analyzer reports for the reference arm's n.

MDE — stated up front, and it is LARGE

At 14 runs/arm over the frozen 15-opponent panel this design resolves about 0.28 wins/run (and the corresponding damage/run MDE the analyzer prints). A two-arm run is 420 battles, ~1 hour. Resolving 0.10 wins/run would need ~2.2 h and ~2,900 battles — which we are NOT doing.

Consequences, recorded before any data:

  • A null is the likely outcome. This would be the fifth consecutive mechanism-positive / outcome-null result in this campaign (after j144, j145, j146, j147).
  • A null here excludes only a LARGE effect (>= ~0.28 wins/run). It does not show the knob does nothing, and it does not retract the offline geometric measurement.
  • Because the offline sweep also measured a diversity regression (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity cost is an argument against shipping, not for it.

Verdict rule (fixed now, not re-read later)

  • Adopt only if BOTH primaries move in B's favour with p(sign-flip) < 0.05 and the effect is at or above the reported MDE. One primary at p<0.05 with the other not down is reported as a partial signal, not a win.
  • Otherwise do not ship; the knob stays default off.
  • The mechanism is reported as measured, with no vote in the verdict.
  • No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result.

MEASURED

(appended after the battles — everything above was committed first)

MEASURED — the live A/B, 420 battles (j159)

  • Provenance. Session /tmp/ab/j159_geo, commit 7c5bc9c, frozen binary sha256 9f116e7a9eb9…, panel tools/ab/panel_movement.txt (FROZEN, 15 opponents), 15 x 2 x 14 runs x 3 rounds = 420 battles, 0 failed, 0 never started, 1110 s. Per-arm env files /tmp/j159_geo/env/{A_off,B_geo}.env.
  • Env verification. All 420 runs carry their intended arm: 210/210 A_off show TR_TFIL_GEO_MODE=off / TR_TFIL_GEO_TAU=0 (parsed off/0.0), 210/210 B_geo show both-rej/60 (parsed both/60.0), every run reports env file: /tmp/j159_geo/env/<arm>.env (source: TR_ENV_FILE) and move.effective = tfil. No .env exists anywhere in the session dir, and the loader's fallbacks are ./.env then .env next to the executable — it does not walk up parents — so the owner's file is unreachable. 0 runs mis-set.
  • Record correction (no battle re-run). The arms file declared only TR_ENV_FILE, which the analyzer's liveness guard reads from session.json and treats as an undeclared TR_MOVEMENT (fatal contamination). session.json and a corrected arms file were rewritten to declare the effective env — the original is kept as /tmp/j159_geo/{session.json.orig,arms_tfil_geo.txt.orig}. The guard then re-verified all 420 declared values verbatim in the boot reports: liveness: 0 run(s) excluded (420 total).

Pooled dashboard (descriptive, NOT the verdict):

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
A_off 210 112.5 195.2 1.21 255/630 40.5% 17.64% 393
B_geo 210 103.7 188.0 1.16 244/630 38.7% 17.42% 419

Verdict layer (per-opponent paired deltas, arm − reference; sign-flip is the exact 2^15 permutation the pre-registration names as the decision test):

arm metric mean Δ 95% CI sign test p(sign) p(sign-flip) Wilcoxon p MDE
B_geo damage/run -8.83 [-14.69, -2.97] 5/15 0.3018 0.006104 0.0115 7.65
B_geo round wins -0.05 [-0.18, +0.08] 4/11 0.5488 0.4619 0.3496 0.17
B_geo damage taken -7.24 [-20.84, +6.36] 8/15 1 0.2786 0.4777 17.77
B_geo incoming hit rate -1.19 pp [-3.88, +1.51] 8/15 1 0.3962 0.5895 3.51
B_geo mean distance +26.3 px [+16.3, +36.4] 15/15 6.1e-05 6.1e-05 0.0007 13.15

Per-opponent damage/run (the pattern/ram rows carry the loss): Coriantumr -25.1, CassiusClay -24.2, SpinBot -26.2, WallAvoider -13.3, BlitzBat -11.2, HawkOnFire -11.7, Diamond -10.9, TripHammer -9.0, YersiniaPestis -8.9, Dookious -7.3 vs GresSuffurd +2.3, DiamondStealer +3.2, Ascendant +3.7, DrussGT +0.1. Round wins: HawkOnFire +0.50 and WallAvoider +0.14 (both closer opponents) against Coriantumr -0.57, TripHammer -0.21, YersiniaPestis -0.21.

Mechanism, offline (the committed ruler, re-run unchanged for this doc): measure_tfil_pick_defects reproduces the pre-registered prediction exactly — REACH 4.5% -> 29.4%, hot-on-arrival 31.0% -> 24.0%, and the diversity cost top-1 tile share 7.0% -> 10.6% (normalised entropy 0.87 -> 0.85). The live mean distance +26 px on 15/15 opponents is the same mechanism seen end to end: the geometry weight prefers tiles that are far better to arrive in, and the bot sits further out and deals less damage.

VERDICT — DO NOT ADOPT

  1. H1 is rejected. Round wins are flat (-0.05/run, p(sign-flip) = 0.46, under the 0.17 MDE) and damage/run is down 8.83 (p(sign-flip) = 0.0061, above the 7.65 MDE, Wilcoxon p = 0.011). The pre-registration required BOTH primaries up; one is down and significant by the test it named.
  2. The MDE, restated. 14 runs/arm on the frozen 15-opponent panel resolves 0.17 wins/run and 7.65 damage/run (better than the 0.28 pre-registered estimate). So this run excludes a large benefit; it also positively measures a small harm in damage. It says nothing about effects below those numbers.
  3. The mechanism moved exactly as predicted, and that is what makes it bad. REACH/arrival 4.5% -> 29.4% and hot-on-arrival 31.0% -> 24.0% are real and reproducible, but they bought +26 px of distance and fewer damage points, not survival: incoming hit rate moved -1.19 pp, a fifth of its own 3.51 MDE. "Arrive at a safe tile" turned out to mean "arrive further away".
  4. Diversity worsened, as the offline sweep warned. Top-1 tile share 7.0% -> 10.6%, and the live losses concentrate against the opponents that punish a long-range mover (SpinBot -26.2 damage with a -14.5 pp hit-rate shift, the pattern guns -9 to -25). Outcomes are null-to-negative AND diversity is worse: that is the argument against shipping, exactly the pre-registered case.
  5. A null on wins lets us claim only "no large win". It does not show the knob is inert, and it does not retract the geometric measurement — but the geometry measurement is not an argument for shipping when the live consequence of it is measurably less damage from measurably further away.

Ship state: TR_TFIL_GEO_MODE stays default off. Nothing changes. Not adopted, not adopted default-off. This is the fifth mechanism-positive / outcome-not-positive result in the movement campaign (j144, j145, j146, j147, j159) — and the first one where the mechanism is anti-correlated with damage.