11 KiB
j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?
Written and committed BEFORE a single battle of this experiment ran. No result in the "MEASURED" section below existed when this section was written.
The question
TR_TFIL_GEO_MODE / TR_TFIL_GEO_TAU (shipped default off, j152) shape the
draw over the safe tiles the tfil picker chooses from. The offline sweep on
recorded fixtures predicted a real geometric improvement —
offline metric (measure_tfil_pick_defects) |
off |
both-rej, tau 60 |
|---|---|---|
| REACH / arrival within feasible time | 4.5% | 29.4% |
hot on arrival (hotAtTta) |
31.0% | 24.0% |
| top-1 tile share (pick DIVERSITY) | 7.0% | 10.6% |
— and a diversity cost, because the geometry weight concentrates the draw. The offline ruler replays the real picker on real fixtures; it says nothing about whether a tile that is geometrically reachable and less hot actually wins a round. That is what this run measures.
Hypothesis H1. Weighting the draw by tile geometry (both-rej, tau 60)
raises damage/run and round-win rate over the shipped uniform draw, because the
mover arrives at safe tiles instead of merely picking them.
Direction is pre-registered as two-sided. A regression is as interesting as a win (the diversity cost makes one plausible) and re-deciding the direction after seeing the data is exactly what this document exists to prevent.
Arms — identical except the geo knob
Both arms: TR_MOVEMENT=tfil, everything else at the shipped defaults, one
frozen binary built once from git archive HEAD (session 2223ca6).
| arm | env | role |
|---|---|---|
A_off |
TR_TFIL_GEO_MODE=off (TAU=0) |
REFERENCE — the shipped uniform draw |
B_geo |
TR_TFIL_GEO_MODE=both-rej TR_TFIL_GEO_TAU=60 |
treatment |
Contamination control (the main risk). The owner has both-rej in a personal
.env (ModularBot_garage/out/.env, a copy at tr_bots/ModularBot_geo/.env),
and this bot's dotenv loader gives the FILE priority over shell exports. If the
tournament's bot instances resolved that file, arm A would silently become arm B
and the whole run would be void. Three guarantees, all verifiable from the logs:
- each arm is launched with
TR_ENV_FILEpointing at a per-arm file this job generated (/tmp/j159_geo/env/A_off.env,/tmp/j159_geo/env/B_geo.env) in this job's own outdir, so the ONLY.envthe loader can resolve is mine; - the tournament's per-run botdir (
$OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot) contains onlyModularBot.json,ModularBot.shand a symlink to the frozen binary — no.env, and the loader's fallback is./.envthen.envnext to the executable (it does not walk up parent directories), so the owner's file is not reachable; - every single run's
[env]boot report is checked for its intendedTR_TFIL_GEO_MODE/TR_TFIL_GEO_TAUbefore any number is read. Any run whose[env]disagrees with its arm invalidates the session and the run is reported as void rather than analysed.
Design
- Harness:
tools/ab/tournament_run.sh+tools/ab/tournament_analyze.py(unmodified). - Panel:
tools/ab/panel_movement.txt— the FROZEN 15-opponent movement panel, unchanged. Unit of evidence is the opponent, not the battle. - 15 opponents x 2 arms x 14 runs x 3 rounds = 420 battles.
- Battles serialised:
--wait-arena 45, one session at a time.
Primary metrics (pre-registered, fixed)
- damage/run (our damage dealt per run)
- round-win rate (rounds won / rounds fought)
Hit rate is NOT a primary metric — it hid a survival regression once already.
Secondary / mechanism (reported, never a verdict)
- incoming hit rate (the survival channel the mechanism actually runs through);
- damage taken/run;
- the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile share). The live battle logs do not contain the per-pick tile or the arrival state, so the live run cannot re-measure them; that is stated in the verdict rather than papered over. No new instrumentation is built for this.
Statistical treatment
Same as every previous movement gate: per-opponent paired deltas (arm −
reference), mean delta, SD, SE, 95% CI, a sign test and a sign-flip
permutation test (exact when 2^n <= 2^20, else Monte-Carlo), Wilcoxon as a
cross-check, plus the MDE the analyzer reports for the reference arm's n.
MDE — stated up front, and it is LARGE
At 14 runs/arm over the frozen 15-opponent panel this design resolves about 0.28 wins/run (and the corresponding damage/run MDE the analyzer prints). A two-arm run is 420 battles, ~1 hour. Resolving 0.10 wins/run would need ~2.2 h and ~2,900 battles — which we are NOT doing.
Consequences, recorded before any data:
- A null is the likely outcome. This would be the fifth consecutive mechanism-positive / outcome-null result in this campaign (after j144, j145, j146, j147).
- A null here excludes only a LARGE effect (>= ~0.28 wins/run). It does not show the knob does nothing, and it does not retract the offline geometric measurement.
- Because the offline sweep also measured a diversity regression (top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity cost is an argument against shipping, not for it.
Verdict rule (fixed now, not re-read later)
- Adopt only if BOTH primaries move in B's favour with
p(sign-flip) < 0.05and the effect is at or above the reported MDE. One primary at p<0.05 with the other not down is reported as a partial signal, not a win. - Otherwise do not ship; the knob stays default
off. - The mechanism is reported as measured, with no vote in the verdict.
- No subsetting, no dropping opponents, no re-running to chase a p-value. A clean null is a fully acceptable result.
MEASURED
(appended after the battles — everything above was committed first)
MEASURED — the live A/B, 420 battles (j159)
- Provenance. Session
/tmp/ab/j159_geo, commit7c5bc9c, frozen binary sha2569f116e7a9eb9…, paneltools/ab/panel_movement.txt(FROZEN, 15 opponents), 15 x 2 x 14 runs x 3 rounds = 420 battles, 0 failed, 0 never started, 1110 s. Per-arm env files/tmp/j159_geo/env/{A_off,B_geo}.env. - Env verification. All 420 runs carry their intended arm: 210/210
A_offshowTR_TFIL_GEO_MODE=off/TR_TFIL_GEO_TAU=0(parsedoff/0.0), 210/210B_geoshowboth-rej/60(parsedboth/60.0), every run reportsenv file: /tmp/j159_geo/env/<arm>.env (source: TR_ENV_FILE)andmove.effective = tfil. No.envexists anywhere in the session dir, and the loader's fallbacks are./.envthen.envnext to the executable — it does not walk up parents — so the owner's file is unreachable. 0 runs mis-set. - Record correction (no battle re-run). The arms file declared only
TR_ENV_FILE, which the analyzer's liveness guard reads fromsession.jsonand treats as an undeclaredTR_MOVEMENT(fatal contamination).session.jsonand a corrected arms file were rewritten to declare the effective env — the original is kept as/tmp/j159_geo/{session.json.orig,arms_tfil_geo.txt.orig}. The guard then re-verified all 420 declared values verbatim in the boot reports:liveness: 0 run(s) excluded (420 total).
Pooled dashboard (descriptive, NOT the verdict):
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---|---|---|---|---|---|---|---|
A_off |
210 | 112.5 | 195.2 | 1.21 | 255/630 | 40.5% | 17.64% | 393 |
B_geo |
210 | 103.7 | 188.0 | 1.16 | 244/630 | 38.7% | 17.42% | 419 |
Verdict layer (per-opponent paired deltas, arm − reference; sign-flip is the exact 2^15 permutation the pre-registration names as the decision test):
| arm | metric | mean Δ | 95% CI | sign test | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---|---|---|---|---|---|---|
B_geo |
damage/run | -8.83 | [-14.69, -2.97] | 5/15 | 0.3018 | 0.006104 | 0.0115 | 7.65 |
B_geo |
round wins | -0.05 | [-0.18, +0.08] | 4/11 | 0.5488 | 0.4619 | 0.3496 | 0.17 |
B_geo |
damage taken | -7.24 | [-20.84, +6.36] | 8/15 | 1 | 0.2786 | 0.4777 | 17.77 |
B_geo |
incoming hit rate | -1.19 pp | [-3.88, +1.51] | 8/15 | 1 | 0.3962 | 0.5895 | 3.51 |
B_geo |
mean distance | +26.3 px | [+16.3, +36.4] | 15/15 | 6.1e-05 | 6.1e-05 | 0.0007 | 13.15 |
Per-opponent damage/run (the pattern/ram rows carry the loss): Coriantumr -25.1, CassiusClay -24.2, SpinBot -26.2, WallAvoider -13.3, BlitzBat -11.2, HawkOnFire -11.7, Diamond -10.9, TripHammer -9.0, YersiniaPestis -8.9, Dookious -7.3 vs GresSuffurd +2.3, DiamondStealer +3.2, Ascendant +3.7, DrussGT +0.1. Round wins: HawkOnFire +0.50 and WallAvoider +0.14 (both closer opponents) against Coriantumr -0.57, TripHammer -0.21, YersiniaPestis -0.21.
Mechanism, offline (the committed ruler, re-run unchanged for this doc):
measure_tfil_pick_defects reproduces the pre-registered prediction exactly —
REACH 4.5% -> 29.4%, hot-on-arrival 31.0% -> 24.0%, and the diversity cost
top-1 tile share 7.0% -> 10.6% (normalised entropy 0.87 -> 0.85). The live
mean distance +26 px on 15/15 opponents is the same mechanism seen end to
end: the geometry weight prefers tiles that are far better to arrive in, and
the bot sits further out and deals less damage.
VERDICT — DO NOT ADOPT
- H1 is rejected. Round wins are flat (-0.05/run, p(sign-flip) = 0.46, under the 0.17 MDE) and damage/run is down 8.83 (p(sign-flip) = 0.0061, above the 7.65 MDE, Wilcoxon p = 0.011). The pre-registration required BOTH primaries up; one is down and significant by the test it named.
- The MDE, restated. 14 runs/arm on the frozen 15-opponent panel resolves 0.17 wins/run and 7.65 damage/run (better than the 0.28 pre-registered estimate). So this run excludes a large benefit; it also positively measures a small harm in damage. It says nothing about effects below those numbers.
- The mechanism moved exactly as predicted, and that is what makes it bad. REACH/arrival 4.5% -> 29.4% and hot-on-arrival 31.0% -> 24.0% are real and reproducible, but they bought +26 px of distance and fewer damage points, not survival: incoming hit rate moved -1.19 pp, a fifth of its own 3.51 MDE. "Arrive at a safe tile" turned out to mean "arrive further away".
- Diversity worsened, as the offline sweep warned. Top-1 tile share 7.0% -> 10.6%, and the live losses concentrate against the opponents that punish a long-range mover (SpinBot -26.2 damage with a -14.5 pp hit-rate shift, the pattern guns -9 to -25). Outcomes are null-to-negative AND diversity is worse: that is the argument against shipping, exactly the pre-registered case.
- A null on wins lets us claim only "no large win". It does not show the knob is inert, and it does not retract the geometric measurement — but the geometry measurement is not an argument for shipping when the live consequence of it is measurably less damage from measurably further away.
Ship state: TR_TFIL_GEO_MODE stays default off. Nothing changes. Not
adopted, not adopted default-off. This is the fifth mechanism-positive /
outcome-not-positive result in the movement campaign (j144, j145, j146, j147,
j159) — and the first one where the mechanism is anti-correlated with damage.