210 lines
11 KiB
Markdown
210 lines
11 KiB
Markdown
# j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?
|
||
|
||
**Written and committed BEFORE a single battle of this experiment ran.** No
|
||
result in the "MEASURED" section below existed when this section was written.
|
||
|
||
## The question
|
||
|
||
`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` (shipped default `off`, j152) shape the
|
||
**draw** over the safe tiles the tfil picker chooses from. The offline sweep on
|
||
recorded fixtures predicted a real geometric improvement —
|
||
|
||
| offline metric (`measure_tfil_pick_defects`) | `off` | `both-rej`, tau 60 |
|
||
|---|---:|---:|
|
||
| REACH / arrival within feasible time | 4.5% | 29.4% |
|
||
| hot on arrival (`hotAtTta`) | 31.0% | 24.0% |
|
||
| **top-1 tile share (pick DIVERSITY)** | **7.0%** | **10.6%** |
|
||
|
||
— and a **diversity cost**, because the geometry weight concentrates the draw.
|
||
The offline ruler replays the real picker on real fixtures; it says nothing
|
||
about whether a tile that is geometrically reachable and less hot actually wins
|
||
a round. That is what this run measures.
|
||
|
||
**Hypothesis H1.** Weighting the draw by tile geometry (`both-rej`, tau 60)
|
||
raises damage/run and round-win rate over the shipped uniform draw, because the
|
||
mover arrives at safe tiles instead of merely picking them.
|
||
|
||
**Direction is pre-registered as two-sided.** A regression is as interesting as
|
||
a win (the diversity cost makes one plausible) and re-deciding the direction
|
||
after seeing the data is exactly what this document exists to prevent.
|
||
|
||
## Arms — identical except the geo knob
|
||
|
||
Both arms: `TR_MOVEMENT=tfil`, everything else at the shipped defaults, one
|
||
frozen binary built once from `git archive HEAD` (session 2223ca6).
|
||
|
||
| arm | env | role |
|
||
|---|---|---|
|
||
| `A_off` | `TR_TFIL_GEO_MODE=off` (`TAU=0`) | **REFERENCE** — the shipped uniform draw |
|
||
| `B_geo` | `TR_TFIL_GEO_MODE=both-rej` `TR_TFIL_GEO_TAU=60` | treatment |
|
||
|
||
**Contamination control (the main risk).** The owner has `both-rej` in a personal
|
||
`.env` (`ModularBot_garage/out/.env`, a copy at `tr_bots/ModularBot_geo/.env`),
|
||
and this bot's dotenv loader gives the FILE priority over shell exports. If the
|
||
tournament's bot instances resolved that file, arm A would silently become arm B
|
||
and the whole run would be void. Three guarantees, all verifiable from the logs:
|
||
|
||
1. each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this
|
||
job generated** (`/tmp/j159_geo/env/A_off.env`, `/tmp/j159_geo/env/B_geo.env`)
|
||
in this job's own outdir, so the ONLY `.env` the loader can resolve is mine;
|
||
2. the tournament's per-run botdir (`$OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot`)
|
||
contains only `ModularBot.json`, `ModularBot.sh` and a symlink to the frozen
|
||
binary — **no `.env`**, and the loader's fallback is `./.env` then `.env` next
|
||
to the executable (it does not walk up parent directories), so the owner's file
|
||
is not reachable;
|
||
3. every single run's `[env]` boot report is checked for its intended
|
||
`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` before any number is read. **Any run
|
||
whose `[env]` disagrees with its arm invalidates the session** and the run is
|
||
reported as void rather than analysed.
|
||
|
||
## Design
|
||
|
||
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`
|
||
(unmodified).
|
||
* Panel: `tools/ab/panel_movement.txt` — the **FROZEN 15-opponent movement
|
||
panel**, unchanged. Unit of evidence is the opponent, not the battle.
|
||
* 15 opponents x 2 arms x **14 runs** x 3 rounds = **420 battles**.
|
||
* Battles serialised: `--wait-arena 45`, one session at a time.
|
||
|
||
### Primary metrics (pre-registered, fixed)
|
||
|
||
1. **damage/run** (our damage dealt per run)
|
||
2. **round-win rate** (rounds won / rounds fought)
|
||
|
||
**Hit rate is NOT a primary metric** — it hid a survival regression once already.
|
||
|
||
### Secondary / mechanism (reported, never a verdict)
|
||
|
||
* incoming hit rate (the survival channel the mechanism actually runs through);
|
||
* damage taken/run;
|
||
* the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile
|
||
share). The live battle logs do not contain the per-pick tile or the arrival
|
||
state, so the live run **cannot** re-measure them; that is stated in the
|
||
verdict rather than papered over. No new instrumentation is built for this.
|
||
|
||
### Statistical treatment
|
||
|
||
Same as every previous movement gate: per-opponent paired deltas (arm −
|
||
reference), mean delta, SD, SE, 95% CI, a sign test and a **sign-flip
|
||
permutation test** (exact when `2^n <= 2^20`, else Monte-Carlo), Wilcoxon as a
|
||
cross-check, plus the MDE the analyzer reports for the reference arm's n.
|
||
|
||
### MDE — stated up front, and it is LARGE
|
||
|
||
At **14 runs/arm** over the frozen 15-opponent panel this design resolves about
|
||
**0.28 wins/run** (and the corresponding damage/run MDE the analyzer prints).
|
||
A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need
|
||
~2.2 h and ~2,900 battles — **which we are NOT doing.**
|
||
|
||
**Consequences, recorded before any data:**
|
||
|
||
* **A null is the likely outcome.** This would be the *fifth* consecutive
|
||
mechanism-positive / outcome-null result in this campaign (after j144, j145,
|
||
j146, j147).
|
||
* A null here **excludes only a LARGE effect** (>= ~0.28 wins/run). It does not
|
||
show the knob does nothing, and it does not retract the offline geometric
|
||
measurement.
|
||
* Because the offline sweep also measured a **diversity regression**
|
||
(top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity
|
||
cost is an argument **against** shipping, not for it.
|
||
|
||
### Verdict rule (fixed now, not re-read later)
|
||
|
||
* **Adopt** only if BOTH primaries move in B's favour with `p(sign-flip) < 0.05`
|
||
and the effect is at or above the reported MDE. One primary at p<0.05 with
|
||
the other not down is reported as a partial signal, not a win.
|
||
* Otherwise **do not ship**; the knob stays default `off`.
|
||
* The mechanism is reported as measured, with no vote in the verdict.
|
||
* No subsetting, no dropping opponents, no re-running to chase a p-value. A
|
||
clean null is a fully acceptable result.
|
||
|
||
---
|
||
|
||
## MEASURED
|
||
|
||
*(appended after the battles — everything above was committed first)*
|
||
### MEASURED — the live A/B, 420 battles (j159)
|
||
|
||
* **Provenance.** Session `/tmp/ab/j159_geo`, commit `7c5bc9c`, frozen binary
|
||
sha256 `9f116e7a9eb9…`, panel `tools/ab/panel_movement.txt` (FROZEN, 15
|
||
opponents), 15 x 2 x **14 runs** x 3 rounds = **420 battles, 0 failed, 0 never
|
||
started, 1110 s**. Per-arm env files `/tmp/j159_geo/env/{A_off,B_geo}.env`.
|
||
* **Env verification.** All **420** runs carry their intended arm: 210/210
|
||
`A_off` show `TR_TFIL_GEO_MODE=off` / `TR_TFIL_GEO_TAU=0` (parsed `off`/`0.0`),
|
||
210/210 `B_geo` show `both-rej`/`60` (parsed `both`/`60.0`), every run reports
|
||
`env file: /tmp/j159_geo/env/<arm>.env (source: TR_ENV_FILE)` and
|
||
`move.effective = tfil`. No `.env` exists anywhere in the session dir, and the
|
||
loader's fallbacks are `./.env` then `.env` next to the executable — it does
|
||
not walk up parents — so the owner's file is unreachable. **0 runs mis-set.**
|
||
* **Record correction (no battle re-run).** The arms file declared only
|
||
`TR_ENV_FILE`, which the analyzer's liveness guard reads from `session.json`
|
||
and treats as an undeclared `TR_MOVEMENT` (fatal contamination). `session.json`
|
||
and a corrected arms file were rewritten to declare the effective env — the
|
||
original is kept as `/tmp/j159_geo/{session.json.orig,arms_tfil_geo.txt.orig}`.
|
||
The guard then re-verified all 420 declared values verbatim in the boot
|
||
reports: `liveness: 0 run(s) excluded (420 total)`.
|
||
|
||
**Pooled dashboard (descriptive, NOT the verdict):**
|
||
|
||
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---:|
|
||
| `A_off` | 210 | 112.5 | 195.2 | 1.21 | 255/630 | 40.5% | 17.64% | 393 |
|
||
| `B_geo` | 210 | 103.7 | 188.0 | 1.16 | 244/630 | 38.7% | 17.42% | 419 |
|
||
|
||
**Verdict layer** (per-opponent paired deltas, arm − reference; sign-flip is
|
||
the exact 2^15 permutation the pre-registration names as the decision test):
|
||
|
||
| arm | metric | mean Δ | 95% CI | sign test | p(sign) | **p(sign-flip)** | Wilcoxon p | MDE |
|
||
|---|---|---:|---|---:|---:|---:|---:|---:|
|
||
| `B_geo` | **damage/run** | **-8.83** | [-14.69, -2.97] | 5/15 | 0.3018 | **0.006104** | 0.0115 | 7.65 |
|
||
| `B_geo` | round wins | -0.05 | [-0.18, +0.08] | 4/11 | 0.5488 | 0.4619 | 0.3496 | 0.17 |
|
||
| `B_geo` | damage taken | -7.24 | [-20.84, +6.36] | 8/15 | 1 | 0.2786 | 0.4777 | 17.77 |
|
||
| `B_geo` | incoming hit rate | -1.19 pp | [-3.88, +1.51] | 8/15 | 1 | 0.3962 | 0.5895 | 3.51 |
|
||
| `B_geo` | mean distance | **+26.3 px** | [+16.3, +36.4] | **15/15** | 6.1e-05 | 6.1e-05 | 0.0007 | 13.15 |
|
||
|
||
**Per-opponent damage/run** (the pattern/ram rows carry the loss): Coriantumr
|
||
-25.1, CassiusClay -24.2, SpinBot -26.2, WallAvoider -13.3, BlitzBat -11.2,
|
||
HawkOnFire -11.7, Diamond -10.9, TripHammer -9.0, YersiniaPestis -8.9,
|
||
Dookious -7.3 vs GresSuffurd +2.3, DiamondStealer +3.2, Ascendant +3.7,
|
||
DrussGT +0.1. Round wins: HawkOnFire +0.50 and WallAvoider +0.14 (both closer
|
||
opponents) against Coriantumr -0.57, TripHammer -0.21, YersiniaPestis -0.21.
|
||
|
||
**Mechanism, offline (the committed ruler, re-run unchanged for this doc):**
|
||
`measure_tfil_pick_defects` reproduces the pre-registered prediction exactly —
|
||
REACH **4.5% -> 29.4%**, hot-on-arrival **31.0% -> 24.0%**, and the diversity cost
|
||
**top-1 tile share 7.0% -> 10.6%** (normalised entropy 0.87 -> 0.85). The live
|
||
mean distance **+26 px on 15/15 opponents** is the same mechanism seen end to
|
||
end: the geometry weight prefers tiles that are far better to *arrive* in, and
|
||
the bot sits further out and deals **less** damage.
|
||
|
||
### VERDICT — DO NOT ADOPT
|
||
|
||
1. **H1 is rejected.** Round wins are flat (-0.05/run, p(sign-flip) = 0.46, under
|
||
the 0.17 MDE) and damage/run is **down 8.83** (p(sign-flip) = 0.0061, above
|
||
the 7.65 MDE, Wilcoxon p = 0.011). The pre-registration required BOTH
|
||
primaries up; one is down and significant by the test it named.
|
||
2. **The MDE, restated.** 14 runs/arm on the frozen 15-opponent panel resolves
|
||
**0.17 wins/run** and **7.65 damage/run** (better than the 0.28 pre-registered
|
||
estimate). So this run excludes a large *benefit*; it also positively measures
|
||
a small *harm* in damage. It says nothing about effects below those numbers.
|
||
3. **The mechanism moved exactly as predicted, and that is what makes it bad.**
|
||
REACH/arrival 4.5% -> 29.4% and hot-on-arrival 31.0% -> 24.0% are real and
|
||
reproducible, but they bought **+26 px of distance** and fewer damage points,
|
||
not survival: incoming hit rate moved -1.19 pp, a fifth of its own 3.51 MDE.
|
||
"Arrive at a safe tile" turned out to mean "arrive further away".
|
||
4. **Diversity worsened, as the offline sweep warned.** Top-1 tile share
|
||
7.0% -> 10.6%, and the live losses concentrate against the opponents that
|
||
punish a long-range mover (SpinBot -26.2 damage with a -14.5 pp hit-rate
|
||
shift, the pattern guns -9 to -25). Outcomes are null-to-negative AND
|
||
diversity is worse: that is the argument against shipping, exactly the
|
||
pre-registered case.
|
||
5. **A null on wins lets us claim only "no large win".** It does not show the
|
||
knob is inert, and it does not retract the geometric measurement — but the
|
||
geometry measurement is not an argument for shipping when the live
|
||
consequence of it is measurably less damage from measurably further away.
|
||
|
||
**Ship state: `TR_TFIL_GEO_MODE` stays default `off`. Nothing changes.** Not
|
||
adopted, not adopted default-off. This is the **fifth** mechanism-positive /
|
||
outcome-not-positive result in the movement campaign (j144, j145, j146, j147,
|
||
j159) — and the first one where the mechanism is *anti*-correlated with damage.
|