Files
SirRoboGarage/docs/tfil_geo_ab.md
T

210 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# j159 — PRE-REGISTRATION: does tile-geometry weighting in the tfil picker win?
**Written and committed BEFORE a single battle of this experiment ran.** No
result in the "MEASURED" section below existed when this section was written.
## The question
`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` (shipped default `off`, j152) shape the
**draw** over the safe tiles the tfil picker chooses from. The offline sweep on
recorded fixtures predicted a real geometric improvement —
| offline metric (`measure_tfil_pick_defects`) | `off` | `both-rej`, tau 60 |
|---|---:|---:|
| REACH / arrival within feasible time | 4.5% | 29.4% |
| hot on arrival (`hotAtTta`) | 31.0% | 24.0% |
| **top-1 tile share (pick DIVERSITY)** | **7.0%** | **10.6%** |
— and a **diversity cost**, because the geometry weight concentrates the draw.
The offline ruler replays the real picker on real fixtures; it says nothing
about whether a tile that is geometrically reachable and less hot actually wins
a round. That is what this run measures.
**Hypothesis H1.** Weighting the draw by tile geometry (`both-rej`, tau 60)
raises damage/run and round-win rate over the shipped uniform draw, because the
mover arrives at safe tiles instead of merely picking them.
**Direction is pre-registered as two-sided.** A regression is as interesting as
a win (the diversity cost makes one plausible) and re-deciding the direction
after seeing the data is exactly what this document exists to prevent.
## Arms — identical except the geo knob
Both arms: `TR_MOVEMENT=tfil`, everything else at the shipped defaults, one
frozen binary built once from `git archive HEAD` (session 2223ca6).
| arm | env | role |
|---|---|---|
| `A_off` | `TR_TFIL_GEO_MODE=off` (`TAU=0`) | **REFERENCE** — the shipped uniform draw |
| `B_geo` | `TR_TFIL_GEO_MODE=both-rej` `TR_TFIL_GEO_TAU=60` | treatment |
**Contamination control (the main risk).** The owner has `both-rej` in a personal
`.env` (`ModularBot_garage/out/.env`, a copy at `tr_bots/ModularBot_geo/.env`),
and this bot's dotenv loader gives the FILE priority over shell exports. If the
tournament's bot instances resolved that file, arm A would silently become arm B
and the whole run would be void. Three guarantees, all verifiable from the logs:
1. each arm is launched with `TR_ENV_FILE` pointing at a **per-arm file this
job generated** (`/tmp/j159_geo/env/A_off.env`, `/tmp/j159_geo/env/B_geo.env`)
in this job's own outdir, so the ONLY `.env` the loader can resolve is mine;
2. the tournament's per-run botdir (`$OUTDIR/.work/<opp>/<arm>/run<N>/bots/ModularBot`)
contains only `ModularBot.json`, `ModularBot.sh` and a symlink to the frozen
binary — **no `.env`**, and the loader's fallback is `./.env` then `.env` next
to the executable (it does not walk up parent directories), so the owner's file
is not reachable;
3. every single run's `[env]` boot report is checked for its intended
`TR_TFIL_GEO_MODE` / `TR_TFIL_GEO_TAU` before any number is read. **Any run
whose `[env]` disagrees with its arm invalidates the session** and the run is
reported as void rather than analysed.
## Design
* Harness: `tools/ab/tournament_run.sh` + `tools/ab/tournament_analyze.py`
(unmodified).
* Panel: `tools/ab/panel_movement.txt` — the **FROZEN 15-opponent movement
panel**, unchanged. Unit of evidence is the opponent, not the battle.
* 15 opponents x 2 arms x **14 runs** x 3 rounds = **420 battles**.
* Battles serialised: `--wait-arena 45`, one session at a time.
### Primary metrics (pre-registered, fixed)
1. **damage/run** (our damage dealt per run)
2. **round-win rate** (rounds won / rounds fought)
**Hit rate is NOT a primary metric** — it hid a survival regression once already.
### Secondary / mechanism (reported, never a verdict)
* incoming hit rate (the survival channel the mechanism actually runs through);
* damage taken/run;
* the offline geometric numbers above (REACH%, hot-on-arrival%, top-1 tile
share). The live battle logs do not contain the per-pick tile or the arrival
state, so the live run **cannot** re-measure them; that is stated in the
verdict rather than papered over. No new instrumentation is built for this.
### Statistical treatment
Same as every previous movement gate: per-opponent paired deltas (arm −
reference), mean delta, SD, SE, 95% CI, a sign test and a **sign-flip
permutation test** (exact when `2^n <= 2^20`, else Monte-Carlo), Wilcoxon as a
cross-check, plus the MDE the analyzer reports for the reference arm's n.
### MDE — stated up front, and it is LARGE
At **14 runs/arm** over the frozen 15-opponent panel this design resolves about
**0.28 wins/run** (and the corresponding damage/run MDE the analyzer prints).
A two-arm run is 420 battles, ~1 hour. Resolving **0.10 wins/run** would need
~2.2 h and ~2,900 battles — **which we are NOT doing.**
**Consequences, recorded before any data:**
* **A null is the likely outcome.** This would be the *fifth* consecutive
mechanism-positive / outcome-null result in this campaign (after j144, j145,
j146, j147).
* A null here **excludes only a LARGE effect** (>= ~0.28 wins/run). It does not
show the knob does nothing, and it does not retract the offline geometric
measurement.
* Because the offline sweep also measured a **diversity regression**
(top-1 tile share 7.0% -> 10.6%), a null combined with a confirmed diversity
cost is an argument **against** shipping, not for it.
### Verdict rule (fixed now, not re-read later)
* **Adopt** only if BOTH primaries move in B's favour with `p(sign-flip) < 0.05`
and the effect is at or above the reported MDE. One primary at p<0.05 with
the other not down is reported as a partial signal, not a win.
* Otherwise **do not ship**; the knob stays default `off`.
* The mechanism is reported as measured, with no vote in the verdict.
* No subsetting, no dropping opponents, no re-running to chase a p-value. A
clean null is a fully acceptable result.
---
## MEASURED
*(appended after the battles — everything above was committed first)*
### MEASURED — the live A/B, 420 battles (j159)
* **Provenance.** Session `/tmp/ab/j159_geo`, commit `7c5bc9c`, frozen binary
sha256 `9f116e7a9eb9…`, panel `tools/ab/panel_movement.txt` (FROZEN, 15
opponents), 15 x 2 x **14 runs** x 3 rounds = **420 battles, 0 failed, 0 never
started, 1110 s**. Per-arm env files `/tmp/j159_geo/env/{A_off,B_geo}.env`.
* **Env verification.** All **420** runs carry their intended arm: 210/210
`A_off` show `TR_TFIL_GEO_MODE=off` / `TR_TFIL_GEO_TAU=0` (parsed `off`/`0.0`),
210/210 `B_geo` show `both-rej`/`60` (parsed `both`/`60.0`), every run reports
`env file: /tmp/j159_geo/env/<arm>.env (source: TR_ENV_FILE)` and
`move.effective = tfil`. No `.env` exists anywhere in the session dir, and the
loader's fallbacks are `./.env` then `.env` next to the executable — it does
not walk up parents — so the owner's file is unreachable. **0 runs mis-set.**
* **Record correction (no battle re-run).** The arms file declared only
`TR_ENV_FILE`, which the analyzer's liveness guard reads from `session.json`
and treats as an undeclared `TR_MOVEMENT` (fatal contamination). `session.json`
and a corrected arms file were rewritten to declare the effective env — the
original is kept as `/tmp/j159_geo/{session.json.orig,arms_tfil_geo.txt.orig}`.
The guard then re-verified all 420 declared values verbatim in the boot
reports: `liveness: 0 run(s) excluded (420 total)`.
**Pooled dashboard (descriptive, NOT the verdict):**
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `A_off` | 210 | 112.5 | 195.2 | 1.21 | 255/630 | 40.5% | 17.64% | 393 |
| `B_geo` | 210 | 103.7 | 188.0 | 1.16 | 244/630 | 38.7% | 17.42% | 419 |
**Verdict layer** (per-opponent paired deltas, arm − reference; sign-flip is
the exact 2^15 permutation the pre-registration names as the decision test):
| arm | metric | mean Δ | 95% CI | sign test | p(sign) | **p(sign-flip)** | Wilcoxon p | MDE |
|---|---|---:|---|---:|---:|---:|---:|---:|
| `B_geo` | **damage/run** | **-8.83** | [-14.69, -2.97] | 5/15 | 0.3018 | **0.006104** | 0.0115 | 7.65 |
| `B_geo` | round wins | -0.05 | [-0.18, +0.08] | 4/11 | 0.5488 | 0.4619 | 0.3496 | 0.17 |
| `B_geo` | damage taken | -7.24 | [-20.84, +6.36] | 8/15 | 1 | 0.2786 | 0.4777 | 17.77 |
| `B_geo` | incoming hit rate | -1.19 pp | [-3.88, +1.51] | 8/15 | 1 | 0.3962 | 0.5895 | 3.51 |
| `B_geo` | mean distance | **+26.3 px** | [+16.3, +36.4] | **15/15** | 6.1e-05 | 6.1e-05 | 0.0007 | 13.15 |
**Per-opponent damage/run** (the pattern/ram rows carry the loss): Coriantumr
-25.1, CassiusClay -24.2, SpinBot -26.2, WallAvoider -13.3, BlitzBat -11.2,
HawkOnFire -11.7, Diamond -10.9, TripHammer -9.0, YersiniaPestis -8.9,
Dookious -7.3 vs GresSuffurd +2.3, DiamondStealer +3.2, Ascendant +3.7,
DrussGT +0.1. Round wins: HawkOnFire +0.50 and WallAvoider +0.14 (both closer
opponents) against Coriantumr -0.57, TripHammer -0.21, YersiniaPestis -0.21.
**Mechanism, offline (the committed ruler, re-run unchanged for this doc):**
`measure_tfil_pick_defects` reproduces the pre-registered prediction exactly —
REACH **4.5% -> 29.4%**, hot-on-arrival **31.0% -> 24.0%**, and the diversity cost
**top-1 tile share 7.0% -> 10.6%** (normalised entropy 0.87 -> 0.85). The live
mean distance **+26 px on 15/15 opponents** is the same mechanism seen end to
end: the geometry weight prefers tiles that are far better to *arrive* in, and
the bot sits further out and deals **less** damage.
### VERDICT — DO NOT ADOPT
1. **H1 is rejected.** Round wins are flat (-0.05/run, p(sign-flip) = 0.46, under
the 0.17 MDE) and damage/run is **down 8.83** (p(sign-flip) = 0.0061, above
the 7.65 MDE, Wilcoxon p = 0.011). The pre-registration required BOTH
primaries up; one is down and significant by the test it named.
2. **The MDE, restated.** 14 runs/arm on the frozen 15-opponent panel resolves
**0.17 wins/run** and **7.65 damage/run** (better than the 0.28 pre-registered
estimate). So this run excludes a large *benefit*; it also positively measures
a small *harm* in damage. It says nothing about effects below those numbers.
3. **The mechanism moved exactly as predicted, and that is what makes it bad.**
REACH/arrival 4.5% -> 29.4% and hot-on-arrival 31.0% -> 24.0% are real and
reproducible, but they bought **+26 px of distance** and fewer damage points,
not survival: incoming hit rate moved -1.19 pp, a fifth of its own 3.51 MDE.
"Arrive at a safe tile" turned out to mean "arrive further away".
4. **Diversity worsened, as the offline sweep warned.** Top-1 tile share
7.0% -> 10.6%, and the live losses concentrate against the opponents that
punish a long-range mover (SpinBot -26.2 damage with a -14.5 pp hit-rate
shift, the pattern guns -9 to -25). Outcomes are null-to-negative AND
diversity is worse: that is the argument against shipping, exactly the
pre-registered case.
5. **A null on wins lets us claim only "no large win".** It does not show the
knob is inert, and it does not retract the geometric measurement — but the
geometry measurement is not an argument for shipping when the live
consequence of it is measurably less damage from measurably further away.
**Ship state: `TR_TFIL_GEO_MODE` stays default `off`. Nothing changes.** Not
adopted, not adopted default-off. This is the **fifth** mechanism-positive /
outcome-not-positive result in the movement campaign (j144, j145, j146, j147,
j159) — and the first one where the mechanism is *anti*-correlated with damage.