Files
SirRoboGarage/docs/tfil_heat_pillar_ab.md

282 lines
15 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TFIL heat-time bullet model + virtual-pillar removal — LIVE A/B (negative)
> ## DECISION — the virtual centre pillar STAYS REMOVED
>
> **Decided by the bot's owner on 2026-09-25, after this A/B.** The recommendation
> below ("revert the default to pillar-ON") was **NOT adopted**. Rationale, in the
> owner's terms: the pillar is an *invented* hazard with no physical object behind
> it, and this test cannot distinguish "no benefit" from "small benefit" — the
> pillar contrast is **inside the minimum detectable effect** (32.6 dmg/run,
> 1.22 wins/run at n=10), and the raised damage-taken point estimate (p=0.040) is
> not corrected for the multiple arms tested.
>
> **So: do not "fix" this by reverting.** The shipped default keeps
> `PillarHotness = 0.0` / `PillarRadiance = 0.0`; `TR_TFIL_PILLAR_ON=1` restores the
> old field for any future experiment. If someone later wants to reverse this
> decision, it needs **its own pre-registered A/B at a sample size that can actually
> resolve a contrast this small** — not a re-reading of this one.
**Question.** Two movement changes shipped in HEAD — (a) the time-indexed bullet
heat (`TR_TFIL_HEAT_TIME`, job j105) and (b) the removal of the invented virtual
centre pillar (`TR_TFIL_PILLAR_ON=1` restores it, job j106) — judged on
**damage/run** and **ROUND WINS**, never on hit rate.
**Setup.** One frozen ModularBot from `git archive HEAD`
(commit `f58d65d2e8206cebd5b7d0a9465950dd9e6d2c28`, binary sha256
`74fd010ed1ffc4bb8a8de5569f58843479e1f50164a8a3f61e9c81cef108eaef`) vs the real
DrussGT through `tools/robocode_shim/run_bridge_battle.sh`: **6 arms × 10 runs ×
7 rounds = 60 battles, 420 rounds**, `--conc 7`, 0 failed. Raw per-tick captures
(~306 MB) live at `/tmp/ab/heatshift/` and are **not** committed; every per-run
value and every test below is in
`common_libs/tests/fixtures/tfil_heat_pillar_ab_results.json` (raw tool output in
`..._report.txt`, `..._mechanism.txt`, `..._bands.txt`).
`HEAD` already has the pillar REMOVED and heat-time OFF, so the PRE-change mover
is reconstructed with env: `old = TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0`.
> **THE METRIC RULE.** Movement arms are judged on **damage/run** and **round
> wins**. Hit rate and hits-taken are reported **as context only**, and neither
> decides anything here. This is not stylistic: in the previous TFIL A/B an arm
> took significantly FEWER hits (87.2 vs 96.5, p=0.0022) and still dealt the
> least damage and won the fewest rounds. It has happened **again below** —
> `tau3` has the *best* pooled hit rate of all six arms (11.56% vs `old` 10.99%)
> and the *fewest* round wins (20/70 vs 35/70). Hit rate would have ranked this
> experiment exactly backwards.
## 1. Live result table (10 runs × 7 rounds per arm vs real DrussGT) `[MEASURED]`
| arm | env | damage/run | damage taken/run | ROUND WINS | win% | shots/run | hits taken/run | hit rate (context) |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| **old** (pre-change mover) | `TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0` | **287** | **202** | **35/70** | **50.0%** | 783 | 93.5 | 10.99% |
| **pillaoff** (shipped default) | *(none)* | 282 | 232 | 30/70 | 42.9% | 729 | 95.0 | 11.14% |
| tau3 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3` | 261 | 299 | 20/70 | 28.6% | 666 | 106.2 | 11.56% |
| tau5 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5` | 250 | 294 | 18/70 | 25.7% | 689 | 105.3 | 10.83% |
| tau9 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9` | 265 | 231 | 22/70 | 31.4% | 750 | 99.5 | 10.51% |
| tau15 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15` | 265 | 185 | 32/70 | 45.7% | 812 | 89.5 | 9.93% |
`old` is the best arm on both verdict metrics (damage/run and round wins); only
`tau15` takes less damage/run, and it wins 3 fewer rounds and deals 22 less
damage per run. Every heat-time arm is worse or equal on both verdict columns.
The two "isolation" comparisons:
* **pillar**: `old` vs `pillaoff` — +5.7 damage/run, +0.5 wins/run, **−30.8
damage taken/run** (i.e. `pillaoff` takes MORE).
* **heat-time**: `pillaoff` vs `tau9` — +16.4 damage/run, +0.8 wins/run (i.e.
`tau9` is worse even against the pillar-removed control).
## 2. Per-run values `[MEASURED]`
| arm | damage/run by run (r1..r10) | wins by run (r1..r10) |
|---|---|---|
| old | 271 273 255 270 282 309 338 305 266 306 | 3 4 4 2 2 4 4 4 3 5 /7 |
| pillaoff | 207 332 263 332 286 280 244 269 311 294 | 1 5 2 4 4 3 2 2 4 3 /7 |
| tau3 | 213 296 182 285 283 220 290 304 245 293 | 1 3 1 3 3 1 4 1 2 1 /7 |
| tau5 | 272 235 241 236 281 295 248 243 218 225 | 2 2 2 1 3 1 2 2 2 1 /7 |
| tau9 | 295 283 275 239 268 258 280 262 228 266 | 3 3 2 2 2 3 2 3 1 1 /7 |
| tau15 | 266 264 258 239 278 291 230 286 283 261 | 3 4 3 2 4 4 2 4 4 2 /7 |
The tau3/tau5 win counts are *consistently* low (all 10 runs ≤3/7 for tau5, 9 of
10 ≤3/7 for tau3) — not one lucky bad run.
## 3. Tests vs `old` (per-run values, two-sided; exact permutation at 10v10) `[MEASURED]`
| metric | arm | diff (old − arm) | perm p | method | Mann-Whitney p | U |
|---|---|---:|---:|---|---:|---:|
| damage/run | pillaoff | +5.66 | 0.7096 | exact | 0.8501 | 47.0 |
| damage/run | tau3 | +26.34 | 0.1137 | exact | 0.3075 | 36.0 |
| damage/run | tau5 | +37.89 | **0.0042** | exact | **0.0091** | 15.0 |
| damage/run | tau9 | +22.06 | **0.0468** | exact | 0.1041 | 28.0 |
| damage/run | tau15 | +22.04 | **0.0460** | exact | 0.1041 | 28.0 |
| round wins | pillaoff | +0.50 | 0.4317 | exact | 0.3631 | 38.0 |
| round wins | tau3 | +1.50 | **0.0125** | exact | **0.0101** | 16.5 |
| round wins | tau5 | +1.70 | **0.0010** | exact | **0.0014** | 9.0 |
| round wins | tau9 | +1.30 | **0.0097** | exact | **0.0087** | 16.0 |
| round wins | tau15 | +0.30 | 0.6369 | exact | 0.5127 | 41.5 |
| damage taken/run | pillaoff | −30.82 | **0.0398** | exact | 0.1041 | — |
| damage taken/run | tau3 | −97.27 | **<0.0001** | exact | **0.0002** | — |
| damage taken/run | tau5 | −92.03 | **0.0004** | exact | **0.0022** | — |
| damage taken/run | tau9 | −29.59 | 0.1036 | exact | 0.2413 | — |
| damage taken/run | tau15 | +16.13 | 0.2213 | exact | 0.3075 | — |
Round-level pooled Fisher vs `old` (anti-conservative — rounds cluster within
runs): pillaoff 0.4980, tau3 0.0150, tau5 0.0051, tau9 0.0386, tau15 0.7352.
## 4. What the test can and cannot see `[MEASURED]`
MDE (two-sample, α=0.05 two-sided, 80% power, n=10/arm, from the `old` per-run SD):
| metric | sd(old) | MDE (absolute) | MDE vs `old` mean |
|---|---:|---:|---:|
| damage/run | 25.98 | **32.55** | 11.3% of 287.5 |
| round wins/run | 0.97 | **1.22** | 34.8% of 3.5 win/run |
* **Heat-time is a visible effect**: tau3/tau5/tau9 lose 1.3–1.7 wins/run, at or
above the 1.22 win MDE, with p≤0.013. tau5/tau9/tau15 lose 22–38 damage/run,
around the 32.55 damage MDE, p≤0.047.
* **The pillar result is NOT decidable at this n**: the whole observed `old`
advantage is +5.7 damage and +0.5 wins/run, both *well inside* the MDE. This
test only rules out the pillar removing ≥33 damage/run or ≥1.2 wins/run; it
cannot see anything smaller. The `pillaoff` damage-taken regression (−30.8/run,
p=0.040) is right at the MDE edge and its rank-sum cross-check is only
p=0.104, so treat it as a weak-but-consistent signal, not a proven loss.
## 5. Mechanism checks — did the knob actually change behaviour? `[MEASURED]`
Raw per-tick worldstate (both tanks' positions every tick) + the fire/hit event
sidecar. Aggregated per round, then averaged over rounds and runs.
### 5a. Central-box occupancy — the 144×144 px box the pillar covered (x 328–472, y 228–372)
Spawns are bottom-left (us) / top (enemy), never in the box, so occupancy is
genuine transit. **The pillar removal DID make us use the centre.**
| arm | ticks inside box | diff vs `old` | perm p |
|---|---:|---:|---:|
| old | 0.09% | — | — |
| pillaoff | 2.37% | +2.29pp | **<0.0001** |
| tau3 | 26.12% | +26.04pp | **<0.0001** |
| tau5 | 21.30% | +21.21pp | **<0.0001** |
| tau9 | 10.40% | +10.31pp | **<0.0001** |
| tau15 | 4.68% | +4.60pp | **<0.0001** |
MDE for this metric is 0.11pp, so all the shifts are real. Note the ordering:
`old` 0.09% → `pillaoff` 2.4% → `tau15` 4.7% → `tau9` 10.4% → `tau5` 21.3% →
`tau3` 26.1%. Removing the pillar opens the centre; the time-indexed heat (which
stops the bullet corridor at `speed·tau` instead of the wall) opens it much more,
and monotonically more as tau shrinks.
### 5b. Distance to the enemy (px, per-tick)
| arm | mean | p10 | median | p90 | 400+ px % of ticks | diff vs `old` | perm p |
|---|---:|---:|---:|---:|---:|---:|---:|
| old | 469.3 | 433 | 465 | 511 | 83.0 | — | — |
| pillaoff | 443.5 | 396 | 446 | 490 | 72.1 | −25.7 | **0.0002** |
| tau3 | 439.4 | 406 | 440 | 471 | 70.4 | −29.9 | **0.0004** |
| tau5 | 447.8 | 410 | 451 | 484 | 73.9 | −21.4 | **0.0014** |
| tau9 | 481.1 | 440 | 482 | 512 | 85.5 | +11.8 | 0.0813 |
| tau15 | 500.1 | 467 | 502 | 533 | 88.2 | +30.8 | **0.0001** |
MDE = 20.8 px. `old` fights at ~469 px; every opening of the centre pulls the
engagement 21–30 px closer (and 9–13pp more of the battle is inside 400 px), and
tau15 pushes it 31 px further out. `tau9` (the shipped tau) is essentially `old`
(+11.8 px, p=0.08).
### 5c. Bullet proximity — how much time we spend near live bullet paths
Nearest **live enemy** bullet (our own bullets are excluded: a bullet is born at
its own tank, so "any bullet" is dominated by our own just-fired shot; the
any-bullet version, per the task text, is in the ANY-BULLET PROXIMITY table of
`common_libs/tests/fixtures/tfil_heat_pillar_ab_mechanism.txt`).
| arm | ≤50 px | ≤100 px | ≤150 px | mean min-dist px | ≤100px diff vs `old` | perm p |
|---|---:|---:|---:|---:|---:|---:|
| old | 4.76% | 27.99% | 62.48% | 136.1 | — | — |
| pillaoff | 4.97% | 29.10% | 64.98% | 132.8 | +1.11pp | 0.1381 |
| tau3 | 6.19% | 30.86% | 61.39% | 134.8 | +2.87pp | **0.0014** |
| tau5 | 5.88% | 30.60% | 62.59% | 134.1 | +2.61pp | **0.0005** |
| tau9 | 4.99% | 28.31% | 60.75% | 137.8 | +0.32pp | 0.6066 |
| tau15 | 4.12% | 25.57% | 58.43% | 141.4 | −2.43pp | **0.0001** |
MDE = 1.74pp. tau3/tau5 spend significantly more time within 100 px of a live
enemy bullet; tau15 significantly less.
## 6. Liveness `[MEASURED]`
Every arm's declared env appears verbatim in the bot's own boot report
(`<arm>/run<N>.bot.stdout.log`, section A "raw process environment") in all 10
runs; no arm env was ignored as unrecognised. This is a *measurement*, not a
skip:
```
old OK (10/10 runs: TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 applied)
pillaoff OK (10/10 runs: no arm env; report present)
tau3 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 applied)
tau5 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 applied)
tau9 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 applied)
tau15 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 applied)
```
The effective-value section confirms the reconstruction: `old` reports
`TR_TFIL_PILLAR_ON = on (source: env)` / `TR_TFIL_HEAT_TIME = off`, `pillaoff`
reports `TR_TFIL_PILLAR_ON = off (source: default)` / `TR_TFIL_HEAT_TIME = off
(source: default)`, and each tau arm reports the requested tau
(`TR_TFIL_HEAT_TAU = 3.0/5.0/9.0/15.0 (source: env)`).
## 7. Hit-rate context (NOT a verdict metric) `[MEASURED]`
| arm | pooled hit rate | round wins | our hits taken/run | our shots/run |
|---|---:|---:|---:|---:|
| old | 10.99% | 35/70 | 93.5 | 783 |
| pillaoff | 11.14% | 30/70 | 95.0 | 729 |
| tau3 | **11.56%** (best) | **20/70** (worst) | 106.2 | 666 |
| tau5 | 10.83% | 18/70 | 105.3 | 689 |
| tau9 | 10.51% | 22/70 | 99.5 | 750 |
| tau15 | **9.93%** (worst) | **32/70** (2nd best) | 89.5 | 812 |
The inversion is exact at both ends: the best-accuracy arm (`tau3`, 11.56%) wins
the fewest rounds; the worst-accuracy arm (`tau15`, 9.93%) wins the second most.
Hit rate is monotone in the WRONG direction here. Range-banded hit rates
(`..._bands.txt`) are flat in the 300–450 px band that holds 40% of our shots
(11.3–13.4%) and in 450+ (8.8–10.1%); no band rescues any heat-time arm.
## 8. Direct answers
**Does the heat-time model help, hurt, or do nothing? `[MEASURED] → HURTS.`**
* At the shipped tau (9) and at 3/5 it **significantly reduces round wins**
(20–22/70 vs 35/70, p=0.0010–0.0125) and reduces damage/run (p=0.004–0.047).
* `tau15` is a **wash on wins** (32/70, p=0.637) and 22 damage/run **lower**
(p=0.046) — not better, mildly worse.
* No tau improves either verdict metric. The mechanism is exactly what the
change claims (shorter corridor → centre opens, engagement closes by 20–30 px,
more time within 100 px of a live enemy bullet) — and the closer arms take
≈12 more hits/run (tau3 12.7, tau5 11.8) and lose 13–17 more rounds per 10
runs. **Inferred:** the free space this frees is the arena centre, and standing
there costs more than the wall-hugging flat model costs. The offline
"largest safe region" ruler rewarded exactly the region that shorter tau opens
(tau 2 → 162 px, tau 9 → 140 px, tau 15 → 127 px), and that region is a proxy
anti-correlated with survival against DrussGT.
**Does removing the pillar help, hurt, or do nothing? `[MEASURED] → no
measurable benefit; point estimates are worse, and it is undecidable at this n.`**
* `old` vs `pillaoff`: damage +5.7/run (p=0.710), wins +0.5/run (35 vs 30,
p=0.432), round-level Fisher p=0.498 — **no significant difference**.
* Damage taken is **30.8/run higher** without the pillar (p=0.040 permutation;
rank-sum cross-check p=0.104).
* The mechanism check proves the change **did** alter movement: box occupancy
0.09% → 2.37% (p<0.0001), mean range 469 → 443 px (p=0.0002). So this is a real
behaviour change, not a dead knob — it simply does not buy anything.
## 9. Which arm should be the shipped default?
**`old` — the pre-change mover (`TR_TFIL_PILLAR_ON=1` behaviour, heat-time off).
Nothing beats it: it is the best arm on damage/run (287) and round wins
(35/70), and second only to `tau15` on damage taken/run (202 vs 185), where
`tau15` pays for its 3 fewer round wins and 22 less damage per run.**
* **Heat-time: keep it OFF (as shipped).** The default is already off (`fca8993`);
no action needed, and `TR_TFIL_HEAT_TIME=1` at any tested tau should stay off.
* **Pillar removal: this test recommends reverting the default** — see the **DECISION** banner at the top of this document: the owner **overruled** this recommendation and the pillar **stays removed**. The shipped default
since `d0750ab` is `pillaoff`; it does **not** beat `old` on any verdict metric
and is directionally worse on all three (Δdamage −5.7, Δwins −0.5, Δdamage
taken +30.8). This is a decision for the user: restore the pre-change default
(`PillarHotness`/`PillarRadiance` back to 30/10, i.e. the `TR_TFIL_PILLAR_ON`
behaviour by default, keeping the env knob for the off state). **Honesty
caveat:** at n=10/arm the pillar contrast is inside the MDE (33 damage/run,
1.22 wins/run), so this is not a statistically significant "removal is worse";
it is "removal bought nothing measurable, and every point estimate moved the
wrong way". The case for reverting is the absence of evidence of benefit plus
parsimony, not a proven regression.
**Reproduce** (raw captures are not committed):
```sh
tools/ab/ab_run.sh --arms tools/ab/arms_heat_pillar.txt --runs 10 --rounds 7 \
--conc 7 --outdir /tmp/ab/heatshift
python3 tools/ab/ab_analyze.py /tmp/ab/heatshift --reference old
python3 tools/ab/ab_mechanism.py /tmp/ab/heatshift --reference old
python3 tools/ab/ab_range_bands.py /tmp/ab/heatshift --reference old
```