TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover

Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
This commit is contained in:
2026-09-25 22:21:41 +02:00
parent f58d65d2e8
commit 48f38b80e7
7 changed files with 3121 additions and 0 deletions
+264
View File
@@ -0,0 +1,264 @@
# TFIL heat-time bullet model + virtual-pillar removal — LIVE A/B (negative)
**Question.** Two movement changes shipped in HEAD — (a) the time-indexed bullet
heat (`TR_TFIL_HEAT_TIME`, job j105) and (b) the removal of the invented virtual
centre pillar (`TR_TFIL_PILLAR_ON=1` restores it, job j106) — judged on
**damage/run** and **ROUND WINS**, never on hit rate.
**Setup.** One frozen ModularBot from `git archive HEAD`
(commit `f58d65d2e8206cebd5b7d0a9465950dd9e6d2c28`, binary sha256
`74fd010ed1ffc4bb8a8de5569f58843479e1f50164a8a3f61e9c81cef108eaef`) vs the real
DrussGT through `tools/robocode_shim/run_bridge_battle.sh`: **6 arms × 10 runs ×
7 rounds = 60 battles, 420 rounds**, `--conc 7`, 0 failed. Raw per-tick captures
(~306 MB) live at `/tmp/ab/heatshift/` and are **not** committed; every per-run
value and every test below is in
`common_libs/tests/fixtures/tfil_heat_pillar_ab_results.json` (raw tool output in
`..._report.txt`, `..._mechanism.txt`, `..._bands.txt`).
`HEAD` already has the pillar REMOVED and heat-time OFF, so the PRE-change mover
is reconstructed with env: `old = TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0`.
> **THE METRIC RULE.** Movement arms are judged on **damage/run** and **round
> wins**. Hit rate and hits-taken are reported **as context only**, and neither
> decides anything here. This is not stylistic: in the previous TFIL A/B an arm
> took significantly FEWER hits (87.2 vs 96.5, p=0.0022) and still dealt the
> least damage and won the fewest rounds. It has happened **again below** —
> `tau3` has the *best* pooled hit rate of all six arms (11.56% vs `old` 10.99%)
> and the *fewest* round wins (20/70 vs 35/70). Hit rate would have ranked this
> experiment exactly backwards.
## 1. Live result table (10 runs × 7 rounds per arm vs real DrussGT) `[MEASURED]`
| arm | env | damage/run | damage taken/run | ROUND WINS | win% | shots/run | hits taken/run | hit rate (context) |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| **old** (pre-change mover) | `TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0` | **287** | **202** | **35/70** | **50.0%** | 783 | 93.5 | 10.99% |
| **pillaoff** (shipped default) | *(none)* | 282 | 232 | 30/70 | 42.9% | 729 | 95.0 | 11.14% |
| tau3 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3` | 261 | 299 | 20/70 | 28.6% | 666 | 106.2 | 11.56% |
| tau5 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5` | 250 | 294 | 18/70 | 25.7% | 689 | 105.3 | 10.83% |
| tau9 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9` | 265 | 231 | 22/70 | 31.4% | 750 | 99.5 | 10.51% |
| tau15 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15` | 265 | 185 | 32/70 | 45.7% | 812 | 89.5 | 9.93% |
`old` is the best arm on both verdict metrics (damage/run and round wins); only
`tau15` takes less damage/run, and it wins 3 fewer rounds and deals 22 less
damage per run. Every heat-time arm is worse or equal on both verdict columns.
The two "isolation" comparisons:
* **pillar**: `old` vs `pillaoff` — +5.7 damage/run, +0.5 wins/run, **−30.8
damage taken/run** (i.e. `pillaoff` takes MORE).
* **heat-time**: `pillaoff` vs `tau9` — +16.4 damage/run, +0.8 wins/run (i.e.
`tau9` is worse even against the pillar-removed control).
## 2. Per-run values `[MEASURED]`
| arm | damage/run by run (r1..r10) | wins by run (r1..r10) |
|---|---|---|
| old | 271 273 255 270 282 309 338 305 266 306 | 3 4 4 2 2 4 4 4 3 5 /7 |
| pillaoff | 207 332 263 332 286 280 244 269 311 294 | 1 5 2 4 4 3 2 2 4 3 /7 |
| tau3 | 213 296 182 285 283 220 290 304 245 293 | 1 3 1 3 3 1 4 1 2 1 /7 |
| tau5 | 272 235 241 236 281 295 248 243 218 225 | 2 2 2 1 3 1 2 2 2 1 /7 |
| tau9 | 295 283 275 239 268 258 280 262 228 266 | 3 3 2 2 2 3 2 3 1 1 /7 |
| tau15 | 266 264 258 239 278 291 230 286 283 261 | 3 4 3 2 4 4 2 4 4 2 /7 |
The tau3/tau5 win counts are *consistently* low (all 10 runs ≤3/7 for tau5, 9 of
10 ≤3/7 for tau3) — not one lucky bad run.
## 3. Tests vs `old` (per-run values, two-sided; exact permutation at 10v10) `[MEASURED]`
| metric | arm | diff (old − arm) | perm p | method | Mann-Whitney p | U |
|---|---|---:|---:|---|---:|---:|
| damage/run | pillaoff | +5.66 | 0.7096 | exact | 0.8501 | 47.0 |
| damage/run | tau3 | +26.34 | 0.1137 | exact | 0.3075 | 36.0 |
| damage/run | tau5 | +37.89 | **0.0042** | exact | **0.0091** | 15.0 |
| damage/run | tau9 | +22.06 | **0.0468** | exact | 0.1041 | 28.0 |
| damage/run | tau15 | +22.04 | **0.0460** | exact | 0.1041 | 28.0 |
| round wins | pillaoff | +0.50 | 0.4317 | exact | 0.3631 | 38.0 |
| round wins | tau3 | +1.50 | **0.0125** | exact | **0.0101** | 16.5 |
| round wins | tau5 | +1.70 | **0.0010** | exact | **0.0014** | 9.0 |
| round wins | tau9 | +1.30 | **0.0097** | exact | **0.0087** | 16.0 |
| round wins | tau15 | +0.30 | 0.6369 | exact | 0.5127 | 41.5 |
| damage taken/run | pillaoff | −30.82 | **0.0398** | exact | 0.1041 | — |
| damage taken/run | tau3 | −97.27 | **<0.0001** | exact | **0.0002** | — |
| damage taken/run | tau5 | −92.03 | **0.0004** | exact | **0.0022** | — |
| damage taken/run | tau9 | −29.59 | 0.1036 | exact | 0.2413 | — |
| damage taken/run | tau15 | +16.13 | 0.2213 | exact | 0.3075 | — |
Round-level pooled Fisher vs `old` (anti-conservative — rounds cluster within
runs): pillaoff 0.4980, tau3 0.0150, tau5 0.0051, tau9 0.0386, tau15 0.7352.
## 4. What the test can and cannot see `[MEASURED]`
MDE (two-sample, α=0.05 two-sided, 80% power, n=10/arm, from the `old` per-run SD):
| metric | sd(old) | MDE (absolute) | MDE vs `old` mean |
|---|---:|---:|---:|
| damage/run | 25.98 | **32.55** | 11.3% of 287.5 |
| round wins/run | 0.97 | **1.22** | 34.8% of 3.5 win/run |
* **Heat-time is a visible effect**: tau3/tau5/tau9 lose 1.3–1.7 wins/run, at or
above the 1.22 win MDE, with p≤0.013. tau5/tau9/tau15 lose 22–38 damage/run,
around the 32.55 damage MDE, p≤0.047.
* **The pillar result is NOT decidable at this n**: the whole observed `old`
advantage is +5.7 damage and +0.5 wins/run, both *well inside* the MDE. This
test only rules out the pillar removing ≥33 damage/run or ≥1.2 wins/run; it
cannot see anything smaller. The `pillaoff` damage-taken regression (−30.8/run,
p=0.040) is right at the MDE edge and its rank-sum cross-check is only
p=0.104, so treat it as a weak-but-consistent signal, not a proven loss.
## 5. Mechanism checks — did the knob actually change behaviour? `[MEASURED]`
Raw per-tick worldstate (both tanks' positions every tick) + the fire/hit event
sidecar. Aggregated per round, then averaged over rounds and runs.
### 5a. Central-box occupancy — the 144×144 px box the pillar covered (x 328–472, y 228–372)
Spawns are bottom-left (us) / top (enemy), never in the box, so occupancy is
genuine transit. **The pillar removal DID make us use the centre.**
| arm | ticks inside box | diff vs `old` | perm p |
|---|---:|---:|---:|
| old | 0.09% | — | — |
| pillaoff | 2.37% | +2.29pp | **<0.0001** |
| tau3 | 26.12% | +26.04pp | **<0.0001** |
| tau5 | 21.30% | +21.21pp | **<0.0001** |
| tau9 | 10.40% | +10.31pp | **<0.0001** |
| tau15 | 4.68% | +4.60pp | **<0.0001** |
MDE for this metric is 0.11pp, so all the shifts are real. Note the ordering:
`old` 0.09% → `pillaoff` 2.4% → `tau15` 4.7% → `tau9` 10.4% → `tau5` 21.3% →
`tau3` 26.1%. Removing the pillar opens the centre; the time-indexed heat (which
stops the bullet corridor at `speed·tau` instead of the wall) opens it much more,
and monotonically more as tau shrinks.
### 5b. Distance to the enemy (px, per-tick)
| arm | mean | p10 | median | p90 | 400+ px % of ticks | diff vs `old` | perm p |
|---|---:|---:|---:|---:|---:|---:|---:|
| old | 469.3 | 433 | 465 | 511 | 83.0 | — | — |
| pillaoff | 443.5 | 396 | 446 | 490 | 72.1 | −25.7 | **0.0002** |
| tau3 | 439.4 | 406 | 440 | 471 | 70.4 | −29.9 | **0.0004** |
| tau5 | 447.8 | 410 | 451 | 484 | 73.9 | −21.4 | **0.0014** |
| tau9 | 481.1 | 440 | 482 | 512 | 85.5 | +11.8 | 0.0813 |
| tau15 | 500.1 | 467 | 502 | 533 | 88.2 | +30.8 | **0.0001** |
MDE = 20.8 px. `old` fights at ~469 px; every opening of the centre pulls the
engagement 21–30 px closer (and 9–13pp more of the battle is inside 400 px), and
tau15 pushes it 31 px further out. `tau9` (the shipped tau) is essentially `old`
(+11.8 px, p=0.08).
### 5c. Bullet proximity — how much time we spend near live bullet paths
Nearest **live enemy** bullet (our own bullets are excluded: a bullet is born at
its own tank, so "any bullet" is dominated by our own just-fired shot; the
any-bullet version, per the task text, is in the ANY-BULLET PROXIMITY table of
`common_libs/tests/fixtures/tfil_heat_pillar_ab_mechanism.txt`).
| arm | ≤50 px | ≤100 px | ≤150 px | mean min-dist px | ≤100px diff vs `old` | perm p |
|---|---:|---:|---:|---:|---:|---:|
| old | 4.76% | 27.99% | 62.48% | 136.1 | — | — |
| pillaoff | 4.97% | 29.10% | 64.98% | 132.8 | +1.11pp | 0.1381 |
| tau3 | 6.19% | 30.86% | 61.39% | 134.8 | +2.87pp | **0.0014** |
| tau5 | 5.88% | 30.60% | 62.59% | 134.1 | +2.61pp | **0.0005** |
| tau9 | 4.99% | 28.31% | 60.75% | 137.8 | +0.32pp | 0.6066 |
| tau15 | 4.12% | 25.57% | 58.43% | 141.4 | −2.43pp | **0.0001** |
MDE = 1.74pp. tau3/tau5 spend significantly more time within 100 px of a live
enemy bullet; tau15 significantly less.
## 6. Liveness `[MEASURED]`
Every arm's declared env appears verbatim in the bot's own boot report
(`<arm>/run<N>.bot.stdout.log`, section A "raw process environment") in all 10
runs; no arm env was ignored as unrecognised. This is a *measurement*, not a
skip:
```
old OK (10/10 runs: TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 applied)
pillaoff OK (10/10 runs: no arm env; report present)
tau3 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 applied)
tau5 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 applied)
tau9 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 applied)
tau15 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 applied)
```
The effective-value section confirms the reconstruction: `old` reports
`TR_TFIL_PILLAR_ON = on (source: env)` / `TR_TFIL_HEAT_TIME = off`, `pillaoff`
reports `TR_TFIL_PILLAR_ON = off (source: default)` / `TR_TFIL_HEAT_TIME = off
(source: default)`, and each tau arm reports the requested tau
(`TR_TFIL_HEAT_TAU = 3.0/5.0/9.0/15.0 (source: env)`).
## 7. Hit-rate context (NOT a verdict metric) `[MEASURED]`
| arm | pooled hit rate | round wins | our hits taken/run | our shots/run |
|---|---:|---:|---:|---:|
| old | 10.99% | 35/70 | 93.5 | 783 |
| pillaoff | 11.14% | 30/70 | 95.0 | 729 |
| tau3 | **11.56%** (best) | **20/70** (worst) | 106.2 | 666 |
| tau5 | 10.83% | 18/70 | 105.3 | 689 |
| tau9 | 10.51% | 22/70 | 99.5 | 750 |
| tau15 | **9.93%** (worst) | **32/70** (2nd best) | 89.5 | 812 |
The inversion is exact at both ends: the best-accuracy arm (`tau3`, 11.56%) wins
the fewest rounds; the worst-accuracy arm (`tau15`, 9.93%) wins the second most.
Hit rate is monotone in the WRONG direction here. Range-banded hit rates
(`..._bands.txt`) are flat in the 300–450 px band that holds 40% of our shots
(11.3–13.4%) and in 450+ (8.8–10.1%); no band rescues any heat-time arm.
## 8. Direct answers
**Does the heat-time model help, hurt, or do nothing? `[MEASURED] → HURTS.`**
* At the shipped tau (9) and at 3/5 it **significantly reduces round wins**
(20–22/70 vs 35/70, p=0.0010–0.0125) and reduces damage/run (p=0.004–0.047).
* `tau15` is a **wash on wins** (32/70, p=0.637) and 22 damage/run **lower**
(p=0.046) — not better, mildly worse.
* No tau improves either verdict metric. The mechanism is exactly what the
change claims (shorter corridor → centre opens, engagement closes by 20–30 px,
more time within 100 px of a live enemy bullet) — and the closer arms take
≈12 more hits/run (tau3 12.7, tau5 11.8) and lose 13–17 more rounds per 10
runs. **Inferred:** the free space this frees is the arena centre, and standing
there costs more than the wall-hugging flat model costs. The offline
"largest safe region" ruler rewarded exactly the region that shorter tau opens
(tau 2 → 162 px, tau 9 → 140 px, tau 15 → 127 px), and that region is a proxy
anti-correlated with survival against DrussGT.
**Does removing the pillar help, hurt, or do nothing? `[MEASURED] → no
measurable benefit; point estimates are worse, and it is undecidable at this n.`**
* `old` vs `pillaoff`: damage +5.7/run (p=0.710), wins +0.5/run (35 vs 30,
p=0.432), round-level Fisher p=0.498 — **no significant difference**.
* Damage taken is **30.8/run higher** without the pillar (p=0.040 permutation;
rank-sum cross-check p=0.104).
* The mechanism check proves the change **did** alter movement: box occupancy
0.09% → 2.37% (p<0.0001), mean range 469 → 443 px (p=0.0002). So this is a real
behaviour change, not a dead knob — it simply does not buy anything.
## 9. Which arm should be the shipped default?
**`old` — the pre-change mover (`TR_TFIL_PILLAR_ON=1` behaviour, heat-time off).
Nothing beats it: it is the best arm on damage/run (287) and round wins
(35/70), and second only to `tau15` on damage taken/run (202 vs 185), where
`tau15` pays for its 3 fewer round wins and 22 less damage per run.**
* **Heat-time: keep it OFF (as shipped).** The default is already off (`fca8993`);
no action needed, and `TR_TFIL_HEAT_TIME=1` at any tested tau should stay off.
* **Pillar removal: the shipped default should be REVERTED.** The shipped default
since `d0750ab` is `pillaoff`; it does **not** beat `old` on any verdict metric
and is directionally worse on all three (Δdamage −5.7, Δwins −0.5, Δdamage
taken +30.8). This is a decision for the user: restore the pre-change default
(`PillarHotness`/`PillarRadiance` back to 30/10, i.e. the `TR_TFIL_PILLAR_ON`
behaviour by default, keeping the env knob for the off state). **Honesty
caveat:** at n=10/arm the pillar contrast is inside the MDE (33 damage/run,
1.22 wins/run), so this is not a statistically significant "removal is worse";
it is "removal bought nothing measurable, and every point estimate moved the
wrong way". The case for reverting is the absence of evidence of benefit plus
parsimony, not a proven regression.
**Reproduce** (raw captures are not committed):
```sh
tools/ab/ab_run.sh --arms tools/ab/arms_heat_pillar.txt --runs 10 --rounds 7 \
--conc 7 --outdir /tmp/ab/heatshift
python3 tools/ab/ab_analyze.py /tmp/ab/heatshift --reference old
python3 tools/ab/ab_mechanism.py /tmp/ab/heatshift --reference old
python3 tools/ab/ab_range_bands.py /tmp/ab/heatshift --reference old
```