282 lines
15 KiB
Markdown
282 lines
15 KiB
Markdown
# TFIL heat-time bullet model + virtual-pillar removal — LIVE A/B (negative)
|
||
|
||
> ## DECISION — the virtual centre pillar STAYS REMOVED
|
||
>
|
||
> **Decided by the bot's owner on 2026-09-25, after this A/B.** The recommendation
|
||
> below ("revert the default to pillar-ON") was **NOT adopted**. Rationale, in the
|
||
> owner's terms: the pillar is an *invented* hazard with no physical object behind
|
||
> it, and this test cannot distinguish "no benefit" from "small benefit" — the
|
||
> pillar contrast is **inside the minimum detectable effect** (32.6 dmg/run,
|
||
> 1.22 wins/run at n=10), and the raised damage-taken point estimate (p=0.040) is
|
||
> not corrected for the multiple arms tested.
|
||
>
|
||
> **So: do not "fix" this by reverting.** The shipped default keeps
|
||
> `PillarHotness = 0.0` / `PillarRadiance = 0.0`; `TR_TFIL_PILLAR_ON=1` restores the
|
||
> old field for any future experiment. If someone later wants to reverse this
|
||
> decision, it needs **its own pre-registered A/B at a sample size that can actually
|
||
> resolve a contrast this small** — not a re-reading of this one.
|
||
|
||
|
||
**Question.** Two movement changes shipped in HEAD — (a) the time-indexed bullet
|
||
heat (`TR_TFIL_HEAT_TIME`, job j105) and (b) the removal of the invented virtual
|
||
centre pillar (`TR_TFIL_PILLAR_ON=1` restores it, job j106) — judged on
|
||
**damage/run** and **ROUND WINS**, never on hit rate.
|
||
|
||
**Setup.** One frozen ModularBot from `git archive HEAD`
|
||
(commit `f58d65d2e8206cebd5b7d0a9465950dd9e6d2c28`, binary sha256
|
||
`74fd010ed1ffc4bb8a8de5569f58843479e1f50164a8a3f61e9c81cef108eaef`) vs the real
|
||
DrussGT through `tools/robocode_shim/run_bridge_battle.sh`: **6 arms × 10 runs ×
|
||
7 rounds = 60 battles, 420 rounds**, `--conc 7`, 0 failed. Raw per-tick captures
|
||
(~306 MB) live at `/tmp/ab/heatshift/` and are **not** committed; every per-run
|
||
value and every test below is in
|
||
`common_libs/tests/fixtures/tfil_heat_pillar_ab_results.json` (raw tool output in
|
||
`..._report.txt`, `..._mechanism.txt`, `..._bands.txt`).
|
||
|
||
`HEAD` already has the pillar REMOVED and heat-time OFF, so the PRE-change mover
|
||
is reconstructed with env: `old = TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0`.
|
||
|
||
> **THE METRIC RULE.** Movement arms are judged on **damage/run** and **round
|
||
> wins**. Hit rate and hits-taken are reported **as context only**, and neither
|
||
> decides anything here. This is not stylistic: in the previous TFIL A/B an arm
|
||
> took significantly FEWER hits (87.2 vs 96.5, p=0.0022) and still dealt the
|
||
> least damage and won the fewest rounds. It has happened **again below** —
|
||
> `tau3` has the *best* pooled hit rate of all six arms (11.56% vs `old` 10.99%)
|
||
> and the *fewest* round wins (20/70 vs 35/70). Hit rate would have ranked this
|
||
> experiment exactly backwards.
|
||
|
||
## 1. Live result table (10 runs × 7 rounds per arm vs real DrussGT) `[MEASURED]`
|
||
|
||
| arm | env | damage/run | damage taken/run | ROUND WINS | win% | shots/run | hits taken/run | hit rate (context) |
|
||
|---|---|---:|---:|---:|---:|---:|---:|---:|
|
||
| **old** (pre-change mover) | `TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0` | **287** | **202** | **35/70** | **50.0%** | 783 | 93.5 | 10.99% |
|
||
| **pillaoff** (shipped default) | *(none)* | 282 | 232 | 30/70 | 42.9% | 729 | 95.0 | 11.14% |
|
||
| tau3 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3` | 261 | 299 | 20/70 | 28.6% | 666 | 106.2 | 11.56% |
|
||
| tau5 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5` | 250 | 294 | 18/70 | 25.7% | 689 | 105.3 | 10.83% |
|
||
| tau9 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9` | 265 | 231 | 22/70 | 31.4% | 750 | 99.5 | 10.51% |
|
||
| tau15 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15` | 265 | 185 | 32/70 | 45.7% | 812 | 89.5 | 9.93% |
|
||
|
||
`old` is the best arm on both verdict metrics (damage/run and round wins); only
|
||
`tau15` takes less damage/run, and it wins 3 fewer rounds and deals 22 less
|
||
damage per run. Every heat-time arm is worse or equal on both verdict columns.
|
||
The two "isolation" comparisons:
|
||
|
||
* **pillar**: `old` vs `pillaoff` — +5.7 damage/run, +0.5 wins/run, **−30.8
|
||
damage taken/run** (i.e. `pillaoff` takes MORE).
|
||
* **heat-time**: `pillaoff` vs `tau9` — +16.4 damage/run, +0.8 wins/run (i.e.
|
||
`tau9` is worse even against the pillar-removed control).
|
||
|
||
## 2. Per-run values `[MEASURED]`
|
||
|
||
| arm | damage/run by run (r1..r10) | wins by run (r1..r10) |
|
||
|---|---|---|
|
||
| old | 271 273 255 270 282 309 338 305 266 306 | 3 4 4 2 2 4 4 4 3 5 /7 |
|
||
| pillaoff | 207 332 263 332 286 280 244 269 311 294 | 1 5 2 4 4 3 2 2 4 3 /7 |
|
||
| tau3 | 213 296 182 285 283 220 290 304 245 293 | 1 3 1 3 3 1 4 1 2 1 /7 |
|
||
| tau5 | 272 235 241 236 281 295 248 243 218 225 | 2 2 2 1 3 1 2 2 2 1 /7 |
|
||
| tau9 | 295 283 275 239 268 258 280 262 228 266 | 3 3 2 2 2 3 2 3 1 1 /7 |
|
||
| tau15 | 266 264 258 239 278 291 230 286 283 261 | 3 4 3 2 4 4 2 4 4 2 /7 |
|
||
|
||
The tau3/tau5 win counts are *consistently* low (all 10 runs ≤3/7 for tau5, 9 of
|
||
10 ≤3/7 for tau3) — not one lucky bad run.
|
||
|
||
## 3. Tests vs `old` (per-run values, two-sided; exact permutation at 10v10) `[MEASURED]`
|
||
|
||
| metric | arm | diff (old − arm) | perm p | method | Mann-Whitney p | U |
|
||
|---|---|---:|---:|---|---:|---:|
|
||
| damage/run | pillaoff | +5.66 | 0.7096 | exact | 0.8501 | 47.0 |
|
||
| damage/run | tau3 | +26.34 | 0.1137 | exact | 0.3075 | 36.0 |
|
||
| damage/run | tau5 | +37.89 | **0.0042** | exact | **0.0091** | 15.0 |
|
||
| damage/run | tau9 | +22.06 | **0.0468** | exact | 0.1041 | 28.0 |
|
||
| damage/run | tau15 | +22.04 | **0.0460** | exact | 0.1041 | 28.0 |
|
||
| round wins | pillaoff | +0.50 | 0.4317 | exact | 0.3631 | 38.0 |
|
||
| round wins | tau3 | +1.50 | **0.0125** | exact | **0.0101** | 16.5 |
|
||
| round wins | tau5 | +1.70 | **0.0010** | exact | **0.0014** | 9.0 |
|
||
| round wins | tau9 | +1.30 | **0.0097** | exact | **0.0087** | 16.0 |
|
||
| round wins | tau15 | +0.30 | 0.6369 | exact | 0.5127 | 41.5 |
|
||
| damage taken/run | pillaoff | −30.82 | **0.0398** | exact | 0.1041 | — |
|
||
| damage taken/run | tau3 | −97.27 | **<0.0001** | exact | **0.0002** | — |
|
||
| damage taken/run | tau5 | −92.03 | **0.0004** | exact | **0.0022** | — |
|
||
| damage taken/run | tau9 | −29.59 | 0.1036 | exact | 0.2413 | — |
|
||
| damage taken/run | tau15 | +16.13 | 0.2213 | exact | 0.3075 | — |
|
||
|
||
Round-level pooled Fisher vs `old` (anti-conservative — rounds cluster within
|
||
runs): pillaoff 0.4980, tau3 0.0150, tau5 0.0051, tau9 0.0386, tau15 0.7352.
|
||
|
||
## 4. What the test can and cannot see `[MEASURED]`
|
||
|
||
MDE (two-sample, α=0.05 two-sided, 80% power, n=10/arm, from the `old` per-run SD):
|
||
|
||
| metric | sd(old) | MDE (absolute) | MDE vs `old` mean |
|
||
|---|---:|---:|---:|
|
||
| damage/run | 25.98 | **32.55** | 11.3% of 287.5 |
|
||
| round wins/run | 0.97 | **1.22** | 34.8% of 3.5 win/run |
|
||
|
||
* **Heat-time is a visible effect**: tau3/tau5/tau9 lose 1.3–1.7 wins/run, at or
|
||
above the 1.22 win MDE, with p≤0.013. tau5/tau9/tau15 lose 22–38 damage/run,
|
||
around the 32.55 damage MDE, p≤0.047.
|
||
* **The pillar result is NOT decidable at this n**: the whole observed `old`
|
||
advantage is +5.7 damage and +0.5 wins/run, both *well inside* the MDE. This
|
||
test only rules out the pillar removing ≥33 damage/run or ≥1.2 wins/run; it
|
||
cannot see anything smaller. The `pillaoff` damage-taken regression (−30.8/run,
|
||
p=0.040) is right at the MDE edge and its rank-sum cross-check is only
|
||
p=0.104, so treat it as a weak-but-consistent signal, not a proven loss.
|
||
|
||
## 5. Mechanism checks — did the knob actually change behaviour? `[MEASURED]`
|
||
|
||
Raw per-tick worldstate (both tanks' positions every tick) + the fire/hit event
|
||
sidecar. Aggregated per round, then averaged over rounds and runs.
|
||
|
||
### 5a. Central-box occupancy — the 144×144 px box the pillar covered (x 328–472, y 228–372)
|
||
|
||
Spawns are bottom-left (us) / top (enemy), never in the box, so occupancy is
|
||
genuine transit. **The pillar removal DID make us use the centre.**
|
||
|
||
| arm | ticks inside box | diff vs `old` | perm p |
|
||
|---|---:|---:|---:|
|
||
| old | 0.09% | — | — |
|
||
| pillaoff | 2.37% | +2.29pp | **<0.0001** |
|
||
| tau3 | 26.12% | +26.04pp | **<0.0001** |
|
||
| tau5 | 21.30% | +21.21pp | **<0.0001** |
|
||
| tau9 | 10.40% | +10.31pp | **<0.0001** |
|
||
| tau15 | 4.68% | +4.60pp | **<0.0001** |
|
||
|
||
MDE for this metric is 0.11pp, so all the shifts are real. Note the ordering:
|
||
`old` 0.09% → `pillaoff` 2.4% → `tau15` 4.7% → `tau9` 10.4% → `tau5` 21.3% →
|
||
`tau3` 26.1%. Removing the pillar opens the centre; the time-indexed heat (which
|
||
stops the bullet corridor at `speed·tau` instead of the wall) opens it much more,
|
||
and monotonically more as tau shrinks.
|
||
|
||
### 5b. Distance to the enemy (px, per-tick)
|
||
|
||
| arm | mean | p10 | median | p90 | 400+ px % of ticks | diff vs `old` | perm p |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|
|
||
| old | 469.3 | 433 | 465 | 511 | 83.0 | — | — |
|
||
| pillaoff | 443.5 | 396 | 446 | 490 | 72.1 | −25.7 | **0.0002** |
|
||
| tau3 | 439.4 | 406 | 440 | 471 | 70.4 | −29.9 | **0.0004** |
|
||
| tau5 | 447.8 | 410 | 451 | 484 | 73.9 | −21.4 | **0.0014** |
|
||
| tau9 | 481.1 | 440 | 482 | 512 | 85.5 | +11.8 | 0.0813 |
|
||
| tau15 | 500.1 | 467 | 502 | 533 | 88.2 | +30.8 | **0.0001** |
|
||
|
||
MDE = 20.8 px. `old` fights at ~469 px; every opening of the centre pulls the
|
||
engagement 21–30 px closer (and 9–13pp more of the battle is inside 400 px), and
|
||
tau15 pushes it 31 px further out. `tau9` (the shipped tau) is essentially `old`
|
||
(+11.8 px, p=0.08).
|
||
|
||
### 5c. Bullet proximity — how much time we spend near live bullet paths
|
||
|
||
Nearest **live enemy** bullet (our own bullets are excluded: a bullet is born at
|
||
its own tank, so "any bullet" is dominated by our own just-fired shot; the
|
||
any-bullet version, per the task text, is in the ANY-BULLET PROXIMITY table of
|
||
`common_libs/tests/fixtures/tfil_heat_pillar_ab_mechanism.txt`).
|
||
|
||
| arm | ≤50 px | ≤100 px | ≤150 px | mean min-dist px | ≤100px diff vs `old` | perm p |
|
||
|---|---:|---:|---:|---:|---:|---:|
|
||
| old | 4.76% | 27.99% | 62.48% | 136.1 | — | — |
|
||
| pillaoff | 4.97% | 29.10% | 64.98% | 132.8 | +1.11pp | 0.1381 |
|
||
| tau3 | 6.19% | 30.86% | 61.39% | 134.8 | +2.87pp | **0.0014** |
|
||
| tau5 | 5.88% | 30.60% | 62.59% | 134.1 | +2.61pp | **0.0005** |
|
||
| tau9 | 4.99% | 28.31% | 60.75% | 137.8 | +0.32pp | 0.6066 |
|
||
| tau15 | 4.12% | 25.57% | 58.43% | 141.4 | −2.43pp | **0.0001** |
|
||
|
||
MDE = 1.74pp. tau3/tau5 spend significantly more time within 100 px of a live
|
||
enemy bullet; tau15 significantly less.
|
||
|
||
## 6. Liveness `[MEASURED]`
|
||
|
||
Every arm's declared env appears verbatim in the bot's own boot report
|
||
(`<arm>/run<N>.bot.stdout.log`, section A "raw process environment") in all 10
|
||
runs; no arm env was ignored as unrecognised. This is a *measurement*, not a
|
||
skip:
|
||
|
||
```
|
||
old OK (10/10 runs: TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 applied)
|
||
pillaoff OK (10/10 runs: no arm env; report present)
|
||
tau3 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 applied)
|
||
tau5 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 applied)
|
||
tau9 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 applied)
|
||
tau15 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 applied)
|
||
```
|
||
|
||
The effective-value section confirms the reconstruction: `old` reports
|
||
`TR_TFIL_PILLAR_ON = on (source: env)` / `TR_TFIL_HEAT_TIME = off`, `pillaoff`
|
||
reports `TR_TFIL_PILLAR_ON = off (source: default)` / `TR_TFIL_HEAT_TIME = off
|
||
(source: default)`, and each tau arm reports the requested tau
|
||
(`TR_TFIL_HEAT_TAU = 3.0/5.0/9.0/15.0 (source: env)`).
|
||
|
||
## 7. Hit-rate context (NOT a verdict metric) `[MEASURED]`
|
||
|
||
| arm | pooled hit rate | round wins | our hits taken/run | our shots/run |
|
||
|---|---:|---:|---:|---:|
|
||
| old | 10.99% | 35/70 | 93.5 | 783 |
|
||
| pillaoff | 11.14% | 30/70 | 95.0 | 729 |
|
||
| tau3 | **11.56%** (best) | **20/70** (worst) | 106.2 | 666 |
|
||
| tau5 | 10.83% | 18/70 | 105.3 | 689 |
|
||
| tau9 | 10.51% | 22/70 | 99.5 | 750 |
|
||
| tau15 | **9.93%** (worst) | **32/70** (2nd best) | 89.5 | 812 |
|
||
|
||
The inversion is exact at both ends: the best-accuracy arm (`tau3`, 11.56%) wins
|
||
the fewest rounds; the worst-accuracy arm (`tau15`, 9.93%) wins the second most.
|
||
Hit rate is monotone in the WRONG direction here. Range-banded hit rates
|
||
(`..._bands.txt`) are flat in the 300–450 px band that holds 40% of our shots
|
||
(11.3–13.4%) and in 450+ (8.8–10.1%); no band rescues any heat-time arm.
|
||
|
||
## 8. Direct answers
|
||
|
||
**Does the heat-time model help, hurt, or do nothing? `[MEASURED] → HURTS.`**
|
||
|
||
* At the shipped tau (9) and at 3/5 it **significantly reduces round wins**
|
||
(20–22/70 vs 35/70, p=0.0010–0.0125) and reduces damage/run (p=0.004–0.047).
|
||
* `tau15` is a **wash on wins** (32/70, p=0.637) and 22 damage/run **lower**
|
||
(p=0.046) — not better, mildly worse.
|
||
* No tau improves either verdict metric. The mechanism is exactly what the
|
||
change claims (shorter corridor → centre opens, engagement closes by 20–30 px,
|
||
more time within 100 px of a live enemy bullet) — and the closer arms take
|
||
≈12 more hits/run (tau3 12.7, tau5 11.8) and lose 13–17 more rounds per 10
|
||
runs. **Inferred:** the free space this frees is the arena centre, and standing
|
||
there costs more than the wall-hugging flat model costs. The offline
|
||
"largest safe region" ruler rewarded exactly the region that shorter tau opens
|
||
(tau 2 → 162 px, tau 9 → 140 px, tau 15 → 127 px), and that region is a proxy
|
||
anti-correlated with survival against DrussGT.
|
||
|
||
**Does removing the pillar help, hurt, or do nothing? `[MEASURED] → no
|
||
measurable benefit; point estimates are worse, and it is undecidable at this n.`**
|
||
|
||
* `old` vs `pillaoff`: damage +5.7/run (p=0.710), wins +0.5/run (35 vs 30,
|
||
p=0.432), round-level Fisher p=0.498 — **no significant difference**.
|
||
* Damage taken is **30.8/run higher** without the pillar (p=0.040 permutation;
|
||
rank-sum cross-check p=0.104).
|
||
* The mechanism check proves the change **did** alter movement: box occupancy
|
||
0.09% → 2.37% (p<0.0001), mean range 469 → 443 px (p=0.0002). So this is a real
|
||
behaviour change, not a dead knob — it simply does not buy anything.
|
||
|
||
## 9. Which arm should be the shipped default?
|
||
|
||
**`old` — the pre-change mover (`TR_TFIL_PILLAR_ON=1` behaviour, heat-time off).
|
||
Nothing beats it: it is the best arm on damage/run (287) and round wins
|
||
(35/70), and second only to `tau15` on damage taken/run (202 vs 185), where
|
||
`tau15` pays for its 3 fewer round wins and 22 less damage per run.**
|
||
|
||
* **Heat-time: keep it OFF (as shipped).** The default is already off (`fca8993`);
|
||
no action needed, and `TR_TFIL_HEAT_TIME=1` at any tested tau should stay off.
|
||
* **Pillar removal: this test recommends reverting the default** — see the **DECISION** banner at the top of this document: the owner **overruled** this recommendation and the pillar **stays removed**. The shipped default
|
||
since `d0750ab` is `pillaoff`; it does **not** beat `old` on any verdict metric
|
||
and is directionally worse on all three (Δdamage −5.7, Δwins −0.5, Δdamage
|
||
taken +30.8). This is a decision for the user: restore the pre-change default
|
||
(`PillarHotness`/`PillarRadiance` back to 30/10, i.e. the `TR_TFIL_PILLAR_ON`
|
||
behaviour by default, keeping the env knob for the off state). **Honesty
|
||
caveat:** at n=10/arm the pillar contrast is inside the MDE (33 damage/run,
|
||
1.22 wins/run), so this is not a statistically significant "removal is worse";
|
||
it is "removal bought nothing measurable, and every point estimate moved the
|
||
wrong way". The case for reverting is the absence of evidence of benefit plus
|
||
parsimony, not a proven regression.
|
||
|
||
**Reproduce** (raw captures are not committed):
|
||
|
||
```sh
|
||
tools/ab/ab_run.sh --arms tools/ab/arms_heat_pillar.txt --runs 10 --rounds 7 \
|
||
--conc 7 --outdir /tmp/ab/heatshift
|
||
python3 tools/ab/ab_analyze.py /tmp/ab/heatshift --reference old
|
||
python3 tools/ab/ab_mechanism.py /tmp/ab/heatshift --reference old
|
||
python3 tools/ab/ab_range_bands.py /tmp/ab/heatshift --reference old
|
||
```
|