Files
SirRoboGarage/docs/tfil_heat_pillar_ab.md
T
SirStone 48f38b80e7 TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
2026-09-25 22:23:11 +02:00

14 KiB
Raw Blame History

TFIL heat-time bullet model + virtual-pillar removal — LIVE A/B (negative)

Question. Two movement changes shipped in HEAD — (a) the time-indexed bullet heat (TR_TFIL_HEAT_TIME, job j105) and (b) the removal of the invented virtual centre pillar (TR_TFIL_PILLAR_ON=1 restores it, job j106) — judged on damage/run and ROUND WINS, never on hit rate.

Setup. One frozen ModularBot from git archive HEAD (commit f58d65d2e8206cebd5b7d0a9465950dd9e6d2c28, binary sha256 74fd010ed1ffc4bb8a8de5569f58843479e1f50164a8a3f61e9c81cef108eaef) vs the real DrussGT through tools/robocode_shim/run_bridge_battle.sh: 6 arms × 10 runs × 7 rounds = 60 battles, 420 rounds, --conc 7, 0 failed. Raw per-tick captures (~306 MB) live at /tmp/ab/heatshift/ and are not committed; every per-run value and every test below is in common_libs/tests/fixtures/tfil_heat_pillar_ab_results.json (raw tool output in ..._report.txt, ..._mechanism.txt, ..._bands.txt).

HEAD already has the pillar REMOVED and heat-time OFF, so the PRE-change mover is reconstructed with env: old = TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0.

THE METRIC RULE. Movement arms are judged on damage/run and round wins. Hit rate and hits-taken are reported as context only, and neither decides anything here. This is not stylistic: in the previous TFIL A/B an arm took significantly FEWER hits (87.2 vs 96.5, p=0.0022) and still dealt the least damage and won the fewest rounds. It has happened again below — tau3 has the best pooled hit rate of all six arms (11.56% vs old 10.99%) and the fewest round wins (20/70 vs 35/70). Hit rate would have ranked this experiment exactly backwards.

1. Live result table (10 runs × 7 rounds per arm vs real DrussGT) [MEASURED]

arm env damage/run damage taken/run ROUND WINS win% shots/run hits taken/run hit rate (context)
old (pre-change mover) TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 287 202 35/70 50.0% 783 93.5 10.99%
pillaoff (shipped default) (none) 282 232 30/70 42.9% 729 95.0 11.14%
tau3 TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 261 299 20/70 28.6% 666 106.2 11.56%
tau5 TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 250 294 18/70 25.7% 689 105.3 10.83%
tau9 TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 265 231 22/70 31.4% 750 99.5 10.51%
tau15 TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 265 185 32/70 45.7% 812 89.5 9.93%

old is the best arm on both verdict metrics (damage/run and round wins); only tau15 takes less damage/run, and it wins 3 fewer rounds and deals 22 less damage per run. Every heat-time arm is worse or equal on both verdict columns. The two "isolation" comparisons:

  • pillar: old vs pillaoff — +5.7 damage/run, +0.5 wins/run, −30.8 damage taken/run (i.e. pillaoff takes MORE).
  • heat-time: pillaoff vs tau9 — +16.4 damage/run, +0.8 wins/run (i.e. tau9 is worse even against the pillar-removed control).

2. Per-run values [MEASURED]

arm damage/run by run (r1..r10) wins by run (r1..r10)
old 271 273 255 270 282 309 338 305 266 306 3 4 4 2 2 4 4 4 3 5 /7
pillaoff 207 332 263 332 286 280 244 269 311 294 1 5 2 4 4 3 2 2 4 3 /7
tau3 213 296 182 285 283 220 290 304 245 293 1 3 1 3 3 1 4 1 2 1 /7
tau5 272 235 241 236 281 295 248 243 218 225 2 2 2 1 3 1 2 2 2 1 /7
tau9 295 283 275 239 268 258 280 262 228 266 3 3 2 2 2 3 2 3 1 1 /7
tau15 266 264 258 239 278 291 230 286 283 261 3 4 3 2 4 4 2 4 4 2 /7

The tau3/tau5 win counts are consistently low (all 10 runs ≤3/7 for tau5, 9 of 10 ≤3/7 for tau3) — not one lucky bad run.

3. Tests vs old (per-run values, two-sided; exact permutation at 10v10) [MEASURED]

metric arm diff (old − arm) perm p method Mann-Whitney p U
damage/run pillaoff +5.66 0.7096 exact 0.8501 47.0
damage/run tau3 +26.34 0.1137 exact 0.3075 36.0
damage/run tau5 +37.89 0.0042 exact 0.0091 15.0
damage/run tau9 +22.06 0.0468 exact 0.1041 28.0
damage/run tau15 +22.04 0.0460 exact 0.1041 28.0
round wins pillaoff +0.50 0.4317 exact 0.3631 38.0
round wins tau3 +1.50 0.0125 exact 0.0101 16.5
round wins tau5 +1.70 0.0010 exact 0.0014 9.0
round wins tau9 +1.30 0.0097 exact 0.0087 16.0
round wins tau15 +0.30 0.6369 exact 0.5127 41.5
damage taken/run pillaoff −30.82 0.0398 exact 0.1041 —
damage taken/run tau3 −97.27 <0.0001 exact 0.0002 —
damage taken/run tau5 −92.03 0.0004 exact 0.0022 —
damage taken/run tau9 −29.59 0.1036 exact 0.2413 —
damage taken/run tau15 +16.13 0.2213 exact 0.3075 —

Round-level pooled Fisher vs old (anti-conservative — rounds cluster within runs): pillaoff 0.4980, tau3 0.0150, tau5 0.0051, tau9 0.0386, tau15 0.7352.

4. What the test can and cannot see [MEASURED]

MDE (two-sample, α=0.05 two-sided, 80% power, n=10/arm, from the old per-run SD):

metric sd(old) MDE (absolute) MDE vs old mean
damage/run 25.98 32.55 11.3% of 287.5
round wins/run 0.97 1.22 34.8% of 3.5 win/run
  • Heat-time is a visible effect: tau3/tau5/tau9 lose 1.3–1.7 wins/run, at or above the 1.22 win MDE, with p≤0.013. tau5/tau9/tau15 lose 22–38 damage/run, around the 32.55 damage MDE, p≤0.047.
  • The pillar result is NOT decidable at this n: the whole observed old advantage is +5.7 damage and +0.5 wins/run, both well inside the MDE. This test only rules out the pillar removing ≥33 damage/run or ≥1.2 wins/run; it cannot see anything smaller. The pillaoff damage-taken regression (−30.8/run, p=0.040) is right at the MDE edge and its rank-sum cross-check is only p=0.104, so treat it as a weak-but-consistent signal, not a proven loss.

5. Mechanism checks — did the knob actually change behaviour? [MEASURED]

Raw per-tick worldstate (both tanks' positions every tick) + the fire/hit event sidecar. Aggregated per round, then averaged over rounds and runs.

5a. Central-box occupancy — the 144×144 px box the pillar covered (x 328–472, y 228–372)

Spawns are bottom-left (us) / top (enemy), never in the box, so occupancy is genuine transit. The pillar removal DID make us use the centre.

arm ticks inside box diff vs old perm p
old 0.09% — —
pillaoff 2.37% +2.29pp <0.0001
tau3 26.12% +26.04pp <0.0001
tau5 21.30% +21.21pp <0.0001
tau9 10.40% +10.31pp <0.0001
tau15 4.68% +4.60pp <0.0001

MDE for this metric is 0.11pp, so all the shifts are real. Note the ordering: old 0.09% → pillaoff 2.4% → tau15 4.7% → tau9 10.4% → tau5 21.3% → tau3 26.1%. Removing the pillar opens the centre; the time-indexed heat (which stops the bullet corridor at speed·tau instead of the wall) opens it much more, and monotonically more as tau shrinks.

5b. Distance to the enemy (px, per-tick)

arm mean p10 median p90 400+ px % of ticks diff vs old perm p
old 469.3 433 465 511 83.0 — —
pillaoff 443.5 396 446 490 72.1 −25.7 0.0002
tau3 439.4 406 440 471 70.4 −29.9 0.0004
tau5 447.8 410 451 484 73.9 −21.4 0.0014
tau9 481.1 440 482 512 85.5 +11.8 0.0813
tau15 500.1 467 502 533 88.2 +30.8 0.0001

MDE = 20.8 px. old fights at ~469 px; every opening of the centre pulls the engagement 21–30 px closer (and 9–13pp more of the battle is inside 400 px), and tau15 pushes it 31 px further out. tau9 (the shipped tau) is essentially old (+11.8 px, p=0.08).

5c. Bullet proximity — how much time we spend near live bullet paths

Nearest live enemy bullet (our own bullets are excluded: a bullet is born at its own tank, so "any bullet" is dominated by our own just-fired shot; the any-bullet version, per the task text, is in the ANY-BULLET PROXIMITY table of common_libs/tests/fixtures/tfil_heat_pillar_ab_mechanism.txt).

arm ≤50 px ≤100 px ≤150 px mean min-dist px ≤100px diff vs old perm p
old 4.76% 27.99% 62.48% 136.1 — —
pillaoff 4.97% 29.10% 64.98% 132.8 +1.11pp 0.1381
tau3 6.19% 30.86% 61.39% 134.8 +2.87pp 0.0014
tau5 5.88% 30.60% 62.59% 134.1 +2.61pp 0.0005
tau9 4.99% 28.31% 60.75% 137.8 +0.32pp 0.6066
tau15 4.12% 25.57% 58.43% 141.4 −2.43pp 0.0001

MDE = 1.74pp. tau3/tau5 spend significantly more time within 100 px of a live enemy bullet; tau15 significantly less.

6. Liveness [MEASURED]

Every arm's declared env appears verbatim in the bot's own boot report (<arm>/run<N>.bot.stdout.log, section A "raw process environment") in all 10 runs; no arm env was ignored as unrecognised. This is a measurement, not a skip:

old       OK   (10/10 runs: TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 applied)
pillaoff  OK   (10/10 runs: no arm env; report present)
tau3      OK   (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 applied)
tau5      OK   (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 applied)
tau9      OK   (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 applied)
tau15     OK   (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 applied)

The effective-value section confirms the reconstruction: old reports TR_TFIL_PILLAR_ON = on (source: env) / TR_TFIL_HEAT_TIME = off, pillaoff reports TR_TFIL_PILLAR_ON = off (source: default) / TR_TFIL_HEAT_TIME = off (source: default), and each tau arm reports the requested tau (TR_TFIL_HEAT_TAU = 3.0/5.0/9.0/15.0 (source: env)).

7. Hit-rate context (NOT a verdict metric) [MEASURED]

arm pooled hit rate round wins our hits taken/run our shots/run
old 10.99% 35/70 93.5 783
pillaoff 11.14% 30/70 95.0 729
tau3 11.56% (best) 20/70 (worst) 106.2 666
tau5 10.83% 18/70 105.3 689
tau9 10.51% 22/70 99.5 750
tau15 9.93% (worst) 32/70 (2nd best) 89.5 812

The inversion is exact at both ends: the best-accuracy arm (tau3, 11.56%) wins the fewest rounds; the worst-accuracy arm (tau15, 9.93%) wins the second most. Hit rate is monotone in the WRONG direction here. Range-banded hit rates (..._bands.txt) are flat in the 300–450 px band that holds 40% of our shots (11.3–13.4%) and in 450+ (8.8–10.1%); no band rescues any heat-time arm.

8. Direct answers

Does the heat-time model help, hurt, or do nothing? [MEASURED] → HURTS.

  • At the shipped tau (9) and at 3/5 it significantly reduces round wins (20–22/70 vs 35/70, p=0.0010–0.0125) and reduces damage/run (p=0.004–0.047).
  • tau15 is a wash on wins (32/70, p=0.637) and 22 damage/run lower (p=0.046) — not better, mildly worse.
  • No tau improves either verdict metric. The mechanism is exactly what the change claims (shorter corridor → centre opens, engagement closes by 20–30 px, more time within 100 px of a live enemy bullet) — and the closer arms take ≈12 more hits/run (tau3 12.7, tau5 11.8) and lose 13–17 more rounds per 10 runs. Inferred: the free space this frees is the arena centre, and standing there costs more than the wall-hugging flat model costs. The offline "largest safe region" ruler rewarded exactly the region that shorter tau opens (tau 2 → 162 px, tau 9 → 140 px, tau 15 → 127 px), and that region is a proxy anti-correlated with survival against DrussGT.

Does removing the pillar help, hurt, or do nothing? [MEASURED] → no measurable benefit; point estimates are worse, and it is undecidable at this n.

  • old vs pillaoff: damage +5.7/run (p=0.710), wins +0.5/run (35 vs 30, p=0.432), round-level Fisher p=0.498 — no significant difference.
  • Damage taken is 30.8/run higher without the pillar (p=0.040 permutation; rank-sum cross-check p=0.104).
  • The mechanism check proves the change did alter movement: box occupancy 0.09% → 2.37% (p<0.0001), mean range 469 → 443 px (p=0.0002). So this is a real behaviour change, not a dead knob — it simply does not buy anything.

9. Which arm should be the shipped default?

old — the pre-change mover (TR_TFIL_PILLAR_ON=1 behaviour, heat-time off). Nothing beats it: it is the best arm on damage/run (287) and round wins (35/70), and second only to tau15 on damage taken/run (202 vs 185), where tau15 pays for its 3 fewer round wins and 22 less damage per run.

  • Heat-time: keep it OFF (as shipped). The default is already off (fca8993); no action needed, and TR_TFIL_HEAT_TIME=1 at any tested tau should stay off.
  • Pillar removal: the shipped default should be REVERTED. The shipped default since d0750ab is pillaoff; it does not beat old on any verdict metric and is directionally worse on all three (Δdamage −5.7, Δwins −0.5, Δdamage taken +30.8). This is a decision for the user: restore the pre-change default (PillarHotness/PillarRadiance back to 30/10, i.e. the TR_TFIL_PILLAR_ON behaviour by default, keeping the env knob for the off state). Honesty caveat: at n=10/arm the pillar contrast is inside the MDE (33 damage/run, 1.22 wins/run), so this is not a statistically significant "removal is worse"; it is "removal bought nothing measurable, and every point estimate moved the wrong way". The case for reverting is the absence of evidence of benefit plus parsimony, not a proven regression.

Reproduce (raw captures are not committed):

tools/ab/ab_run.sh --arms tools/ab/arms_heat_pillar.txt --runs 10 --rounds 7 \
    --conc 7 --outdir /tmp/ab/heatshift
python3 tools/ab/ab_analyze.py    /tmp/ab/heatshift --reference old
python3 tools/ab/ab_mechanism.py  /tmp/ab/heatshift --reference old
python3 tools/ab/ab_range_bands.py /tmp/ab/heatshift --reference old