TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover

Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
This commit is contained in:
2026-09-25 22:21:41 +02:00
parent f58d65d2e8
commit 48f38b80e7
7 changed files with 3121 additions and 0 deletions
@@ -0,0 +1,90 @@
====================================================================================================
HIT RATE BY RANGE BAND (our shots; band = shooter->target distance px at the fire tick)
session: /tmp/ab/heatshift arms: old, pillaoff, tau15, tau3, tau5, tau9 reference: old
====================================================================================================
band old pillaoff tau15 tau3 tau5 tau9
----------------------------------------------------------------------------------------------------------------------------------------------------------------------
0-100 2 1 50.0% 3 2 66.7% 2 2 100.0% 6 5 83.3% 6 3 50.0% 8 3 37.5%
100-200 21 3 14.3% 38 11 28.9% 16 3 18.8% 22 8 36.4% 31 14 45.2% 24 5 20.8%
200-300 134 30 22.4% 214 37 17.3% 103 20 19.4% 196 42 21.4% 216 39 18.1% 122 24 19.7%
300-450 3016 342 11.3% 3663 414 11.3% 1768 204 11.5% 3134 367 11.7% 2851 343 12.0% 1939 260 13.4%
450+ 4468 436 9.8% 3196 312 9.8% 6045 539 8.9% 3122 314 10.1% 3619 319 8.8% 5220 469 9.0%
ALL 7641 812 10.6% 7114 776 10.9% 7934 768 9.7% 6480 736 11.4% 6723 718 10.7% 7313 761 10.4%
PER-RUN BAND RATES (shows the spread behind the pooled numbers)
band 0-100
old n= 2 0 100
pillaoff n= 3 100 0 100
tau15 n= 2 100 100
tau3 n= 5 100 100 100 100 50
tau5 n= 5 100 100 50 0 0
tau9 n= 5 0 0 50 100 33
band 100-200
old n= 9 0 0 0 0 0 0 20 100 0
pillaoff n=10 0 0 50 0 50 43 56 0 25 0
tau15 n= 8 0 0 33 50 0 0 0 0
tau3 n= 8 33 0 0 40 33 67 0 0
tau5 n=10 38 100 50 50 0 33 0 0 60 100
tau9 n= 8 25 60 0 0 0 0 0 25
band 200-300
old n=10 15 14 17 26 0 25 60 18 40 30
pillaoff n=10 29 10 14 16 19 24 18 25 10 19
tau15 n=10 50 11 25 12 17 33 17 33 7 0
tau3 n=10 17 17 12 0 39 21 22 27 20 17
tau5 n=10 21 29 6 21 19 26 24 29 9 6
tau9 n=10 33 23 25 12 8 14 33 40 17 13
band 300-450
old n=10 11 12 13 9 12 10 9 14 11 12
pillaoff n=10 9 11 11 12 13 10 12 11 13 12
tau15 n=10 13 11 9 11 8 14 13 9 15 13
tau3 n=10 11 11 13 11 12 11 9 13 13 11
tau5 n=10 14 11 11 12 10 14 12 11 14 12
tau9 n=10 14 13 12 14 15 11 14 14 14 12
band 450+
old n=10 10 10 10 9 9 12 11 9 9 9
pillaoff n=10 8 13 13 7 10 10 11 8 9 10
tau15 n=10 8 10 10 9 9 8 9 8 9 9
tau3 n=10 8 11 11 7 9 12 11 11 12 10
tau5 n=10 10 8 9 10 10 7 9 9 8 8
tau9 n=10 9 8 10 9 8 10 10 9 8 9
band ALL
old n=10 11 11 11 9 10 11 11 11 11 10
pillaoff n=10 9 12 12 10 12 11 13 10 11 12
tau15 n=10 9 10 10 9 9 10 11 8 11 10
tau3 n=10 10 11 12 9 12 12 11 13 13 11
tau5 n=10 12 10 10 11 10 11 11 10 11 10
tau9 n=10 10 11 11 10 9 11 11 10 10 10
PER-BAND PERMUTATION TEST vs `old` (per-run rates, two-sided)
band arm d(pp) p method MCse MDE(pp)
------------------------------------------------------------------
0-100 pillaoff +16.67 1.0000 exact 0.0000 198.10
0-100 tau15 +50.00 1.0000 exact 0.0000 198.10
0-100 tau3 +40.00 0.2857 exact 0.0000 198.10
0-100 tau5 +0.00 1.0000 exact 0.0000 198.10
0-100 tau9 -13.33 0.9524 exact 0.0000 198.10
100-200 pillaoff +9.01 0.5427 exact 0.0000 43.80
100-200 tau15 -2.92 0.8765 exact 0.0000 43.80
100-200 tau3 +8.33 0.6226 exact 0.0000 43.80
100-200 tau5 +29.75 0.0880 exact 0.0000 43.80
100-200 tau9 +0.42 1.0000 exact 0.0000 43.80
200-300 pillaoff -6.15 0.3045 exact 0.0000 20.62
200-300 tau15 -3.96 0.5841 exact 0.0000 20.62
200-300 tau3 -5.29 0.4112 exact 0.0000 20.62
200-300 tau5 -5.54 0.3777 exact 0.0000 20.62
200-300 tau9 -2.57 0.6937 exact 0.0000 20.62
300-450 pillaoff -0.14 0.8270 exact 0.0000 2.03
300-450 tau15 -0.04 0.9677 exact 0.0000 2.03
300-450 tau3 +0.11 0.8729 exact 0.0000 2.03
300-450 tau5 +0.54 0.4566 exact 0.0000 2.03
300-450 tau9 +1.98 0.0069 exact 0.0000 2.03
450+ pillaoff +0.04 0.9606 exact 0.0000 1.21
450+ tau15 -0.87 0.0454 exact 0.0000 1.21
450+ tau3 +0.36 0.5427 exact 0.0000 1.21
450+ tau5 -1.05 0.0290 exact 0.0000 1.21
450+ tau9 -0.88 0.0456 exact 0.0000 1.21
ALL pillaoff +0.32 0.4681 exact 0.0000 0.76
ALL tau15 -0.95 0.0066 exact 0.0000 0.76
ALL tau3 +0.71 0.1325 exact 0.0000 0.76
ALL tau5 +0.06 0.8627 exact 0.0000 0.76
ALL tau9 -0.25 0.3258 exact 0.0000 0.76
@@ -0,0 +1,74 @@
====================================================================================================
MECHANISM CHECK — did the arm actually move differently?
session /tmp/ab/heatshift arms old, pillaoff, tau15, tau3, tau5, tau9 reference old
box = 144x144 px centred on the arena centre (x 328-472, y 228-372)
====================================================================================================
PER-ARM MECHANISM (mean of the per-round values, then over runs)
arm runs box% dist px p10 med p90 enB<=100px% enBmin px
-------------------------------------------------------------------------------
old 10 0.09 469.3 433 465 511 27.99 136.1
pillaoff 10 2.37 443.5 396 446 490 29.10 132.8
tau15 10 4.68 500.1 467 502 533 25.57 141.4
tau3 10 26.12 439.4 406 440 471 30.86 134.8
tau5 10 21.30 447.8 410 451 484 30.60 134.1
tau9 10 10.40 481.1 440 482 512 28.31 137.8
DISTANCE-TO-ENEMY BANDS (fraction of ticks, mean over rounds/runs, %)
arm 0-100 100-200 200-300 300-400 400+
------------------------------------------------------------
old 0.11 0.51 2.15 14.24 82.99
pillaoff 0.15 0.95 3.33 23.46 72.12
tau15 0.18 0.56 1.64 9.47 88.15
tau3 0.18 0.89 3.73 24.80 70.39
tau5 0.23 0.84 3.80 21.20 73.93
tau9 0.21 0.71 2.04 11.53 85.51
ENEMY-BULLET PROXIMITY (nearest LIVE ENEMY bullet; our own bullets are excluded because a bullet is born at its own tank, so 'any bullet' is dominated by our just-fired shot)
arm <=50px <=100px <=150px min px enemy%
-------------------------------------------------------
old 4.76 27.99 62.48 136.1 97.04
pillaoff 4.97 29.10 64.98 132.8 97.25
tau15 4.12 25.57 58.43 141.4 97.06
tau3 6.19 30.86 61.39 134.8 97.03
tau5 5.88 30.60 62.59 134.1 97.31
tau9 4.99 28.31 60.75 137.8 97.45
ANY-BULLET PROXIMITY (per the task text: nearest live bullet, ours included; context only)
arm <=50px <=100px <=150px min px any%
-------------------------------------------------------
old 24.96 56.38 83.24 92.4 98.17
pillaoff 25.33 57.33 84.58 90.5 98.01
tau15 24.39 54.95 81.71 94.3 98.04
tau3 26.25 57.89 82.66 90.7 97.65
tau5 25.88 57.83 83.50 90.6 97.84
tau9 25.05 56.42 82.41 92.7 98.05
PER-RUN MECHANISM vs `old` (two-sided permutation on per-run means; exact when C(n,na)<=2e7)
metric arm delta p method MDE
------------------------------------------------------------------------------
central-box occupancy pillaoff +2.29 0.0000 exact 0.11%
central-box occupancy tau15 +4.60 0.0000 exact 0.11%
central-box occupancy tau3 +26.04 0.0000 exact 0.11%
central-box occupancy tau5 +21.21 0.0000 exact 0.11%
central-box occupancy tau9 +10.31 0.0000 exact 0.11%
mean dist to enemy pillaoff -25.74 0.0002 exact 20.81px
mean dist to enemy tau15 +30.82 0.0001 exact 20.81px
mean dist to enemy tau3 -29.90 0.0004 exact 20.81px
mean dist to enemy tau5 -21.44 0.0014 exact 20.81px
mean dist to enemy tau9 +11.82 0.0813 exact 20.81px
ticks enemy bullet<=100px pillaoff +1.11 0.1381 exact 1.74%
ticks enemy bullet<=100px tau15 -2.43 0.0001 exact 1.74%
ticks enemy bullet<=100px tau3 +2.87 0.0014 exact 1.74%
ticks enemy bullet<=100px tau5 +2.61 0.0005 exact 1.74%
ticks enemy bullet<=100px tau9 +0.32 0.6066 exact 1.74%
mean dist nearest enemy bullet pillaoff -3.29 0.0155 exact 3.03px
mean dist nearest enemy bullet tau15 +5.28 0.0000 exact 3.03px
mean dist nearest enemy bullet tau3 -1.24 0.3106 exact 3.03px
mean dist nearest enemy bullet tau5 -2.01 0.1130 exact 3.03px
mean dist nearest enemy bullet tau9 +1.68 0.1133 exact 3.03px
ticks ANY bullet<=100px pillaoff +0.95 0.0394 exact 1.39%
ticks ANY bullet<=100px tau15 -1.43 0.0040 exact 1.39%
ticks ANY bullet<=100px tau3 +1.51 0.0228 exact 1.39%
ticks ANY bullet<=100px tau5 +1.45 0.0147 exact 1.39%
ticks ANY bullet<=100px tau9 +0.04 0.9388 exact 1.39%
@@ -0,0 +1,101 @@
# session /tmp/ab/heatshift
# commit=f58d65d2e8206cebd5b7d0a9465950dd9e6d2c28 binary_sha256=74fd010ed1ffc4bb8a8de5569f58843479e1f50164a8a3f61e9c81cef108eaef rounds=7 runs=10 conc=7 ts=2026-09-25T22:09:42+02:00
ARM SUMMARY
arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run
--------------------------------------------------------------------------
old 10 287 202 35/70 50.0 783 93.5
pillaoff 10 282 232 30/70 42.9 729 95.0
tau3 10 261 299 20/70 28.6 666 106.2
tau5 10 250 294 18/70 25.7 689 105.3
tau9 10 265 231 22/70 31.4 750 99.5
tau15 10 265 185 32/70 45.7 812 89.5
PER-RUN (never just the mean)
old dmg: r1=271 r2=273 r3=255 r4=270 r5=282 r6=309 r7=338 r8=305 r9=266 r10=306
wins: r1=3/7 r2=4/7 r3=4/7 r4=2/7 r5=2/7 r6=4/7 r7=4/7 r8=4/7 r9=3/7 r10=5/7
pillaoff dmg: r1=207 r2=332 r3=263 r4=332 r5=286 r6=280 r7=244 r8=269 r9=311 r10=294
wins: r1=1/7 r2=5/7 r3=2/7 r4=4/7 r5=4/7 r6=3/7 r7=2/7 r8=2/7 r9=4/7 r10=3/7
tau3 dmg: r1=213 r2=296 r3=182 r4=285 r5=283 r6=220 r7=290 r8=304 r9=245 r10=293
wins: r1=1/7 r2=3/7 r3=1/7 r4=3/7 r5=3/7 r6=1/7 r7=4/7 r8=1/7 r9=2/7 r10=1/7
tau5 dmg: r1=272 r2=235 r3=241 r4=236 r5=281 r6=295 r7=248 r8=243 r9=218 r10=225
wins: r1=2/7 r2=2/7 r3=2/7 r4=1/7 r5=3/7 r6=1/7 r7=2/7 r8=2/7 r9=2/7 r10=1/7
tau9 dmg: r1=295 r2=283 r3=275 r4=239 r5=268 r6=258 r7=280 r8=262 r9=228 r10=266
wins: r1=3/7 r2=3/7 r3=2/7 r4=2/7 r5=2/7 r6=3/7 r7=2/7 r8=3/7 r9=1/7 r10=1/7
tau15 dmg: r1=266 r2=264 r3=258 r4=239 r5=278 r6=291 r7=230 r8=286 r9=283 r10=261
wins: r1=3/7 r2=4/7 r3=3/7 r4=2/7 r5=4/7 r6=4/7 r7=2/7 r8=4/7 r9=4/7 r10=2/7
PAIRWISE PERMUTATION TEST (per-run values) + MANN-WHITNEY CROSS-CHECK
permutation: exact when C(n,na) <= 20,000,000; otherwise Monte-Carlo 1,000,000 draws, seed=0x5eed5eed, p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))
metric A B diff(A-B) perm p method MC se MW p MW U
-------------------------------------------------------------------------------------------------
dmg/run old pillaoff +5.658 0.7096 exact - 0.8501 47.0
round wins old pillaoff +0.500 0.4317 exact - 0.3631 38.0
dmg/run old tau3 +26.343 0.1137 exact - 0.3075 36.0
round wins old tau3 +1.500 0.0125 exact - 0.0101 16.5
dmg/run old tau5 +37.893 0.0042 exact - 0.0091 15.0
round wins old tau5 +1.700 0.0010 exact - 0.0014 9.0
dmg/run old tau9 +22.056 0.0468 exact - 0.1041 28.0
round wins old tau9 +1.300 0.0097 exact - 0.0087 16.0
dmg/run old tau15 +22.039 0.0460 exact - 0.1041 28.0
round wins old tau15 +0.300 0.6369 exact - 0.5127 41.5
dmg/run pillaoff tau3 +20.686 0.2721 exact - 0.4727 40.0
round wins pillaoff tau3 +1.000 0.1149 exact - 0.0869 27.5
dmg/run pillaoff tau5 +32.236 0.0429 exact - 0.0539 24.0
round wins pillaoff tau5 +1.200 0.0262 exact - 0.0254 21.5
dmg/run pillaoff tau9 +16.398 0.2543 exact - 0.1859 32.0
round wins pillaoff tau9 +0.800 0.1548 exact - 0.1461 31.0
dmg/run pillaoff tau15 +16.382 0.2531 exact - 0.1859 32.0
round wins pillaoff tau15 -0.200 0.8378 exact - 0.7204 45.0
dmg/run tau3 tau5 +11.550 0.4674 exact - 0.3447 37.0
round wins tau3 tau5 +0.200 0.8121 exact - 0.9042 48.0
dmg/run tau3 tau9 -4.287 0.7791 exact - 0.5708 42.0
round wins tau3 tau9 -0.200 0.8197 exact - 0.6047 43.0
dmg/run tau3 tau15 -4.304 0.7776 exact - 0.6232 43.0
round wins tau3 tau15 -1.200 0.0356 exact - 0.0260 21.0
dmg/run tau5 tau9 -15.837 0.1359 exact - 0.1859 32.0
round wins tau5 tau9 -0.400 0.3567 exact - 0.2333 35.0
dmg/run tau5 tau15 -15.854 0.1337 exact - 0.1620 31.0
round wins tau5 tau15 -1.400 0.0034 exact - 0.0034 13.0
dmg/run tau9 tau15 -0.017 0.9984 exact - 1.0000 50.0
round wins tau9 tau15 -1.000 0.0356 exact - 0.0298 22.0
MINIMUM DETECTABLE EFFECT (two-sample, alpha=0.05 two-sided, 80% power; MDE = 2.8016*sd*sqrt(2/n))
metric n/arm sd(control) MDE(abs) MDE vs control mean
----------------------------------------------------------------
dmg/run 10 25.982 32.553 11.3% of 287.5
round wins 10 0.972 1.218 34.8% of 3.5
ROUND-LEVEL TEST (pooled rounds, Fisher exact) vs `old` — ANTI-CONSERVATIVE: rounds cluster within runs
arm ref wins arm wins p
----------------------------------------------
pillaoff 35/70 30/70 0.4980
tau3 35/70 20/70 0.0150
tau5 35/70 18/70 0.0051
tau9 35/70 22/70 0.0386
tau15 35/70 32/70 0.7352
LIVENESS (arm env applied in the bot's own boot report)
old OK (10/10 runs: TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 applied)
pillaoff OK (10/10 runs: no arm env; report present)
tau3 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 applied)
tau5 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 applied)
tau9 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 applied)
tau15 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 applied)
[bb] APPLIED-SHIFT CHECK (from bot stdout; needs TR_BITBRAIN_LOG=1). A provably-zero placebo emits ZERO [bb] lines.
arm runs w/log lines min max zeros
old 0/10 0 - - -
pillaoff 0/10 0 - - -
tau3 0/10 0 - - -
tau5 0/10 0 - - -
tau9 0/10 0 - - -
tau15 0/10 0 - - -
ROUND-WIN ATTRIBUTION (events primary; score tie-break for mutual-kill / timeout rounds)
old wins==firstPlaces 10/10 runs OK; single-death rounds agree with score 70/70 (0 tie-broken)
pillaoff wins==firstPlaces 10/10 runs OK; single-death rounds agree with score 70/70 (0 tie-broken)
tau3 wins==firstPlaces 10/10 runs OK; single-death rounds agree with score 69/69 (1 tie-broken)
tau5 wins==firstPlaces 10/10 runs OK; single-death rounds agree with score 69/69 (1 tie-broken)
tau9 wins==firstPlaces 10/10 runs OK; single-death rounds agree with score 70/70 (0 tie-broken)
tau15 wins==firstPlaces 10/10 runs OK; single-death rounds agree with score 69/69 (1 tie-broken)
File diff suppressed because it is too large Load Diff
+264
View File
@@ -0,0 +1,264 @@
# TFIL heat-time bullet model + virtual-pillar removal — LIVE A/B (negative)
**Question.** Two movement changes shipped in HEAD — (a) the time-indexed bullet
heat (`TR_TFIL_HEAT_TIME`, job j105) and (b) the removal of the invented virtual
centre pillar (`TR_TFIL_PILLAR_ON=1` restores it, job j106) — judged on
**damage/run** and **ROUND WINS**, never on hit rate.
**Setup.** One frozen ModularBot from `git archive HEAD`
(commit `f58d65d2e8206cebd5b7d0a9465950dd9e6d2c28`, binary sha256
`74fd010ed1ffc4bb8a8de5569f58843479e1f50164a8a3f61e9c81cef108eaef`) vs the real
DrussGT through `tools/robocode_shim/run_bridge_battle.sh`: **6 arms × 10 runs ×
7 rounds = 60 battles, 420 rounds**, `--conc 7`, 0 failed. Raw per-tick captures
(~306 MB) live at `/tmp/ab/heatshift/` and are **not** committed; every per-run
value and every test below is in
`common_libs/tests/fixtures/tfil_heat_pillar_ab_results.json` (raw tool output in
`..._report.txt`, `..._mechanism.txt`, `..._bands.txt`).
`HEAD` already has the pillar REMOVED and heat-time OFF, so the PRE-change mover
is reconstructed with env: `old = TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0`.
> **THE METRIC RULE.** Movement arms are judged on **damage/run** and **round
> wins**. Hit rate and hits-taken are reported **as context only**, and neither
> decides anything here. This is not stylistic: in the previous TFIL A/B an arm
> took significantly FEWER hits (87.2 vs 96.5, p=0.0022) and still dealt the
> least damage and won the fewest rounds. It has happened **again below** —
> `tau3` has the *best* pooled hit rate of all six arms (11.56% vs `old` 10.99%)
> and the *fewest* round wins (20/70 vs 35/70). Hit rate would have ranked this
> experiment exactly backwards.
## 1. Live result table (10 runs × 7 rounds per arm vs real DrussGT) `[MEASURED]`
| arm | env | damage/run | damage taken/run | ROUND WINS | win% | shots/run | hits taken/run | hit rate (context) |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| **old** (pre-change mover) | `TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0` | **287** | **202** | **35/70** | **50.0%** | 783 | 93.5 | 10.99% |
| **pillaoff** (shipped default) | *(none)* | 282 | 232 | 30/70 | 42.9% | 729 | 95.0 | 11.14% |
| tau3 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3` | 261 | 299 | 20/70 | 28.6% | 666 | 106.2 | 11.56% |
| tau5 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5` | 250 | 294 | 18/70 | 25.7% | 689 | 105.3 | 10.83% |
| tau9 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9` | 265 | 231 | 22/70 | 31.4% | 750 | 99.5 | 10.51% |
| tau15 | `TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15` | 265 | 185 | 32/70 | 45.7% | 812 | 89.5 | 9.93% |
`old` is the best arm on both verdict metrics (damage/run and round wins); only
`tau15` takes less damage/run, and it wins 3 fewer rounds and deals 22 less
damage per run. Every heat-time arm is worse or equal on both verdict columns.
The two "isolation" comparisons:
* **pillar**: `old` vs `pillaoff` — +5.7 damage/run, +0.5 wins/run, **−30.8
damage taken/run** (i.e. `pillaoff` takes MORE).
* **heat-time**: `pillaoff` vs `tau9` — +16.4 damage/run, +0.8 wins/run (i.e.
`tau9` is worse even against the pillar-removed control).
## 2. Per-run values `[MEASURED]`
| arm | damage/run by run (r1..r10) | wins by run (r1..r10) |
|---|---|---|
| old | 271 273 255 270 282 309 338 305 266 306 | 3 4 4 2 2 4 4 4 3 5 /7 |
| pillaoff | 207 332 263 332 286 280 244 269 311 294 | 1 5 2 4 4 3 2 2 4 3 /7 |
| tau3 | 213 296 182 285 283 220 290 304 245 293 | 1 3 1 3 3 1 4 1 2 1 /7 |
| tau5 | 272 235 241 236 281 295 248 243 218 225 | 2 2 2 1 3 1 2 2 2 1 /7 |
| tau9 | 295 283 275 239 268 258 280 262 228 266 | 3 3 2 2 2 3 2 3 1 1 /7 |
| tau15 | 266 264 258 239 278 291 230 286 283 261 | 3 4 3 2 4 4 2 4 4 2 /7 |
The tau3/tau5 win counts are *consistently* low (all 10 runs ≤3/7 for tau5, 9 of
10 ≤3/7 for tau3) — not one lucky bad run.
## 3. Tests vs `old` (per-run values, two-sided; exact permutation at 10v10) `[MEASURED]`
| metric | arm | diff (old − arm) | perm p | method | Mann-Whitney p | U |
|---|---|---:|---:|---|---:|---:|
| damage/run | pillaoff | +5.66 | 0.7096 | exact | 0.8501 | 47.0 |
| damage/run | tau3 | +26.34 | 0.1137 | exact | 0.3075 | 36.0 |
| damage/run | tau5 | +37.89 | **0.0042** | exact | **0.0091** | 15.0 |
| damage/run | tau9 | +22.06 | **0.0468** | exact | 0.1041 | 28.0 |
| damage/run | tau15 | +22.04 | **0.0460** | exact | 0.1041 | 28.0 |
| round wins | pillaoff | +0.50 | 0.4317 | exact | 0.3631 | 38.0 |
| round wins | tau3 | +1.50 | **0.0125** | exact | **0.0101** | 16.5 |
| round wins | tau5 | +1.70 | **0.0010** | exact | **0.0014** | 9.0 |
| round wins | tau9 | +1.30 | **0.0097** | exact | **0.0087** | 16.0 |
| round wins | tau15 | +0.30 | 0.6369 | exact | 0.5127 | 41.5 |
| damage taken/run | pillaoff | −30.82 | **0.0398** | exact | 0.1041 | — |
| damage taken/run | tau3 | −97.27 | **<0.0001** | exact | **0.0002** | — |
| damage taken/run | tau5 | −92.03 | **0.0004** | exact | **0.0022** | — |
| damage taken/run | tau9 | −29.59 | 0.1036 | exact | 0.2413 | — |
| damage taken/run | tau15 | +16.13 | 0.2213 | exact | 0.3075 | — |
Round-level pooled Fisher vs `old` (anti-conservative — rounds cluster within
runs): pillaoff 0.4980, tau3 0.0150, tau5 0.0051, tau9 0.0386, tau15 0.7352.
## 4. What the test can and cannot see `[MEASURED]`
MDE (two-sample, α=0.05 two-sided, 80% power, n=10/arm, from the `old` per-run SD):
| metric | sd(old) | MDE (absolute) | MDE vs `old` mean |
|---|---:|---:|---:|
| damage/run | 25.98 | **32.55** | 11.3% of 287.5 |
| round wins/run | 0.97 | **1.22** | 34.8% of 3.5 win/run |
* **Heat-time is a visible effect**: tau3/tau5/tau9 lose 1.3–1.7 wins/run, at or
above the 1.22 win MDE, with p≤0.013. tau5/tau9/tau15 lose 22–38 damage/run,
around the 32.55 damage MDE, p≤0.047.
* **The pillar result is NOT decidable at this n**: the whole observed `old`
advantage is +5.7 damage and +0.5 wins/run, both *well inside* the MDE. This
test only rules out the pillar removing ≥33 damage/run or ≥1.2 wins/run; it
cannot see anything smaller. The `pillaoff` damage-taken regression (−30.8/run,
p=0.040) is right at the MDE edge and its rank-sum cross-check is only
p=0.104, so treat it as a weak-but-consistent signal, not a proven loss.
## 5. Mechanism checks — did the knob actually change behaviour? `[MEASURED]`
Raw per-tick worldstate (both tanks' positions every tick) + the fire/hit event
sidecar. Aggregated per round, then averaged over rounds and runs.
### 5a. Central-box occupancy — the 144×144 px box the pillar covered (x 328–472, y 228–372)
Spawns are bottom-left (us) / top (enemy), never in the box, so occupancy is
genuine transit. **The pillar removal DID make us use the centre.**
| arm | ticks inside box | diff vs `old` | perm p |
|---|---:|---:|---:|
| old | 0.09% | — | — |
| pillaoff | 2.37% | +2.29pp | **<0.0001** |
| tau3 | 26.12% | +26.04pp | **<0.0001** |
| tau5 | 21.30% | +21.21pp | **<0.0001** |
| tau9 | 10.40% | +10.31pp | **<0.0001** |
| tau15 | 4.68% | +4.60pp | **<0.0001** |
MDE for this metric is 0.11pp, so all the shifts are real. Note the ordering:
`old` 0.09% → `pillaoff` 2.4% → `tau15` 4.7% → `tau9` 10.4% → `tau5` 21.3% →
`tau3` 26.1%. Removing the pillar opens the centre; the time-indexed heat (which
stops the bullet corridor at `speed·tau` instead of the wall) opens it much more,
and monotonically more as tau shrinks.
### 5b. Distance to the enemy (px, per-tick)
| arm | mean | p10 | median | p90 | 400+ px % of ticks | diff vs `old` | perm p |
|---|---:|---:|---:|---:|---:|---:|---:|
| old | 469.3 | 433 | 465 | 511 | 83.0 | — | — |
| pillaoff | 443.5 | 396 | 446 | 490 | 72.1 | −25.7 | **0.0002** |
| tau3 | 439.4 | 406 | 440 | 471 | 70.4 | −29.9 | **0.0004** |
| tau5 | 447.8 | 410 | 451 | 484 | 73.9 | −21.4 | **0.0014** |
| tau9 | 481.1 | 440 | 482 | 512 | 85.5 | +11.8 | 0.0813 |
| tau15 | 500.1 | 467 | 502 | 533 | 88.2 | +30.8 | **0.0001** |
MDE = 20.8 px. `old` fights at ~469 px; every opening of the centre pulls the
engagement 21–30 px closer (and 9–13pp more of the battle is inside 400 px), and
tau15 pushes it 31 px further out. `tau9` (the shipped tau) is essentially `old`
(+11.8 px, p=0.08).
### 5c. Bullet proximity — how much time we spend near live bullet paths
Nearest **live enemy** bullet (our own bullets are excluded: a bullet is born at
its own tank, so "any bullet" is dominated by our own just-fired shot; the
any-bullet version, per the task text, is in the ANY-BULLET PROXIMITY table of
`common_libs/tests/fixtures/tfil_heat_pillar_ab_mechanism.txt`).
| arm | ≤50 px | ≤100 px | ≤150 px | mean min-dist px | ≤100px diff vs `old` | perm p |
|---|---:|---:|---:|---:|---:|---:|
| old | 4.76% | 27.99% | 62.48% | 136.1 | — | — |
| pillaoff | 4.97% | 29.10% | 64.98% | 132.8 | +1.11pp | 0.1381 |
| tau3 | 6.19% | 30.86% | 61.39% | 134.8 | +2.87pp | **0.0014** |
| tau5 | 5.88% | 30.60% | 62.59% | 134.1 | +2.61pp | **0.0005** |
| tau9 | 4.99% | 28.31% | 60.75% | 137.8 | +0.32pp | 0.6066 |
| tau15 | 4.12% | 25.57% | 58.43% | 141.4 | −2.43pp | **0.0001** |
MDE = 1.74pp. tau3/tau5 spend significantly more time within 100 px of a live
enemy bullet; tau15 significantly less.
## 6. Liveness `[MEASURED]`
Every arm's declared env appears verbatim in the bot's own boot report
(`<arm>/run<N>.bot.stdout.log`, section A "raw process environment") in all 10
runs; no arm env was ignored as unrecognised. This is a *measurement*, not a
skip:
```
old OK (10/10 runs: TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 applied)
pillaoff OK (10/10 runs: no arm env; report present)
tau3 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 applied)
tau5 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 applied)
tau9 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 applied)
tau15 OK (10/10 runs: TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 applied)
```
The effective-value section confirms the reconstruction: `old` reports
`TR_TFIL_PILLAR_ON = on (source: env)` / `TR_TFIL_HEAT_TIME = off`, `pillaoff`
reports `TR_TFIL_PILLAR_ON = off (source: default)` / `TR_TFIL_HEAT_TIME = off
(source: default)`, and each tau arm reports the requested tau
(`TR_TFIL_HEAT_TAU = 3.0/5.0/9.0/15.0 (source: env)`).
## 7. Hit-rate context (NOT a verdict metric) `[MEASURED]`
| arm | pooled hit rate | round wins | our hits taken/run | our shots/run |
|---|---:|---:|---:|---:|
| old | 10.99% | 35/70 | 93.5 | 783 |
| pillaoff | 11.14% | 30/70 | 95.0 | 729 |
| tau3 | **11.56%** (best) | **20/70** (worst) | 106.2 | 666 |
| tau5 | 10.83% | 18/70 | 105.3 | 689 |
| tau9 | 10.51% | 22/70 | 99.5 | 750 |
| tau15 | **9.93%** (worst) | **32/70** (2nd best) | 89.5 | 812 |
The inversion is exact at both ends: the best-accuracy arm (`tau3`, 11.56%) wins
the fewest rounds; the worst-accuracy arm (`tau15`, 9.93%) wins the second most.
Hit rate is monotone in the WRONG direction here. Range-banded hit rates
(`..._bands.txt`) are flat in the 300–450 px band that holds 40% of our shots
(11.3–13.4%) and in 450+ (8.8–10.1%); no band rescues any heat-time arm.
## 8. Direct answers
**Does the heat-time model help, hurt, or do nothing? `[MEASURED] → HURTS.`**
* At the shipped tau (9) and at 3/5 it **significantly reduces round wins**
(20–22/70 vs 35/70, p=0.0010–0.0125) and reduces damage/run (p=0.004–0.047).
* `tau15` is a **wash on wins** (32/70, p=0.637) and 22 damage/run **lower**
(p=0.046) — not better, mildly worse.
* No tau improves either verdict metric. The mechanism is exactly what the
change claims (shorter corridor → centre opens, engagement closes by 20–30 px,
more time within 100 px of a live enemy bullet) — and the closer arms take
≈12 more hits/run (tau3 12.7, tau5 11.8) and lose 13–17 more rounds per 10
runs. **Inferred:** the free space this frees is the arena centre, and standing
there costs more than the wall-hugging flat model costs. The offline
"largest safe region" ruler rewarded exactly the region that shorter tau opens
(tau 2 → 162 px, tau 9 → 140 px, tau 15 → 127 px), and that region is a proxy
anti-correlated with survival against DrussGT.
**Does removing the pillar help, hurt, or do nothing? `[MEASURED] → no
measurable benefit; point estimates are worse, and it is undecidable at this n.`**
* `old` vs `pillaoff`: damage +5.7/run (p=0.710), wins +0.5/run (35 vs 30,
p=0.432), round-level Fisher p=0.498 — **no significant difference**.
* Damage taken is **30.8/run higher** without the pillar (p=0.040 permutation;
rank-sum cross-check p=0.104).
* The mechanism check proves the change **did** alter movement: box occupancy
0.09% → 2.37% (p<0.0001), mean range 469 → 443 px (p=0.0002). So this is a real
behaviour change, not a dead knob — it simply does not buy anything.
## 9. Which arm should be the shipped default?
**`old` — the pre-change mover (`TR_TFIL_PILLAR_ON=1` behaviour, heat-time off).
Nothing beats it: it is the best arm on damage/run (287) and round wins
(35/70), and second only to `tau15` on damage taken/run (202 vs 185), where
`tau15` pays for its 3 fewer round wins and 22 less damage per run.**
* **Heat-time: keep it OFF (as shipped).** The default is already off (`fca8993`);
no action needed, and `TR_TFIL_HEAT_TIME=1` at any tested tau should stay off.
* **Pillar removal: the shipped default should be REVERTED.** The shipped default
since `d0750ab` is `pillaoff`; it does **not** beat `old` on any verdict metric
and is directionally worse on all three (Δdamage −5.7, Δwins −0.5, Δdamage
taken +30.8). This is a decision for the user: restore the pre-change default
(`PillarHotness`/`PillarRadiance` back to 30/10, i.e. the `TR_TFIL_PILLAR_ON`
behaviour by default, keeping the env knob for the off state). **Honesty
caveat:** at n=10/arm the pillar contrast is inside the MDE (33 damage/run,
1.22 wins/run), so this is not a statistically significant "removal is worse";
it is "removal bought nothing measurable, and every point estimate moved the
wrong way". The case for reverting is the absence of evidence of benefit plus
parsimony, not a proven regression.
**Reproduce** (raw captures are not committed):
```sh
tools/ab/ab_run.sh --arms tools/ab/arms_heat_pillar.txt --runs 10 --rounds 7 \
--conc 7 --outdir /tmp/ab/heatshift
python3 tools/ab/ab_analyze.py /tmp/ab/heatshift --reference old
python3 tools/ab/ab_mechanism.py /tmp/ab/heatshift --reference old
python3 tools/ab/ab_range_bands.py /tmp/ab/heatshift --reference old
```
+404
View File
@@ -0,0 +1,404 @@
#!/usr/bin/env python3
"""ab_mechanism.py — MECHANISM CHECK for a movement A/B session (ab_run.sh).
python3 tools/ab/ab_mechanism.py <session_dir> [--reference ARM]
ab_analyze.py answers "did the arm win more?", this answers "did the arm
actually MOVE differently?" — because a win/loss difference is uninterpretable
if the knob never changed the trajectory.
Three measurements, all from the captured per-tick worldstate (both bots'
positions every tick) plus the fire/hit event sidecar:
1. CENTRAL-BOX OCCUPANCY. Fraction of ticks our tank spends inside the
144x144 px box centred on the arena centre (x 328-472, y 228-372) — the box
the removed virtual pillar covered. Spawns are bottom-left (us) / top (enemy),
never in the box, so any occupancy is genuine transit.
2. DISTANCE TO THE ENEMY. Per-tick shooter->target distance: mean, p10/median/
p90 and the fraction of ticks in 0-100 / 100-200 / 200-300 / 300-400 / 400+ px.
The flat heat model was measured to keep us at ~400 px; a change that works
should move this distribution.
3. BULLET PROXIMITY. Reconstruct every bullet from its fire event (x, y, dir,
power -> speed = 20-3p) and the resolving hit/hitwall/hitbullet event, and
take the per-tick distance from our tank to the NEAREST live bullet. Reported
as mean min-distance and the fraction of ticks with a bullet within 50/100/150
px. This is how much time we spend near live bullet paths.
All three are aggregated PER ROUND and then averaged over rounds (rounds differ
in length and in how early somebody dies, so a pooled mean would weight a long
lost round more than a short won one), and a two-sided permutation test on the
analysed runs' per-run means says whether an arm's behaviour differs from the
reference beyond run-to-run noise.
MEASURED = every number printed. INFERRED = the causal reading in the doc.
"""
from __future__ import annotations
import argparse
import itertools
import json
import math
import os
import random
import sys
ARENA_W, ARENA_H = 800.0, 600.0
BOX_X0, BOX_X1 = 328.0, 472.0 # 144x144 px centred on (400, 300)
BOX_Y0, BOX_Y1 = 228.0, 372.0
DIST_BANDS = [(0, 100), (100, 200), (200, 300), (300, 400), (400, 1e9)]
DIST_LABELS = ["0-100", "100-200", "200-300", "300-400", "400+"]
NEAR_PX = [50.0, 100.0, 150.0]
SPEED_A, SPEED_B = 20.0, 3.0
US_SIDE = "s" # adversary (ModularBot) = s* (see the capture)
MC_SEED = 0x5EED5EED
MC_DRAWS = 200_000
Z_ALPHA_POWER = 1.959963984540054 + 0.8416212335729143
def rows_of(path):
out = []
with open(path) as f:
for line in f:
line = line.strip()
if not line:
continue
o = json.loads(line)
if "tick" in o:
out.append(o)
return out
def events_of(path):
return [json.loads(l) for l in open(path) if l.strip()]
def resolve_owner_side(rows_by_tick, rounds, events):
"""owner id -> 's' (us) / 'e' (DrussGT), by matching each fire event's
(x, y) to a capture row within +-8 ticks of `startTick + ev.tick`."""
start = {r["round"]: r["startTick"] for r in rounds}
votes = {}
for ev in events:
if ev.get("type") != "fire":
continue
g = start.get(ev["round"], 0) + ev["tick"]
for t in range(g - 8, g + 9):
r = rows_by_tick.get(t)
if r is None:
continue
for side in ("s", "e"):
if (abs(r[side + "x"] - ev["x"]) <= 0.02
and abs(r[side + "y"] - ev["y"]) <= 0.02):
d = votes.setdefault(ev["owner"], {"s": 0, "e": 0})
d[side] += 1
return {o: ("s" if d["s"] >= d["e"] else "e") for o, d in votes.items()}
def analyse_run(cap_path, ev_path, rj_path):
rows = rows_of(cap_path)
if not rows:
return None
by_tick = {r["tick"]: r for r in rows}
rounds = json.load(open(rj_path))["rounds"]
events = events_of(ev_path)
oside = resolve_owner_side(by_tick, rounds, events)
# bullet tracks: (round, owner, bullet) -> start tick / direction / speed.
# A live window is [t0, resolving event] (hit / hitwall / hitbullet); a
# bullet with no resolver (round ended mid-flight) is bounded at t0 + 400.
start = {r["round"]: r["startTick"] for r in rounds}
bullets = {}
for ev in events:
key = (ev["round"], ev.get("owner"), ev.get("bullet"))
if ev.get("type") == "fire":
p = ev["power"]
v = SPEED_A - SPEED_B * p
th = math.radians(ev["dir"])
bullets[key] = {
"t0": start[ev["round"]] + ev["tick"],
"x": ev["x"], "y": ev["y"],
"ux": math.cos(th), "uy": math.sin(th), "v": v,
"side": oside.get(ev["owner"]), "end": None,
}
elif ev.get("type") in ("hit", "hitwall", "hitbullet"):
if key in bullets:
bullets[key]["end"] = start[ev["round"]] + ev["tick"]
# per-round accumulators
per_round = []
for rd in rounds:
a, n = rd["startTick"], rd["count"]
ticks = [t for t in range(a, a + n) if t in by_tick]
if not ticks:
continue
box = near = near_e = 0
dsum = 0.0
mindist_sum = 0.0
mindist_e_sum = 0.0
band_cnt = [0] * len(DIST_BANDS)
near_cnt = [0] * len(NEAR_PX)
near_e_cnt = [0] * len(NEAR_PX)
live = [(k, b) for k, b in bullets.items() if k[0] == rd["round"]]
for t in ticks:
r = by_tick[t]
sx, sy = r[US_SIDE + "x"], r[US_SIDE + "y"]
ox, oy = r["ex"], r["ey"]
if BOX_X0 <= sx <= BOX_X1 and BOX_Y0 <= sy <= BOX_Y1:
box += 1
d = math.hypot(ox - sx, oy - sy)
dsum += d
for i, (lo, hi) in enumerate(DIST_BANDS):
if lo <= d < hi:
band_cnt[i] += 1
break
md = None
mde_ = None
for _k, b in live:
if b["t0"] is None or t < b["t0"]:
continue
te = b["end"] if b["end"] is not None else b["t0"] + 400
if t > te:
continue
dt = t - b["t0"]
bx = b["x"] + b["v"] * dt * b["ux"]
by = b["y"] + b["v"] * dt * b["uy"]
if not (0.0 <= bx <= ARENA_W and 0.0 <= by <= ARENA_H):
continue
dd = math.hypot(bx - sx, by - sy)
if md is None or dd < md:
md = dd
if b["side"] == "e" and (mde_ is None or dd < mde_):
mde_ = dd
if md is not None:
near += 1
mindist_sum += md
for i, x in enumerate(NEAR_PX):
if md <= x:
near_cnt[i] += 1
if mde_ is not None:
near_e += 1
mindist_e_sum += mde_
for i, x in enumerate(NEAR_PX):
if mde_ <= x:
near_e_cnt[i] += 1
nt = len(ticks)
per_round.append({
"round": rd["round"], "ticks": nt,
"box_frac": box / nt,
"mean_dist": dsum / nt,
"dist_bands": [c / nt for c in band_cnt],
"mean_minbullet": (mindist_sum / near if near else float("nan")),
"near_frac": [c / nt for c in near_cnt],
"mean_minbullet_enemy": (mindist_e_sum / near_e if near_e
else float("nan")),
"near_enemy_frac": [c / nt for c in near_e_cnt],
"ticks_with_bullet_frac": near / nt,
"ticks_with_enemy_bullet_frac": near_e / nt,
})
return {"cap": cap_path, "per_round": per_round, "owner_side": oside,
"n_rounds": len(per_round)}
def _mean(xs):
return sum(xs) / len(xs) if xs else float("nan")
def _sd(xs):
if len(xs) < 2:
return 0.0
m = _mean(xs)
return math.sqrt(sum((v - m) ** 2 for v in xs) / (len(xs) - 1))
def perm_test(xa, xb):
na, nb = len(xa), len(xb)
if na == 0 or nb == 0:
return None
obs = abs(_mean(xa) - _mean(xb))
pooled = list(xa) + list(xb)
n = na + nb
total = sum(pooled)
ncomb = math.comb(n, na)
if ncomb <= 20_000_000:
cnt = 0
for combo in itertools.combinations(range(n), na):
sa = sum(pooled[i] for i in combo)
if abs(sa / na - (total - sa) / nb) >= obs - 1e-9:
cnt += 1
return obs, cnt / ncomb, "exact"
rng = random.Random(MC_SEED)
cnt = 0
for _ in range(MC_DRAWS):
sa = sum(pooled[i] for i in rng.sample(range(n), na))
if abs(sa / na - (total - sa) / nb) >= obs - 1e-9:
cnt += 1
return obs, (cnt + 1) / (MC_DRAWS + 1), f"MC/B={MC_DRAWS:,}"
def main():
ap = argparse.ArgumentParser()
ap.add_argument("session_dir")
ap.add_argument("--reference", default=None)
args = ap.parse_args()
root = args.session_dir
arms = [d for d in sorted(os.listdir(root))
if os.path.isdir(os.path.join(root, d)) and d != "frozen"
and not d.startswith(".") and any(
f.endswith(".jsonl") and not f.endswith(".events.jsonl")
for f in os.listdir(os.path.join(root, d)))]
ref = args.reference or arms[0]
data = {}
for arm in arms:
d = os.path.join(root, arm)
runs = []
for fn in sorted(os.listdir(d)):
if not fn.endswith(".jsonl") or fn.endswith(".events.jsonl"):
continue
cap = os.path.join(d, fn)
ev, rj = cap[:-6] + ".events.jsonl", cap + ".rounds.json"
if os.path.exists(ev) and os.path.exists(rj):
a = analyse_run(cap, ev, rj)
if a:
runs.append(a)
data[arm] = runs
print("=" * 100)
print("MECHANISM CHECK — did the arm actually move differently?")
print(f"session {root} arms {', '.join(arms)} reference {ref}")
print("box = 144x144 px centred on the arena centre (x 328-472, y 228-372)")
print("=" * 100)
metrics = [
("box_frac", "central-box occupancy", 100.0, "%"),
("mean_dist", "mean dist to enemy", 1.0, "px"),
("near_enemy_frac[1]", "ticks enemy bullet<=100px", 100.0, "%"),
("mean_minbullet_enemy", "mean dist nearest enemy bullet", 1.0, "px"),
("near_frac[1]", "ticks ANY bullet<=100px", 100.0, "%"),
]
def per_run(arm, key, band=None):
out = []
for a in data[arm]:
vals = []
for pr in a["per_round"]:
if key.startswith("near_enemy_frac["):
vals.append(pr["near_enemy_frac"][int(key[16])])
elif key.startswith("near_frac["):
vals.append(pr["near_frac"][int(key[10])])
else:
vals.append(pr[key])
out.append(_mean(vals))
return out
print("\nPER-ARM MECHANISM (mean of the per-round values, then over runs)")
hdr = (f"{'arm':<10} {'runs':>4} {'box%':>7} {'dist px':>8} "
f"{'p10':>6} {'med':>6} {'p90':>6} {'enB<=100px%':>13} "
f"{'enBmin px':>11}")
print(hdr)
print("-" * len(hdr))
for arm in arms:
runs = data[arm]
if not runs:
print(f"{arm:<10} {0:>4} (no runs)")
continue
boxf = [_mean([pr["box_frac"] for pr in a["per_round"]]) for a in runs]
md = [_mean([pr["mean_dist"] for pr in a["per_round"]]) for a in runs]
nb = [_mean([pr["near_enemy_frac"][1] for pr in a["per_round"]])
for a in runs]
mb = [_mean([pr["mean_minbullet_enemy"] for pr in a["per_round"]])
for a in runs]
allmd = sorted(v for a in runs for v in
[pr["mean_dist"] for pr in a["per_round"]])
p10 = allmd[int(0.10 * (len(allmd) - 1))]
med = allmd[len(allmd) // 2]
p90 = allmd[int(0.90 * (len(allmd) - 1))]
print(f"{arm:<10} {len(runs):>4} {100*_mean(boxf):>7.2f} "
f"{_mean(md):>8.1f} {p10:>6.0f} {med:>6.0f} {p90:>6.0f} "
f"{100*_mean(nb):>13.2f} {_mean(mb):>11.1f}")
print("\nDISTANCE-TO-ENEMY BANDS (fraction of ticks, mean over rounds/runs, %)")
hdr = f"{'arm':<10}" + "".join(f"{lb:>10}" for lb in DIST_LABELS)
print(hdr)
print("-" * len(hdr))
for arm in arms:
runs = data[arm]
if not runs:
continue
cells = []
for i in range(len(DIST_BANDS)):
v = _mean([_mean([pr["dist_bands"][i] for pr in a["per_round"]])
for a in runs])
cells.append(f"{100*v:>10.2f}")
print(f"{arm:<10}" + "".join(cells))
print("\nENEMY-BULLET PROXIMITY (nearest LIVE ENEMY bullet; our own bullets are "
"excluded because a bullet is born at its own tank, so 'any bullet' "
"is dominated by our just-fired shot)")
hdr = (f"{'arm':<10}" + "".join(f"{('<=%gpx' % x):>9}" for x in NEAR_PX)
+ f"{'min px':>9}{'enemy%':>9}")
print(hdr)
print("-" * len(hdr))
for arm in arms:
runs = data[arm]
if not runs:
continue
cells = []
for i in range(len(NEAR_PX)):
v = _mean([_mean([pr["near_enemy_frac"][i] for pr in a["per_round"]])
for a in runs])
cells.append(f"{100*v:>9.2f}")
mb = _mean([_mean([pr["mean_minbullet_enemy"] for pr in a["per_round"]])
for a in runs])
anyv = _mean([_mean([pr["ticks_with_enemy_bullet_frac"]
for pr in a["per_round"]]) for a in runs])
print(f"{arm:<10}" + "".join(cells) + f"{mb:>9.1f}{100*anyv:>9.2f}")
print("\nANY-BULLET PROXIMITY (per the task text: nearest live bullet, ours "
"included; context only)")
hdr = (f"{'arm':<10}" + "".join(f"{('<=%gpx' % x):>9}" for x in NEAR_PX)
+ f"{'min px':>9}{'any%':>9}")
print(hdr)
print("-" * len(hdr))
for arm in arms:
runs = data[arm]
if not runs:
continue
cells = []
for i in range(len(NEAR_PX)):
v = _mean([_mean([pr["near_frac"][i] for pr in a["per_round"]])
for a in runs])
cells.append(f"{100*v:>9.2f}")
anyv = _mean([_mean([pr["ticks_with_bullet_frac"] for pr in a["per_round"]])
for a in runs])
mb = _mean([_mean([pr["mean_minbullet"] for pr in a["per_round"]])
for a in runs])
print(f"{arm:<10}" + "".join(cells) + f"{mb:>9.1f}{100*anyv:>9.2f}")
print(f"\nPER-RUN MECHANISM vs `{ref}` (two-sided permutation on per-run means; "
f"exact when C(n,na)<=2e7)")
print(f"{'metric':<24} {'arm':<10} {'delta':>9} {'p':>9} {'method':<12} "
f"{'MDE':>9}")
print("-" * 78)
for key, label, scale, unit in metrics:
xa = per_run(ref, key)
if len(xa) < 2:
continue
sd = _sd(xa)
mde = Z_ALPHA_POWER * sd * math.sqrt(2.0 / len(xa)) * scale
for arm in arms:
if arm == ref:
continue
xb = per_run(arm, key)
res = perm_test(xa, xb)
if res is None:
continue
obs, p, method = res
signed = (_mean(xb) - _mean(xa)) * scale
print(f"{label:<24} {arm:<10} {signed:>+9.2f} {p:>9.4f} "
f"{method:<12} {mde:>8.2f}{unit}")
return 0
if __name__ == "__main__":
sys.exit(main())
+18
View File
@@ -0,0 +1,18 @@
# Heat-time bullet model + virtual-pillar removal A/B (6 arms x 10 runs x 7 rounds).
#
# `git archive HEAD` now has BOTH changes in their shipped state: the virtual
# centre pillar is REMOVED (d0750ab) and the time-indexed bullet heat is OFF
# (fca8993). So the PRE-change mover has to be reconstructed with env.
#
# old = pre-change mover: virtual centre pillar ON, flat bullet heat
# pillaoff = shipped default as of HEAD: pillar removed, flat bullet heat
# tauN = pillar removed + time-indexed heat with tau = N ticks
#
# old vs pillaoff isolates the PILLAR removal.
# pillaoff vs tau9 isolates the HEAT-TIME change (tau9 is the default tau).
old | TR_TFIL_PILLAR_ON=1 TR_TFIL_HEAT_TIME=0 | pre-change mover (virtual centre pillar ON)
pillaoff | | shipped default at HEAD (pillar removed, flat bullet heat)
tau3 | TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=3 | time-indexed heat, tau=3 (short corridor)
tau5 | TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=5 | time-indexed heat, tau=5
tau9 | TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=9 | time-indexed heat, tau=9 (shipped default tau)
tau15 | TR_TFIL_HEAT_TIME=1 TR_TFIL_HEAT_TAU=15 | time-indexed heat, tau=15 (long corridor)