Files
SirRoboGarage/docs/movement_campaign.md
T

81 KiB
Raw Blame History

Movement campaign — ledger

Goal (owner's mandate, 2026-09-26 overnight): find the best 1v1 movement by measurement, then do the same for the gun. This file is the campaign's single source of truth: every later job appends a ## Batch N section and never edits an earlier one (a wrong earlier number gets a correction line, not a rewrite).

Owner's words: "I want you to do all tests and checks with the goal to have the best 1vs1 movement. You have all night, you can change every parameter. Continue until you found an amazing movement. When found do the same over for a gun."


0. The one caveat this campaign exists to close

Everything measured about movement before this campaign is DrussGT-only: docs/surfer_wiring_ab.md, the j107 range drift, the j113 BitBrain movement notes. The standing lesson of the night is that a one-opponent result is not a result:

an arm can take fewer hits and win fewer rounds (j107 / strafe): the verdict lives in damage/run + ROUND WINS, and hit rate is only ever an explanation.

So from here on the unit of evidence is the number of opponents, not the number of runs: the same arm must win on many opponents before it is called better.


1. Protocol (how every batch must be run)

Element Rule
Subject ONE frozen binary, built from git archive HEAD (tools/ab/tournament_run.sh does this; the commit sha and binary sha256 are recorded in session.json)
Arms env dicts only — no per-arm rebuild, ever; the arm file is a committed file, not a shell history
Panel the frozen panel tools/ab/panel_movement.txt. Adding/removing an opponent starts a new batch number
Pairing per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate
Isolation per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir
Serialization one battle fleet at a time. tournament_run.sh --wait-arena N refuses/stalls while another job's run_bridge_battle/TrBattleCapture/ModularBot_bin is alive (bracketed pgrep; never a broad pkill)
Liveness every declared env token must appear verbatim in OUR bot's own [env] boot report, else the run is excluded and named in the report; an undeclared TR_MOVEMENT in the process env is a fatal FAIL for the reference arm
Never shipped this is a measurement + design campaign: git status clean, defaults untouched, .gitignore untouched

Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit 1984a78)

Provenance note: the harness and these rules were written and staged before Batch 1 was fought, but a parallel job's git commit (j116, same working tree / same index) swept the staged files into its commit 1984a78 ("melee A/B doc…"). The rules are therefore committed under a neighbour's message — they are nonetheless dated before the data: no battle of Batch 1 had been launched when they were written, and Batch 1's session.json records the same commit 1984a78 as the frozen-binary source.

  1. Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean distance.
  2. BETTER than the reference iff one primary metric is up with a cross-opponent sign test p < 0.05 while the other does not go down; or the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (which is a real answer, not a failure).
  3. A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as not detectable, never as absent and never as a win.
  4. Somewhere to stop: if no arm beats the shipped tfil by rule 2 in Batch 1 and no arm shows a ≥ +MDE damage gain with p<0.10, the movement stage's first phase is closed with "the shipped tfil is the best movement we have measured" — that is a successful outcome, and the campaign moves to the gun axis rather than inventing more movement arms. See What would make us stop at the end.
  5. No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
  6. Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.

2. Stage 0 — what we already know (given, not re-derived)

Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary (docs/surfer_wiring_ab.md, commit 0f5cfe3):

arm dmg/run dmg taken round wins incoming hit rate
tfil (SHIPPED) 293 224 45/105 10.40%
strafe (range 325) 250 198 37/105 9.40%
surf 255 259 37/105 13.51%

Read: the shipped tfil deals the most damage and wins the most rounds while being hit the most; strafe dodges best and wins least. Plus j107: drifting 25–30 px closer made damage and wins worse, so the lever is not simply "get closer". Hypothesis entering the campaign: the 325 px range preference of strafe costs wins (INFERRED from DrussGT-only data — this is exactly what Batch 1 tests across a panel).


3. Batch 1 — isolating the range / aggression axis

Design. One frozen binary, five env-only arms, one frozen panel (tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 2 wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas), 3 runs × 3 rounds per (opponent, arm). Arm file: tools/ab/arms_movement_b1.txt.

# arm env what it isolates
1 tfil (none — shipped defaults) the arm to beat
2 strafe_notilt TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999 the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with no range steering at all
3 strafe_325 TR_MOVEMENT=strafe the current strafe default (range 325, tol 25, tilt 15/0.10)
4 ring TR_MOVEMENT=tfil_ring TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) with the range-weighted tile draw (band 100–200)
5 ring_notemp TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0 the control for #4: same retuned heat field, range weighting switched OFF (rand(candidates.high) path)

ring − ring_notemp is therefore the range-weighting lever alone, on a heat field that is already retuned. The originally-suggested 5th arm ("tfil with less saturated heat") is not buildable in this campaign: in common_libs/movements/the_floor_is_lava.nim CorridorHeat/WallHotness are Nim consts (env_report only reports them); only the tfil_ring copy reads them from the env. #5 is the honest substitute.

Pre-registered prediction (written before the battles finished): tfil still wins the panel on damage and round wins; strafe_notilt will beat strafe_325 on round wins (the range tilt is a net cost), and the ring arms will land between them. If instead the range-steering arms beat tfil on wins, the "range preference costs wins" hypothesis is confirmed across bots, not just against DrussGT.

Outcome — direct answer

Batch 1 is a NULL for the hypothesis that the shipped tfil is the best movement. It is not. Measured on the frozen 15-opponent panel, one frozen binary, 225 battles, 0 invalid runs, 0 liveness failures, 0 failed starts:

arm dmg/run wins/run round wins incoming hit rate dmg taken/run mean distance
tfil (SHIPPED) 118.9 1.22 55/135 (40.7%) 18.17% 199.8 382 px
strafe_notilt 108.8 1.60 72/135 (53.3%) 12.24% 150.3 456 px
strafe_325 111.8 1.56 70/135 (51.9%) 13.14% 155.7 436 px
ring_notemp 108.2 1.29 58/135 (43.0%) 16.67% 193.4 395 px
ring 150.1 1.18 53/135 (39.3%) 29.42% 225.1 236 px

Paired across opponents, the winner is strafe_notilt (pure perpendicular strafe, range steering provably off): Δwins/run +0.38 [95% CI +0.16, +0.60], positive on 9 of 9 decisive opponents (exact sign test p = 0.0039, sign-flip permutation p = 0.0039, Wilcoxon p = 0.0090), and Δdmg/run −10.2 [−25.8, +5.5], p = 0.61, MDE 20.4 ⇒ not detectable — i.e. +17 rounds out of 135 won, at no detectable damage cost, with a third fewer incoming hits (hit rate −7.3 pp, p = 6e-5, and 0/15 opponents in favour of tfil) and 50 less damage taken per run. strafe_325 is the same effect, slightly smaller (Δwins/run +0.33, [0.04, +0.63], p = 0.039, 10/12) — the two strafe arms are not separable from each other by this batch.

ring is the opposite trade and must not be read as a movement win: it deals +31.2 dmg/run (+26%, p = 0.0074, 13/15) but wins no more rounds (Δwins −0.04, p = 1.00) and pays for the damage with the panel's worst dodging (hit rate 29.42% vs 18.17%, +25 dmg taken/run) because it fights at a mean 236 px (vs 382/456). ring_notemp — the same retuned heat field with the range weighting switched off — is indistinguishable from tfil on both primaries, so the heat-field retune alone is not what makes strafe win (INFERRED: ring_notemp also differs from tfil in commit ticks and wall radiance, so this is evidence against, not a clean isolation).

The cleanest aggression isolation in the batch is ring − ring_notemp (same engine, same retuned heat field, only the range-weighted tile draw differs, band 100–200): that lever alone is worth +41.9 dmg/run (150.1 vs 108.2), −0.11 wins/run (1.18 vs 1.29) and +12.8 pp incoming hit rate (29.42% vs 16.67%) at 236 vs 395 px. Engaging harder converts into damage, never into wins, and pays with hits.

Mechanism (MEASURED, and the reason the win is a movement win): in 216 of the 219 attributable runs, our round-win count equals exactly the number of rounds in which the opponent's death event appears — round wins in this harness are survival wins. The winning arm survives by taking fewer, weaker hits at longer range, not by dealing more damage (its damage is unchanged).

DIRECT ANSWER. The best 1v1 movement measured across this panel is TR_MOVEMENT=strafe with the range tilt disabled (pure perpendicular strafe, no range steering). It beats the shipped tfil on round wins by an effect that survives the between-opponent spread (observed +0.38 vs MDE 0.29; 9/9 opponents; CI excludes 0) with no detectable damage cost, and it dodges substantially better. strafe_325 (the current strafe default) is essentially the same arm. The shipped tfil is 4th of the five on round wins (only ring is nominally lower, and tfil vs ring on wins is a dead heat, p = 1.00): the hypothesis in §2 that its win came from the DrussGT-only measurement is supported — on a panel it loses to both strafe arms.

Correction (added after the Batch-1 commit 0776630, whose message says "last of five"): tfil is 4th of five, not last — ring is nominally 0.04 wins/run lower and that difference is not significant. The batch message overstates one word; the numbers it quotes are the measured ones.

The pre-registered prediction for this batch was WRONG and is recorded as wrong: I predicted tfil would still win the panel (it came 4th of five on wins) and that strafe_notilt would beat strafe_325 on wins (it does by +0.05 wins/run, which this batch cannot resolve).

Honest readings of the pre-registered rule (both printed by the analyzer; the strict reading is the literal one and it is NOT satisfied by anything):

  • strict (the other metric's mean delta is not negative at all): no arm is BETTER than tfil. The two strafe arms win more rounds but their mean damage is 7–10/run lower (inside the MDE, but negative).
  • substantive (the other primary metric is not detectably down — sign test not significant and |Δ| < its MDE, per rule 3): strafe_notilt, strafe_325 and ring are each BETTER than tfil on one primary metric.
  • The ordering is identical under both readings, and under the standing rule (round wins first, then damage) the winner is strafe_notilt.

The analyzer's full report (verbatim)

MEASURED: session

  • commit 1984a780f494ce246e0f916934b9581e07c89ed2, frozen binary sha256 1817c75ab1d0…

  • 15 opponents × 5 arms × 3 runs × 3 rounds = 225 battles, conc=6

  • arms file arms_movement_b1.txt, panel file panel_movement.txt

  • reference arm: tfil — every delta below is (arm − tfil), opponent by opponent

  • liveness: 0 run(s) excluded (225 total)

MEASURED: per-opponent paired table (per arm)

tfil — shipped baseline (movement engine tfil, every knob at its default) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→124.7 +0.0 1.67→1.67 +0.00 +0.0 +0.00 452→452
Diamond dodger 39.5→39.5 +0.0 0.00→0.00 +0.00 +0.0 +0.00 458→458
Dookious dodger 130.1→130.1 +0.0 1.00→1.00 +0.00 +0.0 +0.00 410→410
GresSuffurd dodger 129.2→129.2 +0.0 1.33→1.33 +0.00 +0.0 +0.00 413→413
CassiusClay dodger 73.4→73.4 +0.0 0.33→0.33 +0.00 +0.0 +0.00 339→339
RetroGirl pattern 181.0→181.0 +0.0 2.33→2.33 +0.00 +0.0 +0.00 402→402
TripHammer pattern 59.5→59.5 +0.0 0.33→0.33 +0.00 +0.0 +0.00 418→418
Coriantumr pattern 67.8→67.8 +0.0 1.00→1.00 +0.00 +0.0 +0.00 424→424
WallAvoider wallfollower 229.1→229.1 +0.0 2.67→2.67 +0.00 +0.0 +0.00 277→277
HawkOnFire cornercamper 150.7→150.7 +0.0 1.67→1.67 +0.00 +0.0 +0.00 410→410
SpinBot spinner 279.3→279.3 +0.0 3.00→3.00 +0.00 +0.0 +0.00 351→351
DiamondStealer rammer 139.4→139.4 +0.0 0.67→0.67 +0.00 +0.0 +0.00 235→235
BlitzBat brawler 54.3→54.3 +0.0 2.00→2.00 +0.00 +0.0 +0.00 422→422
YersiniaPestis aggressive 52.3→52.3 +0.0 0.33→0.33 +0.00 +0.0 +0.00 401→401
Ascendant aggressive 73.3→73.3 +0.0 0.00→0.00 +0.00 +0.0 +0.00 317→317

strafe_notilt — strafe, range steering OFF (tilt always 0) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→115.7 -9.0 1.67→1.67 +0.00 -31.6 -2.95 452→535
Diamond dodger 39.5→56.6 +17.1 0.00→0.00 +0.00 -62.9 -7.97 458→543
Dookious dodger 130.1→105.8 -24.3 1.00→1.00 +0.00 -60.1 -6.08 410→484
GresSuffurd dodger 129.2→111.3 -17.8 1.33→2.33 +1.00 -36.3 -8.61 413→482
CassiusClay dodger 73.4→92.1 +18.7 0.33→1.33 +1.00 -55.5 -7.24 339→397
RetroGirl pattern 181.0→133.2 -47.8 2.33→2.67 +0.33 -0.9 -2.46 402→447
TripHammer pattern 59.5→57.5 -2.1 0.33→0.67 +0.33 -45.0 -3.84 418→547
Coriantumr pattern 67.8→77.2 +9.4 1.00→1.67 +0.67 -54.2 -3.35 424→554
WallAvoider wallfollower 229.1→162.1 -66.9 2.67→2.67 +0.00 -54.9 -9.49 277→361
HawkOnFire cornercamper 150.7→95.7 -55.0 1.67→1.67 +0.00 -76.2 -7.27 410→567
SpinBot spinner 279.3→271.3 -8.0 3.00→3.00 +0.00 -42.7 -14.27 351→335
DiamondStealer rammer 139.4→154.7 +15.2 0.67→1.00 +0.33 -33.7 -2.80 235→243
BlitzBat brawler 54.3→41.5 -12.8 2.00→3.00 +1.00 -105.6 -13.42 422→588
YersiniaPestis aggressive 52.3→78.4 +26.0 0.33→1.00 +0.67 -45.0 -6.09 401→397
Ascendant aggressive 73.3→78.2 +4.9 0.00→0.33 +0.33 -36.9 -14.30 317→364

strafe_325 — strafe default (range 325, tol 25, tilt 15/0.10) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→118.5 -6.2 1.67→1.33 -0.33 -17.5 -1.77 452→494
Diamond dodger 39.5→71.8 +32.2 0.00→0.00 +0.00 -38.7 -5.93 458→506
Dookious dodger 130.1→84.0 -46.0 1.00→2.00 +1.00 -108.4 -9.12 410→457
GresSuffurd dodger 129.2→112.3 -16.9 1.33→2.00 +0.67 -26.1 -5.44 413→430
CassiusClay dodger 73.4→81.5 +8.1 0.33→1.33 +1.00 -68.4 -7.31 339→386
RetroGirl pattern 181.0→169.4 -11.6 2.33→2.33 +0.00 -7.9 -3.52 402→425
TripHammer pattern 59.5→43.6 -15.9 0.33→0.67 +0.33 -43.0 -4.74 418→500
Coriantumr pattern 67.8→96.5 +28.7 1.00→1.67 +0.67 -46.2 -3.37 424→494
WallAvoider wallfollower 229.1→163.8 -65.2 2.67→1.67 -1.00 -12.1 -7.90 277→332
HawkOnFire cornercamper 150.7→119.8 -30.9 1.67→2.00 +0.33 -78.8 -5.38 410→517
SpinBot spinner 279.3→259.3 -20.0 3.00→3.00 +0.00 -37.3 -12.96 351→436
DiamondStealer rammer 139.4→135.3 -4.1 0.67→1.33 +0.67 -25.4 -3.50 235→274
BlitzBat brawler 54.3→60.5 +6.3 2.00→2.67 +0.67 -88.1 -10.88 422→531
YersiniaPestis aggressive 52.3→69.3 +16.9 0.33→0.67 +0.33 -15.7 -2.95 401→383
Ascendant aggressive 73.3→90.4 +17.1 0.00→0.67 +0.67 -47.0 -12.83 317→373

ring — tfil_ring (retuned heat field + range weighting 100-200) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→127.8 +3.1 1.67→0.67 -1.00 +105.6 +12.67 452→244
Diamond dodger 39.5→46.4 +6.9 0.00→0.00 +0.00 +40.9 +17.58 458→240
Dookious dodger 130.1→149.2 +19.2 1.00→0.33 -0.67 +50.0 +8.21 410→267
GresSuffurd dodger 129.2→222.7 +93.6 1.33→1.33 +0.00 +83.9 +12.13 413→234
CassiusClay dodger 73.4→102.5 +29.1 0.33→0.33 +0.00 +21.8 +5.21 339→242
RetroGirl pattern 181.0→249.8 +68.8 2.33→3.00 +0.67 -41.9 +0.85 402→209
TripHammer pattern 59.5→76.2 +16.7 0.33→0.00 -0.33 +47.8 +13.66 418→266
Coriantumr pattern 67.8→119.4 +51.7 1.00→1.00 +0.00 +32.1 +10.39 424→258
WallAvoider wallfollower 229.1→219.5 -9.6 2.67→1.67 -1.00 +57.2 +7.21 277→232
HawkOnFire cornercamper 150.7→176.4 +25.7 1.67→2.33 +0.67 -41.8 +11.76 410→229
SpinBot spinner 279.3→336.0 +56.7 3.00→3.00 +0.00 +5.3 +24.20 351→172
DiamondStealer rammer 139.4→148.7 +9.3 0.67→1.33 +0.67 -39.4 -1.78 235→214
BlitzBat brawler 54.3→156.7 +102.5 2.00→2.67 +0.67 +8.4 +12.71 422→221
YersiniaPestis aggressive 52.3→43.3 -9.0 0.33→0.00 -0.33 +36.2 +10.82 401→266
Ascendant aggressive 73.3→76.7 +3.4 0.00→0.00 +0.00 +14.0 +13.54 317→243

ring_notemp — tfil_ring, range weighting OFF (temp 0) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→108.0 -16.7 1.67→1.33 -0.33 -0.8 -0.72 452→436
Diamond dodger 39.5→42.9 +3.4 0.00→0.00 +0.00 +9.2 -1.47 458→447
Dookious dodger 130.1→115.2 -14.8 1.00→1.33 +0.33 -24.4 -3.35 410→418
GresSuffurd dodger 129.2→118.3 -10.8 1.33→1.33 +0.00 +20.9 -1.75 413→433
CassiusClay dodger 73.4→92.6 +19.2 0.33→1.67 +1.33 -51.0 -5.72 339→388
RetroGirl pattern 181.0→126.2 -54.8 2.33→1.33 -1.00 +55.3 +3.00 402→405
TripHammer pattern 59.5→44.3 -15.2 0.33→0.33 +0.00 -4.9 +0.65 418→453
Coriantumr pattern 67.8→84.9 +17.1 1.00→1.00 +0.00 -1.7 +0.24 424→430
WallAvoider wallfollower 229.1→141.1 -88.0 2.67→3.00 +0.33 -124.4 -14.57 277→339
HawkOnFire cornercamper 150.7→134.2 -16.5 1.67→1.33 -0.33 +28.8 +1.77 410→436
SpinBot spinner 279.3→287.7 +8.3 3.00→3.00 +0.00 +0.0 +0.45 351→308
DiamondStealer rammer 139.4→154.5 +15.1 0.67→2.00 +1.33 -57.4 -6.08 235→254
BlitzBat brawler 54.3→57.3 +3.1 2.00→1.67 -0.33 +57.7 +2.97 422→436
YersiniaPestis aggressive 52.3→45.2 -7.1 0.33→0.00 -0.33 +21.2 +2.66 401→376
Ascendant aggressive 73.3→70.3 -3.0 0.00→0.00 +0.00 -23.2 -8.01 317→371

MEASURED: pooled dashboard (all valid runs, NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
tfil 45 118.9 199.8 1.22 55/135 40.7% 18.17% 382
strafe_notilt 45 108.8 150.3 1.60 72/135 53.3% 12.24% 456
strafe_325 45 111.8 155.7 1.56 70/135 51.9% 13.14% 436
ring 45 150.1 225.1 1.18 53/135 39.3% 29.42% 236
ring_notemp 45 108.2 193.4 1.29 58/135 43.0% 16.67% 395

MEASURED: cross-opponent aggregation (the verdict layer)

Deltas are per-opponent (arm − reference). spread is the SD of those deltas ACROSS opponents; SE = spread/√n; 95% CI = mean ± t·SE. Sign test = how many opponents the arm wins (ties dropped), exact binomial; sign-flip = permutation test on the mean of the deltas.

arm metric mean Δ spread (SD) SE 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
strafe_notilt damage -10.15 28.19 7.28 [-25.76, +5.46] 6/15 0.6072 0.1887 (exact 2^15) 0.3787 20.39
strafe_notilt wins +0.38 0.40 0.10 [+0.16, +0.60] 9/9 0.003906 0.003906 (exact 2^15) 0.008969 0.29
strafe_notilt damage_taken -49.44 23.28 6.01 [-62.33, -36.54] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 16.84
strafe_notilt hit_rate -7.34 4.10 1.06 [-9.61, -5.07] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 2.96
strafe_notilt dist +74.36 54.83 14.16 [+44.00, +104.73] 13/15 0.007385 0.0003662 (exact 2^15) 0.001621 39.66
strafe_325 damage -7.15 27.04 6.98 [-22.13, +7.82] 6/15 0.6072 0.3276 (exact 2^15) 0.5137 19.56
strafe_325 wins +0.33 0.53 0.14 [+0.04, +0.63] 10/12 0.03857 0.04688 (exact 2^15) 0.05424 0.39
strafe_325 damage_taken -44.05 29.86 7.71 [-60.58, -27.51] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 21.60
strafe_325 hit_rate -6.51 3.58 0.92 [-8.49, -4.53] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 2.59
strafe_325 dist +53.77 33.71 8.70 [+35.10, +72.44] 14/15 0.0009766 0.0001831 (exact 2^15) 0.001092 24.38
ring damage +31.19 35.61 9.19 [+11.47, +50.91] 13/15 0.007385 0.001587 (exact 2^15) 0.004932 25.76
ring wins -0.04 0.56 0.14 [-0.36, +0.27] 4/9 1 0.8828 (exact 2^15) 0.6776 0.41
ring damage_taken +25.35 43.46 11.22 [+1.27, +49.42] 12/15 0.03516 0.04059 (exact 2^15) 0.05708 31.44
ring hit_rate +10.61 6.32 1.63 [+7.11, +14.11] 14/15 0.0009766 0.0001831 (exact 2^15) 0.001092 4.57
ring dist -146.25 60.84 15.71 [-179.94, -112.56] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 44.01
ring_notemp damage -10.73 28.24 7.29 [-26.37, +4.91] 6/15 0.6072 0.1772 (exact 2^15) 0.3487 20.43
ring_notemp wins +0.07 0.61 0.16 [-0.27, +0.40] 4/9 1 0.8086 (exact 2^15) 0.9525 0.44
ring_notemp damage_taken -6.31 46.39 11.98 [-32.00, +19.38] 6/14 0.7905 0.632 (exact 2^15) 0.8017 33.56
ring_notemp hit_rate -2.00 4.87 1.26 [-4.69, +0.70] 7/15 1 0.1341 (exact 2^15) 0.2681 3.52
ring_notemp dist +13.34 29.63 7.65 [-3.07, +29.75] 11/15 0.1185 0.1037 (exact 2^15) 0.1055 21.43

By inferred style (explanation only, never the verdict)

arm style n mean Δdmg mean Δwins mean Δhit rate (pp)
strafe_notilt aggressive 2 +15.4 +0.50 -10.19
strafe_notilt brawler 1 -12.8 +1.00 -13.42
strafe_notilt cornercamper 1 -55.0 +0.00 -7.27
strafe_notilt dodger 5 -3.1 +0.40 -6.57
strafe_notilt pattern 3 -13.5 +0.44 -3.21
strafe_notilt rammer 1 +15.2 +0.33 -2.80
strafe_notilt spinner 1 -8.0 +0.00 -14.27
strafe_notilt wallfollower 1 -66.9 +0.00 -9.49
strafe_325 aggressive 2 +17.0 +0.50 -7.89
strafe_325 brawler 1 +6.3 +0.67 -10.88
strafe_325 cornercamper 1 -30.9 +0.33 -5.38
strafe_325 dodger 5 -5.7 +0.47 -5.91
strafe_325 pattern 3 +0.4 +0.33 -3.88
strafe_325 rammer 1 -4.1 +0.67 -3.50
strafe_325 spinner 1 -20.0 +0.00 -12.96
strafe_325 wallfollower 1 -65.2 -1.00 -7.90
ring aggressive 2 -2.8 -0.17 +12.18
ring brawler 1 +102.5 +0.67 +12.71
ring cornercamper 1 +25.7 +0.67 +11.76
ring dodger 5 +30.4 -0.33 +11.16
ring pattern 3 +45.7 +0.11 +8.30
ring rammer 1 +9.3 +0.67 -1.78
ring spinner 1 +56.7 +0.00 +24.20
ring wallfollower 1 -9.6 -1.00 +7.21
ring_notemp aggressive 2 -5.1 -0.17 -2.67
ring_notemp brawler 1 +3.1 -0.33 +2.97
ring_notemp cornercamper 1 -16.5 -0.33 +1.77
ring_notemp dodger 5 -4.0 +0.27 -2.60
ring_notemp pattern 3 -17.6 -0.33 +1.30
ring_notemp rammer 1 +15.1 +1.33 -6.08
ring_notemp spinner 1 +8.3 +0.00 +0.45
ring_notemp wallfollower 1 -88.0 +0.33 -14.57

The pre-registered verdict table, as printed by the analyzer

PRIMARY metrics are dmg/run and wins/run; hit rate is never the verdict. The pre-registered rule says an arm is BETTER when one primary metric is UP at sign-test p<0.05 while the other does not go down. That phrase has two readings and BOTH are printed:

  • strict — the other metric's mean delta is not negative at all (Δ >= 0). Nothing can be BETTER while it costs any mean damage.
  • substantive — the other metric's delta is not detectably down: the sign test is not significant and the delta is smaller than that metric's MDE (the pre-registered rule 3 says an effect under the MDE is not detectable, so it cannot count as a loss).
rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 strafe_notilt +0.38 -10.2 9/9 p=0.003906 6/15 p=0.6072 not distinguishable BETTER
2 strafe_325 +0.33 -7.2 10/12 p=0.03857 6/15 p=0.6072 not distinguishable BETTER
3 ring_notemp +0.07 -10.7 4/9 p=1 6/15 p=0.6072 not distinguishable not distinguishable
4 ring -0.04 +31.2 4/9 p=1 13/15 p=0.007385 not distinguishable BETTER

Reference tfil: 118.9 dmg/run, 1.22 wins/run, 18.17% incoming, 382 px.

Highest wins delta: strafe_notilt (+0.38 wins/run, -10.2 dmg/run) — strict: not distinguishable, substantive: BETTER.


4. Batch 2 — the range axis ON the winning engine (replication)

Design. Same frozen panel, same 3 runs × 3 rounds, new session /tmp/ab/j118_b2 (commit 8efa627, 225 battles, 0 invalid runs, 0 failed starts; no source file changed between 1984a78 and 8efa627 — only a parallel job's new docs/tools — so this is the same code). Arms (tools/ab/arms_movement_b2.txt): the winner and the strafe default from Batch 1 (replication), plus the tilt re-armed at 600 px and at 250 px, i.e. strafe_notilt has no range control and drifts to ~456 px, so these two separate "the range value is the lever" from "the tilt mechanism is the cost".

Pre-registered prediction (written before the battles): if the range value drives the win, tilt_600 should beat strafe_notilt; if the tilt mechanism itself is the cost, both tilt arms should lose to strafe_notilt. Both halves turned out wrong, and that is the useful part:

arm target / emergent range dmg/run wins/run round wins win rate incoming hit rate dmg taken/run mean distance
tfil (SHIPPED) none 114.1 1.18 53/135 39.3% 17.63% 196.5 394 px
strafe_325 325 111.8 1.76 79/135 58.5% 12.52% 144.2 434 px
strafe_notilt none (drifts) 101.5 1.64 74/135 54.8% 12.05% 148.6 459 px
tilt_600 600 102.3 1.58 71/135 52.6% 11.67% 146.4 478 px
tilt_250 250 112.0 1.56 70/135 51.9% 13.71% 162.2 415 px

Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27, +0.89], 11/12 decisive opponents, p = 0.0063; strafe_notilt +0.47 [+0.22, +0.72], 10/11, p = 0.0117; tilt_600 +0.40 [+0.04, +0.76] (sign test 8/11 p = 0.23, sign-flip p = 0.049); tilt_250 +0.38 [+0.07, +0.69], 10/12, p = 0.0386. Damage deltas are −2.1 … −12.5 (10% of the mean at worst) and never positive; incoming-hit-rate deltas are −5.2 … −6.9 pp with 0/15 opponents favouring tfil.

What this batch actually establishes

  1. The strafe engine's win over the shipped tfil replicates. Batch 1: +0.33 / +0.38 wins/run for the two strafe arms; Batch 2: +0.58 / +0.47 — the same direction, the same magnitude band, in an independent session, with 0/15 opponents going the other way on incoming hit rate in either session. Pooled descriptively, the four strafe-family arms won 52–58% of rounds in Batch 2 and 52–53% in Batch 1, against tfil's 39–41%.
  2. The baseline is reproducible across sessions: tfil won 40.7% of rounds in Batch 1 and 39.3% in Batch 2 (Δ 1.4 pp), and dealt 118.9 vs 114.1 dmg/run. The harness gives the same answer twice, which is why the win delta above is believable.
  3. The range TARGET is not the lever. Re-arming the tilt at 600 px moved the achieved distance to 478 px and at 250 px to 415 px (vs 459 px with no steering), and none of the three was separable from the others on wins. The win comes from the engine, at any of these distances; the range value within 415–478 px does not decide it. This overturns the Batch-1 reading that "the tilt costs wins" (Batch 1: no-tilt > 325; Batch 2: 325 > no-tilt, both inside noise) — the honest statement is the tilt's effect on wins is below this design's resolution (MDE ≈ 0.3–0.4 wins/run).
  4. dmg/run and wins/run remain different questions. The arm that dealt the most damage in Batch 1 (ring, +31) won nothing extra; the arms that win in Batch 2 are not the high-damage ones (strafe_325 111.8 dmg/run vs tilt_250 112.0). The win is bought with survival — 50 fewer damage taken per run, −5…−7 pp incoming hit rate — not with output.

DIRECT ANSWER after two batches (unchanged, now replicated). The best 1v1 movement measured on this panel is the strafe engine: TR_MOVEMENT=strafe. Its two Batch-1/2 configs are statistically tied with each other; if a config must be named, TR_MOVEMENT=strafe at its shipped range (325 px) has the best pooled round-win rate of the five arms in Batch 2 (58.5%) and ties strafe_notilt in Batch 1, while strafe_notilt is the simpler arm (it has no range steering to mis-tune). It is better than the shipped tfil by a margin that survives the between-opponent spread: +0.33…+0.58 wins/run, all four measurements with a 95% CI excluding 0 ([+0.04,+0.63], [+0.16,+0.60], [+0.27,+0.89], [+0.22,+0.72]), and 9/9, 10/12, 11/12 and 10/11 decisive opponents in favour, against an MDE of 0.29–0.40 — i.e. every measurement sits at or above its own detection threshold. Rejecting "no change": tfil's win share of 39–41% is not the best movement we have measured.

The analyzer's full report (verbatim)

MEASURED: session

  • commit 8efa627c05137d5a949d5a899c71fc55b5a1daf5, frozen binary sha256 005d010d8593…

  • 15 opponents × 5 arms × 3 runs × 3 rounds = 225 battles, conc=6

  • arms file arms_movement_b2.txt, panel file panel_movement.txt

  • reference arm: tfil — every delta below is (arm − tfil), opponent by opponent

  • liveness: 0 run(s) excluded (225 total)

MEASURED: per-opponent paired table (per arm)

tfil — shipped baseline, re-measured in this session (replication) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 120.8→120.8 +0.0 0.67→0.67 +0.00 +0.0 +0.00 445→445
Diamond dodger 55.9→55.9 +0.0 0.00→0.00 +0.00 +0.0 +0.00 457→457
Dookious dodger 86.5→86.5 +0.0 1.33→1.33 +0.00 +0.0 +0.00 451→451
GresSuffurd dodger 114.0→114.0 +0.0 1.33→1.33 +0.00 +0.0 +0.00 417→417
CassiusClay dodger 85.3→85.3 +0.0 0.67→0.67 +0.00 +0.0 +0.00 388→388
RetroGirl pattern 181.7→181.7 +0.0 2.00→2.00 +0.00 +0.0 +0.00 391→391
TripHammer pattern 55.8→55.8 +0.0 0.00→0.00 +0.00 +0.0 +0.00 469→469
Coriantumr pattern 100.9→100.9 +0.0 1.67→1.67 +0.00 +0.0 +0.00 444→444
WallAvoider wallfollower 150.0→150.0 +0.0 2.00→2.00 +0.00 +0.0 +0.00 317→317
HawkOnFire cornercamper 115.1→115.1 +0.0 1.67→1.67 +0.00 +0.0 +0.00 419→419
SpinBot spinner 302.0→302.0 +0.0 3.00→3.00 +0.00 +0.0 +0.00 316→316
DiamondStealer rammer 140.1→140.1 +0.0 1.00→1.00 +0.00 +0.0 +0.00 236→236
BlitzBat brawler 74.5→74.5 +0.0 2.00→2.00 +0.00 +0.0 +0.00 420→420
YersiniaPestis aggressive 65.6→65.6 +0.0 0.33→0.33 +0.00 +0.0 +0.00 377→377
Ascendant aggressive 62.8→62.8 +0.0 0.00→0.00 +0.00 +0.0 +0.00 362→362

strafe_notilt — Batch-1 winner, replication (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 120.8→92.2 -28.6 0.67→0.67 +0.00 -44.7 -4.02 445→523
Diamond dodger 55.9→63.5 +7.6 0.00→0.33 +0.33 -43.4 -7.28 457→533
Dookious dodger 86.5→88.5 +2.1 1.33→1.33 +0.00 -31.1 -3.65 451→486
GresSuffurd dodger 114.0→115.8 +1.8 1.33→2.33 +1.00 -79.2 -5.44 417→470
CassiusClay dodger 85.3→91.2 +5.9 0.67→1.33 +0.67 -43.8 -5.85 388→390
RetroGirl pattern 181.7→131.3 -50.4 2.00→2.67 +0.67 -3.1 -7.61 391→444
TripHammer pattern 55.8→65.7 +9.9 0.00→0.67 +0.67 -39.0 -4.46 469→553
Coriantumr pattern 100.9→60.5 -40.4 1.67→1.33 -0.33 -16.9 -1.32 444→572
WallAvoider wallfollower 150.0→166.5 +16.4 2.00→2.67 +0.67 -42.3 +0.54 317→327
HawkOnFire cornercamper 115.1→109.8 -5.3 1.67→2.67 +1.00 -79.0 -8.87 419→556
SpinBot spinner 302.0→259.5 -42.5 3.00→3.00 +0.00 -48.0 -23.48 316→406
DiamondStealer rammer 140.1→117.9 -22.2 1.00→1.00 +0.00 -29.9 -3.78 236→260
BlitzBat brawler 74.5→43.3 -31.2 2.00→2.33 +0.33 -94.9 -9.35 420→571
YersiniaPestis aggressive 65.6→49.7 -15.9 0.33→1.33 +1.00 -73.3 -7.89 377→414
Ascendant aggressive 62.8→67.8 +5.0 0.00→1.00 +1.00 -49.5 -10.67 362→384

strafe_325 — strafe default (range 325), replication (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 120.8→129.4 +8.6 0.67→1.67 +1.00 -23.7 -2.26 445→477
Diamond dodger 55.9→58.0 +2.1 0.00→0.00 +0.00 -47.9 -5.92 457→495
Dookious dodger 86.5→115.7 +29.2 1.33→1.67 +0.33 -53.2 -4.12 451→464
GresSuffurd dodger 114.0→112.3 -1.7 1.33→2.67 +1.33 -94.9 -8.82 417→461
CassiusClay dodger 85.3→95.2 +9.9 0.67→2.33 +1.67 -103.7 -8.41 388→366
RetroGirl pattern 181.7→162.2 -19.4 2.00→2.67 +0.67 +2.2 -4.64 391→442
TripHammer pattern 55.8→55.0 -0.8 0.00→0.67 +0.67 -50.5 -5.25 469→485
Coriantumr pattern 100.9→87.7 -13.2 1.67→1.33 -0.33 -3.8 -0.71 444→460
WallAvoider wallfollower 150.0→138.6 -11.5 2.00→3.00 +1.00 -64.0 -4.78 317→383
HawkOnFire cornercamper 115.1→122.4 +7.3 1.67→2.67 +1.00 -124.2 -11.23 419→514
SpinBot spinner 302.0→261.8 -40.2 3.00→3.00 +0.00 -32.0 -17.27 316→411
DiamondStealer rammer 140.1→149.2 +9.2 1.00→1.33 +0.33 -20.1 -2.23 236→266
BlitzBat brawler 74.5→50.1 -24.5 2.00→2.33 +0.33 -103.2 -8.93 420→519
YersiniaPestis aggressive 65.6→56.8 -8.8 0.33→0.33 +0.00 -35.7 -3.96 377→395
Ascendant aggressive 62.8→82.3 +19.6 0.00→0.67 +0.67 -30.2 -8.29 362→374

tilt_600 — tilt ON, target 600 (farther than the emergent 456) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 120.8→110.5 -10.2 0.67→1.00 +0.33 -40.5 -3.48 445→554
Diamond dodger 55.9→63.9 +8.0 0.00→0.00 +0.00 -97.2 -11.33 457→574
Dookious dodger 86.5→83.4 -3.0 1.33→1.00 -0.33 +6.7 +0.01 451→498
GresSuffurd dodger 114.0→114.8 +0.8 1.33→3.00 +1.67 -102.3 -8.16 417→475
CassiusClay dodger 85.3→62.1 -23.2 0.67→0.67 +0.00 -46.4 -5.24 388→449
RetroGirl pattern 181.7→124.3 -57.4 2.00→2.00 +0.00 +25.0 -5.38 391→467
TripHammer pattern 55.8→52.6 -3.2 0.00→1.33 +1.33 -54.1 -7.01 469→552
Coriantumr pattern 100.9→63.1 -37.8 1.67→1.33 -0.33 +4.3 -1.08 444→557
WallAvoider wallfollower 150.0→147.1 -2.9 2.00→1.67 -0.33 -42.1 -2.06 317→373
HawkOnFire cornercamper 115.1→95.0 -20.1 1.67→2.00 +0.33 -98.9 -9.05 419→572
SpinBot spinner 302.0→250.5 -51.5 3.00→3.00 +0.00 -32.0 -18.70 316→444
DiamondStealer rammer 140.1→159.9 +19.9 1.00→1.67 +0.67 -53.3 -1.92 236→274
BlitzBat brawler 74.5→41.0 -33.5 2.00→3.00 +1.00 -91.0 -11.09 420→585
YersiniaPestis aggressive 65.6→75.4 +9.8 0.33→0.67 +0.33 -45.0 -7.18 377→408
Ascendant aggressive 62.8→90.6 +27.8 0.00→1.33 +1.33 -85.0 -12.15 362→387

tilt_250 — tilt ON, target 250 (much nearer than the emergent 456) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 120.8→108.7 -12.1 0.67→1.00 +0.33 -15.7 -2.25 445→470
Diamond dodger 55.9→97.2 +41.3 0.00→1.00 +1.00 -100.7 -7.74 457→482
Dookious dodger 86.5→105.0 +18.6 1.33→2.00 +0.67 -39.4 -2.97 451→460
GresSuffurd dodger 114.0→145.2 +31.2 1.33→2.67 +1.33 -89.4 -6.57 417→409
CassiusClay dodger 85.3→103.1 +17.8 0.67→1.00 +0.33 -30.2 -3.63 388→367
RetroGirl pattern 181.7→142.8 -38.8 2.00→2.33 +0.33 +36.3 -4.18 391→436
TripHammer pattern 55.8→53.4 -2.3 0.00→0.33 +0.33 -20.3 -2.44 469→467
Coriantumr pattern 100.9→75.1 -25.8 1.67→1.33 -0.33 +5.1 -1.04 444→436
WallAvoider wallfollower 150.0→157.0 +7.0 2.00→1.33 -0.67 +42.9 +1.38 317→297
HawkOnFire cornercamper 115.1→130.6 +15.5 1.67→3.00 +1.33 -123.6 -10.75 419→454
SpinBot spinner 302.0→258.0 -44.0 3.00→3.00 +0.00 -32.0 -18.91 316→432
DiamondStealer rammer 140.1→143.7 +3.6 1.00→1.00 +0.00 -12.7 -1.10 236→260
BlitzBat brawler 74.5→51.2 -23.3 2.00→2.67 +0.67 -94.4 -7.50 420→502
YersiniaPestis aggressive 65.6→57.3 -8.3 0.33→0.67 +0.33 -23.5 -5.03 377→396
Ascendant aggressive 62.8→51.1 -11.7 0.00→0.00 +0.00 -17.3 -4.93 362→364

MEASURED: pooled dashboard (all valid runs, NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
tfil 45 114.1 196.5 1.18 53/135 39.3% 17.63% 394
strafe_notilt 45 101.5 148.6 1.64 74/135 54.8% 12.05% 459
strafe_325 45 111.8 144.2 1.76 79/135 58.5% 12.52% 434
tilt_600 45 102.3 146.4 1.58 71/135 52.6% 11.67% 478
tilt_250 45 112.0 162.2 1.56 70/135 51.9% 13.71% 415

MEASURED: cross-opponent aggregation (the verdict layer)

Deltas are per-opponent (arm − reference). spread is the SD of those deltas ACROSS opponents; SE = spread/√n; 95% CI = mean ± t·SE. Sign test = how many opponents the arm wins (ties dropped), exact binomial; sign-flip = permutation test on the mean of the deltas.

arm metric mean Δ spread (SD) SE 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
strafe_notilt damage -12.52 21.86 5.64 [-24.62, -0.41] 7/15 1 0.04456 (exact 2^15) 0.1323 15.81
strafe_notilt wins +0.47 0.45 0.12 [+0.22, +0.72] 10/11 0.01172 0.003906 (exact 2^15) 0.007526 0.33
strafe_notilt damage_taken -47.88 24.70 6.38 [-61.55, -34.20] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 17.86
strafe_notilt hit_rate -6.87 5.51 1.42 [-9.93, -3.82] 1/15 0.0009766 0.0001221 (exact 2^15) 0.0008919 3.99
strafe_notilt dist +65.39 46.33 11.96 [+39.73, +91.05] 15/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 33.52
strafe_325 damage -2.28 17.82 4.60 [-12.15, +7.59] 7/15 1 0.6319 (exact 2^15) 0.712 12.89
strafe_325 wins +0.58 0.56 0.14 [+0.27, +0.89] 11/12 0.006348 0.002441 (exact 2^15) 0.00525 0.40
strafe_325 damage_taken -52.32 38.49 9.94 [-73.64, -31.01] 1/15 0.0009766 0.0001221 (exact 2^15) 0.0008919 27.84
strafe_325 hit_rate -6.45 4.20 1.08 [-8.78, -4.13] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 3.04
strafe_325 dist +40.15 35.18 9.08 [+20.67, +59.64] 14/15 0.0009766 0.0004272 (exact 2^15) 0.002377 25.45
tilt_600 damage -11.78 25.11 6.48 [-25.68, +2.13] 5/15 0.3018 0.09137 (exact 2^15) 0.1055 18.16
tilt_600 wins +0.40 0.66 0.17 [+0.04, +0.76] 8/11 0.2266 0.04883 (exact 2^15) 0.04491 0.48
tilt_600 damage_taken -50.12 40.17 10.37 [-72.36, -27.87] 3/15 0.03516 0.0005493 (exact 2^15) 0.002377 29.06
tilt_600 hit_rate -6.92 5.05 1.30 [-9.72, -4.12] 1/15 0.0009766 0.0001221 (exact 2^15) 0.0008919 3.65
tilt_600 dist +83.94 44.45 11.48 [+59.32, +108.55] 15/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 32.15
tilt_250 damage -2.09 24.78 6.40 [-15.82, +11.63] 7/15 1 0.7453 (exact 2^15) 0.7983 17.92
tilt_250 wins +0.38 0.56 0.14 [+0.07, +0.69] 10/12 0.03857 0.03125 (exact 2^15) 0.05415 0.41
tilt_250 damage_taken -34.32 48.54 12.53 [-61.21, -7.44] 3/15 0.03516 0.01593 (exact 2^15) 0.02877 35.11
tilt_250 hit_rate -5.18 4.89 1.26 [-7.89, -2.47] 1/15 0.0009766 0.0002441 (exact 2^15) 0.001332 3.54
tilt_250 dist +21.42 37.60 9.71 [+0.60, +42.24] 10/15 0.3018 0.03253 (exact 2^15) 0.04377 27.20

By inferred style (explanation only, never the verdict)

arm style n mean Δdmg mean Δwins mean Δhit rate (pp)
strafe_notilt aggressive 2 -5.5 +1.00 -9.28
strafe_notilt brawler 1 -31.2 +0.33 -9.35
strafe_notilt cornercamper 1 -5.3 +1.00 -8.87
strafe_notilt dodger 5 -2.2 +0.40 -5.25
strafe_notilt pattern 3 -27.0 +0.33 -4.46
strafe_notilt rammer 1 -22.2 +0.00 -3.78
strafe_notilt spinner 1 -42.5 +0.00 -23.48
strafe_notilt wallfollower 1 +16.4 +0.67 +0.54
strafe_325 aggressive 2 +5.4 +0.33 -6.12
strafe_325 brawler 1 -24.5 +0.33 -8.93
strafe_325 cornercamper 1 +7.3 +1.00 -11.23
strafe_325 dodger 5 +9.6 +0.87 -5.91
strafe_325 pattern 3 -11.1 +0.33 -3.53
strafe_325 rammer 1 +9.2 +0.33 -2.23
strafe_325 spinner 1 -40.2 +0.00 -17.27
strafe_325 wallfollower 1 -11.5 +1.00 -4.78
tilt_600 aggressive 2 +18.8 +0.83 -9.67
tilt_600 brawler 1 -33.5 +1.00 -11.09
tilt_600 cornercamper 1 -20.1 +0.33 -9.05
tilt_600 dodger 5 -5.5 +0.33 -5.64
tilt_600 pattern 3 -32.8 +0.33 -4.49
tilt_600 rammer 1 +19.9 +0.67 -1.92
tilt_600 spinner 1 -51.5 +0.00 -18.70
tilt_600 wallfollower 1 -2.9 -0.33 -2.06
tilt_250 aggressive 2 -10.0 +0.17 -4.98
tilt_250 brawler 1 -23.3 +0.67 -7.50
tilt_250 cornercamper 1 +15.5 +1.33 -10.75
tilt_250 dodger 5 +19.4 +0.73 -4.63
tilt_250 pattern 3 -22.3 +0.11 -2.55
tilt_250 rammer 1 +3.6 +0.00 -1.10
tilt_250 spinner 1 -44.0 +0.00 -18.91
tilt_250 wallfollower 1 +7.0 -0.67 +1.38

The pre-registered verdict table, as printed by the analyzer

PRIMARY metrics are dmg/run and wins/run; hit rate is never the verdict. The pre-registered rule says an arm is BETTER when one primary metric is UP at sign-test p<0.05 while the other does not go down. That phrase has two readings and BOTH are printed:

  • strict — the other metric's mean delta is not negative at all (Δ >= 0). Nothing can be BETTER while it costs any mean damage.
  • substantive — the other metric's delta is not detectably down: the sign test is not significant and the delta is smaller than that metric's MDE (the pre-registered rule 3 says an effect under the MDE is not detectable, so it cannot count as a loss).
rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 strafe_325 +0.58 -2.3 11/12 p=0.006348 7/15 p=1 not distinguishable BETTER
2 strafe_notilt +0.47 -12.5 10/11 p=0.01172 7/15 p=1 not distinguishable BETTER
3 tilt_600 +0.40 -11.8 8/11 p=0.2266 5/15 p=0.3018 not distinguishable not distinguishable
4 tilt_250 +0.38 -2.1 10/12 p=0.03857 7/15 p=1 not distinguishable BETTER

Reference tfil: 114.1 dmg/run, 1.18 wins/run, 17.63% incoming, 394 px.

Highest wins delta: strafe_325 (+0.58 wins/run, -2.3 dmg/run) — strict: not distinguishable, substantive: BETTER.


5. What to try next (rewritten AFTER Batches 1–2 — these are recommendations, not results)

Post-hoc structure of the win (60 opponent×arm×session points from Batches 1–2, MEASURED). The strafe win is not uniformly distributed and its size is not predicted by the size of the hit-rate improvement across opponents:

  • 56 of 60 points have a non-negative win delta; the 4 negatives are −0.33 (strafe_notilt vs Coriantumr B2, strafe_325 vs Coriantumr B2), −0.33 (strafe_325 vs DrussGT B1) and one WallAvoider B1 point (−1.00) that reverses to +1.00 in Batch 2 — so no opponent family shows a reproducible regression at this n.
  • corr(Δwins, Δincoming-hit-rate) = −0.09 across those points; corr(Δwins, Δdamage/run) = +0.38. Buckets: points whose hit rate improved by ≥5 pp average +0.53 wins/run (n=36); the 4 points with <2 pp of hit-rate improvement average −0.08.
  • Reading: the aggregate win is a survival effect (fewer hits taken, ~50 less damage taken per run), but "this arm dodges better by X pp here" does not mean "it wins more rounds here". Do not use hit-rate improvement as a proxy for a win at the level of a single opponent — that is the sixth-verdict trap this project keeps paying for.

Ranked by value per battle, given what the two batches measured:

  1. The engine is the lever; the range knob is not. Both batches put the strafe arms 12–19 pp above tfil on round-win rate while three different range targets (none/250/600, achieved 415–478 px) made no separable difference. So the next batch should attack the strafe picker itself, not the range: TR_STRAFE_DWELL_MIN/MAX (reversal frequency), TR_STRAFE_BAND + TR_STRAFE_SPREAD (how far the picker hedges), TR_STRAFE_REACH (line length), TR_STRAFE_WALL_BIAS, TR_STRAFE_WALL_MARGIN. 3–4 arms, same panel, one knob family per batch, and look for a plateau, not a peak.
  2. The verdict metric for movement is round wins; the mechanism metric is incoming hit rate. The winner took ~1/3 fewer hits at the same damage output, and round wins in this harness are survival wins. So screen mechanism ideas on incoming hit rate (±1 pp is detectable here: MDE 1.3–4.0 pp) and only then spend a full panel batch confirming the win effect.
  3. Do not chase damage. The one arm that gained damage (ring, +31/run, p = 0.007) won fewer nominal rounds and took +25 damage/run. A movement arm that raises damage but lowers survival is a loss in disguise — the mirror of the six inverted hit-rate verdicts this project has already paid for.
  4. strafe_notilt is the recommendation to ship-test, if a shipping decision is ever taken: it has the same win effect as the range-steered config without an extra tuning surface. Shipping is a separate decision — this campaign does not touch a shipped default.
  5. Then the gun (the owner's next stage, per the mandate): same harness, same panel or a gun-specific one, same paired-with-sign-test statistics. Two facts for the gun job: (a) round wins here are survival wins, so the gun's job is to kill, not merely to out-damage; (b) the panel is 15 opponents wide and its strong dodgers (Diamond 39.5, CassiusClay 73.4, TripHammer 59.5 dmg/run for tfil) are exactly the ones a DrussGT-only gun claim will fail against.
  6. Melee is a different game (j116's finding): it needs its own panel and its own ledger section; the 1v1 panel's verdicts do not transfer.

6. What would make us stop

  • The movement stage has already produced its first winner (TR_MOVEMENT=strafe), and by rule 2 with the substantive reading it beats the shipped default with a margin that survives the between-opponent spread, replicated in two independent sessions. A later job may therefore either (a) keep hunting within the strafe picker (item 1 above) and stop as soon as two consecutive batches fail to improve on it beyond the MDE, or (b) declare it the movement answer and move to the gun. Both are successful outcomes.
  • Stop the movement stage entirely once a batch's best arm cannot beat strafe beyond the MDE, or when a movement arm's win gain is bought with a detectable damage or survival loss. At that point "this is the measured optimum of this design space" is the conclusion, not a failure.
  • Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.

7. How to run a batch (exact commands)

# 1. wait for the arena (this job may not be the only one fighting)
tools/ab/tournament_run.sh \
    --arms    tools/ab/arms_movement_b1.txt \
    --panel   tools/ab/panel_movement.txt \
    --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
    --reference tfil \
    --outdir /tmp/ab/j118_b1

# Batch 2 (the range axis on the winning engine) was the same command with
#   --arms tools/ab/arms_movement_b2.txt --outdir /tmp/ab/j118_b2

# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil

--reference may be ANY arm of the session: re-analyzing /tmp/ab/j118_b1 --reference strafe_325 is a free pairwise comparison with no battles (it is how the "the two strafe configs are not separable" claim was checked: Δwins +0.04, p = 0.75, MDE 0.33).

8. Session log (outdirs are in /tmp and are NOT committed)

session commit battles arms verdict
/tmp/ab/j118_b1 1984a78 225 (0 invalid) tfil, strafe_notilt, strafe_325, ring, ring_notemp strafe_notilt beats tfil on wins (+0.38, 9/9, p=0.0039)
/tmp/ab/j118_b2 8efa627 225 (0 invalid) tfil, strafe_notilt, strafe_325, tilt_600, tilt_250 all four strafe arms beat tfil on wins (+0.38…+0.58); the range target decides nothing

Both sessions can be re-analyzed offline at any time (no arena needed) as long as /tmp/ab/j118_b* still exists; after a reboot only this ledger's tables remain, which is why every number is inlined above.

The runner writes <outdir>/session.json (commit sha, binary sha256, arms, panel) so any later job can re-analyze an old session offline, with no arena.


Batch 3 — the reversal/dwell timing of the strafe picker

Pre-registration (written and committed BEFORE the battles). Commit 7311aae (Task A, the heat field made env-overridable) is the frozen binary. Session /tmp/ab/j119_b3. Arms file tools/ab/arms_movement_b3.txt, panel tools/ab/panel_movement.txt, 6 arms × 15 opponents × 3 runs × 3 rounds = 270 battles, conc 6, --reference strafe.

Why this batch. Batches 1–2 established that the strafe ENGINE wins by survival (+0.33…+0.58 wins/run over the shipped tfil, incoming hit rate −5…−7 pp) and that the RANGE knob is not the lever. The untouched axis is the picker itself. The strafe design flips the SIGN of setForward (a free reversal) and holds a sign for rand(DWELL_MIN..DWELL_MAX) ticks, so the dwell IS the reversal period — the whole premise of the mover is "when to flip".

Reference in this batch is strafe (current defaults), not tfil. Every delta below is (arm − strafe); tfil is carried only as the shipped control.

# arm env what it isolates
1 strafe TR_MOVEMENT=strafe reference: dwell 6-20, spread 1, reach 144
2 tfil (none — shipped) shipped control / cross-batch calibration
3 fast_flip TR_MOVEMENT=strafe TR_STRAFE_DWELL_MIN=2 TR_STRAFE_DWELL_MAX=8 reversal every ~5 ticks
4 slow_flip TR_MOVEMENT=strafe TR_STRAFE_DWELL_MIN=12 TR_STRAFE_DWELL_MAX=40 reversal every ~26 ticks
5 wide_spread TR_MOVEMENT=strafe TR_STRAFE_SPREAD=2 TR_STRAFE_REACH=216 wider hedge (±2 tiles, 216 px)
6 narrow TR_MOVEMENT=strafe TR_STRAFE_SPREAD=0 TR_STRAFE_REACH=108 no hedge, short 108 px reach

Pre-registered prediction (before the battles): reversal timing is a real mechanism lever; the picker hedge geometry is not. Specifically: (a) fast_flip will LOWER incoming hit rate vs strafe (each heading is exposed for less time) and (b) slow_flip will RAISE it (a pattern gun gets a longer straight run); (c) NEITHER extreme is expected to beat strafe on round wins by rule 2 (a sign-test win with no detectable damage loss), because the win effect is bounded by survival that is already high; (d) wide_spread and narrow should not separate from strafe (the Batch-2 lesson that picker-shape knobs sit below the MDE). If an arm DOES beat strafe, the most likely is fast_flip, via survival. I record this as a falsifiable claim; a wrong prediction is recorded as wrong.

Outcome — Batch 3

Direct answer: NOTHING beats the current strafe on round wins. The session ran 270 battles (0 failed, 0 never started) and excluded 1 run on liveness grounds (Ascendant/strafe run1: owner attribution ambiguous), so strafe has 44 valid runs and every other arm 45. The strafe-over-tfil effect replicates a THIRD time: in this session tfil wins 42.2% of its rounds vs strafe's 53.0%.

Pooled dashboard (valid runs, explanation only — NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
strafe (REF) 44 112.3 154.7 1.59 70/132 53.0% 13.14% 434
tfil 45 113.9 197.9 1.27 57/135 42.2% 18.10% 383
fast_flip 45 105.6 176.6 1.44 65/135 48.1% 15.19% 421
slow_flip 45 101.8 147.0 1.38 62/135 45.9% 12.80% 428
wide_spread 45 111.4 160.0 1.67 75/135 55.6% 13.11% 425
narrow 45 104.6 175.1 1.40 63/135 46.7% 14.57% 429

Per-opponent Δwins/run (arm − strafe)

opponent style tfil fast_flip slow_flip wide_spread narrow
DrussGT dodger +1.33 +1.00 +0.00 +1.00 +0.67
Diamond dodger +0.33 +0.33 +0.00 +0.00 +0.00
Dookious dodger -1.00 +0.33 +0.00 +1.33 +0.33
GresSuffurd dodger -1.33 -0.67 -0.67 -0.33 -1.33
CassiusClay dodger -0.33 -0.67 +0.67 -0.33 -0.33
RetroGirl pattern -1.33 -1.00 -0.67 +0.00 -0.67
TripHammer pattern -1.00 -1.00 -1.00 -0.33 -1.00
Coriantumr pattern -1.00 -1.33 -1.33 -2.00 -1.67
WallAvoider wallfollower +0.67 +0.00 +0.33 +0.33 +0.00
HawkOnFire cornercamper -0.67 +0.00 +0.00 +0.33 +0.00
SpinBot spinner +0.00 +0.00 +0.00 +0.00 +0.00
DiamondStealer rammer +0.33 +0.67 -1.00 +1.00 +1.00
BlitzBat brawler -0.33 +0.33 +0.33 +0.00 +0.00
YersiniaPestis aggressive +0.00 +0.33 +0.00 +0.33 +0.67
Ascendant aggressive +0.00 +0.00 +0.67 +0.33 +0.00

Cross-opponent aggregation (the verdict layer, verbatim)

arm metric mean Δ spread (SD) SE 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
tfil damage +2.66 28.32 7.31 [-13.02, +18.35] 6/15 0.6072 0.7092 0.8871 20.49
tfil wins -0.29 0.78 0.20 [-0.72, +0.14] 4/12 0.3877 0.2056 0.1952 0.56
tfil damage_taken +41.18 36.79 9.50 [+20.81, +61.56] 14/15 0.0009766 0.001221 0.003445 26.61
tfil hit_rate +6.00 4.74 1.22 [+3.37, +8.62] 14/15 0.0009766 0.0004883 0.001966 3.43
tfil dist -49.26 47.09 12.16 [-75.33, -23.18] 1/15 0.0009766 0.00116 0.003445 34.06
fast_flip damage -5.61 25.59 6.61 [-19.79, +8.56] 5/15 0.3018 0.4282 0.5137 18.51
fast_flip wins -0.11 0.67 0.17 [-0.48, +0.26] 6/11 1 0.6152 0.5627 0.49
fast_flip damage_taken +19.90 26.12 6.74 [+5.43, +34.36] 11/15 0.1185 0.01245 0.02143 18.89
fast_flip hit_rate +2.12 2.18 0.56 [+0.91, +3.33] 12/15 0.03516 0.002563 0.004932 1.57
fast_flip dist -11.02 32.34 8.35 [-28.93, +6.89] 5/15 0.3018 0.2111 0.222 23.40
slow_flip damage -9.41 20.87 5.39 [-20.97, +2.15] 5/15 0.3018 0.09509 0.09384 15.10
slow_flip wins -0.18 0.62 0.16 [-0.52, +0.16] 4/9 1 0.3477 0.342 0.45
slow_flip damage_taken -9.69 37.56 9.70 [-30.49, +11.11] 6/15 0.6072 0.3287 0.3203 27.17
slow_flip hit_rate -1.64 3.61 0.93 [-3.64, +0.36] 6/15 0.6072 0.1024 0.1055 2.61
slow_flip dist -4.73 20.40 5.27 [-16.03, +6.57] 6/15 0.6072 0.384 0.4432 14.76
wide_spread damage +0.19 19.58 5.06 [-10.66, +11.03] 8/15 1 0.9717 0.7548 14.17
wide_spread wins +0.11 0.77 0.20 [-0.32, +0.54] 7/11 0.5488 0.6738 0.3273 0.56
wide_spread damage_taken +3.30 28.79 7.43 [-12.65, +19.24] 10/15 0.3018 0.6722 0.5895 20.82
wide_spread hit_rate -0.02 2.52 0.65 [-1.42, +1.38] 6/15 0.6072 0.9786 0.6701 1.82
wide_spread dist -7.33 22.89 5.91 [-20.01, +5.35] 5/15 0.3018 0.2528 0.1055 16.56
narrow damage -6.59 22.84 5.90 [-19.24, +6.06] 5/15 0.3018 0.2835 0.3203 16.52
narrow wins -0.16 0.74 0.19 [-0.57, +0.26] 4/9 1 0.5039 0.5139 0.54
narrow damage_taken +18.44 30.99 8.00 [+1.28, +35.60] 9/15 0.6072 0.03699 0.05708 22.41
narrow hit_rate +1.09 2.19 0.56 [-0.12, +2.30] 10/15 0.3018 0.07574 0.1055 1.58
narrow dist -3.54 20.03 5.17 [-14.64, +7.55] 6/15 0.6072 0.4975 0.4777 14.49

The pre-registered verdict (verbatim)

rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 wide_spread +0.11 +0.2 7/11 p=0.5488 8/15 p=1 not distinguishable not distinguishable
2 fast_flip -0.11 -5.6 6/11 p=1 5/15 p=0.3018 not distinguishable not distinguishable
3 narrow -0.16 -6.6 4/9 p=1 5/15 p=0.3018 not distinguishable not distinguishable
4 slow_flip -0.18 -9.4 4/9 p=1 5/15 p=0.3018 not distinguishable not distinguishable
5 tfil -0.29 +2.7 4/12 p=0.3877 6/15 p=0.6072 not distinguishable not distinguishable

Reference strafe: 112.3 dmg/run, 1.59 wins/run, 13.14% incoming, 434 px. Highest wins delta: wide_spread (+0.11 wins/run, +0.2 dmg/run) — strict: not distinguishable, substantive: not distinguishable.

Reading

  • wide_spread (SPREAD=2, REACH=216) is the ONLY arm with a positive point estimate on wins (+0.11/run) and it is damage-neutral (+0.2). It is not distinguishable: positive on 7 of 11 decisive opponents, p = 0.55, MDE 0.56 — the observed effect is ~5× smaller than the design's detection threshold.
  • fast_flip is the one arm with a detectable survival cost: incoming hit rate +2.12 pp (12/15, p = 0.035), +19.9 damage taken/run (sign-flip p = 0.012), and it wins −0.11/run. Faster reversals do NOT dodge better here.
  • slow_flip dodges marginally better (−1.64 pp, NS) and wins −0.18/run; the two dwell extremes do not bracket a win at all.
  • The pre-registered prediction was partly WRONG and is recorded as wrong: (a) fast_flip was predicted to LOWER the hit rate — it RAISED it (+2.12 pp); (b) slow_flip was predicted to RAISE it — it lowered it (−1.64 pp, NS). Predictions (c) "neither extreme beats strafe on wins" and (d) "spread/reach do not separate" were correct.
  • Net: the reversal/dwell axis is a REAL mechanism knob — fast_flip demonstrably hurts dodging (MDE 1.57 pp, observed 2.12 pp) — but it does not convert into a round-win improvement over the current dwell, and the picker hedge geometry does not separate.

Batch 4 — the heat field strength (how strongly strafe treats danger)

Pre-registration (written and committed BEFORE the battles). Same frozen binary (7311aae), session /tmp/ab/j119_b4, arms file tools/ab/arms_movement_b4.txt, 6 arms × 15 opponents × 3 runs × 3 rounds = 270 battles, conc 6, --reference strafe.

Why this batch. The strafe win is a survival effect, and strafe runs a deliberate RETUNE of the shipped heat field: bullet core/aura 20/10 (the core is ABOVE the 10-px path threshold, so the bullet itself is the danger), corridor 10 (== threshold), wall 15/5 (outer ring only), pillar off — vs the shipped field's corridor 20 and wall 30/10. The question is whether the retune (or the strength of any one source) is what buys the survival. One arm per knob family.

# arm env what it isolates
1 strafe TR_MOVEMENT=strafe reference: bullet 20/10, corridor 10, wall 15/5
2 tfil (none — shipped) shipped control
3 bullet_strong TR_MOVEMENT=strafe TR_STRAFE_BULLET_CORE=30 TR_STRAFE_BULLET_AURA=15 the bullet retune
4 field_strong TR_MOVEMENT=strafe TR_STRAFE_CORRIDOR_HEAT=20 TR_STRAFE_WALL_HOTNESS=30 TR_STRAFE_WALL_RADIANCE=10 the shipped corridor/wall shape
5 field_off TR_MOVEMENT=strafe TR_STRAFE_CORRIDOR_HEAT=0 TR_STRAFE_WALL_HOTNESS=0 no corridors, no wall heat
6 wall_tight TR_MOVEMENT=strafe TR_STRAFE_WALL_MARGIN=54 TR_STRAFE_WALL_BIAS=0.7 TR_STRAFE_KAPPA=0.005 TR_STRAFE_WING_MAX=45 the curved-wing geometry family

Note on field_off: wall hotness is set to 0, NOT the radiance — a radiance of 0 paints a FLAT WallHotness field over the whole arena (the falloff multiplies the tile index), which is the opposite of "no walls".

Pre-registered prediction (before the battles): the strafe retune is load-bearing at the corridor/wall end. Specifically: (a) field_strong (the shipped saturated corridor/wall shape) will RAISE incoming hit rate and LOSE round wins vs strafe; (b) field_off will be a wash or slightly worse — corridors and walls are real threats the picker should see; (c) bullet_strong will be a wash or slightly worse (a 30 core is above the 25 danger-replan threshold, so it over-replans); (d) wall_tight will not separate. NET: no arm is expected to BEAT strafe on round wins, and the current retune should rank at or near the top. A wrong prediction is recorded as wrong.

Task A (this job's separate deliverable). The shipped tfil mover's heat shape (CorridorHeat/WallHotness/WallRadiance) was a Nim const and could not be swept by env; commit 7311aae makes them env-overridable vars (TR_TFIL_CORRIDOR_HEAT/TR_TFIL_WALL_HOTNESS/TR_TFIL_WALL_RADIANCE, shipped defaults 20/30/10) and the default path is proven byte-identical by common_libs/tests/test_tfil_commit_env.nim (30 checks). STRAFE's own heat knobs were already env-overridable, which is what this batch sweeps.

Outcome — Batch 4

Direct answer: NOTHING beats the current strafe on round wins — and the batch says something stronger: two arms are DETECTABLY WORSE. 270 battles (0 failed, 0 never started, 0 excluded). The strafe-over-tfil effect replicates a fourth time: tfil wins 38.5% of its rounds vs strafe's 52.6% (Δwins −0.42 [-0.66, −0.19], 1/11 decisive, p = 0.0117).

Pooled dashboard (valid runs, explanation only — NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
strafe (REF) 45 103.5 157.1 1.58 71/135 52.6% 13.16% 435
tfil 45 110.7 194.6 1.16 52/135 38.5% 17.03% 394
bullet_strong 45 107.3 160.9 1.42 64/135 47.4% 12.63% 430
field_strong 45 110.5 171.3 1.29 58/135 43.0% 13.82% 432
field_off 45 92.7 142.3 1.11 50/135 37.0% 12.65% 462
wall_tight 45 103.4 149.5 1.38 62/135 45.9% 12.81% 444

Per-opponent Δwins/run (arm − strafe)

opponent style tfil bullet_strong field_strong field_off wall_tight
DrussGT dodger +0.00 -0.33 -1.33 -1.33 -1.33
Diamond dodger -0.67 -0.33 -0.67 -0.33 -0.33
Dookious dodger +0.00 +1.00 -0.33 +0.00 -0.67
GresSuffurd dodger +0.33 -0.67 -0.33 -1.33 -0.67
CassiusClay dodger -1.00 -0.33 -0.67 -0.33 +0.00
RetroGirl pattern -1.00 -1.67 -0.33 -2.00 +0.00
TripHammer pattern -0.67 +0.67 +0.67 -0.67 +0.67
Coriantumr pattern -0.33 -0.33 +0.33 -0.33 +1.33
WallAvoider wallfollower +0.00 -0.33 -0.33 +0.00 -0.33
HawkOnFire cornercamper -0.67 +0.67 -0.67 -0.33 -0.67
SpinBot spinner +0.00 +0.00 +0.00 +0.00 +0.00
DiamondStealer rammer -0.33 -0.67 -0.33 -0.33 -1.00
BlitzBat brawler -1.00 +0.00 +0.00 -0.33 +0.33
YersiniaPestis aggressive -0.33 +0.67 +0.00 +1.00 +0.33
Ascendant aggressive -0.67 -0.67 -0.33 -0.67 -0.67

Cross-opponent aggregation (the verdict layer, verbatim)

arm metric mean Δ spread (SD) SE 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
tfil damage +7.28 13.50 3.49 [-0.19, +14.76] 10/15 0.3018 0.05011 0.03817 9.77
tfil wins -0.42 0.43 0.11 [-0.66, -0.19] 1/11 0.01172 0.004883 0.01108 0.31
tfil damage_taken +37.57 26.58 6.86 [+22.84, +52.29] 14/15 0.0009766 0.0001221 0.0008919 19.23
tfil hit_rate +5.82 5.18 1.34 [+2.96, +8.69] 15/15 6.104e-05 6.104e-05 0.0007265 3.74
tfil dist -41.25 33.73 8.71 [-59.93, -22.57] 2/15 0.007385 0.0004272 0.001966 24.40
bullet_strong damage +3.85 11.82 3.05 [-2.69, +10.40] 11/15 0.1185 0.2264 0.222 8.55
bullet_strong wins -0.16 0.69 0.18 [-0.54, +0.23] 4/13 0.2668 0.4736 0.5518 0.50
bullet_strong damage_taken +3.84 37.00 9.55 [-16.66, +24.33] 7/15 1 0.6882 0.7548 26.76
bullet_strong hit_rate -0.02 2.67 0.69 [-1.50, +1.46] 7/15 1 0.9787 0.8871 1.93
bullet_strong dist -4.84 23.47 6.06 [-17.84, +8.16] 7/15 1 0.4423 0.6293 16.98
field_strong damage +7.09 17.16 4.43 [-2.41, +16.59] 10/15 0.3018 0.1321 0.1475 12.41
field_strong wins -0.29 0.47 0.12 [-0.55, -0.03] 2/12 0.03857 0.04688 0.0403 0.34
field_strong damage_taken +14.27 26.24 6.78 [-0.27, +28.80] 10/15 0.3018 0.05359 0.05708 18.98
field_strong hit_rate +1.72 2.74 0.71 [+0.20, +3.23] 11/15 0.1185 0.02704 0.02487 1.98
field_strong dist -3.02 21.65 5.59 [-15.01, +8.97] 8/15 1 0.5974 0.8871 15.66
field_off damage -10.76 14.02 3.62 [-18.52, -2.99] 3/15 0.03516 0.006714 0.01149 10.14
field_off wins -0.47 0.70 0.18 [-0.85, -0.08] 1/12 0.006348 0.02783 0.02037 0.51
field_off damage_taken -14.73 34.01 8.78 [-33.57, +4.11] 5/15 0.3018 0.1121 0.1055 24.60
field_off hit_rate -0.52 2.65 0.68 [-1.99, +0.95] 6/15 0.6072 0.4832 0.5509 1.92
field_off dist +27.24 21.37 5.52 [+15.40, +39.07] 14/15 0.0009766 0.0001831 0.001092 15.46
wall_tight damage -0.09 26.39 6.81 [-14.71, +14.52] 8/15 1 0.9894 0.9773 19.09
wall_tight wins -0.20 0.69 0.18 [-0.58, +0.18] 4/12 0.3877 0.3345 0.208 0.50
wall_tight damage_taken -7.56 31.64 8.17 [-25.08, +9.96] 8/15 1 0.3915 0.3787 22.89
wall_tight hit_rate +0.25 3.07 0.79 [-1.45, +1.94] 6/15 0.6072 0.7711 0.9321 2.22
wall_tight dist +8.70 23.18 5.98 [-4.14, +21.53] 10/15 0.3018 0.1666 0.182 16.76

The pre-registered verdict (verbatim)

rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 bullet_strong -0.16 +3.9 4/13 p=0.2668 11/15 p=0.1185 not distinguishable not distinguishable
2 wall_tight -0.20 -0.1 4/12 p=0.3877 8/15 p=1 not distinguishable not distinguishable
3 field_strong -0.29 +7.1 2/12 p=0.03857 10/15 p=0.3018 not distinguishable WORSE
4 tfil -0.42 +7.3 1/11 p=0.01172 10/15 p=0.3018 not distinguishable WORSE
5 field_off -0.47 -10.8 1/12 p=0.006348 3/15 p=0.03516 WORSE not distinguishable

Reference strafe: 103.5 dmg/run, 1.58 wins/run, 13.16% incoming, 435 px. Highest wins delta: bullet_strong (−0.16 wins/run, +3.9 dmg/run) — strict: not distinguishable, substantive: not distinguishable.

Reading

  • The current strafe retune is load-bearing, in both directions. Weakening the corridor/wall treatment is not free and strengthening it back to the shipped shape is not free either:
    • field_strong (corridor 20, wall 30/10 = the SHIPPED saturated shape) is WORSE on wins: Δ −0.29 [-0.55, −0.03], positive on only 2/12 decisive opponents, p = 0.039; incoming hit rate +1.72 pp.
    • field_off (no corridors, no wall heat) is WORSE on wins, Δ −0.47 [-0.85, −0.08], p = 0.0063, and loses damage (Δ −10.8, p = 0.035): killing the wall logic costs ~10 dmg/run for nothing.
  • bullet_strong (core 30 > the 25 danger-replan threshold) and wall_tight (tighter/faster wings) are indistinguishable from strafe, and both nominally negative on wins.
  • The pre-registered prediction was largely CORRECT, one part wrong: (a) field_strong worse — correct (detectably, Δwins p = 0.039); (b) field_off "wash or slightly worse" — correct in direction but WRONG in size: it is detectably worse, not a wash; (c) bullet_strong wash-or-worse — correct; (d) wall_tight no separation — correct.
  • Mechanism note (the campaign's standing lesson, again): field_off has the BEST incoming hit rate of the batch (12.65% vs strafe's 13.16%) yet the WORST round-win rate (37.0%). Dodging better is not winning more — without the corridor/wall gradient the picker drifts to a mean 462 px and trades damage (−10.8) for avoidance it does not cash in.

Batch 3+4 — consolidated direct answer and the ranked shortlist (appended AFTER the results)

MEASURED — direct answer: NOTHING beats the current strafe on round wins. Across the 12 arm-vs-strafe comparisons of Batches 3–4 (8 non-reference arms, 270+270 battles on the frozen panel), zero arms beat strafe beyond the MDE. The only positive point estimate is wide_spread at +0.11 wins/run (95% CI [−0.32, +0.54], 7/11 decisive, p = 0.55, MDE 0.56) — i.e. the observed effect is ~5× smaller than the design can detect, so it is a TIE, not a win. Two arms are detectably worse (field_strong Δwins −0.29, p = 0.039; field_off Δwins −0.47, p = 0.0063 and Δdmg −10.8, p = 0.035). Meanwhile the strafe-over-tfil effect replicated in BOTH sessions a 3rd and 4th time (53.0% vs 42.2% and 52.6% vs 38.5% round-win rate), so the reference is stable.

MEASURED — the shape of the result. The response surface is FLAT around the current defaults on every tested axis: reversal dwell (2–8 / 12–40 / 6–20), picker hedge (spread/reach), bullet core/aura strength, corridor/wall strength, and wall-wing geometry. The one mechanism signal is that a SHORT dwell (fast_flip) hurts dodging (incoming +2.12 pp, sign test 12/15 p = 0.035) — the opposite of the naive "more reversals = harder to hit" story — and a LONG/short hedge both win nominally fewer rounds. Removing the wall/corridor gradient dodges slightly better but wins far less (field_off: best hit rate 12.65%, worst win rate 37.0%). This is a clean negative for "find a better arm by turning the existing knobs", and a positive for "the current retune is a local optimum of this design space".

Ranked shortlist for the final confirmation test (MEASURED/INFERRED):

  1. strafe — current defaults (TR_MOVEMENT=strafe). The measured champion. Confirm it head-to-head against tfil in one more independent session for the eventual ship decision. (MEASURED: it beats tfil by +0.42 wins/run, 95% CI [−0.66, −0.19] from tfil's perspective, 1/11 decisive, in Batch 4.)
  2. wide_spread (TR_STRAFE_SPREAD=2 TR_STRAFE_REACH=216). The ONLY arm of the 8 with a positive wins point estimate (+0.11, damage-neutral). It is currently a TIE, and resolving +0.11 would need far more than one batch (MDE 0.56 at n=15); include it as the single challenger in the confirmation session and expect a tie. (INFERRED: worth one look because it is the only arm on the correct side of zero.)
  3. strafe_notilt (TR_STRAFE_RANGE_TOL=999999, from Batches 1–2). Ties strafe on wins and removes the range-tuning surface; the recommended SHIP candidate if the default is ever flipped (per §5 item 4). Not re-tested here.

Drop (do not carry into the confirmation test): fast_flip (detectably worse dodging), slow_flip, narrow (negative, NS), bullet_strong, wall_tight (negative, NS), field_strong, field_off (detectably worse), and the Batch-2 range arms tilt_600 / tilt_250 (no separation).

Recommendation (INFERRED): by the §6 stop rule — a batch's best arm cannot beat strafe beyond the MDE — the movement hunt is closed: TR_MOVEMENT=strafe at its current defaults is the measured optimum of this design space, and the next stage is the gun (the owner's mandate). If a shipping decision is taken, the candidate is strafe (optionally strafe_notilt to drop the range knob); the default flip is a separate, explicit decision and was NOT made here.

Session log addition

session commit battles arms verdict
/tmp/ab/j119_b3 1256357 270 (0 failed; 1 excluded: Ascendant/strafe r1) strafe, tfil, fast_flip, slow_flip, wide_spread, narrow nothing beats strafe; wide_spread +0.11 NS (p=0.55)
/tmp/ab/j119_b4 1256357 270 (0 failed, 0 excluded) strafe, tfil, bullet_strong, field_strong, field_off, wall_tight nothing beats strafe; field_strong and field_off detectably WORSE

Final confirmation + SHIP — PRE-REGISTRATION

PRE-REGISTERED BEFORE THE BATTLES (this section is committed, and its commit is the frozen binary's source, before the first battle is launched). Session /tmp/ab/j120_final. Arms file tools/ab/arms_movement_final.txt, panel tools/ab/panel_movement.txt (frozen), 3 arms × 15 opponents × 5 runs × 3 rounds = 225 battles, conc 6, --reference tfil.

THE PRE-REGISTERED SHIP CRITERION (fixed before fighting). Ship the flip of TR_MOVEMENT's default from tfil to strafe only if, on this frozen panel and this one frozen binary, strafe beats tfil head-to-head on the PRIMARY metric round wins/run with both:

  1. the 95% CI on Δwins/run excluding 0; and
  2. the cross-opponent sign test favouring strafe (p < 0.05).

This is stricter than the campaign's §1 rule 2 (which would let a survival win carry a damage-neutral arm): the ship decision requires the win metric itself to clear both bars. If it does not pass, do not ship — say so and stop; that is a fully successful outcome. This batch uses 5 runs/arm (vs 3 in every previous batch) for higher power; wide_spread (the only positive point estimate from Batches 3–4, +0.11/run, NS) rides along as the single challenger.

Pre-registered prediction (before the battles): strafe passes both bars and ships; wide_spread ties strafe (no separable win); tfil is not preferred by any primary metric. A wrong prediction is recorded as wrong.

Outcome: pending — filled in below after the session.