Files
SirRoboGarage/docs/movement_campaign.md
T
SirStone 07766303f5 movement Batch 1: pure strafe (range tilt OFF) beats the shipped tfil on round wins across a 15-opponent panel
225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:

  strafe_notilt  wins/run +0.38  [CI +0.16,+0.60]  9/9 opponents p=0.0039
                 dmg/run  -10.2  [CI -25.8,+5.5]   p=0.61, MDE 20.4 (not detectable)
                 incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
  strafe_325     wins/run +0.33  [CI +0.04,+0.63]  10/12 p=0.0386
  ring           dmg/run  +31.2  [CI +11.5,+50.9]  13/15 p=0.0074, wins/run -0.04 (ns)
                 but hit rate 29.4% at 236 px: a damage/survival trade, not a win
  ring_notemp    indistinguishable from tfil on both primaries

Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.

Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
2026-09-26 01:13:29 +02:00

30 KiB
Raw Blame History

Movement campaign — ledger

Goal (owner's mandate, 2026-09-26 overnight): find the best 1v1 movement by measurement, then do the same for the gun. This file is the campaign's single source of truth: every later job appends a ## Batch N section and never edits an earlier one (a wrong earlier number gets a correction line, not a rewrite).

Owner's words: "I want you to do all tests and checks with the goal to have the best 1vs1 movement. You have all night, you can change every parameter. Continue until you found an amazing movement. When found do the same over for a gun."


0. The one caveat this campaign exists to close

Everything measured about movement before this campaign is DrussGT-only: docs/surfer_wiring_ab.md, the j107 range drift, the j113 BitBrain movement notes. The standing lesson of the night is that a one-opponent result is not a result:

an arm can take fewer hits and win fewer rounds (j107 / strafe): the verdict lives in damage/run + ROUND WINS, and hit rate is only ever an explanation.

So from here on the unit of evidence is the number of opponents, not the number of runs: the same arm must win on many opponents before it is called better.


1. Protocol (how every batch must be run)

Element Rule
Subject ONE frozen binary, built from git archive HEAD (tools/ab/tournament_run.sh does this; the commit sha and binary sha256 are recorded in session.json)
Arms env dicts only — no per-arm rebuild, ever; the arm file is a committed file, not a shell history
Panel the frozen panel tools/ab/panel_movement.txt. Adding/removing an opponent starts a new batch number
Pairing per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate
Isolation per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir
Serialization one battle fleet at a time. tournament_run.sh --wait-arena N refuses/stalls while another job's run_bridge_battle/TrBattleCapture/ModularBot_bin is alive (bracketed pgrep; never a broad pkill)
Liveness every declared env token must appear verbatim in OUR bot's own [env] boot report, else the run is excluded and named in the report; an undeclared TR_MOVEMENT in the process env is a fatal FAIL for the reference arm
Never shipped this is a measurement + design campaign: git status clean, defaults untouched, .gitignore untouched

Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit 1984a78)

Provenance note: the harness and these rules were written and staged before Batch 1 was fought, but a parallel job's git commit (j116, same working tree / same index) swept the staged files into its commit 1984a78 ("melee A/B doc…"). The rules are therefore committed under a neighbour's message — they are nonetheless dated before the data: no battle of Batch 1 had been launched when they were written, and Batch 1's session.json records the same commit 1984a78 as the frozen-binary source.`)

  1. Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean distance.
  2. BETTER than the reference iff one primary metric is up with a cross-opponent sign test p < 0.05 while the other does not go down; or the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (which is a real answer, not a failure).
  3. A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as not detectable, never as absent and never as a win.
  4. Somewhere to stop: if no arm beats the shipped tfil by rule 2 in Batch 1 and no arm shows a ≥ +MDE damage gain with p<0.10, the movement stage's first phase is closed with "the shipped tfil is the best movement we have measured" — that is a successful outcome, and the campaign moves to the gun axis rather than inventing more movement arms. See What would make us stop at the end.
  5. No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
  6. Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.

2. Stage 0 — what we already know (given, not re-derived)

Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary (docs/surfer_wiring_ab.md, commit 0f5cfe3):

arm dmg/run dmg taken round wins incoming hit rate
tfil (SHIPPED) 293 224 45/105 10.40%
strafe (range 325) 250 198 37/105 9.40%
surf 255 259 37/105 13.51%

Read: the shipped tfil deals the most damage and wins the most rounds while being hit the most; strafe dodges best and wins least. Plus j107: drifting 25–30 px closer made damage and wins worse, so the lever is not simply "get closer". Hypothesis entering the campaign: the 325 px range preference of strafe costs wins (INFERRED from DrussGT-only data — this is exactly what Batch 1 tests across a panel).


3. Batch 1 — isolating the range / aggression axis

Design. One frozen binary, five env-only arms, one frozen panel (tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 2 wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas), 3 runs × 3 rounds per (opponent, arm). Arm file: tools/ab/arms_movement_b1.txt.

# arm env what it isolates
1 tfil (none — shipped defaults) the arm to beat
2 strafe_notilt TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999 the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with no range steering at all
3 strafe_325 TR_MOVEMENT=strafe the current strafe default (range 325, tol 25, tilt 15/0.10)
4 ring TR_MOVEMENT=tfil_ring TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) with the range-weighted tile draw (band 100–200)
5 ring_notemp TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0 the control for #4: same retuned heat field, range weighting switched OFF (rand(candidates.high) path)

ring − ring_notemp is therefore the range-weighting lever alone, on a heat field that is already retuned. The originally-suggested 5th arm ("tfil with less saturated heat") is not buildable in this campaign: in common_libs/movements/the_floor_is_lava.nim CorridorHeat/WallHotness are Nim consts (env_report only reports them); only the tfil_ring copy reads them from the env. #5 is the honest substitute.

Pre-registered prediction (written before the battles finished): tfil still wins the panel on damage and round wins; strafe_notilt will beat strafe_325 on round wins (the range tilt is a net cost), and the ring arms will land between them. If instead the range-steering arms beat tfil on wins, the "range preference costs wins" hypothesis is confirmed across bots, not just against DrussGT.

Outcome — direct answer

Batch 1 is a NULL for the hypothesis that the shipped tfil is the best movement. It is not. Measured on the frozen 15-opponent panel, one frozen binary, 225 battles, 0 invalid runs, 0 liveness failures, 0 failed starts:

arm dmg/run wins/run round wins incoming hit rate dmg taken/run mean distance
tfil (SHIPPED) 118.9 1.22 55/135 (40.7%) 18.17% 199.8 382 px
strafe_notilt 108.8 1.60 72/135 (53.3%) 12.24% 150.3 456 px
strafe_325 111.8 1.56 70/135 (51.9%) 13.14% 155.7 436 px
ring_notemp 108.2 1.29 58/135 (43.0%) 16.67% 193.4 395 px
ring 150.1 1.18 53/135 (39.3%) 29.42% 225.1 236 px

Paired across opponents, the winner is strafe_notilt (pure perpendicular strafe, range steering provably off): Δwins/run +0.38 [95% CI +0.16, +0.60], positive on 9 of 9 decisive opponents (exact sign test p = 0.0039, sign-flip permutation p = 0.0039, Wilcoxon p = 0.0090), and Δdmg/run −10.2 [−25.8, +5.5], p = 0.61, MDE 20.4 ⇒ not detectable — i.e. +17 rounds out of 135 won, at no detectable damage cost, with a third fewer incoming hits (hit rate −7.3 pp, p = 6e-5, and 0/15 opponents in favour of tfil) and 50 less damage taken per run. strafe_325 is the same effect, slightly smaller (Δwins/run +0.33, [0.04, +0.63], p = 0.039, 10/12) — the two strafe arms are not separable from each other by this batch.

ring is the opposite trade and must not be read as a movement win: it deals +31.2 dmg/run (+26%, p = 0.0074, 13/15) but wins no more rounds (Δwins −0.04, p = 1.00) and pays for the damage with the panel's worst dodging (hit rate 29.42% vs 18.17%, +25 dmg taken/run) because it fights at a mean 236 px (vs 382/456). ring_notemp — the same retuned heat field with the range weighting switched off — is indistinguishable from tfil on both primaries, so the heat-field retune alone is not what makes strafe win (INFERRED: ring_notemp also differs from tfil in commit ticks and wall radiance, so this is evidence against, not a clean isolation).

Mechanism (MEASURED, and the reason the win is a movement win): in 216 of the 219 attributable runs, our round-win count equals exactly the number of rounds in which the opponent's death event appears — round wins in this harness are survival wins. The winning arm survives by taking fewer, weaker hits at longer range, not by dealing more damage (its damage is unchanged).

DIRECT ANSWER. The best 1v1 movement measured across this panel is TR_MOVEMENT=strafe with the range tilt disabled (pure perpendicular strafe, no range steering). It beats the shipped tfil on round wins by an effect that survives the between-opponent spread (observed +0.38 vs MDE 0.29; 9/9 opponents; CI excludes 0) with no detectable damage cost, and it dodges substantially better. strafe_325 (the current strafe default) is essentially the same arm. The shipped tfil is the worst of the five on round wins: the hypothesis in §2 that its win came from the DrussGT-only measurement is supported — on a panel it loses to both strafe arms.

The pre-registered prediction for this batch was WRONG and is recorded as wrong: I predicted tfil would still win the panel (it came last on wins) and that strafe_notilt would beat strafe_325 on wins (it does by +0.05 wins/run, which this batch cannot resolve).

Honest readings of the pre-registered rule (both printed by the analyzer; the strict reading is the literal one and it is NOT satisfied by anything):

  • strict (the other metric's mean delta is not negative at all): no arm is BETTER than tfil. The two strafe arms win more rounds but their mean damage is 7–10/run lower (inside the MDE, but negative).
  • substantive (the other primary metric is not detectably down — sign test not significant and |Δ| < its MDE, per rule 3): strafe_notilt, strafe_325 and ring are each BETTER than tfil on one primary metric.
  • The ordering is identical under both readings, and under the standing rule (round wins first, then damage) the winner is strafe_notilt.

The analyzer's full report (verbatim)

MEASURED: session

  • commit 1984a780f494ce246e0f916934b9581e07c89ed2, frozen binary sha256 1817c75ab1d0…

  • 15 opponents × 5 arms × 3 runs × 3 rounds = 225 battles, conc=6

  • arms file arms_movement_b1.txt, panel file panel_movement.txt

  • reference arm: tfil — every delta below is (arm − tfil), opponent by opponent

  • liveness: 0 run(s) excluded (225 total)

MEASURED: per-opponent paired table (per arm)

tfil — shipped baseline (movement engine tfil, every knob at its default) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→124.7 +0.0 1.67→1.67 +0.00 +0.0 +0.00 452→452
Diamond dodger 39.5→39.5 +0.0 0.00→0.00 +0.00 +0.0 +0.00 458→458
Dookious dodger 130.1→130.1 +0.0 1.00→1.00 +0.00 +0.0 +0.00 410→410
GresSuffurd dodger 129.2→129.2 +0.0 1.33→1.33 +0.00 +0.0 +0.00 413→413
CassiusClay dodger 73.4→73.4 +0.0 0.33→0.33 +0.00 +0.0 +0.00 339→339
RetroGirl pattern 181.0→181.0 +0.0 2.33→2.33 +0.00 +0.0 +0.00 402→402
TripHammer pattern 59.5→59.5 +0.0 0.33→0.33 +0.00 +0.0 +0.00 418→418
Coriantumr pattern 67.8→67.8 +0.0 1.00→1.00 +0.00 +0.0 +0.00 424→424
WallAvoider wallfollower 229.1→229.1 +0.0 2.67→2.67 +0.00 +0.0 +0.00 277→277
HawkOnFire cornercamper 150.7→150.7 +0.0 1.67→1.67 +0.00 +0.0 +0.00 410→410
SpinBot spinner 279.3→279.3 +0.0 3.00→3.00 +0.00 +0.0 +0.00 351→351
DiamondStealer rammer 139.4→139.4 +0.0 0.67→0.67 +0.00 +0.0 +0.00 235→235
BlitzBat brawler 54.3→54.3 +0.0 2.00→2.00 +0.00 +0.0 +0.00 422→422
YersiniaPestis aggressive 52.3→52.3 +0.0 0.33→0.33 +0.00 +0.0 +0.00 401→401
Ascendant aggressive 73.3→73.3 +0.0 0.00→0.00 +0.00 +0.0 +0.00 317→317

strafe_notilt — strafe, range steering OFF (tilt always 0) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→115.7 -9.0 1.67→1.67 +0.00 -31.6 -2.95 452→535
Diamond dodger 39.5→56.6 +17.1 0.00→0.00 +0.00 -62.9 -7.97 458→543
Dookious dodger 130.1→105.8 -24.3 1.00→1.00 +0.00 -60.1 -6.08 410→484
GresSuffurd dodger 129.2→111.3 -17.8 1.33→2.33 +1.00 -36.3 -8.61 413→482
CassiusClay dodger 73.4→92.1 +18.7 0.33→1.33 +1.00 -55.5 -7.24 339→397
RetroGirl pattern 181.0→133.2 -47.8 2.33→2.67 +0.33 -0.9 -2.46 402→447
TripHammer pattern 59.5→57.5 -2.1 0.33→0.67 +0.33 -45.0 -3.84 418→547
Coriantumr pattern 67.8→77.2 +9.4 1.00→1.67 +0.67 -54.2 -3.35 424→554
WallAvoider wallfollower 229.1→162.1 -66.9 2.67→2.67 +0.00 -54.9 -9.49 277→361
HawkOnFire cornercamper 150.7→95.7 -55.0 1.67→1.67 +0.00 -76.2 -7.27 410→567
SpinBot spinner 279.3→271.3 -8.0 3.00→3.00 +0.00 -42.7 -14.27 351→335
DiamondStealer rammer 139.4→154.7 +15.2 0.67→1.00 +0.33 -33.7 -2.80 235→243
BlitzBat brawler 54.3→41.5 -12.8 2.00→3.00 +1.00 -105.6 -13.42 422→588
YersiniaPestis aggressive 52.3→78.4 +26.0 0.33→1.00 +0.67 -45.0 -6.09 401→397
Ascendant aggressive 73.3→78.2 +4.9 0.00→0.33 +0.33 -36.9 -14.30 317→364

strafe_325 — strafe default (range 325, tol 25, tilt 15/0.10) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→118.5 -6.2 1.67→1.33 -0.33 -17.5 -1.77 452→494
Diamond dodger 39.5→71.8 +32.2 0.00→0.00 +0.00 -38.7 -5.93 458→506
Dookious dodger 130.1→84.0 -46.0 1.00→2.00 +1.00 -108.4 -9.12 410→457
GresSuffurd dodger 129.2→112.3 -16.9 1.33→2.00 +0.67 -26.1 -5.44 413→430
CassiusClay dodger 73.4→81.5 +8.1 0.33→1.33 +1.00 -68.4 -7.31 339→386
RetroGirl pattern 181.0→169.4 -11.6 2.33→2.33 +0.00 -7.9 -3.52 402→425
TripHammer pattern 59.5→43.6 -15.9 0.33→0.67 +0.33 -43.0 -4.74 418→500
Coriantumr pattern 67.8→96.5 +28.7 1.00→1.67 +0.67 -46.2 -3.37 424→494
WallAvoider wallfollower 229.1→163.8 -65.2 2.67→1.67 -1.00 -12.1 -7.90 277→332
HawkOnFire cornercamper 150.7→119.8 -30.9 1.67→2.00 +0.33 -78.8 -5.38 410→517
SpinBot spinner 279.3→259.3 -20.0 3.00→3.00 +0.00 -37.3 -12.96 351→436
DiamondStealer rammer 139.4→135.3 -4.1 0.67→1.33 +0.67 -25.4 -3.50 235→274
BlitzBat brawler 54.3→60.5 +6.3 2.00→2.67 +0.67 -88.1 -10.88 422→531
YersiniaPestis aggressive 52.3→69.3 +16.9 0.33→0.67 +0.33 -15.7 -2.95 401→383
Ascendant aggressive 73.3→90.4 +17.1 0.00→0.67 +0.67 -47.0 -12.83 317→373

ring — tfil_ring (retuned heat field + range weighting 100-200) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→127.8 +3.1 1.67→0.67 -1.00 +105.6 +12.67 452→244
Diamond dodger 39.5→46.4 +6.9 0.00→0.00 +0.00 +40.9 +17.58 458→240
Dookious dodger 130.1→149.2 +19.2 1.00→0.33 -0.67 +50.0 +8.21 410→267
GresSuffurd dodger 129.2→222.7 +93.6 1.33→1.33 +0.00 +83.9 +12.13 413→234
CassiusClay dodger 73.4→102.5 +29.1 0.33→0.33 +0.00 +21.8 +5.21 339→242
RetroGirl pattern 181.0→249.8 +68.8 2.33→3.00 +0.67 -41.9 +0.85 402→209
TripHammer pattern 59.5→76.2 +16.7 0.33→0.00 -0.33 +47.8 +13.66 418→266
Coriantumr pattern 67.8→119.4 +51.7 1.00→1.00 +0.00 +32.1 +10.39 424→258
WallAvoider wallfollower 229.1→219.5 -9.6 2.67→1.67 -1.00 +57.2 +7.21 277→232
HawkOnFire cornercamper 150.7→176.4 +25.7 1.67→2.33 +0.67 -41.8 +11.76 410→229
SpinBot spinner 279.3→336.0 +56.7 3.00→3.00 +0.00 +5.3 +24.20 351→172
DiamondStealer rammer 139.4→148.7 +9.3 0.67→1.33 +0.67 -39.4 -1.78 235→214
BlitzBat brawler 54.3→156.7 +102.5 2.00→2.67 +0.67 +8.4 +12.71 422→221
YersiniaPestis aggressive 52.3→43.3 -9.0 0.33→0.00 -0.33 +36.2 +10.82 401→266
Ascendant aggressive 73.3→76.7 +3.4 0.00→0.00 +0.00 +14.0 +13.54 317→243

ring_notemp — tfil_ring, range weighting OFF (temp 0) (paired on 15 opponents)

opponent style dmg/run ref→arm Δdmg wins/run ref→arm Δwins Δdmg taken Δhit rate (pp) dist ref→arm
DrussGT dodger 124.7→108.0 -16.7 1.67→1.33 -0.33 -0.8 -0.72 452→436
Diamond dodger 39.5→42.9 +3.4 0.00→0.00 +0.00 +9.2 -1.47 458→447
Dookious dodger 130.1→115.2 -14.8 1.00→1.33 +0.33 -24.4 -3.35 410→418
GresSuffurd dodger 129.2→118.3 -10.8 1.33→1.33 +0.00 +20.9 -1.75 413→433
CassiusClay dodger 73.4→92.6 +19.2 0.33→1.67 +1.33 -51.0 -5.72 339→388
RetroGirl pattern 181.0→126.2 -54.8 2.33→1.33 -1.00 +55.3 +3.00 402→405
TripHammer pattern 59.5→44.3 -15.2 0.33→0.33 +0.00 -4.9 +0.65 418→453
Coriantumr pattern 67.8→84.9 +17.1 1.00→1.00 +0.00 -1.7 +0.24 424→430
WallAvoider wallfollower 229.1→141.1 -88.0 2.67→3.00 +0.33 -124.4 -14.57 277→339
HawkOnFire cornercamper 150.7→134.2 -16.5 1.67→1.33 -0.33 +28.8 +1.77 410→436
SpinBot spinner 279.3→287.7 +8.3 3.00→3.00 +0.00 +0.0 +0.45 351→308
DiamondStealer rammer 139.4→154.5 +15.1 0.67→2.00 +1.33 -57.4 -6.08 235→254
BlitzBat brawler 54.3→57.3 +3.1 2.00→1.67 -0.33 +57.7 +2.97 422→436
YersiniaPestis aggressive 52.3→45.2 -7.1 0.33→0.00 -0.33 +21.2 +2.66 401→376
Ascendant aggressive 73.3→70.3 -3.0 0.00→0.00 +0.00 -23.2 -8.01 317→371

MEASURED: pooled dashboard (all valid runs, NOT the verdict)

arm runs dmg/run dmg taken/run wins/run round wins win rate incoming hit rate mean distance
tfil 45 118.9 199.8 1.22 55/135 40.7% 18.17% 382
strafe_notilt 45 108.8 150.3 1.60 72/135 53.3% 12.24% 456
strafe_325 45 111.8 155.7 1.56 70/135 51.9% 13.14% 436
ring 45 150.1 225.1 1.18 53/135 39.3% 29.42% 236
ring_notemp 45 108.2 193.4 1.29 58/135 43.0% 16.67% 395

MEASURED: cross-opponent aggregation (the verdict layer)

Deltas are per-opponent (arm − reference). spread is the SD of those deltas ACROSS opponents; SE = spread/√n; 95% CI = mean ± t·SE. Sign test = how many opponents the arm wins (ties dropped), exact binomial; sign-flip = permutation test on the mean of the deltas.

arm metric mean Δ spread (SD) SE 95% CI sign test (wins/n) p(sign) p(sign-flip) Wilcoxon p MDE
strafe_notilt damage -10.15 28.19 7.28 [-25.76, +5.46] 6/15 0.6072 0.1887 (exact 2^15) 0.3787 20.39
strafe_notilt wins +0.38 0.40 0.10 [+0.16, +0.60] 9/9 0.003906 0.003906 (exact 2^15) 0.008969 0.29
strafe_notilt damage_taken -49.44 23.28 6.01 [-62.33, -36.54] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 16.84
strafe_notilt hit_rate -7.34 4.10 1.06 [-9.61, -5.07] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 2.96
strafe_notilt dist +74.36 54.83 14.16 [+44.00, +104.73] 13/15 0.007385 0.0003662 (exact 2^15) 0.001621 39.66
strafe_325 damage -7.15 27.04 6.98 [-22.13, +7.82] 6/15 0.6072 0.3276 (exact 2^15) 0.5137 19.56
strafe_325 wins +0.33 0.53 0.14 [+0.04, +0.63] 10/12 0.03857 0.04688 (exact 2^15) 0.05424 0.39
strafe_325 damage_taken -44.05 29.86 7.71 [-60.58, -27.51] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 21.60
strafe_325 hit_rate -6.51 3.58 0.92 [-8.49, -4.53] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 2.59
strafe_325 dist +53.77 33.71 8.70 [+35.10, +72.44] 14/15 0.0009766 0.0001831 (exact 2^15) 0.001092 24.38
ring damage +31.19 35.61 9.19 [+11.47, +50.91] 13/15 0.007385 0.001587 (exact 2^15) 0.004932 25.76
ring wins -0.04 0.56 0.14 [-0.36, +0.27] 4/9 1 0.8828 (exact 2^15) 0.6776 0.41
ring damage_taken +25.35 43.46 11.22 [+1.27, +49.42] 12/15 0.03516 0.04059 (exact 2^15) 0.05708 31.44
ring hit_rate +10.61 6.32 1.63 [+7.11, +14.11] 14/15 0.0009766 0.0001831 (exact 2^15) 0.001092 4.57
ring dist -146.25 60.84 15.71 [-179.94, -112.56] 0/15 6.104e-05 6.104e-05 (exact 2^15) 0.0007265 44.01
ring_notemp damage -10.73 28.24 7.29 [-26.37, +4.91] 6/15 0.6072 0.1772 (exact 2^15) 0.3487 20.43
ring_notemp wins +0.07 0.61 0.16 [-0.27, +0.40] 4/9 1 0.8086 (exact 2^15) 0.9525 0.44
ring_notemp damage_taken -6.31 46.39 11.98 [-32.00, +19.38] 6/14 0.7905 0.632 (exact 2^15) 0.8017 33.56
ring_notemp hit_rate -2.00 4.87 1.26 [-4.69, +0.70] 7/15 1 0.1341 (exact 2^15) 0.2681 3.52
ring_notemp dist +13.34 29.63 7.65 [-3.07, +29.75] 11/15 0.1185 0.1037 (exact 2^15) 0.1055 21.43

By inferred style (explanation only, never the verdict)

arm style n mean Δdmg mean Δwins mean Δhit rate (pp)
strafe_notilt aggressive 2 +15.4 +0.50 -10.19
strafe_notilt brawler 1 -12.8 +1.00 -13.42
strafe_notilt cornercamper 1 -55.0 +0.00 -7.27
strafe_notilt dodger 5 -3.1 +0.40 -6.57
strafe_notilt pattern 3 -13.5 +0.44 -3.21
strafe_notilt rammer 1 +15.2 +0.33 -2.80
strafe_notilt spinner 1 -8.0 +0.00 -14.27
strafe_notilt wallfollower 1 -66.9 +0.00 -9.49
strafe_325 aggressive 2 +17.0 +0.50 -7.89
strafe_325 brawler 1 +6.3 +0.67 -10.88
strafe_325 cornercamper 1 -30.9 +0.33 -5.38
strafe_325 dodger 5 -5.7 +0.47 -5.91
strafe_325 pattern 3 +0.4 +0.33 -3.88
strafe_325 rammer 1 -4.1 +0.67 -3.50
strafe_325 spinner 1 -20.0 +0.00 -12.96
strafe_325 wallfollower 1 -65.2 -1.00 -7.90
ring aggressive 2 -2.8 -0.17 +12.18
ring brawler 1 +102.5 +0.67 +12.71
ring cornercamper 1 +25.7 +0.67 +11.76
ring dodger 5 +30.4 -0.33 +11.16
ring pattern 3 +45.7 +0.11 +8.30
ring rammer 1 +9.3 +0.67 -1.78
ring spinner 1 +56.7 +0.00 +24.20
ring wallfollower 1 -9.6 -1.00 +7.21
ring_notemp aggressive 2 -5.1 -0.17 -2.67
ring_notemp brawler 1 +3.1 -0.33 +2.97
ring_notemp cornercamper 1 -16.5 -0.33 +1.77
ring_notemp dodger 5 -4.0 +0.27 -2.60
ring_notemp pattern 3 -17.6 -0.33 +1.30
ring_notemp rammer 1 +15.1 +1.33 -6.08
ring_notemp spinner 1 +8.3 +0.00 +0.45
ring_notemp wallfollower 1 -88.0 +0.33 -14.57

The pre-registered verdict table, as printed by the analyzer

PRIMARY metrics are dmg/run and wins/run; hit rate is never the verdict. The pre-registered rule says an arm is BETTER when one primary metric is UP at sign-test p<0.05 while the other does not go down. That phrase has two readings and BOTH are printed:

  • strict — the other metric's mean delta is not negative at all (Δ >= 0). Nothing can be BETTER while it costs any mean damage.
  • substantive — the other metric's delta is not detectably down: the sign test is not significant and the delta is smaller than that metric's MDE (the pre-registered rule 3 says an effect under the MDE is not detectable, so it cannot count as a loss).
rank arm Δwins/run Δdmg/run sign test wins sign test dmg verdict (strict) verdict (substantive)
1 strafe_notilt +0.38 -10.2 9/9 p=0.003906 6/15 p=0.6072 not distinguishable BETTER
2 strafe_325 +0.33 -7.2 10/12 p=0.03857 6/15 p=0.6072 not distinguishable BETTER
3 ring_notemp +0.07 -10.7 4/9 p=1 6/15 p=0.6072 not distinguishable not distinguishable
4 ring -0.04 +31.2 4/9 p=1 13/15 p=0.007385 not distinguishable BETTER

Reference tfil: 118.9 dmg/run, 1.22 wins/run, 18.17% incoming, 382 px.

Highest wins delta: strafe_notilt (+0.38 wins/run, -10.2 dmg/run) — strict: not distinguishable, substantive: BETTER.


4. What to try next (seeded; every later job adds its own)

  1. If a range-steering arm wins the panel: the win is a range effect, so sweep the band on the winning engine (e.g. TR_TFIL_RANGE_LO/HI on tfil_ring, TR_STRAFE_RANGE on strafe) with the same panel, 4–5 bands, and look for a plateau rather than a peak. A plateau is a result; a peak is a coin flip.
  2. If nothing beats tfil: stop tuning movement by feel. The next honest lever is enemy-model-driven placement (keep the bot where the enemy's expected hit probability is lowest given its gun model), which needs a per-opponent measurement, not a knob.
  3. Melee is a different game (j116's finding): if a melee campaign is opened, it needs its own panel and its own ledger section — do not reuse the 1v1 panel's verdicts.
  4. Close the loop with the enemy's own model: ab_mechanism.py-style instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by range band) is available and cheap — use it to explain a win, never to declare one.
  5. Then the gun (owner's next stage): the same harness, a gun panel, and the same paired-with-sign-test statistics. docs/surfer_wiring_ab.md and the j117 gauntlet are the gun-side priors to beat.

5. What would make us stop

  • Stop the movement stage when a batch produces an arm that is BETTER than the shipped default on the frozen panel by rule 2 and the effect survives the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of the panel). That arm becomes the new default candidate (shipping is a separate decision — this campaign never edits a shipped default).
  • Stop and move to the gun if two consecutive batches fail to produce an arm that beats the shipped tfil beyond the MDE: at that point the honest conclusion is "the shipped movement is the measured optimum of this design space", which is a successful campaign outcome, not a failure.
  • Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.

6. How to run a batch (exact commands)

# 1. wait for the arena (this job may not be the only one fighting)
tools/ab/tournament_run.sh \
    --arms    tools/ab/arms_movement_b1.txt \
    --panel   tools/ab/panel_movement.txt \
    --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
    --reference tfil \
    --outdir /tmp/ab/j118_b1

# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil

The runner writes <outdir>/session.json (commit sha, binary sha256, arms, panel) so any later job can re-analyze an old session offline, with no arena.