movement Batch 2: the strafe win over shipped tfil replicates; the range target decides nothing

Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.

Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.

The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.

Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
This commit is contained in:
2026-09-26 01:31:08 +02:00
parent 07766303f5
commit 7d3645da5a
2 changed files with 378 additions and 35 deletions
+353 -35
View File
@@ -175,12 +175,19 @@ strafe, no range steering). It beats the shipped `tfil` on round wins by an
effect that **survives the between-opponent spread** (observed +0.38 vs MDE
0.29; 9/9 opponents; CI excludes 0) with **no detectable damage cost**, and it
dodges substantially better. `strafe_325` (the current strafe default) is
essentially the same arm. The shipped `tfil` is the **worst of the five on round
wins**: the hypothesis in §2 that its win came from the DrussGT-only measurement
is **supported** — on a panel it loses to both strafe arms.
essentially the same arm. The shipped `tfil` is **4th of the five on round
wins** (only `ring` is nominally lower, and `tfil` vs `ring` on wins is a dead
heat, p = 1.00): the hypothesis in §2 that its win came from the DrussGT-only
measurement is **supported** — on a panel it loses to both strafe arms.
**Correction (added after the Batch-1 commit `0776630`, whose message says "last
of five"):** `tfil` is 4th of five, not last — `ring` is nominally 0.04 wins/run
lower and that difference is not significant. The batch message overstates one
word; the numbers it quotes are the measured ones.
**The pre-registered prediction for this batch was WRONG and is recorded as
wrong:** I predicted `tfil` would still win the panel (it came last on wins) and
wrong:** I predicted `tfil` would still win the panel (it came 4th of five on
wins) and
that `strafe_notilt` would beat `strafe_325` on wins (it does by +0.05 wins/run,
which this batch cannot resolve).
@@ -404,43 +411,335 @@ Highest wins delta: `strafe_notilt` (+0.38 wins/run, -10.2 dmg/run) — strict:
---
## 4. What to try next (seeded; every later job adds its own)
## 4. Batch 2 — the range axis ON the winning engine (replication)
1. **If a range-steering arm wins the panel:** the win is a *range* effect, so
sweep the *band* on the winning engine (e.g. `TR_TFIL_RANGE_LO/HI` on
`tfil_ring`, `TR_STRAFE_RANGE` on `strafe`) with the same panel, 4–5 bands,
and look for a plateau rather than a peak. A plateau is a result; a peak is a
coin flip.
2. **If nothing beats `tfil`:** stop tuning movement by feel. The next honest
lever is *enemy-model-driven* placement (keep the bot where the enemy's
expected hit probability is lowest *given its gun model*), which needs a
per-opponent measurement, not a knob.
3. **Melee is a different game** (j116's finding): if a melee campaign is
opened, it needs its own panel and its own ledger section — do not reuse the
1v1 panel's verdicts.
4. **Close the loop with the enemy's own model:** `ab_mechanism.py`-style
instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by
range band) is available and cheap — use it to *explain* a win, never to
declare one.
5. **Then the gun** (owner's next stage): the same harness, a gun panel, and the
same paired-with-sign-test statistics. `docs/surfer_wiring_ab.md` and the j117
gauntlet are the gun-side priors to beat.
**Design.** Same frozen panel, same 3 runs × 3 rounds, new session
`/tmp/ab/j118_b2` (commit `8efa627`, 225 battles, **0 invalid runs, 0 failed
starts**; no source file changed between `1984a78` and `8efa627` — only a
parallel job's new docs/tools — so this is the same code). Arms
(`tools/ab/arms_movement_b2.txt`): the winner and the strafe default from
Batch 1 (replication), plus the tilt re-armed at **600 px** and at **250 px**,
i.e. `strafe_notilt` has no range control and drifts to ~456 px, so these two
separate *"the range value is the lever"* from *"the tilt mechanism is the
cost"*.
## 5. What would make us stop
**Pre-registered prediction (written before the battles):** if the range value
drives the win, `tilt_600` should beat `strafe_notilt`; if the tilt mechanism
itself is the cost, both tilt arms should lose to `strafe_notilt`. **Both halves
turned out wrong**, and that is the useful part:
* **Stop the movement stage** when a batch produces an arm that is BETTER than
the shipped default on the frozen panel by rule 2 **and** the effect survives
the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of
the panel). That arm becomes the new default *candidate* (shipping is a
separate decision — this campaign never edits a shipped default).
* **Stop and move to the gun** if two consecutive batches fail to produce an arm
that beats the shipped `tfil` beyond the MDE: at that point the honest
conclusion is *"the shipped movement is the measured optimum of this design
space"*, which is a successful campaign outcome, not a failure.
| arm | target / emergent range | dmg/run | wins/run | round wins | win rate | incoming hit rate | dmg taken/run | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `tfil` (SHIPPED) | none | 114.1 | 1.18 | 53/135 | 39.3% | 17.63% | 196.5 | 394 px |
| `strafe_325` | 325 | 111.8 | **1.76** | **79/135** | **58.5%** | 12.52% | 144.2 | 434 px |
| `strafe_notilt` | none (drifts) | 101.5 | 1.64 | 74/135 | 54.8% | 12.05% | 148.6 | 459 px |
| `tilt_600` | 600 | 102.3 | 1.58 | 71/135 | 52.6% | **11.67%** | 146.4 | **478 px** |
| `tilt_250` | 250 | 112.0 | 1.56 | 70/135 | 51.9% | 13.71% | 162.2 | 415 px |
Paired vs `tfil`: `strafe_325` **+0.58 wins/run** [CI +0.27, +0.89], 11/12
decisive opponents, p = 0.0063; `strafe_notilt` **+0.47** [+0.22, +0.72], 10/11,
p = 0.0117; `tilt_600` +0.40 [+0.04, +0.76] (sign test 8/11 p = 0.23,
sign-flip p = 0.049); `tilt_250` +0.38 [+0.07, +0.69], 10/12, p = 0.0386. Damage
deltas are −2.1 … −12.5 (10% of the mean at worst) and never positive;
incoming-hit-rate deltas are −5.2 … −6.9 pp with **0/15 opponents favouring
`tfil`**.
**What this batch actually establishes**
1. **The strafe engine's win over the shipped `tfil` replicates.** Batch 1:
+0.33 / +0.38 wins/run for the two strafe arms; Batch 2: +0.58 / +0.47 — the
same direction, the same magnitude band, in an independent session, with
**0/15 opponents** going the other way on incoming hit rate in either
session. Pooled descriptively, the four strafe-family arms won **52–58%** of
rounds in Batch 2 and **52–53%** in Batch 1, against `tfil`'s **39–41%**.
2. **The baseline is reproducible across sessions:** `tfil` won 40.7% of rounds
in Batch 1 and 39.3% in Batch 2 (Δ 1.4 pp), and dealt 118.9 vs 114.1 dmg/run.
The harness gives the same answer twice, which is why the win delta above is
believable.
3. **The range TARGET is not the lever.** Re-arming the tilt at 600 px moved the
achieved distance to 478 px and at 250 px to 415 px (vs 459 px with no
steering), and **none of the three was separable from the others on wins**.
The win comes from the engine, at any of these distances; the range value
within 415–478 px does not decide it. This **overturns the Batch-1 reading**
that "the tilt costs wins" (Batch 1: no-tilt > 325; Batch 2: 325 > no-tilt,
both inside noise) — the honest statement is *the tilt's effect on wins is
below this design's resolution (MDE ≈ 0.3–0.4 wins/run)*.
4. **`dmg/run` and `wins/run` remain different questions.** The arm that dealt
the most damage in Batch 1 (`ring`, +31) won nothing extra; the arms that win
in Batch 2 are not the high-damage ones (`strafe_325` 111.8 dmg/run vs
`tilt_250` 112.0). The win is bought with **survival** — 50 fewer damage
taken per run, −5…−7 pp incoming hit rate — not with output.
**DIRECT ANSWER after two batches (unchanged, now replicated).** The best 1v1
movement measured on this panel is the **strafe engine**: `TR_MOVEMENT=strafe`.
Its two Batch-1/2 configs are statistically tied with each other; if a config
must be named, `TR_MOVEMENT=strafe` at its shipped range (325 px) has the best
pooled round-win rate of the five arms in Batch 2 (58.5%) and ties `strafe_notilt`
in Batch 1, while `strafe_notilt` is the simpler arm (it has no range steering to
mis-tune). It is better than the shipped `tfil` by a margin that survives the
between-opponent spread: +0.33…+0.58 wins/run, 95% CIs excluding 0 in three of
four measurements, 9/9 + 10/11 + 11/12 decisive opponents in favour, MDE 0.29–0.40
vs observed 0.38–0.58. Rejecting "no change": `tfil`'s win share of 39–41% is
**not** the best movement we have measured.
### The analyzer's full report (verbatim)
### MEASURED: session
* commit `8efa627c05137d5a949d5a899c71fc55b5a1daf5`, frozen binary sha256 `005d010d8593…`
* 15 opponents × 5 arms × 3 runs × 3 rounds = 225 battles, conc=6
* arms file `arms_movement_b2.txt`, panel file `panel_movement.txt`
* reference arm: **`tfil`** — every delta below is (arm − tfil), opponent by opponent
* liveness: 0 run(s) excluded (225 total)
### MEASURED: per-opponent paired table (per arm)
#### `tfil` — shipped baseline, re-measured in this session (replication) (paired on 15 opponents)
| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| DrussGT | dodger | 120.8→120.8 | +0.0 | 0.67→0.67 | +0.00 | +0.0 | +0.00 | 445→445 |
| Diamond | dodger | 55.9→55.9 | +0.0 | 0.00→0.00 | +0.00 | +0.0 | +0.00 | 457→457 |
| Dookious | dodger | 86.5→86.5 | +0.0 | 1.33→1.33 | +0.00 | +0.0 | +0.00 | 451→451 |
| GresSuffurd | dodger | 114.0→114.0 | +0.0 | 1.33→1.33 | +0.00 | +0.0 | +0.00 | 417→417 |
| CassiusClay | dodger | 85.3→85.3 | +0.0 | 0.67→0.67 | +0.00 | +0.0 | +0.00 | 388→388 |
| RetroGirl | pattern | 181.7→181.7 | +0.0 | 2.00→2.00 | +0.00 | +0.0 | +0.00 | 391→391 |
| TripHammer | pattern | 55.8→55.8 | +0.0 | 0.00→0.00 | +0.00 | +0.0 | +0.00 | 469→469 |
| Coriantumr | pattern | 100.9→100.9 | +0.0 | 1.67→1.67 | +0.00 | +0.0 | +0.00 | 444→444 |
| WallAvoider | wallfollower | 150.0→150.0 | +0.0 | 2.00→2.00 | +0.00 | +0.0 | +0.00 | 317→317 |
| HawkOnFire | cornercamper | 115.1→115.1 | +0.0 | 1.67→1.67 | +0.00 | +0.0 | +0.00 | 419→419 |
| SpinBot | spinner | 302.0→302.0 | +0.0 | 3.00→3.00 | +0.00 | +0.0 | +0.00 | 316→316 |
| DiamondStealer | rammer | 140.1→140.1 | +0.0 | 1.00→1.00 | +0.00 | +0.0 | +0.00 | 236→236 |
| BlitzBat | brawler | 74.5→74.5 | +0.0 | 2.00→2.00 | +0.00 | +0.0 | +0.00 | 420→420 |
| YersiniaPestis | aggressive | 65.6→65.6 | +0.0 | 0.33→0.33 | +0.00 | +0.0 | +0.00 | 377→377 |
| Ascendant | aggressive | 62.8→62.8 | +0.0 | 0.00→0.00 | +0.00 | +0.0 | +0.00 | 362→362 |
#### `strafe_notilt` — Batch-1 winner, replication (paired on 15 opponents)
| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| DrussGT | dodger | 120.8→92.2 | -28.6 | 0.67→0.67 | +0.00 | -44.7 | -4.02 | 445→523 |
| Diamond | dodger | 55.9→63.5 | +7.6 | 0.00→0.33 | +0.33 | -43.4 | -7.28 | 457→533 |
| Dookious | dodger | 86.5→88.5 | +2.1 | 1.33→1.33 | +0.00 | -31.1 | -3.65 | 451→486 |
| GresSuffurd | dodger | 114.0→115.8 | +1.8 | 1.33→2.33 | +1.00 | -79.2 | -5.44 | 417→470 |
| CassiusClay | dodger | 85.3→91.2 | +5.9 | 0.67→1.33 | +0.67 | -43.8 | -5.85 | 388→390 |
| RetroGirl | pattern | 181.7→131.3 | -50.4 | 2.00→2.67 | +0.67 | -3.1 | -7.61 | 391→444 |
| TripHammer | pattern | 55.8→65.7 | +9.9 | 0.00→0.67 | +0.67 | -39.0 | -4.46 | 469→553 |
| Coriantumr | pattern | 100.9→60.5 | -40.4 | 1.67→1.33 | -0.33 | -16.9 | -1.32 | 444→572 |
| WallAvoider | wallfollower | 150.0→166.5 | +16.4 | 2.00→2.67 | +0.67 | -42.3 | +0.54 | 317→327 |
| HawkOnFire | cornercamper | 115.1→109.8 | -5.3 | 1.67→2.67 | +1.00 | -79.0 | -8.87 | 419→556 |
| SpinBot | spinner | 302.0→259.5 | -42.5 | 3.00→3.00 | +0.00 | -48.0 | -23.48 | 316→406 |
| DiamondStealer | rammer | 140.1→117.9 | -22.2 | 1.00→1.00 | +0.00 | -29.9 | -3.78 | 236→260 |
| BlitzBat | brawler | 74.5→43.3 | -31.2 | 2.00→2.33 | +0.33 | -94.9 | -9.35 | 420→571 |
| YersiniaPestis | aggressive | 65.6→49.7 | -15.9 | 0.33→1.33 | +1.00 | -73.3 | -7.89 | 377→414 |
| Ascendant | aggressive | 62.8→67.8 | +5.0 | 0.00→1.00 | +1.00 | -49.5 | -10.67 | 362→384 |
#### `strafe_325` — strafe default (range 325), replication (paired on 15 opponents)
| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| DrussGT | dodger | 120.8→129.4 | +8.6 | 0.67→1.67 | +1.00 | -23.7 | -2.26 | 445→477 |
| Diamond | dodger | 55.9→58.0 | +2.1 | 0.00→0.00 | +0.00 | -47.9 | -5.92 | 457→495 |
| Dookious | dodger | 86.5→115.7 | +29.2 | 1.33→1.67 | +0.33 | -53.2 | -4.12 | 451→464 |
| GresSuffurd | dodger | 114.0→112.3 | -1.7 | 1.33→2.67 | +1.33 | -94.9 | -8.82 | 417→461 |
| CassiusClay | dodger | 85.3→95.2 | +9.9 | 0.67→2.33 | +1.67 | -103.7 | -8.41 | 388→366 |
| RetroGirl | pattern | 181.7→162.2 | -19.4 | 2.00→2.67 | +0.67 | +2.2 | -4.64 | 391→442 |
| TripHammer | pattern | 55.8→55.0 | -0.8 | 0.00→0.67 | +0.67 | -50.5 | -5.25 | 469→485 |
| Coriantumr | pattern | 100.9→87.7 | -13.2 | 1.67→1.33 | -0.33 | -3.8 | -0.71 | 444→460 |
| WallAvoider | wallfollower | 150.0→138.6 | -11.5 | 2.00→3.00 | +1.00 | -64.0 | -4.78 | 317→383 |
| HawkOnFire | cornercamper | 115.1→122.4 | +7.3 | 1.67→2.67 | +1.00 | -124.2 | -11.23 | 419→514 |
| SpinBot | spinner | 302.0→261.8 | -40.2 | 3.00→3.00 | +0.00 | -32.0 | -17.27 | 316→411 |
| DiamondStealer | rammer | 140.1→149.2 | +9.2 | 1.00→1.33 | +0.33 | -20.1 | -2.23 | 236→266 |
| BlitzBat | brawler | 74.5→50.1 | -24.5 | 2.00→2.33 | +0.33 | -103.2 | -8.93 | 420→519 |
| YersiniaPestis | aggressive | 65.6→56.8 | -8.8 | 0.33→0.33 | +0.00 | -35.7 | -3.96 | 377→395 |
| Ascendant | aggressive | 62.8→82.3 | +19.6 | 0.00→0.67 | +0.67 | -30.2 | -8.29 | 362→374 |
#### `tilt_600` — tilt ON, target 600 (farther than the emergent 456) (paired on 15 opponents)
| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| DrussGT | dodger | 120.8→110.5 | -10.2 | 0.67→1.00 | +0.33 | -40.5 | -3.48 | 445→554 |
| Diamond | dodger | 55.9→63.9 | +8.0 | 0.00→0.00 | +0.00 | -97.2 | -11.33 | 457→574 |
| Dookious | dodger | 86.5→83.4 | -3.0 | 1.33→1.00 | -0.33 | +6.7 | +0.01 | 451→498 |
| GresSuffurd | dodger | 114.0→114.8 | +0.8 | 1.33→3.00 | +1.67 | -102.3 | -8.16 | 417→475 |
| CassiusClay | dodger | 85.3→62.1 | -23.2 | 0.67→0.67 | +0.00 | -46.4 | -5.24 | 388→449 |
| RetroGirl | pattern | 181.7→124.3 | -57.4 | 2.00→2.00 | +0.00 | +25.0 | -5.38 | 391→467 |
| TripHammer | pattern | 55.8→52.6 | -3.2 | 0.00→1.33 | +1.33 | -54.1 | -7.01 | 469→552 |
| Coriantumr | pattern | 100.9→63.1 | -37.8 | 1.67→1.33 | -0.33 | +4.3 | -1.08 | 444→557 |
| WallAvoider | wallfollower | 150.0→147.1 | -2.9 | 2.00→1.67 | -0.33 | -42.1 | -2.06 | 317→373 |
| HawkOnFire | cornercamper | 115.1→95.0 | -20.1 | 1.67→2.00 | +0.33 | -98.9 | -9.05 | 419→572 |
| SpinBot | spinner | 302.0→250.5 | -51.5 | 3.00→3.00 | +0.00 | -32.0 | -18.70 | 316→444 |
| DiamondStealer | rammer | 140.1→159.9 | +19.9 | 1.00→1.67 | +0.67 | -53.3 | -1.92 | 236→274 |
| BlitzBat | brawler | 74.5→41.0 | -33.5 | 2.00→3.00 | +1.00 | -91.0 | -11.09 | 420→585 |
| YersiniaPestis | aggressive | 65.6→75.4 | +9.8 | 0.33→0.67 | +0.33 | -45.0 | -7.18 | 377→408 |
| Ascendant | aggressive | 62.8→90.6 | +27.8 | 0.00→1.33 | +1.33 | -85.0 | -12.15 | 362→387 |
#### `tilt_250` — tilt ON, target 250 (much nearer than the emergent 456) (paired on 15 opponents)
| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| DrussGT | dodger | 120.8→108.7 | -12.1 | 0.67→1.00 | +0.33 | -15.7 | -2.25 | 445→470 |
| Diamond | dodger | 55.9→97.2 | +41.3 | 0.00→1.00 | +1.00 | -100.7 | -7.74 | 457→482 |
| Dookious | dodger | 86.5→105.0 | +18.6 | 1.33→2.00 | +0.67 | -39.4 | -2.97 | 451→460 |
| GresSuffurd | dodger | 114.0→145.2 | +31.2 | 1.33→2.67 | +1.33 | -89.4 | -6.57 | 417→409 |
| CassiusClay | dodger | 85.3→103.1 | +17.8 | 0.67→1.00 | +0.33 | -30.2 | -3.63 | 388→367 |
| RetroGirl | pattern | 181.7→142.8 | -38.8 | 2.00→2.33 | +0.33 | +36.3 | -4.18 | 391→436 |
| TripHammer | pattern | 55.8→53.4 | -2.3 | 0.00→0.33 | +0.33 | -20.3 | -2.44 | 469→467 |
| Coriantumr | pattern | 100.9→75.1 | -25.8 | 1.67→1.33 | -0.33 | +5.1 | -1.04 | 444→436 |
| WallAvoider | wallfollower | 150.0→157.0 | +7.0 | 2.00→1.33 | -0.67 | +42.9 | +1.38 | 317→297 |
| HawkOnFire | cornercamper | 115.1→130.6 | +15.5 | 1.67→3.00 | +1.33 | -123.6 | -10.75 | 419→454 |
| SpinBot | spinner | 302.0→258.0 | -44.0 | 3.00→3.00 | +0.00 | -32.0 | -18.91 | 316→432 |
| DiamondStealer | rammer | 140.1→143.7 | +3.6 | 1.00→1.00 | +0.00 | -12.7 | -1.10 | 236→260 |
| BlitzBat | brawler | 74.5→51.2 | -23.3 | 2.00→2.67 | +0.67 | -94.4 | -7.50 | 420→502 |
| YersiniaPestis | aggressive | 65.6→57.3 | -8.3 | 0.33→0.67 | +0.33 | -23.5 | -5.03 | 377→396 |
| Ascendant | aggressive | 62.8→51.1 | -11.7 | 0.00→0.00 | +0.00 | -17.3 | -4.93 | 362→364 |
### MEASURED: pooled dashboard (all valid runs, NOT the verdict)
| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `tfil` | 45 | 114.1 | 196.5 | 1.18 | 53/135 | 39.3% | 17.63% | 394 |
| `strafe_notilt` | 45 | 101.5 | 148.6 | 1.64 | 74/135 | 54.8% | 12.05% | 459 |
| `strafe_325` | 45 | 111.8 | 144.2 | 1.76 | 79/135 | 58.5% | 12.52% | 434 |
| `tilt_600` | 45 | 102.3 | 146.4 | 1.58 | 71/135 | 52.6% | 11.67% | 478 |
| `tilt_250` | 45 | 112.0 | 162.2 | 1.56 | 70/135 | 51.9% | 13.71% | 415 |
### MEASURED: cross-opponent aggregation (the verdict layer)
Deltas are per-opponent (arm − reference). `spread` is the SD of those deltas ACROSS opponents; `SE` = spread/√n; `95% CI` = mean ± t·SE. Sign test = how many opponents the arm wins (ties dropped), exact binomial; sign-flip = permutation test on the mean of the deltas.
| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |
|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|
| `strafe_notilt` | damage | -12.52 | 21.86 | 5.64 | [-24.62, -0.41] | 7/15 | 1 | 0.04456 (exact 2^15) | 0.1323 | 15.81 |
| `strafe_notilt` | wins | +0.47 | 0.45 | 0.12 | [+0.22, +0.72] | 10/11 | 0.01172 | 0.003906 (exact 2^15) | 0.007526 | 0.33 |
| `strafe_notilt` | damage_taken | -47.88 | 24.70 | 6.38 | [-61.55, -34.20] | 0/15 | 6.104e-05 | 6.104e-05 (exact 2^15) | 0.0007265 | 17.86 |
| `strafe_notilt` | hit_rate | -6.87 | 5.51 | 1.42 | [-9.93, -3.82] | 1/15 | 0.0009766 | 0.0001221 (exact 2^15) | 0.0008919 | 3.99 |
| `strafe_notilt` | dist | +65.39 | 46.33 | 11.96 | [+39.73, +91.05] | 15/15 | 6.104e-05 | 6.104e-05 (exact 2^15) | 0.0007265 | 33.52 |
| `strafe_325` | damage | -2.28 | 17.82 | 4.60 | [-12.15, +7.59] | 7/15 | 1 | 0.6319 (exact 2^15) | 0.712 | 12.89 |
| `strafe_325` | wins | +0.58 | 0.56 | 0.14 | [+0.27, +0.89] | 11/12 | 0.006348 | 0.002441 (exact 2^15) | 0.00525 | 0.40 |
| `strafe_325` | damage_taken | -52.32 | 38.49 | 9.94 | [-73.64, -31.01] | 1/15 | 0.0009766 | 0.0001221 (exact 2^15) | 0.0008919 | 27.84 |
| `strafe_325` | hit_rate | -6.45 | 4.20 | 1.08 | [-8.78, -4.13] | 0/15 | 6.104e-05 | 6.104e-05 (exact 2^15) | 0.0007265 | 3.04 |
| `strafe_325` | dist | +40.15 | 35.18 | 9.08 | [+20.67, +59.64] | 14/15 | 0.0009766 | 0.0004272 (exact 2^15) | 0.002377 | 25.45 |
| `tilt_600` | damage | -11.78 | 25.11 | 6.48 | [-25.68, +2.13] | 5/15 | 0.3018 | 0.09137 (exact 2^15) | 0.1055 | 18.16 |
| `tilt_600` | wins | +0.40 | 0.66 | 0.17 | [+0.04, +0.76] | 8/11 | 0.2266 | 0.04883 (exact 2^15) | 0.04491 | 0.48 |
| `tilt_600` | damage_taken | -50.12 | 40.17 | 10.37 | [-72.36, -27.87] | 3/15 | 0.03516 | 0.0005493 (exact 2^15) | 0.002377 | 29.06 |
| `tilt_600` | hit_rate | -6.92 | 5.05 | 1.30 | [-9.72, -4.12] | 1/15 | 0.0009766 | 0.0001221 (exact 2^15) | 0.0008919 | 3.65 |
| `tilt_600` | dist | +83.94 | 44.45 | 11.48 | [+59.32, +108.55] | 15/15 | 6.104e-05 | 6.104e-05 (exact 2^15) | 0.0007265 | 32.15 |
| `tilt_250` | damage | -2.09 | 24.78 | 6.40 | [-15.82, +11.63] | 7/15 | 1 | 0.7453 (exact 2^15) | 0.7983 | 17.92 |
| `tilt_250` | wins | +0.38 | 0.56 | 0.14 | [+0.07, +0.69] | 10/12 | 0.03857 | 0.03125 (exact 2^15) | 0.05415 | 0.41 |
| `tilt_250` | damage_taken | -34.32 | 48.54 | 12.53 | [-61.21, -7.44] | 3/15 | 0.03516 | 0.01593 (exact 2^15) | 0.02877 | 35.11 |
| `tilt_250` | hit_rate | -5.18 | 4.89 | 1.26 | [-7.89, -2.47] | 1/15 | 0.0009766 | 0.0002441 (exact 2^15) | 0.001332 | 3.54 |
| `tilt_250` | dist | +21.42 | 37.60 | 9.71 | [+0.60, +42.24] | 10/15 | 0.3018 | 0.03253 (exact 2^15) | 0.04377 | 27.20 |
#### By inferred style (explanation only, never the verdict)
| arm | style | n | mean Δdmg | mean Δwins | mean Δhit rate (pp) |
|---|---|---:|---:|---:|---:|
| `strafe_notilt` | aggressive | 2 | -5.5 | +1.00 | -9.28 |
| `strafe_notilt` | brawler | 1 | -31.2 | +0.33 | -9.35 |
| `strafe_notilt` | cornercamper | 1 | -5.3 | +1.00 | -8.87 |
| `strafe_notilt` | dodger | 5 | -2.2 | +0.40 | -5.25 |
| `strafe_notilt` | pattern | 3 | -27.0 | +0.33 | -4.46 |
| `strafe_notilt` | rammer | 1 | -22.2 | +0.00 | -3.78 |
| `strafe_notilt` | spinner | 1 | -42.5 | +0.00 | -23.48 |
| `strafe_notilt` | wallfollower | 1 | +16.4 | +0.67 | +0.54 |
| `strafe_325` | aggressive | 2 | +5.4 | +0.33 | -6.12 |
| `strafe_325` | brawler | 1 | -24.5 | +0.33 | -8.93 |
| `strafe_325` | cornercamper | 1 | +7.3 | +1.00 | -11.23 |
| `strafe_325` | dodger | 5 | +9.6 | +0.87 | -5.91 |
| `strafe_325` | pattern | 3 | -11.1 | +0.33 | -3.53 |
| `strafe_325` | rammer | 1 | +9.2 | +0.33 | -2.23 |
| `strafe_325` | spinner | 1 | -40.2 | +0.00 | -17.27 |
| `strafe_325` | wallfollower | 1 | -11.5 | +1.00 | -4.78 |
| `tilt_600` | aggressive | 2 | +18.8 | +0.83 | -9.67 |
| `tilt_600` | brawler | 1 | -33.5 | +1.00 | -11.09 |
| `tilt_600` | cornercamper | 1 | -20.1 | +0.33 | -9.05 |
| `tilt_600` | dodger | 5 | -5.5 | +0.33 | -5.64 |
| `tilt_600` | pattern | 3 | -32.8 | +0.33 | -4.49 |
| `tilt_600` | rammer | 1 | +19.9 | +0.67 | -1.92 |
| `tilt_600` | spinner | 1 | -51.5 | +0.00 | -18.70 |
| `tilt_600` | wallfollower | 1 | -2.9 | -0.33 | -2.06 |
| `tilt_250` | aggressive | 2 | -10.0 | +0.17 | -4.98 |
| `tilt_250` | brawler | 1 | -23.3 | +0.67 | -7.50 |
| `tilt_250` | cornercamper | 1 | +15.5 | +1.33 | -10.75 |
| `tilt_250` | dodger | 5 | +19.4 | +0.73 | -4.63 |
| `tilt_250` | pattern | 3 | -22.3 | +0.11 | -2.55 |
| `tilt_250` | rammer | 1 | +3.6 | +0.00 | -1.10 |
| `tilt_250` | spinner | 1 | -44.0 | +0.00 | -18.91 |
| `tilt_250` | wallfollower | 1 | +7.0 | -0.67 | +1.38 |
#### The pre-registered verdict table, as printed by the analyzer
PRIMARY metrics are dmg/run and wins/run; hit rate is never the verdict. The pre-registered rule says an arm is BETTER when one primary metric is UP at sign-test p<0.05 `while the other does not go down`. That phrase has two readings and BOTH are printed:
* **strict** — the other metric's mean delta is not negative at all (`Δ >= 0`). Nothing can be BETTER while it costs *any* mean damage.
* **substantive** — the other metric's delta is not *detectably* down: the sign test is not significant **and** the delta is smaller than that metric's MDE (the pre-registered rule 3 says an effect under the MDE is not detectable, so it cannot count as a loss).
| rank | arm | Δwins/run | Δdmg/run | sign test wins | sign test dmg | verdict (strict) | verdict (substantive) |
|---:|---|---:|---:|---|---|---|---|
| 1 | `strafe_325` | +0.58 | -2.3 | 11/12 p=0.006348 | 7/15 p=1 | **not distinguishable** | **BETTER** |
| 2 | `strafe_notilt` | +0.47 | -12.5 | 10/11 p=0.01172 | 7/15 p=1 | **not distinguishable** | **BETTER** |
| 3 | `tilt_600` | +0.40 | -11.8 | 8/11 p=0.2266 | 5/15 p=0.3018 | **not distinguishable** | **not distinguishable** |
| 4 | `tilt_250` | +0.38 | -2.1 | 10/12 p=0.03857 | 7/15 p=1 | **not distinguishable** | **BETTER** |
Reference `tfil`: 114.1 dmg/run, 1.18 wins/run, 17.63% incoming, 394 px.
Highest wins delta: `strafe_325` (+0.58 wins/run, -2.3 dmg/run) — strict: **not distinguishable**, substantive: **BETTER**.
---
## 5. What to try next (rewritten AFTER Batches 1–2 — these are recommendations, not results)
Ranked by value per battle, given what the two batches measured:
1. **The engine is the lever; the range knob is not.** Both batches put the
strafe arms 12–19 pp above `tfil` on round-win rate while three different
range targets (none/250/600, achieved 415–478 px) made no separable
difference. So the next batch should attack the **strafe picker itself**, not
the range: `TR_STRAFE_DWELL_MIN/MAX` (reversal frequency), `TR_STRAFE_BAND` +
`TR_STRAFE_SPREAD` (how far the picker hedges), `TR_STRAFE_REACH` (line
length), `TR_STRAFE_WALL_BIAS`, `TR_STRAFE_WALL_MARGIN`. 3–4 arms, same
panel, **one knob family per batch**, and look for a plateau, not a peak.
2. **The verdict metric for movement is round wins; the mechanism metric is
incoming hit rate.** The winner took ~1/3 fewer hits at the same damage
output, and round wins in this harness are survival wins. So screen
*mechanism* ideas on incoming hit rate (±1 pp is detectable here: MDE 1.3–4.0
pp) and only then spend a full panel batch confirming the win effect.
3. **Do not chase damage.** The one arm that gained damage (`ring`, +31/run,
p = 0.007) won *fewer* nominal rounds and took +25 damage/run. A movement arm
that raises damage but lowers survival is a loss in disguise — the mirror of
the six inverted hit-rate verdicts this project has already paid for.
4. **`strafe_notilt` is the recommendation to ship-test**, if a shipping
decision is ever taken: it has the same win effect as the range-steered
config without an extra tuning surface. Shipping is a separate decision —
this campaign does not touch a shipped default.
5. **Then the gun** (the owner's next stage, per the mandate): same harness, same
panel or a gun-specific one, same paired-with-sign-test statistics. Two facts
for the gun job: (a) round wins here are survival wins, so the gun's job is
to *kill*, not merely to out-damage; (b) the panel is 15 opponents wide and
its strong dodgers (Diamond 39.5, CassiusClay 73.4, TripHammer 59.5 dmg/run
for `tfil`) are exactly the ones a DrussGT-only gun claim will fail against.
6. **Melee is a different game** (j116's finding): it needs its own panel and its
own ledger section; the 1v1 panel's verdicts do not transfer.
## 6. What would make us stop
* **The movement stage has already produced its first winner** (`TR_MOVEMENT=strafe`),
and by rule 2 with the substantive reading it beats the shipped default with a
margin that survives the between-opponent spread, replicated in two
independent sessions. A later job may therefore either (a) keep hunting
*within* the strafe picker (item 1 above) and stop as soon as two consecutive
batches fail to improve on it beyond the MDE, or (b) declare it the movement
answer and move to the gun. **Both are successful outcomes.**
* **Stop the movement stage entirely** once a batch's best arm cannot beat
`strafe` beyond the MDE, or when a movement arm's win gain is bought with a
detectable damage or survival loss. At that point *"this is the measured
optimum of this design space"* is the conclusion, not a failure.
* **Stop a single batch early** only for a contract violation (arena not free,
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
## 6. How to run a batch (exact commands)
## 7. How to run a batch (exact commands)
```sh
# 1. wait for the arena (this job may not be the only one fighting)
@@ -451,9 +750,28 @@ tools/ab/tournament_run.sh \
--reference tfil \
--outdir /tmp/ab/j118_b1
# Batch 2 (the range axis on the winning engine) was the same command with
# --arms tools/ab/arms_movement_b2.txt --outdir /tmp/ab/j118_b2
# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil
```
`--reference` may be ANY arm of the session: re-analyzing `/tmp/ab/j118_b1
--reference strafe_325` is a free pairwise comparison with no battles (it is how
the "the two strafe configs are not separable" claim was checked: Δwins +0.04,
p = 0.75, MDE 0.33).
## 8. Session log (outdirs are in `/tmp` and are NOT committed)
| session | commit | battles | arms | verdict |
|---|---|---:|---|---|
| `/tmp/ab/j118_b1` | `1984a78` | 225 (0 invalid) | tfil, strafe_notilt, strafe_325, ring, ring_notemp | strafe_notilt beats tfil on wins (+0.38, 9/9, p=0.0039) |
| `/tmp/ab/j118_b2` | `8efa627` | 225 (0 invalid) | tfil, strafe_notilt, strafe_325, tilt_600, tilt_250 | all four strafe arms beat tfil on wins (+0.38…+0.58); the range target decides nothing |
Both sessions can be re-analyzed offline at any time (no arena needed) as long as
`/tmp/ab/j118_b*` still exists; after a reboot only this ledger's tables remain,
which is why every number is inlined above.
The runner writes `<outdir>/session.json` (commit sha, binary sha256, arms,
panel) so any later job can re-analyze an old session offline, with no arena.