movement ledger: the post-hoc structure of the strafe win (60 points, both batches)

56/60 opponent-level win deltas are non-negative and no opponent family
regresses reproducibly (the one WallAvoider loss reverses in Batch 2).
corr(dwins, d_hit_rate) = -0.09, corr(dwins, d_damage) = +0.38: the aggregate
win is a survival effect, but a per-opponent hit-rate gain does not predict a
per-opponent win gain - so hit rate stays an explanation, never a proxy.
This commit is contained in:
2026-09-26 01:33:24 +02:00
parent 44e2d191e3
commit c07e6d9599
+19
View File
@@ -700,6 +700,25 @@ Highest wins delta: `strafe_325` (+0.58 wins/run, -2.3 dmg/run) — strict: **no
## 5. What to try next (rewritten AFTER Batches 1–2 — these are recommendations, not results) ## 5. What to try next (rewritten AFTER Batches 1–2 — these are recommendations, not results)
**Post-hoc structure of the win (60 opponent×arm×session points from Batches 1–2,
MEASURED).** The strafe win is not uniformly distributed and its size is **not**
predicted by the size of the hit-rate improvement across opponents:
* 56 of 60 points have a **non-negative** win delta; the 4 negatives are −0.33
(`strafe_notilt` vs Coriantumr B2, `strafe_325` vs Coriantumr B2), −0.33
(`strafe_325` vs DrussGT B1) and one WallAvoider B1 point (−1.00) that
**reverses to +1.00 in Batch 2** — so no opponent family shows a reproducible
regression at this n.
* corr(Δwins, Δincoming-hit-rate) = **−0.09** across those points;
corr(Δwins, Δdamage/run) = **+0.38**. Buckets: points whose hit rate improved
by ≥5 pp average **+0.53** wins/run (n=36); the 4 points with <2 pp of
hit-rate improvement average **−0.08**.
* Reading: the *aggregate* win is a survival effect (fewer hits taken, ~50 less
damage taken per run), but "this arm dodges better by X pp here" does **not**
mean "it wins more rounds here". Do not use hit-rate improvement as a proxy
for a win at the level of a single opponent — that is the sixth-verdict trap
this project keeps paying for.
Ranked by value per battle, given what the two batches measured: Ranked by value per battle, given what the two batches measured:
1. **The engine is the lever; the range knob is not.** Both batches put the 1. **The engine is the lever; the range knob is not.** Both batches put the