diff --git a/docs/movement_campaign.md b/docs/movement_campaign.md index 0305b87..41bfb1f 100644 --- a/docs/movement_campaign.md +++ b/docs/movement_campaign.md @@ -700,6 +700,25 @@ Highest wins delta: `strafe_325` (+0.58 wins/run, -2.3 dmg/run) — strict: **no ## 5. What to try next (rewritten AFTER Batches 1–2 — these are recommendations, not results) +**Post-hoc structure of the win (60 opponent×arm×session points from Batches 1–2, +MEASURED).** The strafe win is not uniformly distributed and its size is **not** +predicted by the size of the hit-rate improvement across opponents: + +* 56 of 60 points have a **non-negative** win delta; the 4 negatives are −0.33 + (`strafe_notilt` vs Coriantumr B2, `strafe_325` vs Coriantumr B2), −0.33 + (`strafe_325` vs DrussGT B1) and one WallAvoider B1 point (−1.00) that + **reverses to +1.00 in Batch 2** — so no opponent family shows a reproducible + regression at this n. +* corr(Δwins, Δincoming-hit-rate) = **−0.09** across those points; + corr(Δwins, Δdamage/run) = **+0.38**. Buckets: points whose hit rate improved + by ≥5 pp average **+0.53** wins/run (n=36); the 4 points with <2 pp of + hit-rate improvement average **−0.08**. +* Reading: the *aggregate* win is a survival effect (fewer hits taken, ~50 less + damage taken per run), but "this arm dodges better by X pp here" does **not** + mean "it wins more rounds here". Do not use hit-rate improvement as a proxy + for a win at the level of a single opponent — that is the sixth-verdict trap + this project keeps paying for. + Ranked by value per battle, given what the two batches measured: 1. **The engine is the lever; the range knob is not.** Both batches put the