BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE

- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
  provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
  (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
  7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
  (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
  unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
  a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
  standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
  detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
This commit is contained in:
2026-09-24 23:00:20 +02:00
parent f91e121965
commit d93ce444c0
3 changed files with 377 additions and 38 deletions
+14 -3
View File
@@ -36,9 +36,20 @@ python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
```
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits
taken/run, the **per-run** values, an exact two-sided permutation test on
per-run damage and wins vs the reference arm (default: first arm), a
round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line,
taken/run, the **per-run** values, and for **every pair of arms**:
* a two-sided **permutation test** on per-run damage and wins. Full enumeration
when `C(n, na) <= 20,000,000` (7v7 -> C(14,7)=3432, always exact); otherwise a
**Monte-Carlo** permutation test with `MC_DRAWS = 1,000,000` fixed draws and the
fixed seed `MC_SEED = 0x5eed5eed`, reported with its Monte-Carlo standard error
(`p = (cnt+1)/(B+1)`, `se = sqrt(p(1-p)/(B+1))`). Each row says which method
produced its p-value;
* a tie-corrected, continuity-corrected **Mann-Whitney U** cross-check;
plus the **minimum detectable effect** for the reference arm's n and observed
per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test
(labelled anti-conservative), a liveness OK/FAIL line, a `[bb]` applied-shift
check (needs `TR_BITBRAIN_LOG=1`; a zero-shift placebo emits no `[bb]` lines),
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and