BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the 7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS unreachable is used instead. - tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its standard error, a tie-corrected Mann-Whitney U cross-check, a minimum detectable effect line, all-pairs comparisons, and a [bb] shift check. - tools/ab/README.md: document the new analyzer output.
This commit is contained in:
@@ -0,0 +1,185 @@
|
|||||||
|
# BitBrain verdict — CLEAN NEGATIVE against real DrussGT
|
||||||
|
|
||||||
|
**Direct answer (MEASURED).** The BitBrain decay-memory gun does **not** beat the
|
||||||
|
shipped Pattern rack, and it does **not** beat a provably-zero placebo. At
|
||||||
|
30 runs/arm (210 rounds/arm) `bb_decay` is **+3.3 damage/run** over `control`
|
||||||
|
(permutation p = 0.709, Mann-Whitney p = 0.600) and round wins are **dead level,
|
||||||
|
97/210 vs 97/210** (p = 1.000). Every one of the twelve pairwise comparisons is
|
||||||
|
a null (all p >= 0.29). The 7-run "shape" that motivated this test — 294 vs 279
|
||||||
|
damage, 26 vs 22 wins — **did not replicate**: at 30 runs the same difference is
|
||||||
|
+3.3 damage and +0 wins. The one mechanism the offline gate test's diagnosis
|
||||||
|
predicted (a bounded/decaying SBC memory) buys nothing live. **The BitBrain
|
||||||
|
thread closes here**, with a reason instead of an open question.
|
||||||
|
|
||||||
|
## What was run (MEASURED)
|
||||||
|
|
||||||
|
* Frozen from HEAD `795a0e5` via `git archive HEAD`; `ModularBot` binary sha256
|
||||||
|
`cab6083018672168642fd5a3e62259a260a43ff3ebfc0418dd3cbe15ab8e2005`.
|
||||||
|
* Real DrussGT through the `robocode_shim` bridge, 4 arms x 30 runs x 7 rounds
|
||||||
|
(120 battles, 0 failed), `--conc 7`.
|
||||||
|
* Arms (one env knob each, all four verified live in the boot report):
|
||||||
|
* `control` — no env; shipped `onlyPattern` rack.
|
||||||
|
* `bb_decay` — `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay` (the candidate).
|
||||||
|
* `bb_zero` — as `bb_decay` + `TR_BITBRAIN_RANGE=0`; **this is NOT a zero-shift placebo** (see below), so it is reported as a small-dose treatment.
|
||||||
|
* `bb_placebo` — as `bb_decay` + `TR_BITBRAIN_MIN_OBS=100000000`; the **valid placebo** (the whole BitBrain path runs, but the readout branch is never reached, so the applied shift is provably exactly 0).
|
||||||
|
* `TR_BITBRAIN_LOG=1` on the three BitBrain arms so the applied shift is visible and the placebo can be checked.
|
||||||
|
|
||||||
|
## The clear answer to the rubric
|
||||||
|
|
||||||
|
* `bb_decay` beats `control` on damage/run **AND** round wins with p < 0.05? **No**
|
||||||
|
(+3.3 dmg, p = 0.709; +0 wins, p = 1.000).
|
||||||
|
* `bb_decay` ~= `bb_placebo`? **Yes** (dmg p = 0.897, wins p = 0.918). So any
|
||||||
|
difference the BitBrain path makes is **plumbing/noise, not learning** — and
|
||||||
|
here it makes no difference at all.
|
||||||
|
* `bb_zero` ~= `control` **and** `bb_decay` > both? **Not the situation**: `bb_zero`
|
||||||
|
is not a zero-shift placebo, and `bb_decay` beats nothing.
|
||||||
|
* Nothing separates -> **CLEAN NEGATIVE.** Not spun: this is a successful outcome
|
||||||
|
that closes the thread.
|
||||||
|
|
||||||
|
## The placebo: `TR_BITBRAIN_RANGE=0` is clamped, so it is NOT a zero-shift arm (MEASURED)
|
||||||
|
|
||||||
|
`common_libs/guns/bitbrain_gun.nim:207`:
|
||||||
|
|
||||||
|
```nim
|
||||||
|
result.maxDeg = clamp(envFloatBB(BB_RANGE_ENV, BB_RANGE_DEF), 1.0, 180.0)
|
||||||
|
```
|
||||||
|
|
||||||
|
`TR_BITBRAIN_RANGE=0` is therefore clamped to `1.0` degree, and the class centres
|
||||||
|
become +/-0.97 deg, not 0. Confirmed live: the boot report shows
|
||||||
|
`[env] TR_BITBRAIN_RANGE = 1.0 (source: env)` and the `[bb]` log emits
|
||||||
|
`shift=-1.0deg` / `shift=+1.0deg` (edge classes 0 and 31). So `bb_zero` applies a
|
||||||
|
systematic ~1 deg correction and is **not** the intended placebo. Per the task's
|
||||||
|
fallback this arm was kept only as a small-dose treatment; the true placebo is
|
||||||
|
`bb_placebo`, which emits **0 `[bb]` lines in 30/30 runs** — a zero applied shift
|
||||||
|
by construction (`bbLog` is reached only inside the same `trained >= minObs`
|
||||||
|
branch that computes the shift, so a never-firing readout is silent). Nominally
|
||||||
|
`bb_zero` trends *worse* (-5.5 dmg/run vs control, p = 0.51; -7.7 vs the true
|
||||||
|
placebo, p = 0.35), consistent with a tiny systematic mis-aim rather than a
|
||||||
|
neutral probe — and still inside the noise.
|
||||||
|
|
||||||
|
## Raw analyzer output (verbatim, `tools/ab/ab_analyze.py`)
|
||||||
|
|
||||||
|
```
|
||||||
|
# session /tmp/ab/bbverdict
|
||||||
|
# commit=795a0e59febb4d9b9722130ad47cd2dd8f72a315 binary_sha256=cab6083018672168642fd5a3e62259a260a43ff3ebfc0418dd3cbe15ab8e2005 rounds=7 runs=30 conc=7 ts=2026-09-24T22:37:38+02:00
|
||||||
|
|
||||||
|
ARM SUMMARY
|
||||||
|
arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
control 30 280 206 97/210 46.2 796 92.5
|
||||||
|
bb_decay 30 284 206 97/210 46.2 797 93.6
|
||||||
|
bb_zero 30 275 209 93/210 44.3 781 91.8
|
||||||
|
bb_placebo 30 283 209 95/210 45.2 786 92.2
|
||||||
|
|
||||||
|
PER-RUN (never just the mean)
|
||||||
|
control dmg: r1=293 r2=296 r3=352 r4=263 r5=282 r6=292 r7=246 r8=271 r9=246 r10=261 r11=263 r12=175 r13=302 r14=288 r15=280 r16=333 r17=244 r18=293 r19=317 r20=249 r21=283 r22=296 r23=278 r24=281 r25=272 r26=292 r27=238 r28=332 r29=295 r30=299
|
||||||
|
wins: r1=3/7 r2=5/7 r3=4/7 r4=3/7 r5=5/7 r6=3/7 r7=2/7 r8=2/7 r9=3/7 r10=3/7 r11=3/7 r12=0/7 r13=4/7 r14=4/7 r15=3/7 r16=2/7 r17=4/7 r18=3/7 r19=3/7 r20=3/7 r21=2/7 r22=2/7 r23=4/7 r24=5/7 r25=4/7 r26=3/7 r27=3/7 r28=6/7 r29=3/7 r30=3/7
|
||||||
|
bb_decay dmg: r1=284 r2=247 r3=273 r4=300 r5=241 r6=248 r7=276 r8=299 r9=227 r10=279 r11=305 r12=349 r13=302 r14=318 r15=279 r16=244 r17=307 r18=293 r19=328 r20=212 r21=296 r22=296 r23=362 r24=258 r25=245 r26=298 r27=308 r28=294 r29=285 r30=259
|
||||||
|
wins: r1=4/7 r2=2/7 r3=5/7 r4=3/7 r5=3/7 r6=3/7 r7=3/7 r8=3/7 r9=2/7 r10=4/7 r11=2/7 r12=4/7 r13=3/7 r14=6/7 r15=3/7 r16=2/7 r17=5/7 r18=4/7 r19=4/7 r20=2/7 r21=2/7 r22=2/7 r23=6/7 r24=1/7 r25=3/7 r26=3/7 r27=6/7 r28=4/7 r29=1/7 r30=2/7
|
||||||
|
bb_zero dmg: r1=248 r2=295 r3=267 r4=310 r5=277 r6=313 r7=264 r8=276 r9=277 r10=236 r11=302 r12=262 r13=315 r14=251 r15=314 r16=297 r17=231 r18=266 r19=212 r20=224 r21=249 r22=235 r23=314 r24=314 r25=285 r26=262 r27=296 r28=318 r29=271 r30=265
|
||||||
|
wins: r1=1/7 r2=4/7 r3=3/7 r4=2/7 r5=3/7 r6=6/7 r7=3/7 r8=4/7 r9=3/7 r10=3/7 r11=3/7 r12=2/7 r13=4/7 r14=4/7 r15=4/7 r16=2/7 r17=1/7 r18=4/7 r19=3/7 r20=2/7 r21=3/7 r22=2/7 r23=5/7 r24=2/7 r25=4/7 r26=2/7 r27=4/7 r28=3/7 r29=3/7 r30=4/7
|
||||||
|
bb_placebo dmg: r1=279 r2=285 r3=328 r4=253 r5=275 r6=286 r7=348 r8=256 r9=258 r10=303 r11=278 r12=242 r13=291 r14=234 r15=341 r16=274 r17=240 r18=267 r19=312 r20=228 r21=302 r22=337 r23=320 r24=266 r25=308 r26=291 r27=255 r28=275 r29=308 r30=234
|
||||||
|
wins: r1=2/7 r2=3/7 r3=4/7 r4=1/7 r5=4/7 r6=2/7 r7=4/7 r8=2/7 r9=2/7 r10=4/7 r11=4/7 r12=2/7 r13=5/7 r14=2/7 r15=4/7 r16=3/7 r17=3/7 r18=3/7 r19=4/7 r20=2/7 r21=3/7 r22=4/7 r23=4/7 r24=4/7 r25=5/7 r26=2/7 r27=3/7 r28=3/7 r29=5/7 r30=2/7
|
||||||
|
|
||||||
|
PAIRWISE PERMUTATION TEST (per-run values) + MANN-WHITNEY CROSS-CHECK
|
||||||
|
permutation: exact when C(n,na) <= 20,000,000; otherwise Monte-Carlo 1,000,000 draws, seed=0x5eed5eed, p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))
|
||||||
|
metric A B diff(A-B) perm p method MC se MW p MW U
|
||||||
|
-------------------------------------------------------------------------------------------------
|
||||||
|
dmg/run control bb_decay -3.313 0.7091 MC/B=1,000,000 0.0005 0.5997 414.0
|
||||||
|
round wins control bb_decay +0.000 1.0000 MC/B=1,000,000 0.0000 0.7468 428.5
|
||||||
|
dmg/run control bb_zero +5.551 0.5086 MC/B=1,000,000 0.0005 0.5493 409.0
|
||||||
|
round wins control bb_zero +0.133 0.7369 MC/B=1,000,000 0.0004 0.6535 420.5
|
||||||
|
dmg/run control bb_placebo -2.175 0.8030 MC/B=1,000,000 0.0004 0.9117 442.0
|
||||||
|
round wins control bb_placebo +0.067 0.9091 MC/B=1,000,000 0.0003 0.8356 436.0
|
||||||
|
dmg/run bb_decay bb_zero +8.863 0.2947 MC/B=1,000,000 0.0005 0.4376 397.0
|
||||||
|
round wins bb_decay bb_zero +0.133 0.7597 MC/B=1,000,000 0.0004 0.9027 441.5
|
||||||
|
dmg/run bb_decay bb_placebo +1.138 0.8969 MC/B=1,000,000 0.0003 0.7845 431.0
|
||||||
|
round wins bb_decay bb_placebo +0.067 0.9178 MC/B=1,000,000 0.0003 0.9513 445.5
|
||||||
|
dmg/run bb_zero bb_placebo -7.725 0.3525 MC/B=1,000,000 0.0005 0.4733 401.0
|
||||||
|
round wins bb_zero bb_placebo -0.067 0.9072 MC/B=1,000,000 0.0003 0.7763 431.0
|
||||||
|
|
||||||
|
MINIMUM DETECTABLE EFFECT (two-sample, alpha=0.05 two-sided, 80% power; MDE = 2.8016*sd*sqrt(2/n))
|
||||||
|
metric n/arm sd(control) MDE(abs) MDE vs control mean
|
||||||
|
----------------------------------------------------------------
|
||||||
|
dmg/run 30 33.669 24.355 8.7% of 280.4
|
||||||
|
round wins 30 1.165 0.843 26.1% of 3.2
|
||||||
|
|
||||||
|
ROUND-LEVEL TEST (pooled rounds, Fisher exact) vs `control` — ANTI-CONSERVATIVE: rounds cluster within runs
|
||||||
|
arm ref wins arm wins p
|
||||||
|
----------------------------------------------
|
||||||
|
bb_decay 97/210 97/210 1.0000
|
||||||
|
bb_zero 97/210 93/210 0.7687
|
||||||
|
bb_placebo 97/210 95/210 0.9220
|
||||||
|
|
||||||
|
LIVENESS (arm env applied in the bot's own boot report)
|
||||||
|
control OK (30/30 runs: no arm env; report present)
|
||||||
|
bb_decay OK (30/30 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay TR_BITBRAIN_LOG=1 applied)
|
||||||
|
bb_zero OK (30/30 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay TR_BITBRAIN_RANGE=0 TR_BITBRAIN_LOG=1 applied)
|
||||||
|
bb_placebo OK (30/30 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay TR_BITBRAIN_MIN_OBS=100000000 TR_BITBRAIN_LOG=1 applied)
|
||||||
|
|
||||||
|
[bb] APPLIED-SHIFT CHECK (from bot stdout; needs TR_BITBRAIN_LOG=1). A provably-zero placebo emits ZERO [bb] lines.
|
||||||
|
arm runs w/log lines min max zeros
|
||||||
|
control 0/30 0 - - -
|
||||||
|
bb_decay 30/30 905 -38.80 +36.20 0
|
||||||
|
bb_zero 30/30 408 -1.00 +1.00 9
|
||||||
|
bb_placebo 0/30 0 - - -
|
||||||
|
|
||||||
|
ROUND-WIN ATTRIBUTION (events primary; score tie-break for mutual-kill / timeout rounds)
|
||||||
|
control wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 205/205 (5 tie-broken)
|
||||||
|
bb_decay wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 209/209 (1 tie-broken)
|
||||||
|
bb_zero wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 209/209 (1 tie-broken)
|
||||||
|
bb_placebo wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 207/207 (3 tie-broken)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Minimum detectable effect (MEASURED) — what this test CAN and CANNOT see
|
||||||
|
|
||||||
|
From the observed control-arm per-run SD (damage SD = 33.7; wins SD = 1.17) at
|
||||||
|
30 runs/arm, alpha = 0.05 two-sided, 80% power:
|
||||||
|
|
||||||
|
* **damage/run: MDE = 24.4** (8.7% of the 280 control mean).
|
||||||
|
* **round wins: MDE = 0.84 wins/run** (26.1% of the 3.2 mean).
|
||||||
|
|
||||||
|
So this test rules out a `bb_decay` damage gain of **>= 24 dmg/run** (and a win
|
||||||
|
gain of **>= 0.84/run**). A real effect of, say, +10-15 dmg/run — the size the
|
||||||
|
7-run result hinted at — **would not be detectable here**, and is **not** ruled
|
||||||
|
out by the null. That is the honest bound: "no effect >= 24 dmg/run", not "no
|
||||||
|
effect". The placebo isolates the mechanism more sharply: `bb_decay` vs
|
||||||
|
`bb_placebo` is +1.1 dmg/run (p = 0.897) and +0.07 wins (p = 0.918), so the
|
||||||
|
learning-specific effect is bounded by the same ~24 dmg/run and shows no sign.
|
||||||
|
|
||||||
|
## Why the earlier 7-run shape was misleading (MEASURED)
|
||||||
|
|
||||||
|
| metric | 7 runs/arm (commit `795a0e5` session) | 30 runs/arm (this session) |
|
||||||
|
|---|---|---|
|
||||||
|
| `control` dmg/run | 279 | 280 |
|
||||||
|
| `bb_decay` dmg/run | 294 (+15) | 284 (+3.3) |
|
||||||
|
| `control` round wins | 22/49 | 97/210 |
|
||||||
|
| `bb_decay` round wins | 26/49 (+4) | 97/210 (+0) |
|
||||||
|
|
||||||
|
The 7-run difference was inside the noise; it has shrunk to zero, not grown.
|
||||||
|
|
||||||
|
## MEASURED vs INFERRED
|
||||||
|
|
||||||
|
* **MEASURED:** the damage/win table, the pairwise permutation p-values
|
||||||
|
(Monte-Carlo, 1,000,000 draws, seed `0x5eed5eed`, reported with MC SE) and the
|
||||||
|
Mann-Whitney cross-check, the MDE from the observed SD, the liveness lines, the
|
||||||
|
`[bb]` shift check, and the `TR_BITBRAIN_RANGE` clamp (`= 1.0` in the live boot
|
||||||
|
report).
|
||||||
|
* **INFERRED:** that a true effect below ~24 dmg/run would need more runs or a
|
||||||
|
lower-variance opponent to resolve. The offline gate-test diagnosis (decay
|
||||||
|
memory should help) is **not** supported live; whether a *different* memory
|
||||||
|
regime would help is **not** tested here (only `decay` was).
|
||||||
|
|
||||||
|
## Reproduce
|
||||||
|
|
||||||
|
```sh
|
||||||
|
tools/ab/ab_run.sh --arms /tmp/ab/arms_bbverdict.txt --runs 30 \
|
||||||
|
--outdir /tmp/ab/bbverdict --conc 7
|
||||||
|
python3 tools/ab/ab_analyze.py /tmp/ab/bbverdict
|
||||||
|
```
|
||||||
|
|
||||||
|
Analyzer tooling improved in the same change: exact enumeration is kept when
|
||||||
|
`C(n, na) <= 20e6` (7v7), otherwise a seeded Monte-Carlo permutation test
|
||||||
|
(`MC_DRAWS = 1,000,000`, `MC_SEED = 0x5eed5eed`) with its standard error, plus a
|
||||||
|
tie-corrected Mann-Whitney U cross-check and the MDE line. See
|
||||||
|
`tools/ab/README.md`.
|
||||||
+14
-3
@@ -36,9 +36,20 @@ python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
|
|||||||
```
|
```
|
||||||
|
|
||||||
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits
|
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits
|
||||||
taken/run, the **per-run** values, an exact two-sided permutation test on
|
taken/run, the **per-run** values, and for **every pair of arms**:
|
||||||
per-run damage and wins vs the reference arm (default: first arm), a
|
|
||||||
round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line,
|
* a two-sided **permutation test** on per-run damage and wins. Full enumeration
|
||||||
|
when `C(n, na) <= 20,000,000` (7v7 -> C(14,7)=3432, always exact); otherwise a
|
||||||
|
**Monte-Carlo** permutation test with `MC_DRAWS = 1,000,000` fixed draws and the
|
||||||
|
fixed seed `MC_SEED = 0x5eed5eed`, reported with its Monte-Carlo standard error
|
||||||
|
(`p = (cnt+1)/(B+1)`, `se = sqrt(p(1-p)/(B+1))`). Each row says which method
|
||||||
|
produced its p-value;
|
||||||
|
* a tie-corrected, continuity-corrected **Mann-Whitney U** cross-check;
|
||||||
|
|
||||||
|
plus the **minimum detectable effect** for the reference arm's n and observed
|
||||||
|
per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test
|
||||||
|
(labelled anti-conservative), a liveness OK/FAIL line, a `[bb]` applied-shift
|
||||||
|
check (needs `TR_BITBRAIN_LOG=1`; a zero-shift placebo emits no `[bb]` lines),
|
||||||
and a round-win attribution cross-check.
|
and a round-win attribution cross-check.
|
||||||
|
|
||||||
Round wins come from the events sidecar (the bot that does not die wins) and
|
Round wins come from the events sidecar (the bot that does not die wins) and
|
||||||
|
|||||||
+178
-35
@@ -7,10 +7,17 @@ Standard library only, deterministic. For every arm it prints runs, damage/run,
|
|||||||
damage taken/run, ROUND WINS, shots/run, hits taken/run and the PER-RUN values
|
damage taken/run, ROUND WINS, shots/run, hits taken/run and the PER-RUN values
|
||||||
(wins cluster at 0/7 and single-run damage swings ~200, so the mean alone lies).
|
(wins cluster at 0/7 and single-run damage swings ~200, so the mean alone lies).
|
||||||
|
|
||||||
Tests:
|
Tests (all two-sided, on PER-RUN values unless stated):
|
||||||
* exact two-sided permutation test on PER-RUN values (damage/run and round
|
* permutation test on the difference of means. Full enumeration when
|
||||||
wins) vs a reference arm (default: the first arm). C(14,7)=3432 for 7v7 —
|
C(n, na) <= EXACT_CAP (7v7 -> C(14,7)=3432, always exact); otherwise a
|
||||||
enumerated exactly, never sampled, whenever the combination count is small.
|
Monte-Carlo permutation test with a fixed MC_DRAWS draws and the fixed
|
||||||
|
seed MC_SEED, reported with its Monte-Carlo standard error. Every p-value
|
||||||
|
says which method produced it (ncomb is C(60,30) ~ 1.2e17 at 30v30).
|
||||||
|
* Mann-Whitney U (rank-sum) cross-check, normal approximation with the
|
||||||
|
standard tie correction and a continuity correction.
|
||||||
|
* minimum detectable effect for n-per-arm (alpha=0.05 two-sided, 80% power)
|
||||||
|
derived from the observed pooled per-run SD, so a null result can be told
|
||||||
|
apart from an under-powered one.
|
||||||
* a round-level Fisher exact test on pooled rounds, clearly labelled
|
* a round-level Fisher exact test on pooled rounds, clearly labelled
|
||||||
anti-conservative (rounds cluster within runs).
|
anti-conservative (rounds cluster within runs).
|
||||||
|
|
||||||
@@ -34,12 +41,20 @@ import itertools
|
|||||||
import json
|
import json
|
||||||
import math
|
import math
|
||||||
import os
|
import os
|
||||||
|
import random
|
||||||
import re
|
import re
|
||||||
import sys
|
import sys
|
||||||
|
|
||||||
# beyond this many combinations we sample (deterministic seed) and say so;
|
# beyond this many combinations we sample (deterministic seed) and say so;
|
||||||
# 7v7 = C(14,7) = 3432 is always exact.
|
# 7v7 = C(14,7) = 3432 is always exact.
|
||||||
EXACT_CAP = 20_000_000
|
EXACT_CAP = 20_000_000
|
||||||
|
# Monte-Carlo permutation test: fixed draw count and fixed seed, so the p-value
|
||||||
|
# is reproducible byte-for-byte across runs of this analyzer.
|
||||||
|
MC_DRAWS = 1_000_000
|
||||||
|
MC_SEED = 0x5EED5EED
|
||||||
|
# z_{0.975} + z_{0.80}, the constant in the two-sample MDE at alpha=0.05 (two
|
||||||
|
# sided) and 80% power: MDE = 2.8016 * sd * sqrt(2/n).
|
||||||
|
Z_ALPHA_POWER = 1.959963984540054 + 0.8416212335729143
|
||||||
BOT_NAME = "ModularBot"
|
BOT_NAME = "ModularBot"
|
||||||
SUBJECT_NAME = "DrussGT" # the capture subject (our bot is the adversary)
|
SUBJECT_NAME = "DrussGT" # the capture subject (our bot is the adversary)
|
||||||
|
|
||||||
@@ -276,34 +291,98 @@ def _round_count(armdir, run):
|
|||||||
# ── statistics ───────────────────────────────────────────────────────────────
|
# ── statistics ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
def perm_test(xa, xb):
|
def perm_test(xa, xb):
|
||||||
"""Exact two-sided permutation test on the difference of means. Returns
|
"""Two-sided permutation test on the difference of means.
|
||||||
(obs, p, n_perm, exact_bool)."""
|
|
||||||
|
Exact full enumeration when C(n, na) <= EXACT_CAP (7v7 is 3432), otherwise
|
||||||
|
a Monte-Carlo permutation test with MC_DRAWS fixed draws and the fixed seed
|
||||||
|
MC_SEED. Returns None if either arm is empty, else a dict:
|
||||||
|
obs observed |mean(xa) - mean(xb)| (what the p-value tests)
|
||||||
|
signed mean(xa) - mean(xb) (for display; negative means xb is larger)
|
||||||
|
p two-sided p-value
|
||||||
|
draws number of permutations (enumerated or sampled)
|
||||||
|
method 'exact' | 'monte-carlo'
|
||||||
|
se Monte-Carlo standard error of p (0.0 for the exact test)
|
||||||
|
seed MC seed (None for the exact test)
|
||||||
|
The MC p-value uses the (cnt + 1) / (B + 1) estimator, so it is never 0.
|
||||||
|
"""
|
||||||
na, nb = len(xa), len(xb)
|
na, nb = len(xa), len(xb)
|
||||||
if na == 0 or nb == 0:
|
if na == 0 or nb == 0:
|
||||||
return None
|
return None
|
||||||
obs = abs(sum(xa) / na - sum(xb) / nb)
|
obs = abs(sum(xa) / na - sum(xb) / nb)
|
||||||
pooled = list(xa) + list(xb)
|
pooled = list(xa) + list(xb)
|
||||||
n = na + nb
|
n = na + nb
|
||||||
total_sum = sum(pooled)
|
total = sum(pooled)
|
||||||
ncomb = math.comb(n, na)
|
ncomb = math.comb(n, na)
|
||||||
if ncomb <= EXACT_CAP:
|
if ncomb <= EXACT_CAP:
|
||||||
cnt = 0
|
cnt = 0
|
||||||
for combo in itertools.combinations(range(n), na):
|
for combo in itertools.combinations(range(n), na):
|
||||||
sa = sum(pooled[i] for i in combo)
|
sa = sum(pooled[i] for i in combo)
|
||||||
if abs(sa / na - (total_sum - sa) / nb) >= obs - 1e-9:
|
if abs(sa / na - (total - sa) / nb) >= obs - 1e-9:
|
||||||
cnt += 1
|
cnt += 1
|
||||||
return obs, cnt / ncomb, ncomb, True
|
return {"obs": obs, "signed": sum(xa) / na - sum(xb) / nb,
|
||||||
# deterministic fallback for very large n (never hit at the default 7 runs)
|
"p": cnt / ncomb, "draws": ncomb,
|
||||||
import random
|
"method": "exact", "se": 0.0, "seed": None}
|
||||||
rng = random.Random(0xA1B2C3)
|
rng = random.Random(MC_SEED)
|
||||||
B = 200_000
|
sample = rng.sample
|
||||||
|
idx_range = range(n)
|
||||||
|
B = MC_DRAWS
|
||||||
cnt = 0
|
cnt = 0
|
||||||
for _ in range(B):
|
for _ in range(B):
|
||||||
idx = rng.sample(range(n), na)
|
sa = 0
|
||||||
sa = sum(pooled[i] for i in idx)
|
for i in sample(idx_range, na):
|
||||||
if abs(sa / na - (total_sum - sa) / nb) >= obs - 1e-9:
|
sa += pooled[i]
|
||||||
|
if abs(sa / na - (total - sa) / nb) >= obs - 1e-9:
|
||||||
cnt += 1
|
cnt += 1
|
||||||
return obs, (cnt + 1) / (B + 1), ncomb, False
|
p = (cnt + 1) / (B + 1)
|
||||||
|
se = math.sqrt(p * (1.0 - p) / (B + 1))
|
||||||
|
return {"obs": obs, "signed": sum(xa) / na - sum(xb) / nb,
|
||||||
|
"p": p, "draws": B, "method": "monte-carlo",
|
||||||
|
"se": se, "seed": MC_SEED}
|
||||||
|
|
||||||
|
|
||||||
|
def mannwhitney_p(xa, xb):
|
||||||
|
"""Two-sided Mann-Whitney U (rank-sum) test, normal approximation with the
|
||||||
|
standard tie correction and a continuity correction. Returns (p, U) or
|
||||||
|
None when either arm is empty. U is the smaller of U1, U2."""
|
||||||
|
na, nb = len(xa), len(xb)
|
||||||
|
n = na + nb
|
||||||
|
if na == 0 or nb == 0:
|
||||||
|
return None
|
||||||
|
vals = sorted([(v, 0) for v in xa] + [(v, 1) for v in xb])
|
||||||
|
rank = [0.0] * n
|
||||||
|
tie_term = 0.0
|
||||||
|
i = 0
|
||||||
|
while i < n:
|
||||||
|
j = i
|
||||||
|
while j + 1 < n and vals[j + 1][0] == vals[i][0]:
|
||||||
|
j += 1
|
||||||
|
t = j - i + 1
|
||||||
|
avg = (i + j) / 2.0 + 1.0
|
||||||
|
tie_term += t ** 3 - t
|
||||||
|
for k in range(i, j + 1):
|
||||||
|
rank[k] = avg
|
||||||
|
i = j + 1
|
||||||
|
r1 = sum(rank[k] for k in range(n) if vals[k][1] == 0)
|
||||||
|
u1 = r1 - na * (na + 1) / 2.0
|
||||||
|
mu = na * nb / 2.0
|
||||||
|
sigma2 = (na * nb / 12.0) * ((n + 1) - tie_term / (n * (n - 1)))
|
||||||
|
if sigma2 <= 0:
|
||||||
|
return (1.0, min(u1, na * nb - u1))
|
||||||
|
z = (abs(u1 - mu) - 0.5) / math.sqrt(sigma2)
|
||||||
|
if z < 0.0:
|
||||||
|
z = 0.0
|
||||||
|
p = math.erfc(z / math.sqrt(2.0))
|
||||||
|
return (p, min(u1, na * nb - u1))
|
||||||
|
|
||||||
|
|
||||||
|
def _var(x):
|
||||||
|
m = sum(x) / len(x)
|
||||||
|
return sum((v - m) ** 2 for v in x) / (len(x) - 1)
|
||||||
|
|
||||||
|
|
||||||
|
def min_detectable_effect(sd, n_per_arm):
|
||||||
|
"""Two-sample MDE at alpha=0.05 two-sided, 80% power, equal n."""
|
||||||
|
return Z_ALPHA_POWER * sd * math.sqrt(2.0 / n_per_arm)
|
||||||
|
|
||||||
|
|
||||||
def fisher_two_sided(a, b, c, d):
|
def fisher_two_sided(a, b, c, d):
|
||||||
@@ -363,6 +442,28 @@ def liveness(armdir, runs, env):
|
|||||||
return "OK", f"{len(runs)}/{len(runs)} runs: {vars_txt} applied"
|
return "OK", f"{len(runs)}/{len(runs)} runs: {vars_txt} applied"
|
||||||
|
|
||||||
|
|
||||||
|
# ── [bb] shift liveness ──────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
_BB_SHIFT_RE = re.compile(r"\[bb\].*?shift=([+-]?\d+(?:\.\d+)?)deg")
|
||||||
|
|
||||||
|
|
||||||
|
def bb_shifts(armdir, runs):
|
||||||
|
"""All `[bb] ... shift=Xdeg` values in an arm's bot stdout (TR_BITBRAIN_LOG=1).
|
||||||
|
Returns (values, runs_with_lines, runs). A placebo that never applies a
|
||||||
|
correction emits ZERO [bb] lines; `bbLog` is only reached inside the same
|
||||||
|
`trained >= minObs` branch that computes the applied shift."""
|
||||||
|
vals = []
|
||||||
|
runs_with = 0
|
||||||
|
for r in runs:
|
||||||
|
text = "".join(read_lines(os.path.join(
|
||||||
|
armdir, f"run{r}.bot.stdout.log")))
|
||||||
|
found = _BB_SHIFT_RE.findall(text)
|
||||||
|
if found:
|
||||||
|
runs_with += 1
|
||||||
|
vals.extend(float(x) for x in found)
|
||||||
|
return vals, runs_with, len(runs)
|
||||||
|
|
||||||
|
|
||||||
# ── report ───────────────────────────────────────────────────────────────────
|
# ── report ───────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
@@ -438,29 +539,55 @@ def main():
|
|||||||
print(f" {name:<14} dmg: {dmgs}")
|
print(f" {name:<14} dmg: {dmgs}")
|
||||||
print(f" {'':<14} wins: {wins}")
|
print(f" {'':<14} wins: {wins}")
|
||||||
|
|
||||||
# ── exact permutation tests ──────────────────────────────────────────────
|
# ── permutation tests (exact for 7v7, Monte-Carlo for 30v30) ─────────────
|
||||||
if reference not in data:
|
if reference not in data:
|
||||||
print(f"\nWARNING: reference arm '{reference}' not found; skipping tests")
|
print(f"\nWARNING: reference arm '{reference}' not found; skipping tests")
|
||||||
return 1
|
return 1
|
||||||
print(f"\nEXACT TWO-SIDED PERMUTATION TEST (per-run values) vs `{reference}`")
|
|
||||||
print(f"{'metric':<12} {'arm':<14} {'obs(diff)':>12} {'p':>8} permutations")
|
|
||||||
print("-" * 64)
|
|
||||||
ref = data[reference]["per"]
|
ref = data[reference]["per"]
|
||||||
for arm in arms:
|
print("\nPAIRWISE PERMUTATION TEST (per-run values) + MANN-WHITNEY CROSS-CHECK")
|
||||||
name = arm["name"]
|
print(f"permutation: exact when C(n,na) <= {EXACT_CAP:,}; otherwise "
|
||||||
if name == reference:
|
f"Monte-Carlo {MC_DRAWS:,} draws, seed={MC_SEED:#x}, "
|
||||||
|
f"p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))")
|
||||||
|
hdr = (f"{'metric':<11} {'A':<11} {'B':<11} {'diff(A-B)':>11} "
|
||||||
|
f"{'perm p':>9} {'method':<12} {'MC se':>7} {'MW p':>9} {'MW U':>8}")
|
||||||
|
print(hdr)
|
||||||
|
print("-" * len(hdr))
|
||||||
|
for i in range(len(arms)):
|
||||||
|
for j in range(i + 1, len(arms)):
|
||||||
|
ana, bnb = arms[i]["name"], arms[j]["name"]
|
||||||
|
for label, key in (("dmg/run", "damage"), ("round wins", "wins")):
|
||||||
|
xa = [p[key] for p in data[ana]["per"] if p[key] is not None]
|
||||||
|
xb = [p[key] for p in data[bnb]["per"] if p[key] is not None]
|
||||||
|
res = perm_test(xa, xb)
|
||||||
|
mw = mannwhitney_p(xa, xb)
|
||||||
|
if res is None:
|
||||||
|
print(f"{label:<11} {ana:<11} {bnb:<11} {'n/a':>11}")
|
||||||
|
continue
|
||||||
|
mstr = ("exact" if res["method"] == "exact"
|
||||||
|
else f"MC/B={res['draws']:,}")
|
||||||
|
sestr = "-" if res["method"] == "exact" else f"{res['se']:.4f}"
|
||||||
|
mwp = f"{mw[0]:.4f}" if mw else "n/a"
|
||||||
|
mwu = f"{mw[1]:.1f}" if mw else "n/a"
|
||||||
|
print(f"{label:<11} {ana:<11} {bnb:<11} {res['signed']:>+11.3f} "
|
||||||
|
f"{res['p']:>9.4f} {mstr:<12} {sestr:>7} {mwp:>9} {mwu:>8}")
|
||||||
|
|
||||||
|
# ── minimum detectable effect ──────────────────────────────────────────
|
||||||
|
print("\nMINIMUM DETECTABLE EFFECT (two-sample, alpha=0.05 two-sided, "
|
||||||
|
"80% power; MDE = 2.8016*sd*sqrt(2/n))")
|
||||||
|
print(f"{'metric':<11} {'n/arm':>6} {'sd(control)':>12} {'MDE(abs)':>10} "
|
||||||
|
f"{'MDE vs control mean':>22}")
|
||||||
|
print("-" * 64)
|
||||||
|
ctrl = data[reference]["per"]
|
||||||
|
for label, key in (("dmg/run", "damage"), ("round wins", "wins")):
|
||||||
|
vals = [p[key] for p in ctrl if p[key] is not None]
|
||||||
|
if len(vals) < 2:
|
||||||
continue
|
continue
|
||||||
for label, key in (("dmg/run", "damage"), ("round wins", "wins")):
|
sd = math.sqrt(_var(vals))
|
||||||
xa = [p[key] for p in ref if p[key] is not None]
|
mde = min_detectable_effect(sd, len(vals))
|
||||||
xb = [p[key] for p in data[name]["per"] if p[key] is not None]
|
cmean = sum(vals) / len(vals)
|
||||||
res = perm_test(xa, xb)
|
rel = (f"{100.0*mde/abs(cmean):.1f}% of {cmean:.1f}"
|
||||||
if res is None:
|
if cmean else "n/a")
|
||||||
print(f"{label:<12} {name:<14} {'n/a':>12} {'n/a':>8}")
|
print(f"{label:<11} {len(vals):>6} {sd:>12.3f} {mde:>10.3f} {rel:>22}")
|
||||||
continue
|
|
||||||
obs, p, ncomb, exact = res
|
|
||||||
note = f"C({len(xa)+len(xb)},{len(xa)})={ncomb}" + (
|
|
||||||
"" if exact else " SAMPLED")
|
|
||||||
print(f"{label:<12} {name:<14} {obs:>+12.3f} {p:>8.4f} {note}")
|
|
||||||
|
|
||||||
# ── round-level test (anti-conservative) ─────────────────────────────────
|
# ── round-level test (anti-conservative) ─────────────────────────────────
|
||||||
print("\nROUND-LEVEL TEST (pooled rounds, Fisher exact) vs "
|
print("\nROUND-LEVEL TEST (pooled rounds, Fisher exact) vs "
|
||||||
@@ -492,6 +619,22 @@ def main():
|
|||||||
lv_fail += 1
|
lv_fail += 1
|
||||||
print(f" {name:<14} {status:<4} ({why})")
|
print(f" {name:<14} {status:<4} ({why})")
|
||||||
|
|
||||||
|
# ── [bb] applied-shift check (a placebo must apply exactly 0) ────────────
|
||||||
|
print("\n[bb] APPLIED-SHIFT CHECK (from bot stdout; needs TR_BITBRAIN_LOG=1). "
|
||||||
|
"A provably-zero placebo emits ZERO [bb] lines.")
|
||||||
|
print(f" {'arm':<14} {'runs w/log':>11} {'lines':>7} {'min':>9} "
|
||||||
|
f"{'max':>9} {'zeros':>6}")
|
||||||
|
for arm in arms:
|
||||||
|
name = arm["name"]
|
||||||
|
vals, runs_with, nruns = bb_shifts(data[name]["dir"], data[name]["runs"])
|
||||||
|
if not vals:
|
||||||
|
print(f" {name:<14} {runs_with:>4}/{nruns:<4} {0:>7} "
|
||||||
|
f"{'-':>9} {'-':>9} {'-':>6}")
|
||||||
|
else:
|
||||||
|
zeros = sum(1 for v in vals if v == 0.0)
|
||||||
|
print(f" {name:<14} {runs_with:>4}/{nruns:<4} {len(vals):>7} "
|
||||||
|
f"{min(vals):>+9.2f} {max(vals):>+9.2f} {zeros:>6}")
|
||||||
|
|
||||||
# ── round-win attribution cross-check ────────────────────────────────────
|
# ── round-win attribution cross-check ────────────────────────────────────
|
||||||
print("\nROUND-WIN ATTRIBUTION (events primary; score tie-break for "
|
print("\nROUND-WIN ATTRIBUTION (events primary; score tie-break for "
|
||||||
"mutual-kill / timeout rounds)")
|
"mutual-kill / timeout rounds)")
|
||||||
|
|||||||
Reference in New Issue
Block a user