Files
SirRoboGarage/docs/bitbrain_gun_verdict.md
SirStone d93ce444c0 BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
  provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
  (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
  7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
  (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
  unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
  a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
  standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
  detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
2026-09-24 23:02:59 +02:00

12 KiB

BitBrain verdict — CLEAN NEGATIVE against real DrussGT

Direct answer (MEASURED). The BitBrain decay-memory gun does not beat the shipped Pattern rack, and it does not beat a provably-zero placebo. At 30 runs/arm (210 rounds/arm) bb_decay is +3.3 damage/run over control (permutation p = 0.709, Mann-Whitney p = 0.600) and round wins are dead level, 97/210 vs 97/210 (p = 1.000). Every one of the twelve pairwise comparisons is a null (all p >= 0.29). The 7-run "shape" that motivated this test — 294 vs 279 damage, 26 vs 22 wins — did not replicate: at 30 runs the same difference is +3.3 damage and +0 wins. The one mechanism the offline gate test's diagnosis predicted (a bounded/decaying SBC memory) buys nothing live. The BitBrain thread closes here, with a reason instead of an open question.

What was run (MEASURED)

  • Frozen from HEAD 795a0e5 via git archive HEAD; ModularBot binary sha256 cab6083018672168642fd5a3e62259a260a43ff3ebfc0418dd3cbe15ab8e2005.
  • Real DrussGT through the robocode_shim bridge, 4 arms x 30 runs x 7 rounds (120 battles, 0 failed), --conc 7.
  • Arms (one env knob each, all four verified live in the boot report):
    • control — no env; shipped onlyPattern rack.
    • bb_decay — TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay (the candidate).
    • bb_zero — as bb_decay + TR_BITBRAIN_RANGE=0; this is NOT a zero-shift placebo (see below), so it is reported as a small-dose treatment.
    • bb_placebo — as bb_decay + TR_BITBRAIN_MIN_OBS=100000000; the valid placebo (the whole BitBrain path runs, but the readout branch is never reached, so the applied shift is provably exactly 0).
  • TR_BITBRAIN_LOG=1 on the three BitBrain arms so the applied shift is visible and the placebo can be checked.

The clear answer to the rubric

  • bb_decay beats control on damage/run AND round wins with p < 0.05? No (+3.3 dmg, p = 0.709; +0 wins, p = 1.000).
  • bb_decay ~= bb_placebo? Yes (dmg p = 0.897, wins p = 0.918). So any difference the BitBrain path makes is plumbing/noise, not learning — and here it makes no difference at all.
  • bb_zero ~= control and bb_decay > both? Not the situation: bb_zero is not a zero-shift placebo, and bb_decay beats nothing.
  • Nothing separates -> CLEAN NEGATIVE. Not spun: this is a successful outcome that closes the thread.

The placebo: TR_BITBRAIN_RANGE=0 is clamped, so it is NOT a zero-shift arm (MEASURED)

common_libs/guns/bitbrain_gun.nim:207:

result.maxDeg = clamp(envFloatBB(BB_RANGE_ENV, BB_RANGE_DEF), 1.0, 180.0)

TR_BITBRAIN_RANGE=0 is therefore clamped to 1.0 degree, and the class centres become +/-0.97 deg, not 0. Confirmed live: the boot report shows [env] TR_BITBRAIN_RANGE = 1.0 (source: env) and the [bb] log emits shift=-1.0deg / shift=+1.0deg (edge classes 0 and 31). So bb_zero applies a systematic ~1 deg correction and is not the intended placebo. Per the task's fallback this arm was kept only as a small-dose treatment; the true placebo is bb_placebo, which emits 0 [bb] lines in 30/30 runs — a zero applied shift by construction (bbLog is reached only inside the same trained >= minObs branch that computes the shift, so a never-firing readout is silent). Nominally bb_zero trends worse (-5.5 dmg/run vs control, p = 0.51; -7.7 vs the true placebo, p = 0.35), consistent with a tiny systematic mis-aim rather than a neutral probe — and still inside the noise.

Raw analyzer output (verbatim, tools/ab/ab_analyze.py)

# session /tmp/ab/bbverdict
# commit=795a0e59febb4d9b9722130ad47cd2dd8f72a315 binary_sha256=cab6083018672168642fd5a3e62259a260a43ff3ebfc0418dd3cbe15ab8e2005 rounds=7 runs=30 conc=7 ts=2026-09-24T22:37:38+02:00

ARM SUMMARY
arm            runs  dmg/run dmgtk/run    wins   win% shots/run hitstk/run
--------------------------------------------------------------------------
control          30      280       206  97/210   46.2       796       92.5
bb_decay         30      284       206  97/210   46.2       797       93.6
bb_zero          30      275       209  93/210   44.3       781       91.8
bb_placebo       30      283       209  95/210   45.2       786       92.2

PER-RUN (never just the mean)
  control        dmg:  r1=293 r2=296 r3=352 r4=263 r5=282 r6=292 r7=246 r8=271 r9=246 r10=261 r11=263 r12=175 r13=302 r14=288 r15=280 r16=333 r17=244 r18=293 r19=317 r20=249 r21=283 r22=296 r23=278 r24=281 r25=272 r26=292 r27=238 r28=332 r29=295 r30=299
                 wins: r1=3/7 r2=5/7 r3=4/7 r4=3/7 r5=5/7 r6=3/7 r7=2/7 r8=2/7 r9=3/7 r10=3/7 r11=3/7 r12=0/7 r13=4/7 r14=4/7 r15=3/7 r16=2/7 r17=4/7 r18=3/7 r19=3/7 r20=3/7 r21=2/7 r22=2/7 r23=4/7 r24=5/7 r25=4/7 r26=3/7 r27=3/7 r28=6/7 r29=3/7 r30=3/7
  bb_decay       dmg:  r1=284 r2=247 r3=273 r4=300 r5=241 r6=248 r7=276 r8=299 r9=227 r10=279 r11=305 r12=349 r13=302 r14=318 r15=279 r16=244 r17=307 r18=293 r19=328 r20=212 r21=296 r22=296 r23=362 r24=258 r25=245 r26=298 r27=308 r28=294 r29=285 r30=259
                 wins: r1=4/7 r2=2/7 r3=5/7 r4=3/7 r5=3/7 r6=3/7 r7=3/7 r8=3/7 r9=2/7 r10=4/7 r11=2/7 r12=4/7 r13=3/7 r14=6/7 r15=3/7 r16=2/7 r17=5/7 r18=4/7 r19=4/7 r20=2/7 r21=2/7 r22=2/7 r23=6/7 r24=1/7 r25=3/7 r26=3/7 r27=6/7 r28=4/7 r29=1/7 r30=2/7
  bb_zero        dmg:  r1=248 r2=295 r3=267 r4=310 r5=277 r6=313 r7=264 r8=276 r9=277 r10=236 r11=302 r12=262 r13=315 r14=251 r15=314 r16=297 r17=231 r18=266 r19=212 r20=224 r21=249 r22=235 r23=314 r24=314 r25=285 r26=262 r27=296 r28=318 r29=271 r30=265
                 wins: r1=1/7 r2=4/7 r3=3/7 r4=2/7 r5=3/7 r6=6/7 r7=3/7 r8=4/7 r9=3/7 r10=3/7 r11=3/7 r12=2/7 r13=4/7 r14=4/7 r15=4/7 r16=2/7 r17=1/7 r18=4/7 r19=3/7 r20=2/7 r21=3/7 r22=2/7 r23=5/7 r24=2/7 r25=4/7 r26=2/7 r27=4/7 r28=3/7 r29=3/7 r30=4/7
  bb_placebo     dmg:  r1=279 r2=285 r3=328 r4=253 r5=275 r6=286 r7=348 r8=256 r9=258 r10=303 r11=278 r12=242 r13=291 r14=234 r15=341 r16=274 r17=240 r18=267 r19=312 r20=228 r21=302 r22=337 r23=320 r24=266 r25=308 r26=291 r27=255 r28=275 r29=308 r30=234
                 wins: r1=2/7 r2=3/7 r3=4/7 r4=1/7 r5=4/7 r6=2/7 r7=4/7 r8=2/7 r9=2/7 r10=4/7 r11=4/7 r12=2/7 r13=5/7 r14=2/7 r15=4/7 r16=3/7 r17=3/7 r18=3/7 r19=4/7 r20=2/7 r21=3/7 r22=4/7 r23=4/7 r24=4/7 r25=5/7 r26=2/7 r27=3/7 r28=3/7 r29=5/7 r30=2/7

PAIRWISE PERMUTATION TEST (per-run values) + MANN-WHITNEY CROSS-CHECK
permutation: exact when C(n,na) <= 20,000,000; otherwise Monte-Carlo 1,000,000 draws, seed=0x5eed5eed, p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))
metric      A           B             diff(A-B)    perm p method         MC se      MW p     MW U
-------------------------------------------------------------------------------------------------
dmg/run     control     bb_decay         -3.313    0.7091 MC/B=1,000,000  0.0005    0.5997    414.0
round wins  control     bb_decay         +0.000    1.0000 MC/B=1,000,000  0.0000    0.7468    428.5
dmg/run     control     bb_zero          +5.551    0.5086 MC/B=1,000,000  0.0005    0.5493    409.0
round wins  control     bb_zero          +0.133    0.7369 MC/B=1,000,000  0.0004    0.6535    420.5
dmg/run     control     bb_placebo       -2.175    0.8030 MC/B=1,000,000  0.0004    0.9117    442.0
round wins  control     bb_placebo       +0.067    0.9091 MC/B=1,000,000  0.0003    0.8356    436.0
dmg/run     bb_decay    bb_zero          +8.863    0.2947 MC/B=1,000,000  0.0005    0.4376    397.0
round wins  bb_decay    bb_zero          +0.133    0.7597 MC/B=1,000,000  0.0004    0.9027    441.5
dmg/run     bb_decay    bb_placebo       +1.138    0.8969 MC/B=1,000,000  0.0003    0.7845    431.0
round wins  bb_decay    bb_placebo       +0.067    0.9178 MC/B=1,000,000  0.0003    0.9513    445.5
dmg/run     bb_zero     bb_placebo       -7.725    0.3525 MC/B=1,000,000  0.0005    0.4733    401.0
round wins  bb_zero     bb_placebo       -0.067    0.9072 MC/B=1,000,000  0.0003    0.7763    431.0

MINIMUM DETECTABLE EFFECT (two-sample, alpha=0.05 two-sided, 80% power; MDE = 2.8016*sd*sqrt(2/n))
metric       n/arm  sd(control)   MDE(abs)    MDE vs control mean
----------------------------------------------------------------
dmg/run         30       33.669     24.355          8.7% of 280.4
round wins      30        1.165      0.843           26.1% of 3.2

ROUND-LEVEL TEST (pooled rounds, Fisher exact) vs `control` — ANTI-CONSERVATIVE: rounds cluster within runs
arm              ref wins   arm wins        p
----------------------------------------------
bb_decay           97/210     97/210   1.0000
bb_zero            97/210     93/210   0.7687
bb_placebo         97/210     95/210   0.9220

LIVENESS (arm env applied in the bot's own boot report)
  control        OK   (30/30 runs: no arm env; report present)
  bb_decay       OK   (30/30 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay TR_BITBRAIN_LOG=1 applied)
  bb_zero        OK   (30/30 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay TR_BITBRAIN_RANGE=0 TR_BITBRAIN_LOG=1 applied)
  bb_placebo     OK   (30/30 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_MEM=decay TR_BITBRAIN_MIN_OBS=100000000 TR_BITBRAIN_LOG=1 applied)

[bb] APPLIED-SHIFT CHECK (from bot stdout; needs TR_BITBRAIN_LOG=1). A provably-zero placebo emits ZERO [bb] lines.
  arm             runs w/log   lines       min       max  zeros
  control           0/30         0         -         -      -
  bb_decay         30/30       905    -38.80    +36.20      0
  bb_zero          30/30       408     -1.00     +1.00      9
  bb_placebo        0/30         0         -         -      -

ROUND-WIN ATTRIBUTION (events primary; score tie-break for mutual-kill / timeout rounds)
  control        wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 205/205 (5 tie-broken)
  bb_decay       wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 209/209 (1 tie-broken)
  bb_zero        wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 209/209 (1 tie-broken)
  bb_placebo     wins==firstPlaces 30/30 runs OK; single-death rounds agree with score 207/207 (3 tie-broken)

Minimum detectable effect (MEASURED) — what this test CAN and CANNOT see

From the observed control-arm per-run SD (damage SD = 33.7; wins SD = 1.17) at 30 runs/arm, alpha = 0.05 two-sided, 80% power:

  • damage/run: MDE = 24.4 (8.7% of the 280 control mean).
  • round wins: MDE = 0.84 wins/run (26.1% of the 3.2 mean).

So this test rules out a bb_decay damage gain of >= 24 dmg/run (and a win gain of >= 0.84/run). A real effect of, say, +10-15 dmg/run — the size the 7-run result hinted at — would not be detectable here, and is not ruled out by the null. That is the honest bound: "no effect >= 24 dmg/run", not "no effect". The placebo isolates the mechanism more sharply: bb_decay vs bb_placebo is +1.1 dmg/run (p = 0.897) and +0.07 wins (p = 0.918), so the learning-specific effect is bounded by the same ~24 dmg/run and shows no sign.

Why the earlier 7-run shape was misleading (MEASURED)

metric 7 runs/arm (commit 795a0e5 session) 30 runs/arm (this session)
control dmg/run 279 280
bb_decay dmg/run 294 (+15) 284 (+3.3)
control round wins 22/49 97/210
bb_decay round wins 26/49 (+4) 97/210 (+0)

The 7-run difference was inside the noise; it has shrunk to zero, not grown.

MEASURED vs INFERRED

  • MEASURED: the damage/win table, the pairwise permutation p-values (Monte-Carlo, 1,000,000 draws, seed 0x5eed5eed, reported with MC SE) and the Mann-Whitney cross-check, the MDE from the observed SD, the liveness lines, the [bb] shift check, and the TR_BITBRAIN_RANGE clamp (= 1.0 in the live boot report).
  • INFERRED: that a true effect below ~24 dmg/run would need more runs or a lower-variance opponent to resolve. The offline gate-test diagnosis (decay memory should help) is not supported live; whether a different memory regime would help is not tested here (only decay was).

Reproduce

tools/ab/ab_run.sh --arms /tmp/ab/arms_bbverdict.txt --runs 30 \
  --outdir /tmp/ab/bbverdict --conc 7
python3 tools/ab/ab_analyze.py /tmp/ab/bbverdict

Analyzer tooling improved in the same change: exact enumeration is kept when C(n, na) <= 20e6 (7v7), otherwise a seeded Monte-Carlo permutation test (MC_DRAWS = 1,000,000, MC_SEED = 0x5eed5eed) with its standard error, plus a tie-corrected Mann-Whitney U cross-check and the MDE line. See tools/ab/README.md.