4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
13 KiB
BitBrain vs TMHorizon vs Pattern — LIVE A/B (null result; the identity check passes)
Question (the owner's claim, 2026-09-25). "Bitbrain gun is better than
TMhorizon, i won drussgt with it and tfil move 6 to 4." This is the second
6-4 report (the earlier one was a TMHorizon+BitBrain mix). With the measured
baseline of ~49% round wins vs real DrussGT, P(>=6 of 10) = 0.353 — one
round in three by chance, so 6-4 is not evidence. This test replaces the
impression with a measurement on the shipped TFIL default (no TR_MOVEMENT
override), real DrussGT.
Prior knowledge (not re-derived here). At 30 runs/arm BitBrain was
statistically IDENTICAL to Pattern (docs/bitbrain_gun_verdict.md, d93ce44:
damage 283.7 vs 280.4, round wins 97/210 vs 97/210, p=1.0, MDE 24.4 dmg),
and Pattern is in general >= TMHorizon. So "BitBrain > TMHorizon" is what those
facts predict — which is exactly why it needed a direct test.
DIRECT ANSWERS (MEASURED, 15 runs/arm)
- Does BitBrain beat TMHorizon? No. The identity config (
bb_id) vstmhis indistinguishable on both verdict metrics (+2.6 damage/run, p=0.82; +0.267 wins/run, p=0.56). The learned config (bb_learn) — the one the owner most likely ran — is worse thantmhin point estimate (-17.3 damage/run, p=0.16; -0.47 wins/run, p=0.30), though not significantly.- Does BitBrain beat Pattern? No.
bb_idis statistically identical to Pattern (the validity check below).bb_learnis the worst arm on both verdict metrics: -25.9 damage/run vs Pattern (p=0.048, the only comparison in the whole table at alpha=0.05) and -0.60 wins/run (p=0.22).- Does the learned-gain arm differ from the identity? Yes, weakly, in the wrong direction.
bb_learnvsbb_id: -19.9 damage/run (p=0.139) and -0.733 wins/run (p=0.085, 95% CI [-1.46, -0.01]). The learned gains (all >= 1.0, gated to range >= 300 px) over-lead and lose long-range hits.- What this test CAN see: an effect of ~29.5 damage/run (9.8%) or ~1.235 wins/run (17.6 pp) at 80% power. The claimed 6-4 (+11 pp, ~0.77 wins/run) needs ~39 runs/arm to resolve, so this test is underpowered for the claimed effect size — absence of a small positive BitBrain effect is not excluded. What it does exclude is a large one, and it does detect that the learned arm is ~10% worse than Pattern.
Setup [MEASURED]
One frozen ModularBot from git archive HEAD at the run-time commit
ed25ce29ad28838eaf726137e60c036e8ed70de3, binary sha256
0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93; 4 arms ×
15 runs × 7 rounds = 60 battles, 420 rounds, --conc 7, 0 failed. (Another
job committed ccff7e3 after this session started; that commit does not touch the
shipped TFIL default, and the frozen binary is pinned to ed25ce2.) Raw per-tick
captures live at /tmp/ab/j113_bb_vs_tmh/ and are not committed; every table
below is in common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt,
..._bands.txt, ..._session.json.
| arm | env | role |
|---|---|---|
pattern |
(none) | shipped default (onlyPattern rack) — reference |
bb_id |
TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 |
gain 1.0 identity — validity check |
bb_learn |
... TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay |
the config the owner likely ran |
tmh |
TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both |
TMHorizon-only rack |
1. Validity check — bb_id vs pattern [MEASURED]
PASS. TR_BITBRAIN_GAINS=1.0 is a single candidate, so the gun applies a
FIXED gain 1.0 (aim = LOS + 1.0*(patternAim - LOS), no learning) in the long
bands — the identity to Pattern. It is statistically indistinguishable:
| comparison | diff (arm - ref) | 95% CI (Welch) | perm p | MW p |
|---|---|---|---|---|
bb_id vs pattern, damage/run |
-6.0 | [-27.9, +15.9] | 0.597 | 0.772 |
bb_id vs pattern, wins/run |
+0.133 | [-0.63, +0.90] | 0.865 | 0.845 |
Both differences are a small fraction of the MDE, and the round-level Fisher test is 50/105 vs 48/105 (p=0.89). The rack swap (Pattern id 5 -> BitBrain id 16) and the TmHorizon observation ring introduce no detectable systematic offset. The remaining arms therefore sit on validated plumbing.
2. Primary result — damage/run and ROUND WINS [MEASURED]
Analyzer ARM SUMMARY (verbatim, tools/ab/ab_analyze.py):
arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run
--------------------------------------------------------------------------
pattern 15 302 238 48/105 45.7 761 99.4
bb_id 15 296 217 50/105 47.6 787 97.5
bb_learn 15 276 217 39/105 37.1 761 100.3
tmh 15 294 220 46/105 43.8 753 96.0
pattern wins 48/105 = 45.7% of rounds, consistent with the ~49% measured
baseline. bb_learn wins 39/105 = 37.1% — the fewest of all four arms.
hitstk/run (hits taken) and shots/run are context only; per the standing rule
nothing here is judged on hit rate.
Per-run values (never just the mean), from the analyzer:
pattern dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321
wins: r1=3 r2=2 r3=5 r4=1 r5=1 r6=4 r7=5 r8=3 r9=3 r10=4 r11=3 r12=4 r13=3 r14=3 r15=4
bb_id dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292
wins: r1=3 r2=4 r3=4 r4=3 r5=4 r6=3 r7=4 r8=5 r9=2 r10=4 r11=3 r12=2 r13=2 r14=4 r15=3
bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345
wins: r1=2 r2=4 r3=1 r4=2 r5=3 r6=3 r7=3 r8=3 r9=1 r10=2 r11=4 r12=4 r13=1 r14=2 r15=4
tmh dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267
wins: r1=2 r2=4 r3=3 r4=3 r5=3 r6=2 r7=2 r8=3 r9=5 r10=4 r11=4 r12=3 r13=2 r14=4 r15=2
Every pairwise test the analyzer emits, with the Welch 95% CI for the same
difference (sign is A - B, so negative means B is better):
metric A B diff(A-B) perm p MW p
dmg/run pattern bb_id +5.967 0.5968 0.7716
round wins pattern bb_id -0.133 0.8647 0.8450
dmg/run pattern bb_learn +25.894 0.0483 0.0620
round wins pattern bb_learn +0.600 0.2197 0.1837
dmg/run pattern tmh +8.448 0.4038 0.4306
round wins pattern tmh +0.133 0.8671 0.6196
dmg/run bb_id bb_learn +19.927 0.1386 0.1844
round wins bb_id bb_learn +0.733 0.0853 0.0842
dmg/run bb_id tmh +2.481 0.8174 0.7089
round wins bb_id tmh +0.267 0.5574 0.4206
dmg/run bb_learn tmh -17.446 0.1600 0.1585
round wins bb_learn tmh -0.467 0.3019 0.3012
| comparison (arm - ref) | damage/run | 95% CI | wins/run | 95% CI |
|---|---|---|---|---|
bb_id - pattern |
-6.0 | [-27.9, +15.9] | +0.13 | [-0.63, +0.90] |
bb_learn - pattern |
-25.9 | [-50.5, -1.3] | -0.60 | [-1.43, +0.23] |
tmh - pattern |
-8.6 | [-28.3, +11.1] | -0.13 | [-0.91, +0.65] |
bb_id - tmh |
+2.6 | [-18.4, +23.6] | +0.27 | [-0.40, +0.93] |
bb_learn - tmh |
-17.3 | [-41.0, +6.5] | -0.47 | [-1.21, +0.28] |
bb_learn - bb_id |
-19.9 | [-45.5, +5.8] | -0.73 | [-1.46, -0.01] |
The single comparison that crosses alpha=0.05 is Pattern beating bb_learn on
damage (p=0.048; MW p=0.062). Every BitBrain-vs-TMHorizon and BitBrain-vs-Pattern
comparison is a null or a point estimate in the wrong direction for the claim.
Round-level Fisher (anti-conservative; rounds cluster within runs):
bb_id 50/105 vs pattern 48/105 p=0.89; bb_learn 39/105 p=0.26;
tmh 46/105 p=0.89.
3. Minimum detectable effect and power [MEASURED]
Analyzer MDE (alpha=0.05 two-sided, 80% power):
metric n/arm sd(control) MDE(abs) MDE vs control mean
dmg/run 15 28.854 29.517 9.8% of 302.2
round wins 15 1.207 1.235 38.6% of 3.2
At 15 runs/arm this test resolves a 29.5 damage/run (9.8%) or 1.235
wins/run (17.6 pp) effect. The observed bb_id-vs-pattern difference (-6.0
damage) and the bb_id-vs-tmh difference (+2.6) are far inside that, so
"no difference detectable" is the honest reading for BitBrain vs TMHorizon —
not "no difference exists." For the claimed effect, post-hoc power:
the measured sd(wins/run) is 1.207, so 6-4 (60% vs the ~49% baseline, i.e.
+11 pp = +0.77 wins/run) needs ~39 runs/arm and +0.5 wins/run needs ~92
runs/arm. This 15-run/arm test is underpowered for a small positive BitBrain
effect; what it is powered for is a large one, and it rules out the learned
arm being better by ~1.2 wins/run.
4. Hit rate by range band [MEASURED]
BitBrain's only moving part is gated to range >= 300 px (BB_GAIN_BAND_MIN=3);
its gain is pinned to 1.0 below that, so bands 0-300 are a second identity check
and 300+ is where any effect must appear. Pooled hits/shots per arm
(tools/ab/ab_range_bands.py):
band bb_id(s/h) bb_learn(s/h) pattern(s/h) tmh(s/h)
0-100 2/0 0.0% 3/2 66.7% 10/4 40.0% 11/5 45.5%
100-200 30/6 20.0% 46/11 23.9% 65/13 20.0% 62/14 22.6%
200-300 287/41 14.3% 303/50 16.5% 287/45 15.7% 298/51 17.1%
300-450 5708/711 12.5% 5673/659 11.6% 5679/753 13.3% 5726/647 11.3%
450+ 5489/532 9.7% 5109/440 8.6% 5076/470 9.3% 4940/494 10.0%
ALL 11516/1290 11.2% 11134/1162 10.4% 11117/1285 11.6% 11037/1211 11.0%
Per-band permutation test on per-run band rates vs pattern:
band arm d(pp) p method MDE(pp)
300-450 bb_id -0.72 0.1996 MC/B=200,000 1.46
300-450 bb_learn -1.61 0.0035 MC/B=200,000 1.46
300-450 tmh -1.88 0.0010 MC/B=200,000 1.46
450+ bb_id +0.42 0.3927 MC/B=200,000 1.39
450+ bb_learn -0.57 0.2987 MC/B=200,000 1.39
450+ tmh +0.85 0.1511 MC/B=200,000 1.39
ALL bb_learn -1.13 0.0021 MC/B=200,000 0.74
ALL tmh -0.54 0.1397 MC/B=200,000 0.74
The identity arm bb_id is flat in both moving bands (p=0.20, 0.39) — the
second identity check passes. The learned arm loses 1.61 pp of hit rate at
300-450 px (p=0.0035), exactly where its gain fires: because the candidate set
is {1.0, 1.25, 1.5, 2.0} it can only over-lead, never shrink the lead, and
the campaign ledger's Phase 1 measured the optimal long-range gain to be
below 1.0. That is the mechanism behind the bb_learn damage loss.
5. Liveness [MEASURED]
pattern OK (15/15 runs: no arm env; report present)
bb_id OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied)
bb_learn OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied)
tmh OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied)
Round-win attribution cross-check: wins==firstPlaces 15/15 runs for every arm.
(TR_BITBRAIN_LOG was deliberately not set, so the [bb] applied-shift section
is empty; gains liveness is read from the boot env report instead.)
6. Verdict
Null result, honestly stated. On the shipped TFIL mover and real DrussGT, at
15 runs/arm: the BitBrain identity is statistically identical to Pattern (as
predicted by d93ce44), and no BitBrain arm beats either TMHorizon or
Pattern. The learned-gain config the owner most likely ran is the worst arm
(276 damage/run, 39/105 round wins), significantly worse than Pattern on damage
(p=0.048) and weakly worse than the identity on round wins (p=0.085), via an
over-lead at 300-450 px. P(>=6 of 10) = 0.353 at the ~49% baseline means the
6-4 anecdote carries no information, and a 6-4-sized positive effect would need
~39 runs/arm to be seen at all.
MEASURED vs INFERRED
- MEASURED: the 4-arm table (60 battles, 0 failed, frozen commit
ed25ce2, binary0fea927b…); the validity check; all pairwise permutation/MW p-values, MDEs and the Welch CIs; the per-band hit rates and permutation tests; the round-win attribution; liveness. - INFERRED: that the over-lead mechanism (
candidates >= 1.0+ range gate) causes thebb_learnlong-range loss (consistent with the ledger's Phase 1 measured gain curve, not proven by this session); thatbb_learnis the exact config the owner ran; that the owner's "6-4" means round wins per the task framing. - NOT EXCLUDED: a small positive BitBrain effect below the MDE (~30 damage/run, ~1.2 wins/run). The test is underpowered for the claimed effect size; do not read this as "BitBrain is provably not better," read it as "no difference detectable at 15 runs/arm."