Files
SirRoboGarage/docs/bitbrain_vs_tmhorizon_ab.md
T
SirStone 4829f9ca13 BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result)
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
2026-09-25 23:55:21 +02:00

13 KiB
Raw Blame History

BitBrain vs TMHorizon vs Pattern — LIVE A/B (null result; the identity check passes)

Question (the owner's claim, 2026-09-25). "Bitbrain gun is better than TMhorizon, i won drussgt with it and tfil move 6 to 4." This is the second 6-4 report (the earlier one was a TMHorizon+BitBrain mix). With the measured baseline of ~49% round wins vs real DrussGT, P(>=6 of 10) = 0.353 — one round in three by chance, so 6-4 is not evidence. This test replaces the impression with a measurement on the shipped TFIL default (no TR_MOVEMENT override), real DrussGT.

Prior knowledge (not re-derived here). At 30 runs/arm BitBrain was statistically IDENTICAL to Pattern (docs/bitbrain_gun_verdict.md, d93ce44: damage 283.7 vs 280.4, round wins 97/210 vs 97/210, p=1.0, MDE 24.4 dmg), and Pattern is in general >= TMHorizon. So "BitBrain > TMHorizon" is what those facts predict — which is exactly why it needed a direct test.

DIRECT ANSWERS (MEASURED, 15 runs/arm)

  • Does BitBrain beat TMHorizon? No. The identity config (bb_id) vs tmh is indistinguishable on both verdict metrics (+2.6 damage/run, p=0.82; +0.267 wins/run, p=0.56). The learned config (bb_learn) — the one the owner most likely ran — is worse than tmh in point estimate (-17.3 damage/run, p=0.16; -0.47 wins/run, p=0.30), though not significantly.
  • Does BitBrain beat Pattern? No. bb_id is statistically identical to Pattern (the validity check below). bb_learn is the worst arm on both verdict metrics: -25.9 damage/run vs Pattern (p=0.048, the only comparison in the whole table at alpha=0.05) and -0.60 wins/run (p=0.22).
  • Does the learned-gain arm differ from the identity? Yes, weakly, in the wrong direction. bb_learn vs bb_id: -19.9 damage/run (p=0.139) and -0.733 wins/run (p=0.085, 95% CI [-1.46, -0.01]). The learned gains (all >= 1.0, gated to range >= 300 px) over-lead and lose long-range hits.
  • What this test CAN see: an effect of ~29.5 damage/run (9.8%) or ~1.235 wins/run (17.6 pp) at 80% power. The claimed 6-4 (+11 pp, ~0.77 wins/run) needs ~39 runs/arm to resolve, so this test is underpowered for the claimed effect size — absence of a small positive BitBrain effect is not excluded. What it does exclude is a large one, and it does detect that the learned arm is ~10% worse than Pattern.

Setup [MEASURED]

One frozen ModularBot from git archive HEAD at the run-time commit ed25ce29ad28838eaf726137e60c036e8ed70de3, binary sha256 0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93; 4 arms × 15 runs × 7 rounds = 60 battles, 420 rounds, --conc 7, 0 failed. (Another job committed ccff7e3 after this session started; that commit does not touch the shipped TFIL default, and the frozen binary is pinned to ed25ce2.) Raw per-tick captures live at /tmp/ab/j113_bb_vs_tmh/ and are not committed; every table below is in common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt, ..._bands.txt, ..._session.json.

arm env role
pattern (none) shipped default (onlyPattern rack) — reference
bb_id TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 gain 1.0 identity — validity check
bb_learn ... TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay the config the owner likely ran
tmh TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both TMHorizon-only rack

1. Validity check — bb_id vs pattern [MEASURED]

PASS. TR_BITBRAIN_GAINS=1.0 is a single candidate, so the gun applies a FIXED gain 1.0 (aim = LOS + 1.0*(patternAim - LOS), no learning) in the long bands — the identity to Pattern. It is statistically indistinguishable:

comparison diff (arm - ref) 95% CI (Welch) perm p MW p
bb_id vs pattern, damage/run -6.0 [-27.9, +15.9] 0.597 0.772
bb_id vs pattern, wins/run +0.133 [-0.63, +0.90] 0.865 0.845

Both differences are a small fraction of the MDE, and the round-level Fisher test is 50/105 vs 48/105 (p=0.89). The rack swap (Pattern id 5 -> BitBrain id 16) and the TmHorizon observation ring introduce no detectable systematic offset. The remaining arms therefore sit on validated plumbing.

2. Primary result — damage/run and ROUND WINS [MEASURED]

Analyzer ARM SUMMARY (verbatim, tools/ab/ab_analyze.py):

arm            runs  dmg/run dmgtk/run    wins   win% shots/run hitstk/run
--------------------------------------------------------------------------
pattern          15      302       238  48/105   45.7       761       99.4
bb_id            15      296       217  50/105   47.6       787       97.5
bb_learn         15      276       217  39/105   37.1       761      100.3
tmh              15      294       220   46/105   43.8       753       96.0

pattern wins 48/105 = 45.7% of rounds, consistent with the ~49% measured baseline. bb_learn wins 39/105 = 37.1% — the fewest of all four arms. hitstk/run (hits taken) and shots/run are context only; per the standing rule nothing here is judged on hit rate.

Per-run values (never just the mean), from the analyzer:

pattern  dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321
         wins: r1=3 r2=2 r3=5 r4=1 r5=1 r6=4 r7=5 r8=3 r9=3 r10=4 r11=3 r12=4 r13=3 r14=3 r15=4
bb_id    dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292
         wins: r1=3 r2=4 r3=4 r4=3 r5=4 r6=3 r7=4 r8=5 r9=2 r10=4 r11=3 r12=2 r13=2 r14=4 r15=3
bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345
         wins: r1=2 r2=4 r3=1 r4=2 r5=3 r6=3 r7=3 r8=3 r9=1 r10=2 r11=4 r12=4 r13=1 r14=2 r15=4
tmh      dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267
         wins: r1=2 r2=4 r3=3 r4=3 r5=3 r6=2 r7=2 r8=3 r9=5 r10=4 r11=4 r12=3 r13=2 r14=4 r15=2

Every pairwise test the analyzer emits, with the Welch 95% CI for the same difference (sign is A - B, so negative means B is better):

metric      A           B           diff(A-B)   perm p   MW p
dmg/run     pattern     bb_id         +5.967    0.5968   0.7716
round wins  pattern     bb_id         -0.133    0.8647   0.8450
dmg/run     pattern     bb_learn     +25.894    0.0483   0.0620
round wins  pattern     bb_learn      +0.600    0.2197   0.1837
dmg/run     pattern     tmh           +8.448    0.4038   0.4306
round wins  pattern     tmh           +0.133    0.8671   0.6196
dmg/run     bb_id       bb_learn     +19.927    0.1386   0.1844
round wins  bb_id       bb_learn      +0.733    0.0853   0.0842
dmg/run     bb_id       tmh           +2.481    0.8174   0.7089
round wins  bb_id       tmh           +0.267    0.5574   0.4206
dmg/run     bb_learn    tmh          -17.446    0.1600   0.1585
round wins  bb_learn    tmh           -0.467    0.3019   0.3012
comparison (arm - ref) damage/run 95% CI wins/run 95% CI
bb_id - pattern -6.0 [-27.9, +15.9] +0.13 [-0.63, +0.90]
bb_learn - pattern -25.9 [-50.5, -1.3] -0.60 [-1.43, +0.23]
tmh - pattern -8.6 [-28.3, +11.1] -0.13 [-0.91, +0.65]
bb_id - tmh +2.6 [-18.4, +23.6] +0.27 [-0.40, +0.93]
bb_learn - tmh -17.3 [-41.0, +6.5] -0.47 [-1.21, +0.28]
bb_learn - bb_id -19.9 [-45.5, +5.8] -0.73 [-1.46, -0.01]

The single comparison that crosses alpha=0.05 is Pattern beating bb_learn on damage (p=0.048; MW p=0.062). Every BitBrain-vs-TMHorizon and BitBrain-vs-Pattern comparison is a null or a point estimate in the wrong direction for the claim. Round-level Fisher (anti-conservative; rounds cluster within runs): bb_id 50/105 vs pattern 48/105 p=0.89; bb_learn 39/105 p=0.26; tmh 46/105 p=0.89.

3. Minimum detectable effect and power [MEASURED]

Analyzer MDE (alpha=0.05 two-sided, 80% power):

metric       n/arm  sd(control)   MDE(abs)    MDE vs control mean
dmg/run         15       28.854     29.517          9.8% of 302.2
round wins      15        1.207      1.235           38.6% of 3.2

At 15 runs/arm this test resolves a 29.5 damage/run (9.8%) or 1.235 wins/run (17.6 pp) effect. The observed bb_id-vs-pattern difference (-6.0 damage) and the bb_id-vs-tmh difference (+2.6) are far inside that, so "no difference detectable" is the honest reading for BitBrain vs TMHorizon — not "no difference exists." For the claimed effect, post-hoc power: the measured sd(wins/run) is 1.207, so 6-4 (60% vs the ~49% baseline, i.e. +11 pp = +0.77 wins/run) needs ~39 runs/arm and +0.5 wins/run needs ~92 runs/arm. This 15-run/arm test is underpowered for a small positive BitBrain effect; what it is powered for is a large one, and it rules out the learned arm being better by ~1.2 wins/run.

4. Hit rate by range band [MEASURED]

BitBrain's only moving part is gated to range >= 300 px (BB_GAIN_BAND_MIN=3); its gain is pinned to 1.0 below that, so bands 0-300 are a second identity check and 300+ is where any effect must appear. Pooled hits/shots per arm (tools/ab/ab_range_bands.py):

band        bb_id(s/h)      bb_learn(s/h)   pattern(s/h)    tmh(s/h)
0-100            2/0  0.0%      3/2 66.7%      10/4 40.0%     11/5 45.5%
100-200        30/6 20.0%     46/11 23.9%      65/13 20.0%    62/14 22.6%
200-300      287/41 14.3%    303/50 16.5%     287/45 15.7%   298/51 17.1%
300-450    5708/711 12.5%  5673/659 11.6%   5679/753 13.3%  5726/647 11.3%
450+       5489/532  9.7%  5109/440  8.6%   5076/470  9.3%  4940/494 10.0%
ALL       11516/1290 11.2% 11134/1162 10.4% 11117/1285 11.6% 11037/1211 11.0%

Per-band permutation test on per-run band rates vs pattern:

band       arm          d(pp)         p method        MDE(pp)
300-450    bb_id        -0.72    0.1996 MC/B=200,000     1.46
300-450    bb_learn    -1.61    0.0035 MC/B=200,000     1.46
300-450    tmh         -1.88    0.0010 MC/B=200,000     1.46
450+       bb_id        +0.42    0.3927 MC/B=200,000     1.39
450+       bb_learn    -0.57    0.2987 MC/B=200,000     1.39
450+       tmh         +0.85    0.1511 MC/B=200,000     1.39
ALL        bb_learn    -1.13    0.0021 MC/B=200,000     0.74
ALL        tmh         -0.54    0.1397 MC/B=200,000     0.74

The identity arm bb_id is flat in both moving bands (p=0.20, 0.39) — the second identity check passes. The learned arm loses 1.61 pp of hit rate at 300-450 px (p=0.0035), exactly where its gain fires: because the candidate set is {1.0, 1.25, 1.5, 2.0} it can only over-lead, never shrink the lead, and the campaign ledger's Phase 1 measured the optimal long-range gain to be below 1.0. That is the mechanism behind the bb_learn damage loss.

5. Liveness [MEASURED]

pattern        OK   (15/15 runs: no arm env; report present)
bb_id          OK   (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied)
bb_learn       OK   (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied)
tmh            OK   (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied)

Round-win attribution cross-check: wins==firstPlaces 15/15 runs for every arm. (TR_BITBRAIN_LOG was deliberately not set, so the [bb] applied-shift section is empty; gains liveness is read from the boot env report instead.)

6. Verdict

Null result, honestly stated. On the shipped TFIL mover and real DrussGT, at 15 runs/arm: the BitBrain identity is statistically identical to Pattern (as predicted by d93ce44), and no BitBrain arm beats either TMHorizon or Pattern. The learned-gain config the owner most likely ran is the worst arm (276 damage/run, 39/105 round wins), significantly worse than Pattern on damage (p=0.048) and weakly worse than the identity on round wins (p=0.085), via an over-lead at 300-450 px. P(>=6 of 10) = 0.353 at the ~49% baseline means the 6-4 anecdote carries no information, and a 6-4-sized positive effect would need ~39 runs/arm to be seen at all.

MEASURED vs INFERRED

  • MEASURED: the 4-arm table (60 battles, 0 failed, frozen commit ed25ce2, binary 0fea927b…); the validity check; all pairwise permutation/MW p-values, MDEs and the Welch CIs; the per-band hit rates and permutation tests; the round-win attribution; liveness.
  • INFERRED: that the over-lead mechanism (candidates >= 1.0 + range gate) causes the bb_learn long-range loss (consistent with the ledger's Phase 1 measured gain curve, not proven by this session); that bb_learn is the exact config the owner ran; that the owner's "6-4" means round wins per the task framing.
  • NOT EXCLUDED: a small positive BitBrain effect below the MDE (~30 damage/run, ~1.2 wins/run). The test is underpowered for the claimed effect size; do not read this as "BitBrain is provably not better," read it as "no difference detectable at 15 runs/arm."