# BitBrain vs TMHorizon vs Pattern — LIVE A/B (null result; the identity check passes) **Question (the owner's claim, 2026-09-25).** *"Bitbrain gun is better than TMhorizon, i won drussgt with it and tfil move 6 to 4."* This is the **second** 6-4 report (the earlier one was a TMHorizon+BitBrain mix). With the measured baseline of **~49% round wins vs real DrussGT**, `P(>=6 of 10) = 0.353` — one round in three by chance, so 6-4 is not evidence. This test replaces the impression with a measurement on the **shipped TFIL default** (no `TR_MOVEMENT` override), real DrussGT. **Prior knowledge (not re-derived here).** At 30 runs/arm BitBrain was statistically IDENTICAL to Pattern (`docs/bitbrain_gun_verdict.md`, `d93ce44`: damage 283.7 vs 280.4, round wins **97/210 vs 97/210**, p=1.0, MDE 24.4 dmg), and Pattern is in general >= TMHorizon. So "BitBrain > TMHorizon" is what those facts *predict* — which is exactly why it needed a direct test. > ## DIRECT ANSWERS (MEASURED, 15 runs/arm) > > * **Does BitBrain beat TMHorizon?** **No.** The identity config (`bb_id`) vs > `tmh` is indistinguishable on both verdict metrics (+2.6 damage/run, > p=0.82; +0.267 wins/run, p=0.56). The learned config (`bb_learn`) — the one > the owner most likely ran — is *worse* than `tmh` in point estimate > (-17.3 damage/run, p=0.16; -0.47 wins/run, p=0.30), though not significantly. > * **Does BitBrain beat Pattern?** **No.** `bb_id` is statistically identical > to Pattern (the validity check below). `bb_learn` is the **worst arm on both > verdict metrics**: -25.9 damage/run vs Pattern (p=0.048, the only comparison > in the whole table at alpha=0.05) and -0.60 wins/run (p=0.22). > * **Does the learned-gain arm differ from the identity?** **Yes, weakly, in > the wrong direction.** `bb_learn` vs `bb_id`: -19.9 damage/run (p=0.139) and > **-0.733 wins/run (p=0.085**, 95% CI [-1.46, -0.01]). The learned gains > (all >= 1.0, gated to range >= 300 px) over-lead and lose long-range hits. > * **What this test CAN see:** an effect of ~**29.5 damage/run (9.8%)** or > ~**1.235 wins/run (17.6 pp)** at 80% power. The claimed 6-4 (+11 pp, ~0.77 > wins/run) needs **~39 runs/arm** to resolve, so this test is **underpowered > for the claimed effect size** — absence of a small *positive* BitBrain effect > is not excluded. What it does exclude is a *large* one, and it does detect > that the learned arm is ~10% worse than Pattern. ## Setup `[MEASURED]` One frozen `ModularBot` from `git archive HEAD` at the run-time commit `ed25ce29ad28838eaf726137e60c036e8ed70de3`, binary sha256 `0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93`; **4 arms × 15 runs × 7 rounds = 60 battles, 420 rounds**, `--conc 7`, **0 failed**. (Another job committed `ccff7e3` after this session started; that commit does not touch the shipped TFIL default, and the frozen binary is pinned to `ed25ce2`.) Raw per-tick captures live at `/tmp/ab/j113_bb_vs_tmh/` and are **not** committed; every table below is in `common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt`, `..._bands.txt`, `..._session.json`. | arm | env | role | |---|---|---| | `pattern` | *(none)* | shipped default (`onlyPattern` rack) — reference | | `bb_id` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0` | gain 1.0 identity — **validity check** | | `bb_learn` | `... TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | the config the owner likely ran | | `tmh` | `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only rack | ## 1. Validity check — `bb_id` vs `pattern` `[MEASURED]` **PASS.** `TR_BITBRAIN_GAINS=1.0` is a single candidate, so the gun applies a FIXED gain 1.0 (`aim = LOS + 1.0*(patternAim - LOS)`, no learning) in the long bands — the identity to Pattern. It is statistically indistinguishable: | comparison | diff (arm - ref) | 95% CI (Welch) | perm p | MW p | |---|---:|---:|---:|---:| | `bb_id` vs `pattern`, damage/run | **-6.0** | [-27.9, +15.9] | 0.597 | 0.772 | | `bb_id` vs `pattern`, wins/run | **+0.133** | [-0.63, +0.90] | 0.865 | 0.845 | Both differences are a small fraction of the MDE, and the round-level Fisher test is 50/105 vs 48/105 (p=0.89). The rack swap (Pattern id 5 -> BitBrain id 16) and the TmHorizon observation ring introduce no detectable systematic offset. The remaining arms therefore sit on validated plumbing. ## 2. Primary result — damage/run and ROUND WINS `[MEASURED]` Analyzer `ARM SUMMARY` (verbatim, `tools/ab/ab_analyze.py`): ``` arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run -------------------------------------------------------------------------- pattern 15 302 238 48/105 45.7 761 99.4 bb_id 15 296 217 50/105 47.6 787 97.5 bb_learn 15 276 217 39/105 37.1 761 100.3 tmh 15 294 220 46/105 43.8 753 96.0 ``` `pattern` wins 48/105 = 45.7% of rounds, consistent with the ~49% measured baseline. `bb_learn` wins 39/105 = 37.1% — the fewest of all four arms. `hitstk/run` (hits taken) and shots/run are context only; per the standing rule nothing here is judged on hit rate. Per-run values (never just the mean), from the analyzer: ``` pattern dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321 wins: r1=3 r2=2 r3=5 r4=1 r5=1 r6=4 r7=5 r8=3 r9=3 r10=4 r11=3 r12=4 r13=3 r14=3 r15=4 bb_id dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292 wins: r1=3 r2=4 r3=4 r4=3 r5=4 r6=3 r7=4 r8=5 r9=2 r10=4 r11=3 r12=2 r13=2 r14=4 r15=3 bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345 wins: r1=2 r2=4 r3=1 r4=2 r5=3 r6=3 r7=3 r8=3 r9=1 r10=2 r11=4 r12=4 r13=1 r14=2 r15=4 tmh dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267 wins: r1=2 r2=4 r3=3 r4=3 r5=3 r6=2 r7=2 r8=3 r9=5 r10=4 r11=4 r12=3 r13=2 r14=4 r15=2 ``` Every pairwise test the analyzer emits, with the Welch 95% CI for the same difference (sign is `A - B`, so negative means B is better): ``` metric A B diff(A-B) perm p MW p dmg/run pattern bb_id +5.967 0.5968 0.7716 round wins pattern bb_id -0.133 0.8647 0.8450 dmg/run pattern bb_learn +25.894 0.0483 0.0620 round wins pattern bb_learn +0.600 0.2197 0.1837 dmg/run pattern tmh +8.448 0.4038 0.4306 round wins pattern tmh +0.133 0.8671 0.6196 dmg/run bb_id bb_learn +19.927 0.1386 0.1844 round wins bb_id bb_learn +0.733 0.0853 0.0842 dmg/run bb_id tmh +2.481 0.8174 0.7089 round wins bb_id tmh +0.267 0.5574 0.4206 dmg/run bb_learn tmh -17.446 0.1600 0.1585 round wins bb_learn tmh -0.467 0.3019 0.3012 ``` | comparison (arm - ref) | damage/run | 95% CI | wins/run | 95% CI | |---|---:|---:|---:|---:| | `bb_id` - `pattern` | -6.0 | [-27.9, +15.9] | +0.13 | [-0.63, +0.90] | | `bb_learn` - `pattern` | -25.9 | [-50.5, -1.3] | -0.60 | [-1.43, +0.23] | | `tmh` - `pattern` | -8.6 | [-28.3, +11.1] | -0.13 | [-0.91, +0.65] | | `bb_id` - `tmh` | +2.6 | [-18.4, +23.6] | +0.27 | [-0.40, +0.93] | | `bb_learn` - `tmh` | -17.3 | [-41.0, +6.5] | -0.47 | [-1.21, +0.28] | | `bb_learn` - `bb_id` | -19.9 | [-45.5, +5.8] | **-0.73** | **[-1.46, -0.01]** | The single comparison that crosses alpha=0.05 is **Pattern beating `bb_learn` on damage** (p=0.048; MW p=0.062). Every BitBrain-vs-TMHorizon and BitBrain-vs-Pattern comparison is a null or a point estimate in the *wrong* direction for the claim. Round-level Fisher (anti-conservative; rounds cluster within runs): `bb_id` 50/105 vs `pattern` 48/105 p=0.89; `bb_learn` 39/105 p=0.26; `tmh` 46/105 p=0.89. ## 3. Minimum detectable effect and power `[MEASURED]` Analyzer `MDE` (alpha=0.05 two-sided, 80% power): ``` metric n/arm sd(control) MDE(abs) MDE vs control mean dmg/run 15 28.854 29.517 9.8% of 302.2 round wins 15 1.207 1.235 38.6% of 3.2 ``` At 15 runs/arm this test resolves a **29.5 damage/run (9.8%)** or **1.235 wins/run (17.6 pp)** effect. The observed `bb_id`-vs-`pattern` difference (-6.0 damage) and the `bb_id`-vs-`tmh` difference (+2.6) are far inside that, so **"no difference detectable" is the honest reading for BitBrain vs TMHorizon — not "no difference exists."** For the *claimed* effect, post-hoc power: the measured `sd(wins/run)` is 1.207, so 6-4 (60% vs the ~49% baseline, i.e. +11 pp = +0.77 wins/run) needs **~39 runs/arm** and +0.5 wins/run needs ~92 runs/arm. **This 15-run/arm test is underpowered for a small positive BitBrain effect**; what it is powered for is a large one, and it rules out the learned arm being better by ~1.2 wins/run. ## 4. Hit rate by range band `[MEASURED]` BitBrain's only moving part is gated to range >= 300 px (`BB_GAIN_BAND_MIN=3`); its gain is pinned to 1.0 below that, so bands 0-300 are a second identity check and 300+ is where any effect must appear. Pooled hits/shots per arm (`tools/ab/ab_range_bands.py`): ``` band bb_id(s/h) bb_learn(s/h) pattern(s/h) tmh(s/h) 0-100 2/0 0.0% 3/2 66.7% 10/4 40.0% 11/5 45.5% 100-200 30/6 20.0% 46/11 23.9% 65/13 20.0% 62/14 22.6% 200-300 287/41 14.3% 303/50 16.5% 287/45 15.7% 298/51 17.1% 300-450 5708/711 12.5% 5673/659 11.6% 5679/753 13.3% 5726/647 11.3% 450+ 5489/532 9.7% 5109/440 8.6% 5076/470 9.3% 4940/494 10.0% ALL 11516/1290 11.2% 11134/1162 10.4% 11117/1285 11.6% 11037/1211 11.0% ``` Per-band permutation test on **per-run** band rates vs `pattern`: ``` band arm d(pp) p method MDE(pp) 300-450 bb_id -0.72 0.1996 MC/B=200,000 1.46 300-450 bb_learn -1.61 0.0035 MC/B=200,000 1.46 300-450 tmh -1.88 0.0010 MC/B=200,000 1.46 450+ bb_id +0.42 0.3927 MC/B=200,000 1.39 450+ bb_learn -0.57 0.2987 MC/B=200,000 1.39 450+ tmh +0.85 0.1511 MC/B=200,000 1.39 ALL bb_learn -1.13 0.0021 MC/B=200,000 0.74 ALL tmh -0.54 0.1397 MC/B=200,000 0.74 ``` The identity arm `bb_id` is flat in both moving bands (p=0.20, 0.39) — the second identity check passes. The learned arm **loses 1.61 pp of hit rate at 300-450 px (p=0.0035)**, exactly where its gain fires: because the candidate set is `{1.0, 1.25, 1.5, 2.0}` it can only *over-lead*, never shrink the lead, and the campaign ledger's Phase 1 measured the optimal long-range gain to be *below* 1.0. That is the mechanism behind the `bb_learn` damage loss. ## 5. Liveness `[MEASURED]` ``` pattern OK (15/15 runs: no arm env; report present) bb_id OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied) bb_learn OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied) tmh OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied) ``` Round-win attribution cross-check: `wins==firstPlaces` 15/15 runs for every arm. (`TR_BITBRAIN_LOG` was deliberately not set, so the `[bb]` applied-shift section is empty; gains liveness is read from the boot env report instead.) ## 6. Verdict **Null result, honestly stated.** On the shipped TFIL mover and real DrussGT, at **15 runs/arm**: the BitBrain identity is statistically identical to Pattern (as predicted by `d93ce44`), and **no BitBrain arm beats either TMHorizon or Pattern**. The learned-gain config the owner most likely ran is the **worst arm** (276 damage/run, 39/105 round wins), significantly worse than Pattern on damage (p=0.048) and weakly worse than the identity on round wins (p=0.085), via an over-lead at 300-450 px. `P(>=6 of 10) = 0.353` at the ~49% baseline means the 6-4 anecdote carries no information, and a 6-4-sized *positive* effect would need ~39 runs/arm to be seen at all. ### MEASURED vs INFERRED * **MEASURED:** the 4-arm table (60 battles, 0 failed, frozen commit `ed25ce2`, binary `0fea927b…`); the validity check; all pairwise permutation/MW p-values, MDEs and the Welch CIs; the per-band hit rates and permutation tests; the round-win attribution; liveness. * **INFERRED:** that the over-lead mechanism (`candidates >= 1.0` + range gate) *causes* the `bb_learn` long-range loss (consistent with the ledger's Phase 1 measured gain curve, not proven by this session); that `bb_learn` is the exact config the owner ran; that the owner's "6-4" means round wins per the task framing. * **NOT EXCLUDED:** a small positive BitBrain effect below the MDE (~30 damage/run, ~1.2 wins/run). The test is **underpowered** for the claimed effect size; do not read this as "BitBrain is provably not better," read it as "no difference detectable at 15 runs/arm."