BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result)

4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
This commit is contained in:
2026-09-25 23:55:21 +02:00
parent ccff7e3e4a
commit 4829f9ca13
5 changed files with 405 additions and 0 deletions
+238
View File
@@ -0,0 +1,238 @@
# BitBrain vs TMHorizon vs Pattern — LIVE A/B (null result; the identity check passes)
**Question (the owner's claim, 2026-09-25).** *"Bitbrain gun is better than
TMhorizon, i won drussgt with it and tfil move 6 to 4."* This is the **second**
6-4 report (the earlier one was a TMHorizon+BitBrain mix). With the measured
baseline of **~49% round wins vs real DrussGT**, `P(>=6 of 10) = 0.353` — one
round in three by chance, so 6-4 is not evidence. This test replaces the
impression with a measurement on the **shipped TFIL default** (no `TR_MOVEMENT`
override), real DrussGT.
**Prior knowledge (not re-derived here).** At 30 runs/arm BitBrain was
statistically IDENTICAL to Pattern (`docs/bitbrain_gun_verdict.md`, `d93ce44`:
damage 283.7 vs 280.4, round wins **97/210 vs 97/210**, p=1.0, MDE 24.4 dmg),
and Pattern is in general >= TMHorizon. So "BitBrain > TMHorizon" is what those
facts *predict* — which is exactly why it needed a direct test.
> ## DIRECT ANSWERS (MEASURED, 15 runs/arm)
>
> * **Does BitBrain beat TMHorizon?** **No.** The identity config (`bb_id`) vs
> `tmh` is indistinguishable on both verdict metrics (+2.6 damage/run,
> p=0.82; +0.267 wins/run, p=0.56). The learned config (`bb_learn`) — the one
> the owner most likely ran — is *worse* than `tmh` in point estimate
> (-17.3 damage/run, p=0.16; -0.47 wins/run, p=0.30), though not significantly.
> * **Does BitBrain beat Pattern?** **No.** `bb_id` is statistically identical
> to Pattern (the validity check below). `bb_learn` is the **worst arm on both
> verdict metrics**: -25.9 damage/run vs Pattern (p=0.048, the only comparison
> in the whole table at alpha=0.05) and -0.60 wins/run (p=0.22).
> * **Does the learned-gain arm differ from the identity?** **Yes, weakly, in
> the wrong direction.** `bb_learn` vs `bb_id`: -19.9 damage/run (p=0.139) and
> **-0.733 wins/run (p=0.085**, 95% CI [-1.46, -0.01]). The learned gains
> (all >= 1.0, gated to range >= 300 px) over-lead and lose long-range hits.
> * **What this test CAN see:** an effect of ~**29.5 damage/run (9.8%)** or
> ~**1.235 wins/run (17.6 pp)** at 80% power. The claimed 6-4 (+11 pp, ~0.77
> wins/run) needs **~39 runs/arm** to resolve, so this test is **underpowered
> for the claimed effect size** — absence of a small *positive* BitBrain effect
> is not excluded. What it does exclude is a *large* one, and it does detect
> that the learned arm is ~10% worse than Pattern.
## Setup `[MEASURED]`
One frozen `ModularBot` from `git archive HEAD` at the run-time commit
`ed25ce29ad28838eaf726137e60c036e8ed70de3`, binary sha256
`0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93`; **4 arms ×
15 runs × 7 rounds = 60 battles, 420 rounds**, `--conc 7`, **0 failed**. (Another
job committed `ccff7e3` after this session started; that commit does not touch the
shipped TFIL default, and the frozen binary is pinned to `ed25ce2`.) Raw per-tick
captures live at `/tmp/ab/j113_bb_vs_tmh/` and are **not** committed; every table
below is in `common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt`,
`..._bands.txt`, `..._session.json`.
| arm | env | role |
|---|---|---|
| `pattern` | *(none)* | shipped default (`onlyPattern` rack) — reference |
| `bb_id` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0` | gain 1.0 identity — **validity check** |
| `bb_learn` | `... TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | the config the owner likely ran |
| `tmh` | `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only rack |
## 1. Validity check — `bb_id` vs `pattern` `[MEASURED]`
**PASS.** `TR_BITBRAIN_GAINS=1.0` is a single candidate, so the gun applies a
FIXED gain 1.0 (`aim = LOS + 1.0*(patternAim - LOS)`, no learning) in the long
bands — the identity to Pattern. It is statistically indistinguishable:
| comparison | diff (arm - ref) | 95% CI (Welch) | perm p | MW p |
|---|---:|---:|---:|---:|
| `bb_id` vs `pattern`, damage/run | **-6.0** | [-27.9, +15.9] | 0.597 | 0.772 |
| `bb_id` vs `pattern`, wins/run | **+0.133** | [-0.63, +0.90] | 0.865 | 0.845 |
Both differences are a small fraction of the MDE, and the round-level Fisher
test is 50/105 vs 48/105 (p=0.89). The rack swap (Pattern id 5 -> BitBrain id 16)
and the TmHorizon observation ring introduce no detectable systematic offset.
The remaining arms therefore sit on validated plumbing.
## 2. Primary result — damage/run and ROUND WINS `[MEASURED]`
Analyzer `ARM SUMMARY` (verbatim, `tools/ab/ab_analyze.py`):
```
arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run
--------------------------------------------------------------------------
pattern 15 302 238 48/105 45.7 761 99.4
bb_id 15 296 217 50/105 47.6 787 97.5
bb_learn 15 276 217 39/105 37.1 761 100.3
tmh 15 294 220 46/105 43.8 753 96.0
```
`pattern` wins 48/105 = 45.7% of rounds, consistent with the ~49% measured
baseline. `bb_learn` wins 39/105 = 37.1% — the fewest of all four arms.
`hitstk/run` (hits taken) and shots/run are context only; per the standing rule
nothing here is judged on hit rate.
Per-run values (never just the mean), from the analyzer:
```
pattern dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321
wins: r1=3 r2=2 r3=5 r4=1 r5=1 r6=4 r7=5 r8=3 r9=3 r10=4 r11=3 r12=4 r13=3 r14=3 r15=4
bb_id dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292
wins: r1=3 r2=4 r3=4 r4=3 r5=4 r6=3 r7=4 r8=5 r9=2 r10=4 r11=3 r12=2 r13=2 r14=4 r15=3
bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345
wins: r1=2 r2=4 r3=1 r4=2 r5=3 r6=3 r7=3 r8=3 r9=1 r10=2 r11=4 r12=4 r13=1 r14=2 r15=4
tmh dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267
wins: r1=2 r2=4 r3=3 r4=3 r5=3 r6=2 r7=2 r8=3 r9=5 r10=4 r11=4 r12=3 r13=2 r14=4 r15=2
```
Every pairwise test the analyzer emits, with the Welch 95% CI for the same
difference (sign is `A - B`, so negative means B is better):
```
metric A B diff(A-B) perm p MW p
dmg/run pattern bb_id +5.967 0.5968 0.7716
round wins pattern bb_id -0.133 0.8647 0.8450
dmg/run pattern bb_learn +25.894 0.0483 0.0620
round wins pattern bb_learn +0.600 0.2197 0.1837
dmg/run pattern tmh +8.448 0.4038 0.4306
round wins pattern tmh +0.133 0.8671 0.6196
dmg/run bb_id bb_learn +19.927 0.1386 0.1844
round wins bb_id bb_learn +0.733 0.0853 0.0842
dmg/run bb_id tmh +2.481 0.8174 0.7089
round wins bb_id tmh +0.267 0.5574 0.4206
dmg/run bb_learn tmh -17.446 0.1600 0.1585
round wins bb_learn tmh -0.467 0.3019 0.3012
```
| comparison (arm - ref) | damage/run | 95% CI | wins/run | 95% CI |
|---|---:|---:|---:|---:|
| `bb_id` - `pattern` | -6.0 | [-27.9, +15.9] | +0.13 | [-0.63, +0.90] |
| `bb_learn` - `pattern` | -25.9 | [-50.5, -1.3] | -0.60 | [-1.43, +0.23] |
| `tmh` - `pattern` | -8.6 | [-28.3, +11.1] | -0.13 | [-0.91, +0.65] |
| `bb_id` - `tmh` | +2.6 | [-18.4, +23.6] | +0.27 | [-0.40, +0.93] |
| `bb_learn` - `tmh` | -17.3 | [-41.0, +6.5] | -0.47 | [-1.21, +0.28] |
| `bb_learn` - `bb_id` | -19.9 | [-45.5, +5.8] | **-0.73** | **[-1.46, -0.01]** |
The single comparison that crosses alpha=0.05 is **Pattern beating `bb_learn` on
damage** (p=0.048; MW p=0.062). Every BitBrain-vs-TMHorizon and BitBrain-vs-Pattern
comparison is a null or a point estimate in the *wrong* direction for the claim.
Round-level Fisher (anti-conservative; rounds cluster within runs):
`bb_id` 50/105 vs `pattern` 48/105 p=0.89; `bb_learn` 39/105 p=0.26;
`tmh` 46/105 p=0.89.
## 3. Minimum detectable effect and power `[MEASURED]`
Analyzer `MDE` (alpha=0.05 two-sided, 80% power):
```
metric n/arm sd(control) MDE(abs) MDE vs control mean
dmg/run 15 28.854 29.517 9.8% of 302.2
round wins 15 1.207 1.235 38.6% of 3.2
```
At 15 runs/arm this test resolves a **29.5 damage/run (9.8%)** or **1.235
wins/run (17.6 pp)** effect. The observed `bb_id`-vs-`pattern` difference (-6.0
damage) and the `bb_id`-vs-`tmh` difference (+2.6) are far inside that, so
**"no difference detectable" is the honest reading for BitBrain vs TMHorizon —
not "no difference exists."** For the *claimed* effect, post-hoc power:
the measured `sd(wins/run)` is 1.207, so 6-4 (60% vs the ~49% baseline, i.e.
+11 pp = +0.77 wins/run) needs **~39 runs/arm** and +0.5 wins/run needs ~92
runs/arm. **This 15-run/arm test is underpowered for a small positive BitBrain
effect**; what it is powered for is a large one, and it rules out the learned
arm being better by ~1.2 wins/run.
## 4. Hit rate by range band `[MEASURED]`
BitBrain's only moving part is gated to range >= 300 px (`BB_GAIN_BAND_MIN=3`);
its gain is pinned to 1.0 below that, so bands 0-300 are a second identity check
and 300+ is where any effect must appear. Pooled hits/shots per arm
(`tools/ab/ab_range_bands.py`):
```
band bb_id(s/h) bb_learn(s/h) pattern(s/h) tmh(s/h)
0-100 2/0 0.0% 3/2 66.7% 10/4 40.0% 11/5 45.5%
100-200 30/6 20.0% 46/11 23.9% 65/13 20.0% 62/14 22.6%
200-300 287/41 14.3% 303/50 16.5% 287/45 15.7% 298/51 17.1%
300-450 5708/711 12.5% 5673/659 11.6% 5679/753 13.3% 5726/647 11.3%
450+ 5489/532 9.7% 5109/440 8.6% 5076/470 9.3% 4940/494 10.0%
ALL 11516/1290 11.2% 11134/1162 10.4% 11117/1285 11.6% 11037/1211 11.0%
```
Per-band permutation test on **per-run** band rates vs `pattern`:
```
band arm d(pp) p method MDE(pp)
300-450 bb_id -0.72 0.1996 MC/B=200,000 1.46
300-450 bb_learn -1.61 0.0035 MC/B=200,000 1.46
300-450 tmh -1.88 0.0010 MC/B=200,000 1.46
450+ bb_id +0.42 0.3927 MC/B=200,000 1.39
450+ bb_learn -0.57 0.2987 MC/B=200,000 1.39
450+ tmh +0.85 0.1511 MC/B=200,000 1.39
ALL bb_learn -1.13 0.0021 MC/B=200,000 0.74
ALL tmh -0.54 0.1397 MC/B=200,000 0.74
```
The identity arm `bb_id` is flat in both moving bands (p=0.20, 0.39) — the
second identity check passes. The learned arm **loses 1.61 pp of hit rate at
300-450 px (p=0.0035)**, exactly where its gain fires: because the candidate set
is `{1.0, 1.25, 1.5, 2.0}` it can only *over-lead*, never shrink the lead, and
the campaign ledger's Phase 1 measured the optimal long-range gain to be
*below* 1.0. That is the mechanism behind the `bb_learn` damage loss.
## 5. Liveness `[MEASURED]`
```
pattern OK (15/15 runs: no arm env; report present)
bb_id OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied)
bb_learn OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied)
tmh OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied)
```
Round-win attribution cross-check: `wins==firstPlaces` 15/15 runs for every arm.
(`TR_BITBRAIN_LOG` was deliberately not set, so the `[bb]` applied-shift section
is empty; gains liveness is read from the boot env report instead.)
## 6. Verdict
**Null result, honestly stated.** On the shipped TFIL mover and real DrussGT, at
**15 runs/arm**: the BitBrain identity is statistically identical to Pattern (as
predicted by `d93ce44`), and **no BitBrain arm beats either TMHorizon or
Pattern**. The learned-gain config the owner most likely ran is the **worst arm**
(276 damage/run, 39/105 round wins), significantly worse than Pattern on damage
(p=0.048) and weakly worse than the identity on round wins (p=0.085), via an
over-lead at 300-450 px. `P(>=6 of 10) = 0.353` at the ~49% baseline means the
6-4 anecdote carries no information, and a 6-4-sized *positive* effect would need
~39 runs/arm to be seen at all.
### MEASURED vs INFERRED
* **MEASURED:** the 4-arm table (60 battles, 0 failed, frozen commit `ed25ce2`,
binary `0fea927b…`); the validity check; all pairwise permutation/MW p-values,
MDEs and the Welch CIs; the per-band hit rates and permutation tests; the
round-win attribution; liveness.
* **INFERRED:** that the over-lead mechanism (`candidates >= 1.0` + range gate)
*causes* the `bb_learn` long-range loss (consistent with the ledger's Phase 1
measured gain curve, not proven by this session); that `bb_learn` is the exact
config the owner ran; that the owner's "6-4" means round wins per the task
framing.
* **NOT EXCLUDED:** a small positive BitBrain effect below the MDE (~30
damage/run, ~1.2 wins/run). The test is **underpowered** for the claimed
effect size; do not read this as "BitBrain is provably not better," read it as
"no difference detectable at 15 runs/arm."