Files
SirRoboGarage/docs/bitbrain_vs_tmhorizon_ab.md
T
SirStone 4829f9ca13 BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result)
4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the
shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is
statistically indistinguishable from shipped Pattern -> plumbing validity
check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config
the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only
comparison at alpha=0.05 is Pattern beating it on damage. Learned gains
(>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px.
MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm.
2026-09-25 23:55:21 +02:00

239 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# BitBrain vs TMHorizon vs Pattern — LIVE A/B (null result; the identity check passes)
**Question (the owner's claim, 2026-09-25).** *"Bitbrain gun is better than
TMhorizon, i won drussgt with it and tfil move 6 to 4."* This is the **second**
6-4 report (the earlier one was a TMHorizon+BitBrain mix). With the measured
baseline of **~49% round wins vs real DrussGT**, `P(>=6 of 10) = 0.353` — one
round in three by chance, so 6-4 is not evidence. This test replaces the
impression with a measurement on the **shipped TFIL default** (no `TR_MOVEMENT`
override), real DrussGT.
**Prior knowledge (not re-derived here).** At 30 runs/arm BitBrain was
statistically IDENTICAL to Pattern (`docs/bitbrain_gun_verdict.md`, `d93ce44`:
damage 283.7 vs 280.4, round wins **97/210 vs 97/210**, p=1.0, MDE 24.4 dmg),
and Pattern is in general >= TMHorizon. So "BitBrain > TMHorizon" is what those
facts *predict* — which is exactly why it needed a direct test.
> ## DIRECT ANSWERS (MEASURED, 15 runs/arm)
>
> * **Does BitBrain beat TMHorizon?** **No.** The identity config (`bb_id`) vs
> `tmh` is indistinguishable on both verdict metrics (+2.6 damage/run,
> p=0.82; +0.267 wins/run, p=0.56). The learned config (`bb_learn`) — the one
> the owner most likely ran — is *worse* than `tmh` in point estimate
> (-17.3 damage/run, p=0.16; -0.47 wins/run, p=0.30), though not significantly.
> * **Does BitBrain beat Pattern?** **No.** `bb_id` is statistically identical
> to Pattern (the validity check below). `bb_learn` is the **worst arm on both
> verdict metrics**: -25.9 damage/run vs Pattern (p=0.048, the only comparison
> in the whole table at alpha=0.05) and -0.60 wins/run (p=0.22).
> * **Does the learned-gain arm differ from the identity?** **Yes, weakly, in
> the wrong direction.** `bb_learn` vs `bb_id`: -19.9 damage/run (p=0.139) and
> **-0.733 wins/run (p=0.085**, 95% CI [-1.46, -0.01]). The learned gains
> (all >= 1.0, gated to range >= 300 px) over-lead and lose long-range hits.
> * **What this test CAN see:** an effect of ~**29.5 damage/run (9.8%)** or
> ~**1.235 wins/run (17.6 pp)** at 80% power. The claimed 6-4 (+11 pp, ~0.77
> wins/run) needs **~39 runs/arm** to resolve, so this test is **underpowered
> for the claimed effect size** — absence of a small *positive* BitBrain effect
> is not excluded. What it does exclude is a *large* one, and it does detect
> that the learned arm is ~10% worse than Pattern.
## Setup `[MEASURED]`
One frozen `ModularBot` from `git archive HEAD` at the run-time commit
`ed25ce29ad28838eaf726137e60c036e8ed70de3`, binary sha256
`0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93`; **4 arms ×
15 runs × 7 rounds = 60 battles, 420 rounds**, `--conc 7`, **0 failed**. (Another
job committed `ccff7e3` after this session started; that commit does not touch the
shipped TFIL default, and the frozen binary is pinned to `ed25ce2`.) Raw per-tick
captures live at `/tmp/ab/j113_bb_vs_tmh/` and are **not** committed; every table
below is in `common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt`,
`..._bands.txt`, `..._session.json`.
| arm | env | role |
|---|---|---|
| `pattern` | *(none)* | shipped default (`onlyPattern` rack) — reference |
| `bb_id` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0` | gain 1.0 identity — **validity check** |
| `bb_learn` | `... TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | the config the owner likely ran |
| `tmh` | `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only rack |
## 1. Validity check — `bb_id` vs `pattern` `[MEASURED]`
**PASS.** `TR_BITBRAIN_GAINS=1.0` is a single candidate, so the gun applies a
FIXED gain 1.0 (`aim = LOS + 1.0*(patternAim - LOS)`, no learning) in the long
bands — the identity to Pattern. It is statistically indistinguishable:
| comparison | diff (arm - ref) | 95% CI (Welch) | perm p | MW p |
|---|---:|---:|---:|---:|
| `bb_id` vs `pattern`, damage/run | **-6.0** | [-27.9, +15.9] | 0.597 | 0.772 |
| `bb_id` vs `pattern`, wins/run | **+0.133** | [-0.63, +0.90] | 0.865 | 0.845 |
Both differences are a small fraction of the MDE, and the round-level Fisher
test is 50/105 vs 48/105 (p=0.89). The rack swap (Pattern id 5 -> BitBrain id 16)
and the TmHorizon observation ring introduce no detectable systematic offset.
The remaining arms therefore sit on validated plumbing.
## 2. Primary result — damage/run and ROUND WINS `[MEASURED]`
Analyzer `ARM SUMMARY` (verbatim, `tools/ab/ab_analyze.py`):
```
arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run
--------------------------------------------------------------------------
pattern 15 302 238 48/105 45.7 761 99.4
bb_id 15 296 217 50/105 47.6 787 97.5
bb_learn 15 276 217 39/105 37.1 761 100.3
tmh 15 294 220 46/105 43.8 753 96.0
```
`pattern` wins 48/105 = 45.7% of rounds, consistent with the ~49% measured
baseline. `bb_learn` wins 39/105 = 37.1% — the fewest of all four arms.
`hitstk/run` (hits taken) and shots/run are context only; per the standing rule
nothing here is judged on hit rate.
Per-run values (never just the mean), from the analyzer:
```
pattern dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321
wins: r1=3 r2=2 r3=5 r4=1 r5=1 r6=4 r7=5 r8=3 r9=3 r10=4 r11=3 r12=4 r13=3 r14=3 r15=4
bb_id dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292
wins: r1=3 r2=4 r3=4 r4=3 r5=4 r6=3 r7=4 r8=5 r9=2 r10=4 r11=3 r12=2 r13=2 r14=4 r15=3
bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345
wins: r1=2 r2=4 r3=1 r4=2 r5=3 r6=3 r7=3 r8=3 r9=1 r10=2 r11=4 r12=4 r13=1 r14=2 r15=4
tmh dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267
wins: r1=2 r2=4 r3=3 r4=3 r5=3 r6=2 r7=2 r8=3 r9=5 r10=4 r11=4 r12=3 r13=2 r14=4 r15=2
```
Every pairwise test the analyzer emits, with the Welch 95% CI for the same
difference (sign is `A - B`, so negative means B is better):
```
metric A B diff(A-B) perm p MW p
dmg/run pattern bb_id +5.967 0.5968 0.7716
round wins pattern bb_id -0.133 0.8647 0.8450
dmg/run pattern bb_learn +25.894 0.0483 0.0620
round wins pattern bb_learn +0.600 0.2197 0.1837
dmg/run pattern tmh +8.448 0.4038 0.4306
round wins pattern tmh +0.133 0.8671 0.6196
dmg/run bb_id bb_learn +19.927 0.1386 0.1844
round wins bb_id bb_learn +0.733 0.0853 0.0842
dmg/run bb_id tmh +2.481 0.8174 0.7089
round wins bb_id tmh +0.267 0.5574 0.4206
dmg/run bb_learn tmh -17.446 0.1600 0.1585
round wins bb_learn tmh -0.467 0.3019 0.3012
```
| comparison (arm - ref) | damage/run | 95% CI | wins/run | 95% CI |
|---|---:|---:|---:|---:|
| `bb_id` - `pattern` | -6.0 | [-27.9, +15.9] | +0.13 | [-0.63, +0.90] |
| `bb_learn` - `pattern` | -25.9 | [-50.5, -1.3] | -0.60 | [-1.43, +0.23] |
| `tmh` - `pattern` | -8.6 | [-28.3, +11.1] | -0.13 | [-0.91, +0.65] |
| `bb_id` - `tmh` | +2.6 | [-18.4, +23.6] | +0.27 | [-0.40, +0.93] |
| `bb_learn` - `tmh` | -17.3 | [-41.0, +6.5] | -0.47 | [-1.21, +0.28] |
| `bb_learn` - `bb_id` | -19.9 | [-45.5, +5.8] | **-0.73** | **[-1.46, -0.01]** |
The single comparison that crosses alpha=0.05 is **Pattern beating `bb_learn` on
damage** (p=0.048; MW p=0.062). Every BitBrain-vs-TMHorizon and BitBrain-vs-Pattern
comparison is a null or a point estimate in the *wrong* direction for the claim.
Round-level Fisher (anti-conservative; rounds cluster within runs):
`bb_id` 50/105 vs `pattern` 48/105 p=0.89; `bb_learn` 39/105 p=0.26;
`tmh` 46/105 p=0.89.
## 3. Minimum detectable effect and power `[MEASURED]`
Analyzer `MDE` (alpha=0.05 two-sided, 80% power):
```
metric n/arm sd(control) MDE(abs) MDE vs control mean
dmg/run 15 28.854 29.517 9.8% of 302.2
round wins 15 1.207 1.235 38.6% of 3.2
```
At 15 runs/arm this test resolves a **29.5 damage/run (9.8%)** or **1.235
wins/run (17.6 pp)** effect. The observed `bb_id`-vs-`pattern` difference (-6.0
damage) and the `bb_id`-vs-`tmh` difference (+2.6) are far inside that, so
**"no difference detectable" is the honest reading for BitBrain vs TMHorizon —
not "no difference exists."** For the *claimed* effect, post-hoc power:
the measured `sd(wins/run)` is 1.207, so 6-4 (60% vs the ~49% baseline, i.e.
+11 pp = +0.77 wins/run) needs **~39 runs/arm** and +0.5 wins/run needs ~92
runs/arm. **This 15-run/arm test is underpowered for a small positive BitBrain
effect**; what it is powered for is a large one, and it rules out the learned
arm being better by ~1.2 wins/run.
## 4. Hit rate by range band `[MEASURED]`
BitBrain's only moving part is gated to range >= 300 px (`BB_GAIN_BAND_MIN=3`);
its gain is pinned to 1.0 below that, so bands 0-300 are a second identity check
and 300+ is where any effect must appear. Pooled hits/shots per arm
(`tools/ab/ab_range_bands.py`):
```
band bb_id(s/h) bb_learn(s/h) pattern(s/h) tmh(s/h)
0-100 2/0 0.0% 3/2 66.7% 10/4 40.0% 11/5 45.5%
100-200 30/6 20.0% 46/11 23.9% 65/13 20.0% 62/14 22.6%
200-300 287/41 14.3% 303/50 16.5% 287/45 15.7% 298/51 17.1%
300-450 5708/711 12.5% 5673/659 11.6% 5679/753 13.3% 5726/647 11.3%
450+ 5489/532 9.7% 5109/440 8.6% 5076/470 9.3% 4940/494 10.0%
ALL 11516/1290 11.2% 11134/1162 10.4% 11117/1285 11.6% 11037/1211 11.0%
```
Per-band permutation test on **per-run** band rates vs `pattern`:
```
band arm d(pp) p method MDE(pp)
300-450 bb_id -0.72 0.1996 MC/B=200,000 1.46
300-450 bb_learn -1.61 0.0035 MC/B=200,000 1.46
300-450 tmh -1.88 0.0010 MC/B=200,000 1.46
450+ bb_id +0.42 0.3927 MC/B=200,000 1.39
450+ bb_learn -0.57 0.2987 MC/B=200,000 1.39
450+ tmh +0.85 0.1511 MC/B=200,000 1.39
ALL bb_learn -1.13 0.0021 MC/B=200,000 0.74
ALL tmh -0.54 0.1397 MC/B=200,000 0.74
```
The identity arm `bb_id` is flat in both moving bands (p=0.20, 0.39) — the
second identity check passes. The learned arm **loses 1.61 pp of hit rate at
300-450 px (p=0.0035)**, exactly where its gain fires: because the candidate set
is `{1.0, 1.25, 1.5, 2.0}` it can only *over-lead*, never shrink the lead, and
the campaign ledger's Phase 1 measured the optimal long-range gain to be
*below* 1.0. That is the mechanism behind the `bb_learn` damage loss.
## 5. Liveness `[MEASURED]`
```
pattern OK (15/15 runs: no arm env; report present)
bb_id OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied)
bb_learn OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied)
tmh OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied)
```
Round-win attribution cross-check: `wins==firstPlaces` 15/15 runs for every arm.
(`TR_BITBRAIN_LOG` was deliberately not set, so the `[bb]` applied-shift section
is empty; gains liveness is read from the boot env report instead.)
## 6. Verdict
**Null result, honestly stated.** On the shipped TFIL mover and real DrussGT, at
**15 runs/arm**: the BitBrain identity is statistically identical to Pattern (as
predicted by `d93ce44`), and **no BitBrain arm beats either TMHorizon or
Pattern**. The learned-gain config the owner most likely ran is the **worst arm**
(276 damage/run, 39/105 round wins), significantly worse than Pattern on damage
(p=0.048) and weakly worse than the identity on round wins (p=0.085), via an
over-lead at 300-450 px. `P(>=6 of 10) = 0.353` at the ~49% baseline means the
6-4 anecdote carries no information, and a 6-4-sized *positive* effect would need
~39 runs/arm to be seen at all.
### MEASURED vs INFERRED
* **MEASURED:** the 4-arm table (60 battles, 0 failed, frozen commit `ed25ce2`,
binary `0fea927b…`); the validity check; all pairwise permutation/MW p-values,
MDEs and the Welch CIs; the per-band hit rates and permutation tests; the
round-win attribution; liveness.
* **INFERRED:** that the over-lead mechanism (`candidates >= 1.0` + range gate)
*causes* the `bb_learn` long-range loss (consistent with the ledger's Phase 1
measured gain curve, not proven by this session); that `bb_learn` is the exact
config the owner ran; that the owner's "6-4" means round wins per the task
framing.
* **NOT EXCLUDED:** a small positive BitBrain effect below the MDE (~30
damage/run, ~1.2 wins/run). The test is **underpowered** for the claimed
effect size; do not read this as "BitBrain is provably not better," read it as
"no difference detectable at 15 runs/arm."