From 4829f9ca13cf3fcd5a250c2a7c566e111d1afa8c Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Fri, 25 Sep 2026 23:55:21 +0200 Subject: [PATCH] BitBrain vs TMHorizon vs Pattern: live A/B on shipped TFIL (null result) 4 arms x 15 runs x 7 rounds (60 battles, 0 failed) vs real DrussGT on the shipped TFIL default, frozen at ed25ce2. bb_id (gain 1.0 identity) is statistically indistinguishable from shipped Pattern -> plumbing validity check passes. No BitBrain arm beats TMHorizon or Pattern: bb_learn (the config the owner likely ran) is the worst arm (276 dmg/run, 39/105 wins), the only comparison at alpha=0.05 is Pattern beating it on damage. Learned gains (>=1.0, gated >=300px) over-lead and lose 1.61pp of hit rate at 300-450px. MDE 29.5 dmg/run, 1.235 wins/run; a 6-4-sized effect needs ~39 runs/arm. --- .../bitbrain_vs_tmhorizon_ab_bands.txt | 66 +++++ .../bitbrain_vs_tmhorizon_ab_report.txt | 69 +++++ .../bitbrain_vs_tmhorizon_ab_session.json | 16 ++ docs/bitbrain_vs_tmhorizon_ab.md | 238 ++++++++++++++++++ tools/ab/arms_bitbrain_vs_tmh.txt | 16 ++ 5 files changed, 405 insertions(+) create mode 100644 common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_bands.txt create mode 100644 common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt create mode 100644 common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_session.json create mode 100644 docs/bitbrain_vs_tmhorizon_ab.md create mode 100644 tools/ab/arms_bitbrain_vs_tmh.txt diff --git a/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_bands.txt b/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_bands.txt new file mode 100644 index 0000000..6a4aff6 --- /dev/null +++ b/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_bands.txt @@ -0,0 +1,66 @@ +==================================================================================================== +HIT RATE BY RANGE BAND (our shots; band = shooter->target distance px at the fire tick) +session: /tmp/ab/j113_bb_vs_tmh arms: bb_id, bb_learn, pattern, tmh reference: pattern +==================================================================================================== +band bb_id bb_learn pattern tmh +------------------------------------------------------------------------------------------------------------------ +0-100 2 0 0.0% 3 2 66.7% 10 4 40.0% 11 5 45.5% +100-200 30 6 20.0% 46 11 23.9% 65 13 20.0% 62 14 22.6% +200-300 287 41 14.3% 303 50 16.5% 287 45 15.7% 298 51 17.1% +300-450 5708 711 12.5% 5673 659 11.6% 5679 753 13.3% 5726 647 11.3% +450+ 5489 532 9.7% 5109 440 8.6% 5076 470 9.3% 4940 494 10.0% +ALL 11516 1290 11.2% 11134 1162 10.4% 11117 1285 11.6% 11037 1211 11.0% + +PER-RUN BAND RATES (shows the spread behind the pooled numbers) + band 0-100 + bb_id n= 2 0 0 + bb_learn n= 2 50 100 + pattern n= 5 50 50 33 0 50 + tmh n= 6 0 0 100 100 40 50 + band 100-200 + bb_id n=14 0 0 0 50 0 0 20 0 50 25 0 50 0 100 + bb_learn n=13 0 0 17 0 0 50 50 27 0 33 50 40 33 + pattern n=13 25 0 20 14 40 20 33 0 14 0 0 29 20 + tmh n=13 0 0 0 20 25 0 50 0 17 67 29 0 25 + band 200-300 + bb_id n=15 0 20 22 10 14 14 15 10 19 11 25 15 6 19 9 + bb_learn n=15 4 14 16 0 0 14 25 20 9 19 14 39 35 11 18 + pattern n=15 29 0 17 12 19 15 17 10 33 10 0 13 24 12 10 + tmh n=15 11 32 30 12 0 14 12 26 13 14 14 17 18 10 36 + band 300-450 + bb_id n=15 11 13 13 12 11 16 13 12 12 12 10 13 14 13 13 + bb_learn n=15 10 11 11 14 11 13 13 11 11 11 10 13 11 13 10 + pattern n=15 13 14 13 16 12 13 15 11 12 15 15 13 14 11 13 + tmh n=15 10 12 11 12 12 9 10 12 14 11 12 14 10 10 11 + band 450+ + bb_id n=15 12 9 9 10 10 8 8 8 12 9 10 10 10 11 11 + bb_learn n=15 7 8 12 8 7 9 12 10 10 8 10 8 8 7 8 + pattern n=15 11 9 9 8 11 11 10 9 9 9 6 10 8 11 10 + tmh n=15 13 9 9 11 11 10 10 10 9 14 7 12 10 8 12 + band ALL + bb_id n=15 11 11 11 11 10 12 10 10 12 11 10 11 12 12 12 + bb_learn n=15 9 9 12 11 9 11 13 11 10 10 11 11 10 11 9 + pattern n=15 12 12 12 12 12 13 12 10 11 12 11 11 11 11 11 + tmh n=15 11 11 10 12 11 10 10 11 11 12 10 14 11 9 12 + +PER-BAND PERMUTATION TEST vs `pattern` (per-run rates, two-sided) +band arm d(pp) p method MCse MDE(pp) +------------------------------------------------------------------ +0-100 bb_id -36.67 0.1429 exact 0.0000 38.50 +0-100 bb_learn +38.33 0.2381 exact 0.0000 38.50 +0-100 tmh +11.67 0.6017 exact 0.0000 38.50 +100-200 bb_id +4.50 0.6598 MC/B=200,000 0.0011 14.84 +100-200 bb_learn +6.55 0.3520 exact 0.0000 14.84 +100-200 tmh +1.26 0.8610 exact 0.0000 14.84 +200-300 bb_id -0.73 0.8022 MC/B=200,000 0.0009 9.33 +200-300 bb_learn +1.13 0.7635 MC/B=200,000 0.0010 9.33 +200-300 tmh +2.59 0.4574 MC/B=200,000 0.0011 9.33 +300-450 bb_id -0.72 0.1996 MC/B=200,000 0.0009 1.46 +300-450 bb_learn -1.61 0.0035 MC/B=200,000 0.0001 1.46 +300-450 tmh -1.88 0.0010 MC/B=200,000 0.0001 1.46 +450+ bb_id +0.42 0.3927 MC/B=200,000 0.0011 1.39 +450+ bb_learn -0.57 0.2987 MC/B=200,000 0.0010 1.39 +450+ tmh +0.85 0.1511 MC/B=200,000 0.0008 1.39 +ALL bb_id -0.36 0.1921 MC/B=200,000 0.0009 0.74 +ALL bb_learn -1.13 0.0021 MC/B=200,000 0.0001 0.74 +ALL tmh -0.54 0.1397 MC/B=200,000 0.0008 0.74 diff --git a/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt b/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt new file mode 100644 index 0000000..23b2e39 --- /dev/null +++ b/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt @@ -0,0 +1,69 @@ +# session /tmp/ab/j113_bb_vs_tmh +# commit=ed25ce29ad28838eaf726137e60c036e8ed70de3 binary_sha256=0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93 rounds=7 runs=15 conc=7 ts=2026-09-25T23:41:38+02:00 + +ARM SUMMARY +arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run +-------------------------------------------------------------------------- +pattern 15 302 238 48/105 45.7 761 99.4 +bb_id 15 296 217 50/105 47.6 787 97.5 +bb_learn 15 276 217 39/105 37.1 761 100.3 +tmh 15 294 220 46/105 43.8 753 96.0 + +PER-RUN (never just the mean) + pattern dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321 + wins: r1=3/7 r2=2/7 r3=5/7 r4=1/7 r5=1/7 r6=4/7 r7=5/7 r8=3/7 r9=3/7 r10=4/7 r11=3/7 r12=4/7 r13=3/7 r14=3/7 r15=4/7 + bb_id dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292 + wins: r1=3/7 r2=4/7 r3=4/7 r4=3/7 r5=4/7 r6=3/7 r7=4/7 r8=5/7 r9=2/7 r10=4/7 r11=3/7 r12=2/7 r13=2/7 r14=4/7 r15=3/7 + bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345 + wins: r1=2/7 r2=4/7 r3=1/7 r4=2/7 r5=3/7 r6=3/7 r7=3/7 r8=3/7 r9=1/7 r10=2/7 r11=4/7 r12=4/7 r13=1/7 r14=2/7 r15=4/7 + tmh dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267 + wins: r1=2/7 r2=4/7 r3=3/7 r4=3/7 r5=3/7 r6=2/7 r7=2/7 r8=3/7 r9=5/7 r10=4/7 r11=4/7 r12=3/7 r13=2/7 r14=4/7 r15=2/7 + +PAIRWISE PERMUTATION TEST (per-run values) + MANN-WHITNEY CROSS-CHECK +permutation: exact when C(n,na) <= 20,000,000; otherwise Monte-Carlo 1,000,000 draws, seed=0x5eed5eed, p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1)) +metric A B diff(A-B) perm p method MC se MW p MW U +------------------------------------------------------------------------------------------------- +dmg/run pattern bb_id +5.967 0.5968 MC/B=1,000,000 0.0005 0.7716 105.0 +round wins pattern bb_id -0.133 0.8647 MC/B=1,000,000 0.0003 0.8450 107.5 +dmg/run pattern bb_learn +25.894 0.0483 MC/B=1,000,000 0.0002 0.0620 67.0 +round wins pattern bb_learn +0.600 0.2197 MC/B=1,000,000 0.0004 0.1837 81.0 +dmg/run pattern tmh +8.448 0.4038 MC/B=1,000,000 0.0005 0.4306 93.0 +round wins pattern tmh +0.133 0.8671 MC/B=1,000,000 0.0003 0.6196 100.5 +dmg/run bb_id bb_learn +19.927 0.1386 MC/B=1,000,000 0.0003 0.1844 80.0 +round wins bb_id bb_learn +0.733 0.0853 MC/B=1,000,000 0.0003 0.0842 72.0 +dmg/run bb_id tmh +2.481 0.8174 MC/B=1,000,000 0.0004 0.7089 103.0 +round wins bb_id tmh +0.267 0.5574 MC/B=1,000,000 0.0005 0.4206 93.5 +dmg/run bb_learn tmh -17.446 0.1600 MC/B=1,000,000 0.0004 0.1585 78.0 +round wins bb_learn tmh -0.467 0.3019 MC/B=1,000,000 0.0005 0.3012 88.0 + +MINIMUM DETECTABLE EFFECT (two-sample, alpha=0.05 two-sided, 80% power; MDE = 2.8016*sd*sqrt(2/n)) +metric n/arm sd(control) MDE(abs) MDE vs control mean +---------------------------------------------------------------- +dmg/run 15 28.854 29.517 9.8% of 302.2 +round wins 15 1.207 1.235 38.6% of 3.2 + +ROUND-LEVEL TEST (pooled rounds, Fisher exact) vs `pattern` — ANTI-CONSERVATIVE: rounds cluster within runs +arm ref wins arm wins p +---------------------------------------------- +bb_id 48/105 50/105 0.8900 +bb_learn 48/105 39/105 0.2624 +tmh 48/105 46/105 0.8897 + +LIVENESS (arm env applied in the bot's own boot report) + pattern OK (15/15 runs: no arm env; report present) + bb_id OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied) + bb_learn OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied) + tmh OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied) + +[bb] APPLIED-SHIFT CHECK (from bot stdout; needs TR_BITBRAIN_LOG=1). A provably-zero placebo emits ZERO [bb] lines. + arm runs w/log lines min max zeros + pattern 0/15 0 - - - + bb_id 0/15 0 - - - + bb_learn 0/15 0 - - - + tmh 0/15 0 - - - + +ROUND-WIN ATTRIBUTION (events primary; score tie-break for mutual-kill / timeout rounds) + pattern wins==firstPlaces 15/15 runs OK; single-death rounds agree with score 104/104 (1 tie-broken) + bb_id wins==firstPlaces 15/15 runs OK; single-death rounds agree with score 104/104 (1 tie-broken) + bb_learn wins==firstPlaces 15/15 runs OK; single-death rounds agree with score 105/105 (0 tie-broken) + tmh wins==firstPlaces 15/15 runs OK; single-death rounds agree with score 104/104 (1 tie-broken) diff --git a/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_session.json b/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_session.json new file mode 100644 index 0000000..d57c6bf --- /dev/null +++ b/common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_session.json @@ -0,0 +1,16 @@ +{ + "commit": "ed25ce29ad28838eaf726137e60c036e8ed70de3", + "binary_sha256": "0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93", + "binary": "frozen/ModularBot/ModularBot_bin", + "rounds": 7, + "runs": 15, + "conc": 7, + "timestamp": "2026-09-25T23:41:38+02:00", + "outdir": "/tmp/ab/j113_bb_vs_tmh", + "arms": [ + {"name": "pattern", "env": "", "label": ""}, + {"name": "bb_id", "env": "TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0", "label": "gain 1.0 identity (validity check)"}, + {"name": "bb_learn", "env": "TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay", "label": "learned gain"}, + {"name": "tmh", "env": "TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both", "label": "TMHorizon-only rack"} + ] +} diff --git a/docs/bitbrain_vs_tmhorizon_ab.md b/docs/bitbrain_vs_tmhorizon_ab.md new file mode 100644 index 0000000..0d5f44a --- /dev/null +++ b/docs/bitbrain_vs_tmhorizon_ab.md @@ -0,0 +1,238 @@ +# BitBrain vs TMHorizon vs Pattern — LIVE A/B (null result; the identity check passes) + +**Question (the owner's claim, 2026-09-25).** *"Bitbrain gun is better than +TMhorizon, i won drussgt with it and tfil move 6 to 4."* This is the **second** +6-4 report (the earlier one was a TMHorizon+BitBrain mix). With the measured +baseline of **~49% round wins vs real DrussGT**, `P(>=6 of 10) = 0.353` — one +round in three by chance, so 6-4 is not evidence. This test replaces the +impression with a measurement on the **shipped TFIL default** (no `TR_MOVEMENT` +override), real DrussGT. + +**Prior knowledge (not re-derived here).** At 30 runs/arm BitBrain was +statistically IDENTICAL to Pattern (`docs/bitbrain_gun_verdict.md`, `d93ce44`: +damage 283.7 vs 280.4, round wins **97/210 vs 97/210**, p=1.0, MDE 24.4 dmg), +and Pattern is in general >= TMHorizon. So "BitBrain > TMHorizon" is what those +facts *predict* — which is exactly why it needed a direct test. + +> ## DIRECT ANSWERS (MEASURED, 15 runs/arm) +> +> * **Does BitBrain beat TMHorizon?** **No.** The identity config (`bb_id`) vs +> `tmh` is indistinguishable on both verdict metrics (+2.6 damage/run, +> p=0.82; +0.267 wins/run, p=0.56). The learned config (`bb_learn`) — the one +> the owner most likely ran — is *worse* than `tmh` in point estimate +> (-17.3 damage/run, p=0.16; -0.47 wins/run, p=0.30), though not significantly. +> * **Does BitBrain beat Pattern?** **No.** `bb_id` is statistically identical +> to Pattern (the validity check below). `bb_learn` is the **worst arm on both +> verdict metrics**: -25.9 damage/run vs Pattern (p=0.048, the only comparison +> in the whole table at alpha=0.05) and -0.60 wins/run (p=0.22). +> * **Does the learned-gain arm differ from the identity?** **Yes, weakly, in +> the wrong direction.** `bb_learn` vs `bb_id`: -19.9 damage/run (p=0.139) and +> **-0.733 wins/run (p=0.085**, 95% CI [-1.46, -0.01]). The learned gains +> (all >= 1.0, gated to range >= 300 px) over-lead and lose long-range hits. +> * **What this test CAN see:** an effect of ~**29.5 damage/run (9.8%)** or +> ~**1.235 wins/run (17.6 pp)** at 80% power. The claimed 6-4 (+11 pp, ~0.77 +> wins/run) needs **~39 runs/arm** to resolve, so this test is **underpowered +> for the claimed effect size** — absence of a small *positive* BitBrain effect +> is not excluded. What it does exclude is a *large* one, and it does detect +> that the learned arm is ~10% worse than Pattern. + +## Setup `[MEASURED]` + +One frozen `ModularBot` from `git archive HEAD` at the run-time commit +`ed25ce29ad28838eaf726137e60c036e8ed70de3`, binary sha256 +`0fea927b5e0e17d1f82791bb620c08f1a460c074f3d21211bddc9c75f3bf1b93`; **4 arms × +15 runs × 7 rounds = 60 battles, 420 rounds**, `--conc 7`, **0 failed**. (Another +job committed `ccff7e3` after this session started; that commit does not touch the +shipped TFIL default, and the frozen binary is pinned to `ed25ce2`.) Raw per-tick +captures live at `/tmp/ab/j113_bb_vs_tmh/` and are **not** committed; every table +below is in `common_libs/tests/fixtures/bitbrain_vs_tmhorizon_ab_report.txt`, +`..._bands.txt`, `..._session.json`. + +| arm | env | role | +|---|---|---| +| `pattern` | *(none)* | shipped default (`onlyPattern` rack) — reference | +| `bb_id` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0` | gain 1.0 identity — **validity check** | +| `bb_learn` | `... TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | the config the owner likely ran | +| `tmh` | `TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both` | TMHorizon-only rack | + +## 1. Validity check — `bb_id` vs `pattern` `[MEASURED]` + +**PASS.** `TR_BITBRAIN_GAINS=1.0` is a single candidate, so the gun applies a +FIXED gain 1.0 (`aim = LOS + 1.0*(patternAim - LOS)`, no learning) in the long +bands — the identity to Pattern. It is statistically indistinguishable: + +| comparison | diff (arm - ref) | 95% CI (Welch) | perm p | MW p | +|---|---:|---:|---:|---:| +| `bb_id` vs `pattern`, damage/run | **-6.0** | [-27.9, +15.9] | 0.597 | 0.772 | +| `bb_id` vs `pattern`, wins/run | **+0.133** | [-0.63, +0.90] | 0.865 | 0.845 | + +Both differences are a small fraction of the MDE, and the round-level Fisher +test is 50/105 vs 48/105 (p=0.89). The rack swap (Pattern id 5 -> BitBrain id 16) +and the TmHorizon observation ring introduce no detectable systematic offset. +The remaining arms therefore sit on validated plumbing. + +## 2. Primary result — damage/run and ROUND WINS `[MEASURED]` + +Analyzer `ARM SUMMARY` (verbatim, `tools/ab/ab_analyze.py`): + +``` +arm runs dmg/run dmgtk/run wins win% shots/run hitstk/run +-------------------------------------------------------------------------- +pattern 15 302 238 48/105 45.7 761 99.4 +bb_id 15 296 217 50/105 47.6 787 97.5 +bb_learn 15 276 217 39/105 37.1 761 100.3 +tmh 15 294 220 46/105 43.8 753 96.0 +``` + +`pattern` wins 48/105 = 45.7% of rounds, consistent with the ~49% measured +baseline. `bb_learn` wins 39/105 = 37.1% — the fewest of all four arms. +`hitstk/run` (hits taken) and shots/run are context only; per the standing rule +nothing here is judged on hit rate. + +Per-run values (never just the mean), from the analyzer: + +``` +pattern dmg: r1=315 r2=262 r3=299 r4=284 r5=266 r6=274 r7=300 r8=280 r9=309 r10=358 r11=273 r12=338 r13=334 r14=320 r15=321 + wins: r1=3 r2=2 r3=5 r4=1 r5=1 r6=4 r7=5 r8=3 r9=3 r10=4 r11=3 r12=4 r13=3 r14=3 r15=4 +bb_id dmg: r1=294 r2=279 r3=320 r4=272 r5=328 r6=305 r7=338 r8=337 r9=299 r10=328 r11=281 r12=226 r13=247 r14=297 r15=292 + wins: r1=3 r2=4 r3=4 r4=3 r5=4 r6=3 r7=4 r8=5 r9=2 r10=4 r11=3 r12=2 r13=2 r14=4 r15=3 +bb_learn dmg: r1=204 r2=297 r3=284 r4=272 r5=257 r6=296 r7=311 r8=295 r9=237 r10=238 r11=320 r12=304 r13=225 r14=260 r15=345 + wins: r1=2 r2=4 r3=1 r4=2 r5=3 r6=3 r7=3 r8=3 r9=1 r10=2 r11=4 r12=4 r13=1 r14=2 r15=4 +tmh dmg: r1=279 r2=307 r3=258 r4=307 r5=262 r6=323 r7=267 r8=260 r9=320 r10=305 r11=318 r12=335 r13=306 r14=290 r15=267 + wins: r1=2 r2=4 r3=3 r4=3 r5=3 r6=2 r7=2 r8=3 r9=5 r10=4 r11=4 r12=3 r13=2 r14=4 r15=2 +``` + +Every pairwise test the analyzer emits, with the Welch 95% CI for the same +difference (sign is `A - B`, so negative means B is better): + +``` +metric A B diff(A-B) perm p MW p +dmg/run pattern bb_id +5.967 0.5968 0.7716 +round wins pattern bb_id -0.133 0.8647 0.8450 +dmg/run pattern bb_learn +25.894 0.0483 0.0620 +round wins pattern bb_learn +0.600 0.2197 0.1837 +dmg/run pattern tmh +8.448 0.4038 0.4306 +round wins pattern tmh +0.133 0.8671 0.6196 +dmg/run bb_id bb_learn +19.927 0.1386 0.1844 +round wins bb_id bb_learn +0.733 0.0853 0.0842 +dmg/run bb_id tmh +2.481 0.8174 0.7089 +round wins bb_id tmh +0.267 0.5574 0.4206 +dmg/run bb_learn tmh -17.446 0.1600 0.1585 +round wins bb_learn tmh -0.467 0.3019 0.3012 +``` + +| comparison (arm - ref) | damage/run | 95% CI | wins/run | 95% CI | +|---|---:|---:|---:|---:| +| `bb_id` - `pattern` | -6.0 | [-27.9, +15.9] | +0.13 | [-0.63, +0.90] | +| `bb_learn` - `pattern` | -25.9 | [-50.5, -1.3] | -0.60 | [-1.43, +0.23] | +| `tmh` - `pattern` | -8.6 | [-28.3, +11.1] | -0.13 | [-0.91, +0.65] | +| `bb_id` - `tmh` | +2.6 | [-18.4, +23.6] | +0.27 | [-0.40, +0.93] | +| `bb_learn` - `tmh` | -17.3 | [-41.0, +6.5] | -0.47 | [-1.21, +0.28] | +| `bb_learn` - `bb_id` | -19.9 | [-45.5, +5.8] | **-0.73** | **[-1.46, -0.01]** | + +The single comparison that crosses alpha=0.05 is **Pattern beating `bb_learn` on +damage** (p=0.048; MW p=0.062). Every BitBrain-vs-TMHorizon and BitBrain-vs-Pattern +comparison is a null or a point estimate in the *wrong* direction for the claim. +Round-level Fisher (anti-conservative; rounds cluster within runs): +`bb_id` 50/105 vs `pattern` 48/105 p=0.89; `bb_learn` 39/105 p=0.26; +`tmh` 46/105 p=0.89. + +## 3. Minimum detectable effect and power `[MEASURED]` + +Analyzer `MDE` (alpha=0.05 two-sided, 80% power): + +``` +metric n/arm sd(control) MDE(abs) MDE vs control mean +dmg/run 15 28.854 29.517 9.8% of 302.2 +round wins 15 1.207 1.235 38.6% of 3.2 +``` + +At 15 runs/arm this test resolves a **29.5 damage/run (9.8%)** or **1.235 +wins/run (17.6 pp)** effect. The observed `bb_id`-vs-`pattern` difference (-6.0 +damage) and the `bb_id`-vs-`tmh` difference (+2.6) are far inside that, so +**"no difference detectable" is the honest reading for BitBrain vs TMHorizon — +not "no difference exists."** For the *claimed* effect, post-hoc power: +the measured `sd(wins/run)` is 1.207, so 6-4 (60% vs the ~49% baseline, i.e. ++11 pp = +0.77 wins/run) needs **~39 runs/arm** and +0.5 wins/run needs ~92 +runs/arm. **This 15-run/arm test is underpowered for a small positive BitBrain +effect**; what it is powered for is a large one, and it rules out the learned +arm being better by ~1.2 wins/run. + +## 4. Hit rate by range band `[MEASURED]` + +BitBrain's only moving part is gated to range >= 300 px (`BB_GAIN_BAND_MIN=3`); +its gain is pinned to 1.0 below that, so bands 0-300 are a second identity check +and 300+ is where any effect must appear. Pooled hits/shots per arm +(`tools/ab/ab_range_bands.py`): + +``` +band bb_id(s/h) bb_learn(s/h) pattern(s/h) tmh(s/h) +0-100 2/0 0.0% 3/2 66.7% 10/4 40.0% 11/5 45.5% +100-200 30/6 20.0% 46/11 23.9% 65/13 20.0% 62/14 22.6% +200-300 287/41 14.3% 303/50 16.5% 287/45 15.7% 298/51 17.1% +300-450 5708/711 12.5% 5673/659 11.6% 5679/753 13.3% 5726/647 11.3% +450+ 5489/532 9.7% 5109/440 8.6% 5076/470 9.3% 4940/494 10.0% +ALL 11516/1290 11.2% 11134/1162 10.4% 11117/1285 11.6% 11037/1211 11.0% +``` + +Per-band permutation test on **per-run** band rates vs `pattern`: + +``` +band arm d(pp) p method MDE(pp) +300-450 bb_id -0.72 0.1996 MC/B=200,000 1.46 +300-450 bb_learn -1.61 0.0035 MC/B=200,000 1.46 +300-450 tmh -1.88 0.0010 MC/B=200,000 1.46 +450+ bb_id +0.42 0.3927 MC/B=200,000 1.39 +450+ bb_learn -0.57 0.2987 MC/B=200,000 1.39 +450+ tmh +0.85 0.1511 MC/B=200,000 1.39 +ALL bb_learn -1.13 0.0021 MC/B=200,000 0.74 +ALL tmh -0.54 0.1397 MC/B=200,000 0.74 +``` + +The identity arm `bb_id` is flat in both moving bands (p=0.20, 0.39) — the +second identity check passes. The learned arm **loses 1.61 pp of hit rate at +300-450 px (p=0.0035)**, exactly where its gain fires: because the candidate set +is `{1.0, 1.25, 1.5, 2.0}` it can only *over-lead*, never shrink the lead, and +the campaign ledger's Phase 1 measured the optimal long-range gain to be +*below* 1.0. That is the mechanism behind the `bb_learn` damage loss. + +## 5. Liveness `[MEASURED]` + +``` +pattern OK (15/15 runs: no arm env; report present) +bb_id OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 applied) +bb_learn OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay applied) +tmh OK (15/15 runs: TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both applied) +``` + +Round-win attribution cross-check: `wins==firstPlaces` 15/15 runs for every arm. +(`TR_BITBRAIN_LOG` was deliberately not set, so the `[bb]` applied-shift section +is empty; gains liveness is read from the boot env report instead.) + +## 6. Verdict + +**Null result, honestly stated.** On the shipped TFIL mover and real DrussGT, at +**15 runs/arm**: the BitBrain identity is statistically identical to Pattern (as +predicted by `d93ce44`), and **no BitBrain arm beats either TMHorizon or +Pattern**. The learned-gain config the owner most likely ran is the **worst arm** +(276 damage/run, 39/105 round wins), significantly worse than Pattern on damage +(p=0.048) and weakly worse than the identity on round wins (p=0.085), via an +over-lead at 300-450 px. `P(>=6 of 10) = 0.353` at the ~49% baseline means the +6-4 anecdote carries no information, and a 6-4-sized *positive* effect would need +~39 runs/arm to be seen at all. + +### MEASURED vs INFERRED + +* **MEASURED:** the 4-arm table (60 battles, 0 failed, frozen commit `ed25ce2`, + binary `0fea927b…`); the validity check; all pairwise permutation/MW p-values, + MDEs and the Welch CIs; the per-band hit rates and permutation tests; the + round-win attribution; liveness. +* **INFERRED:** that the over-lead mechanism (`candidates >= 1.0` + range gate) + *causes* the `bb_learn` long-range loss (consistent with the ledger's Phase 1 + measured gain curve, not proven by this session); that `bb_learn` is the exact + config the owner ran; that the owner's "6-4" means round wins per the task + framing. +* **NOT EXCLUDED:** a small positive BitBrain effect below the MDE (~30 + damage/run, ~1.2 wins/run). The test is **underpowered** for the claimed + effect size; do not read this as "BitBrain is provably not better," read it as + "no difference detectable at 15 runs/arm." diff --git a/tools/ab/arms_bitbrain_vs_tmh.txt b/tools/ab/arms_bitbrain_vs_tmh.txt new file mode 100644 index 0000000..93bfb83 --- /dev/null +++ b/tools/ab/arms_bitbrain_vs_tmh.txt @@ -0,0 +1,16 @@ +# BitBrain vs TMHorizon vs Pattern — live A/B, 4 arms x 15 runs x 7 rounds. +# +# Tests the user's claim "BitBrain gun is better than TMhorizon ... tfil move +# 6 to 4" on the SHIPPED TFIL movement (no TR_MOVEMENT override), real DrussGT. +# +# pattern = shipped default (onlyPattern rack), no env (reference) +# bb_id = Pattern off, BitBrain on, gain pinned to 1.0 (single candidate => +# FIXED gain, no learning) -> `aim = LOS + 1.0*(patternAim-LOS)` is +# the IDENTITY. Built-in PLUMBING VALIDITY CHECK: MUST match pattern. +# bb_learn = Pattern off, BitBrain on, learner over {1.0,1.25,1.5,2.0} with +# decay memory -> the configuration the user most likely ran. +# tmh = Pattern off, TMHorizon on (its own head), shipped gains. +pattern | +bb_id | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0 | gain 1.0 identity (validity check) +bb_learn | TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay | learned gain +tmh | TR_RACK_PATTERN=off TR_RACK_TMHORIZON=both | TMHorizon-only rack