Files
SirRoboGarage/docs/gauntlet_bitbrain_vs_pattern.md
T
SirStone 8efa627c05 gauntlet: BitBrain vs Pattern across 32 legacy opponents (does not generalize)
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh
(per-opponent A/B over the legacy roster, subject = frozen ModularBot),
tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign
test, style split, MDE) and the arm/opponent fixtures.

Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x
3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run
214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480,
p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only
penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not
supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the
measured movement predictability does not correlate with the delta.
2026-09-26 00:55:57 +02:00

221 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# BitBrain vs Pattern across 32 legacy opponents — the generalization gauntlet
**This is the repo's first multi-opponent gun measurement.** Every previous
gun/movement claim in this project was measured against the single opponent
DrussGT; that caveat was flagged repeatedly and never closed until now. j114's
28 validated legacy champions (`tools/robocode_shim/robots.json`) made a
per-adversary gauntlet possible, so this test asks whether the DrussGT-only
conclusions *generalize*.
The claim under test is the bot owner's, from watching the GUI (2026-09-25):
*"BitBrain is a killer gun, especially against regular movements (spinners,
wall-followers), and its fast adaptation is the reason."* The counter-evidence
was 30 runs/arm vs DrussGT where BitBrain and Pattern were **identical**
(97/210 vs 97/210 round wins) and the learned-gain config was the worst arm
(`docs/bitbrain_vs_tmhorizon_ab.md`, `docs/bitbrain_gun_verdict.md`).
> ## DIRECT ANSWERS (MEASURED, 32 opponents × 2 arms × 3 runs × 5 rounds)
>
> * **Does BitBrain generalize beyond DrussGT?** **No.** Pooled over 32
> opponents the two arms are indistinguishable: damage/run **214.5 vs 210.9**
> (`bb - pattern` = **-3.6** ± 5.6, sign-flip p=0.53) and round wins
> **237/480 = 49.4% vs 239/480 = 49.8%** (delta **+0.02** wins/run, p=0.91).
> The cross-opponent sign test favours neither arm: BitBrain wins on **13/32**
> opponents on damage (Pattern on 19) and on **10** opponents on round wins
> (Pattern 9, 13 ties). **The DrussGT-specific penalty does not carry** — on
> DrussGT alone BitBrain was -18.1 damage/run, but across the field the
> estimate collapses to ~0. BitBrain is a wash, not a killer.
> * **Is it specifically stronger against regular/periodic movers?** **No.**
> On the 6 inferred regular movers (wall-followers, campers, rammers,
> flood-fill) the paired delta is **+1.3 damage/run** (SD 24.4); on the 15
> inferred dodgers it is **-0.2** (SD 33.6). Regular-vs-dodger Mann-Whitney
> **p=0.85** (damage) / **p=0.19** (wins). The measured movement predictability
> does not correlate with the delta either (Spearman +0.18, p=0.32). The
> owner's observation is **not** supported: the tournament's MDE at 32
> opponents is **15.7 damage/run** and **0.26 wins/run**, and the point
> estimate sits at zero, so an effect of the claimed (visible) size would have
> shown up.
> * **Caveat that limits the sub-claim:** the legacy roster has **no true
> constant-turn spinner** — the "regular" bucket is wall-followers / corner
> campers / a rammer / a flood-filler. The spinner-specific half of the claim
> is therefore *untested*, not refuted. What is refuted is the general
> "killer vs regular movers" reading.
## Setup `[MEASURED]`
One frozen `ModularBot` built from `git archive HEAD` at run-time commit
`da4a971ca99e81894e86a35a521bbef9ba081d3f` (binary sha256
`f5876e7e8c14389de8f335b1529868da80d987cea5ca298abc058ecaff1925cf`). **33
opponents attempted, 32 retained, 2 arms, 3 runs × 5 rounds each =
192 battles / 960 rounds**, `--conc 4`, **0 failed, 0 retries needed**. Raw
per-tick captures live at `/tmp/ab/j117_gauntlet/` and are **not** committed;
the full machine report is
`common_libs/tests/fixtures/gauntlet_bitbrain_vs_pattern_report.txt` and the
session header (arms, opponents, styles) is
`..._session.json`.
| arm | env | role |
|---|---|---|
| `pattern` | *(none)* | shipped default (`onlyPattern` rack) — reference |
| `bb` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, the learned-gain config the owner most likely ran |
Runner `tools/ab/gauntlet_run.sh` (per-opponent A/B; subject = the frozen
ModularBot itself, adversary iterates over the roster), arm list
`tools/ab/arms_gauntlet.txt`, opponents + inferred movement styles
`tools/ab/opponents_gauntlet.txt`, analyzer `tools/ab/gauntlet_analyze.py`.
## Validity / false-negative guard `[MEASURED]`
Per j114's warning, `run_smoke.sh` has false negatives and only a real battle is
authoritative, so each battle is validated and (if the bot failed to connect
inside the booter's 30 s window) retried up to 4 times. Results:
* **0 of 192 battles needed a retry**; every *retained* opponent fired and was
scored (the roster also includes the 5 `WORKS_WEAK` bots — valid movement
targets — of which only Aurora turned out inert).
* The least-active retained opponent fired **136** times across both arms; the
one inert bot (Aurora — a single fire in 6 battles) was **dropped**.
* Every declared arm env var reached the process, verified from each bot's own
`[env]` boot report: `TR_RACK_BITBRAIN = both`, `TR_RACK_PATTERN = off`,
`TR_BITBRAIN_GAINS = 1.0,1.25,1.5,2.0`, `TR_BITBRAIN_MEM = decay`, and the
active 1v1 rack line reads `BITBRAIN` for `bb` and `PATTERN` for `pattern`.
## 1. Per-opponent result `[MEASURED]`
The unit of evidence is the **number of opponents**, so the design spends its
budget on breadth (3 runs/opponent) rather than depth. Full table in the fixture;
the deltas that matter:
| opponent | style (INFERRED) | Δ dmg/run | Δ wins/run |
|---|---|---:|---:|
| Cigaret | dodger | +68.1 | +1.00 |
| WallAvoider | regular | +44.7 | +0.33 |
| Coriantumr | other | +44.5 | +0.33 |
| Komarious | dodger | +41.9 | +0.67 |
| CigaretBH | dodger | +41.4 | +1.00 |
| WaveSurferPG | dodger | +23.1 | -0.33 |
| DiamondHawk | other | +18.0 | +0.67 |
| Jen | dodger | +8.0 | 0.00 |
| FloodMini | regular | +6.7 | 0.00 |
| Dookious | dodger | +5.5 | 0.00 |
| PatternRobot | regular | +4.6 | 0.00 |
| LionWWSVMvoid | dodger | +4.5 | 0.00 |
| YersiniaPestis | other | +3.8 | -0.33 |
| KurtWaveSurfer | dodger | -3.1 | -1.00 |
| TripHammer | other | -3.4 | 0.00 |
| LightningBug | other | -6.0 | 0.00 |
| CassiusClay | dodger | -6.1 | +0.33 |
| Diamond | dodger | -8.0 | +0.33 |
| DiamondStealer | regular | -8.8 | 0.00 |
| Shadow | other | -9.9 | 0.00 |
| Ascendant | other | -11.5 | 0.00 |
| BlitzBat | regular | -13.6 | -0.33 |
| BrokenSword | other | -15.5 | 0.00 |
| DrussGT | other | -18.1 | -0.67 |
| HawkOnFire | regular | -25.9 | -1.00 |
| RougeDC | dodger | -26.8 | 0.00 |
| GresSuffurd | dodger | -29.1 | +1.00 |
| Aristocles | dodger | -29.3 | 0.00 |
| Phoenix | other | -29.4 | -0.33 |
| Lukious | dodger | -37.1 | +0.33 |
| WaveSurferGF | dodger | -55.9 | -0.33 |
| RetroGirl | other | -93.1 | -1.00 |
The individual deltas swing from -93 to +68 damage/run — with only 3 runs per
opponent, a single battle dominates each cell, so **no individual row is
evidence on its own.** The evidence is the aggregate below.
## 2. Cross-opponent sign test — the headline `[MEASURED]`
On how many opponents does each arm win (paired `bb - pattern`)?
```
metric bb better pattern better tie two-sided sign-test p
damage/run 13 19 0 0.377
wins/run 10 9 13 1.000
```
Neither arm wins on the majority of opponents. On round wins the two arms are
dead even (10 vs 9, 13 exact ties), which is the round-level echo of the
30-run-vs-DrussGT finding (97/210 vs 97/210).
## 3. Pooled paired estimate with between-opponent spread `[MEASURED]`
Per-opponent paired deltas, so one outlier bot cannot carry the result:
```
metric mean Δ between-opp SD SE 95% CI sign-flip p MDE(n=32)
damage/run -3.62 31.65 5.60 [-14.59, +7.35] 0.527 15.68
wins/run +0.02 0.52 0.09 [-0.16, +0.20] 0.911 0.258
```
Pooled arm totals over the 32 opponents:
```
arm damage/run damage taken/run round wins
pattern 214.5 291.5 237/480 = 49.4%
bb 210.9 296.4 239/480 = 49.8%
```
Both verdict metrics (damage/run and round wins, per the standing rule) say the
same thing: **no detectable difference.** The MDE at 32 opponents is 15.7
damage/run (≈7% of the ~214 baseline) and 0.258 wins/run (≈5.2 pp); the point
estimates are ~0. So a *large* BitBrain advantage is excluded, and the observed
effect is at the null. The usual caveat stands: a genuinely small positive
effect (<~8%) is **not** excluded by this design — but there is no evidence of
one.
## 4. Style split and measured movement character `[MEASURED]`
Style labels are **INFERRED** from the bot names/docs in `robots.json`. To
check them against data, the opponent's own captured trajectory (`sx,sy,sh,ss`)
was characterised by its single most-common turn magnitude (`mode%`): a
straight-liner, a spinner and a fixed oscillator all concentrate there, an
adaptive surfer does not.
```
style n mean Δdmg/run (SD) mean Δwins/run (SD) measured mode%
regular 6 +1.3 (24.4) -0.17 (0.46) 0.58
dodger 15 -0.2 (33.6) +0.20 (0.56) 0.65
other 11 -11.0 (33.7) -0.12 (0.45) 0.59
regular vs dodger: Δdmg/run Mann-Whitney p=0.85
Δwins/run Mann-Whitney p=0.19
Spearman(measured mode%, Δdmg/run) = +0.178 (p≈0.32, n=32)
```
**The directional sub-claim is not supported and not even close**: if BitBrain
were a killer against regular movers, the `regular` bucket should sit clearly
above the `dodger` bucket; instead both are ~0 and the difference is noise.
The measured movement character also fails to separate the inferred groups (the
"regular" bucket is not measurably more periodic than the "dodger" bucket), so
the sub-claim is doubly weak: not only is the delta not larger on regular
movers, the label itself is not confirmed by the traces. The several big
per-opponent deltas are scattered across both buckets in both directions
(e.g. +68 Cigaret/dodger, -56 WaveSurferGF/dodger, +45 WallAvoider/regular,
-26 HawkOnFire/regular), which is what pure noise looks like.
## 5. What this changes about the earlier DrussGT-only conclusions `[MEASURED/INFERRED]`
* **MEASURED:** on DrussGT alone in this gauntlet BitBrain is -18.1 damage/run
(1 run-mean of 3); in the larger j113 test the learned arm was -25.9
damage/run vs Pattern (p=0.048). Both are *single-opponent* results.
* **MEASURED:** across 32 opponents the pooled estimate is -3.6 (p=0.53).
* **INFERRED (the lesson):** the DrussGT-only penalty is an opponent-specific
interaction, not a general property of the gun. Likewise the owner's
GUI impression of a "killer gun" is an opponent-specific (or
small-sample) impression that does not survive a broad field. If anything,
the broad field slightly favours Pattern on damage (19/32 opponents).
## MEASURED vs INFERRED
* **MEASURED:** every per-opponent delta and round-win count, the pooled
estimates, the sign test, the sign-flip permutation p-values, the MDE, the
movement traces (`straight%`, `wall%`, `mode%`), the liveness checks (fire
counts, 0 retries, env reached the bot), and the raw captures.
* **INFERRED:** the movement *style* labels (from names/docs; the measured
`mode%` does not confirm them), and the reading of the DrussGT-vs-field
discrepancy as an opponent-specific interaction rather than a code
difference. The spinner-specific half of the owner's claim is untested
because no constant-turn spinner is in the roster.