gauntlet: BitBrain vs Pattern across 32 legacy opponents (does not generalize)
The repo's first multi-opponent gun measurement. Adds tools/ab/gauntlet_run.sh (per-opponent A/B over the legacy roster, subject = frozen ModularBot), tools/ab/gauntlet_analyze.py (paired per-opponent deltas, cross-opponent sign test, style split, MDE) and the arm/opponent fixtures. Result: BitBrain does NOT generalize beyond DrussGT. 32 opponents x 2 arms x 3 runs x 5 rounds = 192 battles / 960 rounds, 0 failed, 0 retries: damage/run 214.5 (pattern) vs 210.9 (bb), sign-flip p=0.53; round wins 237/480 vs 239/480, p=0.91. Sign test: bb better on 13/32 opponents (damage). The DrussGT-only penalty does not carry. The owner's 'killer vs regular movers' sub-claim is not supported: regular bucket +1.3 dmg/run vs dodgers -0.2 (MW p=0.85), and the measured movement predictability does not correlate with the delta.
This commit is contained in:
@@ -0,0 +1,220 @@
|
||||
# BitBrain vs Pattern across 32 legacy opponents — the generalization gauntlet
|
||||
|
||||
**This is the repo's first multi-opponent gun measurement.** Every previous
|
||||
gun/movement claim in this project was measured against the single opponent
|
||||
DrussGT; that caveat was flagged repeatedly and never closed until now. j114's
|
||||
28 validated legacy champions (`tools/robocode_shim/robots.json`) made a
|
||||
per-adversary gauntlet possible, so this test asks whether the DrussGT-only
|
||||
conclusions *generalize*.
|
||||
|
||||
The claim under test is the bot owner's, from watching the GUI (2026-09-25):
|
||||
*"BitBrain is a killer gun, especially against regular movements (spinners,
|
||||
wall-followers), and its fast adaptation is the reason."* The counter-evidence
|
||||
was 30 runs/arm vs DrussGT where BitBrain and Pattern were **identical**
|
||||
(97/210 vs 97/210 round wins) and the learned-gain config was the worst arm
|
||||
(`docs/bitbrain_vs_tmhorizon_ab.md`, `docs/bitbrain_gun_verdict.md`).
|
||||
|
||||
> ## DIRECT ANSWERS (MEASURED, 32 opponents × 2 arms × 3 runs × 5 rounds)
|
||||
>
|
||||
> * **Does BitBrain generalize beyond DrussGT?** **No.** Pooled over 32
|
||||
> opponents the two arms are indistinguishable: damage/run **214.5 vs 210.9**
|
||||
> (`bb - pattern` = **-3.6** ± 5.6, sign-flip p=0.53) and round wins
|
||||
> **237/480 = 49.4% vs 239/480 = 49.8%** (delta **+0.02** wins/run, p=0.91).
|
||||
> The cross-opponent sign test favours neither arm: BitBrain wins on **13/32**
|
||||
> opponents on damage (Pattern on 19) and on **10** opponents on round wins
|
||||
> (Pattern 9, 13 ties). **The DrussGT-specific penalty does not carry** — on
|
||||
> DrussGT alone BitBrain was -18.1 damage/run, but across the field the
|
||||
> estimate collapses to ~0. BitBrain is a wash, not a killer.
|
||||
> * **Is it specifically stronger against regular/periodic movers?** **No.**
|
||||
> On the 6 inferred regular movers (wall-followers, campers, rammers,
|
||||
> flood-fill) the paired delta is **+1.3 damage/run** (SD 24.4); on the 15
|
||||
> inferred dodgers it is **-0.2** (SD 33.6). Regular-vs-dodger Mann-Whitney
|
||||
> **p=0.85** (damage) / **p=0.19** (wins). The measured movement predictability
|
||||
> does not correlate with the delta either (Spearman +0.18, p=0.32). The
|
||||
> owner's observation is **not** supported: the tournament's MDE at 32
|
||||
> opponents is **15.7 damage/run** and **0.26 wins/run**, and the point
|
||||
> estimate sits at zero, so an effect of the claimed (visible) size would have
|
||||
> shown up.
|
||||
> * **Caveat that limits the sub-claim:** the legacy roster has **no true
|
||||
> constant-turn spinner** — the "regular" bucket is wall-followers / corner
|
||||
> campers / a rammer / a flood-filler. The spinner-specific half of the claim
|
||||
> is therefore *untested*, not refuted. What is refuted is the general
|
||||
> "killer vs regular movers" reading.
|
||||
|
||||
## Setup `[MEASURED]`
|
||||
|
||||
One frozen `ModularBot` built from `git archive HEAD` at run-time commit
|
||||
`da4a971ca99e81894e86a35a521bbef9ba081d3f` (binary sha256
|
||||
`f5876e7e8c14389de8f335b1529868da80d987cea5ca298abc058ecaff1925cf`). **33
|
||||
opponents attempted, 32 retained, 2 arms, 3 runs × 5 rounds each =
|
||||
192 battles / 960 rounds**, `--conc 4`, **0 failed, 0 retries needed**. Raw
|
||||
per-tick captures live at `/tmp/ab/j117_gauntlet/` and are **not** committed;
|
||||
the full machine report is
|
||||
`common_libs/tests/fixtures/gauntlet_bitbrain_vs_pattern_report.txt` and the
|
||||
session header (arms, opponents, styles) is
|
||||
`..._session.json`.
|
||||
|
||||
| arm | env | role |
|
||||
|---|---|---|
|
||||
| `pattern` | *(none)* | shipped default (`onlyPattern` rack) — reference |
|
||||
| `bb` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | BitBrain-only, the learned-gain config the owner most likely ran |
|
||||
|
||||
Runner `tools/ab/gauntlet_run.sh` (per-opponent A/B; subject = the frozen
|
||||
ModularBot itself, adversary iterates over the roster), arm list
|
||||
`tools/ab/arms_gauntlet.txt`, opponents + inferred movement styles
|
||||
`tools/ab/opponents_gauntlet.txt`, analyzer `tools/ab/gauntlet_analyze.py`.
|
||||
|
||||
## Validity / false-negative guard `[MEASURED]`
|
||||
|
||||
Per j114's warning, `run_smoke.sh` has false negatives and only a real battle is
|
||||
authoritative, so each battle is validated and (if the bot failed to connect
|
||||
inside the booter's 30 s window) retried up to 4 times. Results:
|
||||
|
||||
* **0 of 192 battles needed a retry**; every *retained* opponent fired and was
|
||||
scored (the roster also includes the 5 `WORKS_WEAK` bots — valid movement
|
||||
targets — of which only Aurora turned out inert).
|
||||
* The least-active retained opponent fired **136** times across both arms; the
|
||||
one inert bot (Aurora — a single fire in 6 battles) was **dropped**.
|
||||
* Every declared arm env var reached the process, verified from each bot's own
|
||||
`[env]` boot report: `TR_RACK_BITBRAIN = both`, `TR_RACK_PATTERN = off`,
|
||||
`TR_BITBRAIN_GAINS = 1.0,1.25,1.5,2.0`, `TR_BITBRAIN_MEM = decay`, and the
|
||||
active 1v1 rack line reads `BITBRAIN` for `bb` and `PATTERN` for `pattern`.
|
||||
|
||||
## 1. Per-opponent result `[MEASURED]`
|
||||
|
||||
The unit of evidence is the **number of opponents**, so the design spends its
|
||||
budget on breadth (3 runs/opponent) rather than depth. Full table in the fixture;
|
||||
the deltas that matter:
|
||||
|
||||
| opponent | style (INFERRED) | Δ dmg/run | Δ wins/run |
|
||||
|---|---|---:|---:|
|
||||
| Cigaret | dodger | +68.1 | +1.00 |
|
||||
| WallAvoider | regular | +44.7 | +0.33 |
|
||||
| Coriantumr | other | +44.5 | +0.33 |
|
||||
| Komarious | dodger | +41.9 | +0.67 |
|
||||
| CigaretBH | dodger | +41.4 | +1.00 |
|
||||
| WaveSurferPG | dodger | +23.1 | -0.33 |
|
||||
| DiamondHawk | other | +18.0 | +0.67 |
|
||||
| Jen | dodger | +8.0 | 0.00 |
|
||||
| FloodMini | regular | +6.7 | 0.00 |
|
||||
| Dookious | dodger | +5.5 | 0.00 |
|
||||
| PatternRobot | regular | +4.6 | 0.00 |
|
||||
| LionWWSVMvoid | dodger | +4.5 | 0.00 |
|
||||
| YersiniaPestis | other | +3.8 | -0.33 |
|
||||
| KurtWaveSurfer | dodger | -3.1 | -1.00 |
|
||||
| TripHammer | other | -3.4 | 0.00 |
|
||||
| LightningBug | other | -6.0 | 0.00 |
|
||||
| CassiusClay | dodger | -6.1 | +0.33 |
|
||||
| Diamond | dodger | -8.0 | +0.33 |
|
||||
| DiamondStealer | regular | -8.8 | 0.00 |
|
||||
| Shadow | other | -9.9 | 0.00 |
|
||||
| Ascendant | other | -11.5 | 0.00 |
|
||||
| BlitzBat | regular | -13.6 | -0.33 |
|
||||
| BrokenSword | other | -15.5 | 0.00 |
|
||||
| DrussGT | other | -18.1 | -0.67 |
|
||||
| HawkOnFire | regular | -25.9 | -1.00 |
|
||||
| RougeDC | dodger | -26.8 | 0.00 |
|
||||
| GresSuffurd | dodger | -29.1 | +1.00 |
|
||||
| Aristocles | dodger | -29.3 | 0.00 |
|
||||
| Phoenix | other | -29.4 | -0.33 |
|
||||
| Lukious | dodger | -37.1 | +0.33 |
|
||||
| WaveSurferGF | dodger | -55.9 | -0.33 |
|
||||
| RetroGirl | other | -93.1 | -1.00 |
|
||||
|
||||
The individual deltas swing from -93 to +68 damage/run — with only 3 runs per
|
||||
opponent, a single battle dominates each cell, so **no individual row is
|
||||
evidence on its own.** The evidence is the aggregate below.
|
||||
|
||||
## 2. Cross-opponent sign test — the headline `[MEASURED]`
|
||||
|
||||
On how many opponents does each arm win (paired `bb - pattern`)?
|
||||
|
||||
```
|
||||
metric bb better pattern better tie two-sided sign-test p
|
||||
damage/run 13 19 0 0.377
|
||||
wins/run 10 9 13 1.000
|
||||
```
|
||||
|
||||
Neither arm wins on the majority of opponents. On round wins the two arms are
|
||||
dead even (10 vs 9, 13 exact ties), which is the round-level echo of the
|
||||
30-run-vs-DrussGT finding (97/210 vs 97/210).
|
||||
|
||||
## 3. Pooled paired estimate with between-opponent spread `[MEASURED]`
|
||||
|
||||
Per-opponent paired deltas, so one outlier bot cannot carry the result:
|
||||
|
||||
```
|
||||
metric mean Δ between-opp SD SE 95% CI sign-flip p MDE(n=32)
|
||||
damage/run -3.62 31.65 5.60 [-14.59, +7.35] 0.527 15.68
|
||||
wins/run +0.02 0.52 0.09 [-0.16, +0.20] 0.911 0.258
|
||||
```
|
||||
|
||||
Pooled arm totals over the 32 opponents:
|
||||
|
||||
```
|
||||
arm damage/run damage taken/run round wins
|
||||
pattern 214.5 291.5 237/480 = 49.4%
|
||||
bb 210.9 296.4 239/480 = 49.8%
|
||||
```
|
||||
|
||||
Both verdict metrics (damage/run and round wins, per the standing rule) say the
|
||||
same thing: **no detectable difference.** The MDE at 32 opponents is 15.7
|
||||
damage/run (≈7% of the ~214 baseline) and 0.258 wins/run (≈5.2 pp); the point
|
||||
estimates are ~0. So a *large* BitBrain advantage is excluded, and the observed
|
||||
effect is at the null. The usual caveat stands: a genuinely small positive
|
||||
effect (<~8%) is **not** excluded by this design — but there is no evidence of
|
||||
one.
|
||||
|
||||
## 4. Style split and measured movement character `[MEASURED]`
|
||||
|
||||
Style labels are **INFERRED** from the bot names/docs in `robots.json`. To
|
||||
check them against data, the opponent's own captured trajectory (`sx,sy,sh,ss`)
|
||||
was characterised by its single most-common turn magnitude (`mode%`): a
|
||||
straight-liner, a spinner and a fixed oscillator all concentrate there, an
|
||||
adaptive surfer does not.
|
||||
|
||||
```
|
||||
style n mean Δdmg/run (SD) mean Δwins/run (SD) measured mode%
|
||||
regular 6 +1.3 (24.4) -0.17 (0.46) 0.58
|
||||
dodger 15 -0.2 (33.6) +0.20 (0.56) 0.65
|
||||
other 11 -11.0 (33.7) -0.12 (0.45) 0.59
|
||||
|
||||
regular vs dodger: Δdmg/run Mann-Whitney p=0.85
|
||||
Δwins/run Mann-Whitney p=0.19
|
||||
Spearman(measured mode%, Δdmg/run) = +0.178 (p≈0.32, n=32)
|
||||
```
|
||||
|
||||
**The directional sub-claim is not supported and not even close**: if BitBrain
|
||||
were a killer against regular movers, the `regular` bucket should sit clearly
|
||||
above the `dodger` bucket; instead both are ~0 and the difference is noise.
|
||||
The measured movement character also fails to separate the inferred groups (the
|
||||
"regular" bucket is not measurably more periodic than the "dodger" bucket), so
|
||||
the sub-claim is doubly weak: not only is the delta not larger on regular
|
||||
movers, the label itself is not confirmed by the traces. The several big
|
||||
per-opponent deltas are scattered across both buckets in both directions
|
||||
(e.g. +68 Cigaret/dodger, -56 WaveSurferGF/dodger, +45 WallAvoider/regular,
|
||||
-26 HawkOnFire/regular), which is what pure noise looks like.
|
||||
|
||||
## 5. What this changes about the earlier DrussGT-only conclusions `[MEASURED/INFERRED]`
|
||||
|
||||
* **MEASURED:** on DrussGT alone in this gauntlet BitBrain is -18.1 damage/run
|
||||
(1 run-mean of 3); in the larger j113 test the learned arm was -25.9
|
||||
damage/run vs Pattern (p=0.048). Both are *single-opponent* results.
|
||||
* **MEASURED:** across 32 opponents the pooled estimate is -3.6 (p=0.53).
|
||||
* **INFERRED (the lesson):** the DrussGT-only penalty is an opponent-specific
|
||||
interaction, not a general property of the gun. Likewise the owner's
|
||||
GUI impression of a "killer gun" is an opponent-specific (or
|
||||
small-sample) impression that does not survive a broad field. If anything,
|
||||
the broad field slightly favours Pattern on damage (19/32 opponents).
|
||||
|
||||
## MEASURED vs INFERRED
|
||||
|
||||
* **MEASURED:** every per-opponent delta and round-win count, the pooled
|
||||
estimates, the sign test, the sign-flip permutation p-values, the MDE, the
|
||||
movement traces (`straight%`, `wall%`, `mode%`), the liveness checks (fire
|
||||
counts, 0 retries, env reached the bot), and the raw captures.
|
||||
* **INFERRED:** the movement *style* labels (from names/docs; the measured
|
||||
`mode%` does not confirm them), and the reading of the DrussGT-vs-field
|
||||
discrepancy as an opponent-specific interaction rather than a code
|
||||
difference. The spinner-specific half of the owner's claim is untested
|
||||
because no constant-turn spinner is in the roster.
|
||||
Reference in New Issue
Block a user