MELEE A/B: BitBrain vs shipped Pattern rack — repo's first melee measurement (null)

Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.

Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.

Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.

Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
This commit is contained in:
2026-09-26 00:35:44 +02:00
parent 02b691dd5f
commit da4a971ca9
70 changed files with 1743 additions and 0 deletions
+262
View File
@@ -0,0 +1,262 @@
# Is BitBrain actually better than Pattern in melee?
**This is the repository's FIRST melee measurement.** Every A/B before this one
was 1v1 vs DrussGT. This doc answers one question with a real number:
> *"btw bitbrain gun is crushing in melee, the fast adaptation is a killer
> feature there"* — the owner.
## TL;DR (direct answer)
**Not detectable at this sample size.** On the metric that matters most in a
melee (the server's round score = damage + survival bonus), the three BitBrain
arms land within ±62 points of the shipped Pattern rack, and every difference is
well inside the noise: permutation **p = 0.16 – 0.71**, all **< the MDE of ±151
score points (~5.1 %)**. Placement is saturated: every arm finishes **rank 1**
(ModularBot's TFIL movement beats this field regardless of gun). Round wins
*hint* in BitBrain's favour — **112/112** (retained, learned) and **111/112**
(perRound) vs Pattern's **109/112** — but that +0.19 rounds/run is not
significant (**perm p = 0.225, Mann-Whitney p = 0.080**) and is half the
**MDE of 0.40 rounds/run**. So:
| claim | verdict |
|---|---|
| BitBrain beats Pattern on **round wins** in melee | **hint only, not significant** (p=0.225; MDE=0.40 wins/run) |
| BitBrain beats Pattern on **survival** in melee | **not detectable** (p=0.23–1.00) |
| BitBrain beats Pattern on **placement** in melee | **no — every arm sweeps** (rank 1 everywhere) |
| BitBrain beats Pattern on **score/damage** in melee | **no — not detectable** (p=0.16–0.71; all diffs << MDE) |
**MEASURED** (this doc): the numbers above. **INFERRED** (not measured): the
mechanism (Pattern loses history across a target switch; BitBrain resets and
re-adapts). The mechanism is real and was exercised — see the target-switching
liveness — but at this sample it buys no detectable score.
---
## Why melee is a different question
Pattern's strength comes from accumulating history with **one** enemy. Melee
rotates targets, so per-enemy histories fragment; a gun that **resets per target
and re-adapts fast** should gain exactly where Pattern loses. BitBrain is built
for that: `TR_BITBRAIN_RESET_ON_TARGET` defaults to **on**, and its memory modes
(`perRound` / `retained` / `decay`) exist for changing targets. That is a
plausible mechanism — but a plausible mechanism is not a measurement, and
**melee is higher-variance than 1v1** (where a 6-4 result already had P=0.353),
so an impression is worth even less here. This is the first real number.
## Melee scoring caveat
Melee is scored differently from 1v1 and the framework does **not** expose raw
damage. All numbers below come from the framework's own fields
(`BattleResult` / `BotResult` / `BotRoundResult`):
* **`firstPlaces`** — rounds won (rank 1 in a round). The run's primary outcome.
* **`survivalCount`** — total ticks survived across the run's 7 rounds.
* **`rank` / round `rank`** — final placement and per-round placement (1 = best).
* **`totalScore`** — the **server's round score**, which is **damage dealt plus
the survival bonus**, *not raw damage*. It is the closest available proxy for
damage; treat it as "score", not "damage". `score share` = ModularBot's score
÷ (all four bots' scores).
## Method
*Field (fixed across arms):* 4-bot Free-For-All — **ModularBot + WaveSurfer +
PatternMover + RandomMover**. Mixed by design: WaveSurfer is the strong reactive
mover (linear gun), PatternMover the deterministic pattern (Pattern's home turf),
RandomMover the erratic one. All three shoot.
*Arms (only ModularBot's gun rack differs):*
| arm | env |
|---|---|
| `pattern` | *(no env — shipped default: Pattern-only rack)* |
| `bb_round` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=perRound TR_BITBRAIN_LOG=1` |
| `bb_ret` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=retained TR_BITBRAIN_LOG=1` |
| `bb_learn` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=decay TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_LOG=1` |
*Protocol:* **16 independent runs × 7 rounds per arm = 112 rounds/arm.** Each
run is its own battle with fresh random initial positions (server
`--enable-initial-position`); runs are not paired. A dirty tree cannot leak in:
the frozen ModularBot is built with `git archive HEAD`.
```
commit = 0f5cfe37b24952ae07fee0a7db526317ae127e6c
binary sha256 = 11bba2735413bc35ab3b61ead711c51d2223183ad9da510799bc7d76b99b22ca
movement = tfil (shipped default) in every run — the A/B isolates the gun rack
n = 16 runs/arm, 7 rounds/run
```
`TR_RACK_BITBRAIN=melee` was verified as *melee-rack-only* (not `both`): the
boot line reads `[env] rack active 1v1 = FULL` and
`[env] rack active melee = BITBRAIN`, i.e. BitBrain is admitted **only** in the
melee rack. The `pattern` arm reads `[env] rack active melee = PATTERN`.
## Results
### Arm summary
```
arm runs wins/rd score surv rank meanrank score share targets tchanges
bb_learn 16 112/112 2945 1050 1.00 1.000 60.2% 3.00 66.75
bb_ret 16 112/112 2903 1047 1.00 1.000 59.4% 3.00 68.06
bb_round 16 111/112 2998 1034 1.00 1.009 60.8% 3.00 68.75
pattern 16 109/112 2965 1031 1.00 1.000 59.5% 3.00 66.44
```
`score` / `surv` are per-run means; `wins/rd` is pooled round wins.
### Per-run values (never just the mean)
```
bb_learn wins: r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7
score: r1=2883 r10=2757 r11=2906 r12=3145 r13=2911 r14=3068 r15=2758 r16=2715 r2=3107 r3=2848 r4=3190 r5=2912 r6=2859 r7=3092 r8=2907 r9=3066
surv: r1=1050 … (all 1050)
mrank: all 1.00
bb_ret wins: r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7
score: r1=2891 r10=2770 r11=2792 r12=2934 r13=2943 r14=2890 r15=2817 r16=2999 r2=2994 r3=2913 r4=2920 r5=2890 r6=3011 r7=2945 r8=2943 r9=2802
surv: r5=1000, all others 1050
mrank: all 1.00
bb_round wins: r1=7 r10=6 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7
score: r1=3040 r10=2944 r11=3125 r12=3110 r13=3037 r14=2993 r15=2985 r16=2842 r2=2835 r3=3045 r4=3180 r5=2952 r6=3126 r7=2927 r8=3021 r9=2804
surv: r10=900, r9=950, all others 1050
mrank: r10=1.14, all others 1.00
pattern wins: r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=6 r4=7 r5=6 r6=7 r7=6 r8=7 r9=7
score: r1=3082 r10=2795 r11=2897 r12=2915 r13=3176 r14=2922 r15=2860 r16=2878 r2=3261 r3=2957 r4=3106 r5=2869 r6=3194 r7=2728 r8=2869 r9=2931
surv: r3=1000, r7=900, r14=950, all others 1050
mrank: all 1.00
```
The only structural difference is that Pattern lost three rounds (r3, r5, r7) and
`bb_round` one (r10); every other round is a ModularBot sweep. `surv` < 1050
marks the runs where ModularBot died at least once. The raw per-run JSON for all
64 runs is committed at
`common_libs/tests/fixtures/melee_bitbrain_ab/<arm>/run<N>.json`.
### Statistics (per-run, two-sided)
Permutation test (Monte-Carlo, 1,000,000 draws, seed `0x5eed5eed`, `p=(cnt+1)/(B+1)`)
and Mann-Whitney U cross-check, reused verbatim from `tools/ab/ab_analyze.py`.
`diff(A-B)` is `pattern − bitbrain`.
```
metric A B diff(A-B) perm p MW p
score pattern bb_learn +19.75 0.7117 0.6109
score pattern bb_ret +61.63 0.1594 0.5590
score pattern bb_round -32.88 0.4900 0.3365
survival pattern bb_learn -18.75 0.2254 0.0800
survival pattern bb_ret -15.63 0.3507 0.2791
survival pattern bb_round -3.13 1.0000 0.6982
wins pattern bb_learn -0.188 0.2251 0.0795
wins pattern bb_ret -0.188 0.2251 0.0795
wins pattern bb_round -0.125 0.5984 0.3080
mean_rank pattern bb_* ±0.000 1.0000 0.35–1.00
rank pattern bb_* ±0.000 1.0000 1.0000
score_share pattern bb_learn -0.007 0.5082 0.6109
score_share pattern bb_ret +0.001 0.9046 0.8653
score_share pattern bb_round -0.013 0.2908 0.2662
```
### Minimum detectable effect (α=0.05 two-sided, 80 % power)
`MDE = 2.8016 · sd(pattern) · sqrt(2/n)`, n=16/arm.
```
metric sd(ref) MDE ref mean MDE as % of mean
score 152.85 151.40 2965.0 5.1 %
survival 44.25 43.83 1031.3 4.3 %
wins 0.40 0.40 6.81 5.9 %
mean_rank 0.00 0.00 1.000 —
```
Every observed effect is **below its MDE**, so this is a *powered null*, not an
under-powered one — at least for effects ≥ ~5 % of the score.
## Liveness — the melee premise WAS exercised
Target switching is the whole mechanism, so it had to be proven, not assumed.
* **Arm applied:** the per-run boot report shows the rack actually flipped —
`pattern`: `[env] rack active melee = PATTERN`; `bb_*`:
`[env] rack active melee = BITBRAIN` with `TR_RACK_PATTERN=off`,
`TR_RACK_BITBRAIN=melee`. All 64 runs passed the liveness check (16/16 per arm).
* **Movement held fixed:** `[env] TR_MOVEMENT = tfil (source: default)` in every
run, so the difference is the gun rack alone.
* **Every run faced all three enemies and switched targets repeatedly.** From
the bot's own `[config] … target=#N` lines: **3.00 distinct targets/run** and
**66–69 target changes/run** (about 10 per round). Example transitions from one
run: `#1 → #4 → #2 → #4 → #1 → #2 → …`.
* **BitBrain really reset on every switch.** `[bb-reset] reason=target_change`
fires once per target change: **66.75** (perRound), **68.19** (retained),
**66.75** (decay) per run — equal, within rounding, to the target-change count.
Example `[bb]` line from a live run:
`[bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound`.
The `pattern` arm logs **zero** `[bb]`/`[bb-reset]` lines (BitBrain never
spawned), confirming the arms are cleanly separated.
**So the premise held** (targets rotated, the gun reset), but the fast adaptation
bought no measurable edge over Pattern's per-round wipe in this field.
## Direct answer
At **16 runs/arm (112 rounds/arm)**, in a 4-bot melee against
WaveSurfer + PatternMover + RandomMover:
* **Round wins:** BitBrain is *slightly* ahead (112/112 and 111/112 vs 109/112),
but the difference is **not significant** (perm p=0.225; MW p=0.080) and is
below the 0.40-wins/run MDE. Do **not** claim a win-rate gain from this.
* **Survival:** **not detectable** (p=0.23–1.00).
* **Placement:** **no difference** — ModularBot is rank 1 in every arm.
* **Score/damage:** **not detectable** — differences of −63 to +33 points
(p=0.16–0.71) against an MDE of ±151 (≈5 %).
The mechanism is plausible and was demonstrably exercised, but the owner's claim
that BitBrain is "crushing in melee" is **not supported by this measurement**.
The honest reading is the opposite of spin in either direction: **the first
melee number is a null.**
### Why the null is weak evidence in one direction
The field is not discriminating enough to separate the guns on *outcome*:
ModularBot's TFIL movement wins 97–100 % of rounds no matter which gun is in the
rack, so wins/rank hit a ceiling. Only the score margin can move, and its MDE is
~5 %. A smaller true gun effect (<5 % score) would be invisible here; a larger
one would not.
## Reproduce
```sh
# build the runner (once)
nim c --nimcache:/tmp/nc_j116 --path:common_libs \
common_libs/tests/measure_melee_bitbrain_ab.nim
# run the session (needs TR_SERVER_JAR + the TR runner jar; ~15–25 min under load)
MELEE_RUNS=16 MELEE_ROUNDS=7 \
common_libs/tests/run_melee_bitbrain_ab.sh /tmp/melee_bitbrain_ab
# analyze the committed fixtures (no Java needed)
python3 common_libs/tests/analyze_melee_ab.py \
common_libs/tests/fixtures/melee_bitbrain_ab
```
## Files
* `common_libs/tests/measure_melee_bitbrain_ab.nim` — the runner (4-bot melee,
per-arm env, per-run liveness/JSON, binary built from `git archive HEAD`).
* `common_libs/tests/run_melee_bitbrain_ab.sh` — parallel arm driver.
* `common_libs/tests/analyze_melee_ab.py` — per-run table + permutation + MW +
MDE (reuses `tools/ab/ab_analyze.py`'s stats functions).
* `common_libs/tests/fixtures/melee_bitbrain_ab/` — the 64 committed per-run
JSONs + `analysis.json` + `analysis_report.txt`.
## MEASURED vs INFERRED
**MEASURED:** commit `0f5cfe3`, binary `11bba27`; 4-bot melee, 4 arms × 16 runs
× 7 rounds; the wins/survival/rank/score means and per-run values; the
permutation p-values, Mann-Whitney p-values and MDEs; the rack boot lines, the
3 distinct targets/run, 66–69 target changes/run and the per-switch `[bb-reset]`.
All of the above are reproducible from the committed per-run JSON.
**INFERRED:** that Pattern's history fragmentation is *the* reason to expect a
melee gain; that `perRound`/`retained`/`decay` differ in adaptation speed in a
way this test could resolve; that a stronger field (or a >5 % effect) is where a
difference would show. None of these are claimed as measured.