# Is BitBrain actually better than Pattern in melee? **This is the repository's FIRST melee measurement.** Every A/B before this one was 1v1 vs DrussGT. This doc answers one question with a real number: > *"btw bitbrain gun is crushing in melee, the fast adaptation is a killer > feature there"* — the owner. ## TL;DR (direct answer) **Not detectable at this sample size.** On the metric that matters most in a melee (the server's round score = damage + survival bonus), the three BitBrain arms land within ±62 points of the shipped Pattern rack, and every difference is well inside the noise: permutation **p = 0.16 – 0.71**, all **< the MDE of ±151 score points (~5.1 %)**. Placement is saturated: every arm finishes **rank 1** (ModularBot's TFIL movement beats this field regardless of gun). Round wins *hint* in BitBrain's favour — **112/112** (retained, learned) and **111/112** (perRound) vs Pattern's **109/112** — but that +0.19 rounds/run is not significant (**perm p = 0.225, Mann-Whitney p = 0.080**) and is half the **MDE of 0.40 rounds/run**. So: | claim | verdict | |---|---| | BitBrain beats Pattern on **round wins** in melee | **hint only, not significant** (p=0.225; MDE=0.40 wins/run) | | BitBrain beats Pattern on **survival** in melee | **not detectable** (p=0.23–1.00) | | BitBrain beats Pattern on **placement** in melee | **no — every arm sweeps** (rank 1 everywhere) | | BitBrain beats Pattern on **score/damage** in melee | **no — not detectable** (p=0.16–0.71; all diffs << MDE) | **MEASURED** (this doc): the numbers above. **INFERRED** (not measured): the mechanism (Pattern loses history across a target switch; BitBrain resets and re-adapts). The mechanism is real and was exercised — see the target-switching liveness — but at this sample it buys no detectable score. --- ## Why melee is a different question Pattern's strength comes from accumulating history with **one** enemy. Melee rotates targets, so per-enemy histories fragment; a gun that **resets per target and re-adapts fast** should gain exactly where Pattern loses. BitBrain is built for that: `TR_BITBRAIN_RESET_ON_TARGET` defaults to **on**, and its memory modes (`perRound` / `retained` / `decay`) exist for changing targets. That is a plausible mechanism — but a plausible mechanism is not a measurement, and **melee is higher-variance than 1v1** (where a 6-4 result already had P=0.353), so an impression is worth even less here. This is the first real number. ## Melee scoring caveat Melee is scored differently from 1v1 and the framework does **not** expose raw damage. All numbers below come from the framework's own fields (`BattleResult` / `BotResult` / `BotRoundResult`): * **`firstPlaces`** — rounds won (rank 1 in a round). The run's primary outcome. * **`survivalCount`** — total ticks survived across the run's 7 rounds. * **`rank` / round `rank`** — final placement and per-round placement (1 = best). * **`totalScore`** — the **server's round score**, which is **damage dealt plus the survival bonus**, *not raw damage*. It is the closest available proxy for damage; treat it as "score", not "damage". `score share` = ModularBot's score ÷ (all four bots' scores). ## Method *Field (fixed across arms):* 4-bot Free-For-All — **ModularBot + WaveSurfer + PatternMover + RandomMover**. Mixed by design: WaveSurfer is the strong reactive mover (linear gun), PatternMover the deterministic pattern (Pattern's home turf), RandomMover the erratic one. All three shoot. *Arms (only ModularBot's gun rack differs):* | arm | env | |---|---| | `pattern` | *(no env — shipped default: Pattern-only rack)* | | `bb_round` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=perRound TR_BITBRAIN_LOG=1` | | `bb_ret` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=retained TR_BITBRAIN_LOG=1` | | `bb_learn` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=decay TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_LOG=1` | *Protocol:* **16 independent runs × 7 rounds per arm = 112 rounds/arm.** Each run is its own battle with fresh random initial positions (server `--enable-initial-position`); runs are not paired. A dirty tree cannot leak in: the frozen ModularBot is built with `git archive HEAD`. ``` commit = 0f5cfe37b24952ae07fee0a7db526317ae127e6c binary sha256 = 11bba2735413bc35ab3b61ead711c51d2223183ad9da510799bc7d76b99b22ca movement = tfil (shipped default) in every run — the A/B isolates the gun rack n = 16 runs/arm, 7 rounds/run ``` `TR_RACK_BITBRAIN=melee` was verified as *melee-rack-only* (not `both`): the boot line reads `[env] rack active 1v1 = FULL` and `[env] rack active melee = BITBRAIN`, i.e. BitBrain is admitted **only** in the melee rack. The `pattern` arm reads `[env] rack active melee = PATTERN`. ## Results ### Arm summary ``` arm runs wins/rd score surv rank meanrank score share targets tchanges bb_learn 16 112/112 2945 1050 1.00 1.000 60.2% 3.00 66.75 bb_ret 16 112/112 2903 1047 1.00 1.000 59.4% 3.00 68.06 bb_round 16 111/112 2998 1034 1.00 1.009 60.8% 3.00 68.75 pattern 16 109/112 2965 1031 1.00 1.000 59.5% 3.00 66.44 ``` `score` / `surv` are per-run means; `wins/rd` is pooled round wins. ### Per-run values (never just the mean) ``` bb_learn wins: r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7 score: r1=2883 r10=2757 r11=2906 r12=3145 r13=2911 r14=3068 r15=2758 r16=2715 r2=3107 r3=2848 r4=3190 r5=2912 r6=2859 r7=3092 r8=2907 r9=3066 surv: r1=1050 … (all 1050) mrank: all 1.00 bb_ret wins: r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7 score: r1=2891 r10=2770 r11=2792 r12=2934 r13=2943 r14=2890 r15=2817 r16=2999 r2=2994 r3=2913 r4=2920 r5=2890 r6=3011 r7=2945 r8=2943 r9=2802 surv: r5=1000, all others 1050 mrank: all 1.00 bb_round wins: r1=7 r10=6 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7 score: r1=3040 r10=2944 r11=3125 r12=3110 r13=3037 r14=2993 r15=2985 r16=2842 r2=2835 r3=3045 r4=3180 r5=2952 r6=3126 r7=2927 r8=3021 r9=2804 surv: r10=900, r9=950, all others 1050 mrank: r10=1.14, all others 1.00 pattern wins: r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=6 r4=7 r5=6 r6=7 r7=6 r8=7 r9=7 score: r1=3082 r10=2795 r11=2897 r12=2915 r13=3176 r14=2922 r15=2860 r16=2878 r2=3261 r3=2957 r4=3106 r5=2869 r6=3194 r7=2728 r8=2869 r9=2931 surv: r3=1000, r7=900, r14=950, all others 1050 mrank: all 1.00 ``` The only structural difference is that Pattern lost three rounds (r3, r5, r7) and `bb_round` one (r10); every other round is a ModularBot sweep. `surv` < 1050 marks the runs where ModularBot died at least once. The raw per-run JSON for all 64 runs is committed at `common_libs/tests/fixtures/melee_bitbrain_ab//run.json`. ### Statistics (per-run, two-sided) Permutation test (Monte-Carlo, 1,000,000 draws, seed `0x5eed5eed`, `p=(cnt+1)/(B+1)`) and Mann-Whitney U cross-check, reused verbatim from `tools/ab/ab_analyze.py`. `diff(A-B)` is `pattern − bitbrain`. ``` metric A B diff(A-B) perm p MW p score pattern bb_learn +19.75 0.7117 0.6109 score pattern bb_ret +61.63 0.1594 0.5590 score pattern bb_round -32.88 0.4900 0.3365 survival pattern bb_learn -18.75 0.2254 0.0800 survival pattern bb_ret -15.63 0.3507 0.2791 survival pattern bb_round -3.13 1.0000 0.6982 wins pattern bb_learn -0.188 0.2251 0.0795 wins pattern bb_ret -0.188 0.2251 0.0795 wins pattern bb_round -0.125 0.5984 0.3080 mean_rank pattern bb_* ±0.000 1.0000 0.35–1.00 rank pattern bb_* ±0.000 1.0000 1.0000 score_share pattern bb_learn -0.007 0.5082 0.6109 score_share pattern bb_ret +0.001 0.9046 0.8653 score_share pattern bb_round -0.013 0.2908 0.2662 ``` ### Minimum detectable effect (α=0.05 two-sided, 80 % power) `MDE = 2.8016 · sd(pattern) · sqrt(2/n)`, n=16/arm. ``` metric sd(ref) MDE ref mean MDE as % of mean score 152.85 151.40 2965.0 5.1 % survival 44.25 43.83 1031.3 4.3 % wins 0.40 0.40 6.81 5.9 % mean_rank 0.00 0.00 1.000 — ``` Every observed effect is **below its MDE**, so this is a *powered null*, not an under-powered one — at least for effects ≥ ~5 % of the score. ## Liveness — the melee premise WAS exercised Target switching is the whole mechanism, so it had to be proven, not assumed. * **Arm applied:** the per-run boot report shows the rack actually flipped — `pattern`: `[env] rack active melee = PATTERN`; `bb_*`: `[env] rack active melee = BITBRAIN` with `TR_RACK_PATTERN=off`, `TR_RACK_BITBRAIN=melee`. All 64 runs passed the liveness check (16/16 per arm). * **Movement held fixed:** `[env] TR_MOVEMENT = tfil (source: default)` in every run, so the difference is the gun rack alone. * **Every run faced all three enemies and switched targets repeatedly.** From the bot's own `[config] … target=#N` lines: **3.00 distinct targets/run** and **66–69 target changes/run** (about 10 per round). Example transitions from one run: `#1 → #4 → #2 → #4 → #1 → #2 → …`. * **BitBrain really reset on every switch.** `[bb-reset] reason=target_change` fires once per target change: **66.75** (perRound), **68.19** (retained), **66.75** (decay) per run — equal, within rounding, to the target-change count. Example `[bb]` line from a live run: `[bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound`. The `pattern` arm logs **zero** `[bb]`/`[bb-reset]` lines (BitBrain never spawned), confirming the arms are cleanly separated. **So the premise held** (targets rotated, the gun reset), but the fast adaptation bought no measurable edge over Pattern's per-round wipe in this field. ## Direct answer At **16 runs/arm (112 rounds/arm)**, in a 4-bot melee against WaveSurfer + PatternMover + RandomMover: * **Round wins:** BitBrain is *slightly* ahead (112/112 and 111/112 vs 109/112), but the difference is **not significant** (perm p=0.225; MW p=0.080) and is below the 0.40-wins/run MDE. Do **not** claim a win-rate gain from this. * **Survival:** **not detectable** (p=0.23–1.00). * **Placement:** **no difference** — ModularBot is rank 1 in every arm. * **Score/damage:** **not detectable** — differences of −63 to +33 points (p=0.16–0.71) against an MDE of ±151 (≈5 %). The mechanism is plausible and was demonstrably exercised, but the owner's claim that BitBrain is "crushing in melee" is **not supported by this measurement**. The honest reading is the opposite of spin in either direction: **the first melee number is a null.** ### Why the null is weak evidence in one direction The field is not discriminating enough to separate the guns on *outcome*: ModularBot's TFIL movement wins 97–100 % of rounds no matter which gun is in the rack, so wins/rank hit a ceiling. Only the score margin can move, and its MDE is ~5 %. A smaller true gun effect (<5 % score) would be invisible here; a larger one would not. ## Reproduce ```sh # build the runner (once) nim c --nimcache:/tmp/nc_j116 --path:common_libs \ common_libs/tests/measure_melee_bitbrain_ab.nim # run the session (needs TR_SERVER_JAR + the TR runner jar; ~15–25 min under load) MELEE_RUNS=16 MELEE_ROUNDS=7 \ common_libs/tests/run_melee_bitbrain_ab.sh /tmp/melee_bitbrain_ab # analyze the committed fixtures (no Java needed) python3 common_libs/tests/analyze_melee_ab.py \ common_libs/tests/fixtures/melee_bitbrain_ab ``` ## Files * `common_libs/tests/measure_melee_bitbrain_ab.nim` — the runner (4-bot melee, per-arm env, per-run liveness/JSON, binary built from `git archive HEAD`). * `common_libs/tests/run_melee_bitbrain_ab.sh` — parallel arm driver. * `common_libs/tests/analyze_melee_ab.py` — per-run table + permutation + MW + MDE (reuses `tools/ab/ab_analyze.py`'s stats functions). * `common_libs/tests/fixtures/melee_bitbrain_ab/` — the 64 committed per-run JSONs + `analysis.json` + `analysis_report.txt`. ## MEASURED vs INFERRED **MEASURED:** commit `0f5cfe3`, binary `11bba27`; 4-bot melee, 4 arms × 16 runs × 7 rounds; the wins/survival/rank/score means and per-run values; the permutation p-values, Mann-Whitney p-values and MDEs; the rack boot lines, the 3 distinct targets/run, 66–69 target changes/run and the per-switch `[bb-reset]`. All of the above are reproducible from the committed per-run JSON. **INFERRED:** that Pattern's history fragmentation is *the* reason to expect a melee gain; that `perRound`/`retained`/`decay` differ in adaptation speed in a way this test could resolve; that a stronger field (or a >5 % effect) is where a difference would show. None of these are claimed as measured.