Files
SirRoboGarage/docs/melee_bitbrain_ab.md
T
SirStone da4a971ca9 MELEE A/B: BitBrain vs shipped Pattern rack — repo's first melee measurement (null)
Adds common_libs/tests/measure_melee_bitbrain_ab.nim (+ .sh driver, .py analyzer,
committed per-run fixtures) and docs/melee_bitbrain_ab.md.

Experiment: 4-bot Free-For-All (ModularBot + WaveSurfer + PatternMover +
RandomMover), 4 arms x 16 runs x 7 rounds, frozen ModularBot from git archive
HEAD (commit 0f5cfe3, binary 11bba27), shipped tfil movement in every run.
Arms differ only in the gun rack: pattern (shipped), bb_round, bb_ret, bb_learn.

Result: NOT DETECTABLE. Score (server round score = damage + survival bonus)
differs by -63..+33 pts (perm p=0.16-0.71) against an MDE of 151 (~5.1%).
Every arm finishes rank 1. Round wins hint BitBrain's way (112/112 and 111/112
vs 109/112) but p=0.225 (MW 0.080), half the 0.40-win MDE.

Liveness proven: rack boot lines flip (rack active melee = PATTERN / BITBRAIN),
every run faced 3 distinct targets and ~66-69 target changes, and the bb arms
logged one [bb-reset] reason=target_change per switch. The melee premise was
exercised; the fast adaptation bought no measurable score edge at this sample.
2026-09-26 00:35:44 +02:00

13 KiB
Raw Blame History

Is BitBrain actually better than Pattern in melee?

This is the repository's FIRST melee measurement. Every A/B before this one was 1v1 vs DrussGT. This doc answers one question with a real number:

"btw bitbrain gun is crushing in melee, the fast adaptation is a killer feature there" — the owner.

TL;DR (direct answer)

Not detectable at this sample size. On the metric that matters most in a melee (the server's round score = damage + survival bonus), the three BitBrain arms land within ±62 points of the shipped Pattern rack, and every difference is well inside the noise: permutation p = 0.16 – 0.71, all < the MDE of ±151 score points (~5.1 %). Placement is saturated: every arm finishes rank 1 (ModularBot's TFIL movement beats this field regardless of gun). Round wins hint in BitBrain's favour — 112/112 (retained, learned) and 111/112 (perRound) vs Pattern's 109/112 — but that +0.19 rounds/run is not significant (perm p = 0.225, Mann-Whitney p = 0.080) and is half the MDE of 0.40 rounds/run. So:

claim verdict
BitBrain beats Pattern on round wins in melee hint only, not significant (p=0.225; MDE=0.40 wins/run)
BitBrain beats Pattern on survival in melee not detectable (p=0.23–1.00)
BitBrain beats Pattern on placement in melee no — every arm sweeps (rank 1 everywhere)
BitBrain beats Pattern on score/damage in melee no — not detectable (p=0.16–0.71; all diffs << MDE)

MEASURED (this doc): the numbers above. INFERRED (not measured): the mechanism (Pattern loses history across a target switch; BitBrain resets and re-adapts). The mechanism is real and was exercised — see the target-switching liveness — but at this sample it buys no detectable score.


Why melee is a different question

Pattern's strength comes from accumulating history with one enemy. Melee rotates targets, so per-enemy histories fragment; a gun that resets per target and re-adapts fast should gain exactly where Pattern loses. BitBrain is built for that: TR_BITBRAIN_RESET_ON_TARGET defaults to on, and its memory modes (perRound / retained / decay) exist for changing targets. That is a plausible mechanism — but a plausible mechanism is not a measurement, and melee is higher-variance than 1v1 (where a 6-4 result already had P=0.353), so an impression is worth even less here. This is the first real number.

Melee scoring caveat

Melee is scored differently from 1v1 and the framework does not expose raw damage. All numbers below come from the framework's own fields (BattleResult / BotResult / BotRoundResult):

  • firstPlaces — rounds won (rank 1 in a round). The run's primary outcome.
  • survivalCount — total ticks survived across the run's 7 rounds.
  • rank / round rank — final placement and per-round placement (1 = best).
  • totalScore — the server's round score, which is damage dealt plus the survival bonus, not raw damage. It is the closest available proxy for damage; treat it as "score", not "damage". score share = ModularBot's score ÷ (all four bots' scores).

Method

Field (fixed across arms): 4-bot Free-For-All — ModularBot + WaveSurfer + PatternMover + RandomMover. Mixed by design: WaveSurfer is the strong reactive mover (linear gun), PatternMover the deterministic pattern (Pattern's home turf), RandomMover the erratic one. All three shoot.

Arms (only ModularBot's gun rack differs):

arm env
pattern (no env — shipped default: Pattern-only rack)
bb_round TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=perRound TR_BITBRAIN_LOG=1
bb_ret TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=retained TR_BITBRAIN_LOG=1
bb_learn TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=decay TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_LOG=1

Protocol: 16 independent runs × 7 rounds per arm = 112 rounds/arm. Each run is its own battle with fresh random initial positions (server --enable-initial-position); runs are not paired. A dirty tree cannot leak in: the frozen ModularBot is built with git archive HEAD.

commit        = 0f5cfe37b24952ae07fee0a7db526317ae127e6c
binary sha256 = 11bba2735413bc35ab3b61ead711c51d2223183ad9da510799bc7d76b99b22ca
movement      = tfil (shipped default) in every run — the A/B isolates the gun rack
n             = 16 runs/arm, 7 rounds/run

TR_RACK_BITBRAIN=melee was verified as melee-rack-only (not both): the boot line reads [env] rack active 1v1 = FULL and [env] rack active melee = BITBRAIN, i.e. BitBrain is admitted only in the melee rack. The pattern arm reads [env] rack active melee = PATTERN.

Results

Arm summary

arm        runs   wins/rd    score    surv  rank  meanrank  score share  targets  tchanges
bb_learn     16   112/112     2945    1050  1.00     1.000        60.2%     3.00     66.75
bb_ret       16   112/112     2903    1047  1.00     1.000        59.4%     3.00     68.06
bb_round     16   111/112     2998    1034  1.00     1.009        60.8%     3.00     68.75
pattern      16   109/112     2965    1031  1.00     1.000        59.5%     3.00     66.44

score / surv are per-run means; wins/rd is pooled round wins.

Per-run values (never just the mean)

bb_learn  wins:  r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7
          score: r1=2883 r10=2757 r11=2906 r12=3145 r13=2911 r14=3068 r15=2758 r16=2715 r2=3107 r3=2848 r4=3190 r5=2912 r6=2859 r7=3092 r8=2907 r9=3066
          surv:  r1=1050 … (all 1050)
          mrank: all 1.00
bb_ret    wins:  r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7
          score: r1=2891 r10=2770 r11=2792 r12=2934 r13=2943 r14=2890 r15=2817 r16=2999 r2=2994 r3=2913 r4=2920 r5=2890 r6=3011 r7=2945 r8=2943 r9=2802
          surv:  r5=1000, all others 1050
          mrank: all 1.00
bb_round  wins:  r1=7 r10=6 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=7 r4=7 r5=7 r6=7 r7=7 r8=7 r9=7
          score: r1=3040 r10=2944 r11=3125 r12=3110 r13=3037 r14=2993 r15=2985 r16=2842 r2=2835 r3=3045 r4=3180 r5=2952 r6=3126 r7=2927 r8=3021 r9=2804
          surv:  r10=900, r9=950, all others 1050
          mrank: r10=1.14, all others 1.00
pattern   wins:  r1=7 r10=7 r11=7 r12=7 r13=7 r14=7 r15=7 r16=7 r2=7 r3=6 r4=7 r5=6 r6=7 r7=6 r8=7 r9=7
          score: r1=3082 r10=2795 r11=2897 r12=2915 r13=3176 r14=2922 r15=2860 r16=2878 r2=3261 r3=2957 r4=3106 r5=2869 r6=3194 r7=2728 r8=2869 r9=2931
          surv:  r3=1000, r7=900, r14=950, all others 1050
          mrank: all 1.00

The only structural difference is that Pattern lost three rounds (r3, r5, r7) and bb_round one (r10); every other round is a ModularBot sweep. surv < 1050 marks the runs where ModularBot died at least once. The raw per-run JSON for all 64 runs is committed at common_libs/tests/fixtures/melee_bitbrain_ab/<arm>/run<N>.json.

Statistics (per-run, two-sided)

Permutation test (Monte-Carlo, 1,000,000 draws, seed 0x5eed5eed, p=(cnt+1)/(B+1)) and Mann-Whitney U cross-check, reused verbatim from tools/ab/ab_analyze.py. diff(A-B) is pattern − bitbrain.

metric       A        B          diff(A-B)   perm p   MW p
score        pattern  bb_learn     +19.75     0.7117   0.6109
score        pattern  bb_ret       +61.63     0.1594   0.5590
score        pattern  bb_round     -32.88     0.4900   0.3365
survival     pattern  bb_learn     -18.75     0.2254   0.0800
survival     pattern  bb_ret       -15.63     0.3507   0.2791
survival     pattern  bb_round      -3.13     1.0000   0.6982
wins         pattern  bb_learn      -0.188     0.2251   0.0795
wins         pattern  bb_ret        -0.188     0.2251   0.0795
wins         pattern  bb_round      -0.125     0.5984   0.3080
mean_rank    pattern  bb_*          ±0.000     1.0000   0.35–1.00
rank         pattern  bb_*          ±0.000     1.0000   1.0000
score_share  pattern  bb_learn      -0.007     0.5082   0.6109
score_share  pattern  bb_ret        +0.001     0.9046   0.8653
score_share  pattern  bb_round      -0.013     0.2908   0.2662

Minimum detectable effect (α=0.05 two-sided, 80 % power)

MDE = 2.8016 · sd(pattern) · sqrt(2/n), n=16/arm.

metric      sd(ref)    MDE      ref mean    MDE as % of mean
score       152.85    151.40    2965.0           5.1 %
survival     44.25     43.83    1031.3           4.3 %
wins          0.40      0.40       6.81           5.9 %
mean_rank     0.00      0.00       1.000          —

Every observed effect is below its MDE, so this is a powered null, not an under-powered one — at least for effects ≥ ~5 % of the score.

Liveness — the melee premise WAS exercised

Target switching is the whole mechanism, so it had to be proven, not assumed.

  • Arm applied: the per-run boot report shows the rack actually flipped — pattern: [env] rack active melee = PATTERN; bb_*: [env] rack active melee = BITBRAIN with TR_RACK_PATTERN=off, TR_RACK_BITBRAIN=melee. All 64 runs passed the liveness check (16/16 per arm).
  • Movement held fixed: [env] TR_MOVEMENT = tfil (source: default) in every run, so the difference is the gun rack alone.
  • Every run faced all three enemies and switched targets repeatedly. From the bot's own [config] … target=#N lines: 3.00 distinct targets/run and 66–69 target changes/run (about 10 per round). Example transitions from one run: #1 → #4 → #2 → #4 → #1 → #2 → ….
  • BitBrain really reset on every switch. [bb-reset] reason=target_change fires once per target change: 66.75 (perRound), 68.19 (retained), 66.75 (decay) per run — equal, within rounding, to the target-change count. Example [bb] line from a live run: [bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound. The pattern arm logs zero [bb]/[bb-reset] lines (BitBrain never spawned), confirming the arms are cleanly separated.

So the premise held (targets rotated, the gun reset), but the fast adaptation bought no measurable edge over Pattern's per-round wipe in this field.

Direct answer

At 16 runs/arm (112 rounds/arm), in a 4-bot melee against WaveSurfer + PatternMover + RandomMover:

  • Round wins: BitBrain is slightly ahead (112/112 and 111/112 vs 109/112), but the difference is not significant (perm p=0.225; MW p=0.080) and is below the 0.40-wins/run MDE. Do not claim a win-rate gain from this.
  • Survival: not detectable (p=0.23–1.00).
  • Placement: no difference — ModularBot is rank 1 in every arm.
  • Score/damage: not detectable — differences of −63 to +33 points (p=0.16–0.71) against an MDE of ±151 (≈5 %).

The mechanism is plausible and was demonstrably exercised, but the owner's claim that BitBrain is "crushing in melee" is not supported by this measurement. The honest reading is the opposite of spin in either direction: the first melee number is a null.

Why the null is weak evidence in one direction

The field is not discriminating enough to separate the guns on outcome: ModularBot's TFIL movement wins 97–100 % of rounds no matter which gun is in the rack, so wins/rank hit a ceiling. Only the score margin can move, and its MDE is ~5 %. A smaller true gun effect (<5 % score) would be invisible here; a larger one would not.

Reproduce

# build the runner (once)
nim c --nimcache:/tmp/nc_j116 --path:common_libs \
  common_libs/tests/measure_melee_bitbrain_ab.nim

# run the session (needs TR_SERVER_JAR + the TR runner jar; ~15–25 min under load)
MELEE_RUNS=16 MELEE_ROUNDS=7 \
  common_libs/tests/run_melee_bitbrain_ab.sh /tmp/melee_bitbrain_ab

# analyze the committed fixtures (no Java needed)
python3 common_libs/tests/analyze_melee_ab.py \
  common_libs/tests/fixtures/melee_bitbrain_ab

Files

  • common_libs/tests/measure_melee_bitbrain_ab.nim — the runner (4-bot melee, per-arm env, per-run liveness/JSON, binary built from git archive HEAD).
  • common_libs/tests/run_melee_bitbrain_ab.sh — parallel arm driver.
  • common_libs/tests/analyze_melee_ab.py — per-run table + permutation + MW + MDE (reuses tools/ab/ab_analyze.py's stats functions).
  • common_libs/tests/fixtures/melee_bitbrain_ab/ — the 64 committed per-run JSONs + analysis.json + analysis_report.txt.

MEASURED vs INFERRED

MEASURED: commit 0f5cfe3, binary 11bba27; 4-bot melee, 4 arms × 16 runs × 7 rounds; the wins/survival/rank/score means and per-run values; the permutation p-values, Mann-Whitney p-values and MDEs; the rack boot lines, the 3 distinct targets/run, 66–69 target changes/run and the per-switch [bb-reset]. All of the above are reproducible from the committed per-run JSON.

INFERRED: that Pattern's history fragmentation is the reason to expect a melee gain; that perRound/retained/decay differ in adaptation speed in a way this test could resolve; that a stronger field (or a >5 % effect) is where a difference would show. None of these are claimed as measured.