From cf77d0d6478743560b823ed11cf540a7c554d5f3 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sat, 26 Sep 2026 05:05:09 +0200 Subject: [PATCH] j124: spinner + fair-melee results doc; analyzer convergence tests --- docs/spinner_and_melee_claims.md | 469 +++++++++++++++++++++++++++++++ tools/ab/spinner_analyze.py | 122 ++++++++ 2 files changed, 591 insertions(+) create mode 100644 docs/spinner_and_melee_claims.md diff --git a/docs/spinner_and_melee_claims.md b/docs/spinner_and_melee_claims.md new file mode 100644 index 0000000..47c6f9d --- /dev/null +++ b/docs/spinner_and_melee_claims.md @@ -0,0 +1,469 @@ +# Two untested owner claims: the spinner gun, and a fair melee field + +The owner reported two things from watching the GUI (2026-09-25): + +1. *"I never seen a gun that learns wall movement or circular movement like + spinning bot so fast - like a gun made on purpose for those movements"* — + a **spinner-specific** advantage of BitBrain over the shipped `Pattern` gun. +2. *"bitbrain gun is crushing in melee, the fast adaptation is a killer feature + there"* — a **melee** advantage of BitBrain over the shipped melee rack. + +Both had been *recorded as untested for a specific reason*, not refuted: + +* **Claim 1's strongest case was absent from every panel.** The legacy roster + (`tools/robocode_shim/robots.json`) has no purpose-built constant-turn + spinner. `docs/gauntlet_bitbrain_vs_pattern.md` (j117) found no regular-vs- + dodger difference (Mann-Whitney p=0.85 on damage) but its "regular" bucket was + wall-followers / campers / a rammer / a flood-filler, and + `docs/gun_campaign.md` (j121) explicitly flagged the spinner half as + **UNTESTED**. +* **Claim 2 was tested against the wrong field.** `docs/melee_bitbrain_ab.md` + (j116) ran a 4-bot melee against weak in-repo adversaries and ModularBot won + **97–100 % of rounds regardless of gun** — a ceiling effect. The result was + "not detectable", and the doc itself says the question remains open until a + field exists that can punish a bad gun. + +This document closes both. It adds the missing **true constant-turn spinner** +fixture, reruns the gun A/B on it (with a convergence analysis, because the +claim is about *speed* of adaptation), and reruns the melee A/B on a **strong** +field of battle-validated legacy champions. + +> ## DIRECT ANSWERS (MEASURED) +> +> * **Claim 1 — spinner gun: REFUTED on a finally-fair field.** Against the two +> true constant-turn spinners plus the sample SpinBot, BitBrain (`decay`) is +> **not** better than `Pattern` on any metric: damage **−13.4 dmg/run** +> (0/3 spinners positive; MDE 15.3), round wins **+0.00** (both arms win 5/5 — +> saturated), our gun's hit rate **−4.4 pp** (MDE 10.7). The `retained` memory +> mode is the same (−14.3 dmg/run, 1/3). The convergence trajectory — the +> actual claim — shows BitBrain **starts lower** on round 1 (54.4 % vs +> Pattern's 69.0 %) and catches up to roughly the same level by round 5 +> (75.4 % vs 71.5 %); the per-run cross-round adaptation delta is +7.7 pp for +> BitBrain vs +0.4 pp for Pattern, but **p=0.38** (not significant), and the +> within-round first-vs-second-half metric is the **opposite** (−2.7 vs +> +5.5 pp, p=0.45). There is no measured fast-adaptation advantage. +> * **Claim 2 — melee: NOT SUPPORTED on a finally-fair field.** The j116 ceiling +> is gone: on a strong field of three battle-validated legacy champions +> (Diamond + Dookious + GresSuffurd) ModularBot wins only **30 %** of rounds +> with `pattern` (18/60), versus 97–100 % against the old weak field. On that +> field BitBrain does **not** beat the shipped melee rack: `bb_decay` is +> **−11.8 score/run** and **−0.08 wins/run** vs `pattern` (p=0.90 / p=1.0), +> and `bb_ret` is **−123.5 score/run** and **−0.50 wins/run** (p=0.21 / +> p=0.35). Every point estimate is on the wrong side of the claim. The field +> is fair but the test is underpowered for small effects (score MDE ±312 on a +> mean of 992, ≈31 %; wins MDE 1.34 on a mean of 1.5). A **large** advantage +> (≥ ~31 % score) is excluded; a smaller one is unmeasured at n=12. + +**MEASURED** = the numbers in this doc. **INFERRED** = mechanisms and style +labels, stated as such. Raw captures are not committed (they live under +`/tmp/ab/j124_spinner/` and `/tmp/melee_strong_field/`); every number below is +reproducible from the committed harnesses. + +--- + +## 1. The new adversary: `ConstantSpinner` `[MEASURED]` + +`common_libs/test_framework/adversaries/ConstantSpinner/` — a minimal Nim Tank +Royale bot whose entire `run()` loop is: + +```nim +setRadarTurnRate(45.0) +setTurnRate(bot.turnRate) # constant body turn, deg/tick +setTargetSpeed(bot.speed) # constant forward speed +``` + +Its path is a circle of radius `v / ω` (a single frequency, perfectly +periodic). The gun is plain **head-on** (aim at the enemy's current position, +fire power 1.0 whenever the gun is cool) — it is a movement fixture, not a +fighter. No randomness, no adaptation, no wall logic: the server's wall +collision is the only non-periodic perturbation. + +The requested turn rate is clamped to the maximum reachable at the target speed +(`MAX_TURN_RATE − 0.75·speed`), so the **actual** turn rate stays exactly +constant instead of being server-clamped while the bot accelerates or slides +along a wall. Both rate and speed are env-configurable +(`SPIN_TURN_RATE`, `SPIN_SPEED`, `SPIN_FIRE_POWER`), and `make_variant.sh` +generates named per-rate bot directories that reuse the one committed binary. + +**Why this is not just SpinBot.** The premise that the roster had *no* spinner +is only half true, and the measurement says so: the sample SpinBot's capture +trace turns at a near-constant **−6.2 deg/tick** (94.9 % mode share), because it +requests the speed-dependent *maximum*. What was genuinely missing is a rate +that is (a) **exactly** constant — not varying with speed near walls — and +(b) **not** the maximum. `ConstantSpinner` provides both, at a slow (3 deg/tick) +and a fast (6 deg/tick) rate, both at speed 5. + +### Fixture validation (from the adversary's own capture trace) + +Per-tick body turn and speed of each spinner opponent, pooled over its 12 runs +(3 arms × 4 runs), read from the capture's `s*` fields: + +| opponent | ticks | turn mode (deg/tick) | mode share | turn SD | speed mode | speed mode share | wall-hug frac | +|---|---:|---:|---:|---:|---:|---:|---:| +| ConstantSpinner_s3 | 8176 | **3.0** | 93.9 % | **0.720** | 5.0 | 89.0 % | 21.3 % | +| ConstantSpinner_s6 | 8183 | **6.0** | 94.3 % | 1.391 | 5.0 | 95.0 % | 17.4 % | +| SpinBot (control) | 7891 | −6.2 | 94.9 % | 1.067 | 5.0 | 97.3 % | 6.1 % | + +The `s3` variant has the **lowest turn SD of all three** — the exactly-constant +rate the claim needs. The wall-hug fraction is the only source of non-mode +ticks (the bot grinds a wall, keeps turning, and leaves). + +### It really fights + +Across the 4 runs per arm, the spinner opponents fired **449–478** shots +(`s3`), **421–447** (`s6`) and **120–128** (SpinBot) and were scored by the +server in every round. In a 3-round smoke vs Diamond it fired 59 shots and won +round 1 (`firstPlaces=1`). It is a valid, scored adversary. + +--- + +## 2. Claim 1 — the spinner gun test `[MEASURED]` + +**Design.** One frozen `ModularBot` built from `git archive HEAD` at commit +`6c21dc4c58caaaaf68dc670a28a7c0d6948e3f3e` (binary sha256 +`cfd11b5fdf3bcfe217c5d9033e2b20a618eb3b2e30d9bd41913d88eb6a0cfeef`), movement +**pinned** to `TR_MOVEMENT=strafe` in every arm (the gun-campaign standard), so +every delta is a pure gun delta. + +*Panel* (`tools/ab/panel_spinner.txt`, 9 opponents): 3 spinners +(`ConstantSpinner_s3`, `ConstantSpinner_s6`, sample `SpinBot`) + 3 known regular +movers (`WallAvoider`, `DiamondStealer`, `HawkOnFire`) + 3 known dodgers +(`Diamond`, `CassiusClay`, `GresSuffurd`), all legacy champions from +`/tmp/tr_bots/`. + +*Arms* (`tools/ab/arms_spinner.txt`, 3 arms, same binary): + +| arm | env (beyond `TR_MOVEMENT=strafe`) | role | +|---|---|---| +| `pattern` | *(none)* | shipped `onlyPattern` rack — reference | +| `bitbrain` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=both TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_MEM=decay` | the owner's config | +| `bitbrain_ret` | `... TR_BITBRAIN_MEM=retained` | "adapt across the battle" mode | + +**Protocol.** 9 opponents × 3 arms × 4 runs × 5 rounds = **108 battles / 540 +rounds**, `--conc 6`, **0 failed, 0 never started, 0 liveness exclusions** +(`tournament_run.sh --wait-arena`). Runner `tools/ab/tournament_run.sh`, +analyzer `tools/ab/spinner_analyze.py` (primaries + hit rate + convergence), +cross-checked with `tools/ab/tournament_analyze.py`. + +**Liveness (verified from each run's own boot report).** Every declared env +token reached the process; the rack lines read `rack active 1v1 = PATTERN` for +`pattern` and `= BITBRAIN` for both BitBrain arms; `TR_MOVEMENT = strafe +(source: env)` in every run. Every opponent fired and was scored. + +### 2.1 Per-opponent paired table (deltas are arm − `pattern`) + +`bitbrain` (`decay`): + +| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | our hit% ref→arm | Δhit (pp) | +|---|---|---:|---:|---:|---:|---:|---:| +| ConstantSpinner_s3 | spinner | 397.5→390.5 | −7.0 | 5.00→5.00 | +0.00 | 56.10→55.77 | −0.33 | +| ConstantSpinner_s6 | spinner | 400.1→391.2 | −8.9 | 5.00→5.00 | +0.00 | 69.01→68.14 | −0.87 | +| SpinBot | spinner | 472.0→447.8 | −24.2 | 5.00→5.00 | +0.00 | 84.37→72.27 | −12.10 | +| WallAvoider | regular | 251.3→245.4 | −6.0 | 3.75→3.25 | −0.50 | 30.62→31.22 | +0.60 | +| DiamondStealer | regular | 237.7→234.4 | −3.3 | 2.00→2.25 | +0.25 | 30.51→28.70 | −1.81 | +| HawkOnFire | regular | 172.5→198.6 | +26.1 | 3.50→4.25 | +0.75 | 22.82→25.36 | +2.54 | +| Diamond | dodger | 87.7→106.7 | +19.0 | 0.25→0.50 | +0.25 | 8.85→9.07 | +0.22 | +| CassiusClay | dodger | 134.4→113.2 | −21.2 | 2.75→2.00 | −0.75 | 13.18→12.83 | −0.35 | +| GresSuffurd | dodger | 180.6→167.9 | −12.6 | 3.25→3.00 | −0.25 | 15.58→13.65 | −1.93 | + +`bitbrain_ret` (`retained`): + +| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | our hit% ref→arm | Δhit (pp) | +|---|---|---:|---:|---:|---:|---:|---:| +| ConstantSpinner_s3 | spinner | 397.5→384.8 | −12.8 | 5.00→5.00 | +0.00 | 56.10→55.23 | −0.87 | +| ConstantSpinner_s6 | spinner | 400.1→403.2 | +3.1 | 5.00→5.00 | +0.00 | 69.01→69.04 | +0.04 | +| SpinBot | spinner | 472.0→438.9 | −33.1 | 5.00→5.00 | +0.00 | 84.37→76.13 | −8.23 | +| WallAvoider | regular | 251.3→307.0 | +55.7 | 3.75→3.00 | −0.75 | 30.62→37.28 | +6.66 | +| DiamondStealer | regular | 237.7→215.1 | −22.6 | 2.00→1.75 | −0.25 | 30.51→28.36 | −2.16 | +| HawkOnFire | regular | 172.5→178.2 | +5.8 | 3.50→4.00 | +0.50 | 22.82→21.55 | −1.27 | +| Diamond | dodger | 87.7→96.0 | +8.3 | 0.25→0.50 | +0.25 | 8.85→8.28 | −0.57 | +| CassiusClay | dodger | 134.4→107.9 | −26.5 | 2.75→0.75 | −2.00 | 13.18→12.66 | −0.51 | +| GresSuffurd | dodger | 180.6→156.0 | −24.5 | 3.25→2.75 | −0.50 | 15.58→13.78 | −1.80 | + +### 2.2 Pooled dashboard (explanation, not the verdict) + +| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | our hit rate | incoming hit rate | +|---|---:|---:|---:|---:|---:|---:|---:|---:| +| `pattern` | 36 | 259.3 | 216.8 | 3.39 | 122/180 | 67.8 % | 28.78 % | 12.45 % | +| `bitbrain` | 36 | 255.1 | 215.0 | 3.36 | 121/180 | 67.2 % | 27.81 % | 12.92 % | +| `bitbrain_ret` | 36 | 254.1 | 218.7 | 3.08 | 111/180 | 61.7 % | 27.82 % | 12.69 % | + +### 2.3 Cross-opponent aggregation (the verdict layer, n=9 opponents) + +| arm | metric | mean Δ | spread (SD) | 95 % CI | sign test | p(sign) | p(sign-flip) | p(Wilcoxon) | MDE | +|---|---|---:|---:|---|---:|---:|---:|---:|---:| +| `bitbrain` | damage | −4.23 | 16.78 | [−15.19, +6.74] | 2/9 | 0.180 | 0.457 | 0.407 | 15.67 | +| `bitbrain` | wins | −0.03 | 0.44 | [−0.32, +0.26] | 3/6 | 1 | 1 | 0.915 | 0.41 | +| `bitbrain` | our hit rate (pp) | −1.56 | 4.17 | [−4.28, +1.17] | 3/9 | 0.508 | 0.328 | 0.286 | 3.90 | +| `bitbrain_ret` | damage | −5.17 | 27.50 | [−23.14, +12.79] | 4/9 | 1 | 0.602 | 0.407 | 25.68 | +| `bitbrain_ret` | wins | −0.31 | 0.74 | [−0.79, +0.18] | 2/6 | 0.688 | 0.344 | 0.292 | 0.69 | +| `bitbrain_ret` | our hit rate (pp) | −0.97 | 3.79 | [−3.44, +1.50] | 2/9 | 0.180 | 0.496 | 0.124 | 3.54 | + +### 2.4 The sub-claim's own field: the 3 spinners only + +| arm | metric | mean Δ | spread (SD) | 95 % CI | sign test | p(sign) | p(sign-flip) | MDE | +|---|---|---:|---:|---|---:|---:|---:|---:| +| `bitbrain` | damage | **−13.38** | 9.46 | [−24.09, −2.66] | **0/3** | 0.25 | 0.25 | 15.31 | +| `bitbrain` | wins | +0.00 | 0.00 | [+0.00, +0.00] | 0/0 | 1 | 1 | 0.00 | +| `bitbrain` | our hit rate (pp) | **−4.43** | 6.64 | [−11.95, +3.08] | **0/3** | 0.25 | 0.25 | 10.74 | +| `bitbrain_ret` | damage | **−14.25** | 18.17 | [−34.81, +6.31] | 1/3 | 1 | 0.5 | 29.39 | +| `bitbrain_ret` | wins | +0.00 | 0.00 | [+0.00, +0.00] | 0/0 | 1 | 1 | 0.00 | +| `bitbrain_ret` | our hit rate (pp) | **−3.02** | 4.54 | [−8.16, +2.11] | 1/3 | 1 | 0.5 | 7.34 | + +Both BitBrain arms lose on damage and on our hit rate on the spinner field; +**every** point estimate is on the wrong side of the claim. Round wins are +saturated (5/5 for every arm on every spinner), so the win metric cannot +discriminate there. + +### 2.5 Style split + +| arm | style | n | mean Δdmg | mean Δwins | mean Δour-hit (pp) | +|---|---|---:|---:|---:|---:| +| `bitbrain` | spinner | 3 | **−13.38** | +0.00 | −4.43 | +| `bitbrain` | regular | 3 | +5.62 | +0.17 | +0.44 | +| `bitbrain` | dodger | 3 | −4.92 | −0.25 | −0.69 | +| `bitbrain_ret` | spinner | 3 | **−14.25** | +0.00 | −3.02 | +| `bitbrain_ret` | regular | 3 | +12.96 | −0.17 | +1.08 | +| `bitbrain_ret` | dodger | 3 | −14.23 | −0.75 | −0.96 | + +The spinner bucket is the **worst** bucket for BitBrain, not the best. This is +the exact opposite of the owner's claim, and it agrees with j121's SpinBot +result (Batch 1 +25.3 dmg was a small-panel fluctuation; Batch 2 −4.7 dmg). + +### 2.6 CONVERGENCE — the actual claim + +**Our gun's per-round hit rate, pooled over runs.** R1…R5: + +*all 9 opponents:* + +| arm | R1 | R2 | R3 | R4 | R5 | R1→Rlast | +|---|---:|---:|---:|---:|---:|---:| +| `pattern` | 29.3 % | 29.6 % | 28.9 % | 26.6 % | 29.6 % | +0.3 pp | +| `bitbrain` | 26.6 % | 28.2 % | 27.5 % | 28.6 % | 28.3 % | +1.7 pp | +| `bitbrain_ret` | 28.9 % | 27.5 % | 26.8 % | 26.4 % | 29.6 % | +0.8 pp | + +*true spinners only:* + +| arm | R1 | R2 | R3 | R4 | R5 | R1→Rlast | +|---|---:|---:|---:|---:|---:|---:| +| `pattern` | 69.0 % | 60.4 % | 82.4 % | 68.9 % | 71.5 % | +2.5 pp | +| `bitbrain` | **54.4 %** | 69.8 % | 62.5 % | 68.7 % | **75.4 %** | **+21.0 pp** | +| `bitbrain_ret` | 66.8 % | 63.8 % | 65.6 % | 68.3 % | 67.8 % | +1.1 pp | + +The per-round trajectory *looks* like a BitBrain convergence story — but it is +a **start-lower, catch-up-to-the-same-level** story: BitBrain is 14.6 pp *worse* +in round 1, then reaches Pattern's level by round 5, never above it. A gun that +adapts *faster* should be **ahead** in the early rounds, not behind. + +**Per-run adaptation deltas on the true spinners** (n = 3 spinners × 4 runs = 12 +per arm): + +* `conv` = hit rate over rounds 2..5 minus round 1 (cross-round); +* `within` = hit rate in the second half of a round minus the first half. + +| arm | n | conv mean Δ | conv SD | within mean Δ | within SD | +|---|---:|---:|---:|---:|---:| +| `pattern` | 12 | +0.39 pp | 9.86 | +5.50 pp | 15.65 | +| `bitbrain` | 12 | +7.67 pp | 24.65 | −2.72 pp | 27.92 | +| `bitbrain_ret` | 12 | −4.66 pp | 19.26 | +16.54 pp | 34.43 | + +Two-sample permutation test (200 000 draws, seed `0x5eed5eed`) against +`pattern`: + +| arm | conv Δ − ref Δ | p(conv) | within Δ − ref Δ | p(within) | +|---|---:|---:|---:|---:| +| `bitbrain` | +7.29 pp | **0.379** | −8.22 pp | **0.446** | +| `bitbrain_ret` | −5.05 pp | 0.490 | +11.04 pp | 0.381 | + +The cross-round hint is **not significant** (p=0.38) and the within-round +adaptation — the fastest possible timescale — points the **other way** +(−2.7 pp vs Pattern's +5.5 pp). There is no measured speed advantage. + +### 2.7 Direct answer — claim 1 + +**REFUTED on a finally-fair field.** The field now contains two purpose-built +constant-turn spinners at different rates (3 and 6 deg/tick), the sample SpinBot, +and regular/dodger controls; movement is pinned; the adversary fires and is +scored; 108 battles, no exclusions. On that field: + +* BitBrain (`decay`) is **not** better than `Pattern` on damage + (−4.2 dmg/run overall; **−13.4** on the spinners, MDE 15.3), on round wins + (−0.03 overall; **+0.00** on spinners — saturated), or on our hit rate + (−1.6 pp overall; **−4.4 pp** on the spinners, MDE 10.7). +* The convergence trajectory — the mechanism the claim is about — shows a + **slower start** and a catch-up to parity, not a faster or higher convergence; + the cross-round delta is non-significant (p=0.38) and the within-round delta is + opposite. +* `retained` memory behaves like `decay` (all deltas negative on the spinner + field, none significant). + +The owner's impression is consistent with BitBrain *changing its aim over the +first rounds* (visible in a GUI) — the trajectory is real — but the change does +not buy a hit-rate or damage advantage over `Pattern`; it recovers a deficit +BitBrain itself created. + +--- + +## 3. Claim 2 — the fair melee test `[MEASURED]` + +**The j116 field is replaced.** The old field was ModularBot + WaveSurfer + +PatternMover + RandomMover, against which ModularBot won 97–100 % of rounds +whatever the gun (a ceiling). The new field is **ModularBot + Diamond + +Dookious + GresSuffurd** — three battle-validated legacy champions from +`/tmp/tr_bots/` (wave-surfing dodgers with real guns; `tools/robocode_shim/robots.json`) +that can and do punish a bad gun. + +**Arms** (only ModularBot's melee rack differs): + +| arm | env | role | +|---|---|---| +| `pattern` | *(none)* — shipped default melee rack | reference | +| `bb_ret` | `TR_RACK_PATTERN=off TR_RACK_BITBRAIN=melee TR_BITBRAIN_MEM=retained TR_BITBRAIN_LOG=1` | adapt across the battle | +| `bb_decay` | `... TR_BITBRAIN_MEM=decay TR_BITBRAIN_GAINS=1.0,1.25,1.5,2.0 TR_BITBRAIN_LOG=1` | the owner's config | + +**Protocol.** One frozen `ModularBot` from `git archive HEAD` at commit +`5fe28574abb3af43af8e0bd4e86d3f8f14a54f21` (binary sha256 +`7dc1b7c1f9a29cc4009cba1f2b53d2424cff1a6d1829fe4322df80783764b46c`), +**12 runs × 5 rounds per arm = 36 battles / 180 rounds**, each run a fresh +random-position melee with its own Tank Royale server. Runner +`common_libs/tests/measure_melee_strong_field.nim` + `run_melee_strong_field.sh`; +analyzer `common_libs/tests/analyze_melee_ab.py` (the same permutation + MDE +machinery as the 1v1 A/Bs). 36/36 runs OK (two runs needed the harness's +auto-retry for a slow JVM boot). + +### 3.1 Is the ceiling gone? YES + +ModularBot's `pattern` arm wins **18/60 rounds (30 %)** and its mean per-round +rank is **2.07/4** (never sweeping); against the j116 field the same rack won +97–100 % and ranked 1.0. The field now has enough teeth to separate guns *in +principle*: a bad gun would lose more rounds and score less. (Contrast: j116 +`pattern` score/run was ~2965 against the weak field; here it is 992.) + +### 3.2 Arm summary (melee: `score` = server round score = damage + survival bonus) + +| arm | runs | wins/rd | score/run | survival/run | final rank | mean rank | score share | targets | target changes | +|---|---:|---:|---:|---:|---:|---:|---:|---:|---:| +| `pattern` | 12 | **18/60 (30 %)** | **992** | 483 | 1.83 | 2.067 | 31.9 % | 3.00 | 32.58 | +| `bb_decay` | 12 | 17/60 (28 %)| 980 | 483 | 1.83 | **1.900** | 31.6 % | 3.00 | 31.75 | +| `bb_ret` | 12 | 12/60 (20 %)| 868 | 433 | 2.08 | 2.067 | 28.2 % | 3.00 | 33.33 | + +### 3.3 Per-run values (never just the mean) + +``` +pattern wins: r1=2 r2=0 r3=1 r4=4 r5=1 r6=1 r7=2 r8=1 r9=1 r10=0 r11=2 r12=3 + score: r1=1046 r2=625 r3=892 r4=1404 r5=592 r6=800 r7=1331 r8=1051 r9=769 r10=904 r11=1229 r12=1261 + surv: r1=500 r2=400 r3=500 r4=700 r5=250 r6=450 r7=600 r8=450 r9=350 r10=450 r11=600 r12=550 +bb_decay wins: r1=3 r2=0 r3=1 r4=0 r5=2 r6=0 r7=2 r8=2 r9=1 r10=2 r11=2 r12=2 + score: r1=1100 r2=838 r3=820 r4=741 r5=908 r6=635 r7=1198 r8=1181 r9=1089 r10=1194 r11=960 r12=1099 + surv: r1=500 r2=400 r3=400 r4=300 r5=500 r6=300 r7=600 r8=600 r9=500 r10=600 r11=500 r12=600 +bb_ret wins: r1=3 r2=0 r3=1 r4=2 r5=1 r6=2 r7=0 r8=1 r9=0 r10=1 r11=1 r12=0 + score: r1=1167 r2=779 r3=937 r4=1111 r5=815 r6=942 r7=467 r8=944 r9=862 r10=849 r11=704 r12=845 + surv: r1=600 r2=350 r3=500 r4=500 r5=450 r6=400 r7=250 r8=500 r9=450 r10=450 r11=350 r12=400 +``` + +### 3.4 Permutation tests (per-run, two-sided) and MDE + +`diff(A−B)` is `pattern − bitbrain`; a **positive** diff means Pattern is better. + +| metric | A | B | diff(A−B) | perm p | Mann-Whitney p | MDE (n=12) | +|---|---|---|---:|---:|---:|---:| +| score | pattern | bb_decay | +11.75 | 0.903 | 1.000 | ±312.0 (31 % of mean) | +| score | pattern | bb_ret | +123.50 | 0.205 | 0.341 | ±312.0 | +| survival | pattern | bb_decay | +0.00 | 1.000 | 0.883 | ±138.7 | +| survival | pattern | bb_ret | +50.00 | 0.309 | 0.278 | ±138.7 | +| wins | pattern | bb_decay | +0.083 | 1.000 | 0.952 | ±1.34 | +| wins | pattern | bb_ret | +0.500 | 0.352 | 0.288 | ±1.34 | +| mean rank | pattern | bb_decay | +0.167 | 0.552 | 0.542 | ±0.68 | +| mean rank | pattern | bb_ret | −0.000 | 1.000 | 0.954 | ±0.68 | + +No metric separates the arms; the only arm with a nominal lead anywhere is +`bb_decay` on mean per-round rank (+0.167 in its favour, p=0.55). The point +estimates for score and wins favour `pattern` for both BitBrain arms. + +### 3.5 Liveness — the premise WAS exercised + +* **Rack applied, verified per run:** `pattern` reads + `rack active melee = PATTERN` and `TR_RACK_BITBRAIN = off`; both BitBrain arms + read `rack active melee = BITBRAIN` with `TR_RACK_PATTERN = off` and + `TR_RACK_BITBRAIN = melee`. All 36 runs pass (12/12 per arm). +* **Movement held fixed:** `TR_MOVEMENT = strafe (source: env)` in every run. +* **Targets rotate:** **3.00 distinct targets/run** and **31.8–33.3 target + changes/run**; BitBrain resets on every switch (`bb-reset` = 31.75 / 33.33 per + run, equal to the target-change count; `pattern` logs 0). +* **The field is alive:** ModularBot's score share is only 28–32 %, so the + three champions together take ~68 %; ModularBot dies in many rounds (per-run + survival varies widely, 250–700). + +### 3.6 Direct answer — claim 2 + +**NOT SUPPORTED on a finally-fair field.** The j116 ceiling is removed: the new +field is strong enough that ModularBot wins only 30 % of rounds with the +shipped rack, so a real gun advantage had room to show. It did not show. On +score and round wins the point estimates favour the shipped `pattern` rack for +both BitBrain memory modes (`bb_decay` −11.8 score, −0.08 wins; `bb_ret` +−123.5 score, −0.50 wins), and nothing approaches significance. The mechanism +the owner describes was demonstrably exercised (targets rotate, BitBrain resets +on every switch), but it buys no measurable score or win advantage. + +The honest caveat is **power, not fairness**: at 12 runs/arm the score MDE is +±312 on a mean of 992 (≈31 %). A *large* melee advantage (≥ ~31 % score) is +excluded by this run; a smaller one is simply below the resolution of n=12. The +field is fair; the question is now testable and the answer so far is "no". + +--- + +## MEASURED vs INFERRED + +**MEASURED:** the ConstantSpinner fixture's per-tick turn/speed modes (table in +§1); the 108-battle spinner session (commit `6c21dc4`, binary `cfd11b5`) with 0 +failed / 0 never-started / 0 liveness exclusions; the per-opponent, pooled, +cross-opponent, spinner-only, style and convergence tables in §2; the 36-battle +strong-field melee session (commit `5fe2857`, binary `7dc1b7c`) with the arm +summary, per-run values, permutation p-values and MDEs in §3; the liveness boot +lines and target-switch/reset counts; the permutation/MDE machinery (reused from +`tools/ab` and `common_libs/tests/analyze_melee_ab.py`). + +**INFERRED:** the style labels (from `robots.json` / the fixture's own +constants, never decompiled); the reading that "start-lower, catch-up" is a +deficit recovery rather than a fast adaptation; the mechanism by which BitBrain +reaches parity; the interpretation that a melee score effect smaller than the +n=12 MDE would be invisible. + +--- + +## Reproduce + +```sh +# 1. build the spinner and its two rate presets +cd common_libs/test_framework/adversaries/ConstantSpinner +nim c -d:release --nimcache:/tmp/nc_j124 --out:out/ConstantSpinner src/ConstantSpinner.nim +./make_variant.sh ConstantSpinner_s3 3 5 /tmp/tr_spinners +./make_variant.sh ConstantSpinner_s6 6 5 /tmp/tr_spinners + +# 2. the spinner gun A/B (waits for a free arena; ~8 min) +cd /home/davide/Projects/SirRoboGarage +TOURNAMENT_NIMCACHE=/tmp/nc_j124 tools/ab/tournament_run.sh \ + --arms tools/ab/arms_spinner.txt --panel tools/ab/panel_spinner.txt \ + --runs 4 --rounds 5 --conc 6 --wait-arena 45 --outdir /tmp/ab/j124_spinner +python3 tools/ab/spinner_analyze.py /tmp/ab/j124_spinner --reference pattern + +# 3. the fair melee A/B (waits for a free arena; ~8 min) +nim c --nimcache:/tmp/nc_j124 --path:common_libs \ + common_libs/tests/measure_melee_strong_field.nim +MELEE_RUNS=12 MELEE_ROUNDS=5 MELEE_ARMS="pattern bb_ret bb_decay" \ + MELEE_FIELD=/tmp/tr_bots/Diamond,/tmp/tr_bots/Dookious,/tmp/tr_bots/GresSuffurd \ + common_libs/tests/run_melee_strong_field.sh /tmp/melee_strong_field +python3 common_libs/tests/analyze_melee_ab.py /tmp/melee_strong_field --reference pattern +``` + +## Files + +* `common_libs/test_framework/adversaries/ConstantSpinner/` — the fixture + (`src/ConstantSpinner.nim`, JSON, `.sh`, `make_variant.sh`, committed binary). +* `tools/ab/panel_spinner.txt`, `tools/ab/arms_spinner.txt` — the 9-opponent + panel and 3 arms. +* `tools/ab/spinner_analyze.py` — N-arm analyzer: primaries, hit rate, + spinner-only stats, per-round convergence + permutation tests. +* `common_libs/tests/measure_melee_strong_field.nim`, + `common_libs/tests/run_melee_strong_field.sh` — the strong-field melee + harness/driver (§3). diff --git a/tools/ab/spinner_analyze.py b/tools/ab/spinner_analyze.py index 6bb768f..b8321d5 100644 --- a/tools/ab/spinner_analyze.py +++ b/tools/ab/spinner_analyze.py @@ -363,6 +363,30 @@ def main(): f"| {statistics.mean([e['d_hit'] for e in es]):+.2f} |") out() + # ── spinner-only aggregation (the sub-claim's own field) ──────────────── + out("### MEASURED: spinner-only paired stats (the sub-claim's own field)") + out() + out("| arm | metric | mean Δ | spread (SD) | 95% CI | sign test | p(sign) | p(sign-flip) | MDE |") + out("|---|---|---:|---:|---|---:|---:|---:|---:|") + for a in arms: + if a == ref: + continue + for key, mkey in (("d_damage", "damage"), ("d_wins", "wins"), + ("d_hit", "our hit rate (pp)")): + ds = [per_opp[a][o][key] for o in spinner_opps + if per_opp[a][o] and not math.isnan(per_opp[a][o][key])] + if len(ds) < 2: + continue + desc = ga.describe(ds) + pos, neg, ties, p_sign = ga.sign_test(ds) + sf = ga.signflip_perm(ds) + out(f"| `{a}` | {mkey} | {desc['mean']:+.2f} | {desc['sd']:.2f} " + f"| [{desc['mean'] - 1.96 * desc['se']:+.2f}, " + f"{desc['mean'] + 1.96 * desc['se']:+.2f}] " + f"| {pos}/{pos + neg} | {p_sign:.4g} | {sf['p']:.4g} " + f"| {desc['mde']:.2f} |") + out() + # ── CONVERGENCE ───────────────────────────────────────────────────────── out("### MEASURED: CONVERGENCE — our per-round hit rate") out() @@ -447,6 +471,88 @@ def main(): out(f"| `{a}` | n/a | n/a | n/a |") out() + # per-RUN convergence deltas + a two-sample permutation test + def conv_run_deltas(arm): + ds = [] + for o in spinner_opps: + for r in data[o][arm]: + if 1 not in r["per_round"]: + continue + h1 = r["per_round"][1]["mb_hits"] + f1 = r["per_round"][1]["mb_fired"] + hl = sum(d["mb_hits"] for i, d in r["per_round"].items() if i >= 2) + fl = sum(d["mb_fired"] for i, d in r["per_round"].items() if i >= 2) + if f1 and fl: + ds.append(100.0 * hl / fl - 100.0 * h1 / f1) + return ds + + def within_run_deltas(arm): + ds = [] + for o in spinner_opps: + adir = os.path.join(session_dir, o, arm) + for run in discover_runs(adir): + evs = ga.parse_events(os.path.join(adir, f"run{run}.events.jsonl")) + rpath = os.path.join(adir, f"run{run}.jsonl.rounds.json") + starts = read_round_starts(rpath) + counters = ga.parse_counters("".join(ga.read_lines( + os.path.join(adir, f"run{run}.battle.log")))) + if counters is None or not starts: + continue + subj, _ = ga.attribute_subject(evs, counters) + if subj is None: + continue + counts = {rr["round"]: rr["count"] + for rr in json.load(open(rpath))["rounds"]} + e2 = [0, 0] + l2 = [0, 0] + for e in evs: + rnd = e.get("round", 0) + st = starts.get(rnd) + if st is None: + continue + b = e2 if (e.get("tick", 0) - st) < counts.get(rnd, 0) / 2 else l2 + if e.get("type") == "fire" and e.get("owner") == subj: + b[1] += 1 + elif e.get("type") == "hit" and e.get("owner") == subj: + b[0] += 1 + if e2[1] and l2[1]: + ds.append(100.0 * l2[0] / l2[1] - 100.0 * e2[0] / e2[1]) + return ds + + out("Per-run adaptation deltas on the true spinners (the claim is about") + out("SPEED, so these are per-run deltas, not pooled rates):") + out() + out("* `conv` = our hit rate over rounds 2..R minus round 1 (cross-round)") + out("* `within` = our hit rate in the second half of a round minus the first") + out(" half (same round)") + out() + out("| arm | n | conv mean Δ (pp) | conv SD | within mean Δ (pp) | within SD |") + out("|---|---:|---:|---:|---:|---:|") + conv_lists = {} + within_lists = {} + for a in arms: + cv = conv_run_deltas(a) + wi = within_run_deltas(a) + conv_lists[a] = cv + within_lists[a] = wi + out(f"| `{a}` | {len(cv)} | {statistics.mean(cv):+.2f} | " + f"{statistics.pstdev(cv):.2f} | {statistics.mean(wi):+.2f} | " + f"{statistics.pstdev(wi):.2f} |") + out() + out("Two-sample permutation test (200,000 draws, seed 0x5eed5eed) of each") + out(f"arm's adaptation delta against `{ref}`:") + out() + out("| arm | conv Δ − ref Δ (pp) | p(conv) | within Δ − ref Δ (pp) | p(within) |") + out("|---|---:|---:|---:|---:|") + for a in arms: + if a == ref: + continue + dc = statistics.mean(conv_lists[a]) - statistics.mean(conv_lists[ref]) + dw = statistics.mean(within_lists[a]) - statistics.mean(within_lists[ref]) + out(f"| `{a}` | {dc:+.2f} | {perm_two_sample(conv_lists[a], conv_lists[ref]):.4g} " + f"| {dw:+.2f} | {perm_two_sample(within_lists[a], within_lists[ref]):.4g} |") + out() + # ── verdict ───────────────────────────────────────────────────────────── out("### The pre-registered reading") out() @@ -485,6 +591,22 @@ def main(): return 0 +def perm_two_sample(xa, xb, draws=MC_DRAWS, seed=MC_SEED): + """Two-sided two-sample permutation test on the difference of means.""" + if len(xa) < 2 or len(xb) < 2: + return float("nan") + obs = abs(statistics.mean(xa) - statistics.mean(xb)) + pool = list(xa) + list(xb) + na = len(xa) + rng = random.Random(seed) + cnt = 0 + for _ in range(draws): + rng.shuffle(pool) + if abs(statistics.mean(pool[:na]) - statistics.mean(pool[na:])) >= obs - 1e-12: + cnt += 1 + return (cnt + 1) / (draws + 1) + + def wilcoxon_signed(deltas): """Two-sided Wilcoxon signed-rank normal approximation with tie correction.""" nz = [d for d in deltas if d != 0.0]