From 1984a780f494ce246e0f916934b9581e07c89ed2 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sat, 26 Sep 2026 00:43:04 +0200 Subject: [PATCH] melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75) --- docs/melee_bitbrain_ab.md | 6 +- docs/movement_campaign.md | 179 +++++++++ tools/ab/arms_movement_b1.txt | 44 +++ tools/ab/panel_movement.txt | 71 ++++ tools/ab/tournament_analyze.py | 639 +++++++++++++++++++++++++++++++++ tools/ab/tournament_run.sh | 406 +++++++++++++++++++++ 6 files changed, 1342 insertions(+), 3 deletions(-) create mode 100644 docs/movement_campaign.md create mode 100644 tools/ab/arms_movement_b1.txt create mode 100644 tools/ab/panel_movement.txt create mode 100644 tools/ab/tournament_analyze.py create mode 100755 tools/ab/tournament_run.sh diff --git a/docs/melee_bitbrain_ab.md b/docs/melee_bitbrain_ab.md index 3d21365..7d7dde4 100644 --- a/docs/melee_bitbrain_ab.md +++ b/docs/melee_bitbrain_ab.md @@ -186,9 +186,9 @@ Target switching is the whole mechanism, so it had to be proven, not assumed. **66–69 target changes/run** (about 10 per round). Example transitions from one run: `#1 → #4 → #2 → #4 → #1 → #2 → …`. * **BitBrain really reset on every switch.** `[bb-reset] reason=target_change` - fires once per target change: **66.75** (perRound), **68.19** (retained), - **66.75** (decay) per run — equal, within rounding, to the target-change count. - Example `[bb]` line from a live run: + fires once per target change: **68.75** (perRound), **68.19** (retained), + **66.75** (decay) per run — equal, within rounding, to each arm's target-change + count. Example `[bb]` line from a live run: `[bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound`. The `pattern` arm logs **zero** `[bb]`/`[bb-reset]` lines (BitBrain never spawned), confirming the arms are cleanly separated. diff --git a/docs/movement_campaign.md b/docs/movement_campaign.md new file mode 100644 index 0000000..8b171ff --- /dev/null +++ b/docs/movement_campaign.md @@ -0,0 +1,179 @@ +# Movement campaign — ledger + +**Goal (owner's mandate, 2026-09-26 overnight):** find the *best 1v1 movement* +by measurement, then do the same for the gun. This file is the campaign's +single source of truth: every later job **appends** a `## Batch N` section and +never edits an earlier one (a wrong earlier number gets a correction line, not +a rewrite). + +**Owner's words:** *"I want you to do all tests and checks with the goal to have +the best 1vs1 movement. You have all night, you can change every parameter. +Continue until you found an amazing movement. When found do the same over for a +gun."* + +--- + +## 0. The one caveat this campaign exists to close + +Everything measured about movement before this campaign is **DrussGT-only**: +`docs/surfer_wiring_ab.md`, the j107 range drift, the j113 BitBrain movement +notes. The standing lesson of the night is that a one-opponent result is not a +result: + +> an arm can take fewer hits **and** win fewer rounds (j107 / `strafe`): the +> verdict lives in **damage/run + ROUND WINS**, and hit rate is only ever an +> explanation. + +So from here on **the unit of evidence is the number of opponents**, not the +number of runs: the same arm must win on *many* opponents before it is called +better. + +--- + +## 1. Protocol (how every batch must be run) + +| Element | Rule | +|---|---| +| Subject | ONE frozen binary, built from `git archive HEAD` (`tools/ab/tournament_run.sh` does this; the commit sha and binary sha256 are recorded in `session.json`) | +| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file, not a shell history | +| Panel | the **frozen** panel `tools/ab/panel_movement.txt`. Adding/removing an opponent starts a **new batch number** | +| Pairing | per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate | +| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir | +| Serialization | **one battle fleet at a time.** `tournament_run.sh --wait-arena N` refuses/stalls while another job's `run_bridge_battle`/`TrBattleCapture`/`ModularBot_bin` is alive (bracketed pgrep; never a broad `pkill`) | +| Liveness | every declared env token must appear verbatim in OUR bot's own `[env]` boot report, else the run is excluded and named in the report; an undeclared `TR_MOVEMENT` in the process env is a fatal FAIL for the reference arm | +| Never shipped | this is a measurement + design campaign: `git status` clean, defaults untouched, `.gitignore` untouched | + +### Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit `__PRECOMMIT__`) + +1. **Primary metrics:** damage/run and ROUND WINS. Secondary/explanation only: + damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean + distance. +2. **BETTER than the reference** iff one primary metric is up with a + cross-opponent **sign test p < 0.05** while the other does **not** go down; + or the mirror image for **WORSE**. Anything else is **NOT + DISTINGUISHABLE** (which is a real answer, not a failure). +3. **A verdict must survive the between-opponent spread**: the pooled mean delta + is reported with the SD across opponents, its SE, a 95% CI, and the MDE + (α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as + *not detectable*, never as *absent* and never as a win. +4. **Somewhere to stop:** if no arm beats the shipped `tfil` by rule 2 in + Batch 1 **and** no arm shows a ≥ +MDE damage gain with p<0.10, the movement + stage's first phase is closed with *"the shipped `tfil` is the best movement + we have measured"* — that is a **successful** outcome, and the campaign moves + to the gun axis rather than inventing more movement arms. See + *What would make us stop* at the end. +5. **No promotion off a single metric, a single opponent, or a single run.** + A change that wins damage by losing wins (or vice-versa) is not a win. +6. Every batch is shot with a **pre-registered prediction** stated in its + section *before* the battles finish; a prediction that turns out wrong is + recorded as wrong. + +--- + +## 2. Stage 0 — what we already know (given, not re-derived) + +Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary +(`docs/surfer_wiring_ab.md`, commit `0f5cfe3`): + +| arm | dmg/run | dmg taken | round wins | incoming hit rate | +|---|---:|---:|---:|---:| +| `tfil` (SHIPPED) | **293** | 224 | **45/105** | 10.40% | +| `strafe` (range 325) | 250 | **198** | 37/105 | **9.40%** | +| `surf` | 255 | 259 | 37/105 | 13.51% | + +Read: the shipped `tfil` deals the most damage and wins the most rounds while +being hit the *most*; `strafe` dodges best and wins least. Plus j107: drifting +25–30 px closer made damage **and** wins worse, so the lever is not simply "get +closer". **Hypothesis entering the campaign: the 325 px range preference of +`strafe` costs wins** (INFERRED from DrussGT-only data — this is exactly what +Batch 1 tests across a panel). + +--- + +## 3. Batch 1 — isolating the range / aggression axis + +**Design.** One frozen binary, five env-only arms, one frozen panel +(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 2 +wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas), +3 runs × 3 rounds per (opponent, arm). Arm file: +`tools/ab/arms_movement_b1.txt`. + +| # | arm | env | what it isolates | +|---|---|---|---| +| 1 | `tfil` | *(none — shipped defaults)* | the arm to beat | +| 2 | `strafe_notilt` | `TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999` | the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with **no range steering at all** | +| 3 | `strafe_325` | `TR_MOVEMENT=strafe` | the current strafe default (range 325, tol 25, tilt 15/0.10) | +| 4 | `ring` | `TR_MOVEMENT=tfil_ring` | TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) **with** the range-weighted tile draw (band 100–200) | +| 5 | `ring_notemp` | `TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0` | the control for #4: same retuned heat field, range weighting switched OFF (`rand(candidates.high)` path) | + +`ring` − `ring_notemp` is therefore the range-weighting lever **alone**, on a +heat field that is already retuned. The originally-suggested 5th arm ("`tfil` +with less saturated heat") is **not buildable in this campaign**: in +`common_libs/movements/the_floor_is_lava.nim` `CorridorHeat`/`WallHotness` are +Nim `const`s (env_report only *reports* them); only the `tfil_ring` copy reads +them from the env. #5 is the honest substitute. + +**Pre-registered prediction (written before the battles finished):** `tfil` +still wins the panel on damage and round wins; `strafe_notilt` will beat +`strafe_325` on round wins (the range tilt is a net cost), and the ring arms will +land between them. If instead the range-steering arms beat `tfil` on wins, the +"range preference costs wins" hypothesis is confirmed across bots, not just +against DrussGT. + + + +--- + +## 4. What to try next (seeded; every later job adds its own) + +1. **If a range-steering arm wins the panel:** the win is a *range* effect, so + sweep the *band* on the winning engine (e.g. `TR_TFIL_RANGE_LO/HI` on + `tfil_ring`, `TR_STRAFE_RANGE` on `strafe`) with the same panel, 4–5 bands, + and look for a plateau rather than a peak. A plateau is a result; a peak is a + coin flip. +2. **If nothing beats `tfil`:** stop tuning movement by feel. The next honest + lever is *enemy-model-driven* placement (keep the bot where the enemy's + expected hit probability is lowest *given its gun model*), which needs a + per-opponent measurement, not a knob. +3. **Melee is a different game** (j116's finding): if a melee campaign is + opened, it needs its own panel and its own ledger section — do not reuse the + 1v1 panel's verdicts. +4. **Close the loop with the enemy's own model:** `ab_mechanism.py`-style + instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by + range band) is available and cheap — use it to *explain* a win, never to + declare one. +5. **Then the gun** (owner's next stage): the same harness, a gun panel, and the + same paired-with-sign-test statistics. `docs/surfer_wiring_ab.md` and the j117 + gauntlet are the gun-side priors to beat. + +## 5. What would make us stop + +* **Stop the movement stage** when a batch produces an arm that is BETTER than + the shipped default on the frozen panel by rule 2 **and** the effect survives + the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of + the panel). That arm becomes the new default *candidate* (shipping is a + separate decision — this campaign never edits a shipped default). +* **Stop and move to the gun** if two consecutive batches fail to produce an arm + that beats the shipped `tfil` beyond the MDE: at that point the honest + conclusion is *"the shipped movement is the measured optimum of this design + space"*, which is a successful campaign outcome, not a failure. +* **Stop a single batch early** only for a contract violation (arena not free, + liveness FAIL, non-zero exit rate) — never because the numbers look boring. + +## 6. How to run a batch (exact commands) + +```sh +# 1. wait for the arena (this job may not be the only one fighting) +tools/ab/tournament_run.sh \ + --arms tools/ab/arms_movement_b1.txt \ + --panel tools/ab/panel_movement.txt \ + --runs 3 --rounds 3 --conc 6 --wait-arena 45 \ + --reference tfil \ + --outdir /tmp/ab/j118_b1 + +# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict +python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil +``` + +The runner writes `/session.json` (commit sha, binary sha256, arms, +panel) so any later job can re-analyze an old session offline, with no arena. diff --git a/tools/ab/arms_movement_b1.txt b/tools/ab/arms_movement_b1.txt new file mode 100644 index 0000000..5e35f39 --- /dev/null +++ b/tools/ab/arms_movement_b1.txt @@ -0,0 +1,44 @@ +# ───────────────────────────────────────────────────────────────────────────── +# arms_movement_b1.txt — BATCH 1 of the movement campaign: isolate the +# RANGE / AGGRESSION axis. Five arms, ONE frozen binary, env vars only. +# +# Format: name | ENV=value ENV=value | label +# +# Prior data this batch is built on (all DrussGT-only, all MEASURED): +# * shipped `tfil` 293 dmg/run, 45/105 round wins, 10.40% incoming hit rate +# * `strafe` (range 325) 250 dmg/run, 37/105 wins, 9.40% incoming -> best +# dodger, fewest wins: the 325px range preference may COST wins +# * j107: drifting 25-30 px closer made damage AND wins worse, so "get +# closer" is not the lever by itself +# This batch closes the caveat that all of that is DrussGT-only and decomposes +# the range-steering lever from the engine. +# ───────────────────────────────────────────────────────────────────────────── + +# 1. the arm to beat: shipped default, no env at all +tfil | | shipped baseline (movement engine tfil, every knob at its default) + +# 2. pure perpendicular strafe with the range steering DISABLED: tilt is zero +# whenever |dist - StrafeRange| <= TOL, so a huge TOL makes `tilt` provably +# 0 every tick (and the approach-direction branch that depends on tiltMag is +# skipped). This isolates the COST of the 325px preference: same engine, same +# wall logic, no range steering at all. +strafe_notilt | TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999 | strafe, range steering OFF (tilt always 0) + +# 3. the current strafe default: range 325, tol 25, tilt max 15, gain 0.10 +strafe_325 | TR_MOVEMENT=strafe | strafe default (range 325, tol 25, tilt 15/0.10) + +# 4. the owner's own retuned mover: tfil semantics + retuned heat field +# (CorridorHeat 10 = the safety threshold, WallHotness 15, WallRadiance 5, +# BulletCore/Aura 20/10, commit 5 ticks) + range-weighted tile draw +# (band 100-200, temp 0.4, K 60). Range steering ON, different heat field. +ring | TR_MOVEMENT=tfil_ring | tfil_ring (retuned heat field + range weighting 100-200) + +# 5. the control for arm 4: TR_TFIL_RANGE_TEMP=0 makes the tile draw UNIFORM +# again (the documented OFF path), so arm 5 = the retuned heat field with NO +# range steering. (4) - (5) is therefore the range-weighting lever alone, on +# a heat field that is already retuned. NOTE: there is no env knob for the +# heat constants of the SHIPPED tfil mover (CorridorHeat/WallHotness are +# Nim `const` there, only REPORTED by env_report.nim), so the suggested +# "tfil with less saturated heat" arm cannot be built without a rebuild, +# which this campaign forbids. This is the honest substitute. +ring_notemp | TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0 | tfil_ring, range weighting OFF (temp 0) diff --git a/tools/ab/panel_movement.txt b/tools/ab/panel_movement.txt new file mode 100644 index 0000000..aa60c0d --- /dev/null +++ b/tools/ab/panel_movement.txt @@ -0,0 +1,71 @@ +# ───────────────────────────────────────────────────────────────────────────── +# panel_movement.txt — THE FROZEN MOVEMENT PANEL (Batch 1, 2026-09-26) +# +# DO NOT ADD, REMOVE OR REORDER AN OPPONENT without starting a new batch number +# in docs/movement_campaign.md. Every later job in the movement campaign must +# measure against THIS list: the unit of evidence is the number of opponents, so +# changing the panel changes what "better movement" means. +# +# Format (one opponent per line): +# |