melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75)
This commit is contained in:
@@ -0,0 +1,179 @@
|
||||
# Movement campaign — ledger
|
||||
|
||||
**Goal (owner's mandate, 2026-09-26 overnight):** find the *best 1v1 movement*
|
||||
by measurement, then do the same for the gun. This file is the campaign's
|
||||
single source of truth: every later job **appends** a `## Batch N` section and
|
||||
never edits an earlier one (a wrong earlier number gets a correction line, not
|
||||
a rewrite).
|
||||
|
||||
**Owner's words:** *"I want you to do all tests and checks with the goal to have
|
||||
the best 1vs1 movement. You have all night, you can change every parameter.
|
||||
Continue until you found an amazing movement. When found do the same over for a
|
||||
gun."*
|
||||
|
||||
---
|
||||
|
||||
## 0. The one caveat this campaign exists to close
|
||||
|
||||
Everything measured about movement before this campaign is **DrussGT-only**:
|
||||
`docs/surfer_wiring_ab.md`, the j107 range drift, the j113 BitBrain movement
|
||||
notes. The standing lesson of the night is that a one-opponent result is not a
|
||||
result:
|
||||
|
||||
> an arm can take fewer hits **and** win fewer rounds (j107 / `strafe`): the
|
||||
> verdict lives in **damage/run + ROUND WINS**, and hit rate is only ever an
|
||||
> explanation.
|
||||
|
||||
So from here on **the unit of evidence is the number of opponents**, not the
|
||||
number of runs: the same arm must win on *many* opponents before it is called
|
||||
better.
|
||||
|
||||
---
|
||||
|
||||
## 1. Protocol (how every batch must be run)
|
||||
|
||||
| Element | Rule |
|
||||
|---|---|
|
||||
| Subject | ONE frozen binary, built from `git archive HEAD` (`tools/ab/tournament_run.sh` does this; the commit sha and binary sha256 are recorded in `session.json`) |
|
||||
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file, not a shell history |
|
||||
| Panel | the **frozen** panel `tools/ab/panel_movement.txt`. Adding/removing an opponent starts a **new batch number** |
|
||||
| Pairing | per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate |
|
||||
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
|
||||
| Serialization | **one battle fleet at a time.** `tournament_run.sh --wait-arena N` refuses/stalls while another job's `run_bridge_battle`/`TrBattleCapture`/`ModularBot_bin` is alive (bracketed pgrep; never a broad `pkill`) |
|
||||
| Liveness | every declared env token must appear verbatim in OUR bot's own `[env]` boot report, else the run is excluded and named in the report; an undeclared `TR_MOVEMENT` in the process env is a fatal FAIL for the reference arm |
|
||||
| Never shipped | this is a measurement + design campaign: `git status` clean, defaults untouched, `.gitignore` untouched |
|
||||
|
||||
### Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit `__PRECOMMIT__`)
|
||||
|
||||
1. **Primary metrics:** damage/run and ROUND WINS. Secondary/explanation only:
|
||||
damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean
|
||||
distance.
|
||||
2. **BETTER than the reference** iff one primary metric is up with a
|
||||
cross-opponent **sign test p < 0.05** while the other does **not** go down;
|
||||
or the mirror image for **WORSE**. Anything else is **NOT
|
||||
DISTINGUISHABLE** (which is a real answer, not a failure).
|
||||
3. **A verdict must survive the between-opponent spread**: the pooled mean delta
|
||||
is reported with the SD across opponents, its SE, a 95% CI, and the MDE
|
||||
(α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as
|
||||
*not detectable*, never as *absent* and never as a win.
|
||||
4. **Somewhere to stop:** if no arm beats the shipped `tfil` by rule 2 in
|
||||
Batch 1 **and** no arm shows a ≥ +MDE damage gain with p<0.10, the movement
|
||||
stage's first phase is closed with *"the shipped `tfil` is the best movement
|
||||
we have measured"* — that is a **successful** outcome, and the campaign moves
|
||||
to the gun axis rather than inventing more movement arms. See
|
||||
*What would make us stop* at the end.
|
||||
5. **No promotion off a single metric, a single opponent, or a single run.**
|
||||
A change that wins damage by losing wins (or vice-versa) is not a win.
|
||||
6. Every batch is shot with a **pre-registered prediction** stated in its
|
||||
section *before* the battles finish; a prediction that turns out wrong is
|
||||
recorded as wrong.
|
||||
|
||||
---
|
||||
|
||||
## 2. Stage 0 — what we already know (given, not re-derived)
|
||||
|
||||
Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary
|
||||
(`docs/surfer_wiring_ab.md`, commit `0f5cfe3`):
|
||||
|
||||
| arm | dmg/run | dmg taken | round wins | incoming hit rate |
|
||||
|---|---:|---:|---:|---:|
|
||||
| `tfil` (SHIPPED) | **293** | 224 | **45/105** | 10.40% |
|
||||
| `strafe` (range 325) | 250 | **198** | 37/105 | **9.40%** |
|
||||
| `surf` | 255 | 259 | 37/105 | 13.51% |
|
||||
|
||||
Read: the shipped `tfil` deals the most damage and wins the most rounds while
|
||||
being hit the *most*; `strafe` dodges best and wins least. Plus j107: drifting
|
||||
25–30 px closer made damage **and** wins worse, so the lever is not simply "get
|
||||
closer". **Hypothesis entering the campaign: the 325 px range preference of
|
||||
`strafe` costs wins** (INFERRED from DrussGT-only data — this is exactly what
|
||||
Batch 1 tests across a panel).
|
||||
|
||||
---
|
||||
|
||||
## 3. Batch 1 — isolating the range / aggression axis
|
||||
|
||||
**Design.** One frozen binary, five env-only arms, one frozen panel
|
||||
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 2
|
||||
wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas),
|
||||
3 runs × 3 rounds per (opponent, arm). Arm file:
|
||||
`tools/ab/arms_movement_b1.txt`.
|
||||
|
||||
| # | arm | env | what it isolates |
|
||||
|---|---|---|---|
|
||||
| 1 | `tfil` | *(none — shipped defaults)* | the arm to beat |
|
||||
| 2 | `strafe_notilt` | `TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999` | the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with **no range steering at all** |
|
||||
| 3 | `strafe_325` | `TR_MOVEMENT=strafe` | the current strafe default (range 325, tol 25, tilt 15/0.10) |
|
||||
| 4 | `ring` | `TR_MOVEMENT=tfil_ring` | TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) **with** the range-weighted tile draw (band 100–200) |
|
||||
| 5 | `ring_notemp` | `TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0` | the control for #4: same retuned heat field, range weighting switched OFF (`rand(candidates.high)` path) |
|
||||
|
||||
`ring` − `ring_notemp` is therefore the range-weighting lever **alone**, on a
|
||||
heat field that is already retuned. The originally-suggested 5th arm ("`tfil`
|
||||
with less saturated heat") is **not buildable in this campaign**: in
|
||||
`common_libs/movements/the_floor_is_lava.nim` `CorridorHeat`/`WallHotness` are
|
||||
Nim `const`s (env_report only *reports* them); only the `tfil_ring` copy reads
|
||||
them from the env. #5 is the honest substitute.
|
||||
|
||||
**Pre-registered prediction (written before the battles finished):** `tfil`
|
||||
still wins the panel on damage and round wins; `strafe_notilt` will beat
|
||||
`strafe_325` on round wins (the range tilt is a net cost), and the ring arms will
|
||||
land between them. If instead the range-steering arms beat `tfil` on wins, the
|
||||
"range preference costs wins" hypothesis is confirmed across bots, not just
|
||||
against DrussGT.
|
||||
|
||||
<!-- BATCH1-RESULTS -->
|
||||
|
||||
---
|
||||
|
||||
## 4. What to try next (seeded; every later job adds its own)
|
||||
|
||||
1. **If a range-steering arm wins the panel:** the win is a *range* effect, so
|
||||
sweep the *band* on the winning engine (e.g. `TR_TFIL_RANGE_LO/HI` on
|
||||
`tfil_ring`, `TR_STRAFE_RANGE` on `strafe`) with the same panel, 4–5 bands,
|
||||
and look for a plateau rather than a peak. A plateau is a result; a peak is a
|
||||
coin flip.
|
||||
2. **If nothing beats `tfil`:** stop tuning movement by feel. The next honest
|
||||
lever is *enemy-model-driven* placement (keep the bot where the enemy's
|
||||
expected hit probability is lowest *given its gun model*), which needs a
|
||||
per-opponent measurement, not a knob.
|
||||
3. **Melee is a different game** (j116's finding): if a melee campaign is
|
||||
opened, it needs its own panel and its own ledger section — do not reuse the
|
||||
1v1 panel's verdicts.
|
||||
4. **Close the loop with the enemy's own model:** `ab_mechanism.py`-style
|
||||
instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by
|
||||
range band) is available and cheap — use it to *explain* a win, never to
|
||||
declare one.
|
||||
5. **Then the gun** (owner's next stage): the same harness, a gun panel, and the
|
||||
same paired-with-sign-test statistics. `docs/surfer_wiring_ab.md` and the j117
|
||||
gauntlet are the gun-side priors to beat.
|
||||
|
||||
## 5. What would make us stop
|
||||
|
||||
* **Stop the movement stage** when a batch produces an arm that is BETTER than
|
||||
the shipped default on the frozen panel by rule 2 **and** the effect survives
|
||||
the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of
|
||||
the panel). That arm becomes the new default *candidate* (shipping is a
|
||||
separate decision — this campaign never edits a shipped default).
|
||||
* **Stop and move to the gun** if two consecutive batches fail to produce an arm
|
||||
that beats the shipped `tfil` beyond the MDE: at that point the honest
|
||||
conclusion is *"the shipped movement is the measured optimum of this design
|
||||
space"*, which is a successful campaign outcome, not a failure.
|
||||
* **Stop a single batch early** only for a contract violation (arena not free,
|
||||
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
|
||||
|
||||
## 6. How to run a batch (exact commands)
|
||||
|
||||
```sh
|
||||
# 1. wait for the arena (this job may not be the only one fighting)
|
||||
tools/ab/tournament_run.sh \
|
||||
--arms tools/ab/arms_movement_b1.txt \
|
||||
--panel tools/ab/panel_movement.txt \
|
||||
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
|
||||
--reference tfil \
|
||||
--outdir /tmp/ab/j118_b1
|
||||
|
||||
# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
|
||||
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil
|
||||
```
|
||||
|
||||
The runner writes `<outdir>/session.json` (commit sha, binary sha256, arms,
|
||||
panel) so any later job can re-analyze an old session offline, with no arena.
|
||||
Reference in New Issue
Block a user