Files
SirRoboGarage/docs/movement_campaign.md
T

180 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Movement campaign — ledger
**Goal (owner's mandate, 2026-09-26 overnight):** find the *best 1v1 movement*
by measurement, then do the same for the gun. This file is the campaign's
single source of truth: every later job **appends** a `## Batch N` section and
never edits an earlier one (a wrong earlier number gets a correction line, not
a rewrite).
**Owner's words:** *"I want you to do all tests and checks with the goal to have
the best 1vs1 movement. You have all night, you can change every parameter.
Continue until you found an amazing movement. When found do the same over for a
gun."*
---
## 0. The one caveat this campaign exists to close
Everything measured about movement before this campaign is **DrussGT-only**:
`docs/surfer_wiring_ab.md`, the j107 range drift, the j113 BitBrain movement
notes. The standing lesson of the night is that a one-opponent result is not a
result:
> an arm can take fewer hits **and** win fewer rounds (j107 / `strafe`): the
> verdict lives in **damage/run + ROUND WINS**, and hit rate is only ever an
> explanation.
So from here on **the unit of evidence is the number of opponents**, not the
number of runs: the same arm must win on *many* opponents before it is called
better.
---
## 1. Protocol (how every batch must be run)
| Element | Rule |
|---|---|
| Subject | ONE frozen binary, built from `git archive HEAD` (`tools/ab/tournament_run.sh` does this; the commit sha and binary sha256 are recorded in `session.json`) |
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file, not a shell history |
| Panel | the **frozen** panel `tools/ab/panel_movement.txt`. Adding/removing an opponent starts a **new batch number** |
| Pairing | per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | **one battle fleet at a time.** `tournament_run.sh --wait-arena N` refuses/stalls while another job's `run_bridge_battle`/`TrBattleCapture`/`ModularBot_bin` is alive (bracketed pgrep; never a broad `pkill`) |
| Liveness | every declared env token must appear verbatim in OUR bot's own `[env]` boot report, else the run is excluded and named in the report; an undeclared `TR_MOVEMENT` in the process env is a fatal FAIL for the reference arm |
| Never shipped | this is a measurement + design campaign: `git status` clean, defaults untouched, `.gitignore` untouched |
### Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit `__PRECOMMIT__`)
1. **Primary metrics:** damage/run and ROUND WINS. Secondary/explanation only:
damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean
distance.
2. **BETTER than the reference** iff one primary metric is up with a
cross-opponent **sign test p < 0.05** while the other does **not** go down;
or the mirror image for **WORSE**. Anything else is **NOT
DISTINGUISHABLE** (which is a real answer, not a failure).
3. **A verdict must survive the between-opponent spread**: the pooled mean delta
is reported with the SD across opponents, its SE, a 95% CI, and the MDE
(α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as
*not detectable*, never as *absent* and never as a win.
4. **Somewhere to stop:** if no arm beats the shipped `tfil` by rule 2 in
Batch 1 **and** no arm shows a ≥ +MDE damage gain with p<0.10, the movement
stage's first phase is closed with *"the shipped `tfil` is the best movement
we have measured"* — that is a **successful** outcome, and the campaign moves
to the gun axis rather than inventing more movement arms. See
*What would make us stop* at the end.
5. **No promotion off a single metric, a single opponent, or a single run.**
A change that wins damage by losing wins (or vice-versa) is not a win.
6. Every batch is shot with a **pre-registered prediction** stated in its
section *before* the battles finish; a prediction that turns out wrong is
recorded as wrong.
---
## 2. Stage 0 — what we already know (given, not re-derived)
Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary
(`docs/surfer_wiring_ab.md`, commit `0f5cfe3`):
| arm | dmg/run | dmg taken | round wins | incoming hit rate |
|---|---:|---:|---:|---:|
| `tfil` (SHIPPED) | **293** | 224 | **45/105** | 10.40% |
| `strafe` (range 325) | 250 | **198** | 37/105 | **9.40%** |
| `surf` | 255 | 259 | 37/105 | 13.51% |
Read: the shipped `tfil` deals the most damage and wins the most rounds while
being hit the *most*; `strafe` dodges best and wins least. Plus j107: drifting
25–30 px closer made damage **and** wins worse, so the lever is not simply "get
closer". **Hypothesis entering the campaign: the 325 px range preference of
`strafe` costs wins** (INFERRED from DrussGT-only data — this is exactly what
Batch 1 tests across a panel).
---
## 3. Batch 1 — isolating the range / aggression axis
**Design.** One frozen binary, five env-only arms, one frozen panel
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 2
wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas),
3 runs × 3 rounds per (opponent, arm). Arm file:
`tools/ab/arms_movement_b1.txt`.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | `tfil` | *(none — shipped defaults)* | the arm to beat |
| 2 | `strafe_notilt` | `TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999` | the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with **no range steering at all** |
| 3 | `strafe_325` | `TR_MOVEMENT=strafe` | the current strafe default (range 325, tol 25, tilt 15/0.10) |
| 4 | `ring` | `TR_MOVEMENT=tfil_ring` | TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) **with** the range-weighted tile draw (band 100–200) |
| 5 | `ring_notemp` | `TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0` | the control for #4: same retuned heat field, range weighting switched OFF (`rand(candidates.high)` path) |
`ring` − `ring_notemp` is therefore the range-weighting lever **alone**, on a
heat field that is already retuned. The originally-suggested 5th arm ("`tfil`
with less saturated heat") is **not buildable in this campaign**: in
`common_libs/movements/the_floor_is_lava.nim` `CorridorHeat`/`WallHotness` are
Nim `const`s (env_report only *reports* them); only the `tfil_ring` copy reads
them from the env. #5 is the honest substitute.
**Pre-registered prediction (written before the battles finished):** `tfil`
still wins the panel on damage and round wins; `strafe_notilt` will beat
`strafe_325` on round wins (the range tilt is a net cost), and the ring arms will
land between them. If instead the range-steering arms beat `tfil` on wins, the
"range preference costs wins" hypothesis is confirmed across bots, not just
against DrussGT.
<!-- BATCH1-RESULTS -->
---
## 4. What to try next (seeded; every later job adds its own)
1. **If a range-steering arm wins the panel:** the win is a *range* effect, so
sweep the *band* on the winning engine (e.g. `TR_TFIL_RANGE_LO/HI` on
`tfil_ring`, `TR_STRAFE_RANGE` on `strafe`) with the same panel, 4–5 bands,
and look for a plateau rather than a peak. A plateau is a result; a peak is a
coin flip.
2. **If nothing beats `tfil`:** stop tuning movement by feel. The next honest
lever is *enemy-model-driven* placement (keep the bot where the enemy's
expected hit probability is lowest *given its gun model*), which needs a
per-opponent measurement, not a knob.
3. **Melee is a different game** (j116's finding): if a melee campaign is
opened, it needs its own panel and its own ledger section — do not reuse the
1v1 panel's verdicts.
4. **Close the loop with the enemy's own model:** `ab_mechanism.py`-style
instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by
range band) is available and cheap — use it to *explain* a win, never to
declare one.
5. **Then the gun** (owner's next stage): the same harness, a gun panel, and the
same paired-with-sign-test statistics. `docs/surfer_wiring_ab.md` and the j117
gauntlet are the gun-side priors to beat.
## 5. What would make us stop
* **Stop the movement stage** when a batch produces an arm that is BETTER than
the shipped default on the frozen panel by rule 2 **and** the effect survives
the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of
the panel). That arm becomes the new default *candidate* (shipping is a
separate decision — this campaign never edits a shipped default).
* **Stop and move to the gun** if two consecutive batches fail to produce an arm
that beats the shipped `tfil` beyond the MDE: at that point the honest
conclusion is *"the shipped movement is the measured optimum of this design
space"*, which is a successful campaign outcome, not a failure.
* **Stop a single batch early** only for a contract violation (arena not free,
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
## 6. How to run a batch (exact commands)
```sh
# 1. wait for the arena (this job may not be the only one fighting)
tools/ab/tournament_run.sh \
--arms tools/ab/arms_movement_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference tfil \
--outdir /tmp/ab/j118_b1
# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil
```
The runner writes `<outdir>/session.json` (commit sha, binary sha256, arms,
panel) so any later job can re-analyze an old session offline, with no arena.