Files
SirRoboGarage/docs/movement_campaign.md
T

9.7 KiB
Raw Blame History

Movement campaign — ledger

Goal (owner's mandate, 2026-09-26 overnight): find the best 1v1 movement by measurement, then do the same for the gun. This file is the campaign's single source of truth: every later job appends a ## Batch N section and never edits an earlier one (a wrong earlier number gets a correction line, not a rewrite).

Owner's words: "I want you to do all tests and checks with the goal to have the best 1vs1 movement. You have all night, you can change every parameter. Continue until you found an amazing movement. When found do the same over for a gun."


0. The one caveat this campaign exists to close

Everything measured about movement before this campaign is DrussGT-only: docs/surfer_wiring_ab.md, the j107 range drift, the j113 BitBrain movement notes. The standing lesson of the night is that a one-opponent result is not a result:

an arm can take fewer hits and win fewer rounds (j107 / strafe): the verdict lives in damage/run + ROUND WINS, and hit rate is only ever an explanation.

So from here on the unit of evidence is the number of opponents, not the number of runs: the same arm must win on many opponents before it is called better.


1. Protocol (how every batch must be run)

Element Rule
Subject ONE frozen binary, built from git archive HEAD (tools/ab/tournament_run.sh does this; the commit sha and binary sha256 are recorded in session.json)
Arms env dicts only — no per-arm rebuild, ever; the arm file is a committed file, not a shell history
Panel the frozen panel tools/ab/panel_movement.txt. Adding/removing an opponent starts a new batch number
Pairing per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate
Isolation per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir
Serialization one battle fleet at a time. tournament_run.sh --wait-arena N refuses/stalls while another job's run_bridge_battle/TrBattleCapture/ModularBot_bin is alive (bracketed pgrep; never a broad pkill)
Liveness every declared env token must appear verbatim in OUR bot's own [env] boot report, else the run is excluded and named in the report; an undeclared TR_MOVEMENT in the process env is a fatal FAIL for the reference arm
Never shipped this is a measurement + design campaign: git status clean, defaults untouched, .gitignore untouched

Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit __PRECOMMIT__)

  1. Primary metrics: damage/run and ROUND WINS. Secondary/explanation only: damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean distance.
  2. BETTER than the reference iff one primary metric is up with a cross-opponent sign test p < 0.05 while the other does not go down; or the mirror image for WORSE. Anything else is NOT DISTINGUISHABLE (which is a real answer, not a failure).
  3. A verdict must survive the between-opponent spread: the pooled mean delta is reported with the SD across opponents, its SE, a 95% CI, and the MDE (α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as not detectable, never as absent and never as a win.
  4. Somewhere to stop: if no arm beats the shipped tfil by rule 2 in Batch 1 and no arm shows a ≥ +MDE damage gain with p<0.10, the movement stage's first phase is closed with "the shipped tfil is the best movement we have measured" — that is a successful outcome, and the campaign moves to the gun axis rather than inventing more movement arms. See What would make us stop at the end.
  5. No promotion off a single metric, a single opponent, or a single run. A change that wins damage by losing wins (or vice-versa) is not a win.
  6. Every batch is shot with a pre-registered prediction stated in its section before the battles finish; a prediction that turns out wrong is recorded as wrong.

2. Stage 0 — what we already know (given, not re-derived)

Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary (docs/surfer_wiring_ab.md, commit 0f5cfe3):

arm dmg/run dmg taken round wins incoming hit rate
tfil (SHIPPED) 293 224 45/105 10.40%
strafe (range 325) 250 198 37/105 9.40%
surf 255 259 37/105 13.51%

Read: the shipped tfil deals the most damage and wins the most rounds while being hit the most; strafe dodges best and wins least. Plus j107: drifting 25–30 px closer made damage and wins worse, so the lever is not simply "get closer". Hypothesis entering the campaign: the 325 px range preference of strafe costs wins (INFERRED from DrussGT-only data — this is exactly what Batch 1 tests across a panel).


3. Batch 1 — isolating the range / aggression axis

Design. One frozen binary, five env-only arms, one frozen panel (tools/ab/panel_movement.txt, 15 opponents: 5 dodger, 3 pattern, 2 wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas), 3 runs × 3 rounds per (opponent, arm). Arm file: tools/ab/arms_movement_b1.txt.

# arm env what it isolates
1 tfil (none — shipped defaults) the arm to beat
2 strafe_notilt TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999 the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with no range steering at all
3 strafe_325 TR_MOVEMENT=strafe the current strafe default (range 325, tol 25, tilt 15/0.10)
4 ring TR_MOVEMENT=tfil_ring TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) with the range-weighted tile draw (band 100–200)
5 ring_notemp TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0 the control for #4: same retuned heat field, range weighting switched OFF (rand(candidates.high) path)

ring − ring_notemp is therefore the range-weighting lever alone, on a heat field that is already retuned. The originally-suggested 5th arm ("tfil with less saturated heat") is not buildable in this campaign: in common_libs/movements/the_floor_is_lava.nim CorridorHeat/WallHotness are Nim consts (env_report only reports them); only the tfil_ring copy reads them from the env. #5 is the honest substitute.

Pre-registered prediction (written before the battles finished): tfil still wins the panel on damage and round wins; strafe_notilt will beat strafe_325 on round wins (the range tilt is a net cost), and the ring arms will land between them. If instead the range-steering arms beat tfil on wins, the "range preference costs wins" hypothesis is confirmed across bots, not just against DrussGT.


4. What to try next (seeded; every later job adds its own)

  1. If a range-steering arm wins the panel: the win is a range effect, so sweep the band on the winning engine (e.g. TR_TFIL_RANGE_LO/HI on tfil_ring, TR_STRAFE_RANGE on strafe) with the same panel, 4–5 bands, and look for a plateau rather than a peak. A plateau is a result; a peak is a coin flip.
  2. If nothing beats tfil: stop tuning movement by feel. The next honest lever is enemy-model-driven placement (keep the bot where the enemy's expected hit probability is lowest given its gun model), which needs a per-opponent measurement, not a knob.
  3. Melee is a different game (j116's finding): if a melee campaign is opened, it needs its own panel and its own ledger section — do not reuse the 1v1 panel's verdicts.
  4. Close the loop with the enemy's own model: ab_mechanism.py-style instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by range band) is available and cheap — use it to explain a win, never to declare one.
  5. Then the gun (owner's next stage): the same harness, a gun panel, and the same paired-with-sign-test statistics. docs/surfer_wiring_ab.md and the j117 gauntlet are the gun-side priors to beat.

5. What would make us stop

  • Stop the movement stage when a batch produces an arm that is BETTER than the shipped default on the frozen panel by rule 2 and the effect survives the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of the panel). That arm becomes the new default candidate (shipping is a separate decision — this campaign never edits a shipped default).
  • Stop and move to the gun if two consecutive batches fail to produce an arm that beats the shipped tfil beyond the MDE: at that point the honest conclusion is "the shipped movement is the measured optimum of this design space", which is a successful campaign outcome, not a failure.
  • Stop a single batch early only for a contract violation (arena not free, liveness FAIL, non-zero exit rate) — never because the numbers look boring.

6. How to run a batch (exact commands)

# 1. wait for the arena (this job may not be the only one fighting)
tools/ab/tournament_run.sh \
    --arms    tools/ab/arms_movement_b1.txt \
    --panel   tools/ab/panel_movement.txt \
    --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
    --reference tfil \
    --outdir /tmp/ab/j118_b1

# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil

The runner writes <outdir>/session.json (commit sha, binary sha256, arms, panel) so any later job can re-analyze an old session offline, with no arena.