Files
SirRoboGarage/tools/ab
SirStone d2005abee9 j144 TFIL: arrival-based commitment + hysteresis + no mid-flight reversal
The owner's live-GUI report was correct on all four counts, and all four are
one bug: the commitment is cancelled by our own tile-boundary crossing
(96.1% of picks, 3793/3946, mean hold 5.06 ticks) while the bot is still
accelerating, and the picker is an unconstrained uniform draw over every
safe tile, so the new target can land in the mirror direction at |speed| < 4.

New knobs, all env-gated and default = today's behaviour (byte-for-byte
default parity guard re-run and green, 51 checks):
  TR_TFIL_COMMIT_ARRIVAL  hold the committed tile until we are ON it; the
                          tick knob becomes a MINIMUM dwell. 0 = shipped.
  TR_TFIL_COMMIT_MARGIN   leave only if the best alternative is at least
                          this much cooler on the same pathMaxHeat scale.
                          0 = shipped.
  TR_TFIL_NOREV_SPEED     while |speed| is below this, a mid-flight switch
                          may not take a tile >90 deg off the travel
                          direction. 0 = shipped. norevPool() never returns
                          an empty pool: with every candidate behind us it
                          takes the least-bad turn.

Offline gate (recorded DrussGT fixture, 20026 ticks): mean hold 4.1 -> 24.0
ticks, abandoned-before-arrival 92.8% -> 40.5%, committed tile actually
reached 3.3% -> 17.2%, opposite-direction slow mid-flight switches 394 -> 64
(-84%). 'TR_TFIL_TILE_REPLAN=off' alone - what cc11ede's arm B already tried -
only gets the hold to 13.6, which is why that A/B could not find this.

strafe is untouched: it imports only heatDecay/bulletMagScale/Pillar*, none
of which this touches. TR_MOVEMENT default stays strafe. Registered in
env_report.nim + knownEnvNames() + .env.example. Arms pre-registered in
docs/movement_campaign.md and tools/ab/arms_tfil_commit.txt.
2026-09-26 21:07:09 +02:00
..

tools/ab — reusable A/B harness

Two tools, built once and reused for every variant test. Adding an arm costs nothing: the frozen bot is built once per session and every arm reuses it.

1. Run a session

tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
  • builds ONE frozen ModularBot from current HEAD (git archive HEAD + nim c -d:release) and reuses that binary for every arm — a dirty tree cannot leak into the measurement;
  • runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one DrussGT botdir/data per run, one ModularBot botdir per run);
  • writes session.json (commit, binary sha256, arms, runs/rounds, timestamp) and <outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}.

Options: --arms FILE (required) --runs N (default 7) --outdir DIR (required) --conc K (default 7) --rounds R (default 7).

It kills its own children (own process group + outdir-tagged backstop) on EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.

Prerequisites (fails loudly if any is missing): /tmp/robocode/install/libs/robocode.jar, /tmp/drussgt/DrussGT.jar, the Tank Royale runner jar, the bot-API jar, nim, and the shim out/ classes. /tmp/tr_bots/DrussGT is recreated via make_botdir.sh if absent (the actual battles still use per-run copies).

2. Analyze a session

python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]

Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits taken/run, the per-run values, and for every pair of arms:

  • a two-sided permutation test on per-run damage and wins. Full enumeration when C(n, na) <= 20,000,000 (7v7 -> C(14,7)=3432, always exact); otherwise a Monte-Carlo permutation test with MC_DRAWS = 1,000,000 fixed draws and the fixed seed MC_SEED = 0x5eed5eed, reported with its Monte-Carlo standard error (p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))). Each row says which method produced its p-value;
  • a tie-corrected, continuity-corrected Mann-Whitney U cross-check;

plus the minimum detectable effect for the reference arm's n and observed per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line, a [bb] applied-shift check (needs TR_BITBRAIN_LOG=1; a zero-shift placebo emits no [bb] lines), and a round-win attribution cross-check.

Round wins come from the events sidecar (the bot that does not die wins) and are cross-checked against the runner's firstPlaces. The per-round lines in *.results.json are cumulative standings — not round winners.

Arm file

See arms.example.txt:

name | ENV_VAR=value ENV_VAR2=value2 | optional label

Known gotchas

  • Ports: the runner picks ephemeral ports itself; nothing to configure.
  • Races: never share a DrussGT botdir/data or a ModularBot stdout log across parallel runs — ab_run.sh already gives every run its own.
  • pkill -f run_bridge_battle matches the pkill command itself; use the [r]un_bridge_battle trick (as ab_run.sh does).
  • Liveness reads the bot's [env] boot report from <arm>/run<N>.bot.stdout.log; if an arm's variable is missing there it is a FAIL, not a measurement.