Files
SirRoboGarage/tools/ab
SirStone 7d3645da5a movement Batch 2: the strafe win over shipped tfil replicates; the range target decides nothing
Same frozen panel, same 3x3 design, new session on commit 8efa627 (no source file
changed since 1984a78, so the same code), 225 battles, 0 invalid runs. Arms:
tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px.

Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063;
strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76]
(sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12
p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil.

The range TARGET is not the lever: re-arming the tilt moved the achieved
distance from 459px (no steering) to 478px and 415px, and none of the three is
separable on wins. This overturns Batch 1's reading that the tilt costs wins -
the honest statement is that the tilt's win effect is below this design's
resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the
baseline itself is stable across sessions.

Ledger: Batch 2 section, the verbatim analyzer report, a data-driven
what-to-try-next, and the session log.
2026-09-26 01:31:08 +02:00
..

tools/ab — reusable A/B harness

Two tools, built once and reused for every variant test. Adding an arm costs nothing: the frozen bot is built once per session and every arm reuses it.

1. Run a session

tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
  • builds ONE frozen ModularBot from current HEAD (git archive HEAD + nim c -d:release) and reuses that binary for every arm — a dirty tree cannot leak into the measurement;
  • runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one DrussGT botdir/data per run, one ModularBot botdir per run);
  • writes session.json (commit, binary sha256, arms, runs/rounds, timestamp) and <outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}.

Options: --arms FILE (required) --runs N (default 7) --outdir DIR (required) --conc K (default 7) --rounds R (default 7).

It kills its own children (own process group + outdir-tagged backstop) on EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.

Prerequisites (fails loudly if any is missing): /tmp/robocode/install/libs/robocode.jar, /tmp/drussgt/DrussGT.jar, the Tank Royale runner jar, the bot-API jar, nim, and the shim out/ classes. /tmp/tr_bots/DrussGT is recreated via make_botdir.sh if absent (the actual battles still use per-run copies).

2. Analyze a session

python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]

Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits taken/run, the per-run values, and for every pair of arms:

  • a two-sided permutation test on per-run damage and wins. Full enumeration when C(n, na) <= 20,000,000 (7v7 -> C(14,7)=3432, always exact); otherwise a Monte-Carlo permutation test with MC_DRAWS = 1,000,000 fixed draws and the fixed seed MC_SEED = 0x5eed5eed, reported with its Monte-Carlo standard error (p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))). Each row says which method produced its p-value;
  • a tie-corrected, continuity-corrected Mann-Whitney U cross-check;

plus the minimum detectable effect for the reference arm's n and observed per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line, a [bb] applied-shift check (needs TR_BITBRAIN_LOG=1; a zero-shift placebo emits no [bb] lines), and a round-win attribution cross-check.

Round wins come from the events sidecar (the bot that does not die wins) and are cross-checked against the runner's firstPlaces. The per-round lines in *.results.json are cumulative standings — not round winners.

Arm file

See arms.example.txt:

name | ENV_VAR=value ENV_VAR2=value2 | optional label

Known gotchas

  • Ports: the runner picks ephemeral ports itself; nothing to configure.
  • Races: never share a DrussGT botdir/data or a ModularBot stdout log across parallel runs — ab_run.sh already gives every run its own.
  • pkill -f run_bridge_battle matches the pkill command itself; use the [r]un_bridge_battle trick (as ab_run.sh does).
  • Liveness reads the bot's [env] boot report from <arm>/run<N>.bot.stdout.log; if an arm's variable is missing there it is a FAIL, not a measurement.