Same frozen panel, same 3x3 design, new session on commit8efa627(no source file changed since1984a78, so the same code), 225 battles, 0 invalid runs. Arms: tfil, strafe_notilt, strafe_325 + the tilt re-armed at 600px and 250px. Paired vs tfil: strafe_325 +0.58 wins/run [CI +0.27,+0.89] 11/12 p=0.0063; strafe_notilt +0.47 [+0.22,+0.72] 10/11 p=0.0117; tilt_600 +0.40 [+0.04,+0.76] (sign test 8/11 p=0.23, sign-flip p=0.049); tilt_250 +0.38 [+0.07,+0.69] 10/12 p=0.039. Incoming hit rate -5.2..-6.9 pp with 0/15 opponents favouring tfil. The range TARGET is not the lever: re-arming the tilt moved the achieved distance from 459px (no steering) to 478px and 415px, and none of the three is separable on wins. This overturns Batch 1's reading that the tilt costs wins - the honest statement is that the tilt's win effect is below this design's resolution. tfil reproduced to within 1.4 pp (40.7% -> 39.3% of rounds), so the baseline itself is stable across sessions. Ledger: Batch 2 section, the verbatim analyzer report, a data-driven what-to-try-next, and the session log.
tools/ab — reusable A/B harness
Two tools, built once and reused for every variant test. Adding an arm costs nothing: the frozen bot is built once per session and every arm reuses it.
1. Run a session
tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
- builds ONE frozen ModularBot from current HEAD (
git archive HEAD+nim c -d:release) and reuses that binary for every arm — a dirty tree cannot leak into the measurement; - runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one DrussGT botdir/data per run, one ModularBot botdir per run);
- writes
session.json(commit, binary sha256, arms, runs/rounds, timestamp) and<outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}.
Options: --arms FILE (required) --runs N (default 7) --outdir DIR
(required) --conc K (default 7) --rounds R (default 7).
It kills its own children (own process group + outdir-tagged backstop) on EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.
Prerequisites (fails loudly if any is missing):
/tmp/robocode/install/libs/robocode.jar, /tmp/drussgt/DrussGT.jar,
the Tank Royale runner jar, the bot-API jar, nim, and the shim out/ classes.
/tmp/tr_bots/DrussGT is recreated via make_botdir.sh if absent (the actual
battles still use per-run copies).
2. Analyze a session
python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits taken/run, the per-run values, and for every pair of arms:
- a two-sided permutation test on per-run damage and wins. Full enumeration
when
C(n, na) <= 20,000,000(7v7 -> C(14,7)=3432, always exact); otherwise a Monte-Carlo permutation test withMC_DRAWS = 1,000,000fixed draws and the fixed seedMC_SEED = 0x5eed5eed, reported with its Monte-Carlo standard error (p = (cnt+1)/(B+1),se = sqrt(p(1-p)/(B+1))). Each row says which method produced its p-value; - a tie-corrected, continuity-corrected Mann-Whitney U cross-check;
plus the minimum detectable effect for the reference arm's n and observed
per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test
(labelled anti-conservative), a liveness OK/FAIL line, a [bb] applied-shift
check (needs TR_BITBRAIN_LOG=1; a zero-shift placebo emits no [bb] lines),
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and
are cross-checked against the runner's firstPlaces. The per-round lines in
*.results.json are cumulative standings — not round winners.
Arm file
See arms.example.txt:
name | ENV_VAR=value ENV_VAR2=value2 | optional label
Known gotchas
- Ports: the runner picks ephemeral ports itself; nothing to configure.
- Races: never share a DrussGT botdir/data or a ModularBot stdout log across
parallel runs —
ab_run.shalready gives every run its own. pkill -f run_bridge_battlematches the pkill command itself; use the[r]un_bridge_battletrick (asab_run.shdoes).- Liveness reads the bot's
[env]boot report from<arm>/run<N>.bot.stdout.log; if an arm's variable is missing there it is a FAIL, not a measurement.