225 battles, one frozen binary, five env-only arms, the frozen panel, 0 invalid
runs. Paired per opponent vs the shipped tfil:
strafe_notilt wins/run +0.38 [CI +0.16,+0.60] 9/9 opponents p=0.0039
dmg/run -10.2 [CI -25.8,+5.5] p=0.61, MDE 20.4 (not detectable)
incoming hit rate 12.24% vs 18.17%, dmg taken 150 vs 200
strafe_325 wins/run +0.33 [CI +0.04,+0.63] 10/12 p=0.0386
ring dmg/run +31.2 [CI +11.5,+50.9] 13/15 p=0.0074, wins/run -0.04 (ns)
but hit rate 29.4% at 236 px: a damage/survival trade, not a win
ring_notemp indistinguishable from tfil on both primaries
Round wins in this harness are survival wins (in 216/219 attributable runs the
win count equals the rounds the opponent died in), and the winner takes ~1/3
fewer hits while fighting ~74 px farther out. The shipped tfil is last of five
on wins: the DrussGT-only picture did not generalize.
Also: tournament_analyze.py now prints BOTH readings of the pre-registered
'while the other does not go down' clause (strict: nothing is better;
substantive: the two strafe arms and ring are better on one metric each).
tools/ab — reusable A/B harness
Two tools, built once and reused for every variant test. Adding an arm costs nothing: the frozen bot is built once per session and every arm reuses it.
1. Run a session
tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
- builds ONE frozen ModularBot from current HEAD (
git archive HEAD+nim c -d:release) and reuses that binary for every arm — a dirty tree cannot leak into the measurement; - runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one DrussGT botdir/data per run, one ModularBot botdir per run);
- writes
session.json(commit, binary sha256, arms, runs/rounds, timestamp) and<outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}.
Options: --arms FILE (required) --runs N (default 7) --outdir DIR
(required) --conc K (default 7) --rounds R (default 7).
It kills its own children (own process group + outdir-tagged backstop) on EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.
Prerequisites (fails loudly if any is missing):
/tmp/robocode/install/libs/robocode.jar, /tmp/drussgt/DrussGT.jar,
the Tank Royale runner jar, the bot-API jar, nim, and the shim out/ classes.
/tmp/tr_bots/DrussGT is recreated via make_botdir.sh if absent (the actual
battles still use per-run copies).
2. Analyze a session
python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits taken/run, the per-run values, and for every pair of arms:
- a two-sided permutation test on per-run damage and wins. Full enumeration
when
C(n, na) <= 20,000,000(7v7 -> C(14,7)=3432, always exact); otherwise a Monte-Carlo permutation test withMC_DRAWS = 1,000,000fixed draws and the fixed seedMC_SEED = 0x5eed5eed, reported with its Monte-Carlo standard error (p = (cnt+1)/(B+1),se = sqrt(p(1-p)/(B+1))). Each row says which method produced its p-value; - a tie-corrected, continuity-corrected Mann-Whitney U cross-check;
plus the minimum detectable effect for the reference arm's n and observed
per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test
(labelled anti-conservative), a liveness OK/FAIL line, a [bb] applied-shift
check (needs TR_BITBRAIN_LOG=1; a zero-shift placebo emits no [bb] lines),
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and
are cross-checked against the runner's firstPlaces. The per-round lines in
*.results.json are cumulative standings — not round winners.
Arm file
See arms.example.txt:
name | ENV_VAR=value ENV_VAR2=value2 | optional label
Known gotchas
- Ports: the runner picks ephemeral ports itself; nothing to configure.
- Races: never share a DrussGT botdir/data or a ModularBot stdout log across
parallel runs —
ab_run.shalready gives every run its own. pkill -f run_bridge_battlematches the pkill command itself; use the[r]un_bridge_battletrick (asab_run.shdoes).- Liveness reads the bot's
[env]boot report from<arm>/run<N>.bot.stdout.log; if an arm's variable is missing there it is a FAIL, not a measurement.