tools/ab: reusable A/B runner + analyzer (frozen-HEAD build, exact permutation test, liveness check)

This commit is contained in:
2026-09-24 22:07:55 +02:00
parent e670788eee
commit 17c50159ba
4 changed files with 870 additions and 0 deletions
+65
View File
@@ -0,0 +1,65 @@
# tools/ab — reusable A/B harness
Two tools, built once and reused for every variant test. Adding an arm costs
nothing: the frozen bot is built once per session and every arm reuses it.
## 1. Run a session
```sh
tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
```
* builds ONE frozen ModularBot from **current HEAD** (`git archive HEAD` +
`nim c -d:release`) and reuses that binary for every arm — a dirty tree cannot
leak into the measurement;
* runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one
DrussGT botdir/data per run, one ModularBot botdir per run);
* writes `session.json` (commit, binary sha256, arms, runs/rounds, timestamp)
and `<outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}`.
Options: `--arms FILE` (required) `--runs N` (default 7) `--outdir DIR`
(required) `--conc K` (default 7) `--rounds R` (default 7).
It kills its own children (own process group + outdir-tagged backstop) on
EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.
Prerequisites (fails loudly if any is missing):
`/tmp/robocode/install/libs/robocode.jar`, `/tmp/drussgt/DrussGT.jar`,
the Tank Royale runner jar, the bot-API jar, `nim`, and the shim `out/` classes.
`/tmp/tr_bots/DrussGT` is recreated via `make_botdir.sh` if absent (the actual
battles still use per-run copies).
## 2. Analyze a session
```sh
python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
```
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits
taken/run, the **per-run** values, an exact two-sided permutation test on
per-run damage and wins vs the reference arm (default: first arm), a
round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line,
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and
are cross-checked against the runner's `firstPlaces`. The per-round lines in
`*.results.json` are **cumulative** standings — not round winners.
## Arm file
See `arms.example.txt`:
```
name | ENV_VAR=value ENV_VAR2=value2 | optional label
```
## Known gotchas
* Ports: the runner picks ephemeral ports itself; nothing to configure.
* Races: never share a DrussGT botdir/data or a ModularBot stdout log across
parallel runs — `ab_run.sh` already gives every run its own.
* `pkill -f run_bridge_battle` matches the pkill command itself; use the
`[r]un_bridge_battle` trick (as `ab_run.sh` does).
* Liveness reads the bot's `[env]` boot report from
`<arm>/run<N>.bot.stdout.log`; if an arm's variable is missing there it is a
FAIL, not a measurement.