Files
SirRoboGarage/tools/ab/README.md
T

66 lines
2.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# tools/ab — reusable A/B harness
Two tools, built once and reused for every variant test. Adding an arm costs
nothing: the frozen bot is built once per session and every arm reuses it.
## 1. Run a session
```sh
tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
```
* builds ONE frozen ModularBot from **current HEAD** (`git archive HEAD` +
`nim c -d:release`) and reuses that binary for every arm — a dirty tree cannot
leak into the measurement;
* runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one
DrussGT botdir/data per run, one ModularBot botdir per run);
* writes `session.json` (commit, binary sha256, arms, runs/rounds, timestamp)
and `<outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}`.
Options: `--arms FILE` (required) `--runs N` (default 7) `--outdir DIR`
(required) `--conc K` (default 7) `--rounds R` (default 7).
It kills its own children (own process group + outdir-tagged backstop) on
EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.
Prerequisites (fails loudly if any is missing):
`/tmp/robocode/install/libs/robocode.jar`, `/tmp/drussgt/DrussGT.jar`,
the Tank Royale runner jar, the bot-API jar, `nim`, and the shim `out/` classes.
`/tmp/tr_bots/DrussGT` is recreated via `make_botdir.sh` if absent (the actual
battles still use per-run copies).
## 2. Analyze a session
```sh
python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
```
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits
taken/run, the **per-run** values, an exact two-sided permutation test on
per-run damage and wins vs the reference arm (default: first arm), a
round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line,
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and
are cross-checked against the runner's `firstPlaces`. The per-round lines in
`*.results.json` are **cumulative** standings — not round winners.
## Arm file
See `arms.example.txt`:
```
name | ENV_VAR=value ENV_VAR2=value2 | optional label
```
## Known gotchas
* Ports: the runner picks ephemeral ports itself; nothing to configure.
* Races: never share a DrussGT botdir/data or a ModularBot stdout log across
parallel runs — `ab_run.sh` already gives every run its own.
* `pkill -f run_bridge_battle` matches the pkill command itself; use the
`[r]un_bridge_battle` trick (as `ab_run.sh` does).
* Liveness reads the bot's `[env]` boot report from
`<arm>/run<N>.bot.stdout.log`; if an arm's variable is missing there it is a
FAIL, not a measurement.