Files
SirRoboGarage/tools/ab/README.md
T
SirStone d93ce444c0 BitBrain verdict: clean negative at 30 runs/arm; analyzer gets MC + Mann-Whitney + MDE
- docs/bitbrain_gun_verdict.md: control vs bb_decay (decay SBC memory) vs a
  provably-zero placebo, 30 runs/arm vs real DrussGT. Nothing separates
  (bb_decay +3.3 dmg/run, p=0.71; round wins 97/210 vs 97/210, p=1.00); the
  7-run shape does not replicate. TR_BITBRAIN_RANGE=0 is clamped to 1.0 deg
  (bitbrain_gun.nim:207) so it is NOT a zero-shift placebo; TR_BITBRAIN_MIN_OBS
  unreachable is used instead.
- tools/ab/ab_analyze.py: keep exact enumeration for C(n,na)<=20e6 (7v7), add
  a seeded Monte-Carlo permutation test (1e6 draws, 0x5eed5eed) with its
  standard error, a tie-corrected Mann-Whitney U cross-check, a minimum
  detectable effect line, all-pairs comparisons, and a [bb] shift check.
- tools/ab/README.md: document the new analyzer output.
2026-09-24 23:02:59 +02:00

77 lines
3.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# tools/ab — reusable A/B harness
Two tools, built once and reused for every variant test. Adding an arm costs
nothing: the frozen bot is built once per session and every arm reuses it.
## 1. Run a session
```sh
tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
```
* builds ONE frozen ModularBot from **current HEAD** (`git archive HEAD` +
`nim c -d:release`) and reuses that binary for every arm — a dirty tree cannot
leak into the measurement;
* runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one
DrussGT botdir/data per run, one ModularBot botdir per run);
* writes `session.json` (commit, binary sha256, arms, runs/rounds, timestamp)
and `<outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}`.
Options: `--arms FILE` (required) `--runs N` (default 7) `--outdir DIR`
(required) `--conc K` (default 7) `--rounds R` (default 7).
It kills its own children (own process group + outdir-tagged backstop) on
EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.
Prerequisites (fails loudly if any is missing):
`/tmp/robocode/install/libs/robocode.jar`, `/tmp/drussgt/DrussGT.jar`,
the Tank Royale runner jar, the bot-API jar, `nim`, and the shim `out/` classes.
`/tmp/tr_bots/DrussGT` is recreated via `make_botdir.sh` if absent (the actual
battles still use per-run copies).
## 2. Analyze a session
```sh
python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
```
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits
taken/run, the **per-run** values, and for **every pair of arms**:
* a two-sided **permutation test** on per-run damage and wins. Full enumeration
when `C(n, na) <= 20,000,000` (7v7 -> C(14,7)=3432, always exact); otherwise a
**Monte-Carlo** permutation test with `MC_DRAWS = 1,000,000` fixed draws and the
fixed seed `MC_SEED = 0x5eed5eed`, reported with its Monte-Carlo standard error
(`p = (cnt+1)/(B+1)`, `se = sqrt(p(1-p)/(B+1))`). Each row says which method
produced its p-value;
* a tie-corrected, continuity-corrected **Mann-Whitney U** cross-check;
plus the **minimum detectable effect** for the reference arm's n and observed
per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test
(labelled anti-conservative), a liveness OK/FAIL line, a `[bb]` applied-shift
check (needs `TR_BITBRAIN_LOG=1`; a zero-shift placebo emits no `[bb]` lines),
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and
are cross-checked against the runner's `firstPlaces`. The per-round lines in
`*.results.json` are **cumulative** standings — not round winners.
## Arm file
See `arms.example.txt`:
```
name | ENV_VAR=value ENV_VAR2=value2 | optional label
```
## Known gotchas
* Ports: the runner picks ephemeral ports itself; nothing to configure.
* Races: never share a DrussGT botdir/data or a ModularBot stdout log across
parallel runs — `ab_run.sh` already gives every run its own.
* `pkill -f run_bridge_battle` matches the pkill command itself; use the
`[r]un_bridge_battle` trick (as `ab_run.sh` does).
* Liveness reads the bot's `[env]` boot report from
`<arm>/run<N>.bot.stdout.log`; if an arm's variable is missing there it is a
FAIL, not a measurement.