Files
SirRoboGarage/tools/ab
SirStone 48f38b80e7 TFIL heat-time + virtual pillar: live A/B (6 arms x 70 rounds) - neither change beats the pre-change mover
Runs the pre-registered A/B for the two movement changes in HEAD: the
time-indexed bullet heat (TR_TFIL_HEAT_TIME, fca8993) and the removal of the
invented virtual centre pillar (d0750ab). One frozen binary from HEAD vs real
DrussGT: 6 arms x 10 runs x 7 rounds = 60 battles, 420 rounds, 0 failed.

Judged on damage/run and ROUND WINS only (hit rate and hits-taken are context):
hit rate would have inverted the verdict again - tau3 has the best pooled hit
rate of all arms (11.56%) and the fewest round wins (20/70).

RESULT (vs the reconstructed pre-change mover "old"):
  heat-time HURTS. tau3/tau5/tau9 lose 1.3-1.7 wins/run (p=0.0010-0.0125) and
  deal 22-38 less damage/run (p=0.004-0.047); tau15 is a wash on wins (p=0.64)
  and 22 damage/run lower (p=0.046). Nothing improves either metric.
  pillar removal does nothing measurable. old vs pillaoff: +5.7 damage/run
  (p=0.71), +0.5 wins/run (35 vs 30, p=0.43), 30.8 MORE damage taken/run
  without the pillar (p=0.040). The mechanism check proves the knob works
  (centre-box occupancy 0.09% -> 2.37%, p<0.0001; range 469 -> 443 px,
  p=0.0002), so this is a real behaviour change that buys nothing. At n=10 the
  pillar contrast is inside the MDE (33 damage/run, 1.2 wins/run), so this is
  not a proven regression.

Flags that the shipped default (pillar removed) should be reverted to the
TR_TFIL_PILLAR_ON behaviour; heat-time stays off.

Adds tools/ab/arms_heat_pillar.txt and tools/ab/ab_mechanism.py (per-tick
mechanism check: central-box occupancy, range distribution, live enemy-bullet
proximity) plus the captured summary/report fixtures.
2026-09-25 22:23:11 +02:00
..

tools/ab — reusable A/B harness

Two tools, built once and reused for every variant test. Adding an arm costs nothing: the frozen bot is built once per session and every arm reuses it.

1. Run a session

tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
  • builds ONE frozen ModularBot from current HEAD (git archive HEAD + nim c -d:release) and reuses that binary for every arm — a dirty tree cannot leak into the measurement;
  • runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one DrussGT botdir/data per run, one ModularBot botdir per run);
  • writes session.json (commit, binary sha256, arms, runs/rounds, timestamp) and <outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}.

Options: --arms FILE (required) --runs N (default 7) --outdir DIR (required) --conc K (default 7) --rounds R (default 7).

It kills its own children (own process group + outdir-tagged backstop) on EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.

Prerequisites (fails loudly if any is missing): /tmp/robocode/install/libs/robocode.jar, /tmp/drussgt/DrussGT.jar, the Tank Royale runner jar, the bot-API jar, nim, and the shim out/ classes. /tmp/tr_bots/DrussGT is recreated via make_botdir.sh if absent (the actual battles still use per-run copies).

2. Analyze a session

python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]

Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits taken/run, the per-run values, and for every pair of arms:

  • a two-sided permutation test on per-run damage and wins. Full enumeration when C(n, na) <= 20,000,000 (7v7 -> C(14,7)=3432, always exact); otherwise a Monte-Carlo permutation test with MC_DRAWS = 1,000,000 fixed draws and the fixed seed MC_SEED = 0x5eed5eed, reported with its Monte-Carlo standard error (p = (cnt+1)/(B+1), se = sqrt(p(1-p)/(B+1))). Each row says which method produced its p-value;
  • a tie-corrected, continuity-corrected Mann-Whitney U cross-check;

plus the minimum detectable effect for the reference arm's n and observed per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test (labelled anti-conservative), a liveness OK/FAIL line, a [bb] applied-shift check (needs TR_BITBRAIN_LOG=1; a zero-shift placebo emits no [bb] lines), and a round-win attribution cross-check.

Round wins come from the events sidecar (the bot that does not die wins) and are cross-checked against the runner's firstPlaces. The per-round lines in *.results.json are cumulative standings — not round winners.

Arm file

See arms.example.txt:

name | ENV_VAR=value ENV_VAR2=value2 | optional label

Known gotchas

  • Ports: the runner picks ephemeral ports itself; nothing to configure.
  • Races: never share a DrussGT botdir/data or a ModularBot stdout log across parallel runs — ab_run.sh already gives every run its own.
  • pkill -f run_bridge_battle matches the pkill command itself; use the [r]un_bridge_battle trick (as ab_run.sh does).
  • Liveness reads the bot's [env] boot report from <arm>/run<N>.bot.stdout.log; if an arm's variable is missing there it is a FAIL, not a measurement.