MEASURED LIVE (common_libs/tests/measure_fire_ghost_lag.py, 4 sessions, 1777 matched ghost spawns, both movers): the server dispatches a turn's fire AFTER our go() for that same turn, so a turn-T shot's energy drop first reaches our scan at turn T+1 — and a bullet takes its FIRST step during the turn it is fired, so the true bullet is already one whole bullet step (11-20 px) downrange. Both movers place the ghost at the SCANNED enemy position (where the bullet was born), so the whole ghost trajectory is the true one shifted one turn later and the arrival deadline is a full tick late. MEASURED: detection lag +1 tick on 100% of 1777 matched spawns; ghost-vs- observer displacement 19.06 px mean / 22.00 p90 (tfil) and 16.08 / 21.81 (strafe); arrival-deadline error 0.99 / 0.77 ticks. NOT a rendering artefact: the draw/advance order is correct (advanceBullets -> detectFires -> build). THE FIX: TR_FIRE_LAG (int, default 0 = today byte-for-byte) in the shared fire_tracker, applied by both movers at spawn: x = origin + dir*speed*lag. The deadline needs no separate change — both movers derive it from the ghost's own position, so a correct position gives a correct deadline. WITH IT: displacement 19.06 -> 5.37 px mean (the residue is the enemy's own <=8 px scan staleness) and the deadline error 0.99 -> 0.06 ticks. Guards: test_tfil_commit_env 77 -> 87 checks (default golden parity, exact n-step back-date, deadline shortens by exactly lag, junk/negative degrade to 0, reaped exactly one tick earlier); test_env_report + test_env_dotenv green. TR_FIRE_LAG registered in env_report + knownEnvNames + .env.example + docs/env_reference.md. Live A/B pre-registered in docs/movement_campaign.md (Batch 8) with its MDE stated up front; arms tools/ab/arms_fire_lag.txt. TR_FIRE_DIAG gains a per-round ROUND line (the tick->getTurn anchor) and a per-spawn SPAWN line (the ghost's drawn position).
tools/ab — reusable A/B harness
Two tools, built once and reused for every variant test. Adding an arm costs nothing: the frozen bot is built once per session and every arm reuses it.
1. Run a session
tools/ab/ab_run.sh --arms tools/ab/arms.example.txt --runs 7 --outdir /tmp/ab/power --conc 7
- builds ONE frozen ModularBot from current HEAD (
git archive HEAD+nim c -d:release) and reuses that binary for every arm — a dirty tree cannot leak into the measurement; - runs arm × run battles vs real DrussGT in parallel (ephemeral ports, one DrussGT botdir/data per run, one ModularBot botdir per run);
- writes
session.json(commit, binary sha256, arms, runs/rounds, timestamp) and<outdir>/<arm>/run<N>.{jsonl,jsonl.rounds.json,jsonl.results.json,events.jsonl,battle.log,bot.stdout.log}.
Options: --arms FILE (required) --runs N (default 7) --outdir DIR
(required) --conc K (default 7) --rounds R (default 7).
It kills its own children (own process group + outdir-tagged backstop) on EXIT/INT/TERM, so a Ctrl-C does not leave orphan battles.
Prerequisites (fails loudly if any is missing):
/tmp/robocode/install/libs/robocode.jar, /tmp/drussgt/DrussGT.jar,
the Tank Royale runner jar, the bot-API jar, nim, and the shim out/ classes.
/tmp/tr_bots/DrussGT is recreated via make_botdir.sh if absent (the actual
battles still use per-run copies).
2. Analyze a session
python3 tools/ab/ab_analyze.py /tmp/ab/power [--reference control]
Prints per-arm damage/run, damage taken/run, round wins, shots/run, hits taken/run, the per-run values, and for every pair of arms:
- a two-sided permutation test on per-run damage and wins. Full enumeration
when
C(n, na) <= 20,000,000(7v7 -> C(14,7)=3432, always exact); otherwise a Monte-Carlo permutation test withMC_DRAWS = 1,000,000fixed draws and the fixed seedMC_SEED = 0x5eed5eed, reported with its Monte-Carlo standard error (p = (cnt+1)/(B+1),se = sqrt(p(1-p)/(B+1))). Each row says which method produced its p-value; - a tie-corrected, continuity-corrected Mann-Whitney U cross-check;
plus the minimum detectable effect for the reference arm's n and observed
per-run SD (alpha=0.05 two-sided, 80% power), a round-level Fisher test
(labelled anti-conservative), a liveness OK/FAIL line, a [bb] applied-shift
check (needs TR_BITBRAIN_LOG=1; a zero-shift placebo emits no [bb] lines),
and a round-win attribution cross-check.
Round wins come from the events sidecar (the bot that does not die wins) and
are cross-checked against the runner's firstPlaces. The per-round lines in
*.results.json are cumulative standings — not round winners.
Arm file
See arms.example.txt:
name | ENV_VAR=value ENV_VAR2=value2 | optional label
Known gotchas
- Ports: the runner picks ephemeral ports itself; nothing to configure.
- Races: never share a DrussGT botdir/data or a ModularBot stdout log across
parallel runs —
ab_run.shalready gives every run its own. pkill -f run_bridge_battlematches the pkill command itself; use the[r]un_bridge_battletrick (asab_run.shdoes).- Liveness reads the bot's
[env]boot report from<arm>/run<N>.bot.stdout.log; if an arm's variable is missing there it is a FAIL, not a measurement.