melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75)

This commit is contained in:
2026-09-26 00:43:04 +02:00
parent da4a971ca9
commit 1984a780f4
6 changed files with 1342 additions and 3 deletions
+3 -3
View File
@@ -186,9 +186,9 @@ Target switching is the whole mechanism, so it had to be proven, not assumed.
**66–69 target changes/run** (about 10 per round). Example transitions from one
run: `#1 → #4 → #2 → #4 → #1 → #2 → …`.
* **BitBrain really reset on every switch.** `[bb-reset] reason=target_change`
fires once per target change: **66.75** (perRound), **68.19** (retained),
**66.75** (decay) per run — equal, within rounding, to the target-change count.
Example `[bb]` line from a live run:
fires once per target change: **68.75** (perRound), **68.19** (retained),
**66.75** (decay) per run — equal, within rounding, to each arm's target-change
count. Example `[bb]` line from a live run:
`[bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound`.
The `pattern` arm logs **zero** `[bb]`/`[bb-reset]` lines (BitBrain never
spawned), confirming the arms are cleanly separated.
+179
View File
@@ -0,0 +1,179 @@
# Movement campaign — ledger
**Goal (owner's mandate, 2026-09-26 overnight):** find the *best 1v1 movement*
by measurement, then do the same for the gun. This file is the campaign's
single source of truth: every later job **appends** a `## Batch N` section and
never edits an earlier one (a wrong earlier number gets a correction line, not
a rewrite).
**Owner's words:** *"I want you to do all tests and checks with the goal to have
the best 1vs1 movement. You have all night, you can change every parameter.
Continue until you found an amazing movement. When found do the same over for a
gun."*
---
## 0. The one caveat this campaign exists to close
Everything measured about movement before this campaign is **DrussGT-only**:
`docs/surfer_wiring_ab.md`, the j107 range drift, the j113 BitBrain movement
notes. The standing lesson of the night is that a one-opponent result is not a
result:
> an arm can take fewer hits **and** win fewer rounds (j107 / `strafe`): the
> verdict lives in **damage/run + ROUND WINS**, and hit rate is only ever an
> explanation.
So from here on **the unit of evidence is the number of opponents**, not the
number of runs: the same arm must win on *many* opponents before it is called
better.
---
## 1. Protocol (how every batch must be run)
| Element | Rule |
|---|---|
| Subject | ONE frozen binary, built from `git archive HEAD` (`tools/ab/tournament_run.sh` does this; the commit sha and binary sha256 are recorded in `session.json`) |
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file, not a shell history |
| Panel | the **frozen** panel `tools/ab/panel_movement.txt`. Adding/removing an opponent starts a **new batch number** |
| Pairing | per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate |
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
| Serialization | **one battle fleet at a time.** `tournament_run.sh --wait-arena N` refuses/stalls while another job's `run_bridge_battle`/`TrBattleCapture`/`ModularBot_bin` is alive (bracketed pgrep; never a broad `pkill`) |
| Liveness | every declared env token must appear verbatim in OUR bot's own `[env]` boot report, else the run is excluded and named in the report; an undeclared `TR_MOVEMENT` in the process env is a fatal FAIL for the reference arm |
| Never shipped | this is a measurement + design campaign: `git status` clean, defaults untouched, `.gitignore` untouched |
### Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit `__PRECOMMIT__`)
1. **Primary metrics:** damage/run and ROUND WINS. Secondary/explanation only:
damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean
distance.
2. **BETTER than the reference** iff one primary metric is up with a
cross-opponent **sign test p < 0.05** while the other does **not** go down;
or the mirror image for **WORSE**. Anything else is **NOT
DISTINGUISHABLE** (which is a real answer, not a failure).
3. **A verdict must survive the between-opponent spread**: the pooled mean delta
is reported with the SD across opponents, its SE, a 95% CI, and the MDE
(α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as
*not detectable*, never as *absent* and never as a win.
4. **Somewhere to stop:** if no arm beats the shipped `tfil` by rule 2 in
Batch 1 **and** no arm shows a ≥ +MDE damage gain with p<0.10, the movement
stage's first phase is closed with *"the shipped `tfil` is the best movement
we have measured"* — that is a **successful** outcome, and the campaign moves
to the gun axis rather than inventing more movement arms. See
*What would make us stop* at the end.
5. **No promotion off a single metric, a single opponent, or a single run.**
A change that wins damage by losing wins (or vice-versa) is not a win.
6. Every batch is shot with a **pre-registered prediction** stated in its
section *before* the battles finish; a prediction that turns out wrong is
recorded as wrong.
---
## 2. Stage 0 — what we already know (given, not re-derived)
Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary
(`docs/surfer_wiring_ab.md`, commit `0f5cfe3`):
| arm | dmg/run | dmg taken | round wins | incoming hit rate |
|---|---:|---:|---:|---:|
| `tfil` (SHIPPED) | **293** | 224 | **45/105** | 10.40% |
| `strafe` (range 325) | 250 | **198** | 37/105 | **9.40%** |
| `surf` | 255 | 259 | 37/105 | 13.51% |
Read: the shipped `tfil` deals the most damage and wins the most rounds while
being hit the *most*; `strafe` dodges best and wins least. Plus j107: drifting
25–30 px closer made damage **and** wins worse, so the lever is not simply "get
closer". **Hypothesis entering the campaign: the 325 px range preference of
`strafe` costs wins** (INFERRED from DrussGT-only data — this is exactly what
Batch 1 tests across a panel).
---
## 3. Batch 1 — isolating the range / aggression axis
**Design.** One frozen binary, five env-only arms, one frozen panel
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 2
wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas),
3 runs × 3 rounds per (opponent, arm). Arm file:
`tools/ab/arms_movement_b1.txt`.
| # | arm | env | what it isolates |
|---|---|---|---|
| 1 | `tfil` | *(none — shipped defaults)* | the arm to beat |
| 2 | `strafe_notilt` | `TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999` | the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with **no range steering at all** |
| 3 | `strafe_325` | `TR_MOVEMENT=strafe` | the current strafe default (range 325, tol 25, tilt 15/0.10) |
| 4 | `ring` | `TR_MOVEMENT=tfil_ring` | TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) **with** the range-weighted tile draw (band 100–200) |
| 5 | `ring_notemp` | `TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0` | the control for #4: same retuned heat field, range weighting switched OFF (`rand(candidates.high)` path) |
`ring` − `ring_notemp` is therefore the range-weighting lever **alone**, on a
heat field that is already retuned. The originally-suggested 5th arm ("`tfil`
with less saturated heat") is **not buildable in this campaign**: in
`common_libs/movements/the_floor_is_lava.nim` `CorridorHeat`/`WallHotness` are
Nim `const`s (env_report only *reports* them); only the `tfil_ring` copy reads
them from the env. #5 is the honest substitute.
**Pre-registered prediction (written before the battles finished):** `tfil`
still wins the panel on damage and round wins; `strafe_notilt` will beat
`strafe_325` on round wins (the range tilt is a net cost), and the ring arms will
land between them. If instead the range-steering arms beat `tfil` on wins, the
"range preference costs wins" hypothesis is confirmed across bots, not just
against DrussGT.
<!-- BATCH1-RESULTS -->
---
## 4. What to try next (seeded; every later job adds its own)
1. **If a range-steering arm wins the panel:** the win is a *range* effect, so
sweep the *band* on the winning engine (e.g. `TR_TFIL_RANGE_LO/HI` on
`tfil_ring`, `TR_STRAFE_RANGE` on `strafe`) with the same panel, 4–5 bands,
and look for a plateau rather than a peak. A plateau is a result; a peak is a
coin flip.
2. **If nothing beats `tfil`:** stop tuning movement by feel. The next honest
lever is *enemy-model-driven* placement (keep the bot where the enemy's
expected hit probability is lowest *given its gun model*), which needs a
per-opponent measurement, not a knob.
3. **Melee is a different game** (j116's finding): if a melee campaign is
opened, it needs its own panel and its own ledger section — do not reuse the
1v1 panel's verdicts.
4. **Close the loop with the enemy's own model:** `ab_mechanism.py`-style
instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by
range band) is available and cheap — use it to *explain* a win, never to
declare one.
5. **Then the gun** (owner's next stage): the same harness, a gun panel, and the
same paired-with-sign-test statistics. `docs/surfer_wiring_ab.md` and the j117
gauntlet are the gun-side priors to beat.
## 5. What would make us stop
* **Stop the movement stage** when a batch produces an arm that is BETTER than
the shipped default on the frozen panel by rule 2 **and** the effect survives
the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of
the panel). That arm becomes the new default *candidate* (shipping is a
separate decision — this campaign never edits a shipped default).
* **Stop and move to the gun** if two consecutive batches fail to produce an arm
that beats the shipped `tfil` beyond the MDE: at that point the honest
conclusion is *"the shipped movement is the measured optimum of this design
space"*, which is a successful campaign outcome, not a failure.
* **Stop a single batch early** only for a contract violation (arena not free,
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
## 6. How to run a batch (exact commands)
```sh
# 1. wait for the arena (this job may not be the only one fighting)
tools/ab/tournament_run.sh \
--arms tools/ab/arms_movement_b1.txt \
--panel tools/ab/panel_movement.txt \
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
--reference tfil \
--outdir /tmp/ab/j118_b1
# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil
```
The runner writes `<outdir>/session.json` (commit sha, binary sha256, arms,
panel) so any later job can re-analyze an old session offline, with no arena.
+44
View File
@@ -0,0 +1,44 @@
# ─────────────────────────────────────────────────────────────────────────────
# arms_movement_b1.txt — BATCH 1 of the movement campaign: isolate the
# RANGE / AGGRESSION axis. Five arms, ONE frozen binary, env vars only.
#
# Format: name | ENV=value ENV=value | label
#
# Prior data this batch is built on (all DrussGT-only, all MEASURED):
# * shipped `tfil` 293 dmg/run, 45/105 round wins, 10.40% incoming hit rate
# * `strafe` (range 325) 250 dmg/run, 37/105 wins, 9.40% incoming -> best
# dodger, fewest wins: the 325px range preference may COST wins
# * j107: drifting 25-30 px closer made damage AND wins worse, so "get
# closer" is not the lever by itself
# This batch closes the caveat that all of that is DrussGT-only and decomposes
# the range-steering lever from the engine.
# ─────────────────────────────────────────────────────────────────────────────
# 1. the arm to beat: shipped default, no env at all
tfil | | shipped baseline (movement engine tfil, every knob at its default)
# 2. pure perpendicular strafe with the range steering DISABLED: tilt is zero
# whenever |dist - StrafeRange| <= TOL, so a huge TOL makes `tilt` provably
# 0 every tick (and the approach-direction branch that depends on tiltMag is
# skipped). This isolates the COST of the 325px preference: same engine, same
# wall logic, no range steering at all.
strafe_notilt | TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999 | strafe, range steering OFF (tilt always 0)
# 3. the current strafe default: range 325, tol 25, tilt max 15, gain 0.10
strafe_325 | TR_MOVEMENT=strafe | strafe default (range 325, tol 25, tilt 15/0.10)
# 4. the owner's own retuned mover: tfil semantics + retuned heat field
# (CorridorHeat 10 = the safety threshold, WallHotness 15, WallRadiance 5,
# BulletCore/Aura 20/10, commit 5 ticks) + range-weighted tile draw
# (band 100-200, temp 0.4, K 60). Range steering ON, different heat field.
ring | TR_MOVEMENT=tfil_ring | tfil_ring (retuned heat field + range weighting 100-200)
# 5. the control for arm 4: TR_TFIL_RANGE_TEMP=0 makes the tile draw UNIFORM
# again (the documented OFF path), so arm 5 = the retuned heat field with NO
# range steering. (4) - (5) is therefore the range-weighting lever alone, on
# a heat field that is already retuned. NOTE: there is no env knob for the
# heat constants of the SHIPPED tfil mover (CorridorHeat/WallHotness are
# Nim `const` there, only REPORTED by env_report.nim), so the suggested
# "tfil with less saturated heat" arm cannot be built without a rebuild,
# which this campaign forbids. This is the honest substitute.
ring_notemp | TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0 | tfil_ring, range weighting OFF (temp 0)
+71
View File
@@ -0,0 +1,71 @@
# ─────────────────────────────────────────────────────────────────────────────
# panel_movement.txt — THE FROZEN MOVEMENT PANEL (Batch 1, 2026-09-26)
#
# DO NOT ADD, REMOVE OR REORDER AN OPPONENT without starting a new batch number
# in docs/movement_campaign.md. Every later job in the movement campaign must
# measure against THIS list: the unit of evidence is the number of opponents, so
# changing the panel changes what "better movement" means.
#
# Format (one opponent per line):
# <bot dir or name> | <style group> | <why it is in the panel>
# A bare `Name` is expanded to /tmp/tr_bots/Name. Blank lines and #-comments are
# ignored.
#
# `style` is INFERRED (from the bot's name / docs in
# tools/robocode_shim/robots.json, not from decompilation — robots.json says so
# explicitly). The style group is used ONLY to explain a result, never to decide
# one. The bot dirs come from tools/robocode_shim/robots.json (28 battle-validated
# legacy champions, see LEGACY_BOTS.md) except SpinBot, which is the tank-royale
# sample bot shipped with the project.
#
# Panel design (15 opponents, one axis each):
# * numbers matter, style does not: 5-6 opponents per family so a single
# weird bot can never carry an arm's mean;
# * the anchors (DrussGT/Diamond/Dookious) are the only opponents with older
# movement data, so they keep the new numbers comparable to the old ones;
# * every opponent must be able to DAMAGE us (WORKS in robots.json, not
# WORKS_WEAK): an opponent that lands 0 hits cannot reveal a movement
# regression, and an opponent that cannot dodge cannot reveal a range win.
#
# dodger — wave-surfing / adaptive evasion: punish our bullets AND our
# movement at the same time (the hard axis)
# pattern — pattern-matching guns: punish PREDICTABLE movement, which is
# exactly what a pure perpendicular strafe is
# wallfollower — wall-following / wall-avoiding blockers: punish movers that
# do not manage walls and let a corner-camper's corner matter
# cornercamper — corner camper / low bot
# spinner — periodic circle mover with a head-on gun: the classic
# "did you do anything obviously stupid" control
# rammer — closes to point blank: punishes a mover whose range
# preference cannot be enforced (and rewards damage output)
# brawler — micro/close-range brawler: same axis as the rammer, different
# gun quality
# aggressive — strong aggressive 1v1 megas: the best all-round stress test
# ─────────────────────────────────────────────────────────────────────────────
# ── dodger (5) ───────────────────────────────────────────────────────────────
/tmp/tr_bots/DrussGT | dodger | THE anchor. Our only opponent with prior data (docs/surfer_wiring_ab.md, j107, j113) — keeps new results comparable to the old ones.
/tmp/tr_bots/Diamond | dodger | Voidious top-tier: pattern-matching gun + wave-surfing movement.
/tmp/tr_bots/Dookious | dodger | Voidious: wave-surfing movement + pattern/GF gun.
/tmp/tr_bots/GresSuffurd | dodger | Surf gun + surf movement; the most "modern" of the surfers here.
/tmp/tr_bots/CassiusClay | dodger | Pattern-matching gun + surf movement (different gun family from Diamond).
# ── pattern (3) ──────────────────────────────────────────────────────────────
/tmp/tr_bots/RetroGirl | pattern | Perceptual pattern matcher — the strongest pattern gun in the panel (26 hits in 2 rounds in the validation).
/tmp/tr_bots/TripHammer | pattern | Pattern matcher (19 hits in 2 rounds — strong).
/tmp/tr_bots/Coriantumr | pattern | Mini pattern matcher: small/simple bots stress the same axis at low complexity.
# ── wall follower / corner camper (2) ────────────────────────────────────────
/tmp/tr_bots/WallAvoider | wallfollower | Micro wall-avoider, classic blocking-Robot style: the "wall management" control.
/tmp/tr_bots/HawkOnFire | cornercamper | Corner camper / low bot: punishes movers that let the enemy own a corner.
# ── spinner (1, NOT from robots.json — see the note) ─────────────────────────
/home/davide/Projects/tank-royale/sample-bots/java/build/archive/SpinBot | spinner | The tank-royale sample SpinBot: the only true periodic circle-mover available (there is NO spinner among the 28 legacy champions). It is the same bot every legacy validation used, so it is a known quantity; it is the control for "obviously stupid" movement, not a hard test.
# ── rammer / brawler (2) ─────────────────────────────────────────────────────
/tmp/tr_bots/DiamondStealer | rammer | Rammer / close-range brawler: the only pure rammer validated (12 hits at point blank in the validation).
/tmp/tr_bots/BlitzBat | brawler | Micro brawler: close-range aggression with a weaker gun.
# ── aggressive mega (2) ──────────────────────────────────────────────────────
/tmp/tr_bots/YersiniaPestis | aggressive | Strong aggressive 1v1 mega (never defeated in the 2-round validation) — the generalist stress test.
/tmp/tr_bots/Ascendant | aggressive | Aggressive 1v1 mega on a team-derived pattern base — the second generalist.
+639
View File
@@ -0,0 +1,639 @@
#!/usr/bin/env python3
"""tournament_analyze.py — paired multi-opponent ranking of movement arms.
python3 tools/ab/tournament_analyze.py <session_dir> [--reference ARM]
Reads a session produced by `tournament_run.sh` (layout:
<outdir>/<opponent>/<arm>/run<N>.{battle.log,events.jsonl,jsonl,bot.stdout.log})
and answers the only question the movement campaign asks:
DOES ARM A MOVE BETTER THAN THE REFERENCE ARM, ACROSS OPPONENTS?
Method, in one paragraph. Every arm fights the SAME panel with the SAME frozen
binary; for each opponent the arm's metric is averaged over its runs and
subtracted from the reference arm's average for that same opponent. Those
per-opponent deltas are the unit of evidence: the MEAN of the deltas is the
effect, the SPREAD of the deltas across opponents is the honest error bar (one
weird opponent cannot carry it), and a SIGN TEST over the deltas says how many
opponents the arm actually wins. A pooled number over all runs is also printed,
but it is reported as the descriptive dashboard, never as the verdict.
CLI judgment (pre-registered in docs/movement_campaign.md): the primary metrics
are damage/run and ROUND WINS. An arm is BETTER than the reference only if one
of the two improves with the sign test at p<0.05 while the other does not
degrade; hit rate and distance are explanation, never the verdict. MDE
(alpha=0.05 two-sided, 80% power) is printed for every test so a null can be
told apart from an under-powered null.
Standard library only. Deterministic: the exact sign-flip test is enumerated
when n_opponents <= 20, otherwise sampled with a fixed seed.
"""
import json
import math
import os
import random
import re
import sys
BOT_NAME = "ModularBot"
# The metric keys, in report order.
METRICS = ["damage", "damage_taken", "wins", "hit_rate", "dist"]
# z_{0.975} + z_{0.80}: the constant in MDE = C * sd * sqrt(2/n) is for a
# two-SAMPLE design; for the paired per-opponent design used here the analogous
# constant with n = number of opponents is C * sd(deltas) / sqrt(n).
MDE_C = 1.959963984540054 + 0.8416212335729143
EXACT_SIGNCAP = 20 # 2^20 = 1M sign vectors is still instant
MC_DRAWS = 200_000
MC_SEED = 0x5EED5EED
T975 = {1: 12.706, 2: 4.303, 3: 3.182, 4: 2.776, 5: 2.571, 6: 2.447,
7: 2.365, 8: 2.306, 9: 2.262, 10: 2.228, 11: 2.201, 12: 2.179,
13: 2.160, 14: 2.145, 15: 2.131, 16: 2.120, 17: 2.110, 18: 2.101,
19: 2.093, 20: 2.086, 21: 2.080, 22: 2.074, 23: 2.069, 24: 2.064,
25: 2.060, 26: 2.056, 27: 2.052, 28: 2.048, 29: 2.045, 30: 2.042}
# ── small statistics helpers ─────────────────────────────────────────────────
def mean(xs):
return sum(xs) / len(xs) if xs else float("nan")
def sd(xs):
"""Sample standard deviation (n-1). 0.0 for n<2."""
n = len(xs)
if n < 2:
return 0.0
m = mean(xs)
return math.sqrt(sum((x - m) ** 2 for x in xs) / (n - 1))
def median(xs):
s = sorted(xs)
n = len(s)
if n == 0:
return float("nan")
return s[n // 2] if n % 2 else 0.5 * (s[n // 2 - 1] + s[n // 2])
def binom_two_sided(k, n):
"""Exact two-sided sign-test p-value (p=0.5), ties already removed."""
if n == 0:
return 1.0
def c(nn, kk):
return math.comb(nn, kk)
tail = sum(c(n, i) for i in range(0, min(k, n - k) + 1)) / 2 ** n
return min(1.0, 2.0 * tail)
def signflip_p(deltas):
"""Two-sided sign-flip permutation test on the MEAN of the deltas.
Exact (all 2^n sign vectors) for n <= EXACT_SIGNCAP; otherwise a
deterministic Monte-Carlo draw. Returns (p, method_string)."""
n = len(deltas)
if n == 0:
return 1.0, "n/a"
obs = abs(mean(deltas))
if obs == 0.0:
return 1.0, "exact (degenerate)"
tol = 1e-12
if n <= EXACT_SIGNCAP:
total = 1 << n
hits = 0
for mask in range(total):
s = 0.0
for i, d in enumerate(deltas):
s += -d if (mask >> i) & 1 else d
if abs(s / n) >= obs - tol:
hits += 1
return hits / total, f"exact 2^{n}"
rng = random.Random(MC_SEED)
hits = 0
for _ in range(MC_DRAWS):
s = 0.0
for d in deltas:
s += -d if rng.getrandbits(1) else d
if abs(s / n) >= obs - tol:
hits += 1
return hits / MC_DRAWS, f"MC {MC_DRAWS}"
def wilcoxon_p(deltas):
"""Two-sided Wilcoxon signed-rank, normal approximation with tie
correction. Returns (p, W). p=1.0 when there is nothing to test."""
nz = [d for d in deltas if d != 0.0]
n = len(nz)
if n < 3:
return 1.0, float("nan")
order = sorted(range(n), key=lambda i: abs(nz[i]))
ranks = [0.0] * n
i = 0
while i < n:
j = i
while j + 1 < n and abs(nz[order[j + 1]]) == abs(nz[order[i]]):
j += 1
avg = (i + j) / 2.0 + 1.0
for k in range(i, j + 1):
ranks[order[k]] = avg
i = j + 1
w_plus = sum(ranks[i] for i in range(n) if nz[i] > 0)
mu = n * (n + 1) / 4.0
# tie correction for sigma
from collections import Counter
cnt = Counter(abs(d) for d in nz)
tie = sum(c ** 3 - c for c in cnt.values())
sigma2 = n * (n + 1) * (2 * n + 1) / 24.0 - tie / 48.0
if sigma2 <= 0:
return 1.0, w_plus
z = (w_plus - mu - 0.5 * (1 if w_plus > mu else -1)) / math.sqrt(sigma2)
p = 2.0 * 0.5 * math.erfc(abs(z) / math.sqrt(2.0))
return min(1.0, p), w_plus
def t_crit(df):
return T975.get(df, 1.96)
def mde(deltas):
"""Minimum detectable effect for the paired per-opponent design at
alpha=0.05 (two-sided), power 80%, from the observed spread of the deltas."""
n = len(deltas)
if n < 2:
return float("nan")
return MDE_C * sd(deltas) / math.sqrt(n)
# ── parsing ──────────────────────────────────────────────────────────────────
COUNTERS_RE = re.compile(
r"subject event counts: scans=(\d+) bulletsFired=(\d+) bulletHits=(\d+)"
r" bulletMisses=(\d+) bulletHitBullets=(\d+) hitsTaken=(\d+)")
FIRSTPLACES_RE = re.compile(
r"^\s*#\d+\s+(\S+)\s+totalScore=(-?\d+)\s+firstPlaces=(\d+)\s+survival=(\d+)",
re.MULTILINE)
DIST_RE = re.compile(r"DISTANCE: mean=([\d.]+)")
ROWS_RE = re.compile(r"rows=(\d+)")
def read_text(path):
try:
with open(path, "r", errors="replace") as fh:
return fh.read()
except OSError:
return ""
def parse_events(path):
out = []
try:
with open(path, "r", errors="replace") as fh:
for line in fh:
line = line.strip()
if not line:
continue
try:
out.append(json.loads(line))
except json.JSONDecodeError:
continue # a partial line from a killed battle
except OSError:
out = []
return out
def parse_counters(text):
m = COUNTERS_RE.search(text)
if not m:
return None
return {"scans": int(m.group(1)), "fired": int(m.group(2)),
"hits": int(m.group(3)), "misses": int(m.group(4)),
"hit_bullets": int(m.group(5)), "hits_taken": int(m.group(6))}
def parse_first_places(text):
for m in FIRSTPLACES_RE.finditer(text):
if m.group(1) == BOT_NAME:
return int(m.group(3))
return None
def parse_rounds(path):
try:
with open(path) as fh:
return len(json.load(fh).get("rounds", []))
except (OSError, json.JSONDecodeError):
return None
def attribute_subject(evs, counters):
"""Which event `owner` id is our bot? Matched on fire/hit/hits-taken counts,
never on fired power (the power policy is continuous).
Returns (subject_id, other_id, how). `subject_id` is None only when nothing
could be identified; `other_id` is None when the opponent never fired a
single bullet (then the incoming hit rate is simply undefined, and the run
is still a valid damage measurement)."""
fires, hits, victim_hits = {}, {}, {}
owner_ids = set()
for ev in evs:
t = ev.get("type")
o = ev.get("owner")
if o is not None:
owner_ids.add(o)
if t == "fire":
fires[o] = fires.get(o, 0) + 1
elif t == "hit":
hits[o] = hits.get(o, 0) + 1
v = ev.get("victim")
if v is not None:
victim_hits[v] = victim_hits.get(v, 0) + 1
if not fires:
return None, None, "no fire events"
def other_of(subj):
others = [o for o in owner_ids if o != subj]
return others[0] if len(others) == 1 else None
if counters:
strict = [o for o in fires
if fires[o] == counters["fired"]
and hits.get(o, 0) == counters["hits"]
and victim_hits.get(o, 0) == counters["hits_taken"]]
if len(strict) == 1:
return strict[0], other_of(strict[0]), "exact"
cand = [o for o in fires if counters and fires[o] == counters["fired"]]
if len(cand) != 1:
cand = [o for o in fires
if counters and victim_hits.get(o, 0) == counters["hits_taken"]]
if len(cand) != 1:
cand = list(fires)
if len(cand) == 1:
return cand[0], other_of(cand[0]), "fires-only"
return None, None, "ambiguous"
def liveness(arm_env_tokens, stdout_path):
"""Liveness: every declared env token must appear verbatim in OUR bot's own
boot environment report, so an arm whose setting never reached the process
is a loud FAIL instead of a plausible-looking number. A TR_MOVEMENT the arm
did NOT declare is fatal too (the baseline would be contaminated)."""
text = read_text(stdout_path)
if not text:
return False, "no boot env report", False
missing = [t for t in arm_env_tokens if t not in text]
declared_keys = {t.split("=", 1)[0] for t in arm_env_tokens}
leaked_movement = ("TR_MOVEMENT" not in declared_keys
and re.search(r"^\[env\] TR_MOVEMENT=", text, re.MULTILINE)
is not None)
ok = (not missing) and (not leaked_movement)
why = []
if missing:
why.append("not seen in bot env report: " + ", ".join(missing))
if leaked_movement:
why.append("undeclared TR_MOVEMENT leaked into the process")
return ok, ("; ".join(why) if why else "ok"), bool(leaked_movement)
def parse_run(opp_dir, arm_dir, run):
"""One run -> dict of MEASURED numbers, or ok=False with a reason."""
base = os.path.join(opp_dir, arm_dir)
log_path = os.path.join(base, f"run{run}.battle.log")
ev_path = os.path.join(base, f"run{run}.events.jsonl")
rounds_path = os.path.join(base, f"run{run}.jsonl.rounds.json")
log = read_text(log_path)
out = {"run": run, "ok": False, "reason": "no capture",
"damage": 0.0, "damage_taken": 0.0, "wins": None, "rounds": None,
"opp_fired": 0, "opp_hits": 0, "dist": None, "scans": 0,
"liveness": "n/a"}
if not log:
return out
counters = parse_counters(log)
if counters is None:
out["reason"] = "battle never started (no subject counters)"
return out
evs = parse_events(ev_path)
subj, other, how = attribute_subject(evs, counters)
if subj is None:
out["reason"] = f"owner attribution failed ({how})"
return out
dmg = dmg_taken = 0.0
opp_fired = 0
for ev in evs:
t = ev.get("type")
if t == "fire" and ev.get("owner") == other:
opp_fired += 1
elif t == "hit":
d = float(ev.get("damage", 0.0))
if ev.get("owner") == subj:
dmg += d
elif ev.get("owner") == other:
dmg_taken += d
opp_hits = sum(1 for ev in evs
if ev.get("type") == "hit" and ev.get("owner") == other)
wins = parse_first_places(log)
if wins is None:
out["reason"] = "no final standings in the capture"
return out
m = DIST_RE.search(log)
out.update({
"ok": True, "reason": "ok",
"counters": counters, "subject_id": subj, "owner_attribution": how,
"damage": dmg, "damage_taken": dmg_taken,
"wins": wins, "rounds": parse_rounds(rounds_path),
"opp_fired": opp_fired, "opp_hits": opp_hits,
"dist": float(m.group(1)) if m else None,
"scans": counters["scans"],
})
return out
# ── arm/opponent aggregation ─────────────────────────────────────────────────
def arm_metrics(runs):
"""Aggregate a list of valid run dicts into one metric dict."""
n = len(runs)
total_rounds = sum(r["rounds"] or 0 for r in runs)
opp_fired = sum(r["opp_fired"] for r in runs)
opp_hits = sum(r["opp_hits"] for r in runs)
dists = [r["dist"] for r in runs if r["dist"] is not None]
return {
"runs": n,
"damage": mean([r["damage"] for r in runs]),
"damage_taken": mean([r["damage_taken"] for r in runs]),
"wins": mean([r["wins"] for r in runs]),
"win_rate": (sum(r["wins"] for r in runs) / total_rounds
if total_rounds else float("nan")),
"rounds": total_rounds,
"hit_rate": (100.0 * opp_hits / opp_fired) if opp_fired else float("nan"),
"dist": mean(dists) if dists else float("nan"),
"scans": mean([r["scans"] for r in runs]),
}
def collect(session):
"""-> (data, notes); data[opp][arm] = {"valid": [...], "invalid": [...]}"""
data = {}
notes = []
for opp in session["opponents"]:
oname = opp["name"]
data[oname] = {}
opp_dir = os.path.join(session["outdir"], oname)
for arm in session["arms"]:
aname = arm["name"]
tokens = arm["env"].split() if arm["env"] else []
node = {"valid": [], "invalid": [], "tokens": tokens,
"style": opp.get("style", ""), "label": arm.get("label", "")}
for r in range(1, session["runs"] + 1):
rr = parse_run(opp_dir, aname, r)
live_ok, live_why, fatal_leak = liveness(
tokens, os.path.join(opp_dir, aname, f"run{r}.bot.stdout.log"))
rr["liveness"] = live_why
if not rr["ok"]:
node["invalid"].append(rr)
elif not live_ok:
rr["reason"] = "liveness FAIL: " + live_why
node["invalid"].append(rr)
else:
node["valid"].append(rr)
data[oname][aname] = node
return data, notes
# ── the report ───────────────────────────────────────────────────────────────
def fmt(x, nd=1):
return "n/a" if x is None or (isinstance(x, float) and math.isnan(x)) \
else f"{x:.{nd}f}"
def fmt_s(x, nd=1):
return "n/a" if x is None or (isinstance(x, float) and math.isnan(x)) \
else f"{x:+.{nd}f}"
def main():
args = sys.argv[1:]
if not args:
print(__doc__)
return 2
session_dir = args[0]
ref = None
if "--reference" in args:
ref = args[args.index("--reference") + 1]
try:
with open(os.path.join(session_dir, "session.json")) as fh:
session = json.load(fh)
except (OSError, json.JSONDecodeError) as exc:
print(f"ERROR: cannot read {session_dir}/session.json: {exc}", file=sys.stderr)
return 2
session["outdir"] = session_dir
arms = [a["name"] for a in session["arms"]]
if ref is None:
ref = session.get("reference") or arms[0]
if ref not in arms:
print(f"ERROR: reference arm '{ref}' not in session arms {arms}",
file=sys.stderr)
return 2
opps = [o["name"] for o in session["opponents"]]
style_of = {o["name"]: o.get("style", "") for o in session["opponents"]}
data, _ = collect(session)
out = []
def p(s=""):
out.append(s)
print(s)
p("### MEASURED: session")
p()
p(f"* commit `{session['commit']}`, frozen binary sha256 `{session['binary_sha256'][:12]}…`")
p(f"* {len(opps)} opponents × {len(arms)} arms × {session['runs']} runs × "
f"{session['rounds']} rounds = {len(opps) * len(arms) * session['runs']} battles, "
f"conc={session.get('conc', '?')}")
p(f"* arms file `{os.path.basename(session.get('arms_file', '?'))}`, "
f"panel file `{os.path.basename(session.get('panel_file', '?'))}`")
p(f"* reference arm: **`{ref}`** — every delta below is (arm − {ref}), "
f"opponent by opponent")
p()
invalid = [(o, a, r) for o in opps for a in arms for r in data[o][a]["invalid"]]
p(f"* liveness: {len(invalid)} run(s) excluded "
f"({len(opps) * len(arms) * session['runs']} total)")
for o, a, r in invalid[:20]:
p(f" * `{o}/{a}` run{r['run']}: {r['reason']}")
if len(invalid) > 20:
p(f" * … and {len(invalid) - 20} more")
p()
# ── per-opponent × per-arm deltas ───────────────────────────────────────
per_opp = {a: {} for a in arms}
for o in opps:
ref_runs = data[o][ref]["valid"]
ref_m = arm_metrics(ref_runs) if ref_runs else None
for a in arms:
m = arm_metrics(data[o][a]["valid"]) if data[o][a]["valid"] else None
if m is None or ref_m is None:
per_opp[a][o] = None
continue
per_opp[a][o] = {
"m": m, "ref": ref_m,
"d_damage": m["damage"] - ref_m["damage"],
"d_wins": m["wins"] - ref_m["wins"],
"d_damage_taken": m["damage_taken"] - ref_m["damage_taken"],
"d_hit_rate": m["hit_rate"] - ref_m["hit_rate"],
"d_dist": m["dist"] - ref_m["dist"],
}
# ── per-arm tables ──────────────────────────────────────────────────────
p("### MEASURED: per-opponent paired table (per arm)")
p()
for a in arms:
lbl = next((x.get("label") for x in session["arms"] if x["name"] == a), "")
cnt = sum(1 for o in opps if per_opp[a][o])
p(f"#### `{a}`" + (f" — {lbl}" if lbl else "") + f" (paired on {cnt} opponents)")
p()
p("| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |")
p("|---|---|---:|---:|---:|---:|---:|---:|---:|")
for o in opps:
e = per_opp[a][o]
if e is None:
p(f"| {o} | {style_of[o]} | n/a | n/a | n/a | n/a | n/a | n/a | n/a |")
continue
m, rm = e["m"], e["ref"]
p(f"| {o} | {style_of[o]} | {fmt(rm['damage'])}→{fmt(m['damage'])} "
f"| {fmt_s(e['d_damage'])} "
f"| {fmt(rm['wins'], 2)}→{fmt(m['wins'], 2)} | {fmt_s(e['d_wins'], 2)} "
f"| {fmt_s(e['d_damage_taken'])} | {fmt_s(e['d_hit_rate'], 2)} "
f"| {fmt(rm['dist'], 0)}→{fmt(m['dist'], 0)} |")
p()
# ── aggregate dashboard, pooled over every valid run ────────────────────
p("### MEASURED: pooled dashboard (all valid runs, NOT the verdict)")
p()
p("| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |")
p("|---|---:|---:|---:|---:|---:|---:|---:|---:|")
pooled = {}
for a in arms:
runs = [r for o in opps for r in data[o][a]["valid"]]
if not runs:
p(f"| `{a}` | 0 | n/a | n/a | n/a | n/a | n/a | n/a | n/a |")
continue
m = arm_metrics(runs)
pooled[a] = m
p(f"| `{a}` | {m['runs']} | {fmt(m['damage'])} | {fmt(m['damage_taken'])} "
f"| {fmt(m['wins'], 2)} | {int(sum(r['wins'] for r in runs))}/{m['rounds']} "
f"| {fmt(100 * m['win_rate'])}% | {fmt(m['hit_rate'], 2)}% "
f"| {fmt(m['dist'], 0)} |")
p()
# ── cross-opponent aggregation: mean delta, spread, sign tests, MDE ─────
p("### MEASURED: cross-opponent aggregation (the verdict layer)")
p()
p("Deltas are per-opponent (arm − reference). `spread` is the SD of those "
"deltas ACROSS opponents; `SE` = spread/√n; `95% CI` = mean ± t·SE. "
"Sign test = how many opponents the arm wins (ties dropped), exact "
"binomial; sign-flip = permutation test on the mean of the deltas.")
p()
p("| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |")
p("|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|")
stats = {}
for a in arms:
if a == ref:
continue
ds = [per_opp[a][o] for o in opps if per_opp[a][o]]
st = {"n": len(ds)}
for key, mkey in (("d_damage", "damage"), ("d_wins", "wins"),
("d_damage_taken", "damage_taken"),
("d_hit_rate", "hit_rate"), ("d_dist", "dist")):
vals = [e[key] for e in ds if not math.isnan(e[key])]
if len(vals) < 2:
continue
mu = mean(vals)
s = sd(vals)
se = s / math.sqrt(len(vals))
crit = t_crit(len(vals) - 1)
nz = [v for v in vals if v != 0.0]
wins_sign = sum(1 for v in nz if v > 0)
ps = binom_two_sided(wins_sign, len(nz))
pf, method = signflip_p(vals)
pw, _ = wilcoxon_p(vals)
st[key] = {"mean": mu, "sd": s, "se": se,
"ci": (mu - crit * se, mu + crit * se),
"k": wins_sign, "nz": len(nz), "p_sign": ps,
"p_flip": pf, "p_flip_method": method, "p_wilcox": pw,
"mde": mde(vals)}
p(f"| `{a}` | {mkey} | {fmt_s(mu, 2)} | {fmt(s, 2)} | {fmt(se, 2)} "
f"| [{fmt_s(mu - crit * se, 2)}, {fmt_s(mu + crit * se, 2)}] "
f"| {wins_sign}/{len(nz)} | {ps:.4g} | {pf:.4g} ({method}) "
f"| {pw:.4g} | {fmt(st[key]['mde'], 2)} |")
stats[a] = st
p()
# ── style breakdown (leg/governance only) ───────────────────────────────
styles = sorted({style_of[o] for o in opps if style_of[o]})
if len(styles) > 1:
p("#### By inferred style (explanation only, never the verdict)")
p()
p("| arm | style | n | mean Δdmg | mean Δwins | mean Δhit rate (pp) |")
p("|---|---|---:|---:|---:|---:|")
for a in arms:
if a == ref:
continue
for s in styles:
ds = [per_opp[a][o] for o in opps
if per_opp[a][o] and style_of[o] == s]
if not ds:
continue
p(f"| `{a}` | {s} | {len(ds)} "
f"| {fmt_s(mean([e['d_damage'] for e in ds]))} "
f"| {fmt_s(mean([e['d_wins'] for e in ds]), 2)} "
f"| {fmt_s(mean([e['d_hit_rate'] for e in ds]), 2)} |")
p()
# ── the pre-registered verdict rules ────────────────────────────────────
p("### The pre-registered verdict (rules fixed in `docs/movement_campaign.md`)")
p()
p("BETTER = one primary metric (dmg/run, wins/run) up at sign-test p<0.05 "
"with the other not down; WORSE = the mirror image; otherwise NOT "
"DISTINGUISHABLE. Hit rate is never the verdict.")
p()
ranked = []
for a in arms:
if a == ref:
continue
st = stats.get(a, {})
d, w = st.get("d_damage"), st.get("d_wins")
if d is None or w is None:
ranked.append((a, "n/a", float("nan"), float("nan")))
continue
up_d = d["mean"] > 0 and d["p_sign"] < 0.05
dn_d = d["mean"] < 0 and d["p_sign"] < 0.05
up_w = w["mean"] > 0 and w["p_sign"] < 0.05
dn_w = w["mean"] < 0 and w["p_sign"] < 0.05
if (up_d and w["mean"] >= 0) or (up_w and d["mean"] >= 0):
verdict = "BETTER than reference"
elif (dn_d and w["mean"] <= 0) or (dn_w and d["mean"] <= 0):
verdict = "WORSE than reference"
else:
verdict = "NOT DISTINGUISHABLE from reference"
ranked.append((a, verdict, w["mean"], d["mean"]))
ranked.sort(key=lambda t: (-(t[2] if t[2] == t[2] else -1e9),
-(t[3] if t[3] == t[3] else -1e9)))
p("| rank | arm | Δwins/run | Δdmg/run | sign test dmg | sign test wins | verdict |")
p("|---:|---|---:|---:|---|---|---|")
for i, (a, verdict, w, d) in enumerate(ranked, 1):
st = stats.get(a, {})
dd = st.get("d_damage", {})
ww = st.get("d_wins", {})
p(f"| {i} | `{a}` | {fmt_s(w, 2)} | {fmt_s(d)} "
f"| {dd.get('k', 'n/a')}/{dd.get('nz', 'n/a')} p={dd.get('p_sign', float('nan')):.4g} "
f"| {ww.get('k', 'n/a')}/{ww.get('nz', 'n/a')} p={ww.get('p_sign', float('nan')):.4g} "
f"| **{verdict}** |")
p()
p(f"Reference `{ref}`: {fmt(pooled.get(ref, {}).get('damage'))} dmg/run, "
f"{fmt(pooled.get(ref, {}).get('wins'), 2)} wins/run, "
f"{fmt(pooled.get(ref, {}).get('hit_rate'), 2)}% incoming, "
f"{fmt(pooled.get(ref, {}).get('dist'), 0)} px.")
best = ranked[0] if ranked else None
if best:
p()
p(f"Highest wins delta: `{best[0]}` ({fmt_s(best[2], 2)} wins/run, "
f"{fmt_s(best[3])} dmg/run) — **{best[1]}**.")
return 0
if __name__ == "__main__":
sys.exit(main())
+406
View File
@@ -0,0 +1,406 @@
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# tournament_run.sh — THE MULTI-OPPONENT TOURNAMENT HARNESS of the movement
# campaign: play OUR bot (one FROZEN binary from `git archive HEAD`) against a
# FIXED PANEL of opponents, once per (opponent × arm × run), and leave a session
# dir that `tournament_analyze.py` turns into a paired per-opponent ranking.
#
# tools/ab/tournament_run.sh \
# --arms tools/ab/arms_movement_b1.txt \
# --panel tools/ab/panel_movement.txt \
# --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
# --outdir /tmp/ab/j118_b1
#
# Why it exists next to ab_run.sh: ab_run.sh hard-wires exactly ONE adversary
# (DrussGT). Here the SUBJECT is our frozen bot and the ADVERSARY iterates over
# a frozen panel, so the unit of evidence becomes the NUMBER OF OPPONENTS.
# The arm file format is identical to ab_run.sh / gauntlet_run.sh.
#
# Design contract (all four are load-bearing for a valid measurement):
# 1. ONE binary for every arm. The binary is built once from `git archive
# HEAD` (a dirty tree cannot leak in) and every arm differs ONLY by the env
# dict it is launched with. No per-arm rebuild, ever.
# 2. Every run is isolated: own bot dir, own classic data dir, own output
# files, its own process group; the runner picks ephemeral ports.
# 3. A run that never really started is retried, not recorded. A battle that
# did not boot ("Only 1 of 2 bots started") would otherwise show up as a
# 0-0 phantom and poison the paired deltas.
# 4. ARENA SERIALIZATION. Other campaign jobs may be fighting at the same
# time; two concurrent battle fleets on one 16-core box distort each other.
# --wait-arena N polls for foreign bridge processes (bracketed pgrep so it
# cannot match itself) and WAITS until they are gone, up to N minutes.
# With N=0 the session REFUSES to start while the arena is busy.
#
# Arm file format: name | ENV=value ENV=value | optional label
# Panel file format: /tmp/tr_bots/Name | style | why (a bare Name works)
#
# Output layout (what tournament_analyze.py reads):
# <outdir>/session.json commit, sha, arms, panel
# <outdir>/<opponent>/<arm>/run<N>.jsonl tick capture
# <outdir>/<opponent>/<arm>/run<N>.jsonl.rounds.json
# <outdir>/<opponent>/<arm>/run<N>.jsonl.results.json
# <outdir>/<opponent>/<arm>/run<N>.events.jsonl fire/hit/death sidecar
# <outdir>/<opponent>/<arm>/run<N>.battle.log capture stdout + stats
# <outdir>/<opponent>/<arm>/run<N>.bot.stdout.log our bot's boot env report
#
# Cleanup on EXIT/INT/TERM: only THIS session's process groups, and only
# processes whose command line references THIS outdir. Never a broad pkill.
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
RUN_BRIDGE="$REPO/tools/robocode_shim/run_bridge_battle.sh"
ARMS_FILE=""
PANEL_FILE=""
RUNS=3
ROUNDS=3
CONC=6
OUTDIR=""
WAIT_ARENA=0
FORCE_BUSY=0
REFERENCE=""
# A bot that fails to connect inside the booter's window is a FALSE NEGATIVE.
# Retry until the capture shows real rows AND the subject counter line.
MAX_ATTEMPTS="${TOURNAMENT_MAX_ATTEMPTS:-4}"
# 800x600 default arena; recorded for the ledger.
GAME_TYPE="${TOURNAMENT_GAME_TYPE:-1v1}"
usage() {
sed -n '2,45p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
exit "${1:-0}"
}
while [[ $# -gt 0 ]]; do
case "$1" in
--arms) ARMS_FILE="$2"; shift 2;;
--panel|--opponents) PANEL_FILE="$2"; shift 2;;
--runs) RUNS="$2"; shift 2;;
--rounds) ROUNDS="$2"; shift 2;;
--conc) CONC="$2"; shift 2;;
--outdir) OUTDIR="$2"; shift 2;;
--wait-arena) WAIT_ARENA="$2"; shift 2;;
--force-when-busy) FORCE_BUSY=1; shift;;
--reference) REFERENCE="$2"; shift 2;;
-h|--help) usage 0;;
*) echo "unknown argument: $1" >&2; usage 1;;
esac
done
[[ -n "$ARMS_FILE" ]] || { echo "ERROR: --arms <file> is required" >&2; usage 1; }
[[ -n "$PANEL_FILE" ]] || { echo "ERROR: --panel <file> is required" >&2; usage 1; }
[[ -f "$ARMS_FILE" ]] || { echo "ERROR: arm file not found: $ARMS_FILE" >&2; exit 1; }
[[ -f "$PANEL_FILE" ]] || { echo "ERROR: panel file not found: $PANEL_FILE" >&2; exit 1; }
[[ -n "$OUTDIR" ]] || { echo "ERROR: --outdir <dir> is required" >&2; exit 1; }
[[ "$OUTDIR" != "/" && "$OUTDIR" != "" ]] || { echo "ERROR: refusing outdir '$OUTDIR'" >&2; exit 1; }
mkdir -p "$OUTDIR"
OUTDIR="$(cd "$OUTDIR" && pwd)"
# never nuke a directory that is not one of ours (non-empty without a session.json)
if [[ -n "$(ls -A "$OUTDIR" 2>/dev/null)" && ! -f "$OUTDIR/session.json" ]]; then
echo "ERROR: refusing to clear non-empty, non-session dir: $OUTDIR" >&2
echo " (delete it by hand or pick another --outdir)" >&2
exit 1
fi
# ── prerequisites ────────────────────────────────────────────────────────────
ROBOCODE_JAR="${ROBOCODE_JAR:-/tmp/robocode/install/libs/robocode.jar}"
RUNNER_JAR="${TR_RUNNER_JAR:-/home/davide/Projects/tank-royale/runner/examples/lib/robocode-tankroyale-runner.jar}"
BOT_API_JAR="${TR_BOT_API_JAR:-$HOME/Downloads/sample-bots-java-1.0.2/lib/robocode-tankroyale-bot-api-1.0.2.jar}"
missing=0
for f in "$ROBOCODE_JAR" "$RUNNER_JAR" "$BOT_API_JAR"; do
[[ -f "$f" ]] || { echo "ERROR: missing prerequisite: $f" >&2; missing=1; }
done
command -v nim >/dev/null 2>&1 || { echo "ERROR: nim not on PATH" >&2; missing=1; }
command -v git >/dev/null 2>&1 || { echo "ERROR: git not on PATH" >&2; missing=1; }
[[ -f "$RUN_BRIDGE" ]] || { echo "ERROR: missing $RUN_BRIDGE" >&2; missing=1; }
(( missing == 0 )) || { echo "ERROR: prerequisites missing — not starting a broken session" >&2; exit 1; }
# ── ARENA SERIALIZATION ──────────────────────────────────────────────────────
# Bracketed patterns: pgrep -f must not match this script's own command line.
FOREIGN_PATTERN='[r]un_bridge_battle.sh|[r]obocode_shim.TrBattleCapture|[r]obocode_shim.LegacyBotBridge|[M]odularBot_bin'
foreign_battles() {
pgrep -af "$FOREIGN_PATTERN" 2>/dev/null | grep -v "$OUTDIR" || true
}
check_arena() {
local found
found="$(foreign_battles)"
[[ -z "$found" ]]
}
if ! check_arena; then
if (( FORCE_BUSY == 1 )); then
echo "[tournament] WARNING: arena busy but --force-when-busy was given:" >&2
foreign_battles | sed 's/^/ /' >&2
elif (( WAIT_ARENA > 0 )); then
echo "[tournament] arena busy — waiting up to ${WAIT_ARENA} min for foreign battles:"
foreign_battles | sed 's/^/ /'
deadline=$(( $(date +%s) + WAIT_ARENA * 60 ))
while :; do
sleep 120
if check_arena; then echo "[tournament] arena free."; break; fi
if (( $(date +%s) >= deadline )); then
echo "[tournament] ERROR: arena still busy after ${WAIT_ARENA} min; NOT starting." >&2
foreign_battles | sed 's/^/ /' >&2
exit 3
fi
echo "[tournament] still busy at $(date +%H:%M:%S); $(foreign_battles | wc -l) foreign process(es)"
done
else
echo "[tournament] ERROR: the arena is busy with another job's battles:" >&2
foreign_battles | sed 's/^/ /' >&2
echo "[tournament] rerun with --wait-arena <minutes> to wait for it, or --force-when-busy." >&2
exit 3
fi
fi
# ── parse the arm file ───────────────────────────────────────────────────────
trim() { local s="$1"; s="${s#"${s%%[![:space:]]*}"}"; s="${s%"${s##*[![:space:]]}"}"; printf '%s' "$s"; }
ARM_NAMES=(); ARM_ENVS=(); ARM_LABELS=()
while IFS= read -r line || [[ -n "$line" ]]; do
[[ -z "${line//[[:space:]]/}" ]] && continue
[[ "$line" =~ ^[[:space:]]*# ]] && continue
name="$(trim "${line%%|*}")"
rest="${line#*|}"
if [[ "$line" == *"|"* ]]; then
envspec="$(trim "${rest%%|*}")"
label="$(trim "${rest#*|}")"
else
envspec=""; label=""
fi
[[ -n "$name" ]] || { echo "ERROR: arm with empty name in $ARMS_FILE" >&2; exit 1; }
# every env token must be VAR=value (a typo'd arm must fail loudly, not quietly)
for tok in $envspec; do
[[ "$tok" =~ ^[A-Za-z_][A-Za-z0-9_]*=.*$ ]] \
|| { echo "ERROR: arm '$name' has a malformed env token: '$tok'" >&2; exit 1; }
done
ARM_NAMES+=("$name"); ARM_ENVS+=("$envspec"); ARM_LABELS+=("$label")
done < "$ARMS_FILE"
(( ${#ARM_NAMES[@]} > 0 )) || { echo "ERROR: no arms parsed from $ARMS_FILE" >&2; exit 1; }
dupes="$(printf '%s\n' "${ARM_NAMES[@]}" | sort | uniq -d)"
[[ -z "$dupes" ]] || { echo "ERROR: duplicate arm name(s): $dupes" >&2; exit 1; }
if [[ -n "$REFERENCE" ]]; then
printf '%s\n' "${ARM_NAMES[@]}" | grep -qx "$REFERENCE" \
|| { echo "ERROR: --reference '$REFERENCE' is not one of the arms" >&2; exit 1; }
fi
# ── parse the panel file ─────────────────────────────────────────────────────
OPP_DIRS=(); OPP_NAMES=(); OPP_STYLES=()
while IFS= read -r line || [[ -n "$line" ]]; do
[[ -z "${line//[[:space:]]/}" ]] && continue
[[ "$line" =~ ^[[:space:]]*# ]] && continue
dir="$(trim "${line%%|*}")"
style=""; note=""
if [[ "$line" == *"|"* ]]; then
rest="${line#*|}"
if [[ "$rest" == *"|"* ]]; then
style="$(trim "${rest%%|*}")"; note="$(trim "${rest#*|}")"
else
style="$(trim "$rest")"
fi
fi
[[ "$dir" == */* ]] || dir="/tmp/tr_bots/$dir"
[[ -d "$dir" ]] || { echo "ERROR: opponent bot dir not found: $dir" >&2; exit 1; }
OPP_DIRS+=("$dir"); OPP_NAMES+=("$(basename "$dir")"); OPP_STYLES+=("$style")
done < "$PANEL_FILE"
(( ${#OPP_NAMES[@]} > 0 )) || { echo "ERROR: no opponents parsed from $PANEL_FILE" >&2; exit 1; }
dupes="$(printf '%s\n' "${OPP_NAMES[@]}" | sort | uniq -d)"
[[ -z "$dupes" ]] || { echo "ERROR: duplicate opponent name(s) in the panel: $dupes" >&2; exit 1; }
(( ${#OPP_NAMES[@]} >= 4 )) || { echo "ERROR: a panel below 4 opponents cannot carry a sign test" >&2; exit 1; }
# ── fresh session dir ────────────────────────────────────────────────────────
rm -rf "$OUTDIR"
mkdir -p "$OUTDIR"
WORK="$OUTDIR/.work"; mkdir -p "$WORK"
# ── build ONE frozen bot from HEAD (single binary for every arm) ─────────────
COMMIT="$(git -C "$REPO" rev-parse HEAD)"
FROZEN_DIR="$OUTDIR/frozen/ModularBot"
FROZEN_BIN="$FROZEN_DIR/ModularBot_bin"
FROZEN_JSON="$FROZEN_DIR/ModularBot.json"
mkdir -p "$FROZEN_DIR"
BUILDDIR="$WORK/head"; mkdir -p "$BUILDDIR"
echo "[tournament] exporting HEAD ($COMMIT) -> $BUILDDIR"
git -C "$REPO" archive HEAD | tar -x -C "$BUILDDIR"
NIMCACHE="${TOURNAMENT_NIMCACHE:-/tmp/nc_j118}"
echo "[tournament] building the frozen ModularBot (nim c -d:release, nimcache=$NIMCACHE)…"
( cd "$BUILDDIR/ModularBot_garage" && \
nim c -d:release --nimcache:"$NIMCACHE" --out:"$FROZEN_BIN" src/ModularBot.nim ) \
> "$OUTDIR/frozen/build.log" 2>&1 \
|| { echo "ERROR: frozen build failed — see $OUTDIR/frozen/build.log" >&2; tail -30 "$OUTDIR/frozen/build.log" >&2; exit 1; }
[[ -x "$FROZEN_BIN" ]] || { echo "ERROR: frozen build produced no binary" >&2; exit 1; }
BINSHA="$(sha256sum "$FROZEN_BIN" | awk '{print $1}')"
echo "[tournament] frozen binary sha256=$BINSHA"
cat > "$FROZEN_JSON" <<'JSON'
{
"name": "ModularBot",
"version": "0.1.0",
"authors": ["Davide Cappellini"],
"description": "frozen tournament build of ModularBot (tournament_run.sh)",
"homepage": "",
"countryCodes": ["IT"],
"gameTypes": ["classic", "1v1"],
"platform": "Nim",
"programmingLang": "Nim"
}
JSON
# ── session.json ─────────────────────────────────────────────────────────────
json_esc() { printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g'; }
{
printf '{\n'
printf ' "commit": "%s",\n' "$COMMIT"
printf ' "binary_sha256": "%s",\n' "$BINSHA"
printf ' "binary": "frozen/ModularBot/ModularBot_bin",\n'
printf ' "rounds": %d,\n' "$ROUNDS"
printf ' "runs": %d,\n' "$RUNS"
printf ' "conc": %d,\n' "$CONC"
printf ' "game_type": "%s",\n' "$GAME_TYPE"
printf ' "reference": "%s",\n' "$(json_esc "$REFERENCE")"
printf ' "timestamp": "%s",\n' "$(date -Is)"
printf ' "outdir": "%s",\n' "$(json_esc "$OUTDIR")"
printf ' "arms_file": "%s",\n' "$(json_esc "$ARMS_FILE")"
printf ' "panel_file": "%s",\n' "$(json_esc "$PANEL_FILE")"
printf ' "arms": [\n'
for i in "${!ARM_NAMES[@]}"; do
printf ' {"name": "%s", "env": "%s", "label": "%s"}' \
"$(json_esc "${ARM_NAMES[$i]}")" "$(json_esc "${ARM_ENVS[$i]}")" "$(json_esc "${ARM_LABELS[$i]}")"
(( i + 1 < ${#ARM_NAMES[@]} )) && printf ',' || true
printf '\n'
done
printf ' ],\n'
printf ' "opponents": [\n'
for i in "${!OPP_NAMES[@]}"; do
printf ' {"name": "%s", "dir": "%s", "style": "%s"}' \
"$(json_esc "${OPP_NAMES[$i]}")" "$(json_esc "${OPP_DIRS[$i]}")" "$(json_esc "${OPP_STYLES[$i]}")"
(( i + 1 < ${#OPP_NAMES[@]} )) && printf ',' || true
printf '\n'
done
printf ' ]\n}\n'
} > "$OUTDIR/session.json"
# ── one job ──────────────────────────────────────────────────────────────────
export REPO RUN_BRIDGE ROUNDS OUTDIR WORK FROZEN_BIN FROZEN_JSON MAX_ATTEMPTS
tournament_run_one() {
local opponent_dir="$1" opponent="$2" arm="$3" run="$4" envspec="$5"
local dir="$OUTDIR/$opponent/$arm"
local work="$WORK/$opponent/$arm/run$run"
local botdir="$work/bots/ModularBot"
mkdir -p "$dir" "$botdir" "$work/data"
[[ -n "$envspec" ]] && printf '%s\n' "$envspec" > "$dir/run$run.arm.env"
cp "$FROZEN_JSON" "$botdir/ModularBot.json"
cat > "$botdir/ModularBot.sh" <<SH
#!/bin/sh
cd "\$(dirname "\$0")"
exec ./ModularBot_bin > "$dir/run$run.bot.stdout.log" 2> "$dir/run$run.bot.stderr.log"
SH
chmod +x "$botdir/ModularBot.sh"
ln -sf "$FROZEN_BIN" "$botdir/ModularBot_bin"
local attempt=0
while :; do
attempt=$((attempt + 1))
rm -f "$dir/run$run.jsonl" "$dir/run$run.jsonl.rounds.json" \
"$dir/run$run.jsonl.results.json" "$dir/run$run.events.jsonl" \
"$dir/run$run.status" "$dir/run$run.bot.stdout.log" "$dir/run$run.bot.stderr.log"
local rc=0
# shellcheck disable=SC2086 # envspec is intentionally word-split
env $envspec \
SHIM_BOTDIR="$botdir" \
SHIM_BOT_NAME="ModularBot" \
SHIM_DATA="$work/data" \
TR_EVENTS_OUT="$dir/run$run.events.jsonl" \
timeout 900 "$RUN_BRIDGE" "$opponent_dir" "$ROUNDS" "$dir/run$run.jsonl" \
> "$dir/run$run.battle.log" 2>&1 || rc=$?
echo "$rc" > "$dir/run$run.status"
# liveness: a real battle has captured rows AND the subject counter line
if grep -q "subject event counts" "$dir/run$run.battle.log" 2>/dev/null \
&& grep -Eq "rows=[1-9]" "$dir/run$run.battle.log" 2>/dev/null; then
break
fi
if (( attempt >= MAX_ATTEMPTS )); then
echo "[tournament] WARNING: ${opponent}/${arm} run${run} never started after ${attempt} attempts" >&2
break
fi
sleep $((attempt * 2))
done
echo "$attempt" > "$dir/run$run.attempts"
return 0
}
export -f tournament_run_one
# ── cleanup: only this session's process groups, only this outdir ────────────
JOB_PIDS=()
cleanup() {
local rc=$?
trap - EXIT INT TERM
for p in "${JOB_PIDS[@]:-}"; do
[[ -n "$p" ]] || continue
kill -TERM -- "-$p" 2>/dev/null || kill -TERM "$p" 2>/dev/null || true
done
sleep 0.5
for p in "${JOB_PIDS[@]:-}"; do
[[ -n "$p" ]] || continue
kill -KILL -- "-$p" 2>/dev/null || true
done
pkill -f "[r]un_bridge_battle.sh.*$OUTDIR" 2>/dev/null || true
pkill -f "[r]obocode_shim.TrBattleCapture.*$OUTDIR" 2>/dev/null || true
exit "$rc"
}
trap cleanup EXIT INT TERM
# ── launch, K at a time ──────────────────────────────────────────────────────
NJOBS=$(( ${#OPP_NAMES[@]} * ${#ARM_NAMES[@]} * RUNS ))
SECONDS=0
echo "[tournament] session: ${#OPP_NAMES[@]} opponents × ${#ARM_NAMES[@]} arms × $RUNS runs = $NJOBS battles, rounds=$ROUNDS, conc=$CONC"
echo "[tournament] outdir: $OUTDIR"
RUNNING=0; DONE=0
for oi in "${!OPP_NAMES[@]}"; do
for ai in "${!ARM_NAMES[@]}"; do
for (( r=1; r<=RUNS; r++ )); do
setsid bash -c 'tournament_run_one "$1" "$2" "$3" "$4" "$5"' _ \
"${OPP_DIRS[$oi]}" "${OPP_NAMES[$oi]}" "${ARM_NAMES[$ai]}" "$r" "${ARM_ENVS[$ai]}" &
JOB_PIDS+=("$!")
RUNNING=$((RUNNING + 1))
if (( RUNNING >= CONC )); then
wait -n || true
RUNNING=$((RUNNING - 1)); DONE=$((DONE + 1))
printf '[tournament] %3d/%3d done (%ds)\n' "$DONE" "$NJOBS" "$SECONDS"
fi
done
done
done
wait || true
DONE=$NJOBS
printf '[tournament] %3d/%3d done (%ds)\n' "$DONE" "$NJOBS" "$SECONDS"
JOB_PIDS=()
# ── aggregate status ─────────────────────────────────────────────────────────
FAILED=0; STARTFAIL=0
for oi in "${!OPP_NAMES[@]}"; do
for ai in "${!ARM_NAMES[@]}"; do
for (( r=1; r<=RUNS; r++ )); do
st="$(cat "$OUTDIR/${OPP_NAMES[$oi]}/${ARM_NAMES[$ai]}/run$r.status" 2>/dev/null || echo 999)"
if [[ "$st" != "0" ]]; then
echo "[tournament] FAIL ${OPP_NAMES[$oi]}/${ARM_NAMES[$ai]} run$r (rc=$st)" >&2
FAILED=$((FAILED + 1))
fi
if ! grep -q "subject event counts" \
"$OUTDIR/${OPP_NAMES[$oi]}/${ARM_NAMES[$ai]}/run$r.battle.log" 2>/dev/null; then
STARTFAIL=$((STARTFAIL + 1))
fi
done
done
done
echo "[tournament] done: $(( NJOBS - FAILED )) ok, $FAILED failed, $STARTFAIL never started ($(( SECONDS ))s)"
echo "[tournament] wall clock: ${SECONDS}s for $NJOBS battles"
echo "[tournament] analyze: python3 $REPO/tools/ab/tournament_analyze.py $OUTDIR${REFERENCE:+ --reference $REFERENCE}"
(( FAILED == 0 )) || exit 1