melee A/B doc: correct the per-arm [bb-reset] counts (perRound 68.75, retained 68.19, decay 66.75)
This commit is contained in:
@@ -186,9 +186,9 @@ Target switching is the whole mechanism, so it had to be proven, not assumed.
|
|||||||
**66–69 target changes/run** (about 10 per round). Example transitions from one
|
**66–69 target changes/run** (about 10 per round). Example transitions from one
|
||||||
run: `#1 → #4 → #2 → #4 → #1 → #2 → …`.
|
run: `#1 → #4 → #2 → #4 → #1 → #2 → …`.
|
||||||
* **BitBrain really reset on every switch.** `[bb-reset] reason=target_change`
|
* **BitBrain really reset on every switch.** `[bb-reset] reason=target_change`
|
||||||
fires once per target change: **66.75** (perRound), **68.19** (retained),
|
fires once per target change: **68.75** (perRound), **68.19** (retained),
|
||||||
**66.75** (decay) per run — equal, within rounding, to the target-change count.
|
**66.75** (decay) per run — equal, within rounding, to each arm's target-change
|
||||||
Example `[bb]` line from a live run:
|
count. Example `[bb]` line from a live run:
|
||||||
`[bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound`.
|
`[bb] t=440 band=450.+ gain=0.25 shift=-3.39deg rate=0.818 n=11. ncand=5 trained=11 pend=29 dropped=95 mode=perRound`.
|
||||||
The `pattern` arm logs **zero** `[bb]`/`[bb-reset]` lines (BitBrain never
|
The `pattern` arm logs **zero** `[bb]`/`[bb-reset]` lines (BitBrain never
|
||||||
spawned), confirming the arms are cleanly separated.
|
spawned), confirming the arms are cleanly separated.
|
||||||
|
|||||||
@@ -0,0 +1,179 @@
|
|||||||
|
# Movement campaign — ledger
|
||||||
|
|
||||||
|
**Goal (owner's mandate, 2026-09-26 overnight):** find the *best 1v1 movement*
|
||||||
|
by measurement, then do the same for the gun. This file is the campaign's
|
||||||
|
single source of truth: every later job **appends** a `## Batch N` section and
|
||||||
|
never edits an earlier one (a wrong earlier number gets a correction line, not
|
||||||
|
a rewrite).
|
||||||
|
|
||||||
|
**Owner's words:** *"I want you to do all tests and checks with the goal to have
|
||||||
|
the best 1vs1 movement. You have all night, you can change every parameter.
|
||||||
|
Continue until you found an amazing movement. When found do the same over for a
|
||||||
|
gun."*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0. The one caveat this campaign exists to close
|
||||||
|
|
||||||
|
Everything measured about movement before this campaign is **DrussGT-only**:
|
||||||
|
`docs/surfer_wiring_ab.md`, the j107 range drift, the j113 BitBrain movement
|
||||||
|
notes. The standing lesson of the night is that a one-opponent result is not a
|
||||||
|
result:
|
||||||
|
|
||||||
|
> an arm can take fewer hits **and** win fewer rounds (j107 / `strafe`): the
|
||||||
|
> verdict lives in **damage/run + ROUND WINS**, and hit rate is only ever an
|
||||||
|
> explanation.
|
||||||
|
|
||||||
|
So from here on **the unit of evidence is the number of opponents**, not the
|
||||||
|
number of runs: the same arm must win on *many* opponents before it is called
|
||||||
|
better.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Protocol (how every batch must be run)
|
||||||
|
|
||||||
|
| Element | Rule |
|
||||||
|
|---|---|
|
||||||
|
| Subject | ONE frozen binary, built from `git archive HEAD` (`tools/ab/tournament_run.sh` does this; the commit sha and binary sha256 are recorded in `session.json`) |
|
||||||
|
| Arms | env dicts only — **no per-arm rebuild, ever**; the arm file is a committed file, not a shell history |
|
||||||
|
| Panel | the **frozen** panel `tools/ab/panel_movement.txt`. Adding/removing an opponent starts a **new batch number** |
|
||||||
|
| Pairing | per opponent: average the arm's runs, subtract the reference arm's average for that same opponent → one delta per opponent; then aggregate |
|
||||||
|
| Isolation | per-run bot dir + classic data dir, ephemeral ports, own process group; cleanup only by this session's outdir |
|
||||||
|
| Serialization | **one battle fleet at a time.** `tournament_run.sh --wait-arena N` refuses/stalls while another job's `run_bridge_battle`/`TrBattleCapture`/`ModularBot_bin` is alive (bracketed pgrep; never a broad `pkill`) |
|
||||||
|
| Liveness | every declared env token must appear verbatim in OUR bot's own `[env]` boot report, else the run is excluded and named in the report; an undeclared `TR_MOVEMENT` in the process env is a fatal FAIL for the reference arm |
|
||||||
|
| Never shipped | this is a measurement + design campaign: `git status` clean, defaults untouched, `.gitignore` untouched |
|
||||||
|
|
||||||
|
### Pre-registered decision rules (fixed BEFORE Batch 1 ran, commit `__PRECOMMIT__`)
|
||||||
|
|
||||||
|
1. **Primary metrics:** damage/run and ROUND WINS. Secondary/explanation only:
|
||||||
|
damage taken/run, incoming hit rate (enemy hits ÷ enemy shots), achieved mean
|
||||||
|
distance.
|
||||||
|
2. **BETTER than the reference** iff one primary metric is up with a
|
||||||
|
cross-opponent **sign test p < 0.05** while the other does **not** go down;
|
||||||
|
or the mirror image for **WORSE**. Anything else is **NOT
|
||||||
|
DISTINGUISHABLE** (which is a real answer, not a failure).
|
||||||
|
3. **A verdict must survive the between-opponent spread**: the pooled mean delta
|
||||||
|
is reported with the SD across opponents, its SE, a 95% CI, and the MDE
|
||||||
|
(α=0.05 two-sided, 80% power) — an effect smaller than the MDE is reported as
|
||||||
|
*not detectable*, never as *absent* and never as a win.
|
||||||
|
4. **Somewhere to stop:** if no arm beats the shipped `tfil` by rule 2 in
|
||||||
|
Batch 1 **and** no arm shows a ≥ +MDE damage gain with p<0.10, the movement
|
||||||
|
stage's first phase is closed with *"the shipped `tfil` is the best movement
|
||||||
|
we have measured"* — that is a **successful** outcome, and the campaign moves
|
||||||
|
to the gun axis rather than inventing more movement arms. See
|
||||||
|
*What would make us stop* at the end.
|
||||||
|
5. **No promotion off a single metric, a single opponent, or a single run.**
|
||||||
|
A change that wins damage by losing wins (or vice-versa) is not a win.
|
||||||
|
6. Every batch is shot with a **pre-registered prediction** stated in its
|
||||||
|
section *before* the battles finish; a prediction that turns out wrong is
|
||||||
|
recorded as wrong.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Stage 0 — what we already know (given, not re-derived)
|
||||||
|
|
||||||
|
Live A/B vs real DrussGT, 15 runs × 7 rounds, one frozen binary
|
||||||
|
(`docs/surfer_wiring_ab.md`, commit `0f5cfe3`):
|
||||||
|
|
||||||
|
| arm | dmg/run | dmg taken | round wins | incoming hit rate |
|
||||||
|
|---|---:|---:|---:|---:|
|
||||||
|
| `tfil` (SHIPPED) | **293** | 224 | **45/105** | 10.40% |
|
||||||
|
| `strafe` (range 325) | 250 | **198** | 37/105 | **9.40%** |
|
||||||
|
| `surf` | 255 | 259 | 37/105 | 13.51% |
|
||||||
|
|
||||||
|
Read: the shipped `tfil` deals the most damage and wins the most rounds while
|
||||||
|
being hit the *most*; `strafe` dodges best and wins least. Plus j107: drifting
|
||||||
|
25–30 px closer made damage **and** wins worse, so the lever is not simply "get
|
||||||
|
closer". **Hypothesis entering the campaign: the 325 px range preference of
|
||||||
|
`strafe` costs wins** (INFERRED from DrussGT-only data — this is exactly what
|
||||||
|
Batch 1 tests across a panel).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Batch 1 — isolating the range / aggression axis
|
||||||
|
|
||||||
|
**Design.** One frozen binary, five env-only arms, one frozen panel
|
||||||
|
(`tools/ab/panel_movement.txt`, 15 opponents: 5 dodger, 3 pattern, 2
|
||||||
|
wall-follower/corner-camper, 1 spinner, 2 rammer/brawler, 2 aggressive megas),
|
||||||
|
3 runs × 3 rounds per (opponent, arm). Arm file:
|
||||||
|
`tools/ab/arms_movement_b1.txt`.
|
||||||
|
|
||||||
|
| # | arm | env | what it isolates |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 1 | `tfil` | *(none — shipped defaults)* | the arm to beat |
|
||||||
|
| 2 | `strafe_notilt` | `TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999` | the COST of the 325 range preference: tilt is provably 0 every tick, so this is pure perpendicular strafe with **no range steering at all** |
|
||||||
|
| 3 | `strafe_325` | `TR_MOVEMENT=strafe` | the current strafe default (range 325, tol 25, tilt 15/0.10) |
|
||||||
|
| 4 | `ring` | `TR_MOVEMENT=tfil_ring` | TFIL semantics + retuned heat field (corridor 10, wall 15, radiance 5, bullet core/aura 20/10, 5-tick commit) **with** the range-weighted tile draw (band 100–200) |
|
||||||
|
| 5 | `ring_notemp` | `TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0` | the control for #4: same retuned heat field, range weighting switched OFF (`rand(candidates.high)` path) |
|
||||||
|
|
||||||
|
`ring` − `ring_notemp` is therefore the range-weighting lever **alone**, on a
|
||||||
|
heat field that is already retuned. The originally-suggested 5th arm ("`tfil`
|
||||||
|
with less saturated heat") is **not buildable in this campaign**: in
|
||||||
|
`common_libs/movements/the_floor_is_lava.nim` `CorridorHeat`/`WallHotness` are
|
||||||
|
Nim `const`s (env_report only *reports* them); only the `tfil_ring` copy reads
|
||||||
|
them from the env. #5 is the honest substitute.
|
||||||
|
|
||||||
|
**Pre-registered prediction (written before the battles finished):** `tfil`
|
||||||
|
still wins the panel on damage and round wins; `strafe_notilt` will beat
|
||||||
|
`strafe_325` on round wins (the range tilt is a net cost), and the ring arms will
|
||||||
|
land between them. If instead the range-steering arms beat `tfil` on wins, the
|
||||||
|
"range preference costs wins" hypothesis is confirmed across bots, not just
|
||||||
|
against DrussGT.
|
||||||
|
|
||||||
|
<!-- BATCH1-RESULTS -->
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. What to try next (seeded; every later job adds its own)
|
||||||
|
|
||||||
|
1. **If a range-steering arm wins the panel:** the win is a *range* effect, so
|
||||||
|
sweep the *band* on the winning engine (e.g. `TR_TFIL_RANGE_LO/HI` on
|
||||||
|
`tfil_ring`, `TR_STRAFE_RANGE` on `strafe`) with the same panel, 4–5 bands,
|
||||||
|
and look for a plateau rather than a peak. A plateau is a result; a peak is a
|
||||||
|
coin flip.
|
||||||
|
2. **If nothing beats `tfil`:** stop tuning movement by feel. The next honest
|
||||||
|
lever is *enemy-model-driven* placement (keep the bot where the enemy's
|
||||||
|
expected hit probability is lowest *given its gun model*), which needs a
|
||||||
|
per-opponent measurement, not a knob.
|
||||||
|
3. **Melee is a different game** (j116's finding): if a melee campaign is
|
||||||
|
opened, it needs its own panel and its own ledger section — do not reuse the
|
||||||
|
1v1 panel's verdicts.
|
||||||
|
4. **Close the loop with the enemy's own model:** `ab_mechanism.py`-style
|
||||||
|
instrumentation (time spent within 100 px of a live enemy bullet, hit-rate by
|
||||||
|
range band) is available and cheap — use it to *explain* a win, never to
|
||||||
|
declare one.
|
||||||
|
5. **Then the gun** (owner's next stage): the same harness, a gun panel, and the
|
||||||
|
same paired-with-sign-test statistics. `docs/surfer_wiring_ab.md` and the j117
|
||||||
|
gauntlet are the gun-side priors to beat.
|
||||||
|
|
||||||
|
## 5. What would make us stop
|
||||||
|
|
||||||
|
* **Stop the movement stage** when a batch produces an arm that is BETTER than
|
||||||
|
the shipped default on the frozen panel by rule 2 **and** the effect survives
|
||||||
|
the between-opponent spread (|Δ| > MDE, or a sign test that wins on ≥ 2/3 of
|
||||||
|
the panel). That arm becomes the new default *candidate* (shipping is a
|
||||||
|
separate decision — this campaign never edits a shipped default).
|
||||||
|
* **Stop and move to the gun** if two consecutive batches fail to produce an arm
|
||||||
|
that beats the shipped `tfil` beyond the MDE: at that point the honest
|
||||||
|
conclusion is *"the shipped movement is the measured optimum of this design
|
||||||
|
space"*, which is a successful campaign outcome, not a failure.
|
||||||
|
* **Stop a single batch early** only for a contract violation (arena not free,
|
||||||
|
liveness FAIL, non-zero exit rate) — never because the numbers look boring.
|
||||||
|
|
||||||
|
## 6. How to run a batch (exact commands)
|
||||||
|
|
||||||
|
```sh
|
||||||
|
# 1. wait for the arena (this job may not be the only one fighting)
|
||||||
|
tools/ab/tournament_run.sh \
|
||||||
|
--arms tools/ab/arms_movement_b1.txt \
|
||||||
|
--panel tools/ab/panel_movement.txt \
|
||||||
|
--runs 3 --rounds 3 --conc 6 --wait-arena 45 \
|
||||||
|
--reference tfil \
|
||||||
|
--outdir /tmp/ab/j118_b1
|
||||||
|
|
||||||
|
# 2. the paired per-opponent table, sign tests, MDE and the pre-registered verdict
|
||||||
|
python3 tools/ab/tournament_analyze.py /tmp/ab/j118_b1 --reference tfil
|
||||||
|
```
|
||||||
|
|
||||||
|
The runner writes `<outdir>/session.json` (commit sha, binary sha256, arms,
|
||||||
|
panel) so any later job can re-analyze an old session offline, with no arena.
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
|
# arms_movement_b1.txt — BATCH 1 of the movement campaign: isolate the
|
||||||
|
# RANGE / AGGRESSION axis. Five arms, ONE frozen binary, env vars only.
|
||||||
|
#
|
||||||
|
# Format: name | ENV=value ENV=value | label
|
||||||
|
#
|
||||||
|
# Prior data this batch is built on (all DrussGT-only, all MEASURED):
|
||||||
|
# * shipped `tfil` 293 dmg/run, 45/105 round wins, 10.40% incoming hit rate
|
||||||
|
# * `strafe` (range 325) 250 dmg/run, 37/105 wins, 9.40% incoming -> best
|
||||||
|
# dodger, fewest wins: the 325px range preference may COST wins
|
||||||
|
# * j107: drifting 25-30 px closer made damage AND wins worse, so "get
|
||||||
|
# closer" is not the lever by itself
|
||||||
|
# This batch closes the caveat that all of that is DrussGT-only and decomposes
|
||||||
|
# the range-steering lever from the engine.
|
||||||
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
# 1. the arm to beat: shipped default, no env at all
|
||||||
|
tfil | | shipped baseline (movement engine tfil, every knob at its default)
|
||||||
|
|
||||||
|
# 2. pure perpendicular strafe with the range steering DISABLED: tilt is zero
|
||||||
|
# whenever |dist - StrafeRange| <= TOL, so a huge TOL makes `tilt` provably
|
||||||
|
# 0 every tick (and the approach-direction branch that depends on tiltMag is
|
||||||
|
# skipped). This isolates the COST of the 325px preference: same engine, same
|
||||||
|
# wall logic, no range steering at all.
|
||||||
|
strafe_notilt | TR_MOVEMENT=strafe TR_STRAFE_RANGE_TOL=999999 | strafe, range steering OFF (tilt always 0)
|
||||||
|
|
||||||
|
# 3. the current strafe default: range 325, tol 25, tilt max 15, gain 0.10
|
||||||
|
strafe_325 | TR_MOVEMENT=strafe | strafe default (range 325, tol 25, tilt 15/0.10)
|
||||||
|
|
||||||
|
# 4. the owner's own retuned mover: tfil semantics + retuned heat field
|
||||||
|
# (CorridorHeat 10 = the safety threshold, WallHotness 15, WallRadiance 5,
|
||||||
|
# BulletCore/Aura 20/10, commit 5 ticks) + range-weighted tile draw
|
||||||
|
# (band 100-200, temp 0.4, K 60). Range steering ON, different heat field.
|
||||||
|
ring | TR_MOVEMENT=tfil_ring | tfil_ring (retuned heat field + range weighting 100-200)
|
||||||
|
|
||||||
|
# 5. the control for arm 4: TR_TFIL_RANGE_TEMP=0 makes the tile draw UNIFORM
|
||||||
|
# again (the documented OFF path), so arm 5 = the retuned heat field with NO
|
||||||
|
# range steering. (4) - (5) is therefore the range-weighting lever alone, on
|
||||||
|
# a heat field that is already retuned. NOTE: there is no env knob for the
|
||||||
|
# heat constants of the SHIPPED tfil mover (CorridorHeat/WallHotness are
|
||||||
|
# Nim `const` there, only REPORTED by env_report.nim), so the suggested
|
||||||
|
# "tfil with less saturated heat" arm cannot be built without a rebuild,
|
||||||
|
# which this campaign forbids. This is the honest substitute.
|
||||||
|
ring_notemp | TR_MOVEMENT=tfil_ring TR_TFIL_RANGE_TEMP=0 | tfil_ring, range weighting OFF (temp 0)
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
|
# panel_movement.txt — THE FROZEN MOVEMENT PANEL (Batch 1, 2026-09-26)
|
||||||
|
#
|
||||||
|
# DO NOT ADD, REMOVE OR REORDER AN OPPONENT without starting a new batch number
|
||||||
|
# in docs/movement_campaign.md. Every later job in the movement campaign must
|
||||||
|
# measure against THIS list: the unit of evidence is the number of opponents, so
|
||||||
|
# changing the panel changes what "better movement" means.
|
||||||
|
#
|
||||||
|
# Format (one opponent per line):
|
||||||
|
# <bot dir or name> | <style group> | <why it is in the panel>
|
||||||
|
# A bare `Name` is expanded to /tmp/tr_bots/Name. Blank lines and #-comments are
|
||||||
|
# ignored.
|
||||||
|
#
|
||||||
|
# `style` is INFERRED (from the bot's name / docs in
|
||||||
|
# tools/robocode_shim/robots.json, not from decompilation — robots.json says so
|
||||||
|
# explicitly). The style group is used ONLY to explain a result, never to decide
|
||||||
|
# one. The bot dirs come from tools/robocode_shim/robots.json (28 battle-validated
|
||||||
|
# legacy champions, see LEGACY_BOTS.md) except SpinBot, which is the tank-royale
|
||||||
|
# sample bot shipped with the project.
|
||||||
|
#
|
||||||
|
# Panel design (15 opponents, one axis each):
|
||||||
|
# * numbers matter, style does not: 5-6 opponents per family so a single
|
||||||
|
# weird bot can never carry an arm's mean;
|
||||||
|
# * the anchors (DrussGT/Diamond/Dookious) are the only opponents with older
|
||||||
|
# movement data, so they keep the new numbers comparable to the old ones;
|
||||||
|
# * every opponent must be able to DAMAGE us (WORKS in robots.json, not
|
||||||
|
# WORKS_WEAK): an opponent that lands 0 hits cannot reveal a movement
|
||||||
|
# regression, and an opponent that cannot dodge cannot reveal a range win.
|
||||||
|
#
|
||||||
|
# dodger — wave-surfing / adaptive evasion: punish our bullets AND our
|
||||||
|
# movement at the same time (the hard axis)
|
||||||
|
# pattern — pattern-matching guns: punish PREDICTABLE movement, which is
|
||||||
|
# exactly what a pure perpendicular strafe is
|
||||||
|
# wallfollower — wall-following / wall-avoiding blockers: punish movers that
|
||||||
|
# do not manage walls and let a corner-camper's corner matter
|
||||||
|
# cornercamper — corner camper / low bot
|
||||||
|
# spinner — periodic circle mover with a head-on gun: the classic
|
||||||
|
# "did you do anything obviously stupid" control
|
||||||
|
# rammer — closes to point blank: punishes a mover whose range
|
||||||
|
# preference cannot be enforced (and rewards damage output)
|
||||||
|
# brawler — micro/close-range brawler: same axis as the rammer, different
|
||||||
|
# gun quality
|
||||||
|
# aggressive — strong aggressive 1v1 megas: the best all-round stress test
|
||||||
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
# ── dodger (5) ───────────────────────────────────────────────────────────────
|
||||||
|
/tmp/tr_bots/DrussGT | dodger | THE anchor. Our only opponent with prior data (docs/surfer_wiring_ab.md, j107, j113) — keeps new results comparable to the old ones.
|
||||||
|
/tmp/tr_bots/Diamond | dodger | Voidious top-tier: pattern-matching gun + wave-surfing movement.
|
||||||
|
/tmp/tr_bots/Dookious | dodger | Voidious: wave-surfing movement + pattern/GF gun.
|
||||||
|
/tmp/tr_bots/GresSuffurd | dodger | Surf gun + surf movement; the most "modern" of the surfers here.
|
||||||
|
/tmp/tr_bots/CassiusClay | dodger | Pattern-matching gun + surf movement (different gun family from Diamond).
|
||||||
|
|
||||||
|
# ── pattern (3) ──────────────────────────────────────────────────────────────
|
||||||
|
/tmp/tr_bots/RetroGirl | pattern | Perceptual pattern matcher — the strongest pattern gun in the panel (26 hits in 2 rounds in the validation).
|
||||||
|
/tmp/tr_bots/TripHammer | pattern | Pattern matcher (19 hits in 2 rounds — strong).
|
||||||
|
/tmp/tr_bots/Coriantumr | pattern | Mini pattern matcher: small/simple bots stress the same axis at low complexity.
|
||||||
|
|
||||||
|
# ── wall follower / corner camper (2) ────────────────────────────────────────
|
||||||
|
/tmp/tr_bots/WallAvoider | wallfollower | Micro wall-avoider, classic blocking-Robot style: the "wall management" control.
|
||||||
|
/tmp/tr_bots/HawkOnFire | cornercamper | Corner camper / low bot: punishes movers that let the enemy own a corner.
|
||||||
|
|
||||||
|
# ── spinner (1, NOT from robots.json — see the note) ─────────────────────────
|
||||||
|
/home/davide/Projects/tank-royale/sample-bots/java/build/archive/SpinBot | spinner | The tank-royale sample SpinBot: the only true periodic circle-mover available (there is NO spinner among the 28 legacy champions). It is the same bot every legacy validation used, so it is a known quantity; it is the control for "obviously stupid" movement, not a hard test.
|
||||||
|
|
||||||
|
# ── rammer / brawler (2) ─────────────────────────────────────────────────────
|
||||||
|
/tmp/tr_bots/DiamondStealer | rammer | Rammer / close-range brawler: the only pure rammer validated (12 hits at point blank in the validation).
|
||||||
|
/tmp/tr_bots/BlitzBat | brawler | Micro brawler: close-range aggression with a weaker gun.
|
||||||
|
|
||||||
|
# ── aggressive mega (2) ──────────────────────────────────────────────────────
|
||||||
|
/tmp/tr_bots/YersiniaPestis | aggressive | Strong aggressive 1v1 mega (never defeated in the 2-round validation) — the generalist stress test.
|
||||||
|
/tmp/tr_bots/Ascendant | aggressive | Aggressive 1v1 mega on a team-derived pattern base — the second generalist.
|
||||||
@@ -0,0 +1,639 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""tournament_analyze.py — paired multi-opponent ranking of movement arms.
|
||||||
|
|
||||||
|
python3 tools/ab/tournament_analyze.py <session_dir> [--reference ARM]
|
||||||
|
|
||||||
|
Reads a session produced by `tournament_run.sh` (layout:
|
||||||
|
<outdir>/<opponent>/<arm>/run<N>.{battle.log,events.jsonl,jsonl,bot.stdout.log})
|
||||||
|
and answers the only question the movement campaign asks:
|
||||||
|
|
||||||
|
DOES ARM A MOVE BETTER THAN THE REFERENCE ARM, ACROSS OPPONENTS?
|
||||||
|
|
||||||
|
Method, in one paragraph. Every arm fights the SAME panel with the SAME frozen
|
||||||
|
binary; for each opponent the arm's metric is averaged over its runs and
|
||||||
|
subtracted from the reference arm's average for that same opponent. Those
|
||||||
|
per-opponent deltas are the unit of evidence: the MEAN of the deltas is the
|
||||||
|
effect, the SPREAD of the deltas across opponents is the honest error bar (one
|
||||||
|
weird opponent cannot carry it), and a SIGN TEST over the deltas says how many
|
||||||
|
opponents the arm actually wins. A pooled number over all runs is also printed,
|
||||||
|
but it is reported as the descriptive dashboard, never as the verdict.
|
||||||
|
|
||||||
|
CLI judgment (pre-registered in docs/movement_campaign.md): the primary metrics
|
||||||
|
are damage/run and ROUND WINS. An arm is BETTER than the reference only if one
|
||||||
|
of the two improves with the sign test at p<0.05 while the other does not
|
||||||
|
degrade; hit rate and distance are explanation, never the verdict. MDE
|
||||||
|
(alpha=0.05 two-sided, 80% power) is printed for every test so a null can be
|
||||||
|
told apart from an under-powered null.
|
||||||
|
|
||||||
|
Standard library only. Deterministic: the exact sign-flip test is enumerated
|
||||||
|
when n_opponents <= 20, otherwise sampled with a fixed seed.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import os
|
||||||
|
import random
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
|
||||||
|
BOT_NAME = "ModularBot"
|
||||||
|
# The metric keys, in report order.
|
||||||
|
METRICS = ["damage", "damage_taken", "wins", "hit_rate", "dist"]
|
||||||
|
# z_{0.975} + z_{0.80}: the constant in MDE = C * sd * sqrt(2/n) is for a
|
||||||
|
# two-SAMPLE design; for the paired per-opponent design used here the analogous
|
||||||
|
# constant with n = number of opponents is C * sd(deltas) / sqrt(n).
|
||||||
|
MDE_C = 1.959963984540054 + 0.8416212335729143
|
||||||
|
EXACT_SIGNCAP = 20 # 2^20 = 1M sign vectors is still instant
|
||||||
|
MC_DRAWS = 200_000
|
||||||
|
MC_SEED = 0x5EED5EED
|
||||||
|
T975 = {1: 12.706, 2: 4.303, 3: 3.182, 4: 2.776, 5: 2.571, 6: 2.447,
|
||||||
|
7: 2.365, 8: 2.306, 9: 2.262, 10: 2.228, 11: 2.201, 12: 2.179,
|
||||||
|
13: 2.160, 14: 2.145, 15: 2.131, 16: 2.120, 17: 2.110, 18: 2.101,
|
||||||
|
19: 2.093, 20: 2.086, 21: 2.080, 22: 2.074, 23: 2.069, 24: 2.064,
|
||||||
|
25: 2.060, 26: 2.056, 27: 2.052, 28: 2.048, 29: 2.045, 30: 2.042}
|
||||||
|
|
||||||
|
|
||||||
|
# ── small statistics helpers ─────────────────────────────────────────────────
|
||||||
|
def mean(xs):
|
||||||
|
return sum(xs) / len(xs) if xs else float("nan")
|
||||||
|
|
||||||
|
|
||||||
|
def sd(xs):
|
||||||
|
"""Sample standard deviation (n-1). 0.0 for n<2."""
|
||||||
|
n = len(xs)
|
||||||
|
if n < 2:
|
||||||
|
return 0.0
|
||||||
|
m = mean(xs)
|
||||||
|
return math.sqrt(sum((x - m) ** 2 for x in xs) / (n - 1))
|
||||||
|
|
||||||
|
|
||||||
|
def median(xs):
|
||||||
|
s = sorted(xs)
|
||||||
|
n = len(s)
|
||||||
|
if n == 0:
|
||||||
|
return float("nan")
|
||||||
|
return s[n // 2] if n % 2 else 0.5 * (s[n // 2 - 1] + s[n // 2])
|
||||||
|
|
||||||
|
|
||||||
|
def binom_two_sided(k, n):
|
||||||
|
"""Exact two-sided sign-test p-value (p=0.5), ties already removed."""
|
||||||
|
if n == 0:
|
||||||
|
return 1.0
|
||||||
|
def c(nn, kk):
|
||||||
|
return math.comb(nn, kk)
|
||||||
|
tail = sum(c(n, i) for i in range(0, min(k, n - k) + 1)) / 2 ** n
|
||||||
|
return min(1.0, 2.0 * tail)
|
||||||
|
|
||||||
|
|
||||||
|
def signflip_p(deltas):
|
||||||
|
"""Two-sided sign-flip permutation test on the MEAN of the deltas.
|
||||||
|
|
||||||
|
Exact (all 2^n sign vectors) for n <= EXACT_SIGNCAP; otherwise a
|
||||||
|
deterministic Monte-Carlo draw. Returns (p, method_string)."""
|
||||||
|
n = len(deltas)
|
||||||
|
if n == 0:
|
||||||
|
return 1.0, "n/a"
|
||||||
|
obs = abs(mean(deltas))
|
||||||
|
if obs == 0.0:
|
||||||
|
return 1.0, "exact (degenerate)"
|
||||||
|
tol = 1e-12
|
||||||
|
if n <= EXACT_SIGNCAP:
|
||||||
|
total = 1 << n
|
||||||
|
hits = 0
|
||||||
|
for mask in range(total):
|
||||||
|
s = 0.0
|
||||||
|
for i, d in enumerate(deltas):
|
||||||
|
s += -d if (mask >> i) & 1 else d
|
||||||
|
if abs(s / n) >= obs - tol:
|
||||||
|
hits += 1
|
||||||
|
return hits / total, f"exact 2^{n}"
|
||||||
|
rng = random.Random(MC_SEED)
|
||||||
|
hits = 0
|
||||||
|
for _ in range(MC_DRAWS):
|
||||||
|
s = 0.0
|
||||||
|
for d in deltas:
|
||||||
|
s += -d if rng.getrandbits(1) else d
|
||||||
|
if abs(s / n) >= obs - tol:
|
||||||
|
hits += 1
|
||||||
|
return hits / MC_DRAWS, f"MC {MC_DRAWS}"
|
||||||
|
|
||||||
|
|
||||||
|
def wilcoxon_p(deltas):
|
||||||
|
"""Two-sided Wilcoxon signed-rank, normal approximation with tie
|
||||||
|
correction. Returns (p, W). p=1.0 when there is nothing to test."""
|
||||||
|
nz = [d for d in deltas if d != 0.0]
|
||||||
|
n = len(nz)
|
||||||
|
if n < 3:
|
||||||
|
return 1.0, float("nan")
|
||||||
|
order = sorted(range(n), key=lambda i: abs(nz[i]))
|
||||||
|
ranks = [0.0] * n
|
||||||
|
i = 0
|
||||||
|
while i < n:
|
||||||
|
j = i
|
||||||
|
while j + 1 < n and abs(nz[order[j + 1]]) == abs(nz[order[i]]):
|
||||||
|
j += 1
|
||||||
|
avg = (i + j) / 2.0 + 1.0
|
||||||
|
for k in range(i, j + 1):
|
||||||
|
ranks[order[k]] = avg
|
||||||
|
i = j + 1
|
||||||
|
w_plus = sum(ranks[i] for i in range(n) if nz[i] > 0)
|
||||||
|
mu = n * (n + 1) / 4.0
|
||||||
|
# tie correction for sigma
|
||||||
|
from collections import Counter
|
||||||
|
cnt = Counter(abs(d) for d in nz)
|
||||||
|
tie = sum(c ** 3 - c for c in cnt.values())
|
||||||
|
sigma2 = n * (n + 1) * (2 * n + 1) / 24.0 - tie / 48.0
|
||||||
|
if sigma2 <= 0:
|
||||||
|
return 1.0, w_plus
|
||||||
|
z = (w_plus - mu - 0.5 * (1 if w_plus > mu else -1)) / math.sqrt(sigma2)
|
||||||
|
p = 2.0 * 0.5 * math.erfc(abs(z) / math.sqrt(2.0))
|
||||||
|
return min(1.0, p), w_plus
|
||||||
|
|
||||||
|
|
||||||
|
def t_crit(df):
|
||||||
|
return T975.get(df, 1.96)
|
||||||
|
|
||||||
|
|
||||||
|
def mde(deltas):
|
||||||
|
"""Minimum detectable effect for the paired per-opponent design at
|
||||||
|
alpha=0.05 (two-sided), power 80%, from the observed spread of the deltas."""
|
||||||
|
n = len(deltas)
|
||||||
|
if n < 2:
|
||||||
|
return float("nan")
|
||||||
|
return MDE_C * sd(deltas) / math.sqrt(n)
|
||||||
|
|
||||||
|
|
||||||
|
# ── parsing ──────────────────────────────────────────────────────────────────
|
||||||
|
COUNTERS_RE = re.compile(
|
||||||
|
r"subject event counts: scans=(\d+) bulletsFired=(\d+) bulletHits=(\d+)"
|
||||||
|
r" bulletMisses=(\d+) bulletHitBullets=(\d+) hitsTaken=(\d+)")
|
||||||
|
FIRSTPLACES_RE = re.compile(
|
||||||
|
r"^\s*#\d+\s+(\S+)\s+totalScore=(-?\d+)\s+firstPlaces=(\d+)\s+survival=(\d+)",
|
||||||
|
re.MULTILINE)
|
||||||
|
DIST_RE = re.compile(r"DISTANCE: mean=([\d.]+)")
|
||||||
|
ROWS_RE = re.compile(r"rows=(\d+)")
|
||||||
|
|
||||||
|
|
||||||
|
def read_text(path):
|
||||||
|
try:
|
||||||
|
with open(path, "r", errors="replace") as fh:
|
||||||
|
return fh.read()
|
||||||
|
except OSError:
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def parse_events(path):
|
||||||
|
out = []
|
||||||
|
try:
|
||||||
|
with open(path, "r", errors="replace") as fh:
|
||||||
|
for line in fh:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
out.append(json.loads(line))
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue # a partial line from a killed battle
|
||||||
|
except OSError:
|
||||||
|
out = []
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def parse_counters(text):
|
||||||
|
m = COUNTERS_RE.search(text)
|
||||||
|
if not m:
|
||||||
|
return None
|
||||||
|
return {"scans": int(m.group(1)), "fired": int(m.group(2)),
|
||||||
|
"hits": int(m.group(3)), "misses": int(m.group(4)),
|
||||||
|
"hit_bullets": int(m.group(5)), "hits_taken": int(m.group(6))}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_first_places(text):
|
||||||
|
for m in FIRSTPLACES_RE.finditer(text):
|
||||||
|
if m.group(1) == BOT_NAME:
|
||||||
|
return int(m.group(3))
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def parse_rounds(path):
|
||||||
|
try:
|
||||||
|
with open(path) as fh:
|
||||||
|
return len(json.load(fh).get("rounds", []))
|
||||||
|
except (OSError, json.JSONDecodeError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def attribute_subject(evs, counters):
|
||||||
|
"""Which event `owner` id is our bot? Matched on fire/hit/hits-taken counts,
|
||||||
|
never on fired power (the power policy is continuous).
|
||||||
|
|
||||||
|
Returns (subject_id, other_id, how). `subject_id` is None only when nothing
|
||||||
|
could be identified; `other_id` is None when the opponent never fired a
|
||||||
|
single bullet (then the incoming hit rate is simply undefined, and the run
|
||||||
|
is still a valid damage measurement)."""
|
||||||
|
fires, hits, victim_hits = {}, {}, {}
|
||||||
|
owner_ids = set()
|
||||||
|
for ev in evs:
|
||||||
|
t = ev.get("type")
|
||||||
|
o = ev.get("owner")
|
||||||
|
if o is not None:
|
||||||
|
owner_ids.add(o)
|
||||||
|
if t == "fire":
|
||||||
|
fires[o] = fires.get(o, 0) + 1
|
||||||
|
elif t == "hit":
|
||||||
|
hits[o] = hits.get(o, 0) + 1
|
||||||
|
v = ev.get("victim")
|
||||||
|
if v is not None:
|
||||||
|
victim_hits[v] = victim_hits.get(v, 0) + 1
|
||||||
|
if not fires:
|
||||||
|
return None, None, "no fire events"
|
||||||
|
|
||||||
|
def other_of(subj):
|
||||||
|
others = [o for o in owner_ids if o != subj]
|
||||||
|
return others[0] if len(others) == 1 else None
|
||||||
|
|
||||||
|
if counters:
|
||||||
|
strict = [o for o in fires
|
||||||
|
if fires[o] == counters["fired"]
|
||||||
|
and hits.get(o, 0) == counters["hits"]
|
||||||
|
and victim_hits.get(o, 0) == counters["hits_taken"]]
|
||||||
|
if len(strict) == 1:
|
||||||
|
return strict[0], other_of(strict[0]), "exact"
|
||||||
|
cand = [o for o in fires if counters and fires[o] == counters["fired"]]
|
||||||
|
if len(cand) != 1:
|
||||||
|
cand = [o for o in fires
|
||||||
|
if counters and victim_hits.get(o, 0) == counters["hits_taken"]]
|
||||||
|
if len(cand) != 1:
|
||||||
|
cand = list(fires)
|
||||||
|
if len(cand) == 1:
|
||||||
|
return cand[0], other_of(cand[0]), "fires-only"
|
||||||
|
return None, None, "ambiguous"
|
||||||
|
|
||||||
|
|
||||||
|
def liveness(arm_env_tokens, stdout_path):
|
||||||
|
"""Liveness: every declared env token must appear verbatim in OUR bot's own
|
||||||
|
boot environment report, so an arm whose setting never reached the process
|
||||||
|
is a loud FAIL instead of a plausible-looking number. A TR_MOVEMENT the arm
|
||||||
|
did NOT declare is fatal too (the baseline would be contaminated)."""
|
||||||
|
text = read_text(stdout_path)
|
||||||
|
if not text:
|
||||||
|
return False, "no boot env report", False
|
||||||
|
missing = [t for t in arm_env_tokens if t not in text]
|
||||||
|
declared_keys = {t.split("=", 1)[0] for t in arm_env_tokens}
|
||||||
|
leaked_movement = ("TR_MOVEMENT" not in declared_keys
|
||||||
|
and re.search(r"^\[env\] TR_MOVEMENT=", text, re.MULTILINE)
|
||||||
|
is not None)
|
||||||
|
ok = (not missing) and (not leaked_movement)
|
||||||
|
why = []
|
||||||
|
if missing:
|
||||||
|
why.append("not seen in bot env report: " + ", ".join(missing))
|
||||||
|
if leaked_movement:
|
||||||
|
why.append("undeclared TR_MOVEMENT leaked into the process")
|
||||||
|
return ok, ("; ".join(why) if why else "ok"), bool(leaked_movement)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_run(opp_dir, arm_dir, run):
|
||||||
|
"""One run -> dict of MEASURED numbers, or ok=False with a reason."""
|
||||||
|
base = os.path.join(opp_dir, arm_dir)
|
||||||
|
log_path = os.path.join(base, f"run{run}.battle.log")
|
||||||
|
ev_path = os.path.join(base, f"run{run}.events.jsonl")
|
||||||
|
rounds_path = os.path.join(base, f"run{run}.jsonl.rounds.json")
|
||||||
|
log = read_text(log_path)
|
||||||
|
out = {"run": run, "ok": False, "reason": "no capture",
|
||||||
|
"damage": 0.0, "damage_taken": 0.0, "wins": None, "rounds": None,
|
||||||
|
"opp_fired": 0, "opp_hits": 0, "dist": None, "scans": 0,
|
||||||
|
"liveness": "n/a"}
|
||||||
|
if not log:
|
||||||
|
return out
|
||||||
|
counters = parse_counters(log)
|
||||||
|
if counters is None:
|
||||||
|
out["reason"] = "battle never started (no subject counters)"
|
||||||
|
return out
|
||||||
|
evs = parse_events(ev_path)
|
||||||
|
subj, other, how = attribute_subject(evs, counters)
|
||||||
|
if subj is None:
|
||||||
|
out["reason"] = f"owner attribution failed ({how})"
|
||||||
|
return out
|
||||||
|
dmg = dmg_taken = 0.0
|
||||||
|
opp_fired = 0
|
||||||
|
for ev in evs:
|
||||||
|
t = ev.get("type")
|
||||||
|
if t == "fire" and ev.get("owner") == other:
|
||||||
|
opp_fired += 1
|
||||||
|
elif t == "hit":
|
||||||
|
d = float(ev.get("damage", 0.0))
|
||||||
|
if ev.get("owner") == subj:
|
||||||
|
dmg += d
|
||||||
|
elif ev.get("owner") == other:
|
||||||
|
dmg_taken += d
|
||||||
|
opp_hits = sum(1 for ev in evs
|
||||||
|
if ev.get("type") == "hit" and ev.get("owner") == other)
|
||||||
|
wins = parse_first_places(log)
|
||||||
|
if wins is None:
|
||||||
|
out["reason"] = "no final standings in the capture"
|
||||||
|
return out
|
||||||
|
m = DIST_RE.search(log)
|
||||||
|
out.update({
|
||||||
|
"ok": True, "reason": "ok",
|
||||||
|
"counters": counters, "subject_id": subj, "owner_attribution": how,
|
||||||
|
"damage": dmg, "damage_taken": dmg_taken,
|
||||||
|
"wins": wins, "rounds": parse_rounds(rounds_path),
|
||||||
|
"opp_fired": opp_fired, "opp_hits": opp_hits,
|
||||||
|
"dist": float(m.group(1)) if m else None,
|
||||||
|
"scans": counters["scans"],
|
||||||
|
})
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
# ── arm/opponent aggregation ─────────────────────────────────────────────────
|
||||||
|
def arm_metrics(runs):
|
||||||
|
"""Aggregate a list of valid run dicts into one metric dict."""
|
||||||
|
n = len(runs)
|
||||||
|
total_rounds = sum(r["rounds"] or 0 for r in runs)
|
||||||
|
opp_fired = sum(r["opp_fired"] for r in runs)
|
||||||
|
opp_hits = sum(r["opp_hits"] for r in runs)
|
||||||
|
dists = [r["dist"] for r in runs if r["dist"] is not None]
|
||||||
|
return {
|
||||||
|
"runs": n,
|
||||||
|
"damage": mean([r["damage"] for r in runs]),
|
||||||
|
"damage_taken": mean([r["damage_taken"] for r in runs]),
|
||||||
|
"wins": mean([r["wins"] for r in runs]),
|
||||||
|
"win_rate": (sum(r["wins"] for r in runs) / total_rounds
|
||||||
|
if total_rounds else float("nan")),
|
||||||
|
"rounds": total_rounds,
|
||||||
|
"hit_rate": (100.0 * opp_hits / opp_fired) if opp_fired else float("nan"),
|
||||||
|
"dist": mean(dists) if dists else float("nan"),
|
||||||
|
"scans": mean([r["scans"] for r in runs]),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def collect(session):
|
||||||
|
"""-> (data, notes); data[opp][arm] = {"valid": [...], "invalid": [...]}"""
|
||||||
|
data = {}
|
||||||
|
notes = []
|
||||||
|
for opp in session["opponents"]:
|
||||||
|
oname = opp["name"]
|
||||||
|
data[oname] = {}
|
||||||
|
opp_dir = os.path.join(session["outdir"], oname)
|
||||||
|
for arm in session["arms"]:
|
||||||
|
aname = arm["name"]
|
||||||
|
tokens = arm["env"].split() if arm["env"] else []
|
||||||
|
node = {"valid": [], "invalid": [], "tokens": tokens,
|
||||||
|
"style": opp.get("style", ""), "label": arm.get("label", "")}
|
||||||
|
for r in range(1, session["runs"] + 1):
|
||||||
|
rr = parse_run(opp_dir, aname, r)
|
||||||
|
live_ok, live_why, fatal_leak = liveness(
|
||||||
|
tokens, os.path.join(opp_dir, aname, f"run{r}.bot.stdout.log"))
|
||||||
|
rr["liveness"] = live_why
|
||||||
|
if not rr["ok"]:
|
||||||
|
node["invalid"].append(rr)
|
||||||
|
elif not live_ok:
|
||||||
|
rr["reason"] = "liveness FAIL: " + live_why
|
||||||
|
node["invalid"].append(rr)
|
||||||
|
else:
|
||||||
|
node["valid"].append(rr)
|
||||||
|
data[oname][aname] = node
|
||||||
|
return data, notes
|
||||||
|
|
||||||
|
|
||||||
|
# ── the report ───────────────────────────────────────────────────────────────
|
||||||
|
def fmt(x, nd=1):
|
||||||
|
return "n/a" if x is None or (isinstance(x, float) and math.isnan(x)) \
|
||||||
|
else f"{x:.{nd}f}"
|
||||||
|
|
||||||
|
|
||||||
|
def fmt_s(x, nd=1):
|
||||||
|
return "n/a" if x is None or (isinstance(x, float) and math.isnan(x)) \
|
||||||
|
else f"{x:+.{nd}f}"
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
args = sys.argv[1:]
|
||||||
|
if not args:
|
||||||
|
print(__doc__)
|
||||||
|
return 2
|
||||||
|
session_dir = args[0]
|
||||||
|
ref = None
|
||||||
|
if "--reference" in args:
|
||||||
|
ref = args[args.index("--reference") + 1]
|
||||||
|
try:
|
||||||
|
with open(os.path.join(session_dir, "session.json")) as fh:
|
||||||
|
session = json.load(fh)
|
||||||
|
except (OSError, json.JSONDecodeError) as exc:
|
||||||
|
print(f"ERROR: cannot read {session_dir}/session.json: {exc}", file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
session["outdir"] = session_dir
|
||||||
|
arms = [a["name"] for a in session["arms"]]
|
||||||
|
if ref is None:
|
||||||
|
ref = session.get("reference") or arms[0]
|
||||||
|
if ref not in arms:
|
||||||
|
print(f"ERROR: reference arm '{ref}' not in session arms {arms}",
|
||||||
|
file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
opps = [o["name"] for o in session["opponents"]]
|
||||||
|
style_of = {o["name"]: o.get("style", "") for o in session["opponents"]}
|
||||||
|
data, _ = collect(session)
|
||||||
|
|
||||||
|
out = []
|
||||||
|
def p(s=""):
|
||||||
|
out.append(s)
|
||||||
|
print(s)
|
||||||
|
|
||||||
|
p("### MEASURED: session")
|
||||||
|
p()
|
||||||
|
p(f"* commit `{session['commit']}`, frozen binary sha256 `{session['binary_sha256'][:12]}…`")
|
||||||
|
p(f"* {len(opps)} opponents × {len(arms)} arms × {session['runs']} runs × "
|
||||||
|
f"{session['rounds']} rounds = {len(opps) * len(arms) * session['runs']} battles, "
|
||||||
|
f"conc={session.get('conc', '?')}")
|
||||||
|
p(f"* arms file `{os.path.basename(session.get('arms_file', '?'))}`, "
|
||||||
|
f"panel file `{os.path.basename(session.get('panel_file', '?'))}`")
|
||||||
|
p(f"* reference arm: **`{ref}`** — every delta below is (arm − {ref}), "
|
||||||
|
f"opponent by opponent")
|
||||||
|
p()
|
||||||
|
invalid = [(o, a, r) for o in opps for a in arms for r in data[o][a]["invalid"]]
|
||||||
|
p(f"* liveness: {len(invalid)} run(s) excluded "
|
||||||
|
f"({len(opps) * len(arms) * session['runs']} total)")
|
||||||
|
for o, a, r in invalid[:20]:
|
||||||
|
p(f" * `{o}/{a}` run{r['run']}: {r['reason']}")
|
||||||
|
if len(invalid) > 20:
|
||||||
|
p(f" * … and {len(invalid) - 20} more")
|
||||||
|
p()
|
||||||
|
|
||||||
|
# ── per-opponent × per-arm deltas ───────────────────────────────────────
|
||||||
|
per_opp = {a: {} for a in arms}
|
||||||
|
for o in opps:
|
||||||
|
ref_runs = data[o][ref]["valid"]
|
||||||
|
ref_m = arm_metrics(ref_runs) if ref_runs else None
|
||||||
|
for a in arms:
|
||||||
|
m = arm_metrics(data[o][a]["valid"]) if data[o][a]["valid"] else None
|
||||||
|
if m is None or ref_m is None:
|
||||||
|
per_opp[a][o] = None
|
||||||
|
continue
|
||||||
|
per_opp[a][o] = {
|
||||||
|
"m": m, "ref": ref_m,
|
||||||
|
"d_damage": m["damage"] - ref_m["damage"],
|
||||||
|
"d_wins": m["wins"] - ref_m["wins"],
|
||||||
|
"d_damage_taken": m["damage_taken"] - ref_m["damage_taken"],
|
||||||
|
"d_hit_rate": m["hit_rate"] - ref_m["hit_rate"],
|
||||||
|
"d_dist": m["dist"] - ref_m["dist"],
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── per-arm tables ──────────────────────────────────────────────────────
|
||||||
|
p("### MEASURED: per-opponent paired table (per arm)")
|
||||||
|
p()
|
||||||
|
for a in arms:
|
||||||
|
lbl = next((x.get("label") for x in session["arms"] if x["name"] == a), "")
|
||||||
|
cnt = sum(1 for o in opps if per_opp[a][o])
|
||||||
|
p(f"#### `{a}`" + (f" — {lbl}" if lbl else "") + f" (paired on {cnt} opponents)")
|
||||||
|
p()
|
||||||
|
p("| opponent | style | dmg/run ref→arm | Δdmg | wins/run ref→arm | Δwins | Δdmg taken | Δhit rate (pp) | dist ref→arm |")
|
||||||
|
p("|---|---|---:|---:|---:|---:|---:|---:|---:|")
|
||||||
|
for o in opps:
|
||||||
|
e = per_opp[a][o]
|
||||||
|
if e is None:
|
||||||
|
p(f"| {o} | {style_of[o]} | n/a | n/a | n/a | n/a | n/a | n/a | n/a |")
|
||||||
|
continue
|
||||||
|
m, rm = e["m"], e["ref"]
|
||||||
|
p(f"| {o} | {style_of[o]} | {fmt(rm['damage'])}→{fmt(m['damage'])} "
|
||||||
|
f"| {fmt_s(e['d_damage'])} "
|
||||||
|
f"| {fmt(rm['wins'], 2)}→{fmt(m['wins'], 2)} | {fmt_s(e['d_wins'], 2)} "
|
||||||
|
f"| {fmt_s(e['d_damage_taken'])} | {fmt_s(e['d_hit_rate'], 2)} "
|
||||||
|
f"| {fmt(rm['dist'], 0)}→{fmt(m['dist'], 0)} |")
|
||||||
|
p()
|
||||||
|
|
||||||
|
# ── aggregate dashboard, pooled over every valid run ────────────────────
|
||||||
|
p("### MEASURED: pooled dashboard (all valid runs, NOT the verdict)")
|
||||||
|
p()
|
||||||
|
p("| arm | runs | dmg/run | dmg taken/run | wins/run | round wins | win rate | incoming hit rate | mean distance |")
|
||||||
|
p("|---|---:|---:|---:|---:|---:|---:|---:|---:|")
|
||||||
|
pooled = {}
|
||||||
|
for a in arms:
|
||||||
|
runs = [r for o in opps for r in data[o][a]["valid"]]
|
||||||
|
if not runs:
|
||||||
|
p(f"| `{a}` | 0 | n/a | n/a | n/a | n/a | n/a | n/a | n/a |")
|
||||||
|
continue
|
||||||
|
m = arm_metrics(runs)
|
||||||
|
pooled[a] = m
|
||||||
|
p(f"| `{a}` | {m['runs']} | {fmt(m['damage'])} | {fmt(m['damage_taken'])} "
|
||||||
|
f"| {fmt(m['wins'], 2)} | {int(sum(r['wins'] for r in runs))}/{m['rounds']} "
|
||||||
|
f"| {fmt(100 * m['win_rate'])}% | {fmt(m['hit_rate'], 2)}% "
|
||||||
|
f"| {fmt(m['dist'], 0)} |")
|
||||||
|
p()
|
||||||
|
|
||||||
|
# ── cross-opponent aggregation: mean delta, spread, sign tests, MDE ─────
|
||||||
|
p("### MEASURED: cross-opponent aggregation (the verdict layer)")
|
||||||
|
p()
|
||||||
|
p("Deltas are per-opponent (arm − reference). `spread` is the SD of those "
|
||||||
|
"deltas ACROSS opponents; `SE` = spread/√n; `95% CI` = mean ± t·SE. "
|
||||||
|
"Sign test = how many opponents the arm wins (ties dropped), exact "
|
||||||
|
"binomial; sign-flip = permutation test on the mean of the deltas.")
|
||||||
|
p()
|
||||||
|
p("| arm | metric | mean Δ | spread (SD) | SE | 95% CI | sign test (wins/n) | p(sign) | p(sign-flip) | Wilcoxon p | MDE |")
|
||||||
|
p("|---|---|---:|---:|---:|---|---:|---:|---:|---:|---:|")
|
||||||
|
stats = {}
|
||||||
|
for a in arms:
|
||||||
|
if a == ref:
|
||||||
|
continue
|
||||||
|
ds = [per_opp[a][o] for o in opps if per_opp[a][o]]
|
||||||
|
st = {"n": len(ds)}
|
||||||
|
for key, mkey in (("d_damage", "damage"), ("d_wins", "wins"),
|
||||||
|
("d_damage_taken", "damage_taken"),
|
||||||
|
("d_hit_rate", "hit_rate"), ("d_dist", "dist")):
|
||||||
|
vals = [e[key] for e in ds if not math.isnan(e[key])]
|
||||||
|
if len(vals) < 2:
|
||||||
|
continue
|
||||||
|
mu = mean(vals)
|
||||||
|
s = sd(vals)
|
||||||
|
se = s / math.sqrt(len(vals))
|
||||||
|
crit = t_crit(len(vals) - 1)
|
||||||
|
nz = [v for v in vals if v != 0.0]
|
||||||
|
wins_sign = sum(1 for v in nz if v > 0)
|
||||||
|
ps = binom_two_sided(wins_sign, len(nz))
|
||||||
|
pf, method = signflip_p(vals)
|
||||||
|
pw, _ = wilcoxon_p(vals)
|
||||||
|
st[key] = {"mean": mu, "sd": s, "se": se,
|
||||||
|
"ci": (mu - crit * se, mu + crit * se),
|
||||||
|
"k": wins_sign, "nz": len(nz), "p_sign": ps,
|
||||||
|
"p_flip": pf, "p_flip_method": method, "p_wilcox": pw,
|
||||||
|
"mde": mde(vals)}
|
||||||
|
p(f"| `{a}` | {mkey} | {fmt_s(mu, 2)} | {fmt(s, 2)} | {fmt(se, 2)} "
|
||||||
|
f"| [{fmt_s(mu - crit * se, 2)}, {fmt_s(mu + crit * se, 2)}] "
|
||||||
|
f"| {wins_sign}/{len(nz)} | {ps:.4g} | {pf:.4g} ({method}) "
|
||||||
|
f"| {pw:.4g} | {fmt(st[key]['mde'], 2)} |")
|
||||||
|
stats[a] = st
|
||||||
|
p()
|
||||||
|
|
||||||
|
# ── style breakdown (leg/governance only) ───────────────────────────────
|
||||||
|
styles = sorted({style_of[o] for o in opps if style_of[o]})
|
||||||
|
if len(styles) > 1:
|
||||||
|
p("#### By inferred style (explanation only, never the verdict)")
|
||||||
|
p()
|
||||||
|
p("| arm | style | n | mean Δdmg | mean Δwins | mean Δhit rate (pp) |")
|
||||||
|
p("|---|---|---:|---:|---:|---:|")
|
||||||
|
for a in arms:
|
||||||
|
if a == ref:
|
||||||
|
continue
|
||||||
|
for s in styles:
|
||||||
|
ds = [per_opp[a][o] for o in opps
|
||||||
|
if per_opp[a][o] and style_of[o] == s]
|
||||||
|
if not ds:
|
||||||
|
continue
|
||||||
|
p(f"| `{a}` | {s} | {len(ds)} "
|
||||||
|
f"| {fmt_s(mean([e['d_damage'] for e in ds]))} "
|
||||||
|
f"| {fmt_s(mean([e['d_wins'] for e in ds]), 2)} "
|
||||||
|
f"| {fmt_s(mean([e['d_hit_rate'] for e in ds]), 2)} |")
|
||||||
|
p()
|
||||||
|
|
||||||
|
# ── the pre-registered verdict rules ────────────────────────────────────
|
||||||
|
p("### The pre-registered verdict (rules fixed in `docs/movement_campaign.md`)")
|
||||||
|
p()
|
||||||
|
p("BETTER = one primary metric (dmg/run, wins/run) up at sign-test p<0.05 "
|
||||||
|
"with the other not down; WORSE = the mirror image; otherwise NOT "
|
||||||
|
"DISTINGUISHABLE. Hit rate is never the verdict.")
|
||||||
|
p()
|
||||||
|
ranked = []
|
||||||
|
for a in arms:
|
||||||
|
if a == ref:
|
||||||
|
continue
|
||||||
|
st = stats.get(a, {})
|
||||||
|
d, w = st.get("d_damage"), st.get("d_wins")
|
||||||
|
if d is None or w is None:
|
||||||
|
ranked.append((a, "n/a", float("nan"), float("nan")))
|
||||||
|
continue
|
||||||
|
up_d = d["mean"] > 0 and d["p_sign"] < 0.05
|
||||||
|
dn_d = d["mean"] < 0 and d["p_sign"] < 0.05
|
||||||
|
up_w = w["mean"] > 0 and w["p_sign"] < 0.05
|
||||||
|
dn_w = w["mean"] < 0 and w["p_sign"] < 0.05
|
||||||
|
if (up_d and w["mean"] >= 0) or (up_w and d["mean"] >= 0):
|
||||||
|
verdict = "BETTER than reference"
|
||||||
|
elif (dn_d and w["mean"] <= 0) or (dn_w and d["mean"] <= 0):
|
||||||
|
verdict = "WORSE than reference"
|
||||||
|
else:
|
||||||
|
verdict = "NOT DISTINGUISHABLE from reference"
|
||||||
|
ranked.append((a, verdict, w["mean"], d["mean"]))
|
||||||
|
ranked.sort(key=lambda t: (-(t[2] if t[2] == t[2] else -1e9),
|
||||||
|
-(t[3] if t[3] == t[3] else -1e9)))
|
||||||
|
p("| rank | arm | Δwins/run | Δdmg/run | sign test dmg | sign test wins | verdict |")
|
||||||
|
p("|---:|---|---:|---:|---|---|---|")
|
||||||
|
for i, (a, verdict, w, d) in enumerate(ranked, 1):
|
||||||
|
st = stats.get(a, {})
|
||||||
|
dd = st.get("d_damage", {})
|
||||||
|
ww = st.get("d_wins", {})
|
||||||
|
p(f"| {i} | `{a}` | {fmt_s(w, 2)} | {fmt_s(d)} "
|
||||||
|
f"| {dd.get('k', 'n/a')}/{dd.get('nz', 'n/a')} p={dd.get('p_sign', float('nan')):.4g} "
|
||||||
|
f"| {ww.get('k', 'n/a')}/{ww.get('nz', 'n/a')} p={ww.get('p_sign', float('nan')):.4g} "
|
||||||
|
f"| **{verdict}** |")
|
||||||
|
p()
|
||||||
|
p(f"Reference `{ref}`: {fmt(pooled.get(ref, {}).get('damage'))} dmg/run, "
|
||||||
|
f"{fmt(pooled.get(ref, {}).get('wins'), 2)} wins/run, "
|
||||||
|
f"{fmt(pooled.get(ref, {}).get('hit_rate'), 2)}% incoming, "
|
||||||
|
f"{fmt(pooled.get(ref, {}).get('dist'), 0)} px.")
|
||||||
|
best = ranked[0] if ranked else None
|
||||||
|
if best:
|
||||||
|
p()
|
||||||
|
p(f"Highest wins delta: `{best[0]}` ({fmt_s(best[2], 2)} wins/run, "
|
||||||
|
f"{fmt_s(best[3])} dmg/run) — **{best[1]}**.")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
Executable
+406
@@ -0,0 +1,406 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
|
# tournament_run.sh — THE MULTI-OPPONENT TOURNAMENT HARNESS of the movement
|
||||||
|
# campaign: play OUR bot (one FROZEN binary from `git archive HEAD`) against a
|
||||||
|
# FIXED PANEL of opponents, once per (opponent × arm × run), and leave a session
|
||||||
|
# dir that `tournament_analyze.py` turns into a paired per-opponent ranking.
|
||||||
|
#
|
||||||
|
# tools/ab/tournament_run.sh \
|
||||||
|
# --arms tools/ab/arms_movement_b1.txt \
|
||||||
|
# --panel tools/ab/panel_movement.txt \
|
||||||
|
# --runs 3 --rounds 3 --conc 6 --wait-arena 45 \
|
||||||
|
# --outdir /tmp/ab/j118_b1
|
||||||
|
#
|
||||||
|
# Why it exists next to ab_run.sh: ab_run.sh hard-wires exactly ONE adversary
|
||||||
|
# (DrussGT). Here the SUBJECT is our frozen bot and the ADVERSARY iterates over
|
||||||
|
# a frozen panel, so the unit of evidence becomes the NUMBER OF OPPONENTS.
|
||||||
|
# The arm file format is identical to ab_run.sh / gauntlet_run.sh.
|
||||||
|
#
|
||||||
|
# Design contract (all four are load-bearing for a valid measurement):
|
||||||
|
# 1. ONE binary for every arm. The binary is built once from `git archive
|
||||||
|
# HEAD` (a dirty tree cannot leak in) and every arm differs ONLY by the env
|
||||||
|
# dict it is launched with. No per-arm rebuild, ever.
|
||||||
|
# 2. Every run is isolated: own bot dir, own classic data dir, own output
|
||||||
|
# files, its own process group; the runner picks ephemeral ports.
|
||||||
|
# 3. A run that never really started is retried, not recorded. A battle that
|
||||||
|
# did not boot ("Only 1 of 2 bots started") would otherwise show up as a
|
||||||
|
# 0-0 phantom and poison the paired deltas.
|
||||||
|
# 4. ARENA SERIALIZATION. Other campaign jobs may be fighting at the same
|
||||||
|
# time; two concurrent battle fleets on one 16-core box distort each other.
|
||||||
|
# --wait-arena N polls for foreign bridge processes (bracketed pgrep so it
|
||||||
|
# cannot match itself) and WAITS until they are gone, up to N minutes.
|
||||||
|
# With N=0 the session REFUSES to start while the arena is busy.
|
||||||
|
#
|
||||||
|
# Arm file format: name | ENV=value ENV=value | optional label
|
||||||
|
# Panel file format: /tmp/tr_bots/Name | style | why (a bare Name works)
|
||||||
|
#
|
||||||
|
# Output layout (what tournament_analyze.py reads):
|
||||||
|
# <outdir>/session.json commit, sha, arms, panel
|
||||||
|
# <outdir>/<opponent>/<arm>/run<N>.jsonl tick capture
|
||||||
|
# <outdir>/<opponent>/<arm>/run<N>.jsonl.rounds.json
|
||||||
|
# <outdir>/<opponent>/<arm>/run<N>.jsonl.results.json
|
||||||
|
# <outdir>/<opponent>/<arm>/run<N>.events.jsonl fire/hit/death sidecar
|
||||||
|
# <outdir>/<opponent>/<arm>/run<N>.battle.log capture stdout + stats
|
||||||
|
# <outdir>/<opponent>/<arm>/run<N>.bot.stdout.log our bot's boot env report
|
||||||
|
#
|
||||||
|
# Cleanup on EXIT/INT/TERM: only THIS session's process groups, and only
|
||||||
|
# processes whose command line references THIS outdir. Never a broad pkill.
|
||||||
|
# ─────────────────────────────────────────────────────────────────────────────
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||||
|
RUN_BRIDGE="$REPO/tools/robocode_shim/run_bridge_battle.sh"
|
||||||
|
|
||||||
|
ARMS_FILE=""
|
||||||
|
PANEL_FILE=""
|
||||||
|
RUNS=3
|
||||||
|
ROUNDS=3
|
||||||
|
CONC=6
|
||||||
|
OUTDIR=""
|
||||||
|
WAIT_ARENA=0
|
||||||
|
FORCE_BUSY=0
|
||||||
|
REFERENCE=""
|
||||||
|
# A bot that fails to connect inside the booter's window is a FALSE NEGATIVE.
|
||||||
|
# Retry until the capture shows real rows AND the subject counter line.
|
||||||
|
MAX_ATTEMPTS="${TOURNAMENT_MAX_ATTEMPTS:-4}"
|
||||||
|
# 800x600 default arena; recorded for the ledger.
|
||||||
|
GAME_TYPE="${TOURNAMENT_GAME_TYPE:-1v1}"
|
||||||
|
|
||||||
|
usage() {
|
||||||
|
sed -n '2,45p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//'
|
||||||
|
exit "${1:-0}"
|
||||||
|
}
|
||||||
|
|
||||||
|
while [[ $# -gt 0 ]]; do
|
||||||
|
case "$1" in
|
||||||
|
--arms) ARMS_FILE="$2"; shift 2;;
|
||||||
|
--panel|--opponents) PANEL_FILE="$2"; shift 2;;
|
||||||
|
--runs) RUNS="$2"; shift 2;;
|
||||||
|
--rounds) ROUNDS="$2"; shift 2;;
|
||||||
|
--conc) CONC="$2"; shift 2;;
|
||||||
|
--outdir) OUTDIR="$2"; shift 2;;
|
||||||
|
--wait-arena) WAIT_ARENA="$2"; shift 2;;
|
||||||
|
--force-when-busy) FORCE_BUSY=1; shift;;
|
||||||
|
--reference) REFERENCE="$2"; shift 2;;
|
||||||
|
-h|--help) usage 0;;
|
||||||
|
*) echo "unknown argument: $1" >&2; usage 1;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
|
||||||
|
[[ -n "$ARMS_FILE" ]] || { echo "ERROR: --arms <file> is required" >&2; usage 1; }
|
||||||
|
[[ -n "$PANEL_FILE" ]] || { echo "ERROR: --panel <file> is required" >&2; usage 1; }
|
||||||
|
[[ -f "$ARMS_FILE" ]] || { echo "ERROR: arm file not found: $ARMS_FILE" >&2; exit 1; }
|
||||||
|
[[ -f "$PANEL_FILE" ]] || { echo "ERROR: panel file not found: $PANEL_FILE" >&2; exit 1; }
|
||||||
|
[[ -n "$OUTDIR" ]] || { echo "ERROR: --outdir <dir> is required" >&2; exit 1; }
|
||||||
|
[[ "$OUTDIR" != "/" && "$OUTDIR" != "" ]] || { echo "ERROR: refusing outdir '$OUTDIR'" >&2; exit 1; }
|
||||||
|
mkdir -p "$OUTDIR"
|
||||||
|
OUTDIR="$(cd "$OUTDIR" && pwd)"
|
||||||
|
# never nuke a directory that is not one of ours (non-empty without a session.json)
|
||||||
|
if [[ -n "$(ls -A "$OUTDIR" 2>/dev/null)" && ! -f "$OUTDIR/session.json" ]]; then
|
||||||
|
echo "ERROR: refusing to clear non-empty, non-session dir: $OUTDIR" >&2
|
||||||
|
echo " (delete it by hand or pick another --outdir)" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── prerequisites ────────────────────────────────────────────────────────────
|
||||||
|
ROBOCODE_JAR="${ROBOCODE_JAR:-/tmp/robocode/install/libs/robocode.jar}"
|
||||||
|
RUNNER_JAR="${TR_RUNNER_JAR:-/home/davide/Projects/tank-royale/runner/examples/lib/robocode-tankroyale-runner.jar}"
|
||||||
|
BOT_API_JAR="${TR_BOT_API_JAR:-$HOME/Downloads/sample-bots-java-1.0.2/lib/robocode-tankroyale-bot-api-1.0.2.jar}"
|
||||||
|
|
||||||
|
missing=0
|
||||||
|
for f in "$ROBOCODE_JAR" "$RUNNER_JAR" "$BOT_API_JAR"; do
|
||||||
|
[[ -f "$f" ]] || { echo "ERROR: missing prerequisite: $f" >&2; missing=1; }
|
||||||
|
done
|
||||||
|
command -v nim >/dev/null 2>&1 || { echo "ERROR: nim not on PATH" >&2; missing=1; }
|
||||||
|
command -v git >/dev/null 2>&1 || { echo "ERROR: git not on PATH" >&2; missing=1; }
|
||||||
|
[[ -f "$RUN_BRIDGE" ]] || { echo "ERROR: missing $RUN_BRIDGE" >&2; missing=1; }
|
||||||
|
(( missing == 0 )) || { echo "ERROR: prerequisites missing — not starting a broken session" >&2; exit 1; }
|
||||||
|
|
||||||
|
# ── ARENA SERIALIZATION ──────────────────────────────────────────────────────
|
||||||
|
# Bracketed patterns: pgrep -f must not match this script's own command line.
|
||||||
|
FOREIGN_PATTERN='[r]un_bridge_battle.sh|[r]obocode_shim.TrBattleCapture|[r]obocode_shim.LegacyBotBridge|[M]odularBot_bin'
|
||||||
|
foreign_battles() {
|
||||||
|
pgrep -af "$FOREIGN_PATTERN" 2>/dev/null | grep -v "$OUTDIR" || true
|
||||||
|
}
|
||||||
|
|
||||||
|
check_arena() {
|
||||||
|
local found
|
||||||
|
found="$(foreign_battles)"
|
||||||
|
[[ -z "$found" ]]
|
||||||
|
}
|
||||||
|
|
||||||
|
if ! check_arena; then
|
||||||
|
if (( FORCE_BUSY == 1 )); then
|
||||||
|
echo "[tournament] WARNING: arena busy but --force-when-busy was given:" >&2
|
||||||
|
foreign_battles | sed 's/^/ /' >&2
|
||||||
|
elif (( WAIT_ARENA > 0 )); then
|
||||||
|
echo "[tournament] arena busy — waiting up to ${WAIT_ARENA} min for foreign battles:"
|
||||||
|
foreign_battles | sed 's/^/ /'
|
||||||
|
deadline=$(( $(date +%s) + WAIT_ARENA * 60 ))
|
||||||
|
while :; do
|
||||||
|
sleep 120
|
||||||
|
if check_arena; then echo "[tournament] arena free."; break; fi
|
||||||
|
if (( $(date +%s) >= deadline )); then
|
||||||
|
echo "[tournament] ERROR: arena still busy after ${WAIT_ARENA} min; NOT starting." >&2
|
||||||
|
foreign_battles | sed 's/^/ /' >&2
|
||||||
|
exit 3
|
||||||
|
fi
|
||||||
|
echo "[tournament] still busy at $(date +%H:%M:%S); $(foreign_battles | wc -l) foreign process(es)"
|
||||||
|
done
|
||||||
|
else
|
||||||
|
echo "[tournament] ERROR: the arena is busy with another job's battles:" >&2
|
||||||
|
foreign_battles | sed 's/^/ /' >&2
|
||||||
|
echo "[tournament] rerun with --wait-arena <minutes> to wait for it, or --force-when-busy." >&2
|
||||||
|
exit 3
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── parse the arm file ───────────────────────────────────────────────────────
|
||||||
|
trim() { local s="$1"; s="${s#"${s%%[![:space:]]*}"}"; s="${s%"${s##*[![:space:]]}"}"; printf '%s' "$s"; }
|
||||||
|
|
||||||
|
ARM_NAMES=(); ARM_ENVS=(); ARM_LABELS=()
|
||||||
|
while IFS= read -r line || [[ -n "$line" ]]; do
|
||||||
|
[[ -z "${line//[[:space:]]/}" ]] && continue
|
||||||
|
[[ "$line" =~ ^[[:space:]]*# ]] && continue
|
||||||
|
name="$(trim "${line%%|*}")"
|
||||||
|
rest="${line#*|}"
|
||||||
|
if [[ "$line" == *"|"* ]]; then
|
||||||
|
envspec="$(trim "${rest%%|*}")"
|
||||||
|
label="$(trim "${rest#*|}")"
|
||||||
|
else
|
||||||
|
envspec=""; label=""
|
||||||
|
fi
|
||||||
|
[[ -n "$name" ]] || { echo "ERROR: arm with empty name in $ARMS_FILE" >&2; exit 1; }
|
||||||
|
# every env token must be VAR=value (a typo'd arm must fail loudly, not quietly)
|
||||||
|
for tok in $envspec; do
|
||||||
|
[[ "$tok" =~ ^[A-Za-z_][A-Za-z0-9_]*=.*$ ]] \
|
||||||
|
|| { echo "ERROR: arm '$name' has a malformed env token: '$tok'" >&2; exit 1; }
|
||||||
|
done
|
||||||
|
ARM_NAMES+=("$name"); ARM_ENVS+=("$envspec"); ARM_LABELS+=("$label")
|
||||||
|
done < "$ARMS_FILE"
|
||||||
|
(( ${#ARM_NAMES[@]} > 0 )) || { echo "ERROR: no arms parsed from $ARMS_FILE" >&2; exit 1; }
|
||||||
|
dupes="$(printf '%s\n' "${ARM_NAMES[@]}" | sort | uniq -d)"
|
||||||
|
[[ -z "$dupes" ]] || { echo "ERROR: duplicate arm name(s): $dupes" >&2; exit 1; }
|
||||||
|
if [[ -n "$REFERENCE" ]]; then
|
||||||
|
printf '%s\n' "${ARM_NAMES[@]}" | grep -qx "$REFERENCE" \
|
||||||
|
|| { echo "ERROR: --reference '$REFERENCE' is not one of the arms" >&2; exit 1; }
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── parse the panel file ─────────────────────────────────────────────────────
|
||||||
|
OPP_DIRS=(); OPP_NAMES=(); OPP_STYLES=()
|
||||||
|
while IFS= read -r line || [[ -n "$line" ]]; do
|
||||||
|
[[ -z "${line//[[:space:]]/}" ]] && continue
|
||||||
|
[[ "$line" =~ ^[[:space:]]*# ]] && continue
|
||||||
|
dir="$(trim "${line%%|*}")"
|
||||||
|
style=""; note=""
|
||||||
|
if [[ "$line" == *"|"* ]]; then
|
||||||
|
rest="${line#*|}"
|
||||||
|
if [[ "$rest" == *"|"* ]]; then
|
||||||
|
style="$(trim "${rest%%|*}")"; note="$(trim "${rest#*|}")"
|
||||||
|
else
|
||||||
|
style="$(trim "$rest")"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
[[ "$dir" == */* ]] || dir="/tmp/tr_bots/$dir"
|
||||||
|
[[ -d "$dir" ]] || { echo "ERROR: opponent bot dir not found: $dir" >&2; exit 1; }
|
||||||
|
OPP_DIRS+=("$dir"); OPP_NAMES+=("$(basename "$dir")"); OPP_STYLES+=("$style")
|
||||||
|
done < "$PANEL_FILE"
|
||||||
|
(( ${#OPP_NAMES[@]} > 0 )) || { echo "ERROR: no opponents parsed from $PANEL_FILE" >&2; exit 1; }
|
||||||
|
dupes="$(printf '%s\n' "${OPP_NAMES[@]}" | sort | uniq -d)"
|
||||||
|
[[ -z "$dupes" ]] || { echo "ERROR: duplicate opponent name(s) in the panel: $dupes" >&2; exit 1; }
|
||||||
|
(( ${#OPP_NAMES[@]} >= 4 )) || { echo "ERROR: a panel below 4 opponents cannot carry a sign test" >&2; exit 1; }
|
||||||
|
|
||||||
|
# ── fresh session dir ────────────────────────────────────────────────────────
|
||||||
|
rm -rf "$OUTDIR"
|
||||||
|
mkdir -p "$OUTDIR"
|
||||||
|
WORK="$OUTDIR/.work"; mkdir -p "$WORK"
|
||||||
|
|
||||||
|
# ── build ONE frozen bot from HEAD (single binary for every arm) ─────────────
|
||||||
|
COMMIT="$(git -C "$REPO" rev-parse HEAD)"
|
||||||
|
FROZEN_DIR="$OUTDIR/frozen/ModularBot"
|
||||||
|
FROZEN_BIN="$FROZEN_DIR/ModularBot_bin"
|
||||||
|
FROZEN_JSON="$FROZEN_DIR/ModularBot.json"
|
||||||
|
mkdir -p "$FROZEN_DIR"
|
||||||
|
BUILDDIR="$WORK/head"; mkdir -p "$BUILDDIR"
|
||||||
|
echo "[tournament] exporting HEAD ($COMMIT) -> $BUILDDIR"
|
||||||
|
git -C "$REPO" archive HEAD | tar -x -C "$BUILDDIR"
|
||||||
|
|
||||||
|
NIMCACHE="${TOURNAMENT_NIMCACHE:-/tmp/nc_j118}"
|
||||||
|
echo "[tournament] building the frozen ModularBot (nim c -d:release, nimcache=$NIMCACHE)…"
|
||||||
|
( cd "$BUILDDIR/ModularBot_garage" && \
|
||||||
|
nim c -d:release --nimcache:"$NIMCACHE" --out:"$FROZEN_BIN" src/ModularBot.nim ) \
|
||||||
|
> "$OUTDIR/frozen/build.log" 2>&1 \
|
||||||
|
|| { echo "ERROR: frozen build failed — see $OUTDIR/frozen/build.log" >&2; tail -30 "$OUTDIR/frozen/build.log" >&2; exit 1; }
|
||||||
|
[[ -x "$FROZEN_BIN" ]] || { echo "ERROR: frozen build produced no binary" >&2; exit 1; }
|
||||||
|
BINSHA="$(sha256sum "$FROZEN_BIN" | awk '{print $1}')"
|
||||||
|
echo "[tournament] frozen binary sha256=$BINSHA"
|
||||||
|
|
||||||
|
cat > "$FROZEN_JSON" <<'JSON'
|
||||||
|
{
|
||||||
|
"name": "ModularBot",
|
||||||
|
"version": "0.1.0",
|
||||||
|
"authors": ["Davide Cappellini"],
|
||||||
|
"description": "frozen tournament build of ModularBot (tournament_run.sh)",
|
||||||
|
"homepage": "",
|
||||||
|
"countryCodes": ["IT"],
|
||||||
|
"gameTypes": ["classic", "1v1"],
|
||||||
|
"platform": "Nim",
|
||||||
|
"programmingLang": "Nim"
|
||||||
|
}
|
||||||
|
JSON
|
||||||
|
|
||||||
|
# ── session.json ─────────────────────────────────────────────────────────────
|
||||||
|
json_esc() { printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g'; }
|
||||||
|
{
|
||||||
|
printf '{\n'
|
||||||
|
printf ' "commit": "%s",\n' "$COMMIT"
|
||||||
|
printf ' "binary_sha256": "%s",\n' "$BINSHA"
|
||||||
|
printf ' "binary": "frozen/ModularBot/ModularBot_bin",\n'
|
||||||
|
printf ' "rounds": %d,\n' "$ROUNDS"
|
||||||
|
printf ' "runs": %d,\n' "$RUNS"
|
||||||
|
printf ' "conc": %d,\n' "$CONC"
|
||||||
|
printf ' "game_type": "%s",\n' "$GAME_TYPE"
|
||||||
|
printf ' "reference": "%s",\n' "$(json_esc "$REFERENCE")"
|
||||||
|
printf ' "timestamp": "%s",\n' "$(date -Is)"
|
||||||
|
printf ' "outdir": "%s",\n' "$(json_esc "$OUTDIR")"
|
||||||
|
printf ' "arms_file": "%s",\n' "$(json_esc "$ARMS_FILE")"
|
||||||
|
printf ' "panel_file": "%s",\n' "$(json_esc "$PANEL_FILE")"
|
||||||
|
printf ' "arms": [\n'
|
||||||
|
for i in "${!ARM_NAMES[@]}"; do
|
||||||
|
printf ' {"name": "%s", "env": "%s", "label": "%s"}' \
|
||||||
|
"$(json_esc "${ARM_NAMES[$i]}")" "$(json_esc "${ARM_ENVS[$i]}")" "$(json_esc "${ARM_LABELS[$i]}")"
|
||||||
|
(( i + 1 < ${#ARM_NAMES[@]} )) && printf ',' || true
|
||||||
|
printf '\n'
|
||||||
|
done
|
||||||
|
printf ' ],\n'
|
||||||
|
printf ' "opponents": [\n'
|
||||||
|
for i in "${!OPP_NAMES[@]}"; do
|
||||||
|
printf ' {"name": "%s", "dir": "%s", "style": "%s"}' \
|
||||||
|
"$(json_esc "${OPP_NAMES[$i]}")" "$(json_esc "${OPP_DIRS[$i]}")" "$(json_esc "${OPP_STYLES[$i]}")"
|
||||||
|
(( i + 1 < ${#OPP_NAMES[@]} )) && printf ',' || true
|
||||||
|
printf '\n'
|
||||||
|
done
|
||||||
|
printf ' ]\n}\n'
|
||||||
|
} > "$OUTDIR/session.json"
|
||||||
|
|
||||||
|
# ── one job ──────────────────────────────────────────────────────────────────
|
||||||
|
export REPO RUN_BRIDGE ROUNDS OUTDIR WORK FROZEN_BIN FROZEN_JSON MAX_ATTEMPTS
|
||||||
|
|
||||||
|
tournament_run_one() {
|
||||||
|
local opponent_dir="$1" opponent="$2" arm="$3" run="$4" envspec="$5"
|
||||||
|
local dir="$OUTDIR/$opponent/$arm"
|
||||||
|
local work="$WORK/$opponent/$arm/run$run"
|
||||||
|
local botdir="$work/bots/ModularBot"
|
||||||
|
mkdir -p "$dir" "$botdir" "$work/data"
|
||||||
|
[[ -n "$envspec" ]] && printf '%s\n' "$envspec" > "$dir/run$run.arm.env"
|
||||||
|
|
||||||
|
cp "$FROZEN_JSON" "$botdir/ModularBot.json"
|
||||||
|
cat > "$botdir/ModularBot.sh" <<SH
|
||||||
|
#!/bin/sh
|
||||||
|
cd "\$(dirname "\$0")"
|
||||||
|
exec ./ModularBot_bin > "$dir/run$run.bot.stdout.log" 2> "$dir/run$run.bot.stderr.log"
|
||||||
|
SH
|
||||||
|
chmod +x "$botdir/ModularBot.sh"
|
||||||
|
ln -sf "$FROZEN_BIN" "$botdir/ModularBot_bin"
|
||||||
|
|
||||||
|
local attempt=0
|
||||||
|
while :; do
|
||||||
|
attempt=$((attempt + 1))
|
||||||
|
rm -f "$dir/run$run.jsonl" "$dir/run$run.jsonl.rounds.json" \
|
||||||
|
"$dir/run$run.jsonl.results.json" "$dir/run$run.events.jsonl" \
|
||||||
|
"$dir/run$run.status" "$dir/run$run.bot.stdout.log" "$dir/run$run.bot.stderr.log"
|
||||||
|
|
||||||
|
local rc=0
|
||||||
|
# shellcheck disable=SC2086 # envspec is intentionally word-split
|
||||||
|
env $envspec \
|
||||||
|
SHIM_BOTDIR="$botdir" \
|
||||||
|
SHIM_BOT_NAME="ModularBot" \
|
||||||
|
SHIM_DATA="$work/data" \
|
||||||
|
TR_EVENTS_OUT="$dir/run$run.events.jsonl" \
|
||||||
|
timeout 900 "$RUN_BRIDGE" "$opponent_dir" "$ROUNDS" "$dir/run$run.jsonl" \
|
||||||
|
> "$dir/run$run.battle.log" 2>&1 || rc=$?
|
||||||
|
echo "$rc" > "$dir/run$run.status"
|
||||||
|
|
||||||
|
# liveness: a real battle has captured rows AND the subject counter line
|
||||||
|
if grep -q "subject event counts" "$dir/run$run.battle.log" 2>/dev/null \
|
||||||
|
&& grep -Eq "rows=[1-9]" "$dir/run$run.battle.log" 2>/dev/null; then
|
||||||
|
break
|
||||||
|
fi
|
||||||
|
if (( attempt >= MAX_ATTEMPTS )); then
|
||||||
|
echo "[tournament] WARNING: ${opponent}/${arm} run${run} never started after ${attempt} attempts" >&2
|
||||||
|
break
|
||||||
|
fi
|
||||||
|
sleep $((attempt * 2))
|
||||||
|
done
|
||||||
|
echo "$attempt" > "$dir/run$run.attempts"
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
export -f tournament_run_one
|
||||||
|
|
||||||
|
# ── cleanup: only this session's process groups, only this outdir ────────────
|
||||||
|
JOB_PIDS=()
|
||||||
|
cleanup() {
|
||||||
|
local rc=$?
|
||||||
|
trap - EXIT INT TERM
|
||||||
|
for p in "${JOB_PIDS[@]:-}"; do
|
||||||
|
[[ -n "$p" ]] || continue
|
||||||
|
kill -TERM -- "-$p" 2>/dev/null || kill -TERM "$p" 2>/dev/null || true
|
||||||
|
done
|
||||||
|
sleep 0.5
|
||||||
|
for p in "${JOB_PIDS[@]:-}"; do
|
||||||
|
[[ -n "$p" ]] || continue
|
||||||
|
kill -KILL -- "-$p" 2>/dev/null || true
|
||||||
|
done
|
||||||
|
pkill -f "[r]un_bridge_battle.sh.*$OUTDIR" 2>/dev/null || true
|
||||||
|
pkill -f "[r]obocode_shim.TrBattleCapture.*$OUTDIR" 2>/dev/null || true
|
||||||
|
exit "$rc"
|
||||||
|
}
|
||||||
|
trap cleanup EXIT INT TERM
|
||||||
|
|
||||||
|
# ── launch, K at a time ──────────────────────────────────────────────────────
|
||||||
|
NJOBS=$(( ${#OPP_NAMES[@]} * ${#ARM_NAMES[@]} * RUNS ))
|
||||||
|
SECONDS=0
|
||||||
|
echo "[tournament] session: ${#OPP_NAMES[@]} opponents × ${#ARM_NAMES[@]} arms × $RUNS runs = $NJOBS battles, rounds=$ROUNDS, conc=$CONC"
|
||||||
|
echo "[tournament] outdir: $OUTDIR"
|
||||||
|
RUNNING=0; DONE=0
|
||||||
|
for oi in "${!OPP_NAMES[@]}"; do
|
||||||
|
for ai in "${!ARM_NAMES[@]}"; do
|
||||||
|
for (( r=1; r<=RUNS; r++ )); do
|
||||||
|
setsid bash -c 'tournament_run_one "$1" "$2" "$3" "$4" "$5"' _ \
|
||||||
|
"${OPP_DIRS[$oi]}" "${OPP_NAMES[$oi]}" "${ARM_NAMES[$ai]}" "$r" "${ARM_ENVS[$ai]}" &
|
||||||
|
JOB_PIDS+=("$!")
|
||||||
|
RUNNING=$((RUNNING + 1))
|
||||||
|
if (( RUNNING >= CONC )); then
|
||||||
|
wait -n || true
|
||||||
|
RUNNING=$((RUNNING - 1)); DONE=$((DONE + 1))
|
||||||
|
printf '[tournament] %3d/%3d done (%ds)\n' "$DONE" "$NJOBS" "$SECONDS"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
done
|
||||||
|
done
|
||||||
|
wait || true
|
||||||
|
DONE=$NJOBS
|
||||||
|
printf '[tournament] %3d/%3d done (%ds)\n' "$DONE" "$NJOBS" "$SECONDS"
|
||||||
|
JOB_PIDS=()
|
||||||
|
|
||||||
|
# ── aggregate status ─────────────────────────────────────────────────────────
|
||||||
|
FAILED=0; STARTFAIL=0
|
||||||
|
for oi in "${!OPP_NAMES[@]}"; do
|
||||||
|
for ai in "${!ARM_NAMES[@]}"; do
|
||||||
|
for (( r=1; r<=RUNS; r++ )); do
|
||||||
|
st="$(cat "$OUTDIR/${OPP_NAMES[$oi]}/${ARM_NAMES[$ai]}/run$r.status" 2>/dev/null || echo 999)"
|
||||||
|
if [[ "$st" != "0" ]]; then
|
||||||
|
echo "[tournament] FAIL ${OPP_NAMES[$oi]}/${ARM_NAMES[$ai]} run$r (rc=$st)" >&2
|
||||||
|
FAILED=$((FAILED + 1))
|
||||||
|
fi
|
||||||
|
if ! grep -q "subject event counts" \
|
||||||
|
"$OUTDIR/${OPP_NAMES[$oi]}/${ARM_NAMES[$ai]}/run$r.battle.log" 2>/dev/null; then
|
||||||
|
STARTFAIL=$((STARTFAIL + 1))
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
done
|
||||||
|
done
|
||||||
|
echo "[tournament] done: $(( NJOBS - FAILED )) ok, $FAILED failed, $STARTFAIL never started ($(( SECONDS ))s)"
|
||||||
|
echo "[tournament] wall clock: ${SECONDS}s for $NJOBS battles"
|
||||||
|
echo "[tournament] analyze: python3 $REPO/tools/ab/tournament_analyze.py $OUTDIR${REFERENCE:+ --reference $REFERENCE}"
|
||||||
|
(( FAILED == 0 )) || exit 1
|
||||||
Reference in New Issue
Block a user