Build mirror-twin sparring partner + harness readiness check #54

Closed
opened 2026-08-21 22:36:35 +02:00 by SirStone · 2 comments
Owner

Child of the map Overnight Training Campaign — SAC_LSTM_Bot.

Question

Can the harness run the overnight configuration end-to-end? Concretely: create the renamed twin dir (distinct bot name, e.g. SacTwin; own .json/.sh; own weights dir seeded from a frozen copy of the current sac_best.zip) placed so RunTraining.java resolves it alongside the sample bots; verify none of the three identical-name breakages occur (score double-counting, shared liveness counter, checkpoint write races); then dry-run one battle chunk against the planned pool to prove weighted sampling, deterministic eval, checkpoint writing, and crash-restart all behave. The answer records the exact paths and mechanism used, so ticket Lock campaign-v1 config and launch overnight run can reuse them verbatim.

Child of the map *Overnight Training Campaign — SAC_LSTM_Bot*. ## Question Can the harness run the overnight configuration end-to-end? Concretely: create the renamed twin dir (distinct bot name, e.g. `SacTwin`; own `.json`/`.sh`; own weights dir seeded from a frozen copy of the current `sac_best.zip`) placed so `RunTraining.java` resolves it alongside the sample bots; verify none of the three identical-name breakages occur (score double-counting, shared liveness counter, checkpoint write races); then dry-run one battle chunk against the planned pool to prove weighted sampling, deterministic eval, checkpoint writing, and crash-restart all behave. The answer records the exact paths and mechanism used, so ticket *Lock campaign-v1 config and launch overnight run* can reuse them verbatim.
SirStone added the wayfinder:task label 2026-08-21 22:36:35 +02:00
Author
Owner

Claimed for implementation (self-assigned as SirStone — the MCP issue-edit surface exposes no working assignee field, so this comment is the assignment record). Working on branch research/goto-controller.

Claimed for implementation (self-assigned as `SirStone` — the MCP issue-edit surface exposes no working assignee field, so this comment is the assignment record). Working on branch `research/goto-controller`.
Author
Owner

Resolution — mirror-twin sparring partner + readiness check (commit 6a294ad)

Placement decision

Dropped a self-contained twin dir into the external sample-bots archive (SAMPLE_BOTS_DIR default) instead of a symlinked merged-pool dir: RunTraining.java resolves opponents as $SAMPLE_BOTS_DIR/<name> and sac_train.sh samples names against the same dir, so a plain sibling dir needs zero harness/env changes and works for both training chunks and eval. Reproducibility lives in-repo: the committed generator rebuilds the instance anywhere.

Exact paths / mechanism

  • Generator (repo): SAC_LSTM_Bot/make_twin.sh — run after nimble build -d:release. Re-running resets the twin to the frozen baseline (reproducible opponent).
  • Generated instance: /home/davide/Projects/tank-royale/sample-bots/java/build/archive/SacTwin/ containing SAC_LSTM_Bot (copy of the built binary), SacTwin.json ("name": "SacTwin"), SacTwin.sh (exports SACLSTM_BOT_JSON=$DIR/SacTwin.json + SACLSTM_WEIGHTS_PATH=$DIR/weights/sac_latest.zip, then execs the binary), weights/{sac_latest.zip, sac_best.zip} seeded from a frozen copy of the repo-side weights/sac_best.zip (+ fresh round_counter.txt=0).
  • Repo diff: make_twin.sh (new); src/SAC_LSTM_Bot.nim — SACLSTM_BOT_JSON env now overrides the baked-in src json (the API's loadBotInfo gives json total precedence over env, #49 lesson — without this the copied binary would still report SAC_LSTM_Bot = identical-name breakage #1); sac_train.sh — chunk loop converted to while.

Three identical-name breakages — verified absent

  1. Score matching: results matched via getName().equals("SAC_LSTM_Bot"); twin reports SacTwin → no double-count (scores per round were sane throughout).
  2. Liveness counter: bumpRoundCounter() writes next to its own SACLSTM_WEIGHTS_PATH; RunTraining watches only the main dir's file. Main counter advanced exactly +rounds per battle every time; twin counter advanced independently (0→6 over its 3 chunks).
  3. Checkpoints: twin writes only archive/SacTwin/weights/; main sac_best.zip md5/mtime byte-identical before/after the dry-run.

Dry-run evidence (tmux)

  • SAC_OPPONENTS="Corners:1,SacTwin:3" SAC_TOTAL_ROUNDS=8 SAC_CHUNK_SIZE=2 SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=2 SACLSTM_HIDDEN_SIZE=32 SACLSTM_BATCH_SIZE=8 SACLSTM_SAVE_INTERVAL=50 ./sac_train.sh → weighted sampling picked SacTwin in 3/4 chunks, all battles ended Counter check passed (8→20 across train+eval), deterministic eval ran twice (0% vs Corners, parsed from JSONL), harness exit 0.
  • Crash-restart, live: killed the main bot process mid-final-chunk (pkill on the exact binary path) → round_counter 69 < expected 92 … aborting for restart → crash #1 — restarting chunk → rerun battle completed 30/30 rounds, Counter check passed: 139 == expected 139, exit 0. This exposed + fixed a real bug: a crash on the FINAL chunk previously fell through the for chunk in $(seq…) loop (exhausted list) and exited 0 with budget incomplete — now a while loop, proven by the same experiment (old code would have exited after crash #1).

Side effects on main bot state (all expected)

Main round_counter.txt advances by every round of every battle incl. eval; training_log.jsonl/eval_log.jsonl append; weights/sac_latest.zip updates only via interval saves; sac_best.zip/best_score.txt untouched unless eval improves. Nothing twin-related touches the main dir.

Findings #56 must know

  • Checkpoints persist via mid-battle interval saves ONLY: with SACLSTM_SAVE_INTERVAL=999999999 a full clean battle wrote nothing — the shutdown/final-save path never reaches disk in harness context (process teardown wins the race). Rule: keep SACLSTM_SAVE_INTERVAL well below per-chunk gradient-step counts. Values 10–20 fired reliably in smokes; 50 did not fire in short chunks vs SacTwin (twin battles kill the main bot early at the frozen baseline → fewer transitions per chunk). Size accordingly for campaign-v1.
  • The mid-battle liveness guard's System.exit(1) never fires (apparently swallowed by the vendored runner's event dispatch); corpse detection falls to the end-of-battle completeness check — slower (~60s) but reliable. Don't count on fast mid-battle aborts.
  • Cosmetic: a retried first chunk logs === Chunk 0/N === (harmless).
  • SAC_EVAL_OPPONENT defaults to the first pool entry — put the intended eval opponent first or set it explicitly in campaign-v1 config.
  • Overnight pool usage verbatim: generate once via ./make_twin.sh, then e.g. SAC_OPPONENTS="Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1" — no other changes.
## Resolution — mirror-twin sparring partner + readiness check (commit `6a294ad`) **Placement decision** Dropped a self-contained twin dir into the external sample-bots archive (`SAMPLE_BOTS_DIR` default) instead of a symlinked merged-pool dir: `RunTraining.java` resolves opponents as `$SAMPLE_BOTS_DIR/<name>` and `sac_train.sh` samples names against the same dir, so a plain sibling dir needs **zero harness/env changes** and works for both training chunks and eval. Reproducibility lives in-repo: the committed generator rebuilds the instance anywhere. **Exact paths / mechanism** - Generator (repo): `SAC_LSTM_Bot/make_twin.sh` — run after `nimble build -d:release`. Re-running **resets the twin to the frozen baseline** (reproducible opponent). - Generated instance: `/home/davide/Projects/tank-royale/sample-bots/java/build/archive/SacTwin/` containing `SAC_LSTM_Bot` (copy of the built binary), `SacTwin.json` (`"name": "SacTwin"`), `SacTwin.sh` (exports `SACLSTM_BOT_JSON=$DIR/SacTwin.json` + `SACLSTM_WEIGHTS_PATH=$DIR/weights/sac_latest.zip`, then execs the binary), `weights/{sac_latest.zip, sac_best.zip}` seeded from a **frozen copy** of the repo-side `weights/sac_best.zip` (+ fresh `round_counter.txt`=0). - Repo diff: `make_twin.sh` (new); `src/SAC_LSTM_Bot.nim` — `SACLSTM_BOT_JSON` env now overrides the baked-in src json (the API's `loadBotInfo` gives json total precedence over env, #49 lesson — without this the copied binary would still report `SAC_LSTM_Bot` = identical-name breakage #1); `sac_train.sh` — chunk loop converted to `while`. **Three identical-name breakages — verified absent** 1. Score matching: results matched via `getName().equals("SAC_LSTM_Bot")`; twin reports `SacTwin` → no double-count (scores per round were sane throughout). 2. Liveness counter: `bumpRoundCounter()` writes next to *its own* `SACLSTM_WEIGHTS_PATH`; `RunTraining` watches only the main dir's file. Main counter advanced exactly +rounds per battle every time; twin counter advanced independently (0→6 over its 3 chunks). 3. Checkpoints: twin writes only `archive/SacTwin/weights/`; main `sac_best.zip` md5/mtime byte-identical before/after the dry-run. **Dry-run evidence (tmux)** - `SAC_OPPONENTS="Corners:1,SacTwin:3" SAC_TOTAL_ROUNDS=8 SAC_CHUNK_SIZE=2 SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=2 SACLSTM_HIDDEN_SIZE=32 SACLSTM_BATCH_SIZE=8 SACLSTM_SAVE_INTERVAL=50 ./sac_train.sh` → weighted sampling picked **SacTwin in 3/4 chunks**, all battles ended `Counter check passed` (8→20 across train+eval), deterministic eval ran twice (0% vs Corners, parsed from JSONL), harness exit 0. - Crash-restart, live: killed the main bot process mid-final-chunk (`pkill` on the exact binary path) → `round_counter 69 < expected 92 … aborting for restart` → `crash #1 — restarting chunk` → rerun battle completed 30/30 rounds, `Counter check passed: 139 == expected 139`, exit 0. **This exposed + fixed a real bug**: a crash on the FINAL chunk previously fell through the `for chunk in $(seq…)` loop (exhausted list) and exited 0 with budget incomplete — now a `while` loop, proven by the same experiment (old code would have exited after `crash #1`). **Side effects on main bot state (all expected)** Main `round_counter.txt` advances by every round of every battle incl. eval; `training_log.jsonl`/`eval_log.jsonl` append; `weights/sac_latest.zip` updates only via interval saves; `sac_best.zip`/`best_score.txt` untouched unless eval improves. Nothing twin-related touches the main dir. **Findings #56 must know** - **Checkpoints persist via mid-battle interval saves ONLY**: with `SACLSTM_SAVE_INTERVAL=999999999` a full clean battle wrote nothing — the shutdown/final-save path never reaches disk in harness context (process teardown wins the race). Rule: keep `SACLSTM_SAVE_INTERVAL` well below per-chunk gradient-step counts. Values 10–20 fired reliably in smokes; **50 did not fire in short chunks vs SacTwin** (twin battles kill the main bot early at the frozen baseline → fewer transitions per chunk). Size accordingly for campaign-v1. - The mid-battle liveness guard's `System.exit(1)` never fires (apparently swallowed by the vendored runner's event dispatch); corpse detection falls to the end-of-battle completeness check — slower (~60s) but reliable. Don't count on fast mid-battle aborts. - Cosmetic: a retried first chunk logs `=== Chunk 0/N ===` (harmless). - `SAC_EVAL_OPPONENT` defaults to the *first* pool entry — put the intended eval opponent first or set it explicitly in campaign-v1 config. - Overnight pool usage verbatim: generate once via `./make_twin.sh`, then e.g. `SAC_OPPONENTS="Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1"` — no other changes.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#54