Training harness #49

Closed
opened 2026-08-20 23:29:11 +02:00 by SirStone · 2 comments
Owner

Parent

#37

What to build

A training orchestration script sac_train.sh that drives SAC_LSTM_Bot training using RunTraining.java for match management.

Responsibilities:

  • Compile SAC_LSTM_Bot with nimble build -d:release
  • Run training rounds via RunTraining.java (server lifecycle, opponent connection, liveness detection)
  • Opponent sampling: configurable opponent list with weighted selection
  • Periodic evaluation: run deterministic matches (SACLSTM_EVAL_MODE=1), track win rates
  • Checkpoint management: periodic saves, track best checkpoint by evaluation score
  • Configurable total rounds, chunk size, evaluation interval via env vars or script variables

Acceptance criteria

  • sac_train.sh compiles the bot and launches training end-to-end
  • Training runs multiple rounds against configured opponents
  • Evaluation runs periodically in deterministic mode
  • Best checkpoint is saved when evaluation score improves
  • Training recovers from bot crashes (RunTraining.java liveness detection)
  • Configurable: opponents, round counts, evaluation interval

Blocked by

  • #48 (Main bot integration)
## Parent #37 ## What to build A training orchestration script `sac_train.sh` that drives SAC_LSTM_Bot training using `RunTraining.java` for match management. Responsibilities: - Compile SAC_LSTM_Bot with `nimble build -d:release` - Run training rounds via RunTraining.java (server lifecycle, opponent connection, liveness detection) - Opponent sampling: configurable opponent list with weighted selection - Periodic evaluation: run deterministic matches (SACLSTM_EVAL_MODE=1), track win rates - Checkpoint management: periodic saves, track best checkpoint by evaluation score - Configurable total rounds, chunk size, evaluation interval via env vars or script variables ## Acceptance criteria - [ ] `sac_train.sh` compiles the bot and launches training end-to-end - [ ] Training runs multiple rounds against configured opponents - [ ] Evaluation runs periodically in deterministic mode - [ ] Best checkpoint is saved when evaluation score improves - [ ] Training recovers from bot crashes (RunTraining.java liveness detection) - [ ] Configurable: opponents, round counts, evaluation interval ## Blocked by - #48 (Main bot integration)
SirStone added the ready-for-agent label 2026-08-20 23:29:11 +02:00
Author
Owner

Claimed for implementation (self-assigned as SirStone — the MCP issue-edit surface here exposes no assignee field, so this comment is the assignment record). Working on branch research/goto-controller.

Claimed for implementation (self-assigned as `SirStone` — the MCP issue-edit surface here exposes no assignee field, so this comment is the assignment record). Working on branch `research/goto-controller`.
Author
Owner

Resolution — training harness (commit df256b4)

What was built

  • SAC_LSTM_Bot/sac_train.sh (new) — the orchestration script. Compiles the bot (nimble build -d:release) + RunTraining.java, then loops chunks: each chunk runs one RunTraining.java battle (CHUNK_SIZE rounds, one battle per chunk so the bot's training thread survives rounds). Weighted opponent sampling per chunk, deterministic evaluation (SACLSTM_EVAL_MODE=1) every SAC_EVAL_INTERVAL chunks with win-rate parsed from the runner's JSONL log, best checkpoint (weights/sac_best.zip + best_score.txt, persists across harness restarts) updated when the win rate improves, crash-restart loop (rerun chunk from latest checkpoint, abort after SAC_MAX_CRASHES consecutive). All knobs env-configurable: SAC_OPPONENTS ("Name:weight,..."), SAC_TOTAL_ROUNDS, SAC_CHUNK_SIZE, SAC_EVAL_INTERVAL, SAC_EVAL_ROUNDS, SAC_EVAL_OPPONENT, SAC_MAX_CRASHES, SAC_LOG_FILE/SAC_EVAL_LOG_FILE; SACLSTM_* pass through to the bot.
  • SAC_LSTM_Bot/SAC_LSTM_Bot.json + SAC_LSTM_Bot.sh (new) — launch packaging for the Tank Royale booter (<dir>/<dir>.json + <dir>.sh convention).
  • src/SAC_LSTM_Bot/integration.nim — bumpRoundCounter(): one increment per round end to weights/round_counter.txt, the runner's liveness signal (best-effort: never raises into the event handler). opponentKey(): opponent identity helper (see below).
  • src/SAC_LSTM_Bot.nim — onRoundEnded handler calls bumpRoundCounter() (main thread).
  • tools/training_runner/RunTraining.java — parameterized result matching via BOT_NAME env (default PPO_Bot → behavior for PPO_Bot unchanged). Strictly required: name matching was hardcoded.
  • src/SAC_LSTM_Bot.json — self-reported name changed "Recurrent Royalty" → "SAC_LSTM_Bot". Found during smoke testing: the vendored API's loadBotInfo gives the json file total precedence over the booter's BOT_NAME env, so a mismatch between the src json identity and the booter's dir-json identity makes the runner wait forever on a bot that actually connected ("connected 1 of 2, Pending: SAC_LSTM_Bot"). Identities must match; description in the json documents this.

Opponent-ID resolution (Q14 follow-up)

tmkNewBattle still carries the numeric scannedBotId; the training thread now keys the "opponent changed → clear replay buffer" rule on opponentKey(id) = getBotName(id) when non-empty (v1.0.1 BotListUpdate table), falling back to the numeric id as string in the pre-BotListUpdate window. Caveat: until the first BotListUpdate arrives, getBotName returns "" and the fallback key is the raw id — a same-id opponent across battles in that window keeps the buffer (safe, id-stable), but the id→name mapping may lag the very first battle of a process. Lookup happens training-side via the API's lock-guarded table — no strings cross thread boundaries (ORC rule from #48).

How to run

cd SAC_LSTM_Bot && ./sac_train.sh                    # defaults: 100 rounds, chunks of 10
SAC_OPPONENTS="Corners:3,Crazy:1" SAC_TOTAL_ROUNDS=500 ./sac_train.sh

Test status

  • nimble build + nimble test: 8/8 suites green.
  • New asserts in test_integration.nim: name-keyed NewBattle clear/keep (incl. numeric fallback via seeded updateBotNames) and bumpRoundCounter increment/cold-start.
  • End-to-end smoke (tmux, real battles): 4 training rounds in 2-round chunks vs MyFirstBot + 2 deterministic eval battles → counter checks passed every battle, weights/sac_latest.zip written by the bot's I/O thread, eval win rate parsed (0% vs untrained bot, expected), sac_best.zip + best_score.txt created, harness exit 0. Crash-recovery path observed live in an earlier run (bot identity mismatch → runner connect timeout → "crash #1/#2 — restarting chunk" loop engaged and aborted correctly at the configured limit).

Notes / deliberate simplifications

  • Weights path is pinned by the harness to <bot>/weights/sac_latest.zip (not env-overridable): RunTraining hardcodes the counter at $BOT_DIR/weights/, and the bot writes the counter next to its weights.
  • Eval battles still train the bot (no freeze during eval) — the bot is always-on by design (#48); noted as acceptable.
  • Default SACLSTM_HIDDEN_SIZE=256 updates are slow (~seconds/step at batch 16): real runs should tune SACLSTM_SAVE_INTERVAL/SACLSTM_BATCH_SIZE/SACLSTM_HIDDEN_SIZE env knobs; the smoke used hidden=32/batch=8/save=50 to exercise checkpointing in minutes.
  • Budget accounting is chunk-count based (a crash reruns the chunk, losing at most one chunk's progress), replacing run.sh's counter-derived resume since eval rounds also advance the counter.
## Resolution — training harness (commit `df256b4`) **What was built** - **`SAC_LSTM_Bot/sac_train.sh`** (new) — the orchestration script. Compiles the bot (`nimble build -d:release`) + `RunTraining.java`, then loops chunks: each chunk runs one `RunTraining.java` battle (`CHUNK_SIZE` rounds, one battle per chunk so the bot's training thread survives rounds). Weighted opponent sampling per chunk, deterministic evaluation (`SACLSTM_EVAL_MODE=1`) every `SAC_EVAL_INTERVAL` chunks with win-rate parsed from the runner's JSONL log, best checkpoint (`weights/sac_best.zip` + `best_score.txt`, persists across harness restarts) updated when the win rate improves, crash-restart loop (rerun chunk from latest checkpoint, abort after `SAC_MAX_CRASHES` consecutive). All knobs env-configurable: `SAC_OPPONENTS` (`"Name:weight,..."`), `SAC_TOTAL_ROUNDS`, `SAC_CHUNK_SIZE`, `SAC_EVAL_INTERVAL`, `SAC_EVAL_ROUNDS`, `SAC_EVAL_OPPONENT`, `SAC_MAX_CRASHES`, `SAC_LOG_FILE`/`SAC_EVAL_LOG_FILE`; `SACLSTM_*` pass through to the bot. - **`SAC_LSTM_Bot/SAC_LSTM_Bot.json` + `SAC_LSTM_Bot.sh`** (new) — launch packaging for the Tank Royale booter (`<dir>/<dir>.json` + `<dir>.sh` convention). - **`src/SAC_LSTM_Bot/integration.nim`** — `bumpRoundCounter()`: one increment per round end to `weights/round_counter.txt`, the runner's liveness signal (best-effort: never raises into the event handler). `opponentKey()`: opponent identity helper (see below). - **`src/SAC_LSTM_Bot.nim`** — `onRoundEnded` handler calls `bumpRoundCounter()` (main thread). - **`tools/training_runner/RunTraining.java`** — parameterized result matching via `BOT_NAME` env (default `PPO_Bot` → behavior for PPO_Bot unchanged). Strictly required: name matching was hardcoded. - **`src/SAC_LSTM_Bot.json`** — self-reported name changed "Recurrent Royalty" → "SAC_LSTM_Bot". Found during smoke testing: the vendored API's `loadBotInfo` gives the json file total precedence over the booter's `BOT_NAME` env, so a mismatch between the src json identity and the booter's dir-json identity makes the runner wait forever on a bot that actually connected ("connected 1 of 2, Pending: SAC_LSTM_Bot"). Identities must match; description in the json documents this. **Opponent-ID resolution (Q14 follow-up)** `tmkNewBattle` still carries the numeric `scannedBotId`; the training thread now keys the "opponent changed → clear replay buffer" rule on `opponentKey(id)` = `getBotName(id)` when non-empty (v1.0.1 BotListUpdate table), falling back to the numeric id as string in the pre-BotListUpdate window. Caveat: until the first BotListUpdate arrives, `getBotName` returns `""` and the fallback key is the raw id — a same-id opponent across battles in that window keeps the buffer (safe, id-stable), but the id→name mapping may lag the very first battle of a process. Lookup happens training-side via the API's lock-guarded table — no strings cross thread boundaries (ORC rule from #48). **How to run** ``` cd SAC_LSTM_Bot && ./sac_train.sh # defaults: 100 rounds, chunks of 10 SAC_OPPONENTS="Corners:3,Crazy:1" SAC_TOTAL_ROUNDS=500 ./sac_train.sh ``` **Test status** - `nimble build` + `nimble test`: 8/8 suites green. - New asserts in `test_integration.nim`: name-keyed NewBattle clear/keep (incl. numeric fallback via seeded `updateBotNames`) and `bumpRoundCounter` increment/cold-start. - **End-to-end smoke (tmux, real battles)**: 4 training rounds in 2-round chunks vs MyFirstBot + 2 deterministic eval battles → counter checks passed every battle, `weights/sac_latest.zip` written by the bot's I/O thread, eval win rate parsed (0% vs untrained bot, expected), `sac_best.zip` + `best_score.txt` created, harness exit 0. Crash-recovery path observed live in an earlier run (bot identity mismatch → runner connect timeout → "crash #1/#2 — restarting chunk" loop engaged and aborted correctly at the configured limit). **Notes / deliberate simplifications** - Weights path is pinned by the harness to `<bot>/weights/sac_latest.zip` (not env-overridable): RunTraining hardcodes the counter at `$BOT_DIR/weights/`, and the bot writes the counter next to its weights. - Eval battles still train the bot (no freeze during eval) — the bot is always-on by design (#48); noted as acceptable. - Default `SACLSTM_HIDDEN_SIZE=256` updates are slow (~seconds/step at batch 16): real runs should tune `SACLSTM_SAVE_INTERVAL`/`SACLSTM_BATCH_SIZE`/`SACLSTM_HIDDEN_SIZE` env knobs; the smoke used hidden=32/batch=8/save=50 to exercise checkpointing in minutes. - Budget accounting is chunk-count based (a crash reruns the chunk, losing at most one chunk's progress), replacing run.sh's counter-derived resume since eval rounds also advance the counter.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#49