Levers 3, 4, 1 of the #57 sign-off (execution order 3->4->1), tracked in #59.
- Lever 3 (#59): one JSONL line per trainPass in training_metrics.jsonl with
exactly the scalars sacUpdate already exposes (SACMetrics: critic/actor/alpha
losses + alpha, averaged per pass) plus epoch, buffer size (replay_buffer.len),
cumulative steps and drained count. No trainer change needed.
- Lever 4 (#59): sendTrainingMsg drops all training input while SACLSTM_EVAL_MODE=1
(existing #49 harness mechanism) — eval battles can neither pollute the replay
buffer nor trigger gradient updates; one-time stderr notice at bot init.
- Lever 1 (#59): sac_train.sh evaluates every SAC_EVAL_OPPONENTS entry per cycle
(results carry opponent name in eval_log.jsonl); best-gating now uses a
composite = mean over opponents of the last-5-evals moving average per
opponent. best_score.txt format change: float composite replaces the
single-opponent integer win rate semantics (retired).
- Tests: metricsLine JSONL scalars + eval-mode suppression asserts.
Refs: #59, #57
Launch finding during campaign-v1 verification: at production sizes
(hidden 256, ~1s/step, ~13s trainer CPU per ~40s chunk process) the
save check ran only between drain-burst passes, so stepCount never
crossed nextSave before the process died — zero checkpoints persisted
across entire runs (masked at #49/#54 smoke sizes where steps were
sub-millisecond). Check now fires mid-loop; with SAVE_INTERVAL<=5
(within the per-process step budget) every chunk persists its chain.
- make_twin.sh: generates self-contained SacTwin dir in the sample-bots
archive (own json/sh identity, own weights dir seeded from a frozen
sac_best.zip copy, own round_counter) so RunTraining.java resolves it
like any sample bot; re-running resets the twin to the frozen baseline.
- src/SAC_LSTM_Bot.nim: SACLSTM_BOT_JSON env overrides the baked-in bot
json (loadBotInfo gives json total precedence, #49) so the same binary
boots under the twin's name.
- sac_train.sh: chunk loop is a while, not for-over-seq — a crash on the
FINAL chunk previously fell through ((chunk--);continue on an exhausted
seq list) and exited 0 with budget incomplete; observed live vs SacTwin.
Readiness dry-run (#54): weighted pool Corners:1,SacTwin:3 picked the twin
in 3/4 chunks; all battles counter-checked; deterministic eval parsed;
main sac_best.zip/counter untouched by twin (twin counter advanced
independently); crash-restart proven end-to-end incl. final-chunk retry.
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.
Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
(default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
matches the booted identity (mismatch = runner connect timeout).
Save/load all SAC-LSTM tensors (actor, 2 critics, 2 target critics,
alpha, Adam states) into a single .zip of .npy files. Atomic write
via temp path + rename. Adam types (AdamVar, SACAdamStates) defined
here for training.nim to use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>