df256b4d3e
sac_train.sh orchestrates chunked self-play via tools/training_runner/ RunTraining.java: weighted opponent sampling per chunk, deterministic eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on the runner's liveness detection. Supporting changes: - integration.nim: opponentKey() keys the NewBattle buffer-clear rule on getBotName(id) with numeric-id fallback (#49 Q14 follow-up); bumpRoundCounter() emits the per-round liveness signal. - SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter(). - RunTraining.java: BOT_NAME env parameterizes result matching (default PPO_Bot, unchanged behavior for PPO). - Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter; src json name aligned to 'SAC_LSTM_Bot' so self-reported identity matches the booted identity (mismatch = runner connect timeout).
5 lines
198 B
Bash
Executable File
5 lines
198 B
Bash
Executable File
#!/bin/sh
|
|
# Launch config for tools/training_runner/RunTraining.java (#49): the runner
|
|
# executes <json-basename>.sh inside the bot dir (sample-bots convention).
|
|
exec "$(dirname "$0")/SAC_LSTM_Bot"
|