Lever 2 (#59): x1.25 aggression mult on damage dealt, flat +0.5 hit bonus,
-3.0 per bot-bot collision (server deals RAM_DAMAGE=0.6 to both parties but
only notifies the hitter), escalating proximity deterrent below 12% arena
diagonal suppressed while dealing damage. Win/loss terminals unchanged and
dominant. All weights TUNABLE consts marked ponytail. SACLSTM_REWARD_DEBUG=1
env-gated reward_debug.log for calibration greps.
Lever 5 (#59): no code needed — SACLSTM_LR_ACTOR/LR_CRITIC/LR_ALPHA (3e-4)
and SACLSTM_TARGET_ENTROPY (-4.0) were already env-overridable in training.nim.
Smoke vs RamFire+Crazy (hidden=32, random init, isolated weights): 75 ram
penalties, 381 charge events, hit bonuses firing, 0 crashes, metrics JSONL
flowing. Tests: 8/8 suites green incl. new assert-level term math.
Levers 3, 4, 1 of the #57 sign-off (execution order 3->4->1), tracked in #59.
- Lever 3 (#59): one JSONL line per trainPass in training_metrics.jsonl with
exactly the scalars sacUpdate already exposes (SACMetrics: critic/actor/alpha
losses + alpha, averaged per pass) plus epoch, buffer size (replay_buffer.len),
cumulative steps and drained count. No trainer change needed.
- Lever 4 (#59): sendTrainingMsg drops all training input while SACLSTM_EVAL_MODE=1
(existing #49 harness mechanism) — eval battles can neither pollute the replay
buffer nor trigger gradient updates; one-time stderr notice at bot init.
- Lever 1 (#59): sac_train.sh evaluates every SAC_EVAL_OPPONENTS entry per cycle
(results carry opponent name in eval_log.jsonl); best-gating now uses a
composite = mean over opponents of the last-5-evals moving average per
opponent. best_score.txt format change: float composite replaces the
single-opponent integer win rate semantics (retired).
- Tests: metricsLine JSONL scalars + eval-mode suppression asserts.
Refs: #59, #57
sac_train.sh orchestrates chunked self-play via tools/training_runner/
RunTraining.java: weighted opponent sampling per chunk, deterministic
eval (SACLSTM_EVAL_MODE=1) every N chunks with win-rate tracking, best
checkpoint (weights/sac_best.zip) by eval score, crash-restart loop on
the runner's liveness detection.
Supporting changes:
- integration.nim: opponentKey() keys the NewBattle buffer-clear rule on
getBotName(id) with numeric-id fallback (#49 Q14 follow-up);
bumpRoundCounter() emits the per-round liveness signal.
- SAC_LSTM_Bot.nim: onRoundEnded -> bumpRoundCounter().
- RunTraining.java: BOT_NAME env parameterizes result matching
(default PPO_Bot, unchanged behavior for PPO).
- Launch packaging: root SAC_LSTM_Bot.json + .sh for the booter;
src json name aligned to 'SAC_LSTM_Bot' so self-reported identity
matches the booted identity (mismatch = runner connect timeout).
Save/load all SAC-LSTM tensors (actor, 2 critics, 2 target critics,
alpha, Adam states) into a single .zip of .npy files. Atomic write
via temp path + rename. Adam types (AdamVar, SACAdamStates) defined
here for training.nim to use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>