Lock campaign-v1 config and launch overnight run #56
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Blocked by: #54, #55
Child of the map Overnight Training Campaign — SAC_LSTM_Bot.
Question
What exactly runs tonight, and how is it watched? Lock: opponent pool incl. twin weighting (working recommendation
Corners:3,Crazy:2,RamFire:1,Target:1+ twin at weight 2 — adjust freely from the twin-build and observability findings), real (non-smoke) hyperparameters, budget = generous wall-clock CEILING, not a deadline — the run continues while healthy; periodic deterministic eval, and intervention thresholds from the observability inventory. Launch in tmux and leave the monitoring specified so a fresh agent session can execute check-ins without re-deriving anything. The answer records the final config, tmux session names, and a glanceable health-check procedure.Claimed for implementation (self-assigned as
SirStone— MCP assignee field unreliable, this comment is the assignment record). Working on branchresearch/goto-controller. Inputs: observability inventory (#55 comment 461) + twin launch notes (#54 comment 466). Plan: lock config → notebook → regenerate twin → launch tmuxsac_campaign→ verify HEALTHY checklist → record resolution.Resolution — campaign-v1 locked, launched, HEALTHY (commits
2f49cb2,2653671,4b64bf1)Final locked config (verbatim, copy-pasteable)
Rationale per knob in the notebook config table. Deviations from module defaults / prior recommendations, one line each:
SAVE_INTERVAL=5(not 500-default, not #54's 10–20): at hidden 256 a step costs ~1 s and a chunk process fits only ~10–18 steps ⇒ the interval must sit inside the per-process step budget; 5 ⇒ checkpoint every ~5–10 s of active training. Enabled by commit2653671.TOTAL_ROUNDS=25000(map implied smaller): measured throughput ~3400 rounds/h means 2000 was a 1-hour budget; ceiling must bind first.3e-4LRs, γ=0.99, τ=0.005, target entropy −4.0, buffer 500k, burn-in 8, window 16).Launch
sac_campaign, launched 2026-08-22 00:21:55 CEST (epoch 1787350915) fromresearch/goto-controller@2653671../make_twin.sh, seed md5-verified (61521cff…) from the campaign-v1 baseline (fresh random-init hidden-256 checkpoint; old 32-hidden smoke weights archived toweights_smoke49_backup/).First-health-check evidence (t+9 min, per #55 HEALTHY checklist)
campaign_stdout.log; 0 crash/abort bannersEarly 0% evals = near-random policy minutes in; SLOW-LEARNER watch rule applies, never intervene.
Launch incident worth the record (fixed in
2653671)Two launch attempts persisted zero checkpoints: at production sizes the save check ran only between drain-burst passes, so
stepCountnever crossednextSavewithin a process lifetime (~10–18 steps at ~1 s/step vs interval 20). Smoke runs masked it because hidden-32 steps were sub-millisecond. Fix: check moved inside the gradient-step loop; validated (interval=5 fires ~12 s into a battle). Full chain in the notebook's Incident 1.Paths
SAC_LSTM_Bot/docs/campaign_notebook.mdSAC_LSTM_Bot/campaign_stdout.log· training JSONL:SAC_LSTM_Bot/training_log.jsonl· eval JSONL:SAC_LSTM_Bot/eval_log.jsonlSAC_LSTM_Bot/weights/{sac_latest.zip,sac_best.zip,best_score.txt}· twin:$SAMPLE_BOTS_DIR/SacTwin/(frozen)Glanceable check-in (<1 min, for future monitor sessions)
Intervene on: ≥3 consecutive crash banners · counter frozen ≥15 min · zip >10 min stale while counter advances · disk <1 GB. Kill at ceiling. Full procedure + thresholds in the notebook.
Closing #56; #57 unblocks.
Ceiling extended by orchestrator decision (human mandate: no deadlines / let the budget complete): old kill = procedural monitor instruction only (campaign_notebook.md lines 41/110, epoch 1787394115 / 12:22 CEST — sweep found NO os-level watchdog: tmux panes, ps//proc cmdline scan for the epoch, at/cron/systemd timers all clean; sac_train.sh & RunTraining.java have no wall-clock ceiling), inerted by rewriting both instructions to the new ceiling. New safety net: hard transient user-systemd unit
sac-ceiling-netarmed 07:46:04 CEST, fires 16:30:00 CEST today (epoch 1787409000) —tmux kill-session -t sac_campaignthen pgrep-guarded pkill of straggler harness processes, same style as the original stop. Training uninterrupted — counter verified advancing: 21253 → 21352 (+99) over 147 s, sac_latest.zip fresh ≤12 s, training_log.jsonl +69 lines, tmux session alive.