Lock campaign-v1 config and launch overnight run #56

Closed
opened 2026-08-21 22:36:55 +02:00 by SirStone · 3 comments
Owner

Blocked by: #54, #55

Child of the map Overnight Training Campaign — SAC_LSTM_Bot.

Question

What exactly runs tonight, and how is it watched? Lock: opponent pool incl. twin weighting (working recommendation Corners:3,Crazy:2,RamFire:1,Target:1 + twin at weight 2 — adjust freely from the twin-build and observability findings), real (non-smoke) hyperparameters, budget = generous wall-clock CEILING, not a deadline — the run continues while healthy; periodic deterministic eval, and intervention thresholds from the observability inventory. Launch in tmux and leave the monitoring specified so a fresh agent session can execute check-ins without re-deriving anything. The answer records the final config, tmux session names, and a glanceable health-check procedure.

*Blocked by: #54, #55* Child of the map *Overnight Training Campaign — SAC_LSTM_Bot*. ## Question What exactly runs tonight, and how is it watched? Lock: opponent pool incl. twin weighting (working recommendation `Corners:3,Crazy:2,RamFire:1,Target:1` + twin at weight 2 — adjust freely from the twin-build and observability findings), real (non-smoke) hyperparameters, budget = generous wall-clock CEILING, not a deadline — the run continues while healthy; periodic deterministic eval, and intervention thresholds from the observability inventory. Launch in tmux and leave the monitoring specified so a fresh agent session can execute check-ins without re-deriving anything. The answer records the final config, tmux session names, and a glanceable health-check procedure.
SirStone added the wayfinder:task label 2026-08-21 22:36:55 +02:00
Author
Owner

Claimed for implementation (self-assigned as SirStone — MCP assignee field unreliable, this comment is the assignment record). Working on branch research/goto-controller. Inputs: observability inventory (#55 comment 461) + twin launch notes (#54 comment 466). Plan: lock config → notebook → regenerate twin → launch tmux sac_campaign → verify HEALTHY checklist → record resolution.

Claimed for implementation (self-assigned as `SirStone` — MCP assignee field unreliable, this comment is the assignment record). Working on branch `research/goto-controller`. Inputs: observability inventory (#55 comment 461) + twin launch notes (#54 comment 466). Plan: lock config → notebook → regenerate twin → launch tmux `sac_campaign` → verify HEALTHY checklist → record resolution.
Author
Owner

Resolution — campaign-v1 locked, launched, HEALTHY (commits 2f49cb2, 2653671, 4b64bf1)

Final locked config (verbatim, copy-pasteable)

cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=25000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log

Rationale per knob in the notebook config table. Deviations from module defaults / prior recommendations, one line each:

  • SAVE_INTERVAL=5 (not 500-default, not #54's 10–20): at hidden 256 a step costs ~1 s and a chunk process fits only ~10–18 steps ⇒ the interval must sit inside the per-process step budget; 5 ⇒ checkpoint every ~5–10 s of active training. Enabled by commit 2653671.
  • TOTAL_ROUNDS=25000 (map implied smaller): measured throughput ~3400 rounds/h means 2000 was a 1-hour budget; ceiling must bind first.
  • Twin at weight 1 (map suggested 2): #54 showed mirror battles end early ⇒ fewer transitions/chunk; weight 2 starves the buffer.
  • All other bot knobs are unpinned module defaults (3e-4 LRs, γ=0.99, τ=0.005, target entropy −4.0, buffer 500k, burn-in 8, window 16).

Launch

  • tmux session sac_campaign, launched 2026-08-22 00:21:55 CEST (epoch 1787350915) from research/goto-controller @ 2653671.
  • Wall-clock ceiling T+12 h ⇒ 12:22 CEST Aug 22 (epoch 1787394115): a monitor kills the session then if still running. Ceiling is a stop, not a deadline.
  • Twin regenerated pre-launch via ./make_twin.sh, seed md5-verified (61521cff…) from the campaign-v1 baseline (fresh random-init hidden-256 checkpoint; old 32-hidden smoke weights archived to weights_smoke49_backup/).

First-health-check evidence (t+9 min, per #55 HEALTHY checklist)

Signal Reading
S1 stdout teed to campaign_stdout.log; 0 crash/abort banners
S4 round_counter 1890, +460 in 9 min
S6 sac_latest.zip mtime age 6 s (first save at t+54 s)
S2 training_log.jsonl 1492 lines, growing, plausible ticks
S3 eval_log.jsonl fresh atomic replace every 2 chunks vs Corners
S5 best_score 20% (genuine eval best; non-decreasing)
Sampling 112 chunks ≈ 37/22/13/14/13 % vs 3:2:1:1:1 weights ✓
Disk 418 GB free

Early 0% evals = near-random policy minutes in; SLOW-LEARNER watch rule applies, never intervene.

Launch incident worth the record (fixed in 2653671)

Two launch attempts persisted zero checkpoints: at production sizes the save check ran only between drain-burst passes, so stepCount never crossed nextSave within a process lifetime (~10–18 steps at ~1 s/step vs interval 20). Smoke runs masked it because hidden-32 steps were sub-millisecond. Fix: check moved inside the gradient-step loop; validated (interval=5 fires ~12 s into a battle). Full chain in the notebook's Incident 1.

Paths

  • Notebook (config table, phase log, incidents, check-in procedure): SAC_LSTM_Bot/docs/campaign_notebook.md
  • Persisted harness stdout: SAC_LSTM_Bot/campaign_stdout.log · training JSONL: SAC_LSTM_Bot/training_log.jsonl · eval JSONL: SAC_LSTM_Bot/eval_log.jsonl
  • Checkpoints: SAC_LSTM_Bot/weights/{sac_latest.zip,sac_best.zip,best_score.txt} · twin: $SAMPLE_BOTS_DIR/SacTwin/ (frozen)

Glanceable check-in (<1 min, for future monitor sessions)

tmux capture-pane -p -t sac_campaign | tail -5        # banners/crashes
cat  ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip   # <600s old = learning alive
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log

Intervene on: ≥3 consecutive crash banners · counter frozen ≥15 min · zip >10 min stale while counter advances · disk <1 GB. Kill at ceiling. Full procedure + thresholds in the notebook.

Closing #56; #57 unblocks.

## Resolution — campaign-v1 locked, launched, HEALTHY (commits `2f49cb2`, `2653671`, `4b64bf1`) ### Final locked config (verbatim, copy-pasteable) ```bash cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \ SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \ SAC_EVAL_OPPONENT=Corners \ SAC_TOTAL_ROUNDS=25000 \ SAC_CHUNK_SIZE=10 \ SAC_EVAL_INTERVAL=2 \ SAC_EVAL_ROUNDS=10 \ SAC_MAX_CRASHES=5 \ SACLSTM_HIDDEN_SIZE=256 \ SACLSTM_BATCH_SIZE=16 \ SACLSTM_UTD_RATIO=1 \ SACLSTM_SAVE_INTERVAL=5 \ ./sac_train.sh 2>&1 | tee -a campaign_stdout.log ``` Rationale per knob in the notebook config table. Deviations from module defaults / prior recommendations, one line each: - `SAVE_INTERVAL=5` (not 500-default, not #54's 10–20): at hidden 256 a step costs ~1 s and a chunk process fits only ~10–18 steps ⇒ the interval must sit inside the per-process step budget; 5 ⇒ checkpoint every ~5–10 s of active training. Enabled by commit `2653671`. - `TOTAL_ROUNDS=25000` (map implied smaller): measured throughput ~3400 rounds/h means 2000 was a 1-hour budget; ceiling must bind first. - Twin at weight **1** (map suggested 2): #54 showed mirror battles end early ⇒ fewer transitions/chunk; weight 2 starves the buffer. - All other bot knobs are unpinned module defaults (`3e-4` LRs, γ=0.99, τ=0.005, target entropy −4.0, buffer 500k, burn-in 8, window 16). ### Launch - tmux session **`sac_campaign`**, launched **2026-08-22 00:21:55 CEST** (epoch 1787350915) from `research/goto-controller` @ `2653671`. - **Wall-clock ceiling T+12 h ⇒ 12:22 CEST Aug 22 (epoch 1787394115)**: a monitor kills the session then if still running. Ceiling is a stop, not a deadline. - Twin regenerated pre-launch via `./make_twin.sh`, seed md5-verified (`61521cff…`) from the campaign-v1 baseline (fresh random-init hidden-256 checkpoint; old 32-hidden smoke weights archived to `weights_smoke49_backup/`). ### First-health-check evidence (t+9 min, per #55 HEALTHY checklist) | Signal | Reading | |--------|---------| | S1 stdout | teed to `campaign_stdout.log`; **0** crash/abort banners | | S4 round_counter | 1890, +460 in 9 min | | S6 sac_latest.zip | mtime age **6 s** (first save at t+54 s) | | S2 training_log.jsonl | 1492 lines, growing, plausible ticks | | S3 eval_log.jsonl | fresh atomic replace every 2 chunks vs Corners | | S5 best_score | 20% (genuine eval best; non-decreasing) | | Sampling | 112 chunks ≈ 37/22/13/14/13 % vs 3:2:1:1:1 weights ✓ | | Disk | 418 GB free | Early 0% evals = near-random policy minutes in; SLOW-LEARNER watch rule applies, never intervene. ### Launch incident worth the record (fixed in `2653671`) Two launch attempts persisted **zero checkpoints**: at production sizes the save check ran only between drain-burst passes, so `stepCount` never crossed `nextSave` within a process lifetime (~10–18 steps at ~1 s/step vs interval 20). Smoke runs masked it because hidden-32 steps were sub-millisecond. Fix: check moved inside the gradient-step loop; validated (interval=5 fires ~12 s into a battle). Full chain in the notebook's *Incident 1*. ### Paths - Notebook (config table, phase log, incidents, check-in procedure): `SAC_LSTM_Bot/docs/campaign_notebook.md` - Persisted harness stdout: `SAC_LSTM_Bot/campaign_stdout.log` · training JSONL: `SAC_LSTM_Bot/training_log.jsonl` · eval JSONL: `SAC_LSTM_Bot/eval_log.jsonl` - Checkpoints: `SAC_LSTM_Bot/weights/{sac_latest.zip,sac_best.zip,best_score.txt}` · twin: `$SAMPLE_BOTS_DIR/SacTwin/` (frozen) ### Glanceable check-in (<1 min, for future monitor sessions) ```bash tmux capture-pane -p -t sac_campaign | tail -5 # banners/crashes cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip # <600s old = learning alive tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log ``` Intervene on: ≥3 consecutive crash banners · counter frozen ≥15 min · zip >10 min stale while counter advances · disk <1 GB. Kill at ceiling. Full procedure + thresholds in the notebook. Closing #56; #57 unblocks.
Author
Owner

Ceiling extended by orchestrator decision (human mandate: no deadlines / let the budget complete): old kill = procedural monitor instruction only (campaign_notebook.md lines 41/110, epoch 1787394115 / 12:22 CEST — sweep found NO os-level watchdog: tmux panes, ps//proc cmdline scan for the epoch, at/cron/systemd timers all clean; sac_train.sh & RunTraining.java have no wall-clock ceiling), inerted by rewriting both instructions to the new ceiling. New safety net: hard transient user-systemd unit sac-ceiling-net armed 07:46:04 CEST, fires 16:30:00 CEST today (epoch 1787409000) — tmux kill-session -t sac_campaign then pgrep-guarded pkill of straggler harness processes, same style as the original stop. Training uninterrupted — counter verified advancing: 21253 → 21352 (+99) over 147 s, sac_latest.zip fresh ≤12 s, training_log.jsonl +69 lines, tmux session alive.

Ceiling extended by orchestrator decision (human mandate: no deadlines / let the budget complete): old kill = procedural monitor instruction only (campaign_notebook.md lines 41/110, epoch 1787394115 / 12:22 CEST — sweep found NO os-level watchdog: tmux panes, ps//proc cmdline scan for the epoch, at/cron/systemd timers all clean; sac_train.sh & RunTraining.java have no wall-clock ceiling), inerted by rewriting both instructions to the new ceiling. New safety net: hard transient user-systemd unit `sac-ceiling-net` armed 07:46:04 CEST, fires **16:30:00 CEST today (epoch 1787409000)** — `tmux kill-session -t sac_campaign` then pgrep-guarded pkill of straggler harness processes, same style as the original stop. Training uninterrupted — counter verified advancing: 21253 → 21352 (+99) over 147 s, sac_latest.zip fresh ≤12 s, training_log.jsonl +69 lines, tmux session alive.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#56