9.3 KiB
Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
Overnight training campaign on branch research/goto-controller.
Companion tickets: config+launch = #56, morning verdict = #57, map = #53.
Monitoring contract: observability inventory in #55 comment 461 (9 signals, thresholds, four-way discrimination).
Locked config (campaign-v1)
Launch command (tmux session sac_campaign, stdout teed to campaign_stdout.log):
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=25000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
| Knob | Value | Source / rationale |
|---|---|---|
SAC_OPPONENTS |
Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1 |
Working recommendation kept. Twin pinned at 1 not 2: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer. |
SAC_EVAL_OPPONENT |
Corners |
Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun. |
SAC_TOTAL_ROUNDS |
25000 |
Sized so the wall-clock ceiling binds first: measured throughput ~3400 rounds/h early (drops as battles lengthen) ⇒ 2000 would have exhausted in ~1 h. |
SAC_CHUNK_SIZE |
10 |
Harness default. |
SAC_EVAL_INTERVAL |
2 |
Harness default — eval every ~20 rounds. |
SAC_EVAL_ROUNDS |
10 |
Harness default. |
SAC_MAX_CRASHES |
5 |
Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55). |
SACLSTM_HIDDEN_SIZE |
256 |
Module default (network.nim); real capacity vs #49 smoke's 32; under MaxHidden=512 cap. |
SACLSTM_BATCH_SIZE |
16 |
Module default (integration.nim). |
SACLSTM_UTD_RATIO |
1 |
Module default. |
SACLSTM_SAVE_INTERVAL |
5 |
Deviation from default 500 and from #54's "10–20": at hidden 256 a gradient step takes ~1 s and a chunk process fits only ~10–18 steps (see incident below) — interval must sit inside the per-process step budget. 5 ⇒ checkpoint every ~5–10 s of active training; IO trivial (10.5 MB zip, atomic replace). |
| (not pinned) | module defaults | LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4, GAMMA=0.99, TAU=0.005, TARGET_ENTROPY=-4.0, BUFFER_CAPACITY=500000, BURN_IN=8, TRAIN_WINDOW=16. |
Budget: generous wall-clock ceiling, not a deadline — T+12 h from launch 00:21:55 CEST 2026-08-22 ⇒ ceiling 12:22 CEST 2026-08-22 (epoch 1787394115). While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.
Fresh start: pre-campaign weights/ held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to weights_smoke49_backup/; campaign baseline re-established by probe battles vs Corners (random-init hidden-256 checkpoint, best_score.txt reset then re-raised to 20 by a genuine eval). Twin regenerated via ./make_twin.sh from that baseline (md5 61521cff… verified seed).
Phase log
- 2026-08-21 23:22 — Claim posted on #56 (comment 471). Config locked, notebook committed (
2f49cb2). - 2026-08-21 23:27 — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (
sac_latest.zip, 10.5 MB) and twin seed. - 2026-08-21 23:32 — Twin regenerated (md5-verified). Launch attempt 1 (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but zero checkpoints persisted (see incident). Killed 23:48.
- 2026-08-21 23:52–00:10 — Diagnosis (see incident): interval=1 fired, interval=2/20 never; instrumentation + /proc thread forensics ⇒ per-process step budget ~10–18 at ~1 s/step; save check ran only between drain-burst passes.
- 2026-08-22 00:12 — Fix: save check moved inside the gradient-step loop, committed
2653671. Validated: interval=5 save fired ~12 s into a battle. - 2026-08-22 00:21:55 — Launch (final): tmux
sac_campaign, config above. First campaign save on disk at t+54 s; eval #1 on cadence. - 2026-08-22 00:30 — HEALTHY checklist passed (see below).
- 2026-08-22 03:45 — Watch shift 1 (00:34–03:35): liveness flawless (10/10 HEALTHY, zero banners, zip ≤30 s). Learning signal: steady-state eval vs Corners 0–5% with two isolated 10/10 spikes (~01:15) → capability emerged, then lost. Eval-regression intervene threshold fired per #55; intervention DEFERRED to Campaign verdict (#57) — rationale: n=2 evidence, no loss metrics, buffer-loss on restart, run completes ~08:05 anyway. Milestone issue: #58 "Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict".
- 2026-08-22 ~07:15 — Morning audit: policy demonstrably learning off-benchmark (SacTwin 73→100%, Crazy 26→56%) while Corners eval stays ~0–9% with 3 transient 10/10s; wall-clock ceiling extended to let the 25k complete (~14:35 projected); instability-vs-plateau question left to the curve.
Decision-issue index
| Issue | What it decided |
|---|---|
| #37–#48 | Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration. |
| #49 | Training harness + smoke run (toy hyperparams: hidden 32). |
| #54 | Mirror-twin sparring partner; SAVE_INTERVAL persistence rule; eval opponent = first pool entry. |
| #55 | 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds. |
| #56 | This campaign: locked config + launch + the save-check fix (2653671). |
| #57 | Morning verdict — consumes this notebook + logs. |
Incidents & checks
Incident 1 — zero checkpoint persistence at production sizes (launch blockers, fixed)
Symptom: campaign ran 26+ chunk processes across two attempts without a single sac_latest.zip update, while rounds/evals flowed normally. #54's rule ("keep SACLSTM_SAVE_INTERVAL well below per-chunk gradient-step counts, 10–20 fired in smokes") silently broke at hidden 256.
Diagnosis chain (all reproducible):
- Interval=1 saved within seconds; interval=2 and 20 never saved — through the same harness ⇒ not env propagation.
- Temporary step instrumentation (bot stderr via a one-line
SAC_LSTM_Bot.shredirect — the vendored runner swallows bot stderr, #55 gap S9-adjacent): steps cost ~1.06 s each; a drain burst queued 53 steps; logging stopped mid-pass while rounds kept completing. /proc/<pid>/tasksampling: training thread alive and RUNNING (~13 s CPU per ~40 s process) — not deadlocked, just slow ⇒ per-process step budget ≈ 10–18 steps.- The save check lived between drain-burst passes; with bursts queueing minutes of steps,
stepCountnever reachednextSavebefore process teardown. Smoke runs masked this: hidden 32 steps were sub-millisecond, so hundreds of steps fit per chunk.
Fix (commit 2653671): save check relocated inside the step loop (checked every gradient step; packFull+trySend unchanged). Validated: interval=5 save fires ~12 s into a battle; campaign save fired 54 s after launch.
Config consequences: SACLSTM_SAVE_INTERVAL=5 (inside the per-process budget; #54's 10–20 was derived at smoke speeds). SAC_TOTAL_ROUNDS=25000 (throughput measured ~3400 rounds/h, so 2000 was a 1-hour budget, not an overnight one). OMP_NUM_THREADS=1 tested and not needed (hang was step-budget exhaustion, not OpenMP).
Launch health check (t+9 min, 00:30:16) — HEALTHY per #55 checklist
| Signal | Reading | Verdict |
|---|---|---|
| S1 harness stdout | teed to campaign_stdout.log; 0 crash/aborted banners |
✓ |
| S4 round_counter | 1890, +460 in 9 min (~51 rounds/min) | ✓ advancing |
| S6 sac_latest.zip mtime | 6 s old; first save at t+54 s | ✓ fresh |
| S2 training_log.jsonl | 1492 lines, growing; last ticks=668, plausible | ✓ |
| S3 eval_log.jsonl | age 3 s (atomic replace); eval every 2 chunks | ✓ on cadence |
| S5 best_score | 20 (from a genuine campaign-1 eval; non-decreasing) | ✓ |
| Sampling | 112 chunks: Corners 41 / Crazy 25 / RamFire 15 / SacTwin 16 / Target 15 ≈ weights 3:2:1:1:1 | ✓ plausible |
| Disk | 418 GB free | ✓ |
Early evals 0% vs Corners — expected for a near-random policy minutes in; SLOW-LEARNER watch rule (flat ≥5 evals = watch) applies, never intervene.
Check-in procedure (for monitor sessions)
tmux capture-pane -p -t sac_campaign | tail -5 # S1: banners, crashes
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip # age <~600s = training alive
tail -1 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/training_log.jsonl
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log # eval results / new best
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/best_score.txt
Intervene per #55 thresholds: ≥3 consecutive crash #N banners; ΔS4=0 over ≥15 min; zip mtime >10 min stale while S4 advances (STALLED); disk <1 GB. At the ceiling (12:22 CEST Aug 22): tmux kill-session -t sac_campaign if still running — final state is in weights/, logs, and this notebook.
Results
(filled by #57 / morning session)