4.0 KiB
Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
Overnight training campaign on branch research/goto-controller.
Companion tickets: config+launch = #56, morning verdict = #57, map = #53.
Monitoring contract: observability inventory in #55 comment 461 (9 signals, thresholds, four-way discrimination).
Locked config (campaign-v1)
Launch command (tmux session sac_campaign, stdout teed to campaign_stdout.log):
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=2000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=20 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
| Knob | Value | Source / rationale |
|---|---|---|
SAC_OPPONENTS |
Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1 |
Working recommendation kept. Twin pinned at 1 not 2: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer. |
SAC_EVAL_OPPONENT |
Corners |
Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun. |
SAC_TOTAL_ROUNDS |
2000 |
Sized so the wall-clock ceiling binds first (~100–150 rounds/h observed incl. evals ⇒ ~13–20 h of headroom). |
SAC_CHUNK_SIZE |
10 |
Harness default. |
SAC_EVAL_INTERVAL |
2 |
Harness default — eval every ~20 rounds ≈ every ~10 min. |
SAC_EVAL_ROUNDS |
10 |
Harness default. |
SAC_MAX_CRASHES |
5 |
Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55). |
SACLSTM_HIDDEN_SIZE |
256 |
Module default (network.nim); real capacity vs #49 smoke's 32; under MaxHidden=512 cap. |
SACLSTM_BATCH_SIZE |
16 |
Module default (integration.nim). |
SACLSTM_UTD_RATIO |
1 |
Module default. |
SACLSTM_SAVE_INTERVAL |
20 |
Deviation from default 500: #54 proved the shutdown save never reaches disk in harness context — only mid-battle interval saves persist; 10–20 fired reliably, 50 did not in twin chunks. 20 = top of proven range, least IO. |
| (not pinned) | module defaults | LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4, GAMMA=0.99, TAU=0.005, TARGET_ENTROPY=-4.0, BUFFER_CAPACITY=500000, BURN_IN=8, TRAIN_WINDOW=16. |
Budget: generous wall-clock ceiling, not a deadline — T+12 h from launch. While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.
Fresh start: pre-campaign weights/ held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to weights_smoke49_backup/; baseline re-established by a 10-round bootstrap battle vs Corners (also created the seed for make_twin.sh).
Phase log
- 2026-08-21 ~23:30 — Config locked (this file), claim posted on #56 (comment 471). Smoke weights archived.
- 2026-08-21 ~23:4x — Bootstrap battle vs Corners done; baseline checkpoint +
best_score.txtwritten; twin regenerated via./make_twin.shfrom the campaign-v1 baseline. - 2026-08-21 ~23:5x — Launched tmux
sac_campaign. First-health-check: see Incidents & checks below.
Decision-issue index
| Issue | What it decided |
|---|---|
| #37–#48 | Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration. |
| #49 | Training harness + smoke run (toy hyperparams: hidden 32). |
| #54 | Mirror-twin sparring partner; SAVE_INTERVAL ≤20 rule; eval opponent = first pool entry. |
| #55 | 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds. |
| #56 | This campaign: locked config + launch (this notebook). |
| #57 | Morning verdict — consumes this notebook + logs. |
Incidents & checks
(appended during the run)
Launch health check (first chunks)
Pending — filled right after launch.
Results
(filled by #57 / morning session)