Files
SirRoboGarage/SAC_LSTM_Bot/docs/campaign_notebook.md
T

73 lines
4.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
Overnight training campaign on branch `research/goto-controller`.
Companion tickets: config+launch = **#56**, morning verdict = **#57**, map = **#53**.
Monitoring contract: observability inventory in **#55 comment 461** (9 signals, thresholds, four-way discrimination).
## Locked config (campaign-v1)
Launch command (tmux session `sac_campaign`, stdout teed to `campaign_stdout.log`):
```bash
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=2000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=20 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
```
| Knob | Value | Source / rationale |
|------|-------|--------------------|
| `SAC_OPPONENTS` | `Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1` | Working recommendation kept. Twin pinned at **1 not 2**: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer. |
| `SAC_EVAL_OPPONENT` | `Corners` | Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun. |
| `SAC_TOTAL_ROUNDS` | `2000` | Sized so the **wall-clock ceiling binds first** (~100–150 rounds/h observed incl. evals ⇒ ~13–20 h of headroom). |
| `SAC_CHUNK_SIZE` | `10` | Harness default. |
| `SAC_EVAL_INTERVAL` | `2` | Harness default — eval every ~20 rounds ≈ every ~10 min. |
| `SAC_EVAL_ROUNDS` | `10` | Harness default. |
| `SAC_MAX_CRASHES` | `5` | Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55). |
| `SACLSTM_HIDDEN_SIZE` | `256` | Module default (`network.nim`); real capacity vs #49 smoke's 32; under `MaxHidden`=512 cap. |
| `SACLSTM_BATCH_SIZE` | `16` | Module default (`integration.nim`). |
| `SACLSTM_UTD_RATIO` | `1` | Module default. |
| `SACLSTM_SAVE_INTERVAL` | `20` | **Deviation** from default 500: #54 proved the shutdown save never reaches disk in harness context — only mid-battle interval saves persist; 10–20 fired reliably, 50 did not in twin chunks. 20 = top of proven range, least IO. |
| *(not pinned)* | module defaults | `LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4`, `GAMMA=0.99`, `TAU=0.005`, `TARGET_ENTROPY=-4.0`, `BUFFER_CAPACITY=500000`, `BURN_IN=8`, `TRAIN_WINDOW=16`. |
**Budget**: generous wall-clock **ceiling, not a deadline** — **T+12 h** from launch. While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.
**Fresh start**: pre-campaign `weights/` held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to `weights_smoke49_backup/`; baseline re-established by a 10-round bootstrap battle vs Corners (also created the seed for `make_twin.sh`).
## Phase log
- **2026-08-21 ~23:30** — Config locked (this file), claim posted on #56 (comment 471). Smoke weights archived.
- **2026-08-21 ~23:4x** — Bootstrap battle vs Corners done; baseline checkpoint + `best_score.txt` written; twin regenerated via `./make_twin.sh` from the campaign-v1 baseline.
- **2026-08-21 ~23:5x** — Launched tmux `sac_campaign`. First-health-check: see *Incidents & checks* below.
## Decision-issue index
| Issue | What it decided |
|-------|-----------------|
| #37–#48 | Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration. |
| #49 | Training harness + smoke run (toy hyperparams: hidden 32). |
| #54 | Mirror-twin sparring partner; SAVE_INTERVAL ≤20 rule; eval opponent = first pool entry. |
| #55 | 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds. |
| #56 | This campaign: locked config + launch (this notebook). |
| #57 | Morning verdict — consumes this notebook + logs. |
## Incidents & checks
*(appended during the run)*
### Launch health check (first chunks)
Pending — filled right after launch.
## Results
*(filled by #57 / morning session)*