Files
SirRoboGarage/SAC_LSTM_Bot/docs/campaign_notebook.md
T

20 KiB
Raw Blame History

Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)

Overnight training campaign on branch research/goto-controller. Companion tickets: config+launch = #56, morning verdict = #57, map = #53. Monitoring contract: observability inventory in #55 comment 461 (9 signals, thresholds, four-way discrimination).

Locked config (campaign-v1)

Launch command (tmux session sac_campaign, stdout teed to campaign_stdout.log):

SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=25000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
Knob Value Source / rationale
SAC_OPPONENTS Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1 Working recommendation kept. Twin pinned at 1 not 2: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer.
SAC_EVAL_OPPONENT Corners Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun.
SAC_TOTAL_ROUNDS 25000 Sized so the wall-clock ceiling binds first: measured throughput ~3400 rounds/h early (drops as battles lengthen) ⇒ 2000 would have exhausted in ~1 h.
SAC_CHUNK_SIZE 10 Harness default.
SAC_EVAL_INTERVAL 2 Harness default — eval every ~20 rounds.
SAC_EVAL_ROUNDS 10 Harness default.
SAC_MAX_CRASHES 5 Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55).
SACLSTM_HIDDEN_SIZE 256 Module default (network.nim); real capacity vs #49 smoke's 32; under MaxHidden=512 cap.
SACLSTM_BATCH_SIZE 16 Module default (integration.nim).
SACLSTM_UTD_RATIO 1 Module default.
SACLSTM_SAVE_INTERVAL 5 Deviation from default 500 and from #54's "10–20": at hidden 256 a gradient step takes ~1 s and a chunk process fits only ~10–18 steps (see incident below) — interval must sit inside the per-process step budget. 5 ⇒ checkpoint every ~5–10 s of active training; IO trivial (10.5 MB zip, atomic replace).
(not pinned) module defaults LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4, GAMMA=0.99, TAU=0.005, TARGET_ENTROPY=-4.0, BUFFER_CAPACITY=500000, BURN_IN=8, TRAIN_WINDOW=16.

Budget: generous wall-clock ceiling, not a deadline — originally T+12 h from launch 00:21:55 CEST 2026-08-22 ⇒ 12:22 CEST (epoch 1787394115); extended 2026-08-22 ~07:35 by orchestrator decision on human mandate ("no deadlines — let the 25000-round budget complete", ~14:35 projected) ⇒ ceiling now 16:30 CEST 2026-08-22 (epoch 1787409000), enforced by a hard user-systemd net unit sac-ceiling-net (sleeps to the epoch, then kills the tmux session and any straggler harness processes). While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.

Fresh start: pre-campaign weights/ held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to weights_smoke49_backup/; campaign baseline re-established by probe battles vs Corners (random-init hidden-256 checkpoint, best_score.txt reset then re-raised to 20 by a genuine eval). Twin regenerated via ./make_twin.sh from that baseline (md5 61521cff… verified seed).

Phase log

  • 2026-08-21 23:22 — Claim posted on #56 (comment 471). Config locked, notebook committed (2f49cb2).
  • 2026-08-21 23:27 — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (sac_latest.zip, 10.5 MB) and twin seed.
  • 2026-08-21 23:32 — Twin regenerated (md5-verified). Launch attempt 1 (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but zero checkpoints persisted (see incident). Killed 23:48.
  • 2026-08-21 23:52–00:10 — Diagnosis (see incident): interval=1 fired, interval=2/20 never; instrumentation + /proc thread forensics ⇒ per-process step budget ~10–18 at ~1 s/step; save check ran only between drain-burst passes.
  • 2026-08-22 00:12 — Fix: save check moved inside the gradient-step loop, committed 2653671. Validated: interval=5 save fired ~12 s into a battle.
  • 2026-08-22 00:21:55 — Launch (final): tmux sac_campaign, config above. First campaign save on disk at t+54 s; eval #1 on cadence.
  • 2026-08-22 00:30 — HEALTHY checklist passed (see below).
  • 2026-08-22 03:45 — Watch shift 1 (00:34–03:35): liveness flawless (10/10 HEALTHY, zero banners, zip ≤30 s). Learning signal: steady-state eval vs Corners 0–5% with two isolated 10/10 spikes (~01:15) → capability emerged, then lost. Eval-regression intervene threshold fired per #55; intervention DEFERRED to Campaign verdict (#57) — rationale: n=2 evidence, no loss metrics, buffer-loss on restart, run completes ~08:05 anyway. Milestone issue: #58 "Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict".
  • 2026-08-22 ~07:15 — Morning audit: policy demonstrably learning off-benchmark (SacTwin 73→100%, Crazy 26→56%) while Corners eval stays ~0–9% with 3 transient 10/10s; wall-clock ceiling extended to let the 25k complete (~14:35 projected); instability-vs-plateau question left to the curve.
  • 2026-08-22 ~07:50 — score:60 anatomy: Tank Royale survival(50)+last-survivor(10) awarded when opponent dies while we survive; exactly-60 ⇒ zero damage dealt by us that round (opponent self-destructed via wasted shots + 0.1/turn inactivity drain). 3,878 rounds (19%); modal vs SacTwin; vs Corners 986 damageless outlives vs 242 true wins; combined with 43% of rounds being score:0, texture = survivor-not-fighter against walls. RL rewards are event-driven (rewards module), so behavioral evidence, not reward poisoning. Feeds #57 levers: aggression shaping / specialist-vs-generalist.
  • 2026-08-22 ~08:05 — .part debris forensics + sweep: 99 *.zip.tmp.*.part (483 MB) are NOT weights.nim debris (that proc uses fixed .tmp + finally-cleanup, working); naming matches an external write-temp→rename copier killed mid-write, bursts correlating with kill events; possible culprit: a folder-sync client fighting a file that changes every ~20 s (human asked to confirm). Swept with -mmin +10 age guard (protects in-flight writes); 483 MB freed, real zips untouched — note 2 fresh .part reappeared minutes later, copier still active.
  • 2026-08-22 ~08:00 — ceiling defused: the 12:22 'self-kill' was notebook prose instructing watchmen — never an OS mechanism; rewritten to 16:30 CEST AND armed a real systemd --user net sac-ceiling-net firing 16:30:00 (epoch 1787409000); training uninterrupted (counter +99/147 s verified); annotated #56 comment 487.
  • 2026-08-22 ~14:28 — CAMPAIGN ENDED NATURALLY: banner >>> training complete: 2500 chunks after 14 h 07 m (00:21:55 → ~14:28 CEST); round_counter 38912; zero crash banners across the whole run; the sac-ceiling-net backstop never fired — cancelled unneeded. Budget note: the harness loop is chunk-based — SAC_TOTAL_ROUNDS=25000 ÷ CHUNK_SIZE=10 ⇒ 2500 chunk battles of ≤10 rounds each; the "25k-rounds" label was a misnomer (the counter also accrues 1287×10 eval rounds and rerun chunks). Final eval vs Corners: 10%.
  • 2026-08-22 ~14:50 — VERDICT posted (→ #57): RETUNE BEFORE SCALING. Ops layer PROVEN (14 h autonomous, zero crashes, self-healing restarts, natural completion — the harness scales); learning REAL BUT NARROW (within-opponent gains genuine — SacTwin 90.7%, Crazy 26→48% — but specialist-not-generalist, walls untouched; 12 eval spikes ≥8/10 incl. 5×10/10, none retained); benchmark pathology: Corners-only deterministic eval + single-max best gating froze sac_best.zip at 01:01:57 on a fluke 10/10. Five code-level levers staged awaiting human sign-off; v2 NOT launched. Full rationale: Results below + #57 resolution comment.
  • 2026-08-22 ~14:50 — HYGIENE (campaign over, no live writers — safe): sac-ceiling-net stopped + reset-failed (backstop obsolete); .part corpse sweep 48 → 0 (no age guard needed — nothing writes anymore); future-run guard added to sac_train.sh: startup rm -f "$WEIGHTS_DIR"/sac_latest.zip.tmp.*.part "$WEIGHTS_DIR"/sac_latest.zip.tmp so a SIGKILLed run's libzip modify-path corpses can't accumulate again.

Decision-issue index

Issue What it decided
#37–#48 Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration.
#49 Training harness + smoke run (toy hyperparams: hidden 32).
#54 Mirror-twin sparring partner; SAVE_INTERVAL persistence rule; eval opponent = first pool entry.
#55 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds.
#56 This campaign: locked config + launch + the save-check fix (2653671).
#57 Morning verdict — consumes this notebook + logs.

Incidents & checks

Incident 1 — zero checkpoint persistence at production sizes (launch blockers, fixed)

Symptom: campaign ran 26+ chunk processes across two attempts without a single sac_latest.zip update, while rounds/evals flowed normally. #54's rule ("keep SACLSTM_SAVE_INTERVAL well below per-chunk gradient-step counts, 10–20 fired in smokes") silently broke at hidden 256.

Diagnosis chain (all reproducible):

  1. Interval=1 saved within seconds; interval=2 and 20 never saved — through the same harness ⇒ not env propagation.
  2. Temporary step instrumentation (bot stderr via a one-line SAC_LSTM_Bot.sh redirect — the vendored runner swallows bot stderr, #55 gap S9-adjacent): steps cost ~1.06 s each; a drain burst queued 53 steps; logging stopped mid-pass while rounds kept completing.
  3. /proc/<pid>/task sampling: training thread alive and RUNNING (~13 s CPU per ~40 s process) — not deadlocked, just slow ⇒ per-process step budget ≈ 10–18 steps.
  4. The save check lived between drain-burst passes; with bursts queueing minutes of steps, stepCount never reached nextSave before process teardown. Smoke runs masked this: hidden 32 steps were sub-millisecond, so hundreds of steps fit per chunk.

Fix (commit 2653671): save check relocated inside the step loop (checked every gradient step; packFull+trySend unchanged). Validated: interval=5 save fires ~12 s into a battle; campaign save fired 54 s after launch.

Config consequences: SACLSTM_SAVE_INTERVAL=5 (inside the per-process budget; #54's 10–20 was derived at smoke speeds). SAC_TOTAL_ROUNDS=25000 (throughput measured ~3400 rounds/h, so 2000 was a 1-hour budget, not an overnight one). OMP_NUM_THREADS=1 tested and not needed (hang was step-budget exhaustion, not OpenMP).

Launch health check (t+9 min, 00:30:16) — HEALTHY per #55 checklist

Signal Reading Verdict
S1 harness stdout teed to campaign_stdout.log; 0 crash/aborted banners ✓
S4 round_counter 1890, +460 in 9 min (~51 rounds/min) ✓ advancing
S6 sac_latest.zip mtime 6 s old; first save at t+54 s ✓ fresh
S2 training_log.jsonl 1492 lines, growing; last ticks=668, plausible ✓
S3 eval_log.jsonl age 3 s (atomic replace); eval every 2 chunks ✓ on cadence
S5 best_score 20 (from a genuine campaign-1 eval; non-decreasing) ✓
Sampling 112 chunks: Corners 41 / Crazy 25 / RamFire 15 / SacTwin 16 / Target 15 ≈ weights 3:2:1:1:1 ✓ plausible
Disk 418 GB free ✓

Early evals 0% vs Corners — expected for a near-random policy minutes in; SLOW-LEARNER watch rule (flat ≥5 evals = watch) applies, never intervene.

Check-in procedure (for monitor sessions)

tmux capture-pane -p -t sac_campaign | tail -5        # S1: banners, crashes
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip   # age <~600s = training alive
tail -1 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/training_log.jsonl
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log           # eval results / new best
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/best_score.txt

Intervene per #55 thresholds: ≥3 consecutive crash #N banners; ΔS4=0 over ≥15 min; zip mtime >10 min stale while S4 advances (STALLED); disk <1 GB. At the ceiling (16:30 CEST Aug 22, epoch 1787409000 — extended from 12:22 per human mandate): tmux kill-session -t sac_campaign if still running (the hard net unit sac-ceiling-net fires at the same epoch regardless of monitors) — final state is in weights/, logs, and this notebook.

Results (campaign-v1 — filled by #57)

Run: 2500/2500 chunks · 14 h 07 m autonomous (00:21:55 → ~14:28 CEST 2026-08-22) · round_counter 38912 · zero crashes · graceful banner >>> training complete: 2500 chunks · systemd net never fired. Budget was chunk-based (see phase log) — the "25k-rounds" label was a misnomer.

Per-opponent training win rates (run-3 slice of training_log.jsonl):

Opponent Win rate Record Note
SacTwin 90.7% 2693/2970 vs frozen past-self — genuine self-play gain
Crazy 48.3% 2965/6140 doubled from 26% early-run
Target 9.6% — static, barely moved
Corners 7.2% — walls untouched
RamFire 0% 0/2920 mirrors the PPO-era ladder — ram-class needs dedicated pressure

Eval vs Corners (pinned benchmark): 1287 evals · overall mean 7.7% · histogram headline: 0/10 = 867 (67%), spikes ≥8/10 = 12 (incl. 5× perfect 10/10) · final eval 10%. Stdout log carries no timestamps; timing reconstructed from file mtimes.

Best-zip paradox: sac_best.zip frozen since 01:01:57 — a single lucky 10/10 at ~round 3.5k wrote best_score=100, and no later eval could outrank a perfect score (even genuine ~50%-winrate stretches elsewhere). Best checkpoint = lottery ticket, decoupled from the steady-state policy (which sat at 0–10% vs Corners).

Verdict: RETUNE BEFORE SCALING — full rationale in #57 resolution comment. Five code-level levers staged for human sign-off: (1) eval rotation across pool + moving-average best gating; (2) reward shaping toward damage/aggression incl. anti-ram signal; (3) training-loss/step metrics logged from the training thread (#55 gap #1); (4) gate sendTrainingMsg off in eval mode (#55 gap #3); (5) optional stability knobs (lower LR / entropy coeff) once loss curves exist.


Campaign Notebook — campaign-v2 (SAC_LSTM_Bot)

Locked config (campaign-v2)

Launched verbatim from orchestrator mandate (umbrella ticket #59, all five levers approved & implemented in a07e530 + 6fc01eb):

cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
Knob Value Rationale
SAC_OPPONENTS Corners:3, Crazy:2, RamFire:2, Target:1, SacTwin:1 RamFire bumped 1→2 vs v1: anti-ram shaping (lever 2, 6fc01eb) needs exposure to fire; without samples there is no gradient signal against ram-class
SAC_EVAL_OPPONENTS Corners,Crazy,Target Lever-1 rotation set (default); composite = mean of per-opponent MA-5 win rates
SAC_EVAL_INTERVAL/ROUNDS 2 / 10 Unchanged from v1 cadence
SAC_TOTAL_ROUNDS/CHUNK_SIZE/MAX_CRASHES 25000 / 10 / 5 Same budget semantics as v1 (chunk-based)
SACLSTM_HIDDEN_SIZE/BATCH_SIZE/SAVE_INTERVAL 256 / 16 / 5 Architecture + throughput knobs carried over; LR/entropy defaults untouched = lever-5 conditional posture (activation decided by loss-curve evidence, not upfront)

Fresh start & archive

v1 state archived intact (notebook references preserved) into SAC_LSTM_Bot/weights_v1_archive/: sac_latest.zip, sac_best.zip, best_score.txt, round_counter.txt, training_log.jsonl, eval_log.jsonl, campaign_stdout.log, plus the lever-3 smoke leftover training_metrics.jsonl (from src/SAC_LSTM_Bot/, moved so v2 loss curves start clean for lever-5 reading). Main weights/ verified empty afterwards ⇒ main bot takes the genuine random-init path (loadOrInitFull → randomFull()).

Twin reseed: make_twin.sh requires a seed zip, but the fresh-start baseline has none. Generated a fresh random-init checkpoint (hidden=256, alpha=1.0 matching logAlpha=0) via a throwaway Nim script against network.nim/weights.nim, seeded it as sac_best.zip transiently, ran ./make_twin.sh, removed the transient copy. Verified twin dir got byte-identical fresh zips (cmp OK; NOT v1 zips) + round_counter.txt=0. No script changes needed.

Safety net

systemd-run --user --unit=sac-ceiling-net-v2 armed at launch: sleeps 72000 s then tmux kill-session -t sac_campaign_v2; sleep 5; pkill -f sac_train.sh. Unit active at 19:59:40 CEST 2026-08-22, fires 15:59:40 CEST 2026-08-23 (epoch 1787493580).

Launch & health evidence (first ~45 min)

Launched 20:00:19 CEST 2026-08-22 (epoch 1787421619), tmux session sac_campaign_v2.

Check Evidence
Round counter advances round_counter.txt 95→100→465 across polls; RunTraining Counter check passed: N == N every chunk
Metrics JSONL with scalars training_metrics.jsonl growing (20 lines @ t+45m): full {epoch, steps, buffer_size, drained, grad_steps, critic_loss, actor_loss, alpha_loss, alpha} per line
Eval rotation cycles ≥2 All 3 opponents EVERY cycle: Corners→Crazy→Target ×4+ cycles in eval_log.jsonl (10 games each per cycle)
MA files written weights/ma_history_{Corners,Crazy,Target}.txt created at first cycle, appended since
Best-gate on composite only First write exactly when composite 0.0000 > −1 (missing-file default); later 0% cycles correctly did NOT rewrite (strict improvement enforced)
No transitions during eval windows Lever-4 gate active by construction (sendTrainingMsg drops all msgs under SACLSTM_EVAL_MODE=1, unit-tested); metrics epochs cluster at chunk boundaries
Zero crash banners grep -c 'crash #' = 0 through 20 chunks
Sampling distribution plausible 20 chunks: Corners 10, RamFire 5, Crazy 3, SacTwin 2, Target 0 — within small-n noise of weights (3/2/2/1/1)/9; RamFire already sampled (exposure goal met)

Twin liveness: own round_counter.txt advancing, own sac_latest.zip updating during SacTwin chunks, own metrics file separate from the main bot's.

Watch items (not blockers)

  1. Training-throughput signature: metrics lines consistently show steps=1, buffer_size=24 (= burnIn 8 + trainWindow 16, i.e. exact canSample threshold), drained=1 — one gradient step per pass at threshold-crossing moments rather than large drain bursts. Mechanism unexplained by static code read (per-tick sends should yield bigger bursts); v1 learned to its score-60 state under the same integration code without instrumentation, so learning is not obviously broken — but effective grad-steps/hour is THE number to check at first review. This is precisely what lever-3 instrumentation exists to surface.
  2. Early critic-loss spikes: two 1.56e16 outliers (t+6:16, t+12:04) amid otherwise sane values (~8–35) — same class as the random-init artifact flagged in #59 phase-2 notes (~1e12 there); expect decay. If persistent past early chunks, feeds the lever-5 decision.
  3. .part corpses: three sac_latest.zip.tmp.*.part files accumulated mid-run (libzip interrupted-write artifact, #57 forensics); harmless — atomic renames keep the main zips valid, startup sweep clears them next restart.