26 KiB
Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)
Overnight training campaign on branch research/goto-controller.
Companion tickets: config+launch = #56, morning verdict = #57, map = #53.
Monitoring contract: observability inventory in #55 comment 461 (9 signals, thresholds, four-way discrimination).
Locked config (campaign-v1)
Launch command (tmux session sac_campaign, stdout teed to campaign_stdout.log):
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=25000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
| Knob | Value | Source / rationale |
|---|---|---|
SAC_OPPONENTS |
Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1 |
Working recommendation kept. Twin pinned at 1 not 2: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer. |
SAC_EVAL_OPPONENT |
Corners |
Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun. |
SAC_TOTAL_ROUNDS |
25000 |
Sized so the wall-clock ceiling binds first: measured throughput ~3400 rounds/h early (drops as battles lengthen) ⇒ 2000 would have exhausted in ~1 h. |
SAC_CHUNK_SIZE |
10 |
Harness default. |
SAC_EVAL_INTERVAL |
2 |
Harness default — eval every ~20 rounds. |
SAC_EVAL_ROUNDS |
10 |
Harness default. |
SAC_MAX_CRASHES |
5 |
Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55). |
SACLSTM_HIDDEN_SIZE |
256 |
Module default (network.nim); real capacity vs #49 smoke's 32; under MaxHidden=512 cap. |
SACLSTM_BATCH_SIZE |
16 |
Module default (integration.nim). |
SACLSTM_UTD_RATIO |
1 |
Module default. |
SACLSTM_SAVE_INTERVAL |
5 |
Deviation from default 500 and from #54's "10–20": at hidden 256 a gradient step takes ~1 s and a chunk process fits only ~10–18 steps (see incident below) — interval must sit inside the per-process step budget. 5 ⇒ checkpoint every ~5–10 s of active training; IO trivial (10.5 MB zip, atomic replace). |
| (not pinned) | module defaults | LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4, GAMMA=0.99, TAU=0.005, TARGET_ENTROPY=-4.0, BUFFER_CAPACITY=500000, BURN_IN=8, TRAIN_WINDOW=16. |
Budget: generous wall-clock ceiling, not a deadline — originally T+12 h from launch 00:21:55 CEST 2026-08-22 ⇒ 12:22 CEST (epoch 1787394115); extended 2026-08-22 ~07:35 by orchestrator decision on human mandate ("no deadlines — let the 25000-round budget complete", ~14:35 projected) ⇒ ceiling now 16:30 CEST 2026-08-22 (epoch 1787409000), enforced by a hard user-systemd net unit sac-ceiling-net (sleeps to the epoch, then kills the tmux session and any straggler harness processes). While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.
Fresh start: pre-campaign weights/ held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to weights_smoke49_backup/; campaign baseline re-established by probe battles vs Corners (random-init hidden-256 checkpoint, best_score.txt reset then re-raised to 20 by a genuine eval). Twin regenerated via ./make_twin.sh from that baseline (md5 61521cff… verified seed).
Phase log
- 2026-08-21 23:22 — Claim posted on #56 (comment 471). Config locked, notebook committed (
2f49cb2). - 2026-08-21 23:27 — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (
sac_latest.zip, 10.5 MB) and twin seed. - 2026-08-21 23:32 — Twin regenerated (md5-verified). Launch attempt 1 (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but zero checkpoints persisted (see incident). Killed 23:48.
- 2026-08-21 23:52–00:10 — Diagnosis (see incident): interval=1 fired, interval=2/20 never; instrumentation + /proc thread forensics ⇒ per-process step budget ~10–18 at ~1 s/step; save check ran only between drain-burst passes.
- 2026-08-22 00:12 — Fix: save check moved inside the gradient-step loop, committed
2653671. Validated: interval=5 save fired ~12 s into a battle. - 2026-08-22 00:21:55 — Launch (final): tmux
sac_campaign, config above. First campaign save on disk at t+54 s; eval #1 on cadence. - 2026-08-22 00:30 — HEALTHY checklist passed (see below).
- 2026-08-22 03:45 — Watch shift 1 (00:34–03:35): liveness flawless (10/10 HEALTHY, zero banners, zip ≤30 s). Learning signal: steady-state eval vs Corners 0–5% with two isolated 10/10 spikes (~01:15) → capability emerged, then lost. Eval-regression intervene threshold fired per #55; intervention DEFERRED to Campaign verdict (#57) — rationale: n=2 evidence, no loss metrics, buffer-loss on restart, run completes ~08:05 anyway. Milestone issue: #58 "Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict".
- 2026-08-22 ~07:15 — Morning audit: policy demonstrably learning off-benchmark (SacTwin 73→100%, Crazy 26→56%) while Corners eval stays ~0–9% with 3 transient 10/10s; wall-clock ceiling extended to let the 25k complete (~14:35 projected); instability-vs-plateau question left to the curve.
- 2026-08-22 ~07:50 — score:60 anatomy: Tank Royale survival(50)+last-survivor(10) awarded when opponent dies while we survive; exactly-60 ⇒ zero damage dealt by us that round (opponent self-destructed via wasted shots + 0.1/turn inactivity drain). 3,878 rounds (19%); modal vs SacTwin; vs Corners 986 damageless outlives vs 242 true wins; combined with 43% of rounds being score:0, texture = survivor-not-fighter against walls. RL rewards are event-driven (rewards module), so behavioral evidence, not reward poisoning. Feeds #57 levers: aggression shaping / specialist-vs-generalist.
- 2026-08-22 ~08:05 —
.partdebris forensics + sweep: 99*.zip.tmp.*.part(483 MB) are NOT weights.nim debris (that proc uses fixed.tmp+ finally-cleanup, working); naming matches an external write-temp→rename copier killed mid-write, bursts correlating with kill events; possible culprit: a folder-sync client fighting a file that changes every ~20 s (human asked to confirm). Swept with-mmin +10age guard (protects in-flight writes); 483 MB freed, real zips untouched — note 2 fresh.partreappeared minutes later, copier still active. - 2026-08-22 ~08:00 — ceiling defused: the 12:22 'self-kill' was notebook prose instructing watchmen — never an OS mechanism; rewritten to 16:30 CEST AND armed a real systemd --user net
sac-ceiling-netfiring 16:30:00 (epoch 1787409000); training uninterrupted (counter +99/147 s verified); annotated #56 comment 487. - 2026-08-22 ~14:28 — CAMPAIGN ENDED NATURALLY: banner
>>> training complete: 2500 chunksafter 14 h 07 m (00:21:55 → ~14:28 CEST); round_counter 38912; zero crash banners across the whole run; thesac-ceiling-netbackstop never fired — cancelled unneeded. Budget note: the harness loop is chunk-based —SAC_TOTAL_ROUNDS=25000÷CHUNK_SIZE=10⇒ 2500 chunk battles of ≤10 rounds each; the "25k-rounds" label was a misnomer (the counter also accrues 1287×10 eval rounds and rerun chunks). Final eval vs Corners: 10%. - 2026-08-22 ~14:50 — VERDICT posted (→ #57): RETUNE BEFORE SCALING. Ops layer PROVEN (14 h autonomous, zero crashes, self-healing restarts, natural completion — the harness scales); learning REAL BUT NARROW (within-opponent gains genuine — SacTwin 90.7%, Crazy 26→48% — but specialist-not-generalist, walls untouched; 12 eval spikes ≥8/10 incl. 5×10/10, none retained); benchmark pathology: Corners-only deterministic eval + single-max best gating froze
sac_best.zipat 01:01:57 on a fluke 10/10. Five code-level levers staged awaiting human sign-off; v2 NOT launched. Full rationale: Results below + #57 resolution comment. - 2026-08-22 ~14:50 — HYGIENE (campaign over, no live writers — safe):
sac-ceiling-netstopped +reset-failed(backstop obsolete);.partcorpse sweep 48 → 0 (no age guard needed — nothing writes anymore); future-run guard added tosac_train.sh: startuprm -f "$WEIGHTS_DIR"/sac_latest.zip.tmp.*.part "$WEIGHTS_DIR"/sac_latest.zip.tmpso a SIGKILLed run's libzip modify-path corpses can't accumulate again.
Decision-issue index
| Issue | What it decided |
|---|---|
| #37–#48 | Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration. |
| #49 | Training harness + smoke run (toy hyperparams: hidden 32). |
| #54 | Mirror-twin sparring partner; SAVE_INTERVAL persistence rule; eval opponent = first pool entry. |
| #55 | 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds. |
| #56 | This campaign: locked config + launch + the save-check fix (2653671). |
| #57 | Morning verdict — consumes this notebook + logs. |
Incidents & checks
Incident 1 — zero checkpoint persistence at production sizes (launch blockers, fixed)
Symptom: campaign ran 26+ chunk processes across two attempts without a single sac_latest.zip update, while rounds/evals flowed normally. #54's rule ("keep SACLSTM_SAVE_INTERVAL well below per-chunk gradient-step counts, 10–20 fired in smokes") silently broke at hidden 256.
Diagnosis chain (all reproducible):
- Interval=1 saved within seconds; interval=2 and 20 never saved — through the same harness ⇒ not env propagation.
- Temporary step instrumentation (bot stderr via a one-line
SAC_LSTM_Bot.shredirect — the vendored runner swallows bot stderr, #55 gap S9-adjacent): steps cost ~1.06 s each; a drain burst queued 53 steps; logging stopped mid-pass while rounds kept completing. /proc/<pid>/tasksampling: training thread alive and RUNNING (~13 s CPU per ~40 s process) — not deadlocked, just slow ⇒ per-process step budget ≈ 10–18 steps.- The save check lived between drain-burst passes; with bursts queueing minutes of steps,
stepCountnever reachednextSavebefore process teardown. Smoke runs masked this: hidden 32 steps were sub-millisecond, so hundreds of steps fit per chunk.
Fix (commit 2653671): save check relocated inside the step loop (checked every gradient step; packFull+trySend unchanged). Validated: interval=5 save fires ~12 s into a battle; campaign save fired 54 s after launch.
Config consequences: SACLSTM_SAVE_INTERVAL=5 (inside the per-process budget; #54's 10–20 was derived at smoke speeds). SAC_TOTAL_ROUNDS=25000 (throughput measured ~3400 rounds/h, so 2000 was a 1-hour budget, not an overnight one). OMP_NUM_THREADS=1 tested and not needed (hang was step-budget exhaustion, not OpenMP).
Launch health check (t+9 min, 00:30:16) — HEALTHY per #55 checklist
| Signal | Reading | Verdict |
|---|---|---|
| S1 harness stdout | teed to campaign_stdout.log; 0 crash/aborted banners |
✓ |
| S4 round_counter | 1890, +460 in 9 min (~51 rounds/min) | ✓ advancing |
| S6 sac_latest.zip mtime | 6 s old; first save at t+54 s | ✓ fresh |
| S2 training_log.jsonl | 1492 lines, growing; last ticks=668, plausible | ✓ |
| S3 eval_log.jsonl | age 3 s (atomic replace); eval every 2 chunks | ✓ on cadence |
| S5 best_score | 20 (from a genuine campaign-1 eval; non-decreasing) | ✓ |
| Sampling | 112 chunks: Corners 41 / Crazy 25 / RamFire 15 / SacTwin 16 / Target 15 ≈ weights 3:2:1:1:1 | ✓ plausible |
| Disk | 418 GB free | ✓ |
Early evals 0% vs Corners — expected for a near-random policy minutes in; SLOW-LEARNER watch rule (flat ≥5 evals = watch) applies, never intervene.
Check-in procedure (for monitor sessions)
tmux capture-pane -p -t sac_campaign | tail -5 # S1: banners, crashes
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip # age <~600s = training alive
tail -1 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/training_log.jsonl
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log # eval results / new best
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/best_score.txt
Intervene per #55 thresholds: ≥3 consecutive crash #N banners; ΔS4=0 over ≥15 min; zip mtime >10 min stale while S4 advances (STALLED); disk <1 GB. At the ceiling (16:30 CEST Aug 22, epoch 1787409000 — extended from 12:22 per human mandate): tmux kill-session -t sac_campaign if still running (the hard net unit sac-ceiling-net fires at the same epoch regardless of monitors) — final state is in weights/, logs, and this notebook.
Results (campaign-v1 — filled by #57)
Run: 2500/2500 chunks · 14 h 07 m autonomous (00:21:55 → ~14:28 CEST 2026-08-22) · round_counter 38912 · zero crashes · graceful banner >>> training complete: 2500 chunks · systemd net never fired. Budget was chunk-based (see phase log) — the "25k-rounds" label was a misnomer.
Per-opponent training win rates (run-3 slice of training_log.jsonl):
| Opponent | Win rate | Record | Note |
|---|---|---|---|
| SacTwin | 90.7% | 2693/2970 | vs frozen past-self — genuine self-play gain |
| Crazy | 48.3% | 2965/6140 | doubled from 26% early-run |
| Target | 9.6% | — | static, barely moved |
| Corners | 7.2% | — | walls untouched |
| RamFire | 0% | 0/2920 | mirrors the PPO-era ladder — ram-class needs dedicated pressure |
Eval vs Corners (pinned benchmark): 1287 evals · overall mean 7.7% · histogram headline: 0/10 = 867 (67%), spikes ≥8/10 = 12 (incl. 5× perfect 10/10) · final eval 10%. Stdout log carries no timestamps; timing reconstructed from file mtimes.
Best-zip paradox: sac_best.zip frozen since 01:01:57 — a single lucky 10/10 at ~round 3.5k wrote best_score=100, and no later eval could outrank a perfect score (even genuine ~50%-winrate stretches elsewhere). Best checkpoint = lottery ticket, decoupled from the steady-state policy (which sat at 0–10% vs Corners).
Verdict: RETUNE BEFORE SCALING — full rationale in #57 resolution comment. Five code-level levers staged for human sign-off: (1) eval rotation across pool + moving-average best gating; (2) reward shaping toward damage/aggression incl. anti-ram signal; (3) training-loss/step metrics logged from the training thread (#55 gap #1); (4) gate sendTrainingMsg off in eval mode (#55 gap #3); (5) optional stability knobs (lower LR / entropy coeff) once loss curves exist.
Campaign Notebook — campaign-v2 (SAC_LSTM_Bot)
Locked config (campaign-v2)
Launched verbatim from orchestrator mandate (umbrella ticket #59, all five levers approved & implemented in a07e530 + 6fc01eb):
cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
| Knob | Value | Rationale |
|---|---|---|
SAC_OPPONENTS |
Corners:3, Crazy:2, RamFire:2, Target:1, SacTwin:1 | RamFire bumped 1→2 vs v1: anti-ram shaping (lever 2, 6fc01eb) needs exposure to fire; without samples there is no gradient signal against ram-class |
SAC_EVAL_OPPONENTS |
Corners,Crazy,Target | Lever-1 rotation set (default); composite = mean of per-opponent MA-5 win rates |
SAC_EVAL_INTERVAL/ROUNDS |
2 / 10 | Unchanged from v1 cadence |
SAC_TOTAL_ROUNDS/CHUNK_SIZE/MAX_CRASHES |
25000 / 10 / 5 | Same budget semantics as v1 (chunk-based) |
SACLSTM_HIDDEN_SIZE/BATCH_SIZE/SAVE_INTERVAL |
256 / 16 / 5 | Architecture + throughput knobs carried over; LR/entropy defaults untouched = lever-5 conditional posture (activation decided by loss-curve evidence, not upfront) |
Fresh start & archive
v1 state archived intact (notebook references preserved) into SAC_LSTM_Bot/weights_v1_archive/: sac_latest.zip, sac_best.zip, best_score.txt, round_counter.txt, training_log.jsonl, eval_log.jsonl, campaign_stdout.log, plus the lever-3 smoke leftover training_metrics.jsonl (from src/SAC_LSTM_Bot/, moved so v2 loss curves start clean for lever-5 reading). Main weights/ verified empty afterwards ⇒ main bot takes the genuine random-init path (loadOrInitFull → randomFull()).
Twin reseed: make_twin.sh requires a seed zip, but the fresh-start baseline has none. Generated a fresh random-init checkpoint (hidden=256, alpha=1.0 matching logAlpha=0) via a throwaway Nim script against network.nim/weights.nim, seeded it as sac_best.zip transiently, ran ./make_twin.sh, removed the transient copy. Verified twin dir got byte-identical fresh zips (cmp OK; NOT v1 zips) + round_counter.txt=0. No script changes needed.
Safety net
systemd-run --user --unit=sac-ceiling-net-v2 armed at launch: sleeps 72000 s then tmux kill-session -t sac_campaign_v2; sleep 5; pkill -f sac_train.sh. Unit active at 19:59:40 CEST 2026-08-22, fires 15:59:40 CEST 2026-08-23 (epoch 1787493580).
Launch & health evidence (first ~45 min)
Launched 20:00:19 CEST 2026-08-22 (epoch 1787421619), tmux session sac_campaign_v2.
| Check | Evidence |
|---|---|
| Round counter advances | round_counter.txt 95→100→465 across polls; RunTraining Counter check passed: N == N every chunk |
| Metrics JSONL with scalars | training_metrics.jsonl growing (20 lines @ t+45m): full {epoch, steps, buffer_size, drained, grad_steps, critic_loss, actor_loss, alpha_loss, alpha} per line |
| Eval rotation cycles ≥2 | All 3 opponents EVERY cycle: Corners→Crazy→Target ×4+ cycles in eval_log.jsonl (10 games each per cycle) |
| MA files written | weights/ma_history_{Corners,Crazy,Target}.txt created at first cycle, appended since |
| Best-gate on composite only | First write exactly when composite 0.0000 > −1 (missing-file default); later 0% cycles correctly did NOT rewrite (strict improvement enforced) |
| No transitions during eval windows | Lever-4 gate active by construction (sendTrainingMsg drops all msgs under SACLSTM_EVAL_MODE=1, unit-tested); metrics epochs cluster at chunk boundaries |
| Zero crash banners | grep -c 'crash #' = 0 through 20 chunks |
| Sampling distribution plausible | 20 chunks: Corners 10, RamFire 5, Crazy 3, SacTwin 2, Target 0 — within small-n noise of weights (3/2/2/1/1)/9; RamFire already sampled (exposure goal met) |
Twin liveness: own round_counter.txt advancing, own sac_latest.zip updating during SacTwin chunks, own metrics file separate from the main bot's.
Watch items (not blockers)
- Training-throughput signature: metrics lines consistently show
steps=1, buffer_size=24 (= burnIn 8 + trainWindow 16, i.e. exact canSample threshold), drained=1— one gradient step per pass at threshold-crossing moments rather than large drain bursts. Mechanism unexplained by static code read (per-tick sends should yield bigger bursts); v1 learned to its score-60 state under the same integration code without instrumentation, so learning is not obviously broken — but effective grad-steps/hour is THE number to check at first review. This is precisely what lever-3 instrumentation exists to surface. - Early critic-loss spikes: two
1.56e16outliers (t+6:16, t+12:04) amid otherwise sane values (~8–35) — same class as the random-init artifact flagged in #59 phase-2 notes (~1e12 there); expect decay. If persistent past early chunks, feeds the lever-5 decision. .partcorpses: threesac_latest.zip.tmp.*.partfiles accumulated mid-run (libzip interrupted-write artifact, #57 forensics); harmless — atomic renames keep the main zips valid, startup sweep clears them next restart.
~21:30 — v2 attempt-1 checkpoint finding
Attempt-1's rolling zip was dead ~40 min at birth: SACLSTM_SAVE_INTERVAL counts gradient steps, and each training pass inside the short battle processes carries exactly 1 step (the steps=1, drained=1 signature of watch item 1) ⇒ interval-5 never reached its trigger within a process lifetime — no sac_latest.zip roll ever fired despite healthy training. Same class as v1's 23:32 incident, resurfacing through a different seam (per-process step budget vs per-pass step count).
Fix (env-level, no rebuild): SACLSTM_SAVE_INTERVAL=1, restart 21:09:30 CEST. Saves verified flowing: 21:09 / 21:13 / 21:16 (and still flowing at audit time — sac_latest.zip mtime 21:22:28, sac_best.zip 21:12:50). The locked-config table above keeps the original attempt-1 launch line for the record; live relaunch differs only in this knob.
~21:30 — twin-freeze contract corrected
make_twin.sh twin launcher now exports SACLSTM_EVAL_MODE=1 (commit 19f34ab). Without lever-4's gate on the twin side, the v1 SacTwin had been TRAINING throughout, not frozen — every mirror battle updated the twin's own weights. Consequence: v1's headline "90.7% vs twin" is retroactively an arms-race win rate (both policies co-evolving), not a fixed-benchmark score, and the #57 verdict phrasing implying a frozen sparring partner is corrected by this entry. No numbers change; the story does.
~21:30 — net extended + pacing decision
Slowdown attribution complete: 89% of wall clock = eval rotation by design (SAC_EVAL_INTERVAL=2 × 3 opponents × 10 rounds each); trainer exonerated at ~500 ticks/s. Decision recorded in issue #60.
Ceiling net re-armed to match the extended budget: fires Thu 2026-08-27 07:59:40 CEST (epoch 1787810380) — supersedes the 2026-08-23 15:59:40 fire noted under Safety net. Attempt-1 artifacts preserved in /tmp/v2_attempt1_backup/.
Campaign v2 — attempt-3 (2026-08-23)
~06:45 — LEVER-5 TRIGGERED (watchman shift 2, 02:11–06:22 CEST)
Evidence-gated activation of the lever-5 conditional posture (SACLSTM_LR_CRITIC), human signed off. Trigger evidence:
| Signal | Observation |
|---|---|
| actor_loss | |34M| → |106M| monotone (+15M/h), no plateau |
| critic_loss | spikes >1e12 in 100% of NEW metric records; campaign share 80.2% |
| Signature | exact 1.5625e16 recurring (= float32 saturation neighborhood) |
| alpha | decayed 0.951 → 0.711 (entropy collapse under runaway Q scale) |
| best_score.txt | frozen 00:32 (=83.3333) through 5h50m of composite oscillation 3.3–43.3 |
| Per-opponent MAs | whipsawing (Crazy 30→100→90→0→60→90→10→0) |
DECISION — LR_CRITIC=1e-4, single knob, fresh start
- Adjudication: primary suspect is critic/Q value-scale growth. Actor loss inherits Q magnitude through the policy gradient, so the actor explosion is downstream; alpha decay is a symptom (entropy temperature chasing a blown value scale), not a cause. Therefore ONE knob moves:
SACLSTM_LR_CRITIC=1e-4(critic learns slower → Q estimates stay estimable); LR_ACTOR/LR_ALPHA stay at default 3e-4, TARGET_ENTROPY −4. Multi-knob changes would confound attempt-3's attribution. - Fresh start over warm start: attempt-2's weights carry saturated Q representations; warm-starting them under a new LR would inherit the pathology we're trying to escape (contamination risk). Attempt-2 archived intact, nothing discarded.
MA-wipe root cause FOUND & FIXED (167bcc4)
Watchman anomaly explained: ma_history_*.txt held exactly one leading-space value per file (e.g. [ 10 ], [ 20 ], [ 0 ] in the attempt-2 backup) instead of the designed 5-value history, degrading the composite best-gate to last-cycle mean — which fully accounts for the composite whipsaw 3.3–43.3.
Root cause: in eval_checkpoint, the append line was
printf '%s\n' "$(cat "$(ma_hist_file "$opp")")" "$wr" | tail -n 5 | tr '\n' ' ' > "$(ma_hist_file "$opp")"
Bash sets up the > redirect before running the command substitution, so $(cat f) always read the freshly truncated file ⇒ every cycle wiped history to " $wr ". Reproduced standalone on bash 5.3 before touching the script; no other writer exists (grep). Fix hoists the read into its own statement and word-splits it so tail -n 5 keeps exactly the last MA_WINDOW values. Verified live in attempt-3: two eval cycles produced [0 0 ] per opponent (previously impossible).
Attempt-3 launch (2026-08-23 07:01:28 CEST, epoch 1787461288)
- Attempt-2 stopped cleanly (tmux kill; no stray procs); artifacts (76 items, 336 MB incl.
.partcorpses + MA/best/latest/counter) →/tmp/v2_attempt2_backup/; final round_counter 8452;weights/verified empty. - Twin regenerated via documented transient-seed method: throwaway Nim script (
initSACTrainer(35,4)+saveWeights, hidden=256 env, alpha=1.0) seeded as transientweights/sac_best.zip→./make_twin.sh→ twin zips byte-identical to seed (cmpOK), twin counter=0 → transient removed. - Launch line = locked config verbatim plus the two live deltas:
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=1 \
SACLSTM_LR_CRITIC=1e-4 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
- Safety net: existing transient unit
sac-ceiling-net-v2.servicere-verified active; fire epoch start+384587 s = 1787810380 = Thu 2026-08-27 07:59:40 CEST exactly (delta 0 vs mandate) — kept, not duplicated.
Health baseline (t+15 m)
| Check | Evidence |
|---|---|
| Counter | 33 (t+8m) → 125 (t+15m) |
| Metrics JSONL | flowing (6 lines); baseline line: critic_loss=23.09, actor_loss=-3.228, alpha_loss=0.0, alpha=0.99970 (random-init magnitudes — attempt-3's divergence curve starts here); t+15m line: critic 84.5, actor −9.99, alpha 0.9937 |
| Saves | SAVE_INTERVAL=1: sac_latest.zip written from first grad-step onward |
| Eval rotation | full cycle landed: Corners/Crazy/Target ×10 games each in eval_log.jsonl |
| MA files | [0 0 ] per opponent after two cycles — fix confirmed in vivo |
| Best gate | best_score.txt=0.0000 written on first composite (0 > −1 default) |
| Crash banners | 0 |