Files

31 KiB
Raw Permalink Blame History

The story so far, in simple words

This project trains a robot tank. It plays many fights against other tanks. After each fight it changes itself a little. It keeps the changes that helped it win.

Night 1 (run 1) finished without problems. It ran for 14 hours alone. It never crashed. It beat an old copy of itself most of the time. It beat Crazy about half the time. Three things went badly. First, learning was not stable. Good skill appeared, then disappeared again. Second, the saved "best" version came from one lucky perfect score. It was not really its best. Third, the bot learned to hide and survive. It almost never shot back.

The human approved five fixes. All five were put into the code. Run 2 used these fixes. Its error numbers grew far too big. Learning broke. We made one speed number smaller. This number sets how fast one part learns. Then we dropped the broken progress and started clean. This is run 3. It is running now.

Next we watch run 3. One of three doors will open. Door 1: it stays steady. We let it run to the end. Door 2: the numbers grow too big again. We turn the next speed number down. Door 3: it stays steady but still fights badly. We teach aiming as a separate, direct lesson.

Updated: 2026-08-23 — this section is refreshed at every major step.

Small dictionary

  • training: the time when the bot plays fights and changes itself to improve. It learns only during training.
  • battle: one group of fights against one opponent. The bot restarts between groups.
  • round: one single fight. Win it by destroying the enemy tank or outliving it.
  • chunk: one work block: a battle of up to 10 rounds, then some learning from it.
  • eval (test match): a test match. The bot does not learn during these. We use them only to measure.
  • win rate: how many test matches were won, as a percent. 8 wins in 10 matches = 80%.
  • checkpoint: a saved copy of the bot's brain (a zip file). Written every few learning steps.
  • "best" checkpoint: the saved copy we currently call best. Run 1 picked one from a lucky score, hence the quotes.
  • replay buffer: the bot's memory of past moments: what it saw, did, and received. Learning picks old moments from it.
  • loss (critic/actor): a number saying how wrong the bot's inner guesses are. Lower usually means better. Losses growing huge mean trouble.
  • alpha: a dial setting how much the bot tries new moves instead of repeating known good ones.
  • MA / composite score: MA is the average of the last few win rates; it smooths luck. Composite is the average of MAs across all test opponents.
  • twin (SacTwin): a frozen copy of our own bot, used as a practice partner. Beating it proves real improvement.
  • lever: one numbered change we prepared, waiting for approval. There are levers 1 to 5.
  • watchman: a helper who checks the running training at set times and stops it if something breaks.

Campaign Notebook — campaign-v1 (SAC_LSTM_Bot)

Overnight training campaign on branch research/goto-controller. Companion tickets: config+launch = #56, morning verdict = #57, map = #53. Monitoring contract: observability inventory in #55 comment 461 (9 signals, thresholds, four-way discrimination).

Locked config (campaign-v1)

Launch command (tmux session sac_campaign, stdout teed to campaign_stdout.log):

SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENT=Corners \
SAC_TOTAL_ROUNDS=25000 \
SAC_CHUNK_SIZE=10 \
SAC_EVAL_INTERVAL=2 \
SAC_EVAL_ROUNDS=10 \
SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 \
SACLSTM_BATCH_SIZE=16 \
SACLSTM_UTD_RATIO=1 \
SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_stdout.log
Knob Value Source / rationale
SAC_OPPONENTS Corners:3,Crazy:2,RamFire:1,Target:1,SacTwin:1 Working recommendation kept. Twin pinned at 1 not 2: #54 showed mirror battles end early ⇒ fewer transitions per chunk; weight 2 would starve the replay buffer.
SAC_EVAL_OPPONENT Corners Pinned explicitly (= first pool entry default, #54 note) — removes reorder footgun.
SAC_TOTAL_ROUNDS 25000 Sized so the wall-clock ceiling binds first: measured throughput ~3400 rounds/h early (drops as battles lengthen) ⇒ 2000 would have exhausted in ~1 h.
SAC_CHUNK_SIZE 10 Harness default.
SAC_EVAL_INTERVAL 2 Harness default — eval every ~20 rounds.
SAC_EVAL_ROUNDS 10 Harness default.
SAC_MAX_CRASHES 5 Harness self-abort; monitor intervenes earlier at ≥3 consecutive crashes (#55).
SACLSTM_HIDDEN_SIZE 256 Module default (network.nim); real capacity vs #49 smoke's 32; under MaxHidden=512 cap.
SACLSTM_BATCH_SIZE 16 Module default (integration.nim).
SACLSTM_UTD_RATIO 1 Module default.
SACLSTM_SAVE_INTERVAL 5 Deviation from default 500 and from #54's "10–20": at hidden 256 a gradient step takes ~1 s and a chunk process fits only ~10–18 steps (see incident below) — interval must sit inside the per-process step budget. 5 ⇒ checkpoint every ~5–10 s of active training; IO trivial (10.5 MB zip, atomic replace).
(not pinned) module defaults LR_ACTOR/LR_CRITIC/LR_ALPHA=3e-4, GAMMA=0.99, TAU=0.005, TARGET_ENTROPY=-4.0, BUFFER_CAPACITY=500000, BURN_IN=8, TRAIN_WINDOW=16.

Budget: generous wall-clock ceiling, not a deadline — originally T+12 h from launch 00:21:55 CEST 2026-08-22 ⇒ 12:22 CEST (epoch 1787394115); extended 2026-08-22 ~07:35 by orchestrator decision on human mandate ("no deadlines — let the 25000-round budget complete", ~14:35 projected) ⇒ ceiling now 16:30 CEST 2026-08-22 (epoch 1787409000), enforced by a hard user-systemd net unit sac-ceiling-net (sleeps to the epoch, then kills the tmux session and any straggler harness processes). While HEALTHY per #55 discrimination rules the run continues; a monitor kills the tmux session at the ceiling or on an intervene threshold.

Fresh start: pre-campaign weights/ held #49-smoke 32-hidden checkpoints, incompatible with hidden=256. Archived to weights_smoke49_backup/; campaign baseline re-established by probe battles vs Corners (random-init hidden-256 checkpoint, best_score.txt reset then re-raised to 20 by a genuine eval). Twin regenerated via ./make_twin.sh from that baseline (md5 61521cff… verified seed).

Phase log

Every entry below starts with a plain-language first sentence. Technical detail follows for those who want it.

  • 2026-08-21 23:22 — Claim posted on #56 (comment 471). Config locked, notebook committed (2f49cb2).
  • 2026-08-21 23:27 — Smoke weights archived; release build; bootstrap + probe battles vs Corners established a hidden-256 baseline checkpoint (sac_latest.zip, 10.5 MB) and twin seed.
  • 2026-08-21 23:32 — Twin regenerated (md5-verified). Launch attempt 1 (SAVE_INTERVAL=20, TOTAL_ROUNDS=2000): ran 16+ chunks, evals every 2 chunks — but zero checkpoints persisted (see incident). Killed 23:48.
  • 2026-08-21 23:52–00:10 — Diagnosis (see incident): interval=1 fired, interval=2/20 never; instrumentation + /proc thread forensics ⇒ per-process step budget ~10–18 at ~1 s/step; save check ran only between drain-burst passes.
  • 2026-08-22 00:12 — Fix: save check moved inside the gradient-step loop, committed 2653671. Validated: interval=5 save fired ~12 s into a battle.
  • 2026-08-22 00:21:55 — Launch (final): tmux sac_campaign, config above. First campaign save on disk at t+54 s; eval #1 on cadence.
  • 2026-08-22 00:30 — HEALTHY checklist passed (see below).
  • 2026-08-22 03:45 — Watch shift 1 (00:34–03:35): liveness flawless (10/10 HEALTHY, zero banners, zip ≤30 s). Learning signal: steady-state eval vs Corners 0–5% with two isolated 10/10 spikes (~01:15) → capability emerged, then lost. Eval-regression intervene threshold fired per #55; intervention DEFERRED to Campaign verdict (#57) — rationale: n=2 evidence, no loss metrics, buffer-loss on restart, run completes ~08:05 anyway. Milestone issue: #58 "Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict".
  • 2026-08-22 ~07:15 — Morning audit: policy demonstrably learning off-benchmark (SacTwin 73→100%, Crazy 26→56%) while Corners eval stays ~0–9% with 3 transient 10/10s; wall-clock ceiling extended to let the 25k complete (~14:35 projected); instability-vs-plateau question left to the curve.
  • 2026-08-22 ~07:50 — score:60 anatomy: Tank Royale survival(50)+last-survivor(10) awarded when opponent dies while we survive; exactly-60 ⇒ zero damage dealt by us that round (opponent self-destructed via wasted shots + 0.1/turn inactivity drain). 3,878 rounds (19%); modal vs SacTwin; vs Corners 986 damageless outlives vs 242 true wins; combined with 43% of rounds being score:0, texture = survivor-not-fighter against walls. RL rewards are event-driven (rewards module), so behavioral evidence, not reward poisoning. Feeds #57 levers: aggression shaping / specialist-vs-generalist.
  • 2026-08-22 ~08:05 — .part debris forensics + sweep: 99 *.zip.tmp.*.part (483 MB) are NOT weights.nim debris (that proc uses fixed .tmp + finally-cleanup, working); naming matches an external write-temp→rename copier killed mid-write, bursts correlating with kill events; possible culprit: a folder-sync client fighting a file that changes every ~20 s (human asked to confirm). Swept with -mmin +10 age guard (protects in-flight writes); 483 MB freed, real zips untouched — note 2 fresh .part reappeared minutes later, copier still active.
  • 2026-08-22 ~08:00 — ceiling defused: the 12:22 'self-kill' was notebook prose instructing watchmen — never an OS mechanism; rewritten to 16:30 CEST AND armed a real systemd --user net sac-ceiling-net firing 16:30:00 (epoch 1787409000); training uninterrupted (counter +99/147 s verified); annotated #56 comment 487.
  • 2026-08-22 ~14:28 — CAMPAIGN ENDED NATURALLY: banner >>> training complete: 2500 chunks after 14 h 07 m (00:21:55 → ~14:28 CEST); round_counter 38912; zero crash banners across the whole run; the sac-ceiling-net backstop never fired — cancelled unneeded. Budget note: the harness loop is chunk-based — SAC_TOTAL_ROUNDS=25000 ÷ CHUNK_SIZE=10 ⇒ 2500 chunk battles of ≤10 rounds each; the "25k-rounds" label was a misnomer (the counter also accrues 1287×10 eval rounds and rerun chunks). Final eval vs Corners: 10%.
  • 2026-08-22 ~14:50 — VERDICT posted (→ #57): RETUNE BEFORE SCALING. Ops layer PROVEN (14 h autonomous, zero crashes, self-healing restarts, natural completion — the harness scales); learning REAL BUT NARROW (within-opponent gains genuine — SacTwin 90.7%, Crazy 26→48% — but specialist-not-generalist, walls untouched; 12 eval spikes ≥8/10 incl. 5×10/10, none retained); benchmark pathology: Corners-only deterministic eval + single-max best gating froze sac_best.zip at 01:01:57 on a fluke 10/10. Five code-level levers staged awaiting human sign-off; v2 NOT launched. Full rationale: Results below + #57 resolution comment.
  • 2026-08-22 ~14:50 — HYGIENE (campaign over, no live writers — safe): sac-ceiling-net stopped + reset-failed (backstop obsolete); .part corpse sweep 48 → 0 (no age guard needed — nothing writes anymore); future-run guard added to sac_train.sh: startup rm -f "$WEIGHTS_DIR"/sac_latest.zip.tmp.*.part "$WEIGHTS_DIR"/sac_latest.zip.tmp so a SIGKILLed run's libzip modify-path corpses can't accumulate again.

Decision-issue index

Issue What it decided
#37–#48 Bot built: skeleton, state, actions, rewards, LSTM network, weights, SAC+LSTM training, integration.
#49 Training harness + smoke run (toy hyperparams: hidden 32).
#54 Mirror-twin sparring partner; SAVE_INTERVAL persistence rule; eval opponent = first pool entry.
#55 9-signal observability inventory; CRASHED/STALLED/SLOW-LEARNER/HEALTHY discriminators; monitor thresholds.
#56 This campaign: locked config + launch + the save-check fix (2653671).
#57 Morning verdict — consumes this notebook + logs.

Incidents & checks

Incident 1 — zero checkpoint persistence at production sizes (launch blockers, fixed)

Symptom: campaign ran 26+ chunk processes across two attempts without a single sac_latest.zip update, while rounds/evals flowed normally. #54's rule ("keep SACLSTM_SAVE_INTERVAL well below per-chunk gradient-step counts, 10–20 fired in smokes") silently broke at hidden 256.

Diagnosis chain (all reproducible):

  1. Interval=1 saved within seconds; interval=2 and 20 never saved — through the same harness ⇒ not env propagation.
  2. Temporary step instrumentation (bot stderr via a one-line SAC_LSTM_Bot.sh redirect — the vendored runner swallows bot stderr, #55 gap S9-adjacent): steps cost ~1.06 s each; a drain burst queued 53 steps; logging stopped mid-pass while rounds kept completing.
  3. /proc/<pid>/task sampling: training thread alive and RUNNING (~13 s CPU per ~40 s process) — not deadlocked, just slow ⇒ per-process step budget ≈ 10–18 steps.
  4. The save check lived between drain-burst passes; with bursts queueing minutes of steps, stepCount never reached nextSave before process teardown. Smoke runs masked this: hidden 32 steps were sub-millisecond, so hundreds of steps fit per chunk.

Fix (commit 2653671): save check relocated inside the step loop (checked every gradient step; packFull+trySend unchanged). Validated: interval=5 save fires ~12 s into a battle; campaign save fired 54 s after launch.

Config consequences: SACLSTM_SAVE_INTERVAL=5 (inside the per-process budget; #54's 10–20 was derived at smoke speeds). SAC_TOTAL_ROUNDS=25000 (throughput measured ~3400 rounds/h, so 2000 was a 1-hour budget, not an overnight one). OMP_NUM_THREADS=1 tested and not needed (hang was step-budget exhaustion, not OpenMP).

Launch health check (t+9 min, 00:30:16) — HEALTHY per #55 checklist

Signal Reading Verdict
S1 harness stdout teed to campaign_stdout.log; 0 crash/aborted banners ✓
S4 round_counter 1890, +460 in 9 min (~51 rounds/min) ✓ advancing
S6 sac_latest.zip mtime 6 s old; first save at t+54 s ✓ fresh
S2 training_log.jsonl 1492 lines, growing; last ticks=668, plausible ✓
S3 eval_log.jsonl age 3 s (atomic replace); eval every 2 chunks ✓ on cadence
S5 best_score 20 (from a genuine campaign-1 eval; non-decreasing) ✓
Sampling 112 chunks: Corners 41 / Crazy 25 / RamFire 15 / SacTwin 16 / Target 15 ≈ weights 3:2:1:1:1 ✓ plausible
Disk 418 GB free ✓

Early evals 0% vs Corners — expected for a near-random policy minutes in; SLOW-LEARNER watch rule (flat ≥5 evals = watch) applies, never intervene.

Check-in procedure (for monitor sessions)

tmux capture-pane -p -t sac_campaign | tail -5        # S1: banners, crashes
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/round_counter.txt
stat -c '%Y' ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/sac_latest.zip   # age <~600s = training alive
tail -1 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/training_log.jsonl
tail -3 ~/Projects/SirRoboGarage/SAC_LSTM_Bot/campaign_stdout.log           # eval results / new best
cat ~/Projects/SirRoboGarage/SAC_LSTM_Bot/weights/best_score.txt

Intervene per #55 thresholds: ≥3 consecutive crash #N banners; ΔS4=0 over ≥15 min; zip mtime >10 min stale while S4 advances (STALLED); disk <1 GB. At the ceiling (16:30 CEST Aug 22, epoch 1787409000 — extended from 12:22 per human mandate): tmux kill-session -t sac_campaign if still running (the hard net unit sac-ceiling-net fires at the same epoch regardless of monitors) — final state is in weights/, logs, and this notebook.

Results (campaign-v1 — filled by #57)

Run: 2500/2500 chunks · 14 h 07 m autonomous (00:21:55 → ~14:28 CEST 2026-08-22) · round_counter 38912 · zero crashes · graceful banner >>> training complete: 2500 chunks · systemd net never fired. Budget was chunk-based (see phase log) — the "25k-rounds" label was a misnomer.

Per-opponent training win rates (run-3 slice of training_log.jsonl):

Opponent Win rate Record Note
SacTwin 90.7% 2693/2970 vs frozen past-self — genuine self-play gain
Crazy 48.3% 2965/6140 doubled from 26% early-run
Target 9.6% — static, barely moved
Corners 7.2% — walls untouched
RamFire 0% 0/2920 mirrors the PPO-era ladder — ram-class needs dedicated pressure

Eval vs Corners (pinned benchmark): 1287 evals · overall mean 7.7% · histogram headline: 0/10 = 867 (67%), spikes ≥8/10 = 12 (incl. 5× perfect 10/10) · final eval 10%. Stdout log carries no timestamps; timing reconstructed from file mtimes.

Best-zip paradox: sac_best.zip frozen since 01:01:57 — a single lucky 10/10 at ~round 3.5k wrote best_score=100, and no later eval could outrank a perfect score (even genuine ~50%-winrate stretches elsewhere). Best checkpoint = lottery ticket, decoupled from the steady-state policy (which sat at 0–10% vs Corners).

Verdict: RETUNE BEFORE SCALING — full rationale in #57 resolution comment. Five code-level levers staged for human sign-off: (1) eval rotation across pool + moving-average best gating; (2) reward shaping toward damage/aggression incl. anti-ram signal; (3) training-loss/step metrics logged from the training thread (#55 gap #1); (4) gate sendTrainingMsg off in eval mode (#55 gap #3); (5) optional stability knobs (lower LR / entropy coeff) once loss curves exist.


Campaign Notebook — campaign-v2 (SAC_LSTM_Bot)

Locked config (campaign-v2)

Launched verbatim from orchestrator mandate (umbrella ticket #59, all five levers approved & implemented in a07e530 + 6fc01eb):

cd /home/davide/Projects/SirRoboGarage/SAC_LSTM_Bot && \
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=5 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
Knob Value Rationale
SAC_OPPONENTS Corners:3, Crazy:2, RamFire:2, Target:1, SacTwin:1 RamFire bumped 1→2 vs v1: anti-ram shaping (lever 2, 6fc01eb) needs exposure to fire; without samples there is no gradient signal against ram-class
SAC_EVAL_OPPONENTS Corners,Crazy,Target Lever-1 rotation set (default); composite = mean of per-opponent MA-5 win rates
SAC_EVAL_INTERVAL/ROUNDS 2 / 10 Unchanged from v1 cadence
SAC_TOTAL_ROUNDS/CHUNK_SIZE/MAX_CRASHES 25000 / 10 / 5 Same budget semantics as v1 (chunk-based)
SACLSTM_HIDDEN_SIZE/BATCH_SIZE/SAVE_INTERVAL 256 / 16 / 5 Architecture + throughput knobs carried over; LR/entropy defaults untouched = lever-5 conditional posture (activation decided by loss-curve evidence, not upfront)

Fresh start & archive

v1 state archived intact (notebook references preserved) into SAC_LSTM_Bot/weights_v1_archive/: sac_latest.zip, sac_best.zip, best_score.txt, round_counter.txt, training_log.jsonl, eval_log.jsonl, campaign_stdout.log, plus the lever-3 smoke leftover training_metrics.jsonl (from src/SAC_LSTM_Bot/, moved so v2 loss curves start clean for lever-5 reading). Main weights/ verified empty afterwards ⇒ main bot takes the genuine random-init path (loadOrInitFull → randomFull()).

Twin reseed: make_twin.sh requires a seed zip, but the fresh-start baseline has none. Generated a fresh random-init checkpoint (hidden=256, alpha=1.0 matching logAlpha=0) via a throwaway Nim script against network.nim/weights.nim, seeded it as sac_best.zip transiently, ran ./make_twin.sh, removed the transient copy. Verified twin dir got byte-identical fresh zips (cmp OK; NOT v1 zips) + round_counter.txt=0. No script changes needed.

Safety net

systemd-run --user --unit=sac-ceiling-net-v2 armed at launch: sleeps 72000 s then tmux kill-session -t sac_campaign_v2; sleep 5; pkill -f sac_train.sh. Unit active at 19:59:40 CEST 2026-08-22, fires 15:59:40 CEST 2026-08-23 (epoch 1787493580).

Launch & health evidence (first ~45 min)

Launched 20:00:19 CEST 2026-08-22 (epoch 1787421619), tmux session sac_campaign_v2.

Check Evidence
Round counter advances round_counter.txt 95→100→465 across polls; RunTraining Counter check passed: N == N every chunk
Metrics JSONL with scalars training_metrics.jsonl growing (20 lines @ t+45m): full {epoch, steps, buffer_size, drained, grad_steps, critic_loss, actor_loss, alpha_loss, alpha} per line
Eval rotation cycles ≥2 All 3 opponents EVERY cycle: Corners→Crazy→Target ×4+ cycles in eval_log.jsonl (10 games each per cycle)
MA files written weights/ma_history_{Corners,Crazy,Target}.txt created at first cycle, appended since
Best-gate on composite only First write exactly when composite 0.0000 > −1 (missing-file default); later 0% cycles correctly did NOT rewrite (strict improvement enforced)
No transitions during eval windows Lever-4 gate active by construction (sendTrainingMsg drops all msgs under SACLSTM_EVAL_MODE=1, unit-tested); metrics epochs cluster at chunk boundaries
Zero crash banners grep -c 'crash #' = 0 through 20 chunks
Sampling distribution plausible 20 chunks: Corners 10, RamFire 5, Crazy 3, SacTwin 2, Target 0 — within small-n noise of weights (3/2/2/1/1)/9; RamFire already sampled (exposure goal met)

Twin liveness: own round_counter.txt advancing, own sac_latest.zip updating during SacTwin chunks, own metrics file separate from the main bot's.

Watch items (not blockers)

  1. Training-throughput signature: metrics lines consistently show steps=1, buffer_size=24 (= burnIn 8 + trainWindow 16, i.e. exact canSample threshold), drained=1 — one gradient step per pass at threshold-crossing moments rather than large drain bursts. Mechanism unexplained by static code read (per-tick sends should yield bigger bursts); v1 learned to its score-60 state under the same integration code without instrumentation, so learning is not obviously broken — but effective grad-steps/hour is THE number to check at first review. This is precisely what lever-3 instrumentation exists to surface.
  2. Early critic-loss spikes: two 1.56e16 outliers (t+6:16, t+12:04) amid otherwise sane values (~8–35) — same class as the random-init artifact flagged in #59 phase-2 notes (~1e12 there); expect decay. If persistent past early chunks, feeds the lever-5 decision.
  3. .part corpses: three sac_latest.zip.tmp.*.part files accumulated mid-run (libzip interrupted-write artifact, #57 forensics); harmless — atomic renames keep the main zips valid, startup sweep clears them next restart.

~21:30 — v2 attempt-1 checkpoint finding

Attempt-1's rolling zip was dead ~40 min at birth: SACLSTM_SAVE_INTERVAL counts gradient steps, and each training pass inside the short battle processes carries exactly 1 step (the steps=1, drained=1 signature of watch item 1) ⇒ interval-5 never reached its trigger within a process lifetime — no sac_latest.zip roll ever fired despite healthy training. Same class as v1's 23:32 incident, resurfacing through a different seam (per-process step budget vs per-pass step count).

Fix (env-level, no rebuild): SACLSTM_SAVE_INTERVAL=1, restart 21:09:30 CEST. Saves verified flowing: 21:09 / 21:13 / 21:16 (and still flowing at audit time — sac_latest.zip mtime 21:22:28, sac_best.zip 21:12:50). The locked-config table above keeps the original attempt-1 launch line for the record; live relaunch differs only in this knob.

~21:30 — twin-freeze contract corrected

make_twin.sh twin launcher now exports SACLSTM_EVAL_MODE=1 (commit 19f34ab). Without lever-4's gate on the twin side, the v1 SacTwin had been TRAINING throughout, not frozen — every mirror battle updated the twin's own weights. Consequence: v1's headline "90.7% vs twin" is retroactively an arms-race win rate (both policies co-evolving), not a fixed-benchmark score, and the #57 verdict phrasing implying a frozen sparring partner is corrected by this entry. No numbers change; the story does.

~21:30 — net extended + pacing decision

Slowdown attribution complete: 89% of wall clock = eval rotation by design (SAC_EVAL_INTERVAL=2 × 3 opponents × 10 rounds each); trainer exonerated at ~500 ticks/s. Decision recorded in issue #60.

Ceiling net re-armed to match the extended budget: fires Thu 2026-08-27 07:59:40 CEST (epoch 1787810380) — supersedes the 2026-08-23 15:59:40 fire noted under Safety net. Attempt-1 artifacts preserved in /tmp/v2_attempt1_backup/.

Campaign v2 — attempt-3 (2026-08-23)

~06:45 — LEVER-5 TRIGGERED (watchman shift 2, 02:11–06:22 CEST)

Evidence-gated activation of the lever-5 conditional posture (SACLSTM_LR_CRITIC), human signed off. Trigger evidence:

Signal Observation
actor_loss |34M| → |106M| monotone (+15M/h), no plateau
critic_loss spikes >1e12 in 100% of NEW metric records; campaign share 80.2%
Signature exact 1.5625e16 recurring (= float32 saturation neighborhood)
alpha decayed 0.951 → 0.711 (entropy collapse under runaway Q scale)
best_score.txt frozen 00:32 (=83.3333) through 5h50m of composite oscillation 3.3–43.3
Per-opponent MAs whipsawing (Crazy 30→100→90→0→60→90→10→0)

DECISION — LR_CRITIC=1e-4, single knob, fresh start

  • Adjudication: primary suspect is critic/Q value-scale growth. Actor loss inherits Q magnitude through the policy gradient, so the actor explosion is downstream; alpha decay is a symptom (entropy temperature chasing a blown value scale), not a cause. Therefore ONE knob moves: SACLSTM_LR_CRITIC=1e-4 (critic learns slower → Q estimates stay estimable); LR_ACTOR/LR_ALPHA stay at default 3e-4, TARGET_ENTROPY −4. Multi-knob changes would confound attempt-3's attribution.
  • Fresh start over warm start: attempt-2's weights carry saturated Q representations; warm-starting them under a new LR would inherit the pathology we're trying to escape (contamination risk). Attempt-2 archived intact, nothing discarded.

MA-wipe root cause FOUND & FIXED (167bcc4)

Watchman anomaly explained: ma_history_*.txt held exactly one leading-space value per file (e.g. [ 10 ], [ 20 ], [ 0 ] in the attempt-2 backup) instead of the designed 5-value history, degrading the composite best-gate to last-cycle mean — which fully accounts for the composite whipsaw 3.3–43.3.

Root cause: in eval_checkpoint, the append line was

printf '%s\n' "$(cat "$(ma_hist_file "$opp")")" "$wr" | tail -n 5 | tr '\n' ' ' > "$(ma_hist_file "$opp")"

Bash sets up the > redirect before running the command substitution, so $(cat f) always read the freshly truncated file ⇒ every cycle wiped history to " $wr ". Reproduced standalone on bash 5.3 before touching the script; no other writer exists (grep). Fix hoists the read into its own statement and word-splits it so tail -n 5 keeps exactly the last MA_WINDOW values. Verified live in attempt-3: two eval cycles produced [0 0 ] per opponent (previously impossible).

Attempt-3 launch (2026-08-23 07:01:28 CEST, epoch 1787461288)

  • Attempt-2 stopped cleanly (tmux kill; no stray procs); artifacts (76 items, 336 MB incl. .part corpses + MA/best/latest/counter) → /tmp/v2_attempt2_backup/; final round_counter 8452; weights/ verified empty.
  • Twin regenerated via documented transient-seed method: throwaway Nim script (initSACTrainer(35,4) + saveWeights, hidden=256 env, alpha=1.0) seeded as transient weights/sac_best.zip → ./make_twin.sh → twin zips byte-identical to seed (cmp OK), twin counter=0 → transient removed.
  • Launch line = locked config verbatim plus the two live deltas:
SAC_OPPONENTS='Corners:3,Crazy:2,RamFire:2,Target:1,SacTwin:1' \
SAC_EVAL_OPPONENTS='Corners,Crazy,Target' \
SAC_EVAL_INTERVAL=2 SAC_EVAL_ROUNDS=10 \
SAC_TOTAL_ROUNDS=25000 SAC_CHUNK_SIZE=10 SAC_MAX_CRASHES=5 \
SACLSTM_HIDDEN_SIZE=256 SACLSTM_BATCH_SIZE=16 SACLSTM_SAVE_INTERVAL=1 \
SACLSTM_LR_CRITIC=1e-4 \
./sac_train.sh 2>&1 | tee -a campaign_v2_stdout.log
  • Safety net: existing transient unit sac-ceiling-net-v2.service re-verified active; fire epoch start+384587 s = 1787810380 = Thu 2026-08-27 07:59:40 CEST exactly (delta 0 vs mandate) — kept, not duplicated.

Health baseline (t+15 m)

Check Evidence
Counter 33 (t+8m) → 125 (t+15m)
Metrics JSONL flowing (6 lines); baseline line: critic_loss=23.09, actor_loss=-3.228, alpha_loss=0.0, alpha=0.99970 (random-init magnitudes — attempt-3's divergence curve starts here); t+15m line: critic 84.5, actor −9.99, alpha 0.9937
Saves SAVE_INTERVAL=1: sac_latest.zip written from first grad-step onward
Eval rotation full cycle landed: Corners/Crazy/Target ×10 games each in eval_log.jsonl
MA files [0 0 ] per opponent after two cycles — fix confirmed in vivo
Best gate best_score.txt=0.0000 written on first composite (0 > −1 default)
Crash banners 0

Progress graphs

One live dashboard: docs/campaign_dashboard.svg (current run only — test wins, real-fight wins, losses, alpha, throughput; auto-reloads every 60 s when open in Chrome). Keep it fresh with tools/watch_dashboard.sh (regenerates every 60 s), or one-shot python3 tools/plot_progress.py (pure stdlib; paths overridable via argv, --selftest for sanity check).

~10:47 — the dashboard's axes were upside-down since creation

The progress graphs have been lying since they were made: 0% was drawn at the TOP of every panel and the newest games appeared on the LEFT. The cause is a one-line formula bug in tools/plot_progress.py: map_fn interpolated as p1 - t*(p1-p0) instead of p0 + t*(p1-p0), so all five panels plotted 100 - value on y and reversed time on x. Tick labels were computed by separate (correct) code, which is why the numbers on the axes never matched the ink.

Fix + guard: formula corrected; every call site audited (panels 1/2/5 use map_fn for both axes and are fixed by the same line; panels 3/4 already used a correct local x-lambda; no other consumer of map_fn exists in the repo). --selftest now renders a known rising series through the full build path and fails loudly unless higher value = smaller SVG y and newer data = further right — proven to catch this exact bug when the old formula is re-injected. Dashboard regenerated from live logs.

Honest status while reading the now-correct charts: run 3 is only hours old and winning ~0% — recent evals are 0/10 vs Corners, Crazy and Target alike, and real-fight buckets sit at 0–1 wins per 100 games. Expected for a fresh brain. Night-1's gains were real but were intentionally reset by the stability restart that began attempt-3; the curve starts from zero again here.