Campaign-v2 pacing & durability decisions #60
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Decision record for campaign-v2 pacing + checkpoint durability. Relaunch of the campaign occurred 2026-08-22 21:09:30 CEST under the same locked config except
SACLSTM_SAVE_INTERVAL=5→1(see §2). Leave this issue OPEN for the campaign's duration — more pacing decisions may append.1. Slowdown: 5.4× aggregate, eval-tripling dominant, trainer exonerated
Orchestrator measurement (pre-session): campaign-v2 runs at ~5.4× the wall-clock cost per training-round vs the v1 baseline, aggregate. Attribution: eval-tripling dominates — v2 evaluates 3 opponents × 10 deterministic rounds every 2 chunks (30 eval rounds per 20 training rounds) where v1 ran 1×10; eval rounds are also longer (deterministic policy survives: eval tick counts 1200–2900 vs training 400–1500 in
campaign_v2_stdout.log). Trainer exonerated: training battles complete their full round/tick budgets at full speed; the trainer's CPU burn (OMP ~12–16 threads observed at ~85% each) does not starve round completion.DECISION: eval rotation stays AT FULL SPEC. Measurement quality over speed — v1's lesson was that a single-opponent deterministic eval + single-max best gating froze
sac_best.zipon a fluke 10/10 for 13 h (#57 best-zip paradox). Cutting eval quality to buy speed would reintroduce that pathology. Honest ETA at the observed pace: 65–82 h, hence §3.2. Checkpoint cadence: STUCK → root-caused → restarted with
SACLSTM_SAVE_INTERVAL=1Verdict: STUCK (attempt 1, launch 20:00:19, killed 21:07)
weights/sac_latest.zipmtime: 20:17:59 → 20:57:53 — one ~40-min dead gap across ~10 completed chunk processes. Only 2 zip writes in 67 min, neither carrying new learning (see below).training_metrics.jsonl(lever 3): all 36 lines showsteps=1, drained=1, grad_steps=1, buffer_size=24— cumulativestepCountnever exceeded 1 in any process. The save gate (stepCount >= nextSave, interval 5, checked per grad step — the #56 mid-loop fix is present and correct) never fired.sac_latest.zipwith the boot-time snapshot — durability theater, zero gradient work persisted. All gradient work was ephemeral (in-RAM per process).SACLSTM_SAVE_INTERVALcounts gradient steps (integration.nimnextSave/stepCount, inc'd once persacUpdate), not trainPasses, chunks, or rounds.sac_latest.zip.tmp.*.partSIGKILL corpses (20:00:31, 20:14:56, 20:17:48).Root cause
Per-process gradient-step budget collapsed to ~1–2 steps (v1 measured ~10–18/process at the same hidden 256). Every drain burst carries exactly 1 transition (
drained=1) and the buffer sits pinned at exactly thecanSamplethreshold (burnIn 8 + trainWindow 16 = 24) at every logged pass — the producer/consumer dynamics (per-tick sends attps=-1vs ~1 s/step OMP-heavy updates) throttle each process to ~2 passes before teardown. Interval 5 can never fire inside a 1–2-step process. The signature was visible at launch (notebook watch-item #1) but its consequence — zero learning persistence — is confirmed now.Action taken (env-level, per charter preference; restart cheap at ~1% budget)
sac_campaign_v2+ harness (old session id 3800566 tree; new session created 21:09:30 CEST)./tmp/v2_attempt1_backup/(38 MB: weights/, all JSONL logs, stdout log). Nothing archived into the repo — attempt-1 state is noise.weights/emptied ⇒ genuine random init. No counter reset hack needed — fresh dir resetsround_counter.txtsemantics naturally.alpha=1.0) →make_twin.sh→ byte-identical twin zips verified (cmpOK) → transient removed. Twinround_counter.txt=0.SACLSTM_SAVE_INTERVAL=1(env-only, no code change): save gate fires on every gradient step ⇒ ~2 saves/battle ≈ every ~1–3 min at observed pacing. IO trivial (10.5 MB atomic zip).19f34ab):SacTwin.shnow exportsSACLSTM_EVAL_MODE=1— lever-4 gate suppresses all twin training traffic, keeping the twin the frozen reproducible opponent #54 specifies. Without it, interval=1 would have let the twin persist per-battle drift and silently become a co-learner. Verified in production: chunk 3 ran vs SacTwin, twin produced zero new metrics lines, twin zip untouched since seed.round_counter.txtadvancing (850), 4 metrics lines (1/chunk, same pacing as before — fix targets durability, not step rate), zero crash banners.3. Ceiling net extended
Old unit (armed 19:59:40 Aug 22, fire epoch 1787493580 = 15:59:40 Aug 23) stopped; new
sac-ceiling-net-v2armed 21:09:53 CEST:sleep 384587⇒ fires epoch 1787810380 = Thu 2026-08-27 07:59:40 CEST, thentmux kill-session -t sac_campaign_v2; sleep 5; pkill -f '[s]ac_train.sh'(bracket-guarded so the unit can't self-match). Unit verifiedactive (running).⚠️ Epoch-vs-prose discrepancy flagged: the mandate's prose said "~Mon Aug 25 23:59 CEST" but its exact arithmetic (1787493580+316800) yields Aug 27 07:59:40 CEST (+88 h on top of the original net's +20 h ⇒ launch+108 h). Executed per the explicit epoch formula as chartered; re-arm if Mon was the true intent.
4. Critic-spike watch
Attempt-1 logged four
critic_loss=1.5625e+16spikes (20:06:35, 20:12:23, 20:42:35, 20:57:56 — mandate said two; log shows four) amid otherwise sane values (~8–35). One spike already recurred in the fresh random-init run (4th line, 21:0x). Classified random-init transient (same class as #59 phase-2's ~1e12 artifact): expect decay as the critic fits. Monitor via lever-3 curves (training_metrics.jsonl); if persistent past early chunks, feeds the lever-5 (stability knobs) decision.Health snapshot at close of this action (21:15 CEST, t+6 min)
sac_campaign_v2(new), 0 crash bannersround_counter.txtsac_latest.ziptraining_metrics.jsonlsteps=1pacing unchangedsac-ceiling-net-v2active, fires epoch 1787810380 (Aug 27 07:59:40 CEST)Attempt-3 executed 2026-08-23 (lever-5 trigger fired; human signed off)
Lever-5 TRIGGERED — watchman shift-2 evidence (02:11–06:22 CEST)
|actor_loss|34M→106M monotone (+15M/h, no plateau); critic spikes >1e12 in 100% of new metric records (exact1.5625e16signature recurring); alpha decayed 0.951→0.711;best_score.txtfrozen at 83.3333 since 00:32 through 5h50m of composite oscillation 3.3–43.3.DECISION — single knob
SACLSTM_LR_CRITIC=1e-4, fresh startPrimary suspect: critic/Q value-scale growth. Actor explosion is downstream (actor loss inherits Q magnitude via the policy gradient); alpha decay is a symptom (entropy temperature chasing a blown value scale), not a cause ⇒ LR_ACTOR/LR_ALPHA stay 3e-4, TARGET_ENTROPY −4. One knob keeps attempt-3 attributable. Fresh start over warm start: attempt-2 weights carry saturated Q representations — warm-start would inherit the pathology.
MA-wipe anomaly ROOT-CAUSED & FIXED (
167bcc4)The separate watchman anomaly (ma_history files single-valued with leading-space artifacts) was a self-truncation bug in
eval_checkpoint: bash sets up the>redirect before running the command substitution, so$(cat f) … > fread the freshly truncated file every cycle ⇒ history wiped to one value ⇒ composite gate degraded to last-cycle mean (which fully explains the composite whipsaw). Reproduced standalone on bash 5.3 before touching the script; no other writer exists. Fix hoists the read into its own statement; verified live post-launch (two cycles →[0 0 ]per opponent).Attempt-3 launch — 07:01:28 CEST (epoch 1787461288)
/tmp/v2_attempt2_backup/(76 items, 336 MB); final counter 8452;weights/empty ⇒ genuine random init.cmpOK), twin counter=0, EVAL_MODE=1 freeze in force.SACLSTM_SAVE_INTERVAL=1(§2) +SACLSTM_LR_CRITIC=1e-4(this decision).sac-ceiling-net-v2.serviceactive, fire = start+384587 s = epoch 1787810380 = Thu Aug 27 07:59:40 CEST exactly (delta 0).Health baseline t+15m
Counter 33→125; metrics flowing — baseline line
critic_loss=23.09, actor_loss=−3.228, alpha_loss=0.0, alpha=0.9997(attempt-3 divergence curve starts here); t+15m critic 84.5 / alpha 0.9937; saves flowing from first grad step; full eval rotation landed (Corners/Crazy/Target ×10); best gate wrote 0.0000 on first composite; 0 crash banners.Commits:
167bcc4(fix),40e074e(notebook chapter). Issue stays OPEN per header.Lever-5 outcome:
SACLSTM_LR_CRITIC=1e-4does NOT cure divergence — LR ruled outWatchman adjudication confirmed by full-run telemetry. Attempt-3 was cut at 11:48 CEST today rather than burn ~57 h to certain re-divergence.
Key evidence from
training_metrics.jsonl(281 records, archived):Orchestrator decision
sac-ceiling-net-v2cancelled — unit already absent from user systemd (no service/timer), so nothing fires Thursday.Artifacts preserved (nothing deleted)
/tmp/v2_attempt3_backup/(293 MB):weights/—sac_best.zip,sac_latest.zip,best_score.txt(14.6667),round_counter.txt(6951), 3×ma_history_*.txt, + 53 orphanedsac_latest.zip.tmp.*.partpartial writes (kept as durability evidence)campaign_v2_stdout.log(10 161 lines),training_log.jsonl(2 800),training_metrics.jsonl(281),eval_log.jsonl(30) + 2 eval.tmppartialsweights/left empty for the next attempt.Session log — Aug 24 evening (speedup attempt + instability)
Changes deployed
integration.nimenv defaults) — grad_steps confirmed 3 in metricssac_train.sh) — fewer JVM bootsOPENBLAS_NUM_THREADS=1added to top ofsac_train.sh(bot segfaulted insgemv_kernel_4x4during shutdown race)New: training-state persistence (
training_state.nim)weights/training_state.bin(~13.7MB), atomic writeBug found & fixed (needs verification)
lastEnemyKey=""after restart → firsttmkNewBattlecleared the loaded buffer. Fix: persist/restorelastEnemyKey(backward-compatible format). Not yet verified in live run.Open problems
Tomorrow's plan
buf_size > 0and restore log lines appear