Campaign v2 readiness: implement approved levers #59
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Child of the campaign verdict #57 (resolution comment 490; human sign-off recorded 2026-08-22, late evening).
The five approved levers (verbatim from the #57 resolution)
SACLSTM_EVAL_MODE=1path).SAC_EVAL_OPPONENTS) each eval cycle; gatesac_best.zipwrites on a moving-average composite score instead of single-opponent win rate.Execution order
3 → 4 → 1 → 2 → 5
Note: lever 5 is conditional (config support only, activated by evidence from lever 1's loss curves).
Acceptance criteria
nimble testgreen (8 suites)Claiming this issue (self-assigned SirStone) for phase 1 of the v2-readiness work: levers 3 → 4 → 1 in one measurement-hygiene pass on
research/goto-controller. Levers 2 and 5 follow in phase 2.Phase 1 complete — levers 3, 4, 1 implemented (execution order per sign-off), commit
a07e530onresearch/goto-controller.Lever 3 — loss instrumentation:
sacUpdatealready returnedSACMetrics(criticLoss, actorLoss, alphaLoss, alpha) — no trainer change needed.trainPassnow averages those over the pass's gradient steps and appends ONE JSONL line per pass toSAC_LSTM_Bot/training_metrics.jsonl:{"epoch", "steps", "buffer_size" (replay_buffer.len), "drained", "grad_steps", "critic_loss", "actor_loss", "alpha_loss", "alpha"}. Best-effort writes (failure never kills the training thread).Lever 4 — eval-mode gate: mechanism already existed —
sac_train.sheval battles run withSACLSTM_EVAL_MODE=1(#49).sendTrainingMsgnow drops ALL training messages while it is set (transitions AND NewBattle, so eval can't even clear the buffer); one-time stderr notice at init. Unit-tested.Lever 1 — eval rotation + MA best-gating:
SAC_EVAL_OPPONENTS(defaultCorners,Crazy,Target) evaluated per cycle; per-opponent MA over last 5 evals; composite = mean of MAs;sac_best.zipwritten only on strict improvement.best_score.txtFORMAT CHANGE: float composite replaces retired single-opponent integer win rate. Crash-restart/liveness untouched.Verification:
nimble build+ all 8nimble testsuites green (test_integration extended: metricsLine JSONL scalars + suppression asserts). Tmux smoke (SAC_TOTAL_ROUNDS=10, hidden 32): metrics lines appeared; rotation hit Corners+Crazy (eval_log carries both); composite 50.0000 = mean(MA 0%, MA 100%) → new-best write fired under new semantics; zero metrics/buffer activity during eval battles. Note: smoke surfaced that at ~29 ms/gradient-step even tiny nets drain a full 256-msg channel burst in ~7.5 s — throughput headroom is a lever-5/loss-curve question, not a defect in these levers.Levers 2 then 5 remain (phase 2).
Phase 2 complete — lever 2 (aggression/anti-ram shaping) + lever 5 status
Commit:
6fc01eb(part 1 wasa07e530). 3 files, +98/−7.Reward delta table (raw, pre-normalization)
AggressionMult; p=3: 16→20HitBonus, discrete accuracy signalChargePenalty·(1−d/thr)), suppressed while dealing damage that stepAll new weights are TUNABLE consts in rewards.nim marked
# ponytail:with upgrade paths.Smoke evidence (tmux, hidden=32 random-init isolated weights, 3 rounds vs RamFire + 3 vs Crazy)
SACLSTM_REWARD_DEBUG=1→ reward_debug.log (file, because the battle runner swallows bot stderr).Lever 5 — already supported, no diff
training.nim has had env knobs since #48:
SACLSTM_LR_ACTOR/SACLSTM_LR_CRITIC/SACLSTM_LR_ALPHA(default 3e-4 each) andSACLSTM_TARGET_ENTROPY(default −4.0). Mid-campaign LR/entropy tuning needs no code edits; the proposed SACLSTM_LR/SACLSTM_ENTROPY_COEFF aliases were skipped as redundant.Open items for campaign v2 launch
Campaign v2 LAUNCHED — all levers shipped, acceptance met, ticket resolved
Levers: all five shipped — part 1
a07e530(levers 3, 4→gate, 1), part 26fc01eb(lever 2 shaping; lever 5 confirmed already-supported via env knobs, conditional posture preserved). Prereqs verified at launch: both commits onresearch/goto-controller,nimble testall 8 suites green.Fresh start
SAC_LSTM_Bot/weights_v1_archive/(sac_latest/sac_best zips, best_score.txt, round_counter.txt, training_log.jsonl, eval_log.jsonl, campaign_stdout.log, + smoke-leftover training_metrics.jsonl so v2 loss curves start clean). Mainweights/empty ⇒ bot took the genuine random-init path (randomFull()).cmpOK), NOT v1 zips; round_counter=0. No script changes required.Launch
sac_campaign_v2, 2026-08-22 20:00:19 CEST, config exactly as locked by the orchestrator (RamFire weight 2 noted for the anti-ram-exposure rationale).sac-ceiling-net-v2active, fires 2026-08-23 15:59:40 CEST (+20 h).Health evidence (first ~45 min, 20 chunks)
training_metrics.jsonlflowing with full scalar sets (lever 3 live).steps=1/buf=24, early critic spikes ~1e16,.partcorpses): none launch-blocking; they are exactly what lever-3 curves exist to judge at first review.Notebook chapter
## Campaign v2: commitbd58794.Acceptance criteria met (levers implemented ✓, tests green ✓, smoke validated artifacts ✓ per phase comments, v2 launched ✓). Closing.