Campaign verdict: notebook, results table, scale-or-not call #57
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Blocked by: #56
Child of the map Overnight Training Campaign — SAC_LSTM_Bot.
Question
What happened overnight, and does the approach scale? Compile the markdown notebook (the story: config → phases → each milestone-decision issue link → win-rate curves vs pool + twin → incidents and interventions), finalize checkpoints, and write the verdict issue: continue as-is / retune / escalate. Closing this closes the campaign destination: a human-readable morning story ending in a call.
Morning verdict: notebook, results table, scale-or-not callto Campaign verdict: notebook, results table, scale-or-not callClaimed for implementation (self-assigned as
SirStone— MCP assignee field unreliable, this comment is the assignment record). Working on branchresearch/goto-controller. Campaign ended naturally at 2500/2500 chunks (~14:28 CEST, zero crashes) — this ticket now compiles the notebook ending, the results table, and the scale-or-not verdict, then closes the campaign set (#58 retrospective close + map update).Resolution — Campaign verdict: RETUNE BEFORE SCALING
Campaign-v1 completed naturally at ~14:28 CEST 2026-08-22: banner
>>> training complete: 2500 chunks, round_counter 38912, runtime 14 h 07 m (00:21:55 → ~14:28), zero crashes, graceful exit, systemd backstop never fired. Budget was chunk-based —SAC_TOTAL_ROUNDS=25000 ÷ CHUNK_SIZE=10 ⇒ 2500 chunk battles; the "25k-rounds" label was a misnomer (documented in the notebook). Notebook ending + results table:2619ba0(SAC_LSTM_Bot/docs/campaign_notebook.md); hygiene fix:f45e8f2.Ops layer: PROVEN
14 h autonomous run, zero crashes, self-healing crash-recovery restarts (rerun-from-checkpoint), natural completion on its own budget, and the hard ceiling net never needed. The harness scales — chunked self-play + liveness counter + atomic checkpoints held up for a full overnight campaign unattended.
Learning layer: REAL BUT NARROW
Genuine within-opponent gains prove training works: SacTwin 90.7% (2693/2970 vs frozen past-self) and Crazy doubled 26% → 48% (2965/6140). But three caps define v1:
Benchmark pathology
Corners-only deterministic eval + single-max best gating = fragile selection.
sac_best.ziphas been frozen since 01:01:57 — one lucky 10/10 at ~round 3.5k wrotebest_score=100, and nothing could ever outrank a perfect score again (steady-state eval sat at 0–10%, mean 7.7% over 1287 evals; histogram: 0/10 = 867, i.e. 67%). Best checkpoint = lottery ticket, decoupled from the steady-state policy.VERDICT: RETUNE BEFORE SCALING
Do not pour more compute into the v1 config. Five concrete levers for campaign v2, each small and code-level, AWAITING HUMAN SIGN-OFF (v2 has not been launched):
sendTrainingMsgoff in eval mode (#55 gap #3).Deferred-intervention retrospective
Shift-1 deferral (#58) is vindicated: the regression threshold fired on n=2 evidence with zero loss visibility — intervening then would have been knob-twiddling. The run completed cleanly instead and delivered full-curve evidence, which is exactly what this verdict consumes.
Closing #57 → closes the campaign destination. Map updated on #53.
Human sign-off received (2026-08-22, late evening): ALL FIVE levers approved.
Execution order: 3 (loss instrumentation) → 4 (eval-mode training gate) → 1 (eval rotation + MA best-gating) → 2 (aggression/anti-ram reward shaping) → 5 (stability knobs, conditional on loss curves).
Implementation tracked in the v2-readiness issue.