Campaign verdict: notebook, results table, scale-or-not call #57

Closed
opened 2026-08-21 22:37:21 +02:00 by SirStone · 3 comments
Owner

Blocked by: #56

Child of the map Overnight Training Campaign — SAC_LSTM_Bot.

Question

What happened overnight, and does the approach scale? Compile the markdown notebook (the story: config → phases → each milestone-decision issue link → win-rate curves vs pool + twin → incidents and interventions), finalize checkpoints, and write the verdict issue: continue as-is / retune / escalate. Closing this closes the campaign destination: a human-readable morning story ending in a call.

*Blocked by: #56* Child of the map *Overnight Training Campaign — SAC_LSTM_Bot*. ## Question What happened overnight, and does the approach scale? Compile the markdown notebook (the story: config → phases → each milestone-decision issue link → win-rate curves vs pool + twin → incidents and interventions), finalize checkpoints, and write the verdict issue: continue as-is / retune / escalate. Closing this closes the campaign destination: a human-readable morning story ending in a call.
SirStone added the wayfinder:task label 2026-08-21 22:37:21 +02:00
SirStone changed title from Morning verdict: notebook, results table, scale-or-not call to Campaign verdict: notebook, results table, scale-or-not call 2026-08-21 22:42:51 +02:00
Author
Owner

Claimed for implementation (self-assigned as SirStone — MCP assignee field unreliable, this comment is the assignment record). Working on branch research/goto-controller. Campaign ended naturally at 2500/2500 chunks (~14:28 CEST, zero crashes) — this ticket now compiles the notebook ending, the results table, and the scale-or-not verdict, then closes the campaign set (#58 retrospective close + map update).

Claimed for implementation (self-assigned as `SirStone` — MCP assignee field unreliable, this comment is the assignment record). Working on branch `research/goto-controller`. Campaign ended naturally at 2500/2500 chunks (~14:28 CEST, zero crashes) — this ticket now compiles the notebook ending, the results table, and the scale-or-not verdict, then closes the campaign set (#58 retrospective close + map update).
Author
Owner

Resolution — Campaign verdict: RETUNE BEFORE SCALING

Campaign-v1 completed naturally at ~14:28 CEST 2026-08-22: banner >>> training complete: 2500 chunks, round_counter 38912, runtime 14 h 07 m (00:21:55 → ~14:28), zero crashes, graceful exit, systemd backstop never fired. Budget was chunk-based — SAC_TOTAL_ROUNDS=25000 ÷ CHUNK_SIZE=10 ⇒ 2500 chunk battles; the "25k-rounds" label was a misnomer (documented in the notebook). Notebook ending + results table: 2619ba0 (SAC_LSTM_Bot/docs/campaign_notebook.md); hygiene fix: f45e8f2.

Ops layer: PROVEN

14 h autonomous run, zero crashes, self-healing crash-recovery restarts (rerun-from-checkpoint), natural completion on its own budget, and the hard ceiling net never needed. The harness scales — chunked self-play + liveness counter + atomic checkpoints held up for a full overnight campaign unattended.

Learning layer: REAL BUT NARROW

Genuine within-opponent gains prove training works: SacTwin 90.7% (2693/2970 vs frozen past-self) and Crazy doubled 26% → 48% (2965/6140). But three caps define v1:

  • Specialist-not-generalist: walls untouched — Corners 7.2% training win rate; Target 9.6%; RamFire 0/2920.
  • Instability: 12 eval spikes ≥8/10 including 5× perfect 10/10 vs Corners — capability emerged repeatedly and none of it was retained by selection.
  • Texture: survivor-not-fighter against static opponents (score:60 anatomy — 19% of rounds damageless outlives; 43% score:0), consistent with event-driven rewards meeting passive opponents.

Benchmark pathology

Corners-only deterministic eval + single-max best gating = fragile selection. sac_best.zip has been frozen since 01:01:57 — one lucky 10/10 at ~round 3.5k wrote best_score=100, and nothing could ever outrank a perfect score again (steady-state eval sat at 0–10%, mean 7.7% over 1287 evals; histogram: 0/10 = 867, i.e. 67%). Best checkpoint = lottery ticket, decoupled from the steady-state policy.

VERDICT: RETUNE BEFORE SCALING

Do not pour more compute into the v1 config. Five concrete levers for campaign v2, each small and code-level, AWAITING HUMAN SIGN-OFF (v2 has not been launched):

  1. Eval rotation across pool + moving-average best gating instead of single-max on one fixed opponent — kills the lottery-ticket failure mode.
  2. Reward shaping toward damage/aggression incl. an anti-ram signal — RamFire 0/2920 mirrors the PPO-era ladder; ram-class opponents need dedicated pressure, not just exposure.
  3. Training-loss/step metrics logged from the training thread (#55 gap #1) — the next campaign must not be blind.
  4. Gate sendTrainingMsg off in eval mode (#55 gap #3).
  5. Optional stability knobs (lower LR or entropy coefficient) — only once loss curves exist to justify them.

Deferred-intervention retrospective

Shift-1 deferral (#58) is vindicated: the regression threshold fired on n=2 evidence with zero loss visibility — intervening then would have been knob-twiddling. The run completed cleanly instead and delivered full-curve evidence, which is exactly what this verdict consumes.

Closing #57 → closes the campaign destination. Map updated on #53.

## Resolution — Campaign verdict: **RETUNE BEFORE SCALING** Campaign-v1 completed naturally at ~14:28 CEST 2026-08-22: banner `>>> training complete: 2500 chunks`, round_counter 38912, runtime **14 h 07 m** (00:21:55 → ~14:28), **zero crashes**, graceful exit, systemd backstop never fired. Budget was **chunk-based** — `SAC_TOTAL_ROUNDS=25000 ÷ CHUNK_SIZE=10 ⇒ 2500 chunk battles`; the "25k-rounds" label was a misnomer (documented in the notebook). Notebook ending + results table: `2619ba0` (`SAC_LSTM_Bot/docs/campaign_notebook.md`); hygiene fix: `f45e8f2`. ### Ops layer: PROVEN 14 h autonomous run, zero crashes, self-healing crash-recovery restarts (rerun-from-checkpoint), natural completion on its own budget, and the hard ceiling net never needed. The harness scales — chunked self-play + liveness counter + atomic checkpoints held up for a full overnight campaign unattended. ### Learning layer: REAL BUT NARROW Genuine within-opponent gains prove training works: **SacTwin 90.7%** (2693/2970 vs frozen past-self) and **Crazy doubled 26% → 48%** (2965/6140). But three caps define v1: - **Specialist-not-generalist**: walls untouched — Corners 7.2% training win rate; Target 9.6%; RamFire 0/2920. - **Instability**: 12 eval spikes ≥8/10 including 5× perfect 10/10 vs Corners — capability emerged repeatedly and none of it was retained by selection. - **Texture**: survivor-not-fighter against static opponents (score:60 anatomy — 19% of rounds damageless outlives; 43% score:0), consistent with event-driven rewards meeting passive opponents. ### Benchmark pathology Corners-only deterministic eval + single-max best gating = fragile selection. `sac_best.zip` has been **frozen since 01:01:57** — one lucky 10/10 at ~round 3.5k wrote `best_score=100`, and nothing could ever outrank a perfect score again (steady-state eval sat at 0–10%, mean 7.7% over 1287 evals; histogram: 0/10 = 867, i.e. 67%). Best checkpoint = lottery ticket, decoupled from the steady-state policy. ### VERDICT: RETUNE BEFORE SCALING Do **not** pour more compute into the v1 config. Five concrete levers for campaign v2, each small and code-level, **AWAITING HUMAN SIGN-OFF** (v2 has not been launched): 1. **Eval rotation across pool + moving-average best gating** instead of single-max on one fixed opponent — kills the lottery-ticket failure mode. 2. **Reward shaping toward damage/aggression incl. an anti-ram signal** — RamFire 0/2920 mirrors the PPO-era ladder; ram-class opponents need dedicated pressure, not just exposure. 3. **Training-loss/step metrics logged from the training thread** (#55 gap #1) — the next campaign must not be blind. 4. **Gate `sendTrainingMsg` off in eval mode** (#55 gap #3). 5. **Optional stability knobs** (lower LR or entropy coefficient) — only once loss curves exist to justify them. ### Deferred-intervention retrospective Shift-1 deferral (#58) is **vindicated**: the regression threshold fired on n=2 evidence with zero loss visibility — intervening then would have been knob-twiddling. The run completed cleanly instead and delivered full-curve evidence, which is exactly what this verdict consumes. Closing #57 → closes the campaign destination. Map updated on #53.
Author
Owner

Human sign-off received (2026-08-22, late evening): ALL FIVE levers approved.

Execution order: 3 (loss instrumentation) → 4 (eval-mode training gate) → 1 (eval rotation + MA best-gating) → 2 (aggression/anti-ram reward shaping) → 5 (stability knobs, conditional on loss curves).

Implementation tracked in the v2-readiness issue.

Human sign-off received (2026-08-22, late evening): ALL FIVE levers approved. Execution order: 3 (loss instrumentation) → 4 (eval-mode training gate) → 1 (eval rotation + MA best-gating) → 2 (aggression/anti-ram reward shaping) → 5 (stability knobs, conditional on loss curves). Implementation tracked in the v2-readiness issue.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#57