Campaign-v1 watch: eval-regression threshold fired — intervention deferred to verdict #58

Closed
opened 2026-08-22 07:14:15 +02:00 by SirStone · 1 comment
Owner

Observation (watch shift 1, 2026-08-22 00:34–03:35 CEST)

Campaign-v1 live in tmux sac_campaign since 00:21:55 (locked config: #56 resolution comment 475, mirrored in SAC_LSTM_Bot/docs/campaign_notebook.md). Liveness flawless across 10 checks / 3 h: zero crash/abort banners, sac_latest.zip never stale >28 s, round counter ~3250 rounds/h, evals on cadence (341 total by shift end).

Eval timeline (opponent pinned Corners, SAC_EVAL_ROUNDS=10; stdout carries no timestamps — times derived from round-counter rate, #55 gap #5):

Check time (CEST) Round counter Eval state
00:30 (t+9 min health check) 1890 HEALTHY checklist passed; early evals 0% (expected for near-random policy)
00:34 (shift start) 2130 steady-state WR vs Corners 0–5%, evals on cadence
~01:15 3570 eval 10/10 (100%) → >>> [eval] new best (100%) -> sac_best.zip → best_score.txt=100
~02:50 ±15 min 9018 eval 10/10 (100%) again (best already 100, no banner possible)
01:36 onward — SLOW-LEARNER watch flag raised: best_score flat at 100 while typical eval 0–5%
03:35 (shift end) 10520 341 evals total, mean WR 0.6% (last-60-eval mean 0.8%); two isolated perfect spikes

Evals adjacent to the spikes were also elevated (7/10 @ r3720, 8/10 @ r4500, 7/10 @ r6960) before reverting to the 0–5% baseline. best_score.txt=100 is a spike, not steady state.

Fired threshold (source: #55 comment 461)

Eval score: 'watch' = best_score.txt flat over ≥5 consecutive evals (≈10 chunks). 'intervene' = 3 consecutive evals each scoring <50% of the established best after a best ≥10% was reached (real regression), or S9 fired (silent random re-init). Best-not-improving alone is NEVER intervene.

With best=100 established at ~01:15, essentially every subsequent eval (0–5%) scores far below 50% of best ⇒ the intervene condition has been continuously satisfied since shortly after the first spike. Technically fired.

Decision — DEFER intervention to Campaign verdict (#57); no mid-run retune

Rationale:

  1. n=2 evidence: two isolated spikes prove capability emerged, not that it is stable — insufficient evidence to pick a knob.
  2. No loss metrics: #55 gap #1 — no actor/critic loss, alpha, gradient-step or buffer telemetry exists anywhere; knob choice would be blind.
  3. Buffer loss on restart: a restart discards the entire in-memory replay buffer (~50k+ transitions accumulated overnight).
  4. Run self-completes: at ~3250 rounds/h the 25000-round target lands ~08:05 CEST, before the 12:22 CEST kill ceiling — the full-curve evidence arrives regardless.
  5. Human directive: campaign-autonomy mandate — no forcing; the story runs at its own pace.

Decision owner: #57 (Campaign verdict). This issue intentionally stays OPEN until the verdict consumes it with full-curve evidence.

Context: map #53 · observability inventory #55 · campaign config/launch #56 (resolution comment 475) · morning verdict #57.

## Observation (watch shift 1, 2026-08-22 00:34–03:35 CEST) Campaign-v1 live in tmux `sac_campaign` since 00:21:55 (locked config: #56 resolution comment 475, mirrored in `SAC_LSTM_Bot/docs/campaign_notebook.md`). Liveness flawless across 10 checks / 3 h: zero crash/abort banners, `sac_latest.zip` never stale >28 s, round counter ~3250 rounds/h, evals on cadence (341 total by shift end). Eval timeline (opponent pinned Corners, `SAC_EVAL_ROUNDS=10`; stdout carries no timestamps — times derived from round-counter rate, #55 gap #5): | Check time (CEST) | Round counter | Eval state | |-------------------|---------------|------------| | 00:30 (t+9 min health check) | 1890 | HEALTHY checklist passed; early evals 0% (expected for near-random policy) | | 00:34 (shift start) | 2130 | steady-state WR vs Corners 0–5%, evals on cadence | | ~01:15 | 3570 | **eval 10/10 (100%)** → `>>> [eval] new best (100%) -> sac_best.zip` → `best_score.txt`=100 | | ~02:50 ±15 min | 9018 | **eval 10/10 (100%)** again (best already 100, no banner possible) | | 01:36 onward | — | SLOW-LEARNER watch flag raised: best_score flat at 100 while typical eval 0–5% | | 03:35 (shift end) | 10520 | 341 evals total, mean WR **0.6%** (last-60-eval mean 0.8%); two isolated perfect spikes | Evals adjacent to the spikes were also elevated (7/10 @ r3720, 8/10 @ r4500, 7/10 @ r6960) before reverting to the 0–5% baseline. `best_score.txt`=100 is a spike, not steady state. ## Fired threshold (source: [#55 comment 461](https://git.fossellini.top/SirStone/SirRoboGarage/issues/55#issuecomment-461)) > **Eval score**: **'watch'** = `best_score.txt` flat over **≥5 consecutive evals** (≈10 chunks). **'intervene'** = 3 consecutive evals each scoring **<50% of the established best** after a best ≥10% was reached (real regression), or S9 fired (silent random re-init). Best-not-improving alone is NEVER intervene. With best=100 established at ~01:15, essentially every subsequent eval (0–5%) scores far below 50% of best ⇒ the intervene condition has been continuously satisfied since shortly after the first spike. Technically fired. ## Decision — DEFER intervention to Campaign verdict (#57); no mid-run retune Rationale: 1. **n=2 evidence**: two isolated spikes prove capability emerged, not that it is stable — insufficient evidence to pick a knob. 2. **No loss metrics**: #55 gap #1 — no actor/critic loss, alpha, gradient-step or buffer telemetry exists anywhere; knob choice would be blind. 3. **Buffer loss on restart**: a restart discards the entire in-memory replay buffer (~50k+ transitions accumulated overnight). 4. **Run self-completes**: at ~3250 rounds/h the 25000-round target lands ~08:05 CEST, before the 12:22 CEST kill ceiling — the full-curve evidence arrives regardless. 5. **Human directive**: campaign-autonomy mandate — no forcing; the story runs at its own pace. **Decision owner**: #57 (*Campaign verdict*). This issue intentionally stays OPEN until the verdict consumes it with full-curve evidence. Context: map #53 · observability inventory #55 · campaign config/launch #56 (resolution comment 475) · morning verdict #57.
Author
Owner

Outcome recorded — superseded by the #57 resolution (comment 490). The deferral rationale held: intervening at shift-1 would have been knob-twiddling without loss visibility; instead the run completed cleanly (~14:28 CEST, 2500/2500 chunks, zero crashes) and delivered full-curve evidence, consumed by the verdict RETUNE BEFORE SCALING with five code-level levers staged for human sign-off. Closing as resolved-by-verdict.

Outcome recorded — superseded by the #57 resolution (comment 490). The deferral rationale held: intervening at shift-1 would have been knob-twiddling without loss visibility; instead the run completed cleanly (~14:28 CEST, 2500/2500 chunks, zero crashes) and delivered full-curve evidence, consumed by the verdict **RETUNE BEFORE SCALING** with five code-level levers staged for human sign-off. Closing as resolved-by-verdict.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#58