Observability inventory: what can we watch mid-run? #55
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Child of the map Overnight Training Campaign — SAC_LSTM_Bot.
Question
What training/health signals does a running campaign expose, and how do we distinguish "crashed/stalled" from "slow learner" from "healthy"? Inventory from local resources only (
sac_train.sh,integration.nimlogging,tools/training_runner/RunTraining.javaoutput, weights/ artifacts likebest_score.txtandround_counter.txt, eval logs): signal list with format/location/update cadence; then propose the minimal watch-set plus concrete unhealthy thresholds (e.g. battle-chunk crash loop, zero transitions over N minutes, frozen round counter, eval-score regression window). Consumed by Lock campaign-v1 config and launch overnight run to define its monitoring loop.Resolution — observability inventory (local sources only)
Sources read:
SAC_LSTM_Bot/sac_train.sh,src/SAC_LSTM_Bot/integration.nim,src/SAC_LSTM_Bot.nim,network.nim(isEvalMode),weights.nim(save mechanics),tools/training_runner/RunTraining.java, plus real smoke-run artifacts from #49 (weights/*,training_log.jsonl,eval_log.jsonl).Signal table
tee)=== Chunk i/N …,crash #k — restarting chunk,>>> [eval] win rate: W/R (P%),new best (P%),>>> aborted: …new bestaborted: N consecutive crashes,[eval] crashed,[eval] no resultsSAC_LSTM_Bot/training_log.jsonl(append-only){"type":"game","round":N,"ticks":N,"score":N,"total_score":N,"win":bool,"opponent":"str"};roundrestarts at 1 each battleSAC_LSTM_Bot/eval_log.jsonlmv .tmp) per successful evalSAC_EVAL_INTERVALchunks>>> [eval] crashedin S1 (old log kept)weights/round_counter.txtbumpRoundCounter()in onRoundEnded, main thread, fires in eval battles too); monotonic, never auto-resetweights/best_score.txt0+sac_best.zippresent means "best-so-far is 0%", seen in #49 smoke runweights/sac_latest.zip.npy: actor+c1+c2+tc1+tc2+alpha). Atomic write (.tmp+move). Written by I/O thread everySACLSTM_SAVE_INTERVAL=500 gradient steps (UTD=1 ⇒ ≈500 transitions ≈ every ~20–60 s of active play)weights/sac_best.zipcrash #kbanners; harness hard-aborts atSAC_MAX_CRASHES=5[sac] checkpoint load failed (…) — random initMinimal watch-set for the periodic check-in agent (order matters, <1 min)
>>> abortedor dead process? → likely CRASHED, stop here.round_counter.txt, compare to previous check-in: Δ=0 over ≥15 min in active window → CRASHED/STALLED.stat -c %Y weights/sac_latest.zip: fresh ⇒ training thread alive; stale while S4 advanced ⇒ STALLED (learning stopped, motion continues — S4 alone cannot detect this because it's bumped by the bot's event thread, independent of the training thread).dfheadroom on the repo volume.Concrete thresholds (for the monitoring loop)
crash #Nbanners with no intervening successful chunk (harness self-aborts at 5; don't wait for it). Process absent + nonzero exit ⇒ intervene immediately.sac_latest.zipmtime older than max(10 min, 3× median observed inter-save gap) while rounds still complete ⇒ training pipeline stalled. Allow one warm-up grace period after launch (buffer needs ≥ batch×seq-len transitions before first step;canSamplegate).best_score.txtflat over ≥5 consecutive evals (≈10 chunks). 'intervene' = 3 consecutive evals each scoring <50% of the established best after a best ≥10% was reached (real regression), or S9 fired (silent random re-init). Best-not-improving alone is NEVER intervene.sac_*.zipsizes are architecture-fixed and overwritten in place — zero growth, no rotation needed. Only S2/S3 (and the launcher's stdout capture, if redirected) grow, linearly in rounds.Four-way discrimination
Observability gaps worth flagging
RunTraining.java's header claims the bot writes "actorLoss, valueLoss…" to the shared log — true for PPO_Bot, false for SAC_LSTM_Bot (inherited-comment trap). All learning-health inference above is proxy-based. Cheapest future fix: log{stepCount,buf.len,alpha}per save inioThreadEntry.trySendresult ignored) — backpressure is undetectable externally.sendTrainingMsgonisEvalMode()— deterministic-eval rounds feed the replay buffer and consume save-interval budget; eval WR is measured on a moving policy.trainerFromFullcarries no optimizer state) — post-crash-loss transient invisible.… | tee run.out) or S1/S8 are lost to the AFK operator.— Research artifact for Lock campaign-v1 config and launch overnight run (#56): define its monitor loop from the watch-set + thresholds above.