fix(SAC_LSTM_Bot): checkpoint save check inside gradient-step loop (#56)

Launch finding during campaign-v1 verification: at production sizes
(hidden 256, ~1s/step, ~13s trainer CPU per ~40s chunk process) the
save check ran only between drain-burst passes, so stepCount never
crossed nextSave before the process died — zero checkpoints persisted
across entire runs (masked at #49/#54 smoke sizes where steps were
sub-millisecond). Check now fires mid-loop; with SAVE_INTERVAL<=5
(within the per-process step budget) every chunk persists its chain.
This commit is contained in:
2026-08-22 00:21:43 +02:00
parent 2f49cb243f
commit 26536713ba
@@ -290,6 +290,15 @@ proc trainPass*(st: var TrainState; drained: int) =
break
discard sacUpdate(st.trainer, seqs)
inc st.stepCount
# Save check INSIDE the step loop (#56 launch finding): at production sizes
# (hidden 256 ⇒ ~1 s/step) a drain burst queues minutes of steps; checking
# only between passes meant the process died mid-loop before stepCount ever
# reached nextSave — zero checkpoints persisted for the whole campaign.
# Mid-loop checks + SAVE_INTERVAL≤20 (#54) keep saves ~20 s apart.
if st.stepCount >= st.nextSave:
st.nextSave += getSaveInterval()
var full = packFull(st.trainer)
discard gSaveChan.trySend(move(full)) # cap-1: drop if I/O thread is busy (Q5)
# Publish latest actor (Q7): in-place write under the lock, bump version.
withLock(gWeightLock):
assert gSharedSnap.hiddenDim == st.trainer.actor.hiddenDim,
@@ -297,10 +306,6 @@ proc trainPass*(st: var TrainState; drained: int) =
var c = 0
packActor(st.trainer.actor, gSharedSnap.data, c)
inc gSharedSnap.version
if st.stepCount >= st.nextSave:
st.nextSave += getSaveInterval()
var full = packFull(st.trainer)
discard gSaveChan.trySend(move(full)) # cap-1: drop if I/O thread is busy (Q5)
# ── Threads ───────────────────────────────────────────────────────────────────