Block a user
Should eval interval remain at 5, or define a metric to re-evaluate (e.g., "if eval variance > X, drop back to 3")?
Is chunk size=10 final, or should we re-attempt 20 with the persistence fix (using a measurable criterion like 'no corpse stalls for 2h at chunk=10')?
What explicit criteria should define 'verified' persistence fix (e.g.,
buf_size > 0 in first 3 metrics after restart AND training_state.bin rewritten within 2 min)?
Should we configure
ulimit -c unlimited + coredumpctl capture in the service unit, and treat core dump recurrence as a blocker (stop campaign) or data point (continue with monitoring)?
Where and what granularity should we instrument the Nim training thread with chunk-start/end logs + heartbeat to distinguish hang vs crash vs slow for corpse stall diagnosis?
Wayfinding: Recover Hope via Rigorous Check on Open Issues & Training Harness
Campaign-v2 pacing & durability decisions
Session log — Aug 24 evening (speedup attempt + instability)
Changes deployed
- UTD ratio 1→3, batch size 16→32 (
integration.nimenv defaults) — grad_steps confirmed 3 in…
Campaign-v2 pacing & durability decisions
Prototype: GA evolution loop on toy prediction problem
Prototype: GA evolution loop on toy prediction problem
Resolution
Built: prototypes/ga_gun_spike/ga_spike.nim — 111-line pure-Nim GA that evolves a 10-4-1 tanh ANN (49 weights) to predict sin(x) from a sliding window.
**Pipeline validated:*…