Wayfinding: Recover Hope via Rigorous Check on Open Issues & Training Harness #67

Open
opened 2026-08-25 09:44:50 +02:00 by SirStone · 0 comments
Owner

Destination

A state where the SAC_LSTM_Bot training harness runs stably without core dumps or corpse stalls for at least 1 hour, with all technically solvable issues (persistence, stalls, core dumps) resolved, enabling evaluation of learning progress for the scalability verdict.

Notes

Domain = SAC_LSTM_Bot training campaign (Nim/Python, Tank Royale 1v1). Skills every session should consult: diagnosing-bugs (for core dumps and corpse stalls), code-review (for fixes), ponytail (minimal fixes), handoff (if state changes). Standing preferences: root-cause fixes only (no clamp-as-patch), proactive problem detection, short plain sentences for non-native English user, workspace locked to /home/davide/Projects/SirRoboGarage/sac_bot/SAC_LSTM_Bot/, long-running via systemd unit sac-training.service, port 7654 reserved.

Decisions so far

Not yet specified

  • Is the Nim thread actually hanging, or is the Java side not sending data? (Corpse stall cross-process I/O)
  • Does the core dump involve BLAS/OpenBLAS, or something else in the Nim/Atlas stack?
  • Is the actor loss plateau purely from empty-buffer training, or is there a deeper architecture issue?
  • Should we add a watchdog that restarts the training process on corpse stall, or fix the root cause?
  • Is the 50K buffer serialization size appropriate, or should it be configurable/tuned?
  • Do we need a formal "campaign health" dashboard with alerting?
  • Should we parallelize eval (multiple JVMs) instead of just reducing frequency?

Out of scope

## Destination A state where the SAC_LSTM_Bot training harness runs stably without core dumps or corpse stalls for at least 1 hour, with all technically solvable issues (persistence, stalls, core dumps) resolved, enabling evaluation of learning progress for the scalability verdict. ## Notes Domain = SAC_LSTM_Bot training campaign (Nim/Python, Tank Royale 1v1). Skills every session should consult: `diagnosing-bugs` (for core dumps and corpse stalls), `code-review` (for fixes), `ponytail` (minimal fixes), `handoff` (if state changes). Standing preferences: root-cause fixes only (no clamp-as-patch), proactive problem detection, short plain sentences for non-native English user, workspace locked to `/home/davide/Projects/SirRoboGarage/sac_bot/SAC_LSTM_Bot/`, long-running via systemd unit `sac-training.service`, port 7654 reserved. ## Decisions so far <!-- index of closed tickets --> ## Not yet specified - Is the Nim thread actually hanging, or is the Java side not sending data? (Corpse stall cross-process I/O) - Does the core dump involve BLAS/OpenBLAS, or something else in the Nim/Atlas stack? - Is the actor loss plateau purely from empty-buffer training, or is there a deeper architecture issue? - Should we add a watchdog that restarts the training process on corpse stall, or fix the root cause? - Is the 50K buffer serialization size appropriate, or should it be configurable/tuned? - Do we need a formal "campaign health" dashboard with alerting? - Should we parallelize eval (multiple JVMs) instead of just reducing frequency? ## Out of scope <!-- none yet -->
SirStone added the wayfinder:map label 2026-08-25 09:44:50 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#67