Overnight Training Campaign — SAC_LSTM_Bot #53

Open
opened 2026-08-21 22:36:27 +02:00 by SirStone · 0 comments
Owner

Destination

SAC_LSTM_Bot runs a supervised-automation training campaign at its own pace, continuing until the evidence supports a verdict. Deliverables unchanged: real checkpoints saved, a win-rate table vs the opponent pool + mirror twin, a markdown notebook narrating the story as it happens, every milestone decision recorded as its own Gitea issue, and a verdict on whether the RL approach scales. The human reads the story as it unfolds and decides continuation from it.

Notes

  • Domain: Nim RL bot (SAC + LSTM) trained via SAC_LSTM_Bot/sac_train.sh (chunked Tank Royale battles, embedded server, tools/training_runner/RunTraining.java). Repo /home/davide/Projects/SirRoboGarage, branch research/goto-controller. Spec chain #37–#49 complete; vendored botapi v1.0.1 (name-based opponent ID via getBotName).
  • Execution override (human directive): campaign execution happens INSIDE this effort — launching, monitoring, and intervening are steps of the route, not a later hand-off.
  • Standing preferences: long-running commands in tmux; subagents in background; IGNORE PPO_Bot/ (paused project) except its CURRICULUM_STATE.md difficulty-ladder data (easy: Fire/Target/Crazy/MyFirst*; medium: Corners/PaintingBot; hard: RamFire/SpinBot/Walls/TrackFire; boss: VelocityBot).
  • Autonomy grant: campaign-level decisions (pool weights, restarts, hyperparams, phase switch) are taken unsupervised; code-level fixes ONLY when they block training continuation, each in its own issue.
  • Decision convention (human mandate): every milestone decision becomes its own Gitea issue at decision time, cross-linked from the notebook.
  • Tracker conventions: claim a ticket by assigning SirStone BEFORE working it; blocking = *Blocked by: #N* line in ticket body (closed blocker = unblocked); frontier = open, unblocked, unclaimed children.

Decisions so far

  • Observability inventory: what can we watch mid-run? — 9 signals; crashed/stalled/slow-learner/healthy discriminators with concrete thresholds; key gaps: no training-loss metrics (proxy-only inference), harness stdout must be teed at launch.
  • Build mirror-twin sparring partner + harness readiness check — SacTwin lives as a sibling dir in the sample-bots archive (generator committed: make_twin.sh); isolation proven byte-exact under live-kill; final-chunk crash bug found+fixed; launch notes: SAVE_INTERVAL ≤20, corpse detection ~60s delayed, SAC_EVAL_OPPONENT = first pool entry.
  • Lock campaign-v1 config and launch overnight run — LIVE in tmux sac_campaign: pool Corners:3/Crazy:2/RamFire:1/Target:1/SacTwin:1, hidden 256, batch 16, UTD 1, save interval 5, 25k rounds, 12h ceiling; critical launch bug fixed (checkpoints never persisted at hidden-256 — save check moved inside step loop, 2653671); notebook at SAC_LSTM_Bot/docs/campaign_notebook.md.
  • Campaign verdict: notebook, results table, scale-or-not call — ops PROVEN (14h, zero crashes, natural completion); learning REAL BUT NARROW (twin 90.7%, Crazy ×2; walls untouched); verdict RETUNE BEFORE SCALING — five code-level levers staged for human sign-off; v2 not launched.

Not yet specified

  • Mid-run interventions: opponent promotions (hard tier: SpinBot/Walls/TrackFire), hyperparameter retunes, rollback-to-best policy, stall diagnosis — each materializes as its own issue when its trigger fires.
  • The exploratory → real-training phase-switch call.
  • Concrete shape of the verdict if learning stalls before strength shows — it forms around whatever evidence the campaign yields.
  • Observability upgrades (training-loss/step metrics from the training thread, channel-drop counters, eval-mode training gate) — candidates to promote if the campaign verdict blames blind spots; code-level, so out of campaign autonomy unless they block training.

Out of scope

  • Learning-architecture or reward redesign beyond blockage-fixing code changes.
  • PPO_Bot revival or head-to-head comparison runs.
  • VelocityBot (boss tier) — a future effort, not this night.
## Destination SAC_LSTM_Bot runs a supervised-automation training campaign at its own pace, continuing until the evidence supports a verdict. Deliverables unchanged: real checkpoints saved, a win-rate table vs the opponent pool + mirror twin, a markdown notebook narrating the story as it happens, every milestone decision recorded as its own Gitea issue, and a verdict on whether the RL approach scales. The human reads the story as it unfolds and decides continuation from it. ## Notes - Domain: Nim RL bot (SAC + LSTM) trained via `SAC_LSTM_Bot/sac_train.sh` (chunked Tank Royale battles, embedded server, `tools/training_runner/RunTraining.java`). Repo `/home/davide/Projects/SirRoboGarage`, branch `research/goto-controller`. Spec chain #37–#49 complete; vendored botapi v1.0.1 (name-based opponent ID via `getBotName`). - **Execution override (human directive):** campaign execution happens INSIDE this effort — launching, monitoring, and intervening are steps of the route, not a later hand-off. - Standing preferences: long-running commands in tmux; subagents in background; IGNORE `PPO_Bot/` (paused project) except its `CURRICULUM_STATE.md` difficulty-ladder data (easy: Fire/Target/Crazy/MyFirst*; medium: Corners/PaintingBot; hard: RamFire/SpinBot/Walls/TrackFire; boss: VelocityBot). - Autonomy grant: campaign-level decisions (pool weights, restarts, hyperparams, phase switch) are taken unsupervised; code-level fixes ONLY when they block training continuation, each in its own issue. - Decision convention (human mandate): every milestone decision becomes its own Gitea issue at decision time, cross-linked from the notebook. - Tracker conventions: claim a ticket by assigning `SirStone` BEFORE working it; blocking = `*Blocked by: #N*` line in ticket body (closed blocker = unblocked); frontier = open, unblocked, unclaimed children. ## Decisions so far <!-- empty — filled as tickets close --> - [Observability inventory: what can we watch mid-run?](https://git.fossellini.top/SirStone/SirRoboGarage/issues/55) — 9 signals; crashed/stalled/slow-learner/healthy discriminators with concrete thresholds; key gaps: no training-loss metrics (proxy-only inference), harness stdout must be teed at launch. - [Build mirror-twin sparring partner + harness readiness check](https://git.fossellini.top/SirStone/SirRoboGarage/issues/54) — SacTwin lives as a sibling dir in the sample-bots archive (generator committed: make_twin.sh); isolation proven byte-exact under live-kill; final-chunk crash bug found+fixed; launch notes: SAVE_INTERVAL ≤20, corpse detection ~60s delayed, SAC_EVAL_OPPONENT = first pool entry. - [Lock campaign-v1 config and launch overnight run](https://git.fossellini.top/SirStone/SirRoboGarage/issues/56) — LIVE in tmux sac_campaign: pool Corners:3/Crazy:2/RamFire:1/Target:1/SacTwin:1, hidden 256, batch 16, UTD 1, save interval 5, 25k rounds, 12h ceiling; critical launch bug fixed (checkpoints never persisted at hidden-256 — save check moved inside step loop, 2653671); notebook at SAC_LSTM_Bot/docs/campaign_notebook.md. - [Campaign verdict: notebook, results table, scale-or-not call](https://git.fossellini.top/SirStone/SirRoboGarage/issues/57) — ops PROVEN (14h, zero crashes, natural completion); learning REAL BUT NARROW (twin 90.7%, Crazy ×2; walls untouched); verdict RETUNE BEFORE SCALING — five code-level levers staged for human sign-off; v2 not launched. ## Not yet specified - Mid-run interventions: opponent promotions (hard tier: SpinBot/Walls/TrackFire), hyperparameter retunes, rollback-to-best policy, stall diagnosis — each materializes as its own issue when its trigger fires. - The exploratory → real-training phase-switch call. - Concrete shape of the verdict if learning stalls before strength shows — it forms around whatever evidence the campaign yields. - Observability upgrades (training-loss/step metrics from the training thread, channel-drop counters, eval-mode training gate) — candidates to promote if the campaign verdict blames blind spots; code-level, so out of campaign autonomy unless they block training. ## Out of scope - Learning-architecture or reward redesign beyond blockage-fixing code changes. - PPO_Bot revival or head-to-head comparison runs. - VelocityBot (boss tier) — a future effort, not this night.
SirStone added the wayfinder:map label 2026-08-21 22:36:27 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#53