Weight persistence + background training thread #17

Closed
opened 2026-08-16 15:02:38 +02:00 by SirStone · 1 comment
Owner

Parent

#13

What to build

Make training survive restarts and run without blocking play. After this slice, the bot is production-ready: it trains asynchronously, persists weights atomically, and resumes learning after any crash or restart.

End-to-end slice:

  • Weight save/load: .npy per tensor via Arraymancer's write_npy/read_npy. One file per weight tensor per directory.
  • Atomic writes: write to temp directory, then moveDir (rename) to final path. Incomplete temp dirs deleted on startup.
  • weights/latest/: saved every round.
  • 3 rotating checkpoints: weights/checkpoint_1/, checkpoint_2/, checkpoint_3/ — saved every 50 rounds, rotating.
  • Startup loading chain: load latest/ → fall back to newest checkpoint → fall back to random init.
  • Stale temp cleanup: on startup, scan for and delete any weights/tmp_* directories from prior crashes.
  • Background training thread: kick off PPO training (from slice 3) in a separate thread on round end. Bot continues acting with current weights during training.
  • Atomic weight swap: pointer swap when training completes. Bot seamlessly transitions to new weights.
  • Stale pass dropping: if a new round ends before the current training pass finishes, drop the stale pass and start fresh with the newest episode.
  • Always-train mode: no inference-only toggle. Bot trains every round.
  • Assert-based test: weight save → load → tensor equality round-trip.

Acceptance criteria

  • Weights saved as .npy files (one per tensor) to weights/latest/ every round
  • Saves are atomic (temp dir + rename)
  • 3 rotating checkpoints saved every 50 rounds
  • On startup: loads latest → checkpoint → random init
  • Stale temp directories cleaned on startup
  • Training runs in background thread, bot keeps playing
  • Atomic weight swap when training completes
  • Stale training pass dropped when new round ends before training finishes
  • Bot restarts and resumes training with saved weights
  • Assert-based weight round-trip test passes

Blocked by

  • #16 (Reward + trajectory + GAE + PPO training)
## Parent #13 ## What to build Make training survive restarts and run without blocking play. After this slice, the bot is production-ready: it trains asynchronously, persists weights atomically, and resumes learning after any crash or restart. End-to-end slice: - **Weight save/load**: `.npy` per tensor via Arraymancer's `write_npy`/`read_npy`. One file per weight tensor per directory. - **Atomic writes**: write to temp directory, then `moveDir` (rename) to final path. Incomplete temp dirs deleted on startup. - **`weights/latest/`**: saved every round. - **3 rotating checkpoints**: `weights/checkpoint_1/`, `checkpoint_2/`, `checkpoint_3/` — saved every 50 rounds, rotating. - **Startup loading chain**: load `latest/` → fall back to newest checkpoint → fall back to random init. - **Stale temp cleanup**: on startup, scan for and delete any `weights/tmp_*` directories from prior crashes. - **Background training thread**: kick off PPO training (from slice 3) in a separate thread on round end. Bot continues acting with current weights during training. - **Atomic weight swap**: pointer swap when training completes. Bot seamlessly transitions to new weights. - **Stale pass dropping**: if a new round ends before the current training pass finishes, drop the stale pass and start fresh with the newest episode. - **Always-train mode**: no inference-only toggle. Bot trains every round. - **Assert-based test**: weight save → load → tensor equality round-trip. ## Acceptance criteria - [ ] Weights saved as .npy files (one per tensor) to `weights/latest/` every round - [ ] Saves are atomic (temp dir + rename) - [ ] 3 rotating checkpoints saved every 50 rounds - [ ] On startup: loads latest → checkpoint → random init - [ ] Stale temp directories cleaned on startup - [ ] Training runs in background thread, bot keeps playing - [ ] Atomic weight swap when training completes - [ ] Stale training pass dropped when new round ends before training finishes - [ ] Bot restarts and resumes training with saved weights - [ ] Assert-based weight round-trip test passes ## Blocked by - #16 (Reward + trajectory + GAE + PPO training)
SirStone added the ready-for-agent label 2026-08-16 15:02:38 +02:00
Author
Owner

Resolved in f27b023. Atomic .npy weight save/load (temp dir + rename), weights/latest/ every round, 3 rotating checkpoints every 50 rounds, startup loading chain (latest→checkpoint→random), stale temp cleanup, background training thread with atomic pointer swap and stale pass dropping. Training progress shown in game UI per a8ee2a8.

Resolved in f27b023. Atomic .npy weight save/load (temp dir + rename), weights/latest/ every round, 3 rotating checkpoints every 50 rounds, startup loading chain (latest→checkpoint→random), stale temp cleanup, background training thread with atomic pointer swap and stale pass dropping. Training progress shown in game UI per a8ee2a8.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#17