Training loop design #8

Closed
opened 2026-08-15 22:05:40 +02:00 by SirStone · 2 comments
Owner

Question

How does the training loop work between rounds?

Proposed flow:

  1. During round: agent acts, stores (state, action, reward, next_state) tuples in an episode buffer
  2. Round ends (onDeath or win): compute returns/advantages
  3. Between rounds: run N gradient updates on the collected data
  4. Clear episode buffer, start next round with updated weights

Sub-questions:

  • Replay buffer: keep only last episode (on-policy) or accumulate across rounds (off-policy)? Depends on algorithm choice (ticket 2).
  • How many gradient steps between rounds? Fixed N or until convergence?
  • Save weights after every round, every N rounds, or only on shutdown?
  • Memory management: how large can the buffer get in a long match?
## Question How does the training loop work between rounds? Proposed flow: 1. During round: agent acts, stores (state, action, reward, next_state) tuples in an episode buffer 2. Round ends (onDeath or win): compute returns/advantages 3. Between rounds: run N gradient updates on the collected data 4. Clear episode buffer, start next round with updated weights Sub-questions: - Replay buffer: keep only last episode (on-policy) or accumulate across rounds (off-policy)? Depends on algorithm choice (ticket 2). - How many gradient steps between rounds? Fixed N or until convergence? - Save weights after every round, every N rounds, or only on shutdown? - Memory management: how large can the buffer get in a long match?
SirStone added the wayfinder:grilling label 2026-08-15 22:05:40 +02:00
Author
Owner

Blocked by #2 (Arraymancer viability spike) — need confirmed ML stack before implementing training loop.
Blocked by #3 (RL algorithm choice) — algorithm determines on/off-policy, replay buffer strategy, and gradient update schedule.

Blocked by #2 (Arraymancer viability spike) — need confirmed ML stack before implementing training loop. Blocked by #3 (RL algorithm choice) — algorithm determines on/off-policy, replay buffer strategy, and gradient update schedule.
Author
Owner

Resolution

Episode buffer: Seq of (state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32) tuples. Append during round, clear after training. Single episode, on-policy — no accumulation across rounds.

Training pass: 4 PPO epochs, mini-batches of 64, shuffled each epoch. Full episode can be thousands of ticks (no fixed round length — rounds end on death, not a turn limit; inactivity damage kicks in at 450 ticks with no bullet hits).

GAE computation: gamma=0.99, lambda=0.95. Terminal value=0 on natural round end (death/win). Bootstrap with critic(last_state) on abnormal interruption (crash, server stop, user stops game).

Training timing: Background thread, kicked off on round end. Rounds are back-to-back with no pause. Bot continues acting with current weights during training. Atomic weight swap (pointer swap) when training completes. If a new round ends before training finishes, drop the stale training pass and start fresh with the newest episode (on-policy PPO requires current-policy data).

Weight persistence:

  • latest/ directory — atomic write (write to temp dir, then rename) every round
  • 3 rotating checkpoints (checkpoint_1/, checkpoint_2/, checkpoint_3/) every 50 rounds — user-driven rollback if policy collapses. No automatic "best" heuristic (score is opponent-dependent, unreliable).
  • On startup: load latest/. If corrupt/missing, fall back to newest checkpoint. If nothing exists, initialize randomly. Always trains — no inference-only mode.
  • Crash safety: atomic rename protects against mid-write corruption. Incomplete temp dirs are deleted on startup.

Weight format: .npy per tensor via Arraymancer built-in write_npy/read_npy. One file per weight tensor in each directory.

Build note (fog): Static-link OpenBLAS for portable single-binary deployment — build concern, not training loop.

## Resolution **Episode buffer:** Seq of `(state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32)` tuples. Append during round, clear after training. Single episode, on-policy — no accumulation across rounds. **Training pass:** 4 PPO epochs, mini-batches of 64, shuffled each epoch. Full episode can be thousands of ticks (no fixed round length — rounds end on death, not a turn limit; inactivity damage kicks in at 450 ticks with no bullet hits). **GAE computation:** gamma=0.99, lambda=0.95. Terminal value=0 on natural round end (death/win). Bootstrap with critic(last_state) on abnormal interruption (crash, server stop, user stops game). **Training timing:** Background thread, kicked off on round end. Rounds are back-to-back with no pause. Bot continues acting with current weights during training. Atomic weight swap (pointer swap) when training completes. If a new round ends before training finishes, drop the stale training pass and start fresh with the newest episode (on-policy PPO requires current-policy data). **Weight persistence:** - `latest/` directory — atomic write (write to temp dir, then rename) every round - 3 rotating checkpoints (`checkpoint_1/`, `checkpoint_2/`, `checkpoint_3/`) every 50 rounds — user-driven rollback if policy collapses. No automatic "best" heuristic (score is opponent-dependent, unreliable). - On startup: load `latest/`. If corrupt/missing, fall back to newest checkpoint. If nothing exists, initialize randomly. Always trains — no inference-only mode. - Crash safety: atomic rename protects against mid-write corruption. Incomplete temp dirs are deleted on startup. **Weight format:** `.npy` per tensor via Arraymancer built-in `write_npy`/`read_npy`. One file per weight tensor in each directory. **Build note (fog):** Static-link OpenBLAS for portable single-binary deployment — build concern, not training loop.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#8