Training loop design #8
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Question
How does the training loop work between rounds?
Proposed flow:
Sub-questions:
Blocked by #2 (Arraymancer viability spike) — need confirmed ML stack before implementing training loop.
Blocked by #3 (RL algorithm choice) — algorithm determines on/off-policy, replay buffer strategy, and gradient update schedule.
Resolution
Episode buffer: Seq of
(state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32)tuples. Append during round, clear after training. Single episode, on-policy — no accumulation across rounds.Training pass: 4 PPO epochs, mini-batches of 64, shuffled each epoch. Full episode can be thousands of ticks (no fixed round length — rounds end on death, not a turn limit; inactivity damage kicks in at 450 ticks with no bullet hits).
GAE computation: gamma=0.99, lambda=0.95. Terminal value=0 on natural round end (death/win). Bootstrap with critic(last_state) on abnormal interruption (crash, server stop, user stops game).
Training timing: Background thread, kicked off on round end. Rounds are back-to-back with no pause. Bot continues acting with current weights during training. Atomic weight swap (pointer swap) when training completes. If a new round ends before training finishes, drop the stale training pass and start fresh with the newest episode (on-policy PPO requires current-policy data).
Weight persistence:
latest/directory — atomic write (write to temp dir, then rename) every roundcheckpoint_1/,checkpoint_2/,checkpoint_3/) every 50 rounds — user-driven rollback if policy collapses. No automatic "best" heuristic (score is opponent-dependent, unreliable).latest/. If corrupt/missing, fall back to newest checkpoint. If nothing exists, initialize randomly. Always trains — no inference-only mode.Weight format:
.npyper tensor via Arraymancer built-inwrite_npy/read_npy. One file per weight tensor in each directory.Build note (fog): Static-link OpenBLAS for portable single-binary deployment — build concern, not training loop.