Main bot integration #48

Closed
opened 2026-08-20 23:29:01 +02:00 by SirStone · 3 comments
Owner

Parent

#37

What to build

Wire all modules into the working SAC_LSTM_Bot. The skeleton bot (which already connects and has radar lock + colors) gains the full RL pipeline:

  • WebSocket thread (inference): On each tick, build state vector, run actor forward pass, map actions to BotIntent, send intent. Store transition for training.
  • Background thread (training): Consume transitions, add to replay buffer, run SAC updates when buffer has enough data.
  • LSTM state management: Hidden state (h, c) persists across rounds within a battle. Zeros at battle start.
  • Episode boundaries: Battle end = done=true in replay buffer. Round end = NOT an episode boundary.
  • Configuration: All hyperparameters via env vars (hidden size, buffer capacity, burn-in, train window, target entropy, tau, learning rates, eval mode).

The bot should fight and learn — observable through decreasing critic loss and improving match outcomes over training rounds.

Acceptance criteria

  • Bot connects, fights, and completes battles without crashing
  • Inference runs in under 2ms per tick (no skipped turns)
  • LSTM hidden state resets at battle start, persists across rounds
  • Transitions flow to replay buffer with correct battle-boundary episode flags
  • Background training thread runs SAC updates without blocking inference
  • Weights are saved periodically via the weight persistence module
  • All hyperparameters respond to env var configuration
  • Deterministic eval mode works (no sampling noise)

Blocked by

  • #40 (Skeleton bot with radar lock and colors)
  • #41 (LSTM network module)
  • #42 (State vector module)
  • #43 (Action mapping module)
  • #44 (Reward module)
  • #45 (Replay buffer module)
  • #46 (Weight persistence module)
  • #47 (SAC training module)
## Parent #37 ## What to build Wire all modules into the working SAC_LSTM_Bot. The skeleton bot (which already connects and has radar lock + colors) gains the full RL pipeline: - **WebSocket thread (inference):** On each tick, build state vector, run actor forward pass, map actions to BotIntent, send intent. Store transition for training. - **Background thread (training):** Consume transitions, add to replay buffer, run SAC updates when buffer has enough data. - **LSTM state management:** Hidden state (h, c) persists across rounds within a battle. Zeros at battle start. - **Episode boundaries:** Battle end = done=true in replay buffer. Round end = NOT an episode boundary. - **Configuration:** All hyperparameters via env vars (hidden size, buffer capacity, burn-in, train window, target entropy, tau, learning rates, eval mode). The bot should fight and learn — observable through decreasing critic loss and improving match outcomes over training rounds. ## Acceptance criteria - [x] Bot connects, fights, and completes battles without crashing - [x] Inference runs in under 2ms per tick (no skipped turns) - [x] LSTM hidden state resets at battle start, persists across rounds - [x] Transitions flow to replay buffer with correct battle-boundary episode flags - [x] Background training thread runs SAC updates without blocking inference - [x] Weights are saved periodically via the weight persistence module - [x] All hyperparameters respond to env var configuration - [x] Deterministic eval mode works (no sampling noise) ## Blocked by - #40 (Skeleton bot with radar lock and colors) - #41 (LSTM network module) - #42 (State vector module) - #43 (Action mapping module) - #44 (Reward module) - #45 (Replay buffer module) - #46 (Weight persistence module) - #47 (SAC training module)
SirStone added the ready-for-agent label 2026-08-20 23:29:01 +02:00
Author
Owner

Settled architecture decisions (Q1–Q14) — recorded from the #48 wayfinding session (2026-08-21). This thread/channel architecture is the agreed spec for implementation.

# Decision Answer
Q1 Training architecture Dedicated training thread, lives for entire process lifetime. Bot thread does inference only (<2ms). No training on bot thread. Channel-based data flow, no tensors cross threads.
Q2 Gradient steps UTD ratio (update-to-data), configurable via SACLSTM_UTD_RATIO. If round produces 200 ticks, do 200 gradient steps.
Q3 Replay buffer format Tensor[float32] — lives entirely on training thread, no cross-thread issue.
Q4 LSTM hidden state Tensor[float32] on bot thread. Zeros at battle start, persists across rounds. Never crosses to training thread.
Q5 Weight saving Dedicated I/O thread, cap-1 Channel with trySend. Bot/training thread never blocks on disk I/O.
Q6 Tick budget No tick budget needed — training doesn't run on bot thread at all. Bot thread only does inference.
Q7 Weight sync (training → inference) Shared buffer + Lock, "always latest" semantics. Training thread overwrites snapshot under lock. Bot thread copies under lock each tick. Lock held for nanoseconds (memcpy).
Q8 Overshoot protection N/A — training doesn't touch tick.
Q9 Warm-up canSample() already encodes minimum — no separate warm-up.
Q10 Transition channel Cap-256, trySend, drain-then-train loop on training thread.
Q11 Weight sync direction Training thread pushes to shared buffer. Bot thread pulls (copies under lock) each tick.
Q12 Training thread lifecycle Permanent — spawned once at process start, survives across battles.
Q12a Replay buffer across battles Keep if same opponent, clear if opponent changed.
Q12b Battle boundary signaling Via variant message type in transition channel.
Q13 Transition fields state: array[35, float32], action: array[4, float32], reward: float32, nextState: array[35, float32], done: bool. No logProb (SAC recomputes). No LSTM state (burn-in reconstructs).
Q14 Opponent identification Use scannedBotId: int from first scan. Protocol sends no names. If ID changes across battles → buffer clears (safe default).

TrainingMsg variant: Transition(state, action, reward, nextState, done) | NewBattle(enemyId: int) | Shutdown

Channel/shared-state map: transitionChan: Channel[TrainingMsg] (cap 256, bot→training) · weightLock: Lock + latest actor snapshot (training writes, bot reads) · saveChan (cap 1, training→I/O) · existing gTickChan/gIntentChan/gEventChan untouched.

ORC tensor safety: No Arraymancer tensor crosses a thread boundary (PPO_Bot SIGSEGV lesson). Each thread builds its own tensors from plain arrays/seqs/floats.

**Settled architecture decisions (Q1–Q14)** — recorded from the #48 wayfinding session (2026-08-21). This thread/channel architecture is the agreed spec for implementation. | # | Decision | Answer | |---|----------|--------| | Q1 | Training architecture | Dedicated training thread, lives for entire process lifetime. Bot thread does inference only (<2ms). No training on bot thread. Channel-based data flow, no tensors cross threads. | | Q2 | Gradient steps | UTD ratio (update-to-data), configurable via `SACLSTM_UTD_RATIO`. If round produces 200 ticks, do 200 gradient steps. | | Q3 | Replay buffer format | `Tensor[float32]` — lives entirely on training thread, no cross-thread issue. | | Q4 | LSTM hidden state | `Tensor[float32]` on bot thread. Zeros at battle start, persists across rounds. Never crosses to training thread. | | Q5 | Weight saving | Dedicated I/O thread, cap-1 `Channel` with `trySend`. Bot/training thread never blocks on disk I/O. | | Q6 | Tick budget | No tick budget needed — training doesn't run on bot thread at all. Bot thread only does inference. | | Q7 | Weight sync (training → inference) | Shared buffer + Lock, "always latest" semantics. Training thread overwrites snapshot under lock. Bot thread copies under lock each tick. Lock held for nanoseconds (memcpy). | | Q8 | Overshoot protection | N/A — training doesn't touch tick. | | Q9 | Warm-up | `canSample()` already encodes minimum — no separate warm-up. | | Q10 | Transition channel | Cap-256, `trySend`, drain-then-train loop on training thread. | | Q11 | Weight sync direction | Training thread pushes to shared buffer. Bot thread pulls (copies under lock) each tick. | | Q12 | Training thread lifecycle | Permanent — spawned once at process start, survives across battles. | | Q12a | Replay buffer across battles | Keep if same opponent, clear if opponent changed. | | Q12b | Battle boundary signaling | Via variant message type in transition channel. | | Q13 | Transition fields | `state: array[35, float32]`, `action: array[4, float32]`, `reward: float32`, `nextState: array[35, float32]`, `done: bool`. No logProb (SAC recomputes). No LSTM state (burn-in reconstructs). | | Q14 | Opponent identification | Use `scannedBotId: int` from first scan. Protocol sends no names. If ID changes across battles → buffer clears (safe default). | **TrainingMsg variant:** `Transition(state, action, reward, nextState, done) | NewBattle(enemyId: int) | Shutdown` **Channel/shared-state map:** `transitionChan: Channel[TrainingMsg]` (cap 256, bot→training) · `weightLock: Lock` + latest actor snapshot (training writes, bot reads) · `saveChan` (cap 1, training→I/O) · existing `gTickChan`/`gIntentChan`/`gEventChan` untouched. **ORC tensor safety:** No Arraymancer tensor crosses a thread boundary (PPO_Bot SIGSEGV lesson). Each thread builds its own tensors from plain arrays/seqs/floats.
Author
Owner

Claimed for implementation (self-assigned as SirStone — the MCP issue-edit surface here exposes no assignee field, so this comment is the assignment record). Working on branch research/goto-controller.

Claimed for implementation (self-assigned as `SirStone` — the MCP issue-edit surface here exposes no assignee field, so this comment is the assignment record). Working on branch `research/goto-controller`.
Author
Owner

Resolution — main bot integration (commit 32b71d9)

All modules (#38–#47) wired into a working bot per the Q1–Q14 decisions.

What was built

  • src/SAC_LSTM_Bot/integration.nim (new): thread plumbing — TrainingMsg object-variant (Transition(state/action/reward/nextState/done) | NewBattle(enemyId) | Shutdown) carried as plain fixed arrays; flat weight-snapshot pack/unpack (actor + 4 critics + alpha, layout asserts against net shapes); shared actor snapshot (Lock + version, allocated once, written in place, copied out element-wise — zero cross-thread refcount traffic); testable TrainState/handleTrainingMsg/trainPass; training + I/O thread entries; lifecycle initIntegration/shutdownIntegration.
  • src/SAC_LSTM_Bot.nim (rewritten): bot thread inference loop — battle-change detection → weight pull → bullet tracking → 35-dim state → LSTM actor forward → action mapping → intents → go(); transition finalization each tick; event-driven rewards through the Welford normalizer (onBulletHit, onHitByBullet, onHitWall, onBulletHitWall); radar lock preserved; LSTM h/c persist across rounds as plain arrays on the bot object, zeroed at battle start.
  • config.nims: added --path:src (binary now imports submodules; tests already had their own path).
  • replay_buffer.nim: folds in the pre-existing uncommitted clear() used by the NewBattle/opponent-changed rule.

Thread/channel inventory (5 threads total)

  1. Main (bot API receive loop) — unchanged
  2. Bot (inference) — existing API thread, now runs the RL loop
  3. Sender (bot API) — unchanged
  4. Training (new, permanent): drain-then-train over Channel[TrainingMsg] cap-256 (trySend, drops on overflow), UTD ratio via SACLSTM_UTD_RATIO, owns SACTrainer + ReplayBuffer, publishes actor to the locked snapshot, NewBattle clears buffer only when enemyId changed
  5. I/O (new): cap-1 Channel[FullSnap], atomic zip saves via weights.nim every SACLSTM_SAVE_INTERVAL steps + one final save on Shutdown

No Arraymancer tensor crosses a thread boundary; channels/snapshot carry plain arrays/seqs only.

Tests: nimble test — 8/8 suites green (7 existing + new test_integration: channel round-trip incl. closed-channel semantics, NewBattle clear-only-on-opponent-change, trainPass no-op below canSample, flat snapshot pack/unpack round-trip). Additional end-to-end smoke (no server): threads spawn → real sacUpdate on hidden=32 → weight publish v2 → zip saved → clean join. Binary builds clean and starts/exits cleanly without a server.

New env vars: SACLSTM_UTD_RATIO (1), SACLSTM_BATCH_SIZE (16), SACLSTM_SAVE_INTERVAL (500), SACLSTM_WEIGHTS_PATH (/weights/sac_latest.zip).

Deferred: opponent identity is the numeric scannedBotId for now — name-based identification deferred to #49 (protocol exposes no names). Adam momentum is not carried in the periodic flat save (optimizer restarts fresh per process); #49's checkpoint management can switch to saveCheckpoint/loadCheckpoint end-to-end if resume quality matters.

## Resolution — main bot integration (commit `32b71d9`) All modules (#38–#47) wired into a working bot per the Q1–Q14 decisions. **What was built** - `src/SAC_LSTM_Bot/integration.nim` (new): thread plumbing — `TrainingMsg` object-variant (`Transition(state/action/reward/nextState/done)` | `NewBattle(enemyId)` | `Shutdown`) carried as plain fixed arrays; flat weight-snapshot pack/unpack (actor + 4 critics + alpha, layout asserts against net shapes); shared actor snapshot (`Lock` + version, allocated once, written in place, copied out element-wise — zero cross-thread refcount traffic); testable `TrainState`/`handleTrainingMsg`/`trainPass`; training + I/O thread entries; lifecycle `initIntegration`/`shutdownIntegration`. - `src/SAC_LSTM_Bot.nim` (rewritten): bot thread inference loop — battle-change detection → weight pull → bullet tracking → 35-dim state → LSTM actor forward → action mapping → intents → `go()`; transition finalization each tick; event-driven rewards through the Welford normalizer (`onBulletHit`, `onHitByBullet`, `onHitWall`, `onBulletHitWall`); radar lock preserved; LSTM h/c persist across rounds as plain arrays on the bot object, zeroed at battle start. - `config.nims`: added `--path:src` (binary now imports submodules; tests already had their own path). - `replay_buffer.nim`: folds in the pre-existing uncommitted `clear()` used by the NewBattle/opponent-changed rule. **Thread/channel inventory** (5 threads total) 1. Main (bot API receive loop) — unchanged 2. Bot (inference) — existing API thread, now runs the RL loop 3. Sender (bot API) — unchanged 4. Training (new, permanent): drain-then-train over `Channel[TrainingMsg]` cap-256 (`trySend`, drops on overflow), UTD ratio via `SACLSTM_UTD_RATIO`, owns SACTrainer + ReplayBuffer, publishes actor to the locked snapshot, NewBattle clears buffer only when `enemyId` changed 5. I/O (new): cap-1 `Channel[FullSnap]`, atomic zip saves via weights.nim every `SACLSTM_SAVE_INTERVAL` steps + one final save on Shutdown No Arraymancer tensor crosses a thread boundary; channels/snapshot carry plain arrays/seqs only. **Tests**: `nimble test` — 8/8 suites green (7 existing + new `test_integration`: channel round-trip incl. closed-channel semantics, NewBattle clear-only-on-opponent-change, trainPass no-op below canSample, flat snapshot pack/unpack round-trip). Additional end-to-end smoke (no server): threads spawn → real sacUpdate on hidden=32 → weight publish v2 → zip saved → clean join. Binary builds clean and starts/exits cleanly without a server. **New env vars**: `SACLSTM_UTD_RATIO` (1), `SACLSTM_BATCH_SIZE` (16), `SACLSTM_SAVE_INTERVAL` (500), `SACLSTM_WEIGHTS_PATH` (<src>/weights/sac_latest.zip). **Deferred**: opponent identity is the numeric `scannedBotId` for now — name-based identification deferred to #49 (protocol exposes no names). Adam momentum is not carried in the periodic flat save (optimizer restarts fresh per process); #49's checkpoint management can switch to saveCheckpoint/loadCheckpoint end-to-end if resume quality matters.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#48