Spec: PPO RL Bot — full implementation #13

Closed
opened 2026-08-16 14:24:45 +02:00 by SirStone · 1 comment
Owner

Problem Statement

The PPO_Bot scaffold exists but does nothing — it loops go() with no RL logic. The bot needs a brain: a neural network that observes the game, chooses actions, and learns from experience between rounds, so it progressively improves against any 1v1 opponent.

Solution

Implement the full PPO reinforcement learning pipeline inside PPO_Bot: a deterministic enemy tracker (radar lock + dead reckoning), a 42-float state vector, a 5D continuous action space, per-tick reward from energy deltas, per-round training via clipped PPO with GAE, and atomic weight persistence across restarts.

User Stories

  1. As a bot operator, I want the bot to observe its own position, speed, energy, gun state, and nearby walls, so that the RL agent has full awareness of its own situation.
  2. As a bot operator, I want the bot to track the enemy's position, speed, direction, and energy via radar, so that the RL agent knows where the opponent is.
  3. As a bot operator, I want the bot to dead-reckon the enemy position on missed radar ticks, so that the state vector stays smooth and the agent doesn't lose track.
  4. As a bot operator, I want the bot to detect enemy firing from energy drops between scans, so that the agent can learn to dodge.
  5. As a bot operator, I want the bot to maintain a 5-tick sliding window of enemy history, so that the agent can perceive enemy movement patterns.
  6. As a bot operator, I want all state values normalized to roughly [-1, 1], so that the neural network trains stably.
  7. As a bot operator, I want the bot to control body speed, body turn, gun turn, fire decision, and fire power as 5 continuous outputs, so that the agent has full control of combat behavior.
  8. As a bot operator, I want body turn rate dynamically scaled by current speed, so that the agent's full output range is always meaningful regardless of speed.
  9. As a bot operator, I want firing suppressed when the gun is hot, so that the agent doesn't waste actions on impossible shots.
  10. As a bot operator, I want the bot to earn per-tick reward from energy deltas (my change minus enemy change), so that the agent gets dense feedback every tick.
  11. As a bot operator, I want the bot to earn a per-round reward from the game's scoring system (divided by 100), so that the agent optimizes for winning, not just surviving.
  12. As a bot operator, I want the actor network (42→64→tanh→64→tanh→5) to output action means, so that the agent can learn complex behaviors.
  13. As a bot operator, I want the critic network (42→64→tanh→64→tanh→1) to estimate state values, so that GAE can compute advantages.
  14. As a bot operator, I want a learnable log-std parameter vector (5 floats) for the stochastic policy, so that exploration adapts during training.
  15. As a bot operator, I want log-std floored at -3.0, so that the policy never fully collapses and can readapt to new opponents.
  16. As a bot operator, I want the training loop to collect the full round trajectory, then run 4 PPO epochs with mini-batches of 64, so that the agent learns efficiently from each round.
  17. As a bot operator, I want GAE (gamma=0.99, lambda=0.95) for advantage estimation, so that the agent balances bias and variance in credit assignment.
  18. As a bot operator, I want the clipped surrogate objective (epsilon=0.2) with entropy bonus (coefficient=0.01), so that training is stable and exploration is maintained.
  19. As a bot operator, I want Adam optimizer (lr=3e-4) with gradient clipping (max norm 0.5) for both networks, so that updates are smooth and bounded.
  20. As a bot operator, I want training to run in a background thread after each round, so that the bot keeps playing without blocking.
  21. As a bot operator, I want an atomic weight swap (pointer swap) when training completes, so that the bot transitions to new weights cleanly mid-battle.
  22. As a bot operator, I want stale training passes dropped if a new round ends before training finishes, so that the agent always trains on the freshest data.
  23. As a bot operator, I want weights saved atomically (write to temp dir, then rename) to weights/latest/ every round, so that crashes never corrupt saved progress.
  24. As a bot operator, I want 3 rotating checkpoints (weights/checkpoint_1/2/3/) saved every 50 rounds, so that I can recover from bad training runs.
  25. As a bot operator, I want the bot to load weights/latest/ on startup, falling back to the newest checkpoint, falling back to random init, so that training resumes seamlessly.
  26. As a bot operator, I want incomplete temp weight directories cleaned up on startup, so that crash artifacts don't accumulate.
  27. As a bot operator, I want the bot to always train (no inference-only mode), so that it continuously adapts.
  28. As a bot operator, I want weights stored as .npy files (one per tensor) via Arraymancer's write_npy/read_npy, so that the format is standard and simple.

Implementation Decisions

  • Single testing seam at the bot API boundary. All RL logic (network, GAE, PPO, action mapping, state construction, reward computation, weight I/O) operates on plain Tensor[float32] and float values — no bot API types cross in. The bot's event handlers (run, onScannedBot, onRoundEnded) are thin glue: build state tensor → call RL → apply returned floats as commands.
  • Radar is deterministic, not RL-controlled. The enemy tracker fully owns setRadarTurnRate. Round start: 45°/tick sweep until first scan. After contact: lock with overshoot (narrow oscillation). Recovery: widen on 2+ missed ticks. This is hardcoded logic, not a network output.
  • EnemyState struct holds x, y, direction, speed, energy, ticksSinceLastScan, hasFired, lastFirePower. Updated on onScannedBot, dead-reckoned on missed ticks.
  • State vector: 42 normalized floats. 14 current-tick (own bot + enemy), 8 derived (acceleration, turn rate, bearing, distance, 4 wall distances), 20 enemy history (5-tick window × 4 values each).
  • Action space: 5D continuous. targetSpeed (tanh×8), turnRate (tanh×speed-dependent max), gunTurnRate (tanh×20), fireDecision (tanh, threshold 0), firePower (sigmoid×2.9+0.1). Fire masked when gun hot.
  • Reward: per-tick energy delta + per-round score/100. No shaping bonuses for aiming, wall avoidance, or ramming — emergent behavior preferred.
  • Separate actor and critic networks, no shared trunk. ~8K params total. Both: 64→tanh→64→tanh hidden layers. State-independent learnable log-std (5 floats, init 0).
  • PPO training on round end. Single-episode on-policy buffer (Seq of tuples). 4 epochs, mini-batch 64, shuffled. Clipped surrogate ε=0.2, entropy bonus 0.01. Adam lr=3e-4 (both nets), gradient clip max norm 0.5.
  • Background training thread. Atomic weight pointer swap on completion. Stale pass dropped if next round ends first. Terminal value=0 on natural round end; bootstrap critic on abnormal interruption.
  • Weight persistence via .npy. Atomic write (temp dir + rename) to latest/ every round. 3 rotating checkpoints every 50 rounds. Startup: load latest → checkpoint → random init. Clean stale temp dirs on boot.
  • Pure Nim with Arraymancer. No Python, no hybrid pipeline. OpenBLAS linked via config.nims + shell.nix.

Testing Decisions

  • Seam: All tests hit the RL math side of the bot API boundary. No bot process, no server, no WebSocket needed.
  • What good tests look like: Feed known tensors in, assert expected outputs. Test external behavior (given this state, does the network produce actions in valid ranges? given this trajectory, does GAE produce expected advantages?) — not internal implementation details.
  • Modules to test via assert-based self-checks:
    • Network forward pass: known input → output shape and activation ranges correct
    • Action mapping: tanh/sigmoid scaling produces values in valid ranges, gun heat masking works
    • State vector construction: synthetic bot values → correct 42-float normalized vector
    • GAE computation: hand-calculated trajectory → expected advantages
    • Reward computation: known energy deltas → expected reward
    • Weight save/load round-trip: save → load → tensors match
  • No test framework. Assert-based checks in a runnable test file, per project conventions.
  • No prior art for tests in this repo — this will be the first test file.

Out of Scope

  • Melee (multi-enemy) support
  • Radar RL control (radar stays deterministic)
  • Python or hybrid training pipeline
  • Training statistics, visualization, or dashboards
  • Multi-opponent curriculum or opponent pool
  • Performance profiling under tick time constraints
  • Inference-only mode
  • Other bots in this repo
  • CONTEXT.md or ADR creation (no domain glossary exists yet)
  • Non-NixOS deployment / true static linking

Further Notes

  • The two loose ends from the wayfinder map (multi-opponent generalization, tick performance profiling) are explicitly out of scope — they're future concerns once the bot is training and fighting.
  • The ponytail: comment in research notes flags a potential upgrade: small replay buffer seeded with high-reward transitions if sample efficiency becomes the bottleneck after real match testing.
  • Upgrade paths documented across decisions: 128→64 layers, orthogonal init, shared trunk, state-dependent std, RNN/GRU, separate actor/critic learning rates, cosine annealing. None are in scope for v1.
## Problem Statement The PPO_Bot scaffold exists but does nothing — it loops `go()` with no RL logic. The bot needs a brain: a neural network that observes the game, chooses actions, and learns from experience between rounds, so it progressively improves against any 1v1 opponent. ## Solution Implement the full PPO reinforcement learning pipeline inside PPO_Bot: a deterministic enemy tracker (radar lock + dead reckoning), a 42-float state vector, a 5D continuous action space, per-tick reward from energy deltas, per-round training via clipped PPO with GAE, and atomic weight persistence across restarts. ## User Stories 1. As a bot operator, I want the bot to observe its own position, speed, energy, gun state, and nearby walls, so that the RL agent has full awareness of its own situation. 2. As a bot operator, I want the bot to track the enemy's position, speed, direction, and energy via radar, so that the RL agent knows where the opponent is. 3. As a bot operator, I want the bot to dead-reckon the enemy position on missed radar ticks, so that the state vector stays smooth and the agent doesn't lose track. 4. As a bot operator, I want the bot to detect enemy firing from energy drops between scans, so that the agent can learn to dodge. 5. As a bot operator, I want the bot to maintain a 5-tick sliding window of enemy history, so that the agent can perceive enemy movement patterns. 6. As a bot operator, I want all state values normalized to roughly [-1, 1], so that the neural network trains stably. 7. As a bot operator, I want the bot to control body speed, body turn, gun turn, fire decision, and fire power as 5 continuous outputs, so that the agent has full control of combat behavior. 8. As a bot operator, I want body turn rate dynamically scaled by current speed, so that the agent's full output range is always meaningful regardless of speed. 9. As a bot operator, I want firing suppressed when the gun is hot, so that the agent doesn't waste actions on impossible shots. 10. As a bot operator, I want the bot to earn per-tick reward from energy deltas (my change minus enemy change), so that the agent gets dense feedback every tick. 11. As a bot operator, I want the bot to earn a per-round reward from the game's scoring system (divided by 100), so that the agent optimizes for winning, not just surviving. 12. As a bot operator, I want the actor network (42→64→tanh→64→tanh→5) to output action means, so that the agent can learn complex behaviors. 13. As a bot operator, I want the critic network (42→64→tanh→64→tanh→1) to estimate state values, so that GAE can compute advantages. 14. As a bot operator, I want a learnable log-std parameter vector (5 floats) for the stochastic policy, so that exploration adapts during training. 15. As a bot operator, I want log-std floored at -3.0, so that the policy never fully collapses and can readapt to new opponents. 16. As a bot operator, I want the training loop to collect the full round trajectory, then run 4 PPO epochs with mini-batches of 64, so that the agent learns efficiently from each round. 17. As a bot operator, I want GAE (gamma=0.99, lambda=0.95) for advantage estimation, so that the agent balances bias and variance in credit assignment. 18. As a bot operator, I want the clipped surrogate objective (epsilon=0.2) with entropy bonus (coefficient=0.01), so that training is stable and exploration is maintained. 19. As a bot operator, I want Adam optimizer (lr=3e-4) with gradient clipping (max norm 0.5) for both networks, so that updates are smooth and bounded. 20. As a bot operator, I want training to run in a background thread after each round, so that the bot keeps playing without blocking. 21. As a bot operator, I want an atomic weight swap (pointer swap) when training completes, so that the bot transitions to new weights cleanly mid-battle. 22. As a bot operator, I want stale training passes dropped if a new round ends before training finishes, so that the agent always trains on the freshest data. 23. As a bot operator, I want weights saved atomically (write to temp dir, then rename) to `weights/latest/` every round, so that crashes never corrupt saved progress. 24. As a bot operator, I want 3 rotating checkpoints (`weights/checkpoint_1/2/3/`) saved every 50 rounds, so that I can recover from bad training runs. 25. As a bot operator, I want the bot to load `weights/latest/` on startup, falling back to the newest checkpoint, falling back to random init, so that training resumes seamlessly. 26. As a bot operator, I want incomplete temp weight directories cleaned up on startup, so that crash artifacts don't accumulate. 27. As a bot operator, I want the bot to always train (no inference-only mode), so that it continuously adapts. 28. As a bot operator, I want weights stored as `.npy` files (one per tensor) via Arraymancer's `write_npy`/`read_npy`, so that the format is standard and simple. ## Implementation Decisions - **Single testing seam at the bot API boundary.** All RL logic (network, GAE, PPO, action mapping, state construction, reward computation, weight I/O) operates on plain `Tensor[float32]` and `float` values — no bot API types cross in. The bot's event handlers (`run`, `onScannedBot`, `onRoundEnded`) are thin glue: build state tensor → call RL → apply returned floats as commands. - **Radar is deterministic, not RL-controlled.** The enemy tracker fully owns `setRadarTurnRate`. Round start: 45°/tick sweep until first scan. After contact: lock with overshoot (narrow oscillation). Recovery: widen on 2+ missed ticks. This is hardcoded logic, not a network output. - **EnemyState struct** holds x, y, direction, speed, energy, ticksSinceLastScan, hasFired, lastFirePower. Updated on `onScannedBot`, dead-reckoned on missed ticks. - **State vector: 42 normalized floats.** 14 current-tick (own bot + enemy), 8 derived (acceleration, turn rate, bearing, distance, 4 wall distances), 20 enemy history (5-tick window × 4 values each). - **Action space: 5D continuous.** targetSpeed (tanh×8), turnRate (tanh×speed-dependent max), gunTurnRate (tanh×20), fireDecision (tanh, threshold 0), firePower (sigmoid×2.9+0.1). Fire masked when gun hot. - **Reward: per-tick energy delta + per-round score/100.** No shaping bonuses for aiming, wall avoidance, or ramming — emergent behavior preferred. - **Separate actor and critic networks, no shared trunk.** ~8K params total. Both: 64→tanh→64→tanh hidden layers. State-independent learnable log-std (5 floats, init 0). - **PPO training on round end.** Single-episode on-policy buffer (Seq of tuples). 4 epochs, mini-batch 64, shuffled. Clipped surrogate ε=0.2, entropy bonus 0.01. Adam lr=3e-4 (both nets), gradient clip max norm 0.5. - **Background training thread.** Atomic weight pointer swap on completion. Stale pass dropped if next round ends first. Terminal value=0 on natural round end; bootstrap critic on abnormal interruption. - **Weight persistence via `.npy`.** Atomic write (temp dir + rename) to `latest/` every round. 3 rotating checkpoints every 50 rounds. Startup: load latest → checkpoint → random init. Clean stale temp dirs on boot. - **Pure Nim with Arraymancer.** No Python, no hybrid pipeline. OpenBLAS linked via config.nims + shell.nix. ## Testing Decisions - **Seam:** All tests hit the RL math side of the bot API boundary. No bot process, no server, no WebSocket needed. - **What good tests look like:** Feed known tensors in, assert expected outputs. Test external behavior (given this state, does the network produce actions in valid ranges? given this trajectory, does GAE produce expected advantages?) — not internal implementation details. - **Modules to test via assert-based self-checks:** - Network forward pass: known input → output shape and activation ranges correct - Action mapping: tanh/sigmoid scaling produces values in valid ranges, gun heat masking works - State vector construction: synthetic bot values → correct 42-float normalized vector - GAE computation: hand-calculated trajectory → expected advantages - Reward computation: known energy deltas → expected reward - Weight save/load round-trip: save → load → tensors match - **No test framework.** Assert-based checks in a runnable test file, per project conventions. - **No prior art** for tests in this repo — this will be the first test file. ## Out of Scope - Melee (multi-enemy) support - Radar RL control (radar stays deterministic) - Python or hybrid training pipeline - Training statistics, visualization, or dashboards - Multi-opponent curriculum or opponent pool - Performance profiling under tick time constraints - Inference-only mode - Other bots in this repo - CONTEXT.md or ADR creation (no domain glossary exists yet) - Non-NixOS deployment / true static linking ## Further Notes - The two loose ends from the wayfinder map (multi-opponent generalization, tick performance profiling) are explicitly out of scope — they're future concerns once the bot is training and fighting. - The `ponytail:` comment in research notes flags a potential upgrade: small replay buffer seeded with high-reward transitions if sample efficiency becomes the bottleneck after real match testing. - Upgrade paths documented across decisions: 128→64 layers, orthogonal init, shared trunk, state-dependent std, RNN/GRU, separate actor/critic learning rates, cosine annealing. None are in scope for v1.
SirStone added the ready-for-agent label 2026-08-16 14:24:45 +02:00
Author
Owner

All implementation slices (#14–#17) resolved. PPO RL bot is training and fighting in 1v1 matches.

All implementation slices (#14–#17) resolved. PPO RL bot is training and fighting in 1v1 matches.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#13