Spec: PPO RL Bot — full implementation #13
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem Statement
The PPO_Bot scaffold exists but does nothing — it loops
go()with no RL logic. The bot needs a brain: a neural network that observes the game, chooses actions, and learns from experience between rounds, so it progressively improves against any 1v1 opponent.Solution
Implement the full PPO reinforcement learning pipeline inside PPO_Bot: a deterministic enemy tracker (radar lock + dead reckoning), a 42-float state vector, a 5D continuous action space, per-tick reward from energy deltas, per-round training via clipped PPO with GAE, and atomic weight persistence across restarts.
User Stories
weights/latest/every round, so that crashes never corrupt saved progress.weights/checkpoint_1/2/3/) saved every 50 rounds, so that I can recover from bad training runs.weights/latest/on startup, falling back to the newest checkpoint, falling back to random init, so that training resumes seamlessly..npyfiles (one per tensor) via Arraymancer'swrite_npy/read_npy, so that the format is standard and simple.Implementation Decisions
Tensor[float32]andfloatvalues — no bot API types cross in. The bot's event handlers (run,onScannedBot,onRoundEnded) are thin glue: build state tensor → call RL → apply returned floats as commands.setRadarTurnRate. Round start: 45°/tick sweep until first scan. After contact: lock with overshoot (narrow oscillation). Recovery: widen on 2+ missed ticks. This is hardcoded logic, not a network output.onScannedBot, dead-reckoned on missed ticks..npy. Atomic write (temp dir + rename) tolatest/every round. 3 rotating checkpoints every 50 rounds. Startup: load latest → checkpoint → random init. Clean stale temp dirs on boot.Testing Decisions
Out of Scope
Further Notes
ponytail:comment in research notes flags a potential upgrade: small replay buffer seeded with high-reward transitions if sample efficiency becomes the bottleneck after real match testing.All implementation slices (#14–#17) resolved. PPO RL bot is training and fighting in 1v1 matches.