Map cleanup: closed implementation tickets #13–#17, moved #10 to Decisions so far. All decisions made, destination reached. Remaining fog (multi-opponent generalization, performance profiling) is…
All implementation slices (#14–#17) resolved. PPO RL bot is training and fighting in 1v1 matches.
Resolved in f27b023. Atomic .npy weight save/load (temp dir + rename), weights/latest/ every round, 3 rotating checkpoints every 50 rounds, startup loading chain (latest→checkpoint→random),…
Resolved in f27b023. Per-tick energy delta reward + per-round score/100, trajectory buffer, GAE (γ=0.99, λ=0.95), PPO clipped surrogate (ε=0.2) with entropy bonus (0.01), Adam lr=3e-4, gradient…
Resolved in f27b023, refined in fd22535 and bbc9e51. EnemyState struct with dead reckoning, deterministic radar lock with overshoot recovery, enemy fire detection from energy deltas, 5-tick…
Resolved in f27b023. Actor (42→64→64→5) and critic (42→64→64→1) forward pass, stochastic policy with learnable log-std (floored at -3.0), action mapping with speed-dependent turn scaling and…
Resolution
Scaffold created at PPO_Bot/ with 6 files:
PPO_Bot.nim— minimal bot: subclasses Bot,runloops ongo(), starts withstart(bot, botJsonPath). No RL logic yet. -…
Resolution
Optimizer: Adam. Two instances — one for actor, one for critic. Arraymancer has Adam built in (confirmed working in spike).
Learning rate: 3e-4, same for actor and critic.…