SirStone
  • Joined on 2025-05-04
SirStone commented on issue SirStone/SirRoboGarage#1 2026-08-16 16:59:44 +02:00
RL Bot — Wayfinder Map

Map cleanup: closed implementation tickets #13–#17, moved #10 to Decisions so far. All decisions made, destination reached. Remaining fog (multi-opponent generalization, performance profiling) is…

SirStone closed issue SirStone/SirRoboGarage#13 2026-08-16 16:59:07 +02:00
Spec: PPO RL Bot — full implementation
SirStone commented on issue SirStone/SirRoboGarage#13 2026-08-16 16:59:04 +02:00
Spec: PPO RL Bot — full implementation

All implementation slices (#14–#17) resolved. PPO RL bot is training and fighting in 1v1 matches.

SirStone closed issue SirStone/SirRoboGarage#17 2026-08-16 16:58:55 +02:00
Weight persistence + background training thread
SirStone closed issue SirStone/SirRoboGarage#16 2026-08-16 16:58:54 +02:00
Reward + trajectory + GAE + PPO training
SirStone closed issue SirStone/SirRoboGarage#14 2026-08-16 16:58:53 +02:00
Network forward pass → bot moves via neural network
SirStone closed issue SirStone/SirRoboGarage#15 2026-08-16 16:58:53 +02:00
Enemy tracker + full 42-float state vector
SirStone commented on issue SirStone/SirRoboGarage#17 2026-08-16 16:58:49 +02:00
Weight persistence + background training thread

Resolved in f27b023. Atomic .npy weight save/load (temp dir + rename), weights/latest/ every round, 3 rotating checkpoints every 50 rounds, startup loading chain (latest→checkpoint→random),…

SirStone commented on issue SirStone/SirRoboGarage#16 2026-08-16 16:58:47 +02:00
Reward + trajectory + GAE + PPO training

Resolved in f27b023. Per-tick energy delta reward + per-round score/100, trajectory buffer, GAE (γ=0.99, λ=0.95), PPO clipped surrogate (ε=0.2) with entropy bonus (0.01), Adam lr=3e-4, gradient…

SirStone commented on issue SirStone/SirRoboGarage#15 2026-08-16 16:58:43 +02:00
Enemy tracker + full 42-float state vector

Resolved in f27b023, refined in fd22535 and bbc9e51. EnemyState struct with dead reckoning, deterministic radar lock with overshoot recovery, enemy fire detection from energy deltas, 5-tick…

SirStone commented on issue SirStone/SirRoboGarage#14 2026-08-16 16:58:40 +02:00
Network forward pass → bot moves via neural network

Resolved in f27b023. Actor (42→64→64→5) and critic (42→64→64→1) forward pass, stochastic policy with learnable log-std (floored at -3.0), action mapping with speed-dependent turn scaling and…

SirStone opened issue SirStone/SirRoboGarage#17 2026-08-16 15:02:38 +02:00
Weight persistence + background training thread
SirStone opened issue SirStone/SirRoboGarage#16 2026-08-16 15:02:23 +02:00
Reward + trajectory + GAE + PPO training
SirStone opened issue SirStone/SirRoboGarage#15 2026-08-16 15:02:05 +02:00
Enemy tracker + full 42-float state vector
SirStone opened issue SirStone/SirRoboGarage#14 2026-08-16 15:01:49 +02:00
Network forward pass → bot moves via neural network
SirStone opened issue SirStone/SirRoboGarage#13 2026-08-16 14:24:45 +02:00
Spec: PPO RL Bot — full implementation
SirStone closed issue SirStone/SirRoboGarage#11 2026-08-16 13:22:12 +02:00
Bot scaffold and build setup
SirStone commented on issue SirStone/SirRoboGarage#11 2026-08-16 13:22:11 +02:00
Bot scaffold and build setup

Resolution

Scaffold created at PPO_Bot/ with 6 files:

  • PPO_Bot.nim — minimal bot: subclasses Bot, run loops on go(), starts with start(bot, botJsonPath). No RL logic yet. -…
SirStone closed issue SirStone/SirRoboGarage#12 2026-08-16 13:03:06 +02:00
Learning rate and optimizer hyperparameters
SirStone commented on issue SirStone/SirRoboGarage#12 2026-08-16 13:03:02 +02:00
Learning rate and optimizer hyperparameters

Resolution

Optimizer: Adam. Two instances — one for actor, one for critic. Arraymancer has Adam built in (confirmed working in spike).

Learning rate: 3e-4, same for actor and critic.…