RL Bot — Wayfinder Map #1

Closed
opened 2026-08-15 22:05:04 +02:00 by SirStone · 2 comments
Owner

Destination

An RL-trained Tank Royale bot in pure Nim that controls body movement, turret, and firing via continuous actor-critic, trains between rounds, persists weights, and progressively improves against any 1v1 opponent.

Notes

  • Domain: Tank Royale (Robocode successor). Nim bot API in sibling repo /home/davide/Projects/tank-royale branch nim.
  • Skills to consult: /grilling, /domain-modeling, /prototype
  • Arraymancer is the primary ML library; hand-rolled MLP is the fallback if Arraymancer proves unusable.
  • Radar is deterministic (not RL-controlled) — a simple lock/sweep tracker.
  • Bot folder must be self-contained under SirRoboGarage repo root.
  • Physics: tick-based, MAX_SPEED=8, MAX_TURN_RATE=10 (speed-dependent), MAX_GUN_TURN_RATE=20, MAX_RADAR_TURN_RATE=45, BOT_RADIUS=18, RADAR_RADIUS=1200.
  • Action space: targetSpeed ∈ [-8,8], turnRate ∈ [-10,10], gunTurnRate ∈ [-20,20], firePower ∈ {0} ∪ [0.1,3.0]. 4 continuous dimensions.
  • State: own bot state + tracked enemy state (from deterministic radar). Exact vector TBD.

Decisions so far

  • Arraymancer viability spike — GO. Arraymancer 0.7.33 works on Nim 2.2.4 + NixOS, XOR MLP trains correctly.
  • RL algorithm choice — PPO. On-policy fits short episodes, simplest from-scratch implementation (~200-300 lines), clip prevents catastrophic collapse.
  • Enemy tracker design — Decoupled deterministic radar with lock+overshoot, dead reckoning on missed ticks, EnemyState struct with inferred firing detection.
  • State vector design — 42 normalized floats: 14 current tick (own+enemy), 8 derived (acceleration, turn rate, bearings, wall distances), 20 enemy window (5 ticks × 4 features). No own history, RNN upgrade path if needed.
  • Action space mapping — 5D continuous: targetSpeed, turnRate (dynamically scaled by speed), gunTurnRate, fireDecision (tanh, threshold 0, masked when gun hot), firePower (sigmoid [0.1,3.0]). Split fire enables probe/nuke bullet strategy.
  • #6 Reward function design — Per-tick energy delta (raw, no shaping) + per-round game score / 100. Gamma 0.99, GAE lambda 0.95. Minimal shaping philosophy.
  • #8 Training loop design — Background-thread PPO training on round end. 4 epochs, mini-batch 64. Episode buffer cleared each round. Atomic .npy weight saves, 3 rotating checkpoints every 50 rounds. Bootstrap GAE on interruption.
  • #9 Network architecture — Separate actor/critic, both 64→tanh→64→tanh. Actor outputs Gaussian mean, separate learnable log-std vector (5 floats). Critic outputs unbounded scalar. ~8K params. Arraymancer defaults for init.
  • #12 Learning rate and optimizer hyperparameters — Adam, LR 3e-4 fixed, both networks. Gradient clipping max norm 0.5. Default Adam betas/epsilon.
  • #11 Bot scaffold and build setup — PPO_Bot/ folder: nimble project with tankroyale_botapi + arraymancer, shell.nix for OpenBLAS, minimal bot that joins a game. Builds via nix-shell.
  • #10 Exploration noise — Entropy bonus (coefficient 0.01) in PPO loss; no explicit action noise. Sufficient for continuous Gaussian policy with learnable log-std.

Not yet specified

  • Multi-opponent generalization strategy (curriculum, opponent pool)
  • Performance profiling under tick time constraints

Out of scope

  • Melee (multi-enemy) support
  • Radar RL control
  • Python/hybrid training pipeline
  • Training statistics / visualization (handled by external battle software)
  • Other bots in this repo (each gets its own effort)
## Destination An RL-trained Tank Royale bot in pure Nim that controls body movement, turret, and firing via continuous actor-critic, trains between rounds, persists weights, and progressively improves against any 1v1 opponent. ## Notes - Domain: Tank Royale (Robocode successor). Nim bot API in sibling repo `/home/davide/Projects/tank-royale` branch `nim`. - Skills to consult: `/grilling`, `/domain-modeling`, `/prototype` - Arraymancer is the primary ML library; hand-rolled MLP is the fallback if Arraymancer proves unusable. - Radar is deterministic (not RL-controlled) — a simple lock/sweep tracker. - Bot folder must be self-contained under SirRoboGarage repo root. - Physics: tick-based, MAX_SPEED=8, MAX_TURN_RATE=10 (speed-dependent), MAX_GUN_TURN_RATE=20, MAX_RADAR_TURN_RATE=45, BOT_RADIUS=18, RADAR_RADIUS=1200. - Action space: targetSpeed ∈ [-8,8], turnRate ∈ [-10,10], gunTurnRate ∈ [-20,20], firePower ∈ {0} ∪ [0.1,3.0]. 4 continuous dimensions. - State: own bot state + tracked enemy state (from deterministic radar). Exact vector TBD. ## Decisions so far - [Arraymancer viability spike](https://git.fossellini.top/SirStone/SirRoboGarage/issues/2) — GO. Arraymancer 0.7.33 works on Nim 2.2.4 + NixOS, XOR MLP trains correctly. - [RL algorithm choice](https://git.fossellini.top/SirStone/SirRoboGarage/issues/3) — PPO. On-policy fits short episodes, simplest from-scratch implementation (~200-300 lines), clip prevents catastrophic collapse. - [Enemy tracker design](https://git.fossellini.top/SirStone/SirRoboGarage/issues/7) — Decoupled deterministic radar with lock+overshoot, dead reckoning on missed ticks, EnemyState struct with inferred firing detection. - [State vector design](https://git.fossellini.top/SirStone/SirRoboGarage/issues/4) — 42 normalized floats: 14 current tick (own+enemy), 8 derived (acceleration, turn rate, bearings, wall distances), 20 enemy window (5 ticks × 4 features). No own history, RNN upgrade path if needed. - [Action space mapping](https://git.fossellini.top/SirStone/SirRoboGarage/issues/5) — 5D continuous: targetSpeed, turnRate (dynamically scaled by speed), gunTurnRate, fireDecision (tanh, threshold 0, masked when gun hot), firePower (sigmoid [0.1,3.0]). Split fire enables probe/nuke bullet strategy. - [#6 Reward function design](https://git.fossellini.top/SirStone/SirRoboGarage/issues/6) — Per-tick energy delta (raw, no shaping) + per-round game score / 100. Gamma 0.99, GAE lambda 0.95. Minimal shaping philosophy. - [#8 Training loop design](https://git.fossellini.top/SirStone/SirRoboGarage/issues/8) — Background-thread PPO training on round end. 4 epochs, mini-batch 64. Episode buffer cleared each round. Atomic .npy weight saves, 3 rotating checkpoints every 50 rounds. Bootstrap GAE on interruption. - [#9 Network architecture](https://git.fossellini.top/SirStone/SirRoboGarage/issues/9) — Separate actor/critic, both 64→tanh→64→tanh. Actor outputs Gaussian mean, separate learnable log-std vector (5 floats). Critic outputs unbounded scalar. ~8K params. Arraymancer defaults for init. - [#12 Learning rate and optimizer hyperparameters](https://git.fossellini.top/SirStone/SirRoboGarage/issues/12) — Adam, LR 3e-4 fixed, both networks. Gradient clipping max norm 0.5. Default Adam betas/epsilon. - [#11 Bot scaffold and build setup](https://git.fossellini.top/SirStone/SirRoboGarage/issues/11) — PPO_Bot/ folder: nimble project with tankroyale_botapi + arraymancer, shell.nix for OpenBLAS, minimal bot that joins a game. Builds via nix-shell. - [#10 Exploration noise](https://git.fossellini.top/SirStone/SirRoboGarage/issues/10) — Entropy bonus (coefficient 0.01) in PPO loss; no explicit action noise. Sufficient for continuous Gaussian policy with learnable log-std. ## Not yet specified - Multi-opponent generalization strategy (curriculum, opponent pool) - Performance profiling under tick time constraints ## Out of scope - Melee (multi-enemy) support - Radar RL control - Python/hybrid training pipeline - Training statistics / visualization (handled by external battle software) - Other bots in this repo (each gets its own effort)
SirStone added the wayfinder:map label 2026-08-15 22:05:04 +02:00
Author
Owner

Child tickets

# Title Label
#2 Arraymancer viability spike wayfinder:task
#3 RL algorithm choice wayfinder:research
#4 State vector design wayfinder:grilling
#5 Action space mapping wayfinder:grilling
#6 Reward function design wayfinder:grilling
#7 Enemy tracker design wayfinder:grilling
#8 Training loop design wayfinder:grilling

Blocking graph

  • #2 blocks #8 (Arraymancer must work before training loop is built)
  • #3 blocks #8 (algorithm choice determines on/off-policy, replay buffer design)
  • #3 blocks #5 (algorithm may constrain action output style)
  • #4 and #5 are independent of each other
  • #6 is independent
  • #7 is independent
## Child tickets | # | Title | Label | |---|-------|-------| | #2 | Arraymancer viability spike | wayfinder:task | | #3 | RL algorithm choice | wayfinder:research | | #4 | State vector design | wayfinder:grilling | | #5 | Action space mapping | wayfinder:grilling | | #6 | Reward function design | wayfinder:grilling | | #7 | Enemy tracker design | wayfinder:grilling | | #8 | Training loop design | wayfinder:grilling | ## Blocking graph - #2 blocks #8 (Arraymancer must work before training loop is built) - #3 blocks #8 (algorithm choice determines on/off-policy, replay buffer design) - #3 blocks #5 (algorithm may constrain action output style) - #4 and #5 are independent of each other - #6 is independent - #7 is independent
Author
Owner

Map cleanup: closed implementation tickets #13–#17, moved #10 to Decisions so far. All decisions made, destination reached. Remaining fog (multi-opponent generalization, performance profiling) is explicitly out of scope for this effort.

Map cleanup: closed implementation tickets #13–#17, moved #10 to Decisions so far. All decisions made, destination reached. Remaining fog (multi-opponent generalization, performance profiling) is explicitly out of scope for this effort.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#1