Command abstraction layer for PPO action space #18

Closed
opened 2026-08-17 16:57:55 +02:00 by SirStone · 1 comment
Owner

Destination

Replace the PPO bot's per-tick motor outputs (speed, turn rate, gun turn rate) with a command abstraction layer: goto(x, y) for movement and aimTo(angle) for gun aiming. A controller translates these into per-tick motor commands, choosing forward/reverse optimally. Fire decision and fire power stay as raw per-tick outputs. The state vector gains remaining-distance and remaining-gun-angle as inputs so the network learns temporal consistency. The network still runs every tick.

Notes

  • Domain: Tank Royale bot, PPO reinforcement learning in Nim
  • Branch: research/rl-algorithm-choice
  • Key files: PPO_Bot/PPO_Bot.nim, PPO_Bot/training.nim, PPO_Bot/actions.nim
  • Skills: /grilling, /domain-modeling, /prototype
  • The bot currently outputs 5 continuous dims per tick: target speed, body turn rate, gun turn rate, fire decision, fire power. See commit f27b023.

Decisions so far

Not yet specified

  • Reward shaping — may need changes once training with the new action space reveals learning dynamics
  • Episode boundary handling — does the goto command reset between episodes/rounds?
  • Whether the network needs memory (RNN/attention) to work well with this action space, or if remaining-distance-as-input is sufficient
  • Calibration of output ranges and normalization once prototyped

Out of scope

(none yet)

## Destination Replace the PPO bot's per-tick motor outputs (speed, turn rate, gun turn rate) with a **command abstraction layer**: `goto(x, y)` for movement and `aimTo(angle)` for gun aiming. A controller translates these into per-tick motor commands, choosing forward/reverse optimally. Fire decision and fire power stay as raw per-tick outputs. The state vector gains remaining-distance and remaining-gun-angle as inputs so the network learns temporal consistency. The network still runs every tick. ## Notes - Domain: Tank Royale bot, PPO reinforcement learning in Nim - Branch: `research/rl-algorithm-choice` - Key files: `PPO_Bot/PPO_Bot.nim`, `PPO_Bot/training.nim`, `PPO_Bot/actions.nim` - Skills: `/grilling`, `/domain-modeling`, `/prototype` - The bot currently outputs 5 continuous dims per tick: target speed, body turn rate, gun turn rate, fire decision, fire power. See commit f27b023. ## Decisions so far - [Output space design for goto/aimTo](https://git.fossellini.top/SirStone/SirRoboGarage/issues/19) — 6-dim output: goto(x,y) + aimTo(x,y) as sigmoid-bounded battlefield coords, fire decision + fire power explicit - [Goto controller algorithm](https://git.fossellini.top/SirStone/SirRoboGarage/issues/20) — proportional steering with forward/reverse at 90° threshold, reuses `getNewTargetSpeed` from utils.nim - [State vector additions](https://git.fossellini.top/SirStone/SirRoboGarage/issues/21) — 2 new inputs: remaining goto distance (÷ diagonal) and remaining gun angle (÷ 180°), state 42→44 dims - [PPO training loop compatibility](https://git.fossellini.top/SirStone/SirRoboGarage/issues/22) — all mechanical changes (5→6 output dims, 42→44 input dims, discard weights), no structural PPO changes needed - [Prototype goto controller and test in game](https://git.fossellini.top/SirStone/SirRoboGarage/issues/23) — goto controller validated: proportional steering + 90° reverse threshold, smooth arcs, handles wall proximity. Prototype at GotoTest/ ## Not yet specified - Reward shaping — may need changes once training with the new action space reveals learning dynamics - Episode boundary handling — does the goto command reset between episodes/rounds? - Whether the network needs memory (RNN/attention) to work well with this action space, or if remaining-distance-as-input is sufficient - Calibration of output ranges and normalization once prototyped ## Out of scope _(none yet)_
SirStone added the wayfinder:map label 2026-08-17 16:57:55 +02:00
Author
Owner

Map complete

All 5 decision tickets resolved. The command abstraction layer design is fully decided — implementation tracked under PRD #24.

Decisions:

  • #19 — Output space: 6-dim (goto x/y, aimTo x/y, fire decision, fire power)
  • #20 — Goto controller: proportional steering, reverse at 90° bearing threshold
  • #21 — State vector: 42→44 dims (remaining distance + remaining gun angle)
  • #22 — PPO loop: no structural changes needed
  • #23 — Prototype validated in game (GotoTest/)
## Map complete All 5 decision tickets resolved. The command abstraction layer design is fully decided — implementation tracked under PRD #24. Decisions: - #19 — Output space: 6-dim (goto x/y, aimTo x/y, fire decision, fire power) - #20 — Goto controller: proportional steering, reverse at 90° bearing threshold - #21 — State vector: 42→44 dims (remaining distance + remaining gun angle) - #22 — PPO loop: no structural changes needed - #23 — Prototype validated in game (GotoTest/)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#18