PRD: Command abstraction layer for PPO action space #24
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem Statement
The PPO bot outputs raw per-tick motor commands (target speed, body turn rate, gun turn rate) at every game tick. Because the neural network recomputes these from scratch each tick with no concept of ongoing movement, the bot exhibits jerky start-stop behavior — constantly reversing direction, oscillating turn rate, and never committing to fluid multi-tick maneuvers. The action space gives the network maximum control but zero temporal coherence.
Solution
Replace the raw motor outputs with a command abstraction layer. The network outputs high-level spatial commands —
goto(x, y)for movement andaimTo(x, y)for gun aiming — and a deterministic controller translates these into smooth per-tick motor commands. The network still runs every tick and retains explicit fire control. Two new state inputs (remaining goto distance, remaining gun angle) close the feedback loop so the network can see its unfinished commands and learn temporal consistency without architectural tricks like RNNs.User Stories
Implementation Decisions
Action space (decided in wayfinder #19)
6-dimensional continuous Gaussian output, up from 5:
gotoandaimToboth use absolute battlefield coordinates, sigmoid-bounded.aimTois a point (not an angle) — the controller computes the gun angle via bearing calculation, avoiding circular encoding discontinuities.Goto controller (decided in wayfinder #20, validated in #23)
Proportional steering with forward/reverse selection. From the validated prototype:
Key facts:
getNewTargetSpeedalready exists inutils.nim(handles asymmetric accel+1/decel-2).calcMaxTurnRateis fromtankroyale_botapi. No new dependencies.AimTo controller
Gun independent of body via
setAdjustGunForBodyTurn(true).State vector (decided in wayfinder #21)
Two new inputs appended to the 42-dim state vector (now 44):
Network changes (confirmed in wayfinder #22)
All mechanical:
PPO training loop
No structural changes. The trajectory buffer, GAE, PPO update loop, and action log-prob computation are all dimension-agnostic. The controller is non-differentiable but sits on the environment side of the policy gradient — PPO operates on raw network outputs (pre-squash), which is unchanged.
Modules modified
actions.nim— rewritemapActionsfor new 6-dim semantics, add goto and aimTo controller functionsstate_vector.nim— append 2 new inputs, update dimension constantnetwork.nim— update input/output dimension constantstraining.nim— update dimension constant for log-prob computationPPO_Bot.nim— track current goto/aimTo targets for remaining-distance state computationRadar
Unchanged — radar lock is handled by
EnemyTrackeroutside the action space.Testing Decisions
Good tests for this feature verify external behavior at module seams, not internal wiring.
Action decoding (
actions.nim)Goto controller (
actions.nimor new module)getNewTargetSpeed)AimTo controller
State vector (
state_vector.nim)Integration
tools/battle_runner/run.shruns a bot vs Target for 1 roundOut of Scope
Further Notes
docs/research/goto-controller-algorithm.md(branchresearch/goto-controller)docs/research/ppo-training-compatibility.md(branchresearch/ppo-training-compat)GotoTest/— throwaway bot validating goto + aimTo controllers