PRD: Command abstraction layer for PPO action space #24

Closed
opened 2026-08-17 18:51:59 +02:00 by SirStone · 0 comments
Owner

Problem Statement

The PPO bot outputs raw per-tick motor commands (target speed, body turn rate, gun turn rate) at every game tick. Because the neural network recomputes these from scratch each tick with no concept of ongoing movement, the bot exhibits jerky start-stop behavior — constantly reversing direction, oscillating turn rate, and never committing to fluid multi-tick maneuvers. The action space gives the network maximum control but zero temporal coherence.

Solution

Replace the raw motor outputs with a command abstraction layer. The network outputs high-level spatial commands — goto(x, y) for movement and aimTo(x, y) for gun aiming — and a deterministic controller translates these into smooth per-tick motor commands. The network still runs every tick and retains explicit fire control. Two new state inputs (remaining goto distance, remaining gun angle) close the feedback loop so the network can see its unfinished commands and learn temporal consistency without architectural tricks like RNNs.

User Stories

  1. As a bot trainer, I want the bot to move in smooth arcs toward positions, so that movement looks intentional rather than random.
  2. As a bot trainer, I want the bot to commit to a destination for multiple ticks, so that it builds momentum and covers ground efficiently.
  3. As a bot trainer, I want the bot to reverse when that's faster than turning 180°, so that it can dodge and reposition optimally.
  4. As a bot trainer, I want the bot to aim its gun at a point on the battlefield, so that gun tracking is smooth and converges on targets.
  5. As a bot trainer, I want the bot to retain explicit fire control (when and how hard), so that it learns firing tactics independently of aiming.
  6. As a bot trainer, I want the bot to see how far it is from its goto target each tick, so that it can learn to hold course or issue new commands.
  7. As a bot trainer, I want the bot to see how far the gun is from its aimTo target each tick, so that it can learn to time shots when the gun is on target.
  8. As a bot trainer, I want the neural network to still run every tick, so that the bot can react to new threats (getting hit, enemy movement) mid-command.
  9. As a bot trainer, I want the goto controller to decelerate smoothly near the target, so that the bot doesn't overshoot and oscillate.
  10. As a bot trainer, I want the action space bounded to battlefield dimensions, so that the network can never output off-field positions.
  11. As a bot trainer, I want training to resume with the new architecture without structural PPO changes, so that the training infrastructure remains stable.

Implementation Decisions

Action space (decided in wayfinder #19)

6-dimensional continuous Gaussian output, up from 5:

Dim Output Encoding
0 goto x sigmoid × battlefield width
1 goto y sigmoid × battlefield height
2 aimTo x sigmoid × battlefield width
3 aimTo y sigmoid × battlefield height
4 fire decision tanh → fire if ≥ 0
5 fire power sigmoid × 2.9 + 0.1

goto and aimTo both use absolute battlefield coordinates, sigmoid-bounded. aimTo is a point (not an angle) — the controller computes the gun angle via bearing calculation, avoiding circular encoding discontinuities.

Goto controller (decided in wayfinder #20, validated in #23)

Proportional steering with forward/reverse selection. From the validated prototype:

# Per-tick goto controller (from GotoTest prototype)
rawBearing = normalizeRelativeAngle(directionTo(botX, botY, targetX, targetY) - heading)

if abs(rawBearing) > 90.0:
  dirSign = -1.0
  effBearing = normalizeRelativeAngle(rawBearing + 180.0)
else:
  dirSign = 1.0
  effBearing = rawBearing

turnRate = clamp(effBearing, -calcMaxTurnRate(speed), calcMaxTurnRate(speed))
targetSpeed = dirSign * getNewTargetSpeed(MAX_SPEED, abs(speed), distance)

Key facts: getNewTargetSpeed already exists in utils.nim (handles asymmetric accel+1/decel-2). calcMaxTurnRate is from tankroyale_botapi. No new dependencies.

AimTo controller

gunAngle = directionTo(botX, botY, aimX, aimY)
gunTurnNeeded = normalizeRelativeAngle(gunAngle - gunDirection)
gunTurnRate = clamp(gunTurnNeeded, -MAX_GUN_TURN_RATE, MAX_GUN_TURN_RATE)

Gun independent of body via setAdjustGunForBodyTurn(true).

State vector (decided in wayfinder #21)

Two new inputs appended to the 42-dim state vector (now 44):

Input Normalization
Remaining distance to goto target ÷ battlefield diagonal → [0, 1]
Remaining gun angle to aimTo target ÷ 180° → [0, 1]

Network changes (confirmed in wayfinder #22)

All mechanical:

  • Actor output layer: 5 → 6
  • Learnable logStd: 5 → 6
  • Input layer: 42 → 44
  • Old weights must be discarded — shapes change

PPO training loop

No structural changes. The trajectory buffer, GAE, PPO update loop, and action log-prob computation are all dimension-agnostic. The controller is non-differentiable but sits on the environment side of the policy gradient — PPO operates on raw network outputs (pre-squash), which is unchanged.

Modules modified

  • actions.nim — rewrite mapActions for new 6-dim semantics, add goto and aimTo controller functions
  • state_vector.nim — append 2 new inputs, update dimension constant
  • network.nim — update input/output dimension constants
  • training.nim — update dimension constant for log-prob computation
  • PPO_Bot.nim — track current goto/aimTo targets for remaining-distance state computation

Radar

Unchanged — radar lock is handled by EnemyTracker outside the action space.

Testing Decisions

Good tests for this feature verify external behavior at module seams, not internal wiring.

Action decoding (actions.nim)

  • Given raw 6-dim network outputs, verify goto coordinates are sigmoid-bounded to battlefield dimensions
  • Given raw outputs, verify aimTo coordinates are sigmoid-bounded
  • Given raw outputs with dim 4 ≥ 0 and gunHeat ≤ 0, verify fire decision is true
  • Given raw outputs with dim 4 < 0, verify fire decision is false

Goto controller (actions.nim or new module)

  • Given a target ahead (bearing < 90°): verify positive speed (forward), correct turn direction
  • Given a target behind (bearing > 90°): verify negative speed (reverse), adjusted bearing
  • Given a target at close range: verify speed ramps down (deceleration via getNewTargetSpeed)
  • Given a target at current position: verify near-zero speed

AimTo controller

  • Given aimTo target to the right of gun: verify positive gun turn rate
  • Given aimTo target to the left of gun: verify negative gun turn rate
  • Given aimTo target aligned with gun: verify near-zero gun turn rate
  • Verify gun turn rate never exceeds ±20°/tick

State vector (state_vector.nim)

  • Verify output is 44-dim (was 42)
  • Verify remaining goto distance is normalized to [0, 1] using battlefield diagonal
  • Verify remaining gun angle is normalized to [0, 1] using 180°

Integration

  • GotoTest prototype bot in battle runner — validates the full goto+aimTo pipeline in a live game
  • Prior art: tools/battle_runner/run.sh runs a bot vs Target for 1 round

Out of Scope

  • Reward shaping — may need changes once training with the new action space reveals learning dynamics, but not part of this implementation
  • Episode boundary handling — whether goto commands reset between rounds; current behavior (network runs every tick) likely handles this naturally
  • RNN/memory architecture — the remaining-distance feedback loop may be sufficient; evaluate after training
  • Output normalization tuning — calibrate after observing training behavior
  • Stuck detection in PPO_Bot — the GotoTest prototype has a stuck detector; the PPO network should learn to avoid stuck states via reward signal rather than a hardcoded escape heuristic
  • Wall avoidance — the network should learn wall avoidance through negative reward, not a controller-level override

Further Notes

  • Wayfinder map: #18 — all 5 decision tickets resolved
  • Goto controller research: docs/research/goto-controller-algorithm.md (branch research/goto-controller)
  • PPO compatibility research: docs/research/ppo-training-compatibility.md (branch research/ppo-training-compat)
  • Prototype: GotoTest/ — throwaway bot validating goto + aimTo controllers
  • All existing weights must be discarded when this ships — fresh training run required
## Problem Statement The PPO bot outputs raw per-tick motor commands (target speed, body turn rate, gun turn rate) at every game tick. Because the neural network recomputes these from scratch each tick with no concept of ongoing movement, the bot exhibits jerky start-stop behavior — constantly reversing direction, oscillating turn rate, and never committing to fluid multi-tick maneuvers. The action space gives the network maximum control but zero temporal coherence. ## Solution Replace the raw motor outputs with a **command abstraction layer**. The network outputs high-level spatial commands — `goto(x, y)` for movement and `aimTo(x, y)` for gun aiming — and a deterministic controller translates these into smooth per-tick motor commands. The network still runs every tick and retains explicit fire control. Two new state inputs (remaining goto distance, remaining gun angle) close the feedback loop so the network can see its unfinished commands and learn temporal consistency without architectural tricks like RNNs. ## User Stories 1. As a bot trainer, I want the bot to move in smooth arcs toward positions, so that movement looks intentional rather than random. 2. As a bot trainer, I want the bot to commit to a destination for multiple ticks, so that it builds momentum and covers ground efficiently. 3. As a bot trainer, I want the bot to reverse when that's faster than turning 180°, so that it can dodge and reposition optimally. 4. As a bot trainer, I want the bot to aim its gun at a point on the battlefield, so that gun tracking is smooth and converges on targets. 5. As a bot trainer, I want the bot to retain explicit fire control (when and how hard), so that it learns firing tactics independently of aiming. 6. As a bot trainer, I want the bot to see how far it is from its goto target each tick, so that it can learn to hold course or issue new commands. 7. As a bot trainer, I want the bot to see how far the gun is from its aimTo target each tick, so that it can learn to time shots when the gun is on target. 8. As a bot trainer, I want the neural network to still run every tick, so that the bot can react to new threats (getting hit, enemy movement) mid-command. 9. As a bot trainer, I want the goto controller to decelerate smoothly near the target, so that the bot doesn't overshoot and oscillate. 10. As a bot trainer, I want the action space bounded to battlefield dimensions, so that the network can never output off-field positions. 11. As a bot trainer, I want training to resume with the new architecture without structural PPO changes, so that the training infrastructure remains stable. ## Implementation Decisions ### Action space (decided in wayfinder #19) 6-dimensional continuous Gaussian output, up from 5: | Dim | Output | Encoding | |-----|--------|----------| | 0 | goto x | sigmoid × battlefield width | | 1 | goto y | sigmoid × battlefield height | | 2 | aimTo x | sigmoid × battlefield width | | 3 | aimTo y | sigmoid × battlefield height | | 4 | fire decision | tanh → fire if ≥ 0 | | 5 | fire power | sigmoid × 2.9 + 0.1 | `goto` and `aimTo` both use absolute battlefield coordinates, sigmoid-bounded. `aimTo` is a point (not an angle) — the controller computes the gun angle via bearing calculation, avoiding circular encoding discontinuities. ### Goto controller (decided in wayfinder #20, validated in #23) Proportional steering with forward/reverse selection. From the validated prototype: ```nim # Per-tick goto controller (from GotoTest prototype) rawBearing = normalizeRelativeAngle(directionTo(botX, botY, targetX, targetY) - heading) if abs(rawBearing) > 90.0: dirSign = -1.0 effBearing = normalizeRelativeAngle(rawBearing + 180.0) else: dirSign = 1.0 effBearing = rawBearing turnRate = clamp(effBearing, -calcMaxTurnRate(speed), calcMaxTurnRate(speed)) targetSpeed = dirSign * getNewTargetSpeed(MAX_SPEED, abs(speed), distance) ``` Key facts: `getNewTargetSpeed` already exists in `utils.nim` (handles asymmetric accel+1/decel-2). `calcMaxTurnRate` is from `tankroyale_botapi`. No new dependencies. ### AimTo controller ```nim gunAngle = directionTo(botX, botY, aimX, aimY) gunTurnNeeded = normalizeRelativeAngle(gunAngle - gunDirection) gunTurnRate = clamp(gunTurnNeeded, -MAX_GUN_TURN_RATE, MAX_GUN_TURN_RATE) ``` Gun independent of body via `setAdjustGunForBodyTurn(true)`. ### State vector (decided in wayfinder #21) Two new inputs appended to the 42-dim state vector (now 44): | Input | Normalization | |-------|--------------| | Remaining distance to goto target | ÷ battlefield diagonal → [0, 1] | | Remaining gun angle to aimTo target | ÷ 180° → [0, 1] | ### Network changes (confirmed in wayfinder #22) All mechanical: - Actor output layer: 5 → 6 - Learnable logStd: 5 → 6 - Input layer: 42 → 44 - Old weights must be discarded — shapes change ### PPO training loop No structural changes. The trajectory buffer, GAE, PPO update loop, and action log-prob computation are all dimension-agnostic. The controller is non-differentiable but sits on the environment side of the policy gradient — PPO operates on raw network outputs (pre-squash), which is unchanged. ### Modules modified - `actions.nim` — rewrite `mapActions` for new 6-dim semantics, add goto and aimTo controller functions - `state_vector.nim` — append 2 new inputs, update dimension constant - `network.nim` — update input/output dimension constants - `training.nim` — update dimension constant for log-prob computation - `PPO_Bot.nim` — track current goto/aimTo targets for remaining-distance state computation ### Radar Unchanged — radar lock is handled by `EnemyTracker` outside the action space. ## Testing Decisions Good tests for this feature verify **external behavior** at module seams, not internal wiring. ### Action decoding (`actions.nim`) - Given raw 6-dim network outputs, verify goto coordinates are sigmoid-bounded to battlefield dimensions - Given raw outputs, verify aimTo coordinates are sigmoid-bounded - Given raw outputs with dim 4 ≥ 0 and gunHeat ≤ 0, verify fire decision is true - Given raw outputs with dim 4 < 0, verify fire decision is false ### Goto controller (`actions.nim` or new module) - Given a target ahead (bearing < 90°): verify positive speed (forward), correct turn direction - Given a target behind (bearing > 90°): verify negative speed (reverse), adjusted bearing - Given a target at close range: verify speed ramps down (deceleration via `getNewTargetSpeed`) - Given a target at current position: verify near-zero speed ### AimTo controller - Given aimTo target to the right of gun: verify positive gun turn rate - Given aimTo target to the left of gun: verify negative gun turn rate - Given aimTo target aligned with gun: verify near-zero gun turn rate - Verify gun turn rate never exceeds ±20°/tick ### State vector (`state_vector.nim`) - Verify output is 44-dim (was 42) - Verify remaining goto distance is normalized to [0, 1] using battlefield diagonal - Verify remaining gun angle is normalized to [0, 1] using 180° ### Integration - GotoTest prototype bot in battle runner — validates the full goto+aimTo pipeline in a live game - Prior art: `tools/battle_runner/run.sh` runs a bot vs Target for 1 round ## Out of Scope - **Reward shaping** — may need changes once training with the new action space reveals learning dynamics, but not part of this implementation - **Episode boundary handling** — whether goto commands reset between rounds; current behavior (network runs every tick) likely handles this naturally - **RNN/memory architecture** — the remaining-distance feedback loop may be sufficient; evaluate after training - **Output normalization tuning** — calibrate after observing training behavior - **Stuck detection in PPO_Bot** — the GotoTest prototype has a stuck detector; the PPO network should learn to avoid stuck states via reward signal rather than a hardcoded escape heuristic - **Wall avoidance** — the network should learn wall avoidance through negative reward, not a controller-level override ## Further Notes - Wayfinder map: [#18](https://git.fossellini.top/SirStone/SirRoboGarage/issues/18) — all 5 decision tickets resolved - Goto controller research: `docs/research/goto-controller-algorithm.md` (branch `research/goto-controller`) - PPO compatibility research: `docs/research/ppo-training-compatibility.md` (branch `research/ppo-training-compat`) - Prototype: `GotoTest/` — throwaway bot validating goto + aimTo controllers - All existing weights must be discarded when this ships — fresh training run required
SirStone added the ready-for-agent label 2026-08-17 18:51:59 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#24