Action space mapping #5

Closed
opened 2026-08-15 22:05:24 +02:00 by SirStone · 3 comments
Owner

Question

How does the actor network's output map to bot commands each tick?

The network outputs 4 continuous values. Proposed mapping:

  • tanh × 8 → targetSpeed
  • tanh × 10 → turnRate (but actual max depends on speed: 10 - 0.75×|speed|)
  • tanh × 20 → gunTurnRate
  • sigmoid-style → firePower: 0 (don't fire) or [0.1, 3.0] (fire)

Open sub-questions:

  • How to handle the fire threshold? Network output < threshold → don't fire, above → scale to [0.1, 3.0]?
  • Should turnRate be clamped post-hoc to the speed-dependent max, or should the network learn the constraint?
  • Gun heat prevents firing — mask the fire action when gun is hot, or let the network learn to not waste the output?
## Question How does the actor network's output map to bot commands each tick? The network outputs 4 continuous values. Proposed mapping: - tanh × 8 → targetSpeed - tanh × 10 → turnRate (but actual max depends on speed: 10 - 0.75×|speed|) - tanh × 20 → gunTurnRate - sigmoid-style → firePower: 0 (don't fire) or [0.1, 3.0] (fire) Open sub-questions: - How to handle the fire threshold? Network output < threshold → don't fire, above → scale to [0.1, 3.0]? - Should turnRate be clamped post-hoc to the speed-dependent max, or should the network learn the constraint? - Gun heat prevents firing — mask the fire action when gun is hot, or let the network learn to not waste the output?
SirStone added the wayfinder:grilling label 2026-08-15 22:05:24 +02:00
Author
Owner

Blocked by #3 (RL algorithm choice) — the algorithm may constrain action output style (e.g. SAC uses squashed Gaussian, PPO uses clipped policy, DDPG/TD3 use deterministic actor).

Blocked by #3 (RL algorithm choice) — the algorithm may constrain action output style (e.g. SAC uses squashed Gaussian, PPO uses clipped policy, DDPG/TD3 use deterministic actor).
Author
Owner

Resolution

5D continuous action space

# Output Activation Range Meaning
1 targetSpeed tanh × 8 [-8, 8] body speed
2 turnRate tanh × currentMaxTurnRate [-max, max] body turn (dynamically scaled by speed: max = 10 - 0.75 ×
3 gunTurnRate tanh × 20 [-20, 20] gun turn
4 fireDecision tanh [-1, 1] < 0 = hold, ≥ 0 = fire
5 firePower sigmoid × 2.9 + 0.1 [0.1, 3.0] bullet power (only used when fireDecision ≥ 0)

Design decisions

  • Dynamic turnRate scaling — output is tanh × speed-dependent max, so the network always uses full [-1, 1] range meaningfully instead of wasting range on impossible turn rates.
  • Split fire into two outputs — fireDecision (whether) + firePower (how hard). Enables the network to learn advanced techniques: low-power fast bullets (speed 19.7) as probes, high-power slow bullets (speed 11.0) for confirmed hits.
  • Fire threshold at 0 (neutral) — untrained network is equally likely to fire or hold. Reward function teaches restraint.
  • Gun heat masking — when gunHeat > 0, fireDecision is forced to "hold" regardless of network output. No wasted learning signal on impossible actions.
  • Bullet speed reference: firePower 0.1 → speed 19.7, firePower 1.0 → speed 17.0, firePower 3.0 → speed 11.0.
## Resolution ### 5D continuous action space | # | Output | Activation | Range | Meaning | |---|--------|-----------|-------|---------| | 1 | targetSpeed | tanh × 8 | [-8, 8] | body speed | | 2 | turnRate | tanh × currentMaxTurnRate | [-max, max] | body turn (dynamically scaled by speed: max = 10 - 0.75 × |speed|) | | 3 | gunTurnRate | tanh × 20 | [-20, 20] | gun turn | | 4 | fireDecision | tanh | [-1, 1] | < 0 = hold, ≥ 0 = fire | | 5 | firePower | sigmoid × 2.9 + 0.1 | [0.1, 3.0] | bullet power (only used when fireDecision ≥ 0) | ### Design decisions - **Dynamic turnRate scaling** — output is tanh × speed-dependent max, so the network always uses full [-1, 1] range meaningfully instead of wasting range on impossible turn rates. - **Split fire into two outputs** — fireDecision (whether) + firePower (how hard). Enables the network to learn advanced techniques: low-power fast bullets (speed 19.7) as probes, high-power slow bullets (speed 11.0) for confirmed hits. - **Fire threshold at 0 (neutral)** — untrained network is equally likely to fire or hold. Reward function teaches restraint. - **Gun heat masking** — when gunHeat > 0, fireDecision is forced to "hold" regardless of network output. No wasted learning signal on impossible actions. - **Bullet speed reference:** firePower 0.1 → speed 19.7, firePower 1.0 → speed 17.0, firePower 3.0 → speed 11.0.
Author
Owner

Duplicate of #6

Duplicate of #6
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#5