PPO training loop compatibility #22

Closed
opened 2026-08-17 16:58:18 +02:00 by SirStone · 2 comments
Owner

Blocked by: Output space design for goto/aimTo

Parent map: #18

Question

Does the PPO training loop (GAE, trajectory buffer, action log-prob computation) need structural changes for the new action space?

Considerations:

  • Action space is still 5 continuous Gaussian dims — same shape, different semantics. Does PPO care?
  • The controller introduces a non-differentiable layer between network output and environment effect. Does this break anything for policy gradient?
  • Action log-probabilities are computed on raw network output (pre-squash). This shouldn't change — confirm.
  • Reward is still per-tick. The goto command spans multiple ticks. Does temporal credit assignment become harder? Does GAE lambda need tuning?
**Blocked by:** [Output space design for goto/aimTo](#19) Parent map: #18 ## Question Does the PPO training loop (GAE, trajectory buffer, action log-prob computation) need structural changes for the new action space? Considerations: - Action space is still 5 continuous Gaussian dims — same shape, different semantics. Does PPO care? - The controller introduces a non-differentiable layer between network output and environment effect. Does this break anything for policy gradient? - Action log-probabilities are computed on raw network output (pre-squash). This shouldn't change — confirm. - Reward is still per-tick. The goto command spans multiple ticks. Does temporal credit assignment become harder? Does GAE lambda need tuning?
SirStone added the wayfinder:research label 2026-08-17 16:58:18 +02:00
Author
Owner

Research complete. Full findings at docs/research/ppo-training-compatibility.md on branch research/ppo-training-compat.

Key findings

What needs to change — mechanical, low-risk

  • Every hard-coded 5 in network.nim and training.nim becomes 6. That's it for the core PPO math.
  • actions.nim body rewrites for the new 6-dim semantics (goto/aimTo replaces speed/turnRate/gunTurnRate).
  • PPO_Bot.nim needs to call the controller layer instead of passing motor commands directly.
  • Saved weights in weights/latest/ must be discarded — the actor output layer shape changes from [5,64]→[6,64] and log_std from [5]→[6]. No migration path; re-train from scratch.

What does NOT change

  • TrajectoryBuffer — Transition.action is a heap Tensor[float32], holds any size.
  • GAE — operates on scalar (reward, value) sequences, zero dependency on action dims.
  • Adam states — auto-sized from the parameter tensors at init.
  • PPO update loop structure — mini-batch, advantage normalisation, clipped surrogate, value loss, grad clipping are all action-dimension-agnostic.

Controller is opaque to PPO — correct

The goto/aimTo → motor-command controller is non-differentiable and that's fine. PPO only needs logProb(action | state) on the raw 6-dim network output, computed before the controller. The controller is part of the environment from PPO's perspective. No special treatment needed.

No Jacobian correction needed

The current code stores and trains on pre-squash raw network outputs. Squashing (tanh/sigmoid) happens only in mapActions when sending commands, not on the stored action tensor. Because both collection-time and re-evaluation-time log-probs are computed on the same pre-squash values, the Jacobian terms cancel in the importance-sampling ratio exp(newLogP - oldLogP). This holds regardless of whether the squashing function is tanh or sigmoid. Maintain the pre-squash convention and no correction is needed.

Per-tick reward + multi-tick goto — GAE is fine

The network runs every tick and produces a new policy decision every tick. GAE credits per-tick reward to per-tick decisions. The fact that a single goto(x,y) command spans multiple motor ticks is the controller's concern, not GAE's. No structural change to the reward or advantage computation is needed.

One gotcha: fire decision dim 4

Dim 4 is a tanh-threshold (fire if >= 0), not a continuous density target. The Gaussian log-prob for dim 4 is still computed normally — the gradient nudges the mean toward positive/negative territory, which is the correct learning signal for a binary decision via continuous relaxation. Keep it in the loop.

Research complete. Full findings at `docs/research/ppo-training-compatibility.md` on branch `research/ppo-training-compat`. ## Key findings **What needs to change — mechanical, low-risk** - Every hard-coded `5` in `network.nim` and `training.nim` becomes `6`. That's it for the core PPO math. - `actions.nim` body rewrites for the new 6-dim semantics (goto/aimTo replaces speed/turnRate/gunTurnRate). - `PPO_Bot.nim` needs to call the controller layer instead of passing motor commands directly. - Saved weights in `weights/latest/` must be discarded — the actor output layer shape changes from `[5,64]→[6,64]` and `log_std` from `[5]→[6]`. No migration path; re-train from scratch. **What does NOT change** - `TrajectoryBuffer` — `Transition.action` is a heap `Tensor[float32]`, holds any size. - GAE — operates on scalar `(reward, value)` sequences, zero dependency on action dims. - Adam states — auto-sized from the parameter tensors at init. - PPO update loop structure — mini-batch, advantage normalisation, clipped surrogate, value loss, grad clipping are all action-dimension-agnostic. **Controller is opaque to PPO — correct** The goto/aimTo → motor-command controller is non-differentiable and that's fine. PPO only needs `logProb(action | state)` on the raw 6-dim network output, computed before the controller. The controller is part of the environment from PPO's perspective. No special treatment needed. **No Jacobian correction needed** The current code stores and trains on **pre-squash** raw network outputs. Squashing (tanh/sigmoid) happens only in `mapActions` when sending commands, not on the stored `action` tensor. Because both collection-time and re-evaluation-time log-probs are computed on the same pre-squash values, the Jacobian terms cancel in the importance-sampling ratio `exp(newLogP - oldLogP)`. This holds regardless of whether the squashing function is tanh or sigmoid. Maintain the pre-squash convention and no correction is needed. **Per-tick reward + multi-tick goto — GAE is fine** The network runs every tick and produces a new policy decision every tick. GAE credits per-tick reward to per-tick decisions. The fact that a single goto(x,y) command spans multiple motor ticks is the controller's concern, not GAE's. No structural change to the reward or advantage computation is needed. **One gotcha: fire decision dim 4** Dim 4 is a tanh-threshold (fire if >= 0), not a continuous density target. The Gaussian log-prob for dim 4 is still computed normally — the gradient nudges the mean toward positive/negative territory, which is the correct learning signal for a binary decision via continuous relaxation. Keep it in the loop.
Author
Owner

Resolution

All changes are mechanical — no structural surgery needed.

Changes required:

  • 5 → 6 in network.nim (initActorCritic, actorForward, computeLogProb) and training.nim (log-prob loop, dLogP_dMean)
  • actions.nim rewritten for new goto/aimTo semantics
  • State input size 42 → 44 in network.nim
  • Old weights discarded — actor output layer and log_std change shape

No changes needed:

  • TrajectoryBuffer — action is a heap tensor, size-agnostic
  • GAE — scalar reward/value, zero coupling to action dims
  • PPO update loop, gradient clipping, advantage normalization — all dim-agnostic
  • No Jacobian correction needed — pre-squash values used consistently

Research doc: docs/research/ppo-training-compatibility.md on branch research/ppo-training-compat

## Resolution **All changes are mechanical — no structural surgery needed.** Changes required: - `5` → `6` in network.nim (initActorCritic, actorForward, computeLogProb) and training.nim (log-prob loop, dLogP_dMean) - `actions.nim` rewritten for new goto/aimTo semantics - State input size `42` → `44` in network.nim - Old weights discarded — actor output layer and log_std change shape No changes needed: - TrajectoryBuffer — action is a heap tensor, size-agnostic - GAE — scalar reward/value, zero coupling to action dims - PPO update loop, gradient clipping, advantage normalization — all dim-agnostic - No Jacobian correction needed — pre-squash values used consistently Research doc: `docs/research/ppo-training-compatibility.md` on branch `research/ppo-training-compat`
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#22