PPO training loop compatibility #22
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Blocked by: Output space design for goto/aimTo
Parent map: #18
Question
Does the PPO training loop (GAE, trajectory buffer, action log-prob computation) need structural changes for the new action space?
Considerations:
Research complete. Full findings at
docs/research/ppo-training-compatibility.mdon branchresearch/ppo-training-compat.Key findings
What needs to change — mechanical, low-risk
5innetwork.nimandtraining.nimbecomes6. That's it for the core PPO math.actions.nimbody rewrites for the new 6-dim semantics (goto/aimTo replaces speed/turnRate/gunTurnRate).PPO_Bot.nimneeds to call the controller layer instead of passing motor commands directly.weights/latest/must be discarded — the actor output layer shape changes from[5,64]→[6,64]andlog_stdfrom[5]→[6]. No migration path; re-train from scratch.What does NOT change
TrajectoryBuffer—Transition.actionis a heapTensor[float32], holds any size.(reward, value)sequences, zero dependency on action dims.Controller is opaque to PPO — correct
The goto/aimTo → motor-command controller is non-differentiable and that's fine. PPO only needs
logProb(action | state)on the raw 6-dim network output, computed before the controller. The controller is part of the environment from PPO's perspective. No special treatment needed.No Jacobian correction needed
The current code stores and trains on pre-squash raw network outputs. Squashing (tanh/sigmoid) happens only in
mapActionswhen sending commands, not on the storedactiontensor. Because both collection-time and re-evaluation-time log-probs are computed on the same pre-squash values, the Jacobian terms cancel in the importance-sampling ratioexp(newLogP - oldLogP). This holds regardless of whether the squashing function is tanh or sigmoid. Maintain the pre-squash convention and no correction is needed.Per-tick reward + multi-tick goto — GAE is fine
The network runs every tick and produces a new policy decision every tick. GAE credits per-tick reward to per-tick decisions. The fact that a single goto(x,y) command spans multiple motor ticks is the controller's concern, not GAE's. No structural change to the reward or advantage computation is needed.
One gotcha: fire decision dim 4
Dim 4 is a tanh-threshold (fire if >= 0), not a continuous density target. The Gaussian log-prob for dim 4 is still computed normally — the gradient nudges the mean toward positive/negative territory, which is the correct learning signal for a binary decision via continuous relaxation. Keep it in the loop.
Resolution
All changes are mechanical — no structural surgery needed.
Changes required:
5→6in network.nim (initActorCritic, actorForward, computeLogProb) and training.nim (log-prob loop, dLogP_dMean)actions.nimrewritten for new goto/aimTo semantics42→44in network.nimNo changes needed:
Research doc:
docs/research/ppo-training-compatibility.mdon branchresearch/ppo-training-compat