feat(PPO_Bot): reward + trajectory + GAE + PPO training (#16)
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate (4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers computeTickReward/computeRoundReward. Bot wired: tick transitions collected in run loop, ppoUpdate called on onRoundEnded. Fix: add arraymancer import to PPO_Bot.nim so Tensor resolves at top level. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Executable
+4
@@ -0,0 +1,4 @@
|
||||
#!/bin/sh
|
||||
# PPO_Bot — PPO-trained RL bot (compiled native binary)
|
||||
cd -- "$(dirname -- "$0")"
|
||||
exec "./PPO_Bot"
|
||||
Reference in New Issue
Block a user