eadd177d3b2cfba60fa122f237348e662766c625
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate (4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers computeTickReward/computeRoundReward. Bot wired: tick transitions collected in run loop, ppoUpdate called on onRoundEnded. Fix: add arraymancer import to PPO_Bot.nim so Tensor resolves at top level. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Description
No description provided
Languages
Nim
73.7%
Python
18%
Shell
3.7%
Java
3.5%
HTML
1%
Other
0.1%