feat(PPO_Bot): reward + trajectory + GAE + PPO training (#16)

Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate
(4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers
computeTickReward/computeRoundReward. Bot wired: tick transitions
collected in run loop, ppoUpdate called on onRoundEnded. Fix: add
arraymancer import to PPO_Bot.nim so Tensor resolves at top level.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-08-16 15:27:15 +02:00
parent 588c9ebc2f
commit eadd177d3b
10 changed files with 572 additions and 5 deletions
+11
View File
@@ -0,0 +1,11 @@
{
"name": "PPO_Bot",
"version": "0.1.0",
"authors": ["Davide Cappellini"],
"description": "PPO-trained RL bot",
"homepage": "",
"countryCodes": ["IT"],
"gameTypes": ["classic", "melee", "1v1"],
"platform": "Nim",
"programmingLang": "Nim"
}