eadd177d3b
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate (4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers computeTickReward/computeRoundReward. Bot wired: tick transitions collected in run loop, ppoUpdate called on onRoundEnded. Fix: add arraymancer import to PPO_Bot.nim so Tensor resolves at top level. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
12 lines
258 B
JSON
12 lines
258 B
JSON
{
|
|
"name": "PPO_Bot",
|
|
"version": "0.1.0",
|
|
"authors": ["Davide Cappellini"],
|
|
"description": "PPO-trained RL bot",
|
|
"homepage": "",
|
|
"countryCodes": ["IT"],
|
|
"gameTypes": ["classic", "melee", "1v1"],
|
|
"platform": "Nim",
|
|
"programmingLang": "Nim"
|
|
}
|