eadd177d3b
Manual-backprop PPO with Adam: TrajectoryBuffer, computeGAE, ppoUpdate (4 epochs, minibatch 64, clip 0.2, grad norm 0.5). Reward helpers computeTickReward/computeRoundReward. Bot wired: tick transitions collected in run loop, ppoUpdate called on onRoundEnded. Fix: add arraymancer import to PPO_Bot.nim so Tensor resolves at top level. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
12 lines
277 B
Nim
12 lines
277 B
Nim
# Package
|
|
version = "0.1.0"
|
|
author = "Davide Cappellini"
|
|
description = "PPO-trained Tank Royale bot"
|
|
license = "MIT"
|
|
bin = @["PPO_Bot"]
|
|
|
|
# Dependencies
|
|
requires "nim >= 2.0.0"
|
|
requires "tankroyale_botapi >= 1.0.0"
|
|
requires "arraymancer >= 0.7.0"
|