fedab54bc0876c97ef1c1f2a7bb2e432d1e56509
- Accumulate transitions across 10 rounds (~3000) before PPO update (was per-round ~300 — gradient estimates were far too noisy) - training.nim: MAX_TRANSITIONS 4096→8192, done flag on transitions, GAE handles episode boundaries correctly - PPO_Bot.nim: buffer persists across rounds, update every N rounds - training.env: lr 5e-5→1e-4, entropy 0.001, UPDATE_INTERVAL=10
Description
No description provided
Languages
Nim
73.7%
Python
18%
Shell
3.7%
Java
3.5%
HTML
1%
Other
0.1%