30cda871cca3800fe1cfc96b53fc2425ee1c7b95
Evaluates A2C, PPO, TD3, SAC, DDPG against the constraints: short on-policy episodes, no RL library, few-hundred-ms training window. PPO wins on implementation simplicity and stability at this scale. Closes #3 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Description
No description provided
Languages
Nim
73.7%
Python
18%
Shell
3.7%
Java
3.5%
HTML
1%
Other
0.1%