Evaluates A2C, PPO, TD3, SAC, DDPG against the constraints: short
on-policy episodes, no RL library, few-hundred-ms training window.
PPO wins on implementation simplicity and stability at this scale.
Closes#3
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>