Evaluates A2C, PPO, TD3, SAC, DDPG against the constraints: short on-policy episodes, no RL library, few-hundred-ms training window. PPO wins on implementation simplicity and stability at this scale. Closes #3 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
4.1 KiB
RL Algorithm Choice for Tank Royale Bot
Context
- 1v1 Tank Royale bot, pure Nim (Arraymancer or hand-rolled MLP)
- Continuous 4D action space:
targetSpeed ∈ [-8,8],turnRate ∈ [-10,10],gunTurnRate ∈ [-20,20],firePower ∈ {0} ∪ [0.1,3.0] - Episodes: ~30–180 ticks (short)
- Training window: between rounds, ~few hundred milliseconds
- No RL library — everything hand-rolled in Nim
Candidate Summary
| Algorithm | Policy | Replay buffer | Sample efficiency | Stability | Impl. complexity | Fit |
|---|---|---|---|---|---|---|
| A2C | On-policy stochastic | No | Low | Moderate | Low | Poor |
| PPO | On-policy stochastic | No | Low-medium | Good | Medium | Poor |
| DDPG | Off-policy det. | Yes | Medium | Poor (Q overest.) | Medium | Fair |
| TD3 | Off-policy det. | Yes | Medium-high | Good | Medium-high | Good |
| SAC | Off-policy stochastic | Yes | High | Very good | High | Best |
Recommendation: PPO (on-policy, clipped)
Chosen: PPO — not the highest sample-efficiency on paper, but the best fit for this specific setup.
Reasoning
-
No replay buffer needed. With only 10–35 short rounds per match, a replay buffer holding stale transitions from earlier rounds could mislead learning (opponent behavior, bot state drift). On-policy data is always fresh.
-
Simplest correct implementation from scratch. PPO needs: one actor network, one critic network, advantage estimation (GAE or simple TD), and a clipped surrogate loss. That is ~200–300 lines of Nim math. SAC and TD3 require two Q-networks, a target network with soft updates, and (SAC) a learned temperature parameter — doubling the bookkeeping with no proven payoff at this scale.
-
Short episodes suit on-policy collection. Each round provides a complete trajectory. PPO runs one gradient update pass per trajectory batch — exactly matching the between-round training window. No need to manage buffer fill, warm-up, or update frequency scheduling.
-
Stability. The clipping term prevents catastrophic policy collapse, which matters when the episode count per match is too small to recover from a bad update.
-
Sample efficiency concern is overstated here. SAC's replay-buffer advantage shows at millions of environment steps. At 35 rounds × ~100 ticks = ~3500 steps per match, the gap between PPO and SAC is marginal; both learn slowly by ML standards. PPO's stability wins over SAC's asymptotic efficiency.
Why not the others
- A2C: no clipping → less stable, no benefit over PPO, same on-policy cost.
- DDPG: known Q-overestimation instability; superseded by TD3.
- TD3: better than DDPG but replay buffer management adds complexity with no clear gain at this scale.
- SAC: two Q-nets + target nets + temperature tuning → highest impl. cost; entropy bonus helps exploration in long-horizon tasks, overkill for 4D actions.
Key Implementation Notes
- Actor: outputs mean + log-std per action dimension → sample via reparameterization for training, use mean at inference.
- Action bounds: apply
tanh+ scale to map unbounded Gaussian samples into[-8,8],[-10,10],[-20,20]; handlefirePowerseparately (threshold below 0.1 → set to 0, else clamp to[0.1,3.0]). - Critic: single V(s) network, trained with TD(0) or GAE (λ≈0.95 recommended).
- Clipped surrogate: ε = 0.2 is the standard starting point.
- Update cadence: collect full-round trajectory → one pass of K mini-batch updates (K=4–10) → discard trajectory.
- Exploration: stochastic policy provides built-in exploration; no Ornstein-Uhlenbeck needed.
- Compute: one forward + one backward pass on a small MLP per mini-batch fits comfortably in a few hundred milliseconds.