# RL Algorithm Choice for Tank Royale Bot ## Context - 1v1 Tank Royale bot, pure Nim (Arraymancer or hand-rolled MLP) - Continuous 4D action space: `targetSpeed ∈ [-8,8]`, `turnRate ∈ [-10,10]`, `gunTurnRate ∈ [-20,20]`, `firePower ∈ {0} ∪ [0.1,3.0]` - Episodes: ~30–180 ticks (short) - Training window: between rounds, ~few hundred milliseconds - No RL library — everything hand-rolled in Nim ## Candidate Summary | Algorithm | Policy | Replay buffer | Sample efficiency | Stability | Impl. complexity | Fit | |-----------|---------|---------------|-------------------|----------------|------------------|------| | A2C | On-policy stochastic | No | Low | Moderate | Low | Poor | | PPO | On-policy stochastic | No | Low-medium | Good | Medium | Poor | | DDPG | Off-policy det. | Yes | Medium | Poor (Q overest.) | Medium | Fair | | TD3 | Off-policy det. | Yes | Medium-high | Good | Medium-high | Good | | SAC | Off-policy stochastic| Yes | High | Very good | High | Best | ## Recommendation: PPO (on-policy, clipped) **Chosen: PPO** — not the highest sample-efficiency on paper, but the best fit for this specific setup. ### Reasoning 1. **No replay buffer needed.** With only 10–35 short rounds per match, a replay buffer holding stale transitions from earlier rounds could mislead learning (opponent behavior, bot state drift). On-policy data is always fresh. 2. **Simplest correct implementation from scratch.** PPO needs: one actor network, one critic network, advantage estimation (GAE or simple TD), and a clipped surrogate loss. That is ~200–300 lines of Nim math. SAC and TD3 require two Q-networks, a target network with soft updates, and (SAC) a learned temperature parameter — doubling the bookkeeping with no proven payoff at this scale. 3. **Short episodes suit on-policy collection.** Each round provides a complete trajectory. PPO runs one gradient update pass per trajectory batch — exactly matching the between-round training window. No need to manage buffer fill, warm-up, or update frequency scheduling. 4. **Stability.** The clipping term prevents catastrophic policy collapse, which matters when the episode count per match is too small to recover from a bad update. 5. **Sample efficiency concern is overstated here.** SAC's replay-buffer advantage shows at millions of environment steps. At 35 rounds × ~100 ticks = ~3500 steps per match, the gap between PPO and SAC is marginal; both learn slowly by ML standards. PPO's stability wins over SAC's asymptotic efficiency. ### Why not the others - **A2C**: no clipping → less stable, no benefit over PPO, same on-policy cost. - **DDPG**: known Q-overestimation instability; superseded by TD3. - **TD3**: better than DDPG but replay buffer management adds complexity with no clear gain at this scale. - **SAC**: two Q-nets + target nets + temperature tuning → highest impl. cost; entropy bonus helps exploration in long-horizon tasks, overkill for 4D actions. ## Key Implementation Notes - **Actor**: outputs mean + log-std per action dimension → sample via reparameterization for training, use mean at inference. - **Action bounds**: apply `tanh` + scale to map unbounded Gaussian samples into `[-8,8]`, `[-10,10]`, `[-20,20]`; handle `firePower` separately (threshold below 0.1 → set to 0, else clamp to `[0.1,3.0]`). - **Critic**: single V(s) network, trained with TD(0) or GAE (λ≈0.95 recommended). - **Clipped surrogate**: ε = 0.2 is the standard starting point. - **Update cadence**: collect full-round trajectory → one pass of K mini-batch updates (K=4–10) → discard trajectory. - **Exploration**: stochastic policy provides built-in exploration; no Ornstein-Uhlenbeck needed. - **Compute**: one forward + one backward pass on a small MLP per mini-batch fits comfortably in a few hundred milliseconds.