From 30cda871cca3800fe1cfc96b53fc2425ee1c7b95 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sat, 15 Aug 2026 22:07:47 +0200 Subject: [PATCH] =?UTF-8?q?research:=20RL=20algorithm=20choice=20=E2=80=94?= =?UTF-8?q?=20recommend=20PPO=20for=20Tank=20Royale=20bot?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Evaluates A2C, PPO, TD3, SAC, DDPG against the constraints: short on-policy episodes, no RL library, few-hundred-ms training window. PPO wins on implementation simplicity and stability at this scale. Closes #3 Co-Authored-By: Claude Sonnet 4.6 --- research/rl-algorithm-choice.md | 54 +++++++++++++++++++++++++++++++++ 1 file changed, 54 insertions(+) create mode 100644 research/rl-algorithm-choice.md diff --git a/research/rl-algorithm-choice.md b/research/rl-algorithm-choice.md new file mode 100644 index 0000000..c730af5 --- /dev/null +++ b/research/rl-algorithm-choice.md @@ -0,0 +1,54 @@ +# RL Algorithm Choice for Tank Royale Bot + +## Context + +- 1v1 Tank Royale bot, pure Nim (Arraymancer or hand-rolled MLP) +- Continuous 4D action space: `targetSpeed ∈ [-8,8]`, `turnRate ∈ [-10,10]`, `gunTurnRate ∈ [-20,20]`, `firePower ∈ {0} ∪ [0.1,3.0]` +- Episodes: ~30–180 ticks (short) +- Training window: between rounds, ~few hundred milliseconds +- No RL library — everything hand-rolled in Nim + +## Candidate Summary + +| Algorithm | Policy | Replay buffer | Sample efficiency | Stability | Impl. complexity | Fit | +|-----------|---------|---------------|-------------------|----------------|------------------|------| +| A2C | On-policy stochastic | No | Low | Moderate | Low | Poor | +| PPO | On-policy stochastic | No | Low-medium | Good | Medium | Poor | +| DDPG | Off-policy det. | Yes | Medium | Poor (Q overest.) | Medium | Fair | +| TD3 | Off-policy det. | Yes | Medium-high | Good | Medium-high | Good | +| SAC | Off-policy stochastic| Yes | High | Very good | High | Best | + +## Recommendation: PPO (on-policy, clipped) + +**Chosen: PPO** — not the highest sample-efficiency on paper, but the best fit for this specific setup. + +### Reasoning + +1. **No replay buffer needed.** With only 10–35 short rounds per match, a replay buffer holding stale transitions from earlier rounds could mislead learning (opponent behavior, bot state drift). On-policy data is always fresh. + +2. **Simplest correct implementation from scratch.** PPO needs: one actor network, one critic network, advantage estimation (GAE or simple TD), and a clipped surrogate loss. That is ~200–300 lines of Nim math. SAC and TD3 require two Q-networks, a target network with soft updates, and (SAC) a learned temperature parameter — doubling the bookkeeping with no proven payoff at this scale. + +3. **Short episodes suit on-policy collection.** Each round provides a complete trajectory. PPO runs one gradient update pass per trajectory batch — exactly matching the between-round training window. No need to manage buffer fill, warm-up, or update frequency scheduling. + +4. **Stability.** The clipping term prevents catastrophic policy collapse, which matters when the episode count per match is too small to recover from a bad update. + +5. **Sample efficiency concern is overstated here.** SAC's replay-buffer advantage shows at millions of environment steps. At 35 rounds × ~100 ticks = ~3500 steps per match, the gap between PPO and SAC is marginal; both learn slowly by ML standards. PPO's stability wins over SAC's asymptotic efficiency. + +### Why not the others + +- **A2C**: no clipping → less stable, no benefit over PPO, same on-policy cost. +- **DDPG**: known Q-overestimation instability; superseded by TD3. +- **TD3**: better than DDPG but replay buffer management adds complexity with no clear gain at this scale. +- **SAC**: two Q-nets + target nets + temperature tuning → highest impl. cost; entropy bonus helps exploration in long-horizon tasks, overkill for 4D actions. + +## Key Implementation Notes + +- **Actor**: outputs mean + log-std per action dimension → sample via reparameterization for training, use mean at inference. +- **Action bounds**: apply `tanh` + scale to map unbounded Gaussian samples into `[-8,8]`, `[-10,10]`, `[-20,20]`; handle `firePower` separately (threshold below 0.1 → set to 0, else clamp to `[0.1,3.0]`). +- **Critic**: single V(s) network, trained with TD(0) or GAE (λ≈0.95 recommended). +- **Clipped surrogate**: ε = 0.2 is the standard starting point. +- **Update cadence**: collect full-round trajectory → one pass of K mini-batch updates (K=4–10) → discard trajectory. +- **Exploration**: stochastic policy provides built-in exploration; no Ornstein-Uhlenbeck needed. +- **Compute**: one forward + one backward pass on a small MLP per mini-batch fits comfortably in a few hundred milliseconds. + +