RL algorithm choice #3

Closed
opened 2026-08-15 22:05:13 +02:00 by SirStone · 1 comment
Owner

Question

Which actor-critic RL algorithm best fits a 1v1 Tank Royale bot with continuous 4D actions, training between rounds (~30-180 ticks per episode), pure Nim implementation?

Candidates: A2C, PPO, TD3, SAC, DDPG. Evaluate on:

  • Continuous action space support (required)
  • Sample efficiency (limited episodes per match)
  • Implementation complexity in Nim (no RL library exists)
  • Stability with small replay buffers
  • Compute cost fitting between-round training windows

Recommend one algorithm with justification.

## Question Which actor-critic RL algorithm best fits a 1v1 Tank Royale bot with continuous 4D actions, training between rounds (~30-180 ticks per episode), pure Nim implementation? Candidates: A2C, PPO, TD3, SAC, DDPG. Evaluate on: - Continuous action space support (required) - Sample efficiency (limited episodes per match) - Implementation complexity in Nim (no RL library exists) - Stability with small replay buffers - Compute cost fitting between-round training windows Recommend one algorithm with justification.
SirStone added the wayfinder:research label 2026-08-15 22:05:13 +02:00
Author
Owner

Resolution: PPO (Proximal Policy Optimization)

Recommended algorithm: PPO

Why PPO wins for this setup

  • On-policy data is an advantage here, not a weakness. With only 10–35 rounds per match and short episodes (30–180 ticks), a replay buffer would accumulate stale transitions from earlier rounds when both the bot and the opponent were at different skill levels. On-policy trajectories are always fresh.
  • Simplest correct from-scratch implementation. PPO needs one actor net, one critic net, GAE advantage estimation, and a clipped surrogate loss — roughly 200–300 lines of Nim math. SAC and TD3 both require two Q-networks, soft-updated target networks, and (SAC) a learned temperature — significantly more bookkeeping.
  • Clipping prevents catastrophic collapse. With so few rounds to recover from a bad update, PPO's clipping term is a meaningful safeguard.
  • Matches the training window naturally. Collect one round's trajectory → one mini-batch update pass → done. No buffer warm-up, no update frequency scheduling.

Key trade-offs

  • SAC has better asymptotic sample efficiency but that advantage appears at millions of steps, not thousands. At ~3500 ticks/match the gap is marginal.
  • TD3 is more stable than DDPG but still requires a replay buffer and target networks with no proven gain at this scale.
  • A2C is simpler than PPO but lacks the clipping stabilization — strictly worse here.

Research file

Full candidate comparison table and implementation notes: research/rl-algorithm-choice.md

## Resolution: PPO (Proximal Policy Optimization) **Recommended algorithm: PPO** ### Why PPO wins for this setup - **On-policy data is an advantage here, not a weakness.** With only 10–35 rounds per match and short episodes (30–180 ticks), a replay buffer would accumulate stale transitions from earlier rounds when both the bot and the opponent were at different skill levels. On-policy trajectories are always fresh. - **Simplest correct from-scratch implementation.** PPO needs one actor net, one critic net, GAE advantage estimation, and a clipped surrogate loss — roughly 200–300 lines of Nim math. SAC and TD3 both require two Q-networks, soft-updated target networks, and (SAC) a learned temperature — significantly more bookkeeping. - **Clipping prevents catastrophic collapse.** With so few rounds to recover from a bad update, PPO's clipping term is a meaningful safeguard. - **Matches the training window naturally.** Collect one round's trajectory → one mini-batch update pass → done. No buffer warm-up, no update frequency scheduling. ### Key trade-offs - SAC has better asymptotic sample efficiency but that advantage appears at millions of steps, not thousands. At ~3500 ticks/match the gap is marginal. - TD3 is more stable than DDPG but still requires a replay buffer and target networks with no proven gain at this scale. - A2C is simpler than PPO but lacks the clipping stabilization — strictly worse here. ### Research file Full candidate comparison table and implementation notes: [`research/rl-algorithm-choice.md`](../blob/research/rl-algorithm-choice/research/rl-algorithm-choice.md)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#3