chore: rename CLAUDE.md → AGENTS.md, move research doc to docs/

This commit is contained in:
2026-08-27 18:11:56 +02:00
parent f5b1ade48c
commit 8ba4bae21a
3 changed files with 65 additions and 13 deletions
+54
View File
@@ -0,0 +1,54 @@
# RL Algorithm Choice for Tank Royale Bot
## Context
- 1v1 Tank Royale bot, pure Nim (Arraymancer or hand-rolled MLP)
- Continuous 4D action space: `targetSpeed ∈ [-8,8]`, `turnRate ∈ [-10,10]`, `gunTurnRate ∈ [-20,20]`, `firePower ∈ {0} ∪ [0.1,3.0]`
- Episodes: ~30–180 ticks (short)
- Training window: between rounds, ~few hundred milliseconds
- No RL library — everything hand-rolled in Nim
## Candidate Summary
| Algorithm | Policy | Replay buffer | Sample efficiency | Stability | Impl. complexity | Fit |
|-----------|---------|---------------|-------------------|----------------|------------------|------|
| A2C | On-policy stochastic | No | Low | Moderate | Low | Poor |
| PPO | On-policy stochastic | No | Low-medium | Good | Medium | Poor |
| DDPG | Off-policy det. | Yes | Medium | Poor (Q overest.) | Medium | Fair |
| TD3 | Off-policy det. | Yes | Medium-high | Good | Medium-high | Good |
| SAC | Off-policy stochastic| Yes | High | Very good | High | Best |
## Recommendation: PPO (on-policy, clipped)
**Chosen: PPO** — not the highest sample-efficiency on paper, but the best fit for this specific setup.
### Reasoning
1. **No replay buffer needed.** With only 10–35 short rounds per match, a replay buffer holding stale transitions from earlier rounds could mislead learning (opponent behavior, bot state drift). On-policy data is always fresh.
2. **Simplest correct implementation from scratch.** PPO needs: one actor network, one critic network, advantage estimation (GAE or simple TD), and a clipped surrogate loss. That is ~200–300 lines of Nim math. SAC and TD3 require two Q-networks, a target network with soft updates, and (SAC) a learned temperature parameter — doubling the bookkeeping with no proven payoff at this scale.
3. **Short episodes suit on-policy collection.** Each round provides a complete trajectory. PPO runs one gradient update pass per trajectory batch — exactly matching the between-round training window. No need to manage buffer fill, warm-up, or update frequency scheduling.
4. **Stability.** The clipping term prevents catastrophic policy collapse, which matters when the episode count per match is too small to recover from a bad update.
5. **Sample efficiency concern is overstated here.** SAC's replay-buffer advantage shows at millions of environment steps. At 35 rounds × ~100 ticks = ~3500 steps per match, the gap between PPO and SAC is marginal; both learn slowly by ML standards. PPO's stability wins over SAC's asymptotic efficiency.
### Why not the others
- **A2C**: no clipping → less stable, no benefit over PPO, same on-policy cost.
- **DDPG**: known Q-overestimation instability; superseded by TD3.
- **TD3**: better than DDPG but replay buffer management adds complexity with no clear gain at this scale.
- **SAC**: two Q-nets + target nets + temperature tuning → highest impl. cost; entropy bonus helps exploration in long-horizon tasks, overkill for 4D actions.
## Key Implementation Notes
- **Actor**: outputs mean + log-std per action dimension → sample via reparameterization for training, use mean at inference.
- **Action bounds**: apply `tanh` + scale to map unbounded Gaussian samples into `[-8,8]`, `[-10,10]`, `[-20,20]`; handle `firePower` separately (threshold below 0.1 → set to 0, else clamp to `[0.1,3.0]`).
- **Critic**: single V(s) network, trained with TD(0) or GAE (λ≈0.95 recommended).
- **Clipped surrogate**: ε = 0.2 is the standard starting point.
- **Update cadence**: collect full-round trajectory → one pass of K mini-batch updates (K=4–10) → discard trajectory.
- **Exploration**: stochastic policy provides built-in exploration; no Ornstein-Uhlenbeck needed.
- **Compute**: one forward + one backward pass on a small MLP per mini-batch fits comfortably in a few hundred milliseconds.
<!-- ponytail: on-policy data discard means no cross-round transfer learning; add a small replay buffer seeded with high-reward transitions if sample efficiency becomes the bottleneck after real match testing -->