Block a user
Learning rate and optimizer hyperparameters
Exploration noise
Resolution
No explicit exploration noise. PPO's stochastic Gaussian policy (learned log-std per action dimension, decided in #9) provides exploration naturally. Entropy bonus in the PPO…
Network architecture
Resolution
Structure: Separate actor and critic networks (no shared trunk). ~8K parameters total + 5 log-std floats.
Actor:
- Input: 42 floats (state vector from #4)
- Hidden: 64 →…
Bot scaffold and build setup
Training loop design
Resolution
Episode buffer: Seq of (state: Tensor, action: Tensor, log_prob: float32, reward: float32, value: float32) tuples. Append during round, clear after training. Single episode,…
Reward function design
Resolution
Per-tick reward:
reward_tick = my_energy_delta - enemy_energy_delta
- Raw values, no scaling or normalization
- No extra wall penalty (already captured in energy delta)
- No…
State vector design
Resolution
State vector: 42 floats, all normalized to ~[-1, 1]
Current tick (14 floats):
Enemy tracker design
Resolution
Radar strategy (1v1)
- Decoupled radar — tracker fully owns
setRadarTurnRate, compensating for body/gun rotation. - Round start: full-speed sweep (45°/tick) until…
Arraymancer viability spike
Arraymancer viability spike
Resolution: Arraymancer 0.7.33 compiles and trains on Nim 2.2.4 + NixOS. XOR MLP trains in 0.10s, all predictions correct. NixOS requires LD_LIBRARY_PATH for OpenBLAS (standard flake.nix fix).…