From cb1bbc35dc1d2a5e3e6e461030152b783ba65946 Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Mon, 24 Aug 2026 00:02:47 +0200 Subject: [PATCH] docs(adr): Evo_Bot neuroevolution gun architecture (#63) Co-Authored-By: Claude Opus 4.6 --- ...on-gun-fixed-topology-ann-evolved-by-ga.md | 46 +++++++++++++++++++ 1 file changed, 46 insertions(+) create mode 100644 docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md diff --git a/docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md b/docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md new file mode 100644 index 0000000..84900da --- /dev/null +++ b/docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md @@ -0,0 +1,46 @@ +# Neuroevolution gun with fixed-topology ANN evolved by GA + +Evo_Bot needs a gun that adapts to each opponent's dodge patterns during a match, finds nonlinear movement patterns that histograms miss, and is original. We chose a fixed-topology feedforward ANN (91->8->1, 745 weights) whose weights are evolved by a mutation-only GA running on a parallel thread. This beats the alternatives (RL too slow to adapt in-match, Q-learning collapses to histogram for single-shot decisions, guess-factor histograms are unoriginal, transformer/LLM-style prediction is data-starved at ~14k ticks per match) while keeping implementation risk low by deferring topology evolution (NEAT) until the fixed network hits its ceiling. + +## Considered Options + +- **PPO / SAC (end-to-end RL):** Too slow -- thousands of rounds to converge, cannot adapt mid-match. Explored in other bots in this repo. +- **Q-learning for aiming:** Collapses to a histogram. Single-shot aiming has no sequential decision structure for Q-learning to exploit. +- **Guess-factor histogram:** Proven and fast to converge (~15 ticks), but unoriginal -- 20 years of community tuning. +- **Transformer / LLM-style sequence prediction:** Data-starved. ~14k ticks per match vs billions needed for attention-based models. +- **GA with crossover:** Literature uniformly shows crossover is harmful for ANN weight evolution -- it breaks co-adapted weight configurations. Every modern neuroevolution paper (Uber Deep GA, NRA, OpenAI ES) drops it. +- **CMA-ES:** Ideal at d=750 weights (self-adapts sigma and covariance). More complex to implement; upgrade path from simple GA when needed. + +## Decision + +**Architecture:** +- Evo_Bot (1v1) with modular gun interface: `feed(state)` / `aim() -> (angle, power)` +- Gun owns its evolution thread (parallel, never blocks inference) + +**TOPO_Gun (first gun implementation):** +- Network: 91->8->1 (hidden size configurable), 745 weights +- Input: 30 ticks x (lateral_vel, delta_heading, wall_distance_ahead) + current distance = 91 +- Output: guess factor (-1 to +1) +- Bullet power: deterministic distance-based formula (not learned) +- Evolution: population 300, clone loaded champion + small Gaussian mutations (ALL weights, sigma=0.005-0.01), NO crossover, single elite preserved unchanged +- Fitness: virtual bullet hits on 100 randomly sampled replay tape ticks, using real distance-based power +- Replay tape: rolling window ~2000 ticks (configurable) +- Weight persistence: per-opponent -> global fallback -> random init (load order) +- Cold start: first-ever run, don't fire until champion emerges; subsequent runs load weights, fire from tick 1 +- Push new champion weights to inference when it hits better than current + +**Deferred:** +- NEAT_Gun: deferred until TOPO_Gun hits its ceiling +- Virtual Guns: run multiple guns in parallel, fire whichever has best virtual hit rate + +**Boundaries:** +- Bot controls firing discipline (when to shoot, energy management); gun always returns best aim + +## Consequences + +- GA+ANN finds nonlinear patterns histograms miss, but needs more data (~50+ ticks vs ~15 for guess-factor histogram) before it outperforms +- Fixed topology before NEAT reduces implementation risk +- Mutation-only evolution simplifies implementation (no crossover logic) +- Parallel evolution thread reuses the pattern from SAC_LSTM_Bot (#48) +- Per-opponent weight persistence eliminates cold start after first encounter +- CMA-ES is the natural upgrade path if simple GA convergence is too slow (d=750 is CMA-ES sweet spot)