From eae6fc15a283688ffdaf6bfc66db376e3370692e Mon Sep 17 00:00:00 2001 From: Davide Cappellini Date: Sun, 23 Aug 2026 23:32:10 +0200 Subject: [PATCH] research(Evo_Bot): GA/ES parameter recommendations for ~750-weight neuroevolution (#65) Extracted concrete numbers from 13 papers in docs/papers/neuroevolution/. Key findings: use CMA-ES or mutation-only truncation GA, mutate ALL weights (not 5%), sigma=0.005-0.01, pop=64-200, no crossover, single elite. Co-Authored-By: Claude Opus 4.6 --- docs/research/ga-parameters-neuroevolution.md | 173 ++++++++++++++++++ 1 file changed, 173 insertions(+) create mode 100644 docs/research/ga-parameters-neuroevolution.md diff --git a/docs/research/ga-parameters-neuroevolution.md b/docs/research/ga-parameters-neuroevolution.md new file mode 100644 index 0000000..aa976c7 --- /dev/null +++ b/docs/research/ga-parameters-neuroevolution.md @@ -0,0 +1,173 @@ +# GA/ES Parameters for ~750-Weight Neuroevolution + +Research for issue #65. Concrete parameter recommendations extracted from 13 papers in `docs/papers/neuroevolution/`. + +## Context + +Evo_Bot's TOPO_Gun: fixed-topology feedforward ANN, 745-1489 weights depending on layer sizes. Task is predicting enemy dodge behavior from a 30-tick sliding window, outputting a guess factor. Online evolution during Robocode matches. + +--- + +## 1. Population Size + +**Recommendation: 64-256, not 300.** + +| Source | Network size | Population | Notes | +|--------|-------------|------------|-------| +| Uber Deep GA (Such et al. 2017) | 4M params (Atari), 167k (Humanoid) | 1,000 | Massive networks, distributed; overkill for ~750 weights | +| OpenAI ES (Salimans et al. 2017) | 1.7M params | 720-1,440 workers | NES-style, not a population GA | +| Canonical ES (Chrabaszcz et al. 2018) | 1.7M params | 798 (lambda), mu=50 | mu=50 selected as best across games | +| World Models CMA-ES (Ha & Schmidhuber 2018) | 867-1,088 params (controller) | 64 | CMA-ES with 16 evals per individual | +| Challenges paper (Muller & Glasmachers 2018) | 1,352-2,349 weights | default CMA-ES lambda | LM-MA-ES for ~1k-2k weights | +| Evolving Generalists (Triebold & Yaman 2023) | 246-728 weights | xNES default: 4+floor(3*ln(d)) | For d=745 -> ~24; for d=1489 -> ~26 | +| NRA (Le Clei & Bellec 2022) | 3-322 params (dynamic) | 8-512 | 256+elitism best for ~300-param tasks; 16 enough for <100 params | +| Playing Atari 6 Neurons (Cuccu et al. 2018) | ~3k connections, 6-18 neurons | xNES default | Small networks, 100 generations sufficient | + +**Key finding:** For ~750 weights, CMA-ES default lambda = 4+floor(3*ln(745)) = ~24 is a starting floor. The World Models paper (867-1,088 params, closest to our size) used 64 with CMA-ES and solved CarRacing. NRA used 256+elitism for ~300-param Pendulum networks. For a simple truncation-selection GA (not CMA-ES), 100-200 is a reasonable population; 300 is slightly wasteful but not harmful if evaluation is cheap. + +**Verdict: Start with 100-200 for a truncation GA. If using CMA-ES or xNES, use their defaults (~24-26). 300 is too large for the weight count but acceptable if per-evaluation cost is low (Robocode rounds are fast).** + +--- + +## 2. Mutation Rate and Distribution + +**Recommendation: Additive Gaussian on ALL weights, sigma=0.002-0.02, not 5% at sigma=0.1.** + +| Source | Mutation scheme | Notes | +|--------|----------------|-------| +| Uber Deep GA (Such et al. 2017) | theta' = theta + sigma * epsilon, epsilon ~ N(0,I). Sigma determined empirically per task. | Mutates ALL weights every generation, not a fraction. No "mutation rate" — every weight gets noise. | +| OpenAI ES (Salimans et al. 2017) | sigma = fixed hyperparameter (not adapted). Perturbation on full parameter vector. | Full-vector Gaussian perturbation, sigma tuned. | +| Canonical ES (Chrabaszcz et al. 2018) | N(0, sigma^2) added to all params. Network init from N(0, 0.05). | sigma is the step-size, adapted or fixed. | +| World Models (Ha & Schmidhuber 2018) | CMA-ES adapts sigma and full covariance matrix. | Self-adapting sigma — no manual sigma needed. | +| Challenges paper (Muller & Glasmachers 2018) | CSA (cumulative step-size adaptation) essential. Fixed sigma converges as slowly as random search. | Step-size adaptation is critical; fixed sigma is a known failure mode. | +| NRA (Le Clei & Bellec 2022) | N(0, 0.01) perturbation to all weights and biases. Top 50% selection. | sigma=0.01 for small networks. | + +**Key finding:** No paper uses a "5% mutation rate" (mutating only 5% of weights). ALL papers mutate ALL weights simultaneously with small additive Gaussian noise. The "mutation rate" concept from traditional GAs (flip probability per gene) does not apply to real-valued neuroevolution. Instead, the noise magnitude (sigma) controls exploration. + +For ~750 weights: +- NRA uses sigma=0.01 for networks up to ~300 params +- Uber GA uses sigma empirically per task (typical range 0.002-0.02 for Atari) +- CMA-ES/xNES adapt sigma automatically + +**Verdict: Mutate ALL weights every generation. Use sigma=0.005-0.01 as starting point. If using CMA-ES, sigma self-adapts. The "5% of weights mutated" approach is non-standard and likely harmful — it under-explores the search space.** + +--- + +## 3. Selection Pressure + +**Recommendation: Top 10-50% (truncation) or top mu out of lambda. 20% is reasonable.** + +| Source | Selection | Notes | +|--------|-----------|-------| +| Uber Deep GA (Such et al. 2017) | Truncation selection, top T individuals become parents. T not specified as percentage — varies. | Parents chosen uniformly at random from top T. | +| Canonical ES (Chrabaszcz et al. 2018) | Top mu=50 out of lambda=798 (~6%). Weighted mean of top mu. | mu=50 found optimal across games; tested mu in {10,20,50,100,200,400}. | +| NRA (Le Clei & Bellec 2022) | Top 50% duplicated, bottom 50% replaced. | Simple and effective for small populations. | +| Evolving Generalists (Triebold & Yaman 2023) | xNES default selection. | NES uses weighted rank-based update. | + +**Key finding:** Selection pressure varies widely. Canonical ES uses ~6% (mu=50 out of 798). NRA uses 50%. Standard CMA-ES uses mu = lambda/2 (50%). Top 20% is in the middle range and is fine for a truncation GA. + +**Verdict: 20% is reasonable. For small populations (64-100), 50% (top half) may work better. For larger populations (200+), stricter selection (10-20%) is appropriate. The Canonical ES result suggests mu=50 works well regardless of lambda for Atari-scale problems.** + +--- + +## 4. Crossover + +**Recommendation: No crossover. Mutation-only.** + +| Source | Crossover? | Notes | +|--------|-----------|-------| +| Uber Deep GA (Such et al. 2017) | **No crossover.** "Historically, GAs often involve crossover, but for simplicity we did not include it." | Explicitly dropped crossover for DNN weights. | +| OpenAI ES (Salimans et al. 2017) | No crossover. | ES-style: mean update, not recombination of individuals. | +| NRA (Le Clei & Bellec 2022) | **No crossover.** "stripping down many mechanisms popular in traditional evolutionary methods, like agent crossover and speciation" | Crossover explicitly excluded. | +| CMA-ES/xNES | No crossover in the traditional sense. | Weighted recombination of top individuals into distribution mean — not pairwise crossover. | +| NEAT (Stanley & Miikkulainen 2011) | Has crossover via innovation numbers. | But NEAT is for topology evolution, not fixed-topology weight-only GA. | + +**Key finding:** Every modern neuroevolution paper that works with fixed-topology networks drops crossover. Fogel & Stayton (1994, cited by Such et al.) showed crossover is often ineffective for simulated evolutionary optimization. For ANN weight vectors, crossover tends to be destructive because individual weights are not independent genes — they form functional units (layers, pathways) where mixing two different solutions creates non-functional hybrids. + +**Verdict: No crossover. Mutation-only. Uniform crossover of ANN weights is harmful — it breaks co-adapted weight configurations. If recombination is desired, use CMA-ES/xNES-style weighted mean of top solutions, which is mathematically sound.** + +--- + +## 5. Elitism + +**Recommendation: Yes, keep top 1 unchanged (single elite).** + +| Source | Elitism? | Notes | +|--------|---------|-------| +| Uber Deep GA (Such et al. 2017) | **Yes, 1 elite.** "The Nth individual is an unmodified copy of the best individual from the previous generation." Additionally, top 10 re-evaluated 30 times to find the true elite. | Single elite with robust re-evaluation. | +| NRA (Le Clei & Bellec 2022) | **Yes, elitism tested and beneficial.** Population sizes labeled "(elite)" consistently outperform non-elite variants in all figures. | Elitism was the single most impactful improvement for small populations. | +| CMA-ES | Elitist variants exist (mu+lambda). Standard CMA-ES is (mu,lambda) — non-elitist. | Non-elitist CMA-ES relies on distribution adaptation, not individual survival. | + +**Key finding:** For simple truncation GAs, elitism (keeping top 1) prevents regression and is universally recommended. The NRA paper shows that adding elitism to even a population of 16 dramatically improves results. Uber's Deep GA uses elitism with robust re-evaluation (30 episodes to confirm the elite). + +**Verdict: Keep top 1 elite unchanged. In noisy evaluation environments (Robocode), re-evaluate the top few candidates multiple times to find the true elite, following Uber's approach.** + +--- + +## 6. Generations to Convergence + +**Recommendation: 100-1,500 generations for ~750 weights, depending on the algorithm.** + +| Source | Network size | Generations | Notes | +|--------|-------------|-------------|-------| +| Uber Deep GA (Such et al. 2017) | 4M params | 348-1,834 gens (at 1k pop) | Many games: best-in-run found in 1-29 gens | +| World Models CMA-ES (Ha & Schmidhuber 2018) | 867 params | ~1,800 gens | CMA-ES, pop=64, solved CarRacing | +| NRA (Le Clei & Bellec 2022) | 3-322 params (dynamic) | 100-5,000 gens | Simple tasks: <300 gens. Complex (Ant/Humanoid): 5,000+ | +| Playing Atari 6 Neurons (Cuccu et al. 2018) | ~3k connections | 100 gens | Extremely tight budget, still achieved competitive results | +| Challenges paper (Muller & Glasmachers 2018) | 1,352-2,349 weights | 100k-300k evals | LM-MA-ES, ~150k evals for bipedal walker convergence | +| Evolving Generalists (Triebold & Yaman 2023) | 728 weights (Ant) | 5,000 gens max | xNES, some tasks solved in <100 gens | + +**Key finding:** For ~750 weights with a simple truncation GA (pop=100), expect 200-500 generations for a well-tuned sigma. CMA-ES/xNES may converge faster in generations but each generation is more expensive. The Challenges paper warns that halving the distance to the optimum requires O(d) samples, so for d=750, expect ~750 evaluations per halving step. + +For Evo_Bot's online evolution during matches: each Robocode round can evaluate one individual. With 35-round matches (typical), ~5 generations of pop=7 per match, or ~2 generations of pop=15. Convergence within a single match is unlikely; evolution must persist across matches via weight persistence. + +**Verdict: Budget 500-2,000 generations. With pop=100, that's 50k-200k evaluations. Online evolution will need many matches to converge — weight persistence is essential.** + +--- + +## 7. CMA-ES vs Simple GA vs Tournament Selection + +**Recommendation: CMA-ES or xNES for ~750 weights. Simple GA as a simpler fallback.** + +| Algorithm | Sweet spot | Pros | Cons | Source | +|-----------|-----------|------|------|--------| +| **CMA-ES** | d <= 1,000 (ideal), up to ~2,000 (practical) | Self-adapts sigma and covariance, best convergence rate, handles ill-conditioned landscapes | O(d^2) memory/time per generation, needs O(d^2) evals for full covariance learning | Muller & Glasmachers 2018, Ha & Schmidhuber 2018 | +| **LM-MA-ES** | d = 1,000-10,000 | O(d) complexity, adapts fastest-evolving subspace, strong on ~2k weights | More complex to implement | Muller & Glasmachers 2018 | +| **xNES** | d <= 1,000 | Natural gradient, self-adapting, elegant | Similar scaling limits to CMA-ES | Cuccu et al. 2018, Triebold & Yaman 2023 | +| **Simple truncation GA** | Any d | Dead simple, trivially parallel, no internal state beyond population | Needs manual sigma tuning, no adaptation, converges slowly | Such et al. 2017, Le Clei & Bellec 2022 | +| **Canonical (mu,lambda)-ES** | Any d | Step-size adaptation via CSA, simple | mu tuning matters; mu=50 worked well in Chrabaszcz 2018 | Chrabaszcz et al. 2018 | +| **Tournament selection** | Traditional GA context | Tunable selection pressure | No advantage over truncation for ANN weights | Not specifically tested in any of the 13 papers | + +**Key finding at d=750:** CMA-ES is in its sweet spot. The World Models paper (Ha & Schmidhuber 2018) used CMA-ES with pop=64 on 867-1,088 params and solved complex control tasks. The Evolving Generalists paper (Triebold & Yaman 2023) used xNES on 728 weights (Ant controller) with default population sizes. The Challenges paper (Muller & Glasmachers 2018) explicitly shows CMA-ES and LM-MA-ES outperforming simple ES on problems with 769-2,738 weights. + +However, CMA-ES requires O(d^2) = O(560k) memory for the covariance matrix at d=750. This is manageable but not trivial for an online Robocode bot. A simpler option is a (mu,lambda)-ES with CSA for step-size adaptation. + +**Verdict: CMA-ES or xNES is the best fit for 750 weights. If implementation complexity is a concern, a truncation GA with adaptive sigma (or even fixed sigma=0.005) is the pragmatic choice. Tournament selection offers no advantage.** + +--- + +## Summary: Recommended Parameters for Evo_Bot TOPO_Gun + +| Parameter | Current assumption | Recommendation | Rationale | +|-----------|-------------------|----------------|-----------| +| Population | 300 | 64-200 | 300 is oversized for ~750 weights; 64 (CMA-ES) to 200 (truncation GA) | +| Mutation | 5% of weights, gaussian sigma=0.1 | ALL weights, sigma=0.005-0.01 | Every paper mutates all weights. sigma=0.1 is too large. | +| Selection | Top 20% | Top 20-50% | 20% is fine; 50% if pop is small | +| Crossover | Uniform | None | Uniformly dropped in all modern neuroevolution papers | +| Elitism | Not specified | Top 1, re-evaluated | Single elite prevents regression; re-evaluate to handle noise | +| Algorithm | Simple GA | CMA-ES or truncation GA + CSA | CMA-ES is in its sweet spot at d=750; simple GA works but converges slower | +| Generations | Not specified | 500-2,000 | Online evolution needs many matches for convergence | + +--- + +## Sources + +1. Such et al. 2017 — "Deep Neuroevolution: Genetic Algorithms are a Competitive Alternative" (Uber AI Labs) +2. Salimans et al. 2017 — "Evolution Strategies as a Scalable Alternative to Reinforcement Learning" (OpenAI) +3. Chrabaszcz et al. 2018 — "Back to Basics: Benchmarking Canonical Evolution Strategies for Playing Atari" +4. Muller & Glasmachers 2018 — "Challenges in High-dimensional Reinforcement Learning with Evolution Strategies" +5. Ha & Schmidhuber 2018 — "Recurrent World Models Facilitate Policy Evolution" (World Models) +6. Cuccu et al. 2018 — "Playing Atari with Six Neurons" +7. Le Clei & Bellec 2022 — "Neuroevolution of Recurrent Architectures on Control Tasks" +8. Triebold & Yaman 2023 — "Evolving Generalist Controllers to Handle a Wide Range of Morphological Variations" +9. Stanley & Miikkulainen 2011 — "Competitive Coevolution through Evolutionary Complexification" (NEAT)