Research: optimal GA parameters for ~750-weight neuroevolution #65

Closed
opened 2026-08-23 23:26:53 +02:00 by SirStone · 1 comment
Owner

Parent: #61

Question

From the downloaded papers in docs/papers/neuroevolution/, extract concrete GA parameter recommendations for networks in the 745-1489 weight range:

  • Population size (we assumed 300 — is this right?)
  • Mutation rate and distribution (we assumed 5% gaussian σ=0.1)
  • Selection pressure (we assumed top 20%)
  • Crossover type (we assumed uniform — is this best for ANN weights?)
  • Elitism (keep top N unchanged?)
  • Number of generations needed for convergence at this weight count
  • CMA-ES vs simple GA vs tournament selection — which fits best?

Summarize findings as a recommendation document.

Parent: #61 ## Question From the downloaded papers in docs/papers/neuroevolution/, extract concrete GA parameter recommendations for networks in the 745-1489 weight range: - Population size (we assumed 300 — is this right?) - Mutation rate and distribution (we assumed 5% gaussian σ=0.1) - Selection pressure (we assumed top 20%) - Crossover type (we assumed uniform — is this best for ANN weights?) - Elitism (keep top N unchanged?) - Number of generations needed for convergence at this weight count - CMA-ES vs simple GA vs tournament selection — which fits best? Summarize findings as a recommendation document.
SirStone added the wayfinder:research label 2026-08-23 23:26:53 +02:00
Author
Owner

Resolution — GA/ES parameters for ~750-weight neuroevolution

Findings extracted from all 13 papers in docs/papers/neuroevolution/. Full document: docs/research/ga-parameters-neuroevolution.md on branch research/ga-parameters.

Key corrections to our initial assumptions

Parameter We assumed Literature says Change needed?
Population 300 64-200 (CMA-ES: 64; truncation GA: 100-200) Yes — 300 is oversized for ~750 weights
Mutation 5% of weights, gaussian σ=0.1 ALL weights every generation, σ=0.005-0.01 Yes — major fix. No paper uses partial mutation. σ=0.1 is 10-20x too large.
Selection Top 20% Top 20-50% No — 20% is fine
Crossover Uniform None Yes — drop crossover entirely. Every modern neuroevolution paper drops it for fixed-topology ANN weights.
Elitism Unspecified Top 1 elite, re-evaluated for robustness Yes — add single elite

Algorithm choice for d=750

CMA-ES is in its sweet spot at 750 weights. Ha & Schmidhuber (2018) used CMA-ES with pop=64 on 867-1,088 params and solved CarRacing. Triebold & Yaman (2023) used xNES on 728 weights. Muller & Glasmachers (2018) explicitly show CMA-ES/LM-MA-ES outperforming simple ES at 769-2,738 weights.

If CMA-ES is too complex to implement in Nim, a truncation GA with all-weight mutation at σ=0.005 and single elite is the pragmatic fallback.

Convergence budget

Expect 500-2,000 generations (50k-200k evaluations). Online evolution during Robocode matches will need many matches to converge — weight persistence across matches is essential.

Sources (top 5 most relevant)

  1. Such et al. 2017 (Uber) — pop=1000, all-weight mutation, no crossover, 1 elite
  2. Ha & Schmidhuber 2018 (World Models) — CMA-ES pop=64 on 867 params
  3. Chrabaszcz et al. 2018 (Canonical ES) — mu=50 optimal, pop=798
  4. Muller & Glasmachers 2018 (Challenges) — CMA-ES sweet spot at d<1000, step-size adaptation critical
  5. Le Clei & Bellec 2022 (NRA) — σ=0.01, top 50%, elitism essential for small pops
## Resolution — GA/ES parameters for ~750-weight neuroevolution Findings extracted from all 13 papers in `docs/papers/neuroevolution/`. Full document: `docs/research/ga-parameters-neuroevolution.md` on branch `research/ga-parameters`. ### Key corrections to our initial assumptions | Parameter | We assumed | Literature says | Change needed? | |-----------|-----------|----------------|----------------| | **Population** | 300 | 64-200 (CMA-ES: 64; truncation GA: 100-200) | Yes — 300 is oversized for ~750 weights | | **Mutation** | 5% of weights, gaussian σ=0.1 | ALL weights every generation, σ=0.005-0.01 | **Yes — major fix.** No paper uses partial mutation. σ=0.1 is 10-20x too large. | | **Selection** | Top 20% | Top 20-50% | No — 20% is fine | | **Crossover** | Uniform | **None** | **Yes — drop crossover entirely.** Every modern neuroevolution paper drops it for fixed-topology ANN weights. | | **Elitism** | Unspecified | Top 1 elite, re-evaluated for robustness | Yes — add single elite | ### Algorithm choice for d=750 **CMA-ES is in its sweet spot** at 750 weights. Ha & Schmidhuber (2018) used CMA-ES with pop=64 on 867-1,088 params and solved CarRacing. Triebold & Yaman (2023) used xNES on 728 weights. Muller & Glasmachers (2018) explicitly show CMA-ES/LM-MA-ES outperforming simple ES at 769-2,738 weights. If CMA-ES is too complex to implement in Nim, a truncation GA with all-weight mutation at σ=0.005 and single elite is the pragmatic fallback. ### Convergence budget Expect 500-2,000 generations (50k-200k evaluations). Online evolution during Robocode matches will need many matches to converge — weight persistence across matches is essential. ### Sources (top 5 most relevant) 1. Such et al. 2017 (Uber) — pop=1000, all-weight mutation, no crossover, 1 elite 2. Ha & Schmidhuber 2018 (World Models) — CMA-ES pop=64 on 867 params 3. Chrabaszcz et al. 2018 (Canonical ES) — mu=50 optimal, pop=798 4. Muller & Glasmachers 2018 (Challenges) — CMA-ES sweet spot at d<1000, step-size adaptation critical 5. Le Clei & Bellec 2022 (NRA) — σ=0.01, top 50%, elitism essential for small pops
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#65