Document the locked architecture as an ADR #63

Closed
opened 2026-08-23 23:26:44 +02:00 by SirStone · 2 comments
Owner

Parent: #61

Question

Write ADR documenting the Evo_Bot architecture decision. Record:

Architecture:

  • Evo_Bot (1v1) with modular gun interface: feed(state) / aim() → (angle, power)
  • Gun owns its evolution thread (parallel, never blocks inference)
  • TOPO_Gun (first): fixed topology ANN evolved by GA
    • Network: 91→8→1 (hidden size configurable)
    • Input: 30 ticks × (lateral_vel, Δheading, wall_distance_ahead) + current distance = 91
    • Output: guess factor (-1 to +1)
    • Bullet power: deterministic distance-based formula
    • Evolution: population 300, clone loaded weights + small mutations, fitness = hits on 100 sampled replay ticks with real power, push champion when it beats current
    • Replay tape: rolling window ~2000 ticks (configurable)
    • Weight persistence: per-opponent → global → random init (load order)
    • Cold start: first-ever run don't fire until champion emerges; subsequent runs load weights, fire from tick 1
  • NEAT_Gun: deferred until TOPO_Gun hits ceiling
  • Virtual Guns: both run, fire whichever hits better
  • Bot controls firing discipline, gun always returns aim

Alternatives considered: PPO/SAC (too slow to adapt), Q-learning (collapses to histogram for single-shot decisions), guess factor histogram (proven but unoriginal), transformer/LLM-style (data-starved in Robocode).

Trade-offs: GA+ANN is original and finds nonlinear patterns but needs more data than histograms. Fixed topology before NEAT reduces implementation risk.

Parent: #61 ## Question Write ADR documenting the Evo_Bot architecture decision. Record: Architecture: - Evo_Bot (1v1) with modular gun interface: feed(state) / aim() → (angle, power) - Gun owns its evolution thread (parallel, never blocks inference) - TOPO_Gun (first): fixed topology ANN evolved by GA - Network: 91→8→1 (hidden size configurable) - Input: 30 ticks × (lateral_vel, Δheading, wall_distance_ahead) + current distance = 91 - Output: guess factor (-1 to +1) - Bullet power: deterministic distance-based formula - Evolution: population 300, clone loaded weights + small mutations, fitness = hits on 100 sampled replay ticks with real power, push champion when it beats current - Replay tape: rolling window ~2000 ticks (configurable) - Weight persistence: per-opponent → global → random init (load order) - Cold start: first-ever run don't fire until champion emerges; subsequent runs load weights, fire from tick 1 - NEAT_Gun: deferred until TOPO_Gun hits ceiling - Virtual Guns: both run, fire whichever hits better - Bot controls firing discipline, gun always returns aim Alternatives considered: PPO/SAC (too slow to adapt), Q-learning (collapses to histogram for single-shot decisions), guess factor histogram (proven but unoriginal), transformer/LLM-style (data-starved in Robocode). Trade-offs: GA+ANN is original and finds nonlinear patterns but needs more data than histograms. Fixed topology before NEAT reduces implementation risk.
SirStone added the wayfinder:task label 2026-08-23 23:26:44 +02:00
Author
Owner

Blocked by #62 — ADR should use the settled CONTEXT.md vocabulary.

Blocked by #62 — ADR should use the settled CONTEXT.md vocabulary.
Author
Owner

ADR written: docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md on branch research/goto-controller (commit cb1bbc3).

Records the locked Evo_Bot architecture: fixed-topology ANN (91->8->1, 745 weights) evolved by mutation-only GA (all weights, sigma=0.005-0.01, no crossover, single elite, pop 300). Includes considered alternatives (PPO/SAC, Q-learning, GF histogram, transformer, CMA-ES, crossover) and consequences. Uses corrected GA parameters from #65 research.

ADR written: `docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md` on branch `research/goto-controller` (commit cb1bbc3). Records the locked Evo_Bot architecture: fixed-topology ANN (91->8->1, 745 weights) evolved by mutation-only GA (all weights, sigma=0.005-0.01, no crossover, single elite, pop 300). Includes considered alternatives (PPO/SAC, Q-learning, GF histogram, transformer, CMA-ES, crossover) and consequences. Uses corrected GA parameters from #65 research.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#63