Spec: Evo_Bot TOPO_Gun implementation #80

Closed
opened 2026-08-25 19:05:52 +02:00 by SirStone · 0 comments
Owner

Problem Statement

Evo_Bot needs a working gun system that predicts enemy dodge behavior using a neuroevolution approach. The architecture is locked (ADR 0001, wayfinder maps #61 and #74), the GA+ANN pipeline is validated (spike in prototypes/ga_gun_spike/), but no bot code exists yet — there is no src/ directory. The bot needs to be built from scratch, integrating all decided subsystems into a Robocode Tank Royale bot written in Nim.

Solution

Build Evo_Bot as a modular Robocode Tank Royale bot in Nim with a TOPO_Gun (fixed-topology ANN evolved by mutation-only GA), Virtual Guns system, confidence-based firing discipline, and per-opponent weight persistence. The bot predicts enemy dodge behavior via guess factor targeting and adapts mid-match through parallel evolution.

User Stories

  1. As a bot developer, I want a feedforward ANN module (91→8→1, tanh activation) with a pure forward(weights, inputs) → float interface, so that the gun can predict guess factors from the sliding window.
  2. As a bot developer, I want per-feature fixed normalization using game constants (lateral_vel / 8.0, heading_delta / π, wall_distance / 1000.0), so that ANN inputs are bounded without runtime statistics.
  3. As a bot developer, I want a mutation-only GA module that evolves a population of weight vectors, so that the ANN adapts to each opponent mid-match.
  4. As a bot developer, I want the GA to use hybrid population seeding (~25% mutated clones of loaded champion, ~75% random init), so that evolution exploits known-good weights while exploring new strategies.
  5. As a bot developer, I want GA parameters matching the research findings: population 64–200, Gaussian mutation on ALL weights (σ = 0.005–0.01), top 20–50% truncation selection, single elite with re-evaluation, no crossover.
  6. As a bot developer, I want the GA evolution to run on a parallel thread that never blocks the bot's tick loop, so that inference and evolution are decoupled.
  7. As a bot developer, I want a sliding window that records the last 30 ticks of enemy state (lateral_vel, heading_delta, wall_distance) as a flat 90-element buffer, plus current distance as the 91st input.
  8. As a bot developer, I want a replay tape (rolling ~2000-tick buffer) of recorded enemy states, so that the GA can evaluate fitness offline against historical data.
  9. As a bot developer, I want fitness evaluated as virtual bullet hit count on randomly sampled replay tape ticks, so that the GA selects for aiming accuracy.
  10. As a bot developer, I want a bullet power formula power = clamp(3.0 - (distance - 150) / 400, 0.5, 3.0) applied before GF prediction, so that power is mechanical and the ANN only predicts direction.
  11. As a bot developer, I want standard GF aiming geometry: MEA = arcsin(8.0 / bullet_speed), aim_angle = bearing_to_enemy + gf × MEA, so that the ANN's output maps to a firing angle.
  12. As a bot developer, I want a Virtual Guns system where every gun (including the active one) fires a virtual bullet every tick and tracks hit rate via a rolling window of the last N=100 virtual shots.
  13. As a bot developer, I want wave-based virtual hit detection (expanding circle at bullet speed, check if enemy crosses at predicted angle), so that virtual bullet tracking is O(1) per tick per wave.
  14. As a bot developer, I want the active gun to always be whichever virtual gun has the highest rolling hit rate, with no threshold logic, so that gun switching is automatic and smooth.
  15. As a bot developer, I want confidence-based firing discipline: fire only when the active gun's hit rate exceeds a base threshold (0.15), raised when own energy is low (<30 → +0.10), lowered when enemy energy is low (<30 → −0.10).
  16. As a bot developer, I want power scaling with confidence: power = formula_power × clamp(hit_rate / 0.3, 0.3, 1.0), so that uncertain guns fire weaker to conserve energy.
  17. As a bot developer, I want cold-start handling: on true first encounter (no saved weights), the gun is silent until the GA produces a first champion, then fires. On rematch, loaded weights fire from tick 1.
  18. As a bot developer, I want per-opponent weight persistence keyed by enemy bot name (save champion weights after each match), so that rematches start where the last match ended.
  19. As a bot developer, I want a global fallback weight file (overwritten with the latest match champion after every match), so that first encounters against unknown opponents start from a reasonable weight region instead of random init.
  20. As a bot developer, I want the weight load order to be: per-opponent → global fallback → random init, so that the best available starting point is always used.
  21. As a bot developer, I want the bot to compile and run as a Robocode Tank Royale bot, connecting to the game server and executing tick-by-tick with radar lock, movement, and gun firing.
  22. As a bot developer, I want the bot to beat OscillatorBot (the existing sparring partner) as a minimum viability demonstration.

Implementation Decisions

Architecture (from ADR 0001 and wayfinder maps #61, #74)

  • Gun interface: feed(state) / aim() → (angle, power) — black-box module, bot doesn't know internals.
  • ANN topology: 91→8→1, tanh activation, 745 weights. Fixed topology (TOPO_Gun). Flat weight layout: [W1, b1, W2, b2].
  • GA: mutation-only, no crossover. Gaussian mutation on ALL weights (σ = 0.005–0.01). Population 64–200. Top 20–50% truncation selection. Single elite preserved with re-evaluation. Parallel thread, never blocks inference.
  • Input normalization: per-feature fixed scaling using game constants. lateral_vel / 8.0 ([-1,1]), heading_delta / π ([-1,1]), wall_distance / 1000.0 ([0,1]). Mixed ranges are fine — tanh needs bounded magnitude, not symmetry. The GA compensates.
  • Power formula: power = clamp(3.0 - (distance - 150) / 400, 0.5, 3.0). Applied first; ANN predicts GF within the resulting MEA arc. ANN never sees or controls power.
  • Aiming geometry: bullet_speed = 20 - 3 × power, MEA = arcsin(8.0 / bullet_speed), aim_angle = bearing_to_enemy + gf × MEA.
  • Virtual Guns: rolling window hit rate (N=100 virtual shots). Wave-based detection. Active gun = highest hit rate, no threshold logic. The active gun also runs as a virtual gun — competes on equal footing.
  • Energy management: confidence-based firing gated by hit rate + own/enemy energy modulation. Power scaled by confidence. Cold start silence on first encounter only.
  • Weight persistence: per-opponent (keyed by bot name, ~6KB each) + global fallback (copy latest champion). Load order: per-opponent → global → random init. Hybrid population seeding: ~25% mutated clones of loaded champion, ~75% random.

Deferred decisions

  • NEAT_Gun: deferred until TOPO_Gun hits ceiling.
  • CMA-ES: ideal at d=750 but deferred; truncation GA is acceptable starting point.
  • Offline generalist GA for global fallback: deferred until bot beats all sample bots.

Nim-specific

  • Bot follows Robocode Tank Royale Nim bot conventions (see OscillatorBot for reference structure: .nim, .json bot config, .sh launcher, config.nims).
  • GA population parameters, mutation σ, thresholds, and power formula constants are calibration knobs — defined as constants, tuned by testing.

Testing Decisions

Test philosophy

  • Each pure-logic module includes a when isMainModule self-check block with assertions, following the pattern established by prototypes/ga_gun_spike/ga_spike.nim.
  • Use doAssert for checks that must survive -d:release builds. Use assert for development-only checks.
  • Never compile with -d:danger during testing or training — it disables all runtime safety checks including bounds checking.
  • No test framework. No test files. Self-checks live in the module they test.

Module self-checks

Module Self-check
ANN Known weights + known inputs → expected output (forward pass determinism)
GA Evolve population on toy fitness function → fitness improves (spike pattern)
Normalization Game-constant bounds → values in expected ranges
Power formula Distance samples → clamped power values within [0.5, 3.0]
Wave detection Known geometry (bot position, enemy position, wave radius) → expected hit/miss
Weight persistence Write + read back → identical weights (roundtrip)

Integration testing

  • Integration asserts added as they emerge during assembly — not pre-planned.
  • Primary integration validation: run Evo_Bot against OscillatorBot and observe behavior (convergence, firing, adaptation).
  • The seam between pure logic and Robocode integration is the bot's tick loop — everything below it is testable without the game engine.

Out of Scope

  • Wave surfing / bullet dodging (movement system — separate effort)
  • End-to-end RL (PPO/SAC)
  • Melee / multi-enemy support
  • NEAT_Gun design
  • Offline generalist GA for global fallback weights
  • Bot movement strategy (beyond basic movement needed to not be stationary)
  • Radar strategy beyond simple enemy lock

Further Notes

  • The GA spike (prototypes/ga_gun_spike/ga_spike.nim) validates the core GA+ANN pipeline on sin(x) prediction. It achieved MSE < 0.0002 in 200 generations with a 10→4→1 network. The real bot scales this to 91→8→1 with the same mechanics.
  • Predecessor wayfinder maps: #61 (architecture locked), #74 (implementation design decisions resolved).
  • ADR: docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md
  • Domain glossary: CONTEXT.md
  • GA parameter research: docs/research/ga-parameters-neuroevolution.md
  • All calibration knobs (population size, σ, thresholds, N, power formula constants) are tuning values — start with the decided defaults, adjust based on testing against sparring bots.
## Problem Statement Evo_Bot needs a working gun system that predicts enemy dodge behavior using a neuroevolution approach. The architecture is locked (ADR 0001, wayfinder maps #61 and #74), the GA+ANN pipeline is validated (spike in `prototypes/ga_gun_spike/`), but no bot code exists yet — there is no `src/` directory. The bot needs to be built from scratch, integrating all decided subsystems into a Robocode Tank Royale bot written in Nim. ## Solution Build Evo_Bot as a modular Robocode Tank Royale bot in Nim with a TOPO_Gun (fixed-topology ANN evolved by mutation-only GA), Virtual Guns system, confidence-based firing discipline, and per-opponent weight persistence. The bot predicts enemy dodge behavior via guess factor targeting and adapts mid-match through parallel evolution. ## User Stories 1. As a bot developer, I want a feedforward ANN module (91→8→1, tanh activation) with a pure `forward(weights, inputs) → float` interface, so that the gun can predict guess factors from the sliding window. 2. As a bot developer, I want per-feature fixed normalization using game constants (lateral_vel / 8.0, heading_delta / π, wall_distance / 1000.0), so that ANN inputs are bounded without runtime statistics. 3. As a bot developer, I want a mutation-only GA module that evolves a population of weight vectors, so that the ANN adapts to each opponent mid-match. 4. As a bot developer, I want the GA to use hybrid population seeding (~25% mutated clones of loaded champion, ~75% random init), so that evolution exploits known-good weights while exploring new strategies. 5. As a bot developer, I want GA parameters matching the research findings: population 64–200, Gaussian mutation on ALL weights (σ = 0.005–0.01), top 20–50% truncation selection, single elite with re-evaluation, no crossover. 6. As a bot developer, I want the GA evolution to run on a parallel thread that never blocks the bot's tick loop, so that inference and evolution are decoupled. 7. As a bot developer, I want a sliding window that records the last 30 ticks of enemy state (lateral_vel, heading_delta, wall_distance) as a flat 90-element buffer, plus current distance as the 91st input. 8. As a bot developer, I want a replay tape (rolling ~2000-tick buffer) of recorded enemy states, so that the GA can evaluate fitness offline against historical data. 9. As a bot developer, I want fitness evaluated as virtual bullet hit count on randomly sampled replay tape ticks, so that the GA selects for aiming accuracy. 10. As a bot developer, I want a bullet power formula `power = clamp(3.0 - (distance - 150) / 400, 0.5, 3.0)` applied before GF prediction, so that power is mechanical and the ANN only predicts direction. 11. As a bot developer, I want standard GF aiming geometry: `MEA = arcsin(8.0 / bullet_speed)`, `aim_angle = bearing_to_enemy + gf × MEA`, so that the ANN's output maps to a firing angle. 12. As a bot developer, I want a Virtual Guns system where every gun (including the active one) fires a virtual bullet every tick and tracks hit rate via a rolling window of the last N=100 virtual shots. 13. As a bot developer, I want wave-based virtual hit detection (expanding circle at bullet speed, check if enemy crosses at predicted angle), so that virtual bullet tracking is O(1) per tick per wave. 14. As a bot developer, I want the active gun to always be whichever virtual gun has the highest rolling hit rate, with no threshold logic, so that gun switching is automatic and smooth. 15. As a bot developer, I want confidence-based firing discipline: fire only when the active gun's hit rate exceeds a base threshold (0.15), raised when own energy is low (<30 → +0.10), lowered when enemy energy is low (<30 → −0.10). 16. As a bot developer, I want power scaling with confidence: `power = formula_power × clamp(hit_rate / 0.3, 0.3, 1.0)`, so that uncertain guns fire weaker to conserve energy. 17. As a bot developer, I want cold-start handling: on true first encounter (no saved weights), the gun is silent until the GA produces a first champion, then fires. On rematch, loaded weights fire from tick 1. 18. As a bot developer, I want per-opponent weight persistence keyed by enemy bot name (save champion weights after each match), so that rematches start where the last match ended. 19. As a bot developer, I want a global fallback weight file (overwritten with the latest match champion after every match), so that first encounters against unknown opponents start from a reasonable weight region instead of random init. 20. As a bot developer, I want the weight load order to be: per-opponent → global fallback → random init, so that the best available starting point is always used. 21. As a bot developer, I want the bot to compile and run as a Robocode Tank Royale bot, connecting to the game server and executing tick-by-tick with radar lock, movement, and gun firing. 22. As a bot developer, I want the bot to beat OscillatorBot (the existing sparring partner) as a minimum viability demonstration. ## Implementation Decisions ### Architecture (from ADR 0001 and wayfinder maps #61, #74) - **Gun interface**: `feed(state)` / `aim() → (angle, power)` — black-box module, bot doesn't know internals. - **ANN topology**: 91→8→1, tanh activation, 745 weights. Fixed topology (TOPO_Gun). Flat weight layout: [W1, b1, W2, b2]. - **GA**: mutation-only, no crossover. Gaussian mutation on ALL weights (σ = 0.005–0.01). Population 64–200. Top 20–50% truncation selection. Single elite preserved with re-evaluation. Parallel thread, never blocks inference. - **Input normalization**: per-feature fixed scaling using game constants. `lateral_vel / 8.0` ([-1,1]), `heading_delta / π` ([-1,1]), `wall_distance / 1000.0` ([0,1]). Mixed ranges are fine — tanh needs bounded magnitude, not symmetry. The GA compensates. - **Power formula**: `power = clamp(3.0 - (distance - 150) / 400, 0.5, 3.0)`. Applied first; ANN predicts GF within the resulting MEA arc. ANN never sees or controls power. - **Aiming geometry**: `bullet_speed = 20 - 3 × power`, `MEA = arcsin(8.0 / bullet_speed)`, `aim_angle = bearing_to_enemy + gf × MEA`. - **Virtual Guns**: rolling window hit rate (N=100 virtual shots). Wave-based detection. Active gun = highest hit rate, no threshold logic. The active gun also runs as a virtual gun — competes on equal footing. - **Energy management**: confidence-based firing gated by hit rate + own/enemy energy modulation. Power scaled by confidence. Cold start silence on first encounter only. - **Weight persistence**: per-opponent (keyed by bot name, ~6KB each) + global fallback (copy latest champion). Load order: per-opponent → global → random init. Hybrid population seeding: ~25% mutated clones of loaded champion, ~75% random. ### Deferred decisions - **NEAT_Gun**: deferred until TOPO_Gun hits ceiling. - **CMA-ES**: ideal at d=750 but deferred; truncation GA is acceptable starting point. - **Offline generalist GA** for global fallback: deferred until bot beats all sample bots. ### Nim-specific - Bot follows Robocode Tank Royale Nim bot conventions (see OscillatorBot for reference structure: .nim, .json bot config, .sh launcher, config.nims). - GA population parameters, mutation σ, thresholds, and power formula constants are calibration knobs — defined as constants, tuned by testing. ## Testing Decisions ### Test philosophy - Each pure-logic module includes a `when isMainModule` self-check block with assertions, following the pattern established by `prototypes/ga_gun_spike/ga_spike.nim`. - Use `doAssert` for checks that must survive `-d:release` builds. Use `assert` for development-only checks. - **Never compile with `-d:danger` during testing or training** — it disables all runtime safety checks including bounds checking. - No test framework. No test files. Self-checks live in the module they test. ### Module self-checks | Module | Self-check | |--------|-----------| | ANN | Known weights + known inputs → expected output (forward pass determinism) | | GA | Evolve population on toy fitness function → fitness improves (spike pattern) | | Normalization | Game-constant bounds → values in expected ranges | | Power formula | Distance samples → clamped power values within [0.5, 3.0] | | Wave detection | Known geometry (bot position, enemy position, wave radius) → expected hit/miss | | Weight persistence | Write + read back → identical weights (roundtrip) | ### Integration testing - Integration asserts added as they emerge during assembly — not pre-planned. - Primary integration validation: run Evo_Bot against OscillatorBot and observe behavior (convergence, firing, adaptation). - The seam between pure logic and Robocode integration is the bot's tick loop — everything below it is testable without the game engine. ## Out of Scope - Wave surfing / bullet dodging (movement system — separate effort) - End-to-end RL (PPO/SAC) - Melee / multi-enemy support - NEAT_Gun design - Offline generalist GA for global fallback weights - Bot movement strategy (beyond basic movement needed to not be stationary) - Radar strategy beyond simple enemy lock ## Further Notes - The GA spike (`prototypes/ga_gun_spike/ga_spike.nim`) validates the core GA+ANN pipeline on sin(x) prediction. It achieved MSE < 0.0002 in 200 generations with a 10→4→1 network. The real bot scales this to 91→8→1 with the same mechanics. - Predecessor wayfinder maps: [#61](https://git.fossellini.top/SirStone/SirRoboGarage/issues/61) (architecture locked), [#74](https://git.fossellini.top/SirStone/SirRoboGarage/issues/74) (implementation design decisions resolved). - ADR: `docs/adr/0001-neuroevolution-gun-fixed-topology-ann-evolved-by-ga.md` - Domain glossary: `CONTEXT.md` - GA parameter research: `docs/research/ga-parameters-neuroevolution.md` - All calibration knobs (population size, σ, thresholds, N, power formula constants) are tuning values — start with the decided defaults, adjust based on testing against sparring bots.
SirStone added the ready-for-agent label 2026-08-25 19:05:52 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#80