Spec: GA_Gun only — virtual bullets, tick-speed training, hit% fitness #115

Open
opened 2026-08-29 10:41:25 +02:00 by SirStone · 0 comments
Owner

Problem Statement

EvoBot's GA_Gun currently trains in a background thread using a replay tape of past enemy states. This introduces complexity (thread synchronization, replay tape management, staleness) and decouples training from the actual game dynamics. The gun also outputs only a guess factor — bullet power is a hardcoded distance formula, leaving optimization on the table.

Solution

Replace the background-thread evolution with tick-speed training via virtual bullets. Every GA candidate fires a virtual bullet each tick; after enough bullets resolve, hit% becomes the fitness signal. The champion is hot-swapped into the live gun immediately. The ANN gains a second output for bullet power. No multi-gun selector, no CMA-ES, no replay tape — one gun, one optimizer, one loop.

User Stories

  1. As EvoBot, I want my ANN to output both guess factor and bullet power, so that power is optimized alongside aim.
  2. As EvoBot, I want all 64 GA candidates to fire a virtual bullet every tick, so that fitness evaluation uses real-time game data.
  3. As EvoBot, I want virtual bullets to travel at the correct speed (20 − 3 × power), so that hit detection is physically accurate.
  4. As EvoBot, I want virtual bullet hits detected by intersection with the enemy bounding box, so that fitness reflects actual hit geometry.
  5. As EvoBot, I want each candidate's hit% computed after 30 resolved virtual bullets, so that fitness has a statistically meaningful sample.
  6. As EvoBot, I want the GA to run selection + mutation after all candidates have 30 resolved bullets, so that evolution proceeds at tick speed.
  7. As EvoBot, I want the new champion hot-swapped into the live gun immediately when it beats the current best, so that improvements take effect without waiting for round-end.
  8. As EvoBot, I want the live gun to run the champion network on the sliding window each tick, so that aim tracks the enemy in real time.
  9. As EvoBot, I want the live gun to aim at bearing + GF × MEA using the champion's first output, so that aim uses the standard guess-factor geometry.
  10. As EvoBot, I want the live gun's bullet power set by the champion's second output (tanh rescaled to [0.1, 3.0]), so that power is bounded to legal Robocode values.
  11. As EvoBot, I want the gun to fire when gun heat = 0 and energy ≥ power, so that firing respects game constraints.
  12. As EvoBot, I want the gun to fire from tick 1 using whatever weights are available (loaded or random), so that there is no cold-start silence.
  13. As EvoBot, I want champion weights saved to a global file immediately when a new champion emerges, so that progress survives crashes.
  14. As EvoBot, I want weights loaded from the global file at bot start if it exists, otherwise random init, so that training resumes across battles.
  15. As EvoBot, I want no per-enemy weight files, so that the system stays simple and the global prior is strongest.
  16. As EvoBot, I want a battle-wide hit% diagnostic logged at battle end (hits/total), so that I can observe real gun performance.
  17. As EvoBot, I want the diagnostic to be observability only (not fed back into GA fitness), so that fitness remains virtual-bullet-based.
  18. As EvoBot, I want the ANN topology to be 91→16→8→2 with tanh activations, so that the network has enough capacity for both outputs.
  19. As EvoBot, I want the GA to use mutation-only (no crossover), so that co-adapted weight configurations are not disrupted.
  20. As EvoBot, I want the sliding window to remain 30 ticks of (lateral_velocity, heading_delta, wall_distance) + current distance = 91 inputs, so that the input representation is unchanged.
  21. As EvoBot, I want virtual bullets that miss (exit arena or exceed max travel) to count as resolved misses, so that fitness is not biased toward slow bullets.
  22. As EvoBot, I want the GA population fixed at 64 candidates, so that virtual bullet tracking is bounded and predictable per tick.

Implementation Decisions

  • Single gun, no selector. The multi-gun architecture (spec #107) is deferred. EvoBot runs one GA_Gun. Virtual guns selector, GF_Gun, and CMA_Gun are out of scope.
  • No replay tape. Training uses live virtual bullets, not sampled historical states. The replay tape buffer and its management are removed.
  • No background thread. The GA loop runs synchronously in the tick handler. No thread synchronization, no champion channel, no stop flags.
  • ANN topology: 91→16→8→2. Two hidden layers (tanh), two outputs. First output: GF ∈ [-1, +1]. Second output: tanh rescaled to [0.1, 3.0] for bullet power. ~1,634 weights.
  • GA population: 64. Each candidate fires one virtual bullet per tick → 64 virtual bullets tracked simultaneously.
  • Fitness = hit% after 30 resolved bullets. A bullet "resolves" when it hits the enemy bounding box, exits the arena, or exceeds max travel distance. 30 gives a usable signal without waiting too long.
  • Generation trigger. When ALL 64 candidates have 30 resolved bullets, compute hit% for each, run truncation selection + mutation, reset counters, start next generation.
  • Champion hot-swap. If the new generation's best candidate has higher hit% than current champion, copy its weights into the live inference slot immediately. No round-end gating.
  • Persistence: immediate on new champion. Save to weights/ga_gun.weights (or similar) when champion changes. Load at startup. Single global file, no per-opponent.
  • Bullet power output scaling. power = 0.1 + (tanh_output + 1) / 2 × 2.9 maps [-1, +1] → [0.1, 3.0].
  • Virtual bullet physics. Speed = 20 − 3 × power. Travel in straight line from gun position at predicted angle. Hit = circle-rect intersection with 36×36 enemy bounding box (standard Robocode TankRoyale hitbox).
  • Cold start. Random weight init. Champion = first candidate. Fire from tick 1.

Testing Decisions

What makes a good test

Tests assert external behavior of modules, not internal state. Feed inputs → check outputs. No mocking internal functions.

Seam 1: GA_Gun module boundary (primary)

Feed tick data and enemy states into the GA_Gun module. Assert:

  • It produces (aim_angle, bullet_power) each tick.
  • bullet_power is in [0.1, 3.0].
  • aim_angle is a valid bearing offset.
  • After enough ticks for a full generation cycle, fitness tracking has progressed (generation counter incremented).
  • Champion weights change when a better candidate emerges.

This exercises ANN inference, virtual bullet spawning, GA loop, and champion swap end-to-end.

Seam 2: Virtual bullet hit detection (secondary)

Pure geometry test. Feed known bullet trajectories (position, velocity vector) and enemy positions/sizes. Assert:

  • Bullet at enemy center → hit.
  • Bullet 1px outside bounding box → miss.
  • Bullet that exits arena bounds → resolved as miss.
  • Edge case: bullet spawned at point-blank range.

This is isolated because circle-rect intersection is a known bug magnet in Robocode bots.

Prior art

tests/handler_seam_v2.nim — existing integration test pattern using event handler seams. New tests should follow the same assert-based self-check style (no test framework, just assert + runnable main).

Out of Scope

  • Multi-gun selector and GF_Gun (spec #107, deferred).
  • CMA-ES gun and optimizer (spec #107, deferred).
  • Per-enemy weight persistence (eliminated by design).
  • Replay tape (replaced by virtual bullets).
  • Background evolution threads (replaced by tick-speed loop).
  • Movement strategy (EvoBot movement is unchanged).
  • Radar lock (unchanged, uses shared module).
  • Topology evolution / NEAT (Phase 3 per ADR-0002, not this spec).
  • Bullet power as a third optimization axis beyond the ANN output (e.g., game-theoretic energy management).

Further Notes

  • This spec supersedes the training approach in the current evo-bot branch (background thread + replay tape) but NOT the multi-gun spec (#107). The multi-gun architecture can layer on top later — this spec delivers a working single-gun EvoBot that trains in real time.
  • The 15 locked decisions referenced in map issue #108 are incorporated into the Implementation Decisions section above.
  • Tasks #109–#114 map directly to this spec's user stories. Dependency chain: #109 (ANN 2 outputs) → #110 (virtual bullets) → #111 (GA tick loop) → #113 (persistence). #112 (real gun fires) depends on #109. #114 (diagnostics) depends on #112.
## Problem Statement EvoBot's GA_Gun currently trains in a background thread using a replay tape of past enemy states. This introduces complexity (thread synchronization, replay tape management, staleness) and decouples training from the actual game dynamics. The gun also outputs only a guess factor — bullet power is a hardcoded distance formula, leaving optimization on the table. ## Solution Replace the background-thread evolution with **tick-speed training via virtual bullets**. Every GA candidate fires a virtual bullet each tick; after enough bullets resolve, hit% becomes the fitness signal. The champion is hot-swapped into the live gun immediately. The ANN gains a second output for bullet power. No multi-gun selector, no CMA-ES, no replay tape — one gun, one optimizer, one loop. ## User Stories 1. As EvoBot, I want my ANN to output both guess factor and bullet power, so that power is optimized alongside aim. 2. As EvoBot, I want all 64 GA candidates to fire a virtual bullet every tick, so that fitness evaluation uses real-time game data. 3. As EvoBot, I want virtual bullets to travel at the correct speed (20 − 3 × power), so that hit detection is physically accurate. 4. As EvoBot, I want virtual bullet hits detected by intersection with the enemy bounding box, so that fitness reflects actual hit geometry. 5. As EvoBot, I want each candidate's hit% computed after 30 resolved virtual bullets, so that fitness has a statistically meaningful sample. 6. As EvoBot, I want the GA to run selection + mutation after all candidates have 30 resolved bullets, so that evolution proceeds at tick speed. 7. As EvoBot, I want the new champion hot-swapped into the live gun immediately when it beats the current best, so that improvements take effect without waiting for round-end. 8. As EvoBot, I want the live gun to run the champion network on the sliding window each tick, so that aim tracks the enemy in real time. 9. As EvoBot, I want the live gun to aim at bearing + GF × MEA using the champion's first output, so that aim uses the standard guess-factor geometry. 10. As EvoBot, I want the live gun's bullet power set by the champion's second output (tanh rescaled to [0.1, 3.0]), so that power is bounded to legal Robocode values. 11. As EvoBot, I want the gun to fire when gun heat = 0 and energy ≥ power, so that firing respects game constraints. 12. As EvoBot, I want the gun to fire from tick 1 using whatever weights are available (loaded or random), so that there is no cold-start silence. 13. As EvoBot, I want champion weights saved to a global file immediately when a new champion emerges, so that progress survives crashes. 14. As EvoBot, I want weights loaded from the global file at bot start if it exists, otherwise random init, so that training resumes across battles. 15. As EvoBot, I want no per-enemy weight files, so that the system stays simple and the global prior is strongest. 16. As EvoBot, I want a battle-wide hit% diagnostic logged at battle end (hits/total), so that I can observe real gun performance. 17. As EvoBot, I want the diagnostic to be observability only (not fed back into GA fitness), so that fitness remains virtual-bullet-based. 18. As EvoBot, I want the ANN topology to be 91→16→8→2 with tanh activations, so that the network has enough capacity for both outputs. 19. As EvoBot, I want the GA to use mutation-only (no crossover), so that co-adapted weight configurations are not disrupted. 20. As EvoBot, I want the sliding window to remain 30 ticks of (lateral_velocity, heading_delta, wall_distance) + current distance = 91 inputs, so that the input representation is unchanged. 21. As EvoBot, I want virtual bullets that miss (exit arena or exceed max travel) to count as resolved misses, so that fitness is not biased toward slow bullets. 22. As EvoBot, I want the GA population fixed at 64 candidates, so that virtual bullet tracking is bounded and predictable per tick. ## Implementation Decisions - **Single gun, no selector.** The multi-gun architecture (spec #107) is deferred. EvoBot runs one GA_Gun. Virtual guns selector, GF_Gun, and CMA_Gun are out of scope. - **No replay tape.** Training uses live virtual bullets, not sampled historical states. The replay tape buffer and its management are removed. - **No background thread.** The GA loop runs synchronously in the tick handler. No thread synchronization, no champion channel, no stop flags. - **ANN topology: 91→16→8→2.** Two hidden layers (tanh), two outputs. First output: GF ∈ [-1, +1]. Second output: tanh rescaled to [0.1, 3.0] for bullet power. ~1,634 weights. - **GA population: 64.** Each candidate fires one virtual bullet per tick → 64 virtual bullets tracked simultaneously. - **Fitness = hit% after 30 resolved bullets.** A bullet "resolves" when it hits the enemy bounding box, exits the arena, or exceeds max travel distance. 30 gives a usable signal without waiting too long. - **Generation trigger.** When ALL 64 candidates have 30 resolved bullets, compute hit% for each, run truncation selection + mutation, reset counters, start next generation. - **Champion hot-swap.** If the new generation's best candidate has higher hit% than current champion, copy its weights into the live inference slot immediately. No round-end gating. - **Persistence: immediate on new champion.** Save to `weights/ga_gun.weights` (or similar) when champion changes. Load at startup. Single global file, no per-opponent. - **Bullet power output scaling.** `power = 0.1 + (tanh_output + 1) / 2 × 2.9` maps [-1, +1] → [0.1, 3.0]. - **Virtual bullet physics.** Speed = 20 − 3 × power. Travel in straight line from gun position at predicted angle. Hit = circle-rect intersection with 36×36 enemy bounding box (standard Robocode TankRoyale hitbox). - **Cold start.** Random weight init. Champion = first candidate. Fire from tick 1. ## Testing Decisions ### What makes a good test Tests assert external behavior of modules, not internal state. Feed inputs → check outputs. No mocking internal functions. ### Seam 1: GA_Gun module boundary (primary) Feed tick data and enemy states into the GA_Gun module. Assert: - It produces (aim_angle, bullet_power) each tick. - bullet_power is in [0.1, 3.0]. - aim_angle is a valid bearing offset. - After enough ticks for a full generation cycle, fitness tracking has progressed (generation counter incremented). - Champion weights change when a better candidate emerges. This exercises ANN inference, virtual bullet spawning, GA loop, and champion swap end-to-end. ### Seam 2: Virtual bullet hit detection (secondary) Pure geometry test. Feed known bullet trajectories (position, velocity vector) and enemy positions/sizes. Assert: - Bullet at enemy center → hit. - Bullet 1px outside bounding box → miss. - Bullet that exits arena bounds → resolved as miss. - Edge case: bullet spawned at point-blank range. This is isolated because circle-rect intersection is a known bug magnet in Robocode bots. ### Prior art `tests/handler_seam_v2.nim` — existing integration test pattern using event handler seams. New tests should follow the same assert-based self-check style (no test framework, just assert + runnable main). ## Out of Scope - Multi-gun selector and GF_Gun (spec #107, deferred). - CMA-ES gun and optimizer (spec #107, deferred). - Per-enemy weight persistence (eliminated by design). - Replay tape (replaced by virtual bullets). - Background evolution threads (replaced by tick-speed loop). - Movement strategy (EvoBot movement is unchanged). - Radar lock (unchanged, uses shared module). - Topology evolution / NEAT (Phase 3 per ADR-0002, not this spec). - Bullet power as a third optimization axis beyond the ANN output (e.g., game-theoretic energy management). ## Further Notes - This spec supersedes the training approach in the current evo-bot branch (background thread + replay tape) but NOT the multi-gun spec (#107). The multi-gun architecture can layer on top later — this spec delivers a working single-gun EvoBot that trains in real time. - The 15 locked decisions referenced in map issue #108 are incorporated into the Implementation Decisions section above. - Tasks #109–#114 map directly to this spec's user stories. Dependency chain: #109 (ANN 2 outputs) → #110 (virtual bullets) → #111 (GA tick loop) → #113 (persistence). #112 (real gun fires) depends on #109. #114 (diagnostics) depends on #112.
SirStone added the ready-for-agent label 2026-08-29 10:41:27 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#115