Spec: Multi-gun EvoBot — three guns, one selector, global-only weights #107

Open
opened 2026-08-28 22:10:04 +02:00 by SirStone · 0 comments
Owner

part-of: #99

Problem Statement

Evo_Bot currently has a single TOPO_Gun (fixed-topology ANN evolved by mutation-only GA) with per-enemy weight persistence. This approach doesn't generalize across opponents and offers no diversity of aiming strategies. The bot needs multiple competing gun systems with a real-time selector that converges on the best performer, using global-only weights.

Solution

Equip EvoBot with three guns — GF_Gun (guess-factor histogram), GA_Gun (mutation-only GA, current TOPO_Gun renamed), and CMA_Gun (full CMA-ES optimizer) — behind a virtual guns selector that uses angle-delta tracking to pick the best performer after a trust threshold of 10 observations. Strip per-enemy persistence entirely; one global weights file per learning gun. GF_Gun fires from tick 1 as default; after 10 angle-delta observations, the selector picks the best-scoring gun.

User Stories

  1. As a bot operator, I want EvoBot to fire from tick 1 using the GF_Gun, so that the bot is never idle at the start of a match.
  2. As a bot operator, I want a virtual guns selector that tracks angle-delta per gun, so that the best-performing gun is chosen in real-time.
  3. As a bot operator, I want the selector to default to GF_Gun for the first 10 observations, so that gun selection doesn't thrash on insufficient data.
  4. As a bot operator, I want the selector to switch to the best angle-delta scorer after 10 observations, so that aim accuracy improves as the match progresses.
  5. As a bot operator, I want a GF_Gun that builds an unsegmented guess-factor histogram from enemy lateral velocity, so that there is a fast stateless baseline aiming strategy.
  6. As a bot operator, I want GA_Gun to use the same 91→16→8→1 network as CMA_Gun, so that both learning guns are architecturally comparable.
  7. As a bot operator, I want GA_Gun to evolve weights via mutation-only GA (existing logic), so that the proven evolution path is retained.
  8. As a bot operator, I want a CMA_Gun that evolves weights via full CMA-ES with auto-sized λ (~24 for 1625 dims), so that a gradient-free optimizer with better convergence properties is available.
  9. As a bot operator, I want both GA_Gun and CMA_Gun to run their own evolution threads concurrently, so that both optimizers evolve in parallel during a match.
  10. As a bot operator, I want a shared replay tape (one writer from the tick loop, multiple readers for both evo threads and GF_Gun), so that all guns see the same enemy data.
  11. As a bot operator, I want the sliding window to remain 30 ticks (30×3 features + distance = 91 inputs), so that the input contract is stable.
  12. As a bot operator, I want global-only weight persistence — one global_ga.weights and one global_cma.weights — so that learned weights generalize across opponents.
  13. As a bot operator, I want per-enemy weight files and fallback logic removed entirely, so that the persistence model is simple and unambiguous.
  14. As a bot operator, I want the bot to always fire with random-init weights from tick 1 and hot-swap to the latest champion when an evo thread delivers one, so that there is no cold-start silence.
  15. As a bot operator, I want the ANN architecture to be 91→16→8→1 (two hidden layers of 16 and 8, ~1625 weights), so that the network has sufficient capacity for dodge prediction.
  16. As a bot operator, I want HiddenDim defined as a const (recompile to change), so that architecture changes are explicit and intentional.
  17. As a bot operator, I want ADR-0002 to supersede ADR-0001, documenting the three-gun architecture, CMA-ES + GA coexistence, global-only weights rationale, and virtual guns selector design.
  18. As a bot operator, I want CONTEXT.md updated with new glossary entries (CMA_Gun, GF_Gun, Virtual Guns Selector, Angle-Delta Tracking) and the Weight Persistence definition changed to "global-only, one file per learning gun."

Implementation Decisions

  • Three gun modules: GF_Gun (stateless histogram), GA_Gun (mutation-only GA, renamed from TOPO_Gun), CMA_Gun (full CMA-ES). Each is a separate module.
  • CMA-ES module (cmaes.nim): standalone optimizer, does not replace ga.nim — both coexist. Auto-sized λ based on dimensionality.
  • ANN architecture: 91→16→8→1, two hidden layers (16, 8), ~1625 weights. HiddenDim as compile-time const.
  • Virtual guns selector: angle-delta tracking (not wave-based). Each gun's predicted aim angle is compared to where the enemy actually was; running mean of angular error. Trust threshold N=10 shots before selector overrides GF_Gun default.
  • Always fire: random-init weights from tick 1, hot-swap to champion when available. No cold-start waiting.
  • Two evo threads: GA_Gun and CMA_Gun each own an evolution thread, running concurrently against the shared replay tape.
  • Shared replay tape: one writer (tick loop), multiple readers. Existing replay tape mechanism extended for concurrent read access.
  • Global-only persistence: one global_ga.weights, one global_cma.weights. Per-enemy files and load-order fallback removed entirely.
  • ADR-0002 supersedes ADR-0001. CONTEXT.md glossary updated.

Testing Decisions

  • Good tests test external behavior at module boundaries, not internal implementation. Each module gets one behavioral test at its seam — no per-function unit tests.
  • CMA-ES module seam: feed a synthetic fitness function (e.g., sphere function), assert convergence toward known optimum within N generations. Pure math, no game dependency.
  • GF_Gun seam: feed a sequence of enemy states, assert the guess-factor output matches expected histogram peak. No ANN, no evolution.
  • Virtual guns selector seam: feed mock angle-delta observations for multiple guns, assert the selector picks the correct winner after threshold N=10. No real guns needed.
  • ANN forward pass seam: feed known weights + known input vector, assert deterministic output. Pure math.
  • Integration seam (for #105): wire all guns, feed a replay tape, assert the system fires and the selector converges to the best gun. One end-to-end test.
  • Prior art: look at existing test patterns in the repo (if any exist in other *_garage/tests/ directories).

Out of Scope

  • Per-enemy weight specialization (explicitly removed)
  • NEAT_Gun (topology evolution) — deferred until fixed-topology guns hit ceiling
  • Wave surfing / bullet dodging (movement system, separate effort)
  • Wave-based virtual gun detection (angle-delta is sufficient)
  • End-to-end RL (PPO/SAC)
  • Melee / multi-enemy support
  • Training speed optimization (deferred until measured)
  • CMA-ES hyperparameter tuning beyond default λ (deferred until convergence observed)
  • GF_Gun segmentation by lateral velocity or distance (upgrade path if unsegmented baseline is too weak)

Further Notes

  • Child tickets #100–#106 under parent #99 cover the individual implementation tasks. This spec covers the complete feature for agent handoff.
  • Frontier (unblocked) tasks: #100, #101, #102, #103, #104, #106. Task #105 (wire everything together) is blocked by all others.
  • Domain glossary lives in CONTEXT.md at repo root. All implementation should use the glossary terms.
part-of: #99 ## Problem Statement Evo_Bot currently has a single TOPO_Gun (fixed-topology ANN evolved by mutation-only GA) with per-enemy weight persistence. This approach doesn't generalize across opponents and offers no diversity of aiming strategies. The bot needs multiple competing gun systems with a real-time selector that converges on the best performer, using global-only weights. ## Solution Equip EvoBot with three guns — GF_Gun (guess-factor histogram), GA_Gun (mutation-only GA, current TOPO_Gun renamed), and CMA_Gun (full CMA-ES optimizer) — behind a virtual guns selector that uses angle-delta tracking to pick the best performer after a trust threshold of 10 observations. Strip per-enemy persistence entirely; one global weights file per learning gun. GF_Gun fires from tick 1 as default; after 10 angle-delta observations, the selector picks the best-scoring gun. ## User Stories 1. As a bot operator, I want EvoBot to fire from tick 1 using the GF_Gun, so that the bot is never idle at the start of a match. 2. As a bot operator, I want a virtual guns selector that tracks angle-delta per gun, so that the best-performing gun is chosen in real-time. 3. As a bot operator, I want the selector to default to GF_Gun for the first 10 observations, so that gun selection doesn't thrash on insufficient data. 4. As a bot operator, I want the selector to switch to the best angle-delta scorer after 10 observations, so that aim accuracy improves as the match progresses. 5. As a bot operator, I want a GF_Gun that builds an unsegmented guess-factor histogram from enemy lateral velocity, so that there is a fast stateless baseline aiming strategy. 6. As a bot operator, I want GA_Gun to use the same 91→16→8→1 network as CMA_Gun, so that both learning guns are architecturally comparable. 7. As a bot operator, I want GA_Gun to evolve weights via mutation-only GA (existing logic), so that the proven evolution path is retained. 8. As a bot operator, I want a CMA_Gun that evolves weights via full CMA-ES with auto-sized λ (~24 for 1625 dims), so that a gradient-free optimizer with better convergence properties is available. 9. As a bot operator, I want both GA_Gun and CMA_Gun to run their own evolution threads concurrently, so that both optimizers evolve in parallel during a match. 10. As a bot operator, I want a shared replay tape (one writer from the tick loop, multiple readers for both evo threads and GF_Gun), so that all guns see the same enemy data. 11. As a bot operator, I want the sliding window to remain 30 ticks (30×3 features + distance = 91 inputs), so that the input contract is stable. 12. As a bot operator, I want global-only weight persistence — one `global_ga.weights` and one `global_cma.weights` — so that learned weights generalize across opponents. 13. As a bot operator, I want per-enemy weight files and fallback logic removed entirely, so that the persistence model is simple and unambiguous. 14. As a bot operator, I want the bot to always fire with random-init weights from tick 1 and hot-swap to the latest champion when an evo thread delivers one, so that there is no cold-start silence. 15. As a bot operator, I want the ANN architecture to be 91→16→8→1 (two hidden layers of 16 and 8, ~1625 weights), so that the network has sufficient capacity for dodge prediction. 16. As a bot operator, I want HiddenDim defined as a const (recompile to change), so that architecture changes are explicit and intentional. 17. As a bot operator, I want ADR-0002 to supersede ADR-0001, documenting the three-gun architecture, CMA-ES + GA coexistence, global-only weights rationale, and virtual guns selector design. 18. As a bot operator, I want CONTEXT.md updated with new glossary entries (CMA_Gun, GF_Gun, Virtual Guns Selector, Angle-Delta Tracking) and the Weight Persistence definition changed to "global-only, one file per learning gun." ## Implementation Decisions - **Three gun modules**: GF_Gun (stateless histogram), GA_Gun (mutation-only GA, renamed from TOPO_Gun), CMA_Gun (full CMA-ES). Each is a separate module. - **CMA-ES module** (`cmaes.nim`): standalone optimizer, does not replace `ga.nim` — both coexist. Auto-sized λ based on dimensionality. - **ANN architecture**: 91→16→8→1, two hidden layers (16, 8), ~1625 weights. `HiddenDim` as compile-time const. - **Virtual guns selector**: angle-delta tracking (not wave-based). Each gun's predicted aim angle is compared to where the enemy actually was; running mean of angular error. Trust threshold N=10 shots before selector overrides GF_Gun default. - **Always fire**: random-init weights from tick 1, hot-swap to champion when available. No cold-start waiting. - **Two evo threads**: GA_Gun and CMA_Gun each own an evolution thread, running concurrently against the shared replay tape. - **Shared replay tape**: one writer (tick loop), multiple readers. Existing replay tape mechanism extended for concurrent read access. - **Global-only persistence**: one `global_ga.weights`, one `global_cma.weights`. Per-enemy files and load-order fallback removed entirely. - **ADR-0002** supersedes ADR-0001. CONTEXT.md glossary updated. ## Testing Decisions - **Good tests test external behavior at module boundaries, not internal implementation.** Each module gets one behavioral test at its seam — no per-function unit tests. - **CMA-ES module seam**: feed a synthetic fitness function (e.g., sphere function), assert convergence toward known optimum within N generations. Pure math, no game dependency. - **GF_Gun seam**: feed a sequence of enemy states, assert the guess-factor output matches expected histogram peak. No ANN, no evolution. - **Virtual guns selector seam**: feed mock angle-delta observations for multiple guns, assert the selector picks the correct winner after threshold N=10. No real guns needed. - **ANN forward pass seam**: feed known weights + known input vector, assert deterministic output. Pure math. - **Integration seam** (for #105): wire all guns, feed a replay tape, assert the system fires and the selector converges to the best gun. One end-to-end test. - **Prior art**: look at existing test patterns in the repo (if any exist in other `*_garage/tests/` directories). ## Out of Scope - Per-enemy weight specialization (explicitly removed) - NEAT_Gun (topology evolution) — deferred until fixed-topology guns hit ceiling - Wave surfing / bullet dodging (movement system, separate effort) - Wave-based virtual gun detection (angle-delta is sufficient) - End-to-end RL (PPO/SAC) - Melee / multi-enemy support - Training speed optimization (deferred until measured) - CMA-ES hyperparameter tuning beyond default λ (deferred until convergence observed) - GF_Gun segmentation by lateral velocity or distance (upgrade path if unsegmented baseline is too weak) ## Further Notes - Child tickets #100–#106 under parent #99 cover the individual implementation tasks. This spec covers the complete feature for agent handoff. - Frontier (unblocked) tasks: #100, #101, #102, #103, #104, #106. Task #105 (wire everything together) is blocked by all others. - Domain glossary lives in CONTEXT.md at repo root. All implementation should use the glossary terms.
SirStone added the ready-for-agent label 2026-08-28 22:10:06 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SirStone/SirRoboGarage#107