Files
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00

128 lines
8.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gun Predictor Shootout — Final Report
## Executive Summary
WiSARD K=14 (bleach=1, 276-bit input) wins the hyperparameter sweep with MAE=1.37px on the single-enemy multi-frame dataset. In head-to-head battle backtest across 4 enemy types, WiSARD K=12 beats both BNNBot (linear+Hebbian) and TsetlinBot on 3/4 enemies, with an overall MAE of 16.20 vs 17.61 (BNNBot) and 20.69 (TsetlinBot). The Tsetlin Machine is the strongest warm-phase predictor in isolation (MAE(last⅓)=0.27 for best config) but its cold-start overhead costs it in battles with <~300 ticks of data. Ship WiSARD K=14, bleach=1. If cold-start penalty on TM is ever fixed (e.g., pre-training or warm-up from WiSARD), TM becomes worth revisiting.
---
## Methods Tested
**16 method families, ~260 configurations total:**
- WiSARD (192 configs: K=4–20, bleach=0–3, ±XOR features)
- Tsetlin Machine / RTM (57 configs: 10–500 clauses, s=1.5–15, T=10–200, ±weighted)
- Echo State Network (reservoir=512)
- Kanerva Sparse Distributed Memory (addr=2000)
- N-gram Markov predictor (4×69 chunks)
- Bloom Filter (8192 slots)
- Hyperdimensional Computing / VSA (n_hd=2000, 32 classes)
- Random Subspace ensemble (30×50 bits)
- WiSARD + Eligibility Traces (k=12, trace=5)
---
## Results Table
> **Dataset note:** WiSARD sweep and alternatives used a 277–3153 row single-enemy dataset at rep power p1.07. Battle comparison used 73 380 data points across 4 enemy types and all bullet powers. MAE numbers are not directly comparable across the two columns — treat sweep MAE as a relative rank within each sweep, and battle MAE as the real-world number.
| Rank | Method | Best Config | MAE (sweep) | MAE (battle) | Hit% (battle) | Memory | Learning Speed | Verdict |
|------|--------|-------------|-------------|--------------|---------------|--------|----------------|---------|
| 1 | WiSARD | K=14, bleach=1, 276b | 1.37 px | 16.20 | 65.4% | ~276b active | fast | **SHIP IT** |
| 2 | WiSARD+Elig | k=12, trace=5 | 8.24* | — | 90.3%* | ~276b active | fast | promising, untested in battle |
| 3 | RandSubspace | 30×50 bits | 8.31* | — | 90.6%* | ~1500b | fast | similar to WiSARD, no battle test |
| 4 | Echo State Net | res=512 | 8.29* | — | 85.2%* | large (float reservoir) | med | too heavy for Robocode JVM |
| 5 | BNNBot (Lin+Heb) | linear+Hebbian, 24-cell | — | 17.61 | 61.9% | tiny | fast | current baseline |
| 6 | Tsetlin Machine | 500 cl, T=50, s=1.5, W=Y | 2.84† | 20.69 (warm: 18.77) | 55.8% | 531 KB | slow (cold-start) | future work |
| 7 | N-gram Markov | 4×69 chunks | 10.11* | — | 80.5%* | small | med | no improvement on WiSARD |
| 8 | Kanerva SDM | addr=2000 | 10.16* | — | 80.1%* | med | slow | no improvement on baseline |
| 9 | Bloom Filter | 8192 slots | 17.08* | — | 59.2%* | 8 KB | fast | worse than baseline |
| 10 | HDC/VSA | n_hd=2000, 32 cls | 65.33* | — | 0.0%* | large | slow | broken for this task |
*Measured on 277-row alternatives dataset (single enemy, p1.07).
†Measured on 2500-row TM sweep dataset (single enemy).
Battle column from 4-enemy head-to-head (73 380 rows, all powers).
---
## Per-Enemy Analysis
| Enemy | BNNBot MAE | WiSARD MAE | TsetlinBot MAE | Winner | Notes |
|-------|-----------|-----------|---------------|--------|-------|
| target | 2.68 | **2.21** | 4.14 | WiSARD | Predictable bot, all methods work; WiSARD widest margin |
| walls | 15.31 | **13.22** | 18.81 | WiSARD | Wall-bouncing pattern, WiSARD generalises better |
| crazy | 24.72 | **22.29** | 27.10 | WiSARD | Highly random; WiSARD still edges out |
| spinbot | **23.75** | 24.54 | 30.20 | BNNBot | Periodic rotation pattern; Hebbian table captures it |
**Specialist vs generalist:** WiSARD is the best generalist (3/4 wins). BNNBot's Hebbian residual table memorises spinbot's periodic pattern more effectively than WiSARD's RAM-based lookup. TM loses on all 4 — its cold start dominates the all-rows metric in battles of this length.
TM warm-phase on crazy (20.98) nearly matches WiSARD (22.29), suggesting TM catches up once trained but battles are typically too short for it to amortise the cold start.
---
## Key Insights
### WiSARD
- **K=14 is the sweet spot.** MAE plateaus K=12–20; diminishing returns above K=14. K=4–6 measurably worse.
- **Bleaching has marginal effect.** bleach=0 vs bleach=1 vs bleach=2 all within ±0.02 MAE. Skip tuning it; leave at 1.
- **XOR features add nothing.** The +XOR variants (345b input) are uniformly equal or worse than plain 276b for the same K. Extra bits, no gain.
- **Cold-start (MAE f50) still high** (~0.43–0.58 px for best configs), but last-50 converges to ~0.00 px — the model fully memorises the pattern after enough data.
- **Baseline delta:** best WiSARD is -0.26 px vs P1 linear, which sounds small but the absolute MAE of 1.37 vs 1.63 is a ~16% improvement at p1.07. Battle-scale improvement is larger (16.20 vs 17.61, ~8% reduction).
### Tsetlin Machine
- **Weighted clauses are mandatory.** Unweighted configs are consistently 0.5–2.0 MAE higher at identical clause counts.
- **s=1.5 is optimal.** Specificity 3.0 works, 5.0+ degrades, 10–15 is too sparse.
- **T (threshold) barely matters** across 10–200 range — vote clamping is rarely active on these dataset sizes.
- **Cold-start dominates all-rows MAE.** The "MAE(all)" column is misleading — it averages a bad warm-up phase over the whole run. The relevant metric for a real battle is MAE(last⅓): best config gets 0.27 px (!) for 500 clauses, 0.15 px for 100 clauses.
- **Practical config:** 100 clauses, T=100, s=1.5, states=32, weighted=Y. MAE(all)=3.28, MAE(last⅓)=0.20, 106 KB. The 500-clause config is marginally better but 5× the memory and slow.
- **Why it loses in battles:** RTM shared x/y with 60 clauses was used in the battle test — underspecified relative to the sweep's best config. Even so, TM warm-phase (18.77) still loses to WiSARD (16.20) overall.
### Alternatives
- **WiSARD+Eligibility Traces** is the real second-place finisher (MAE=8.24, Hit%=90.3% on the 277-row test). Eligibility traces let the RAM weights decay — this is strictly better than vanilla WiSARD on the small-dataset test. Not yet battle-tested.
- **Random Subspace ensemble** (8.31, 90.6% hit) matches WiSARD+Elig and is simpler to implement, but also untested in the full battle setting.
- **ESN** (8.29) is competitive but requires a float reservoir — doesn't fit the binary/frugal constraint of this project.
- **N-gram and SDM** essentially replicate linear baseline performance. Not worth the complexity.
- **Bloom Filter** — worse than baseline. RAM addressed by pattern ID without content-addressing fails here.
- **HDC/VSA** — completely broken (0% hit rate). The 32-class bucketing is too coarse for continuous angular prediction; this architecture is wrong for regression.
### Eligibility Traces Finding
The WiSARD+Elig variant (trace=5) improves MAE by ~1.7 px and hit rate by +10 pp vs vanilla WiSARD K=12 on the same 277-row dataset. This is the largest single improvement found across all alternatives. Mechanism: eligibility traces propagate reinforcement to recently active RAM addresses, not just the current one — capturing temporal credit assignment that WiSARD's instantaneous lookup misses.
---
## Recommendation
**Ship:** WiSARD K=14, bleach=1, 276-bit input (4 frames × 69 bits), 23 RAM nodes.
```
K = 14
bleach = 1
input_bits = 276 # 4 frames × 69 bits, no XOR
```
This is the best-validated config with a real battle backtest (WiSARD K=12 wins 3/4 enemies; K=14 is marginally better in the sweep). Memory is ~276 bits active per RAM node, negligible for the JVM.
**Explore next (in priority order):**
1. **WiSARD+Eligibility Traces (k=14, trace=5)** — +10 pp hit rate in the alternatives test. Straightforward to implement on top of current WiSARD. Highest expected ROI.
2. **Tsetlin Machine with warm-start** — if you can pre-train TM on a stored trajectory from the prior round, cold-start vanishes and TM's MAE(last⅓)=0.20 becomes competitive. Requires inter-round state persistence.
3. **Random Subspace ensemble** — drop-in if eligibility traces are too complex; similar performance.
4. **Per-enemy specialisation** — spinbot is the one case where Hebbian residuals beat WiSARD. A simple enemy-classifier + model switcher could capture that.
Do not pursue: HDC, Bloom Filter, SDM. None improved on linear.
---
## Raw Data References
| File | Content |
|------|---------|
| `analysis/sweep_wisard_results.txt` | WiSARD sweep: 192 configs (K=4–20, bleach=0–3, ±XOR), 3153 rows, 3 seeds each |
| `analysis/sweep_tsetlin_results.txt` | TM sweep: 57 configs (clauses=10–500, s=1.5–15, T=10–200, ±weighted), 2500 rows |
| `analysis/alternatives_results.txt` | 7 alternative methods vs WiSARD K=12 baseline, 277 rows |
| `analysis/battle_comparison.txt` | Head-to-head battle backtest: BNNBot vs WiSARD K=12 vs TsetlinBot, 4 enemies, 73 380 data points |