1ed7797cb6
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay) - New modules: minimum-risk melee movement, spinning melee radar - New test bots: PatternMover, RandomMover, WaveSurfer - Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning - Fixed: TM gun warmup gating + directional residuals - Fixed: circular gun integrated formula + multi-bin omega cache - Fixed: oscillator wall-bounce lockout - Fixed: phantom meteor perpendicular body orientation - 6/6 battle wins across all enemy types
128 lines
8.8 KiB
Markdown
128 lines
8.8 KiB
Markdown
# Gun Predictor Shootout — Final Report
|
||
|
||
## Executive Summary
|
||
|
||
WiSARD K=14 (bleach=1, 276-bit input) wins the hyperparameter sweep with MAE=1.37px on the single-enemy multi-frame dataset. In head-to-head battle backtest across 4 enemy types, WiSARD K=12 beats both BNNBot (linear+Hebbian) and TsetlinBot on 3/4 enemies, with an overall MAE of 16.20 vs 17.61 (BNNBot) and 20.69 (TsetlinBot). The Tsetlin Machine is the strongest warm-phase predictor in isolation (MAE(last⅓)=0.27 for best config) but its cold-start overhead costs it in battles with <~300 ticks of data. Ship WiSARD K=14, bleach=1. If cold-start penalty on TM is ever fixed (e.g., pre-training or warm-up from WiSARD), TM becomes worth revisiting.
|
||
|
||
---
|
||
|
||
## Methods Tested
|
||
|
||
**16 method families, ~260 configurations total:**
|
||
|
||
- WiSARD (192 configs: K=4–20, bleach=0–3, ±XOR features)
|
||
- Tsetlin Machine / RTM (57 configs: 10–500 clauses, s=1.5–15, T=10–200, ±weighted)
|
||
- Echo State Network (reservoir=512)
|
||
- Kanerva Sparse Distributed Memory (addr=2000)
|
||
- N-gram Markov predictor (4×69 chunks)
|
||
- Bloom Filter (8192 slots)
|
||
- Hyperdimensional Computing / VSA (n_hd=2000, 32 classes)
|
||
- Random Subspace ensemble (30×50 bits)
|
||
- WiSARD + Eligibility Traces (k=12, trace=5)
|
||
|
||
---
|
||
|
||
## Results Table
|
||
|
||
> **Dataset note:** WiSARD sweep and alternatives used a 277–3153 row single-enemy dataset at rep power p1.07. Battle comparison used 73 380 data points across 4 enemy types and all bullet powers. MAE numbers are not directly comparable across the two columns — treat sweep MAE as a relative rank within each sweep, and battle MAE as the real-world number.
|
||
|
||
| Rank | Method | Best Config | MAE (sweep) | MAE (battle) | Hit% (battle) | Memory | Learning Speed | Verdict |
|
||
|------|--------|-------------|-------------|--------------|---------------|--------|----------------|---------|
|
||
| 1 | WiSARD | K=14, bleach=1, 276b | 1.37 px | 16.20 | 65.4% | ~276b active | fast | **SHIP IT** |
|
||
| 2 | WiSARD+Elig | k=12, trace=5 | 8.24* | — | 90.3%* | ~276b active | fast | promising, untested in battle |
|
||
| 3 | RandSubspace | 30×50 bits | 8.31* | — | 90.6%* | ~1500b | fast | similar to WiSARD, no battle test |
|
||
| 4 | Echo State Net | res=512 | 8.29* | — | 85.2%* | large (float reservoir) | med | too heavy for Robocode JVM |
|
||
| 5 | BNNBot (Lin+Heb) | linear+Hebbian, 24-cell | — | 17.61 | 61.9% | tiny | fast | current baseline |
|
||
| 6 | Tsetlin Machine | 500 cl, T=50, s=1.5, W=Y | 2.84† | 20.69 (warm: 18.77) | 55.8% | 531 KB | slow (cold-start) | future work |
|
||
| 7 | N-gram Markov | 4×69 chunks | 10.11* | — | 80.5%* | small | med | no improvement on WiSARD |
|
||
| 8 | Kanerva SDM | addr=2000 | 10.16* | — | 80.1%* | med | slow | no improvement on baseline |
|
||
| 9 | Bloom Filter | 8192 slots | 17.08* | — | 59.2%* | 8 KB | fast | worse than baseline |
|
||
| 10 | HDC/VSA | n_hd=2000, 32 cls | 65.33* | — | 0.0%* | large | slow | broken for this task |
|
||
|
||
*Measured on 277-row alternatives dataset (single enemy, p1.07).
|
||
†Measured on 2500-row TM sweep dataset (single enemy).
|
||
Battle column from 4-enemy head-to-head (73 380 rows, all powers).
|
||
|
||
---
|
||
|
||
## Per-Enemy Analysis
|
||
|
||
| Enemy | BNNBot MAE | WiSARD MAE | TsetlinBot MAE | Winner | Notes |
|
||
|-------|-----------|-----------|---------------|--------|-------|
|
||
| target | 2.68 | **2.21** | 4.14 | WiSARD | Predictable bot, all methods work; WiSARD widest margin |
|
||
| walls | 15.31 | **13.22** | 18.81 | WiSARD | Wall-bouncing pattern, WiSARD generalises better |
|
||
| crazy | 24.72 | **22.29** | 27.10 | WiSARD | Highly random; WiSARD still edges out |
|
||
| spinbot | **23.75** | 24.54 | 30.20 | BNNBot | Periodic rotation pattern; Hebbian table captures it |
|
||
|
||
**Specialist vs generalist:** WiSARD is the best generalist (3/4 wins). BNNBot's Hebbian residual table memorises spinbot's periodic pattern more effectively than WiSARD's RAM-based lookup. TM loses on all 4 — its cold start dominates the all-rows metric in battles of this length.
|
||
|
||
TM warm-phase on crazy (20.98) nearly matches WiSARD (22.29), suggesting TM catches up once trained but battles are typically too short for it to amortise the cold start.
|
||
|
||
---
|
||
|
||
## Key Insights
|
||
|
||
### WiSARD
|
||
|
||
- **K=14 is the sweet spot.** MAE plateaus K=12–20; diminishing returns above K=14. K=4–6 measurably worse.
|
||
- **Bleaching has marginal effect.** bleach=0 vs bleach=1 vs bleach=2 all within ±0.02 MAE. Skip tuning it; leave at 1.
|
||
- **XOR features add nothing.** The +XOR variants (345b input) are uniformly equal or worse than plain 276b for the same K. Extra bits, no gain.
|
||
- **Cold-start (MAE f50) still high** (~0.43–0.58 px for best configs), but last-50 converges to ~0.00 px — the model fully memorises the pattern after enough data.
|
||
- **Baseline delta:** best WiSARD is -0.26 px vs P1 linear, which sounds small but the absolute MAE of 1.37 vs 1.63 is a ~16% improvement at p1.07. Battle-scale improvement is larger (16.20 vs 17.61, ~8% reduction).
|
||
|
||
### Tsetlin Machine
|
||
|
||
- **Weighted clauses are mandatory.** Unweighted configs are consistently 0.5–2.0 MAE higher at identical clause counts.
|
||
- **s=1.5 is optimal.** Specificity 3.0 works, 5.0+ degrades, 10–15 is too sparse.
|
||
- **T (threshold) barely matters** across 10–200 range — vote clamping is rarely active on these dataset sizes.
|
||
- **Cold-start dominates all-rows MAE.** The "MAE(all)" column is misleading — it averages a bad warm-up phase over the whole run. The relevant metric for a real battle is MAE(last⅓): best config gets 0.27 px (!) for 500 clauses, 0.15 px for 100 clauses.
|
||
- **Practical config:** 100 clauses, T=100, s=1.5, states=32, weighted=Y. MAE(all)=3.28, MAE(last⅓)=0.20, 106 KB. The 500-clause config is marginally better but 5× the memory and slow.
|
||
- **Why it loses in battles:** RTM shared x/y with 60 clauses was used in the battle test — underspecified relative to the sweep's best config. Even so, TM warm-phase (18.77) still loses to WiSARD (16.20) overall.
|
||
|
||
### Alternatives
|
||
|
||
- **WiSARD+Eligibility Traces** is the real second-place finisher (MAE=8.24, Hit%=90.3% on the 277-row test). Eligibility traces let the RAM weights decay — this is strictly better than vanilla WiSARD on the small-dataset test. Not yet battle-tested.
|
||
- **Random Subspace ensemble** (8.31, 90.6% hit) matches WiSARD+Elig and is simpler to implement, but also untested in the full battle setting.
|
||
- **ESN** (8.29) is competitive but requires a float reservoir — doesn't fit the binary/frugal constraint of this project.
|
||
- **N-gram and SDM** essentially replicate linear baseline performance. Not worth the complexity.
|
||
- **Bloom Filter** — worse than baseline. RAM addressed by pattern ID without content-addressing fails here.
|
||
- **HDC/VSA** — completely broken (0% hit rate). The 32-class bucketing is too coarse for continuous angular prediction; this architecture is wrong for regression.
|
||
|
||
### Eligibility Traces Finding
|
||
|
||
The WiSARD+Elig variant (trace=5) improves MAE by ~1.7 px and hit rate by +10 pp vs vanilla WiSARD K=12 on the same 277-row dataset. This is the largest single improvement found across all alternatives. Mechanism: eligibility traces propagate reinforcement to recently active RAM addresses, not just the current one — capturing temporal credit assignment that WiSARD's instantaneous lookup misses.
|
||
|
||
---
|
||
|
||
## Recommendation
|
||
|
||
**Ship:** WiSARD K=14, bleach=1, 276-bit input (4 frames × 69 bits), 23 RAM nodes.
|
||
|
||
```
|
||
K = 14
|
||
bleach = 1
|
||
input_bits = 276 # 4 frames × 69 bits, no XOR
|
||
```
|
||
|
||
This is the best-validated config with a real battle backtest (WiSARD K=12 wins 3/4 enemies; K=14 is marginally better in the sweep). Memory is ~276 bits active per RAM node, negligible for the JVM.
|
||
|
||
**Explore next (in priority order):**
|
||
|
||
1. **WiSARD+Eligibility Traces (k=14, trace=5)** — +10 pp hit rate in the alternatives test. Straightforward to implement on top of current WiSARD. Highest expected ROI.
|
||
2. **Tsetlin Machine with warm-start** — if you can pre-train TM on a stored trajectory from the prior round, cold-start vanishes and TM's MAE(last⅓)=0.20 becomes competitive. Requires inter-round state persistence.
|
||
3. **Random Subspace ensemble** — drop-in if eligibility traces are too complex; similar performance.
|
||
4. **Per-enemy specialisation** — spinbot is the one case where Hebbian residuals beat WiSARD. A simple enemy-classifier + model switcher could capture that.
|
||
|
||
Do not pursue: HDC, Bloom Filter, SDM. None improved on linear.
|
||
|
||
---
|
||
|
||
## Raw Data References
|
||
|
||
| File | Content |
|
||
|------|---------|
|
||
| `analysis/sweep_wisard_results.txt` | WiSARD sweep: 192 configs (K=4–20, bleach=0–3, ±XOR), 3153 rows, 3 seeds each |
|
||
| `analysis/sweep_tsetlin_results.txt` | TM sweep: 57 configs (clauses=10–500, s=1.5–15, T=10–200, ±weighted), 2500 rows |
|
||
| `analysis/alternatives_results.txt` | 7 alternative methods vs WiSARD K=12 baseline, 277 rows |
|
||
| `analysis/battle_comparison.txt` | Head-to-head battle backtest: BNNBot vs WiSARD K=12 vs TsetlinBot, 4 enemies, 73 380 data points |
|