# Gun Predictor Shootout — Final Report ## Executive Summary WiSARD K=14 (bleach=1, 276-bit input) wins the hyperparameter sweep with MAE=1.37px on the single-enemy multi-frame dataset. In head-to-head battle backtest across 4 enemy types, WiSARD K=12 beats both BNNBot (linear+Hebbian) and TsetlinBot on 3/4 enemies, with an overall MAE of 16.20 vs 17.61 (BNNBot) and 20.69 (TsetlinBot). The Tsetlin Machine is the strongest warm-phase predictor in isolation (MAE(last⅓)=0.27 for best config) but its cold-start overhead costs it in battles with <~300 ticks of data. Ship WiSARD K=14, bleach=1. If cold-start penalty on TM is ever fixed (e.g., pre-training or warm-up from WiSARD), TM becomes worth revisiting. --- ## Methods Tested **16 method families, ~260 configurations total:** - WiSARD (192 configs: K=4–20, bleach=0–3, ±XOR features) - Tsetlin Machine / RTM (57 configs: 10–500 clauses, s=1.5–15, T=10–200, ±weighted) - Echo State Network (reservoir=512) - Kanerva Sparse Distributed Memory (addr=2000) - N-gram Markov predictor (4×69 chunks) - Bloom Filter (8192 slots) - Hyperdimensional Computing / VSA (n_hd=2000, 32 classes) - Random Subspace ensemble (30×50 bits) - WiSARD + Eligibility Traces (k=12, trace=5) --- ## Results Table > **Dataset note:** WiSARD sweep and alternatives used a 277–3153 row single-enemy dataset at rep power p1.07. Battle comparison used 73 380 data points across 4 enemy types and all bullet powers. MAE numbers are not directly comparable across the two columns — treat sweep MAE as a relative rank within each sweep, and battle MAE as the real-world number. | Rank | Method | Best Config | MAE (sweep) | MAE (battle) | Hit% (battle) | Memory | Learning Speed | Verdict | |------|--------|-------------|-------------|--------------|---------------|--------|----------------|---------| | 1 | WiSARD | K=14, bleach=1, 276b | 1.37 px | 16.20 | 65.4% | ~276b active | fast | **SHIP IT** | | 2 | WiSARD+Elig | k=12, trace=5 | 8.24* | — | 90.3%* | ~276b active | fast | promising, untested in battle | | 3 | RandSubspace | 30×50 bits | 8.31* | — | 90.6%* | ~1500b | fast | similar to WiSARD, no battle test | | 4 | Echo State Net | res=512 | 8.29* | — | 85.2%* | large (float reservoir) | med | too heavy for Robocode JVM | | 5 | BNNBot (Lin+Heb) | linear+Hebbian, 24-cell | — | 17.61 | 61.9% | tiny | fast | current baseline | | 6 | Tsetlin Machine | 500 cl, T=50, s=1.5, W=Y | 2.84† | 20.69 (warm: 18.77) | 55.8% | 531 KB | slow (cold-start) | future work | | 7 | N-gram Markov | 4×69 chunks | 10.11* | — | 80.5%* | small | med | no improvement on WiSARD | | 8 | Kanerva SDM | addr=2000 | 10.16* | — | 80.1%* | med | slow | no improvement on baseline | | 9 | Bloom Filter | 8192 slots | 17.08* | — | 59.2%* | 8 KB | fast | worse than baseline | | 10 | HDC/VSA | n_hd=2000, 32 cls | 65.33* | — | 0.0%* | large | slow | broken for this task | *Measured on 277-row alternatives dataset (single enemy, p1.07). †Measured on 2500-row TM sweep dataset (single enemy). Battle column from 4-enemy head-to-head (73 380 rows, all powers). --- ## Per-Enemy Analysis | Enemy | BNNBot MAE | WiSARD MAE | TsetlinBot MAE | Winner | Notes | |-------|-----------|-----------|---------------|--------|-------| | target | 2.68 | **2.21** | 4.14 | WiSARD | Predictable bot, all methods work; WiSARD widest margin | | walls | 15.31 | **13.22** | 18.81 | WiSARD | Wall-bouncing pattern, WiSARD generalises better | | crazy | 24.72 | **22.29** | 27.10 | WiSARD | Highly random; WiSARD still edges out | | spinbot | **23.75** | 24.54 | 30.20 | BNNBot | Periodic rotation pattern; Hebbian table captures it | **Specialist vs generalist:** WiSARD is the best generalist (3/4 wins). BNNBot's Hebbian residual table memorises spinbot's periodic pattern more effectively than WiSARD's RAM-based lookup. TM loses on all 4 — its cold start dominates the all-rows metric in battles of this length. TM warm-phase on crazy (20.98) nearly matches WiSARD (22.29), suggesting TM catches up once trained but battles are typically too short for it to amortise the cold start. --- ## Key Insights ### WiSARD - **K=14 is the sweet spot.** MAE plateaus K=12–20; diminishing returns above K=14. K=4–6 measurably worse. - **Bleaching has marginal effect.** bleach=0 vs bleach=1 vs bleach=2 all within ±0.02 MAE. Skip tuning it; leave at 1. - **XOR features add nothing.** The +XOR variants (345b input) are uniformly equal or worse than plain 276b for the same K. Extra bits, no gain. - **Cold-start (MAE f50) still high** (~0.43–0.58 px for best configs), but last-50 converges to ~0.00 px — the model fully memorises the pattern after enough data. - **Baseline delta:** best WiSARD is -0.26 px vs P1 linear, which sounds small but the absolute MAE of 1.37 vs 1.63 is a ~16% improvement at p1.07. Battle-scale improvement is larger (16.20 vs 17.61, ~8% reduction). ### Tsetlin Machine - **Weighted clauses are mandatory.** Unweighted configs are consistently 0.5–2.0 MAE higher at identical clause counts. - **s=1.5 is optimal.** Specificity 3.0 works, 5.0+ degrades, 10–15 is too sparse. - **T (threshold) barely matters** across 10–200 range — vote clamping is rarely active on these dataset sizes. - **Cold-start dominates all-rows MAE.** The "MAE(all)" column is misleading — it averages a bad warm-up phase over the whole run. The relevant metric for a real battle is MAE(last⅓): best config gets 0.27 px (!) for 500 clauses, 0.15 px for 100 clauses. - **Practical config:** 100 clauses, T=100, s=1.5, states=32, weighted=Y. MAE(all)=3.28, MAE(last⅓)=0.20, 106 KB. The 500-clause config is marginally better but 5× the memory and slow. - **Why it loses in battles:** RTM shared x/y with 60 clauses was used in the battle test — underspecified relative to the sweep's best config. Even so, TM warm-phase (18.77) still loses to WiSARD (16.20) overall. ### Alternatives - **WiSARD+Eligibility Traces** is the real second-place finisher (MAE=8.24, Hit%=90.3% on the 277-row test). Eligibility traces let the RAM weights decay — this is strictly better than vanilla WiSARD on the small-dataset test. Not yet battle-tested. - **Random Subspace ensemble** (8.31, 90.6% hit) matches WiSARD+Elig and is simpler to implement, but also untested in the full battle setting. - **ESN** (8.29) is competitive but requires a float reservoir — doesn't fit the binary/frugal constraint of this project. - **N-gram and SDM** essentially replicate linear baseline performance. Not worth the complexity. - **Bloom Filter** — worse than baseline. RAM addressed by pattern ID without content-addressing fails here. - **HDC/VSA** — completely broken (0% hit rate). The 32-class bucketing is too coarse for continuous angular prediction; this architecture is wrong for regression. ### Eligibility Traces Finding The WiSARD+Elig variant (trace=5) improves MAE by ~1.7 px and hit rate by +10 pp vs vanilla WiSARD K=12 on the same 277-row dataset. This is the largest single improvement found across all alternatives. Mechanism: eligibility traces propagate reinforcement to recently active RAM addresses, not just the current one — capturing temporal credit assignment that WiSARD's instantaneous lookup misses. --- ## Recommendation **Ship:** WiSARD K=14, bleach=1, 276-bit input (4 frames × 69 bits), 23 RAM nodes. ``` K = 14 bleach = 1 input_bits = 276 # 4 frames × 69 bits, no XOR ``` This is the best-validated config with a real battle backtest (WiSARD K=12 wins 3/4 enemies; K=14 is marginally better in the sweep). Memory is ~276 bits active per RAM node, negligible for the JVM. **Explore next (in priority order):** 1. **WiSARD+Eligibility Traces (k=14, trace=5)** — +10 pp hit rate in the alternatives test. Straightforward to implement on top of current WiSARD. Highest expected ROI. 2. **Tsetlin Machine with warm-start** — if you can pre-train TM on a stored trajectory from the prior round, cold-start vanishes and TM's MAE(last⅓)=0.20 becomes competitive. Requires inter-round state persistence. 3. **Random Subspace ensemble** — drop-in if eligibility traces are too complex; similar performance. 4. **Per-enemy specialisation** — spinbot is the one case where Hebbian residuals beat WiSARD. A simple enemy-classifier + model switcher could capture that. Do not pursue: HDC, Bloom Filter, SDM. None improved on linear. --- ## Raw Data References | File | Content | |------|---------| | `analysis/sweep_wisard_results.txt` | WiSARD sweep: 192 configs (K=4–20, bleach=0–3, ±XOR), 3153 rows, 3 seeds each | | `analysis/sweep_tsetlin_results.txt` | TM sweep: 57 configs (clauses=10–500, s=1.5–15, T=10–200, ±weighted), 2500 rows | | `analysis/alternatives_results.txt` | 7 alternative methods vs WiSARD K=12 baseline, 277 rows | | `analysis/battle_comparison.txt` | Head-to-head battle backtest: BNNBot vs WiSARD K=12 vs TsetlinBot, 4 enemies, 73 380 data points |