Files
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00

8.8 KiB
Raw Permalink Blame History

Gun Predictor Shootout — Final Report

Executive Summary

WiSARD K=14 (bleach=1, 276-bit input) wins the hyperparameter sweep with MAE=1.37px on the single-enemy multi-frame dataset. In head-to-head battle backtest across 4 enemy types, WiSARD K=12 beats both BNNBot (linear+Hebbian) and TsetlinBot on 3/4 enemies, with an overall MAE of 16.20 vs 17.61 (BNNBot) and 20.69 (TsetlinBot). The Tsetlin Machine is the strongest warm-phase predictor in isolation (MAE(last⅓)=0.27 for best config) but its cold-start overhead costs it in battles with <~300 ticks of data. Ship WiSARD K=14, bleach=1. If cold-start penalty on TM is ever fixed (e.g., pre-training or warm-up from WiSARD), TM becomes worth revisiting.


Methods Tested

16 method families, ~260 configurations total:

  • WiSARD (192 configs: K=4–20, bleach=0–3, ±XOR features)
  • Tsetlin Machine / RTM (57 configs: 10–500 clauses, s=1.5–15, T=10–200, ±weighted)
  • Echo State Network (reservoir=512)
  • Kanerva Sparse Distributed Memory (addr=2000)
  • N-gram Markov predictor (4×69 chunks)
  • Bloom Filter (8192 slots)
  • Hyperdimensional Computing / VSA (n_hd=2000, 32 classes)
  • Random Subspace ensemble (30×50 bits)
  • WiSARD + Eligibility Traces (k=12, trace=5)

Results Table

Dataset note: WiSARD sweep and alternatives used a 277–3153 row single-enemy dataset at rep power p1.07. Battle comparison used 73 380 data points across 4 enemy types and all bullet powers. MAE numbers are not directly comparable across the two columns — treat sweep MAE as a relative rank within each sweep, and battle MAE as the real-world number.

Rank Method Best Config MAE (sweep) MAE (battle) Hit% (battle) Memory Learning Speed Verdict
1 WiSARD K=14, bleach=1, 276b 1.37 px 16.20 65.4% ~276b active fast SHIP IT
2 WiSARD+Elig k=12, trace=5 8.24* — 90.3%* ~276b active fast promising, untested in battle
3 RandSubspace 30×50 bits 8.31* — 90.6%* ~1500b fast similar to WiSARD, no battle test
4 Echo State Net res=512 8.29* — 85.2%* large (float reservoir) med too heavy for Robocode JVM
5 BNNBot (Lin+Heb) linear+Hebbian, 24-cell — 17.61 61.9% tiny fast current baseline
6 Tsetlin Machine 500 cl, T=50, s=1.5, W=Y 2.84† 20.69 (warm: 18.77) 55.8% 531 KB slow (cold-start) future work
7 N-gram Markov 4×69 chunks 10.11* — 80.5%* small med no improvement on WiSARD
8 Kanerva SDM addr=2000 10.16* — 80.1%* med slow no improvement on baseline
9 Bloom Filter 8192 slots 17.08* — 59.2%* 8 KB fast worse than baseline
10 HDC/VSA n_hd=2000, 32 cls 65.33* — 0.0%* large slow broken for this task

*Measured on 277-row alternatives dataset (single enemy, p1.07).
†Measured on 2500-row TM sweep dataset (single enemy).
Battle column from 4-enemy head-to-head (73 380 rows, all powers).


Per-Enemy Analysis

Enemy BNNBot MAE WiSARD MAE TsetlinBot MAE Winner Notes
target 2.68 2.21 4.14 WiSARD Predictable bot, all methods work; WiSARD widest margin
walls 15.31 13.22 18.81 WiSARD Wall-bouncing pattern, WiSARD generalises better
crazy 24.72 22.29 27.10 WiSARD Highly random; WiSARD still edges out
spinbot 23.75 24.54 30.20 BNNBot Periodic rotation pattern; Hebbian table captures it

Specialist vs generalist: WiSARD is the best generalist (3/4 wins). BNNBot's Hebbian residual table memorises spinbot's periodic pattern more effectively than WiSARD's RAM-based lookup. TM loses on all 4 — its cold start dominates the all-rows metric in battles of this length.

TM warm-phase on crazy (20.98) nearly matches WiSARD (22.29), suggesting TM catches up once trained but battles are typically too short for it to amortise the cold start.


Key Insights

WiSARD

  • K=14 is the sweet spot. MAE plateaus K=12–20; diminishing returns above K=14. K=4–6 measurably worse.
  • Bleaching has marginal effect. bleach=0 vs bleach=1 vs bleach=2 all within ±0.02 MAE. Skip tuning it; leave at 1.
  • XOR features add nothing. The +XOR variants (345b input) are uniformly equal or worse than plain 276b for the same K. Extra bits, no gain.
  • Cold-start (MAE f50) still high (~0.43–0.58 px for best configs), but last-50 converges to ~0.00 px — the model fully memorises the pattern after enough data.
  • Baseline delta: best WiSARD is -0.26 px vs P1 linear, which sounds small but the absolute MAE of 1.37 vs 1.63 is a ~16% improvement at p1.07. Battle-scale improvement is larger (16.20 vs 17.61, ~8% reduction).

Tsetlin Machine

  • Weighted clauses are mandatory. Unweighted configs are consistently 0.5–2.0 MAE higher at identical clause counts.
  • s=1.5 is optimal. Specificity 3.0 works, 5.0+ degrades, 10–15 is too sparse.
  • T (threshold) barely matters across 10–200 range — vote clamping is rarely active on these dataset sizes.
  • Cold-start dominates all-rows MAE. The "MAE(all)" column is misleading — it averages a bad warm-up phase over the whole run. The relevant metric for a real battle is MAE(last⅓): best config gets 0.27 px (!) for 500 clauses, 0.15 px for 100 clauses.
  • Practical config: 100 clauses, T=100, s=1.5, states=32, weighted=Y. MAE(all)=3.28, MAE(last⅓)=0.20, 106 KB. The 500-clause config is marginally better but 5× the memory and slow.
  • Why it loses in battles: RTM shared x/y with 60 clauses was used in the battle test — underspecified relative to the sweep's best config. Even so, TM warm-phase (18.77) still loses to WiSARD (16.20) overall.

Alternatives

  • WiSARD+Eligibility Traces is the real second-place finisher (MAE=8.24, Hit%=90.3% on the 277-row test). Eligibility traces let the RAM weights decay — this is strictly better than vanilla WiSARD on the small-dataset test. Not yet battle-tested.
  • Random Subspace ensemble (8.31, 90.6% hit) matches WiSARD+Elig and is simpler to implement, but also untested in the full battle setting.
  • ESN (8.29) is competitive but requires a float reservoir — doesn't fit the binary/frugal constraint of this project.
  • N-gram and SDM essentially replicate linear baseline performance. Not worth the complexity.
  • Bloom Filter — worse than baseline. RAM addressed by pattern ID without content-addressing fails here.
  • HDC/VSA — completely broken (0% hit rate). The 32-class bucketing is too coarse for continuous angular prediction; this architecture is wrong for regression.

Eligibility Traces Finding

The WiSARD+Elig variant (trace=5) improves MAE by ~1.7 px and hit rate by +10 pp vs vanilla WiSARD K=12 on the same 277-row dataset. This is the largest single improvement found across all alternatives. Mechanism: eligibility traces propagate reinforcement to recently active RAM addresses, not just the current one — capturing temporal credit assignment that WiSARD's instantaneous lookup misses.


Recommendation

Ship: WiSARD K=14, bleach=1, 276-bit input (4 frames × 69 bits), 23 RAM nodes.

K         = 14
bleach    = 1
input_bits = 276  # 4 frames × 69 bits, no XOR

This is the best-validated config with a real battle backtest (WiSARD K=12 wins 3/4 enemies; K=14 is marginally better in the sweep). Memory is ~276 bits active per RAM node, negligible for the JVM.

Explore next (in priority order):

  1. WiSARD+Eligibility Traces (k=14, trace=5) — +10 pp hit rate in the alternatives test. Straightforward to implement on top of current WiSARD. Highest expected ROI.
  2. Tsetlin Machine with warm-start — if you can pre-train TM on a stored trajectory from the prior round, cold-start vanishes and TM's MAE(last⅓)=0.20 becomes competitive. Requires inter-round state persistence.
  3. Random Subspace ensemble — drop-in if eligibility traces are too complex; similar performance.
  4. Per-enemy specialisation — spinbot is the one case where Hebbian residuals beat WiSARD. A simple enemy-classifier + model switcher could capture that.

Do not pursue: HDC, Bloom Filter, SDM. None improved on linear.


Raw Data References

File Content
analysis/sweep_wisard_results.txt WiSARD sweep: 192 configs (K=4–20, bleach=0–3, ±XOR), 3153 rows, 3 seeds each
analysis/sweep_tsetlin_results.txt TM sweep: 57 configs (clauses=10–500, s=1.5–15, T=10–200, ±weighted), 2500 rows
analysis/alternatives_results.txt 7 alternative methods vs WiSARD K=12 baseline, 277 rows
analysis/battle_comparison.txt Head-to-head battle backtest: BNNBot vs WiSARD K=12 vs TsetlinBot, 4 enemies, 73 380 data points