Files
SirRoboGarage/BNNBot_garage/analysis/backtest_tsetlin_report.txt
T
SirStone 1ed7797cb6 feat(ModularBot): 6 guns, pattern matcher, melee modules, adversarial bots
- New guns: guess-factor (GF histogram), pattern-matcher (movement tape replay)
- New modules: minimum-risk melee movement, spinning melee radar
- New test bots: PatternMover, RandomMover, WaveSurfer
- Fixed: FeedbackEvent now carries actualX/actualY for proper GF learning
- Fixed: TM gun warmup gating + directional residuals
- Fixed: circular gun integrated formula + multi-bin omega cache
- Fixed: oscillator wall-bounce lockout
- Fixed: phantom meteor perpendicular body orientation
- 6/6 battle wins across all enemy types
2026-09-20 00:59:53 +02:00

105 lines
4.2 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
======================================================================
REGRESSION TSETLIN MACHINE BACKTEST REPORT
Rows: 277 Clauses: 60 States: 15 s=3.0 T=30
Input: 4 frames × 68 bits = 272 features (544 literals)
======================================================================
--- TM (Regression Tsetlin Machine) ---
Power MAE RMSE Max %<5u
------------------------------------------
p0.10 12.97 18.57 78.47 33.9%
p0.42 13.45 19.19 81.65 32.5%
p0.74 14.01 19.92 85.14 32.9%
p1.07 14.65 20.78 89.08 28.9%
p1.39 15.38 21.73 93.35 27.4%
p1.71 16.21 22.81 99.09 26.0%
p2.03 17.24 24.15 104.46 23.1%
p2.36 18.42 25.69 111.24 21.7%
p2.68 20.05 27.68 118.79 20.2%
p3.00 22.06 30.17 127.54 17.0%
--- P1 (Linear Extrapolation) ---
Power MAE RMSE Max %<5u
------------------------------------------
p0.10 8.10 11.93 61.09 49.8%
p0.42 8.68 12.70 64.35 49.1%
p0.74 9.37 13.59 67.94 46.6%
p1.07 10.18 14.62 72.04 40.1%
p1.39 11.03 15.72 76.48 38.6%
p1.71 11.97 16.96 81.92 35.7%
p2.03 13.22 18.49 87.57 32.5%
p2.36 14.64 20.24 94.46 29.2%
p2.68 16.46 22.43 102.17 26.4%
p3.00 18.70 25.16 111.15 23.5%
--- P3 (Linear + Hebbian, lr=0.1) ---
Power MAE RMSE Max %<5u
------------------------------------------
p0.10 6.43 9.49 41.84 54.9%
p0.42 6.92 10.09 44.94 52.3%
p0.74 7.57 10.85 48.39 49.5%
p1.07 8.34 11.76 52.47 45.8%
p1.39 9.15 12.77 56.79 39.7%
p1.71 10.14 13.95 61.80 35.0%
p2.03 11.37 15.43 67.33 32.1%
p2.36 12.85 17.18 74.19 29.6%
p2.68 14.72 19.42 81.71 27.8%
p3.00 16.98 22.23 90.49 26.7%
--- LEARNING CURVE (TM, p1.07) ---
First-50 MAE: 35.47
Last-50 MAE: 9.37
Improvement: +26.11 (converging)
--- TM vs BASELINES (avg MAE across all power levels) ---
TM avg MAE (all rows): 16.44
P1 avg MAE (all rows): 12.24 (TM delta: +4.21)
P3 avg MAE (all rows): 10.45 (TM delta: +6.00)
TM avg MAE (last 50): 12.45 (warm TM, best proxy for in-battle perf)
--- PER-POWER TM vs P1 ---
Power TM MAE P1 MAE P3 MAE vs P1 vs P3
----------------------------------------------------
p0.10 12.97 8.10 6.43 +4.87 +6.54
p0.42 13.45 8.68 6.92 +4.77 +6.53
p0.74 14.01 9.37 7.57 +4.63 +6.43
p1.07 14.65 10.18 8.34 +4.46 +6.30
p1.39 15.38 11.03 9.15 +4.35 +6.23
p1.71 16.21 11.97 10.14 +4.24 +6.07
p2.03 17.24 13.22 11.37 +4.02 +5.87
p2.36 18.42 14.64 12.85 +3.78 +5.57
p2.68 20.05 16.46 14.72 +3.59 +5.33
p3.00 22.06 18.70 16.98 +3.37 +5.08
--- ANALYSIS ---
The TM starts cold (zero residual prediction) and converges during the battle.
The all-rows MAE is dominated by early cold-start rows; last-50 MAE is the
better proxy for real in-battle performance after warm-up.
Key observations:
- Learning curve shows strong convergence: first-50 to last-50 MAE drops ~26 units.
- With only 277 rows, the TM sees 277 training steps total (shared RTM).
A real battle (~1000 wave hits) would give ~4x more training signal.
- The TM learns residual correction on top of linear extrapolation, not raw coords.
This is the same structure as P3 (Hebbian residual) but with a more expressive
non-linear function approximator.
- The residual target range is ±30 encoded units. If actual residuals
exceed this (they can for far enemies), the TM clips silently.
Increase RESID_MAX if coverage is needed.
Architectural finding: TM as raw-coordinate predictor fails badly on 277 rows
(MAE ~66). TM as residual corrector over linear extrapolation converges fast
and approaches P1/P3 performance in the warm phase. This matches how P3 works.
Next step: implement in Nim as a residual corrector replacing the Hebbian table,
using M=100 clauses, N_states=20, s=3.0. Expect to match or beat P3 after ~100
battle ticks with a 1000-tick battle.