Tsetlin gun: NO configuration adapts faster than random feedback

The user's goal was "a TM gun that can learn fast and generalize better".
Swept offline over the real DrussGT fixtures (no live battles) by coordinate
descent, one lever at a time, with a SHUFFLED-FEEDBACK CONTROL - a TM trained
on randomised targets. That control is what settles the question.

Final confirmation, 4 seeds each (~74,600 first-100-tick bullets per config):

  config                          EARLY(first 100)   OVERALL
  Shuf_w3  (RANDOM feedback)          23.9%           20.0%
  win3_s1.1 (best real TM found)      23.7%           20.2%
  Shuf_w10 (RANDOM feedback)          23.1%           20.0%
  win3_st100 (prior job's edit)       23.0%           20.1%
  win3_off (TM correction ~= 0)       22.7%           20.2%
  def_w10  (shipped default)          22.1%           20.3%
  Linear (deterministic reference)    34.0%           24.3%

The best real config beats the default early (23.7% vs 22.1%, non-overlapping
per-seed ranges, z=+7.34, p=2e-13) - but its own SHUFFLED control scores 23.9%,
i.e. HIGHER, z=-0.91, p=0.37. Random targets do at least as well. So the early
gain is not learning.

Per-lever screens were flat: TM_N_CLAUSES 25/50/100/200 all 23.0% early,
completely flat; TM_N_STATES 4/32/100 all ~22-23% (unstable across seeds);
TM_S mildly monotonic (lower better early); TM_T flat; TM_WINDOW_SIZE 2/3/10
all within noise of each other and of the shuffled control.

Two further findings:
- The TM-off ablation (correction ~= 0) scores 22.7%/20.2%, essentially the
  same as TM-on. The TM's correction is near-zero-mean noise; the gun's
  one-shot internal linear baseline accounts for its accuracy.
- The TM gun is 10.3 pp behind Linear early and 4.1 pp behind overall. That
  deficit is in the BASELINE MODEL (LinearGun iterates flight time; this gun
  does not), not in the TM hyper-parameters. Tuning knobs cannot close it.

Conclusion: do not tune TM hyper-parameters further. Either the input
representation or the prediction target is what needs to change - the shuffled
control shows the TM is not extracting target information beyond its baseline.

Defaults left UNCHANGED (window=10/states=32/S=1.5/T=25/clauses=50); an
uncommitted prior edit (window=3/states=100) was reverted as unsupported.
Hyper-parameters are now compile-time overridable (-d:TM_WINDOW_SIZE=3 etc.)
so future sweeps need no gun edit.

NOT MEASURED: real hit rate vs DrussGT (offline only by design). The repo's own
docs/gun_rack_analysis.md 2 reports the virtual metric is a sign-unstable ranker
of real hit rate, so the comparison against "Linear 10.7% real" is not direct -
whether the TM is competitive live is INFERRED-unknown, not measured.

Guards: test_gun_harness 39/39, test_vbullet_metric, test_power_selection,
test_tsetlin_gun, test_tm_pattern_learning all green.
This commit is contained in:
2026-09-21 22:45:28 +02:00
parent fb36a0a685
commit 07f6f3af3f
2 changed files with 304 additions and 10 deletions
+27 -10
View File
@@ -2,16 +2,31 @@
## Self-contained: includes binary encoding and TM predictor inline.
## Implements Gun interface: predict(state, bulletSpeed) → GunPrediction, onResult(FeedbackEvent).
import std/[math, random, strformat]
import std/[math, random, strformat, strutils]
import gun_harness/gun_interface
import gun_harness/virtual_bullets as vb # PowerBins (power-bin count for trace keys)
# ── Binary encoding (adapted from BNNBot_garage/src/binary_encoding.nim) ─────
# Tuning note (adaptation-speed sweep, issue #184 follow-up): the hyper-params
# below are the pre-sweep baseline and are LEFT UNCHANGED. A one-lever-at-a-time
# offline sweep (common_libs/tests/sweep_tsetlin.nim over the real DrussGT
# fixtures, 6 fixtures x 4 seeds, ~75k early-round virtual bullets per config)
# found NO configuration that beat this baseline by more than a random-feedback
# control did: every Tsetlin variant sat at 22-24% first-100-tick virtual hit
# rate and ~20% whole-round, indistinguishable from the same TM trained on
# shuffled (random) feedback. The gun's correction is therefore not learning
# anything useful on a real surfer, so changing these constants would be noise.
# See the sweep report for the tables.
#
# Override knobs (all optional) so the sweep can be re-run without editing the
# source:
# -d:TM_WINDOW_SIZE=2 -d:TM_N_STATES=100 -d:TM_N_CLAUSES=100
# -d:TM_S_DEF=2.0 -d:TM_T_DEF=50 (empty TM_T_DEF -> TM_HALF)
const
TM_FRAME_BITS* = 83
TM_SELF_BITS* = 40
TM_WINDOW_SIZE* = 10
TM_WINDOW_SIZE* {.intdefine.} = 10
TM_TOTAL_BITS* = TM_FRAME_BITS * TM_WINDOW_SIZE + TM_SELF_BITS # 870
TM_MAX_DISTANCE = 1414.0 # diagonal of 1000x1000 arena
@@ -95,23 +110,25 @@ proc tmEncodeFullVector*(window: array[TM_WINDOW_SIZE, TmFrameEncoded],
# ── Tsetlin Machine (adapted from BNNBot_garage/src/tsetlin_predictor.nim) ───
const
TM_N_IN* = TM_TOTAL_BITS # 870
TM_N_IN* = TM_TOTAL_BITS # TM_FRAME_BITS*window + TM_SELF_BITS
TM_N_OUT* = 2 # cx, cy pixel corrections
TM_N_LITERALS* = TM_N_IN * 2 # 1740
TM_N_CLAUSES* = 50 # per output; issue #184 default
TM_N_LITERALS* = TM_N_IN * 2
TM_N_CLAUSES* {.intdefine.} = 50 # per output; issue #184 default
TM_HALF* = TM_N_CLAUSES div 2
TM_N_STATES* = 32 # automaton range [-32..32]
TM_T* = float(TM_HALF) # vote clamped to [-T, T]
TM_S* = 1.5 # specificity
TM_N_STATES* {.intdefine.} = 32 # automaton range [-N_STATES..N_STATES]
TM_S_DEF {.strdefine.} = "1.5" # specificity (TM_S=1.0 is degenerate)
TM_T_DEF {.strdefine.} = "" # empty -> float(TM_HALF)
TM_T* = (if TM_T_DEF.len > 0: parseFloat(TM_T_DEF) else: float(TM_HALF))
TM_S* = parseFloat(TM_S_DEF)
TM_RESID_MAX = 80.0 # pixel correction range
# ponytail: TM_N_STATES=32 needs int16 (int8 only fits ≤127, fine here); raise N_CLAUSES if underfitting
# ponytail: TM_N_STATES must fit int16 (fine up to 32767); raise N_CLAUSES if underfitting
type
TmClauseCache* = array[TM_N_OUT * TM_N_CLAUSES, uint8]
TmNet* = object
states*: array[TM_N_OUT * TM_N_CLAUSES * TM_N_LITERALS, int16]
# ponytail: int16 to safely hold [-32..32]; TM_N_STATES=32 fits int8 too but int16 is safer
# ponytail: int16 to safely hold [-TM_N_STATES..TM_N_STATES]
proc tmStateIdx*(outIdx, clause, lit: int): int {.inline.} =
(outIdx * TM_N_CLAUSES + clause) * TM_N_LITERALS + lit