Files
SirRoboGarage/common_libs/tests/tm_pattern_sweep_results.md
T
SirStone ca82053a11 TM gun: the discrete-target diagnosis was RIGHT - it learns now. Still loses to Linear.
The user's goal: a TM gun that is the best 1v1 gun, starting from scratch every
battle but quickly overfitting the current enemy. The previous attempt (knob
tuning) failed: NO configuration beat its own shuffled-feedback control, and the
TM-off ablation scored the same as TM-on, i.e. the TM's correction was
near-zero-mean noise. Diagnosis then: a Tsetlin Machine is a CLASSIFIER, and we
were asking it for an absolute aim point - a regression target. So this attempt
gave it a DISCRETE target (multi-class over guess-factor buckets) with 40
binary/bucketed motion features, and measured it against Linear, the default
Tsetlin gun, and a MANDATORY shuffled control.

THE DIAGNOSIS IS CONFIRMED - THE TM LEARNS, DECISIVELY:
  online class accuracy     46.0%  vs shuffled control 20.0%   (2.3x chance)
  raw ungated argmax        21.2%/18.6% vs shuffled 15.2%/8.3% (18/18, p<0.0001)
  TMPattern > its shuffled control, overall   17/1 runs, p=0.0001
Compare the previous attempt, which could not beat shuffled feedback at all.
TMPattern also beats the default Tsetlin gun early (17/1, p=0.0001), so it is a
strictly better TM gun than the one in the rack.

BUT IT IS NOT COMPETITIVE WITH LINEAR ON REAL SURFERS:
  real DrussGT, bmPath (the shipped metric), 3 seeds, pooled early/overall
    Linear            34.0% (6358/18715)    24.3% (58297/239943)
    TMPattern (gated) 27.9% (15514/55535)   22.0% (158658/719681)
    TMPatternShuf     28.7%                 19.4%
  Linear > TMPattern: 15/18 early p=0.0075, 15/18 overall p=0.0075
  bmPoint: neutral (7.2%/4.6% vs Linear 7.2%/4.7%)
  synthetic controlled motion: matches/edges Linear (66.8%/60.6% vs 66.4%/59.6%,
    shuffled 55.7%/50.1%) - the mechanism works when motion is predictable.

So: the representation fix moved this from "learns nothing" to "learns strongly
but applies its knowledge badly". INFERRED reason for the residual loss: the
linear lead is already the modal GF bucket (the label histogram is centred), so
corrective excursions away from it are net-negative. The measured deficit lives
in the BASELINE and in RANGE, not in the TM knobs - which is why further knob
tuning was never going to work.

Best config: gated hard K=5, TM_CONF_MARGIN=0.25, TM_SHRINK=0.5.
NOT TRIED (time-boxed): the binary-reversal target, and a RADIAL (range-holding)
target - the latter is the top next step.

Adds `common_libs/guns/tm_pattern.nim` (NOT registered in the rack),
`common_libs/tests/sweep_tm_pattern.nim`, and a durable writeup at
`common_libs/tests/tm_pattern_sweep_results.md`.
2026-09-22 00:56:01 +02:00

8.9 KiB
Raw Blame History

TM pattern gun — discrete-target sweep results

Date: 2026-09-21. Author: background worker (executor-heavy). Artifacts implementing this: common_libs/guns/tm_pattern.nim, common_libs/tests/sweep_tm_pattern.nim. Do not commit.

What was built

tm_pattern.nim is a NEW gun (the old guns/tsetlin.nim is untouched). It attacks both the REPRESENTATION and the TARGET as the brief asked:

  • Base: forecastLinear (the exact self-consistent forecast LinearGun uses). GF class 0 (centre) reproduces the Linear gun byte-for-byte, so any measured difference is attributable to the TM.
  • Target: a discrete multi-class GUESS-FACTOR BUCKET — which lateral escape sector (in max-escape-angle units) the enemy occupied at the tick the bullet would have reached the BASE fire distance. 5 or 9 classes.
  • Label: read from a per-tick ring of our own recorded enemy positions at the base arrival tick, NOT from FeedbackEvent.actualXY. Under the shipped bmPath metric actualXY is the closest-approach point on the gun's OWN aim ray, which biases the label toward the gun's own last output; the ring gives a clean, metric-independent label.
  • Features: 40 hand-built binary/bucketed motion features (lateral-velocity sign over 3 ticks, turn-rate sign over 3 ticks, time since reversal, lateral magnitude, speed/distance/flight-time bands, four per-wall proximity bits, radial-fraction band, energy band, heading relative to LOS, approach sign).
  • TM core: compact self-contained Granmo Table 2/3 with the corrected feedback rules and Eq. 6 empty-clause bootstrap (same corrected core as tsetlin.nim / tm_selector.nim, re-derived at 40-bit width).
  • Per-enemy / freshness: a fresh net per gun instance; the net and history reset if the target id changes. Each offline round is replayed with a fresh instance (cold every battle, overfit within the battle).

Config overrides used in the final run: -d:TM_CONF_MARGIN_DEF=0.25 -d:TM_SHRINK_DEF=0.5 (confidence gate + shrink). Defaults are 0.0 / 1.0 (= raw argmax). Compile-time knobs: TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S_DEF, TM_MIN_OBS, TM_CONF_MARGIN_DEF, TM_SHRINK_DEF, TM_GF_MODE (hard|soft), TM_SOFT_BETA_DEF.

How to reproduce

nim c --path:common_libs -d:release \
    -d:TM_CONF_MARGIN_DEF=0.25 -d:TM_SHRINK_DEF=0.5 \
    -o:/tmp/sweep_tm_pattern common_libs/tests/sweep_tm_pattern.nim
/tmp/sweep_tm_pattern --set=real --seeds=3 --metric=path \
    --variants=linear,tsetlin,tmpat,tmpat_shuf
/tmp/sweep_tm_pattern --set=real --seeds=3 --metric=point \
    --variants=linear,tsetlin,tmpat,tmpat_shuf

Raw outputs: /tmp/final_path_s3.txt, /tmp/final_point_s3.txt, /tmp/final_ungated_path_s3.txt, /tmp/syn_*.

Metric

EARLY = resolutions in the first 100 ticks of each round (a cold TM every round). OVERALL = whole fixture. Pooled over all rounds / fixtures / seeds. TMPatternShuf = identical gun/encoding/cadence but the training label is a uniform-random class (the mandatory shuffled-feedback control). Per-run = one fixture × one seed (Linear is deterministic and replicated across seeds for pairing). Significance = exact two-sided paired sign test, 18 pairs.

The core result — the discrete target IS learnable, but does not beat the base

Online classification accuracy of the GF bucket (warm predictions only, seeds=1, n ≈ 1.26 M for each arm):

arm correct/total accuracy
TMPattern (real labels) 578722/1258488 46.0%
TMPatternShuf (random labels) 246733/1231116 20.0% (chance)

So the Tsetlin Machine genuinely learns the discrete target (2.3× chance). The representation mismatch was real and is fixed. The problem is that the target is not aligned with what wins the metric.

Real DrussGT fixtures, bmPath (shipped), seeds=3

variant early overall
Linear 34.0% (6358/18715) 24.3% (58297/239943)
Tsetlin (default) 22.3% (12344/55463) 20.3% (145828/719205)
TMPattern (gated) 27.9% (15514/55535) 22.0% (158658/719681)
TMPatternShuf 28.7% (16049/55969) 19.4% (139639/719790)

Paired sign tests (18 runs; ranges overlap, so the paired test is the test):

  • Linear > TMPattern: early 15/18 p=0.0075; overall 15/18 p=0.0075. Significantly worse than Linear.
  • TMPattern > TMPatternShuf: early 10/8 p=0.81 (tie); overall 17/1 p=0.0001. Learning is real but shows up mainly in the whole-round aggregate, not early.
  • TMPattern > Tsetlin: early 17/1 p=0.0001; overall 12/6 p=0.24. Beats the default TM gun early, ties overall.

Per-run distributions (mean [min,max], 18 runs): Linear early 35.52 [26.03,50.55] / overall 26.86 [9.79,43.36]; Tsetlin 25.76 [19.79,47.68] / 21.99 [9.64,31.34]; TMPattern 31.18 [20.17,55.11] / 24.45 [11.24,37.78]; Shuf 31.42 [20.72,53.49] / 21.07 [9.55,32.65].

Raw ungated hard argmax (margin 0.0, shrink 1.0), bmPath, seeds=3

variant early overall
Linear 34.0% 24.3%
Tsetlin 22.3% 20.3%
TMPattern 21.2% (11783/55664) 18.6% (133465/719432)
TMPatternShuf 15.2% (8710/57169) 8.3% (59903/720583)

TMPattern > Shuf 18/18 p<0.0001 on BOTH early and overall; TMPattern < Linear 3/15 p=0.0075 on both. The raw classifier is a clear, decisive learner and a clear loser to the Linear base: applying an argmax GF bucket costs ~13 pp early.

Real DrussGT fixtures, bmPoint, seeds=3 (gated)

variant early overall
Linear 7.2% (1480/20498) 4.7% (11277/241423)
Tsetlin 7.0% (4403/62617) 4.8% (34588/724717)
TMPattern 7.2% (4441/61842) 4.6% (33341/724556)
TMPatternShuf 6.6% (4112/61979) 3.4% (24847/724655)

Online accuracy 50.8%. TMPattern is statistically indistinguishable from Linear here (early per-run mean 13.67 vs 13.61; overall 5.70 vs 5.77) and beats its control on overall — i.e. on the arrival-time metric the correction is neutral, not harmful.

Synthetic fixtures (known rules) — the mechanism works when motion is predictable

bmPath, seeds=1, soft readout K=9: Linear early 76.3% / overall 71.4%; TMPattern early 76.6% / overall 71.8%; Shuf early 76.6% / overall 67.8%. Per-fixture gains vs Linear: wall-bounce 567 vs 537, energy-threshold-turner 332 vs 319; loss: constant-velocity 417 vs 431.

bmPoint, seeds=1, gated hard K=5: Linear 66.4% / 59.6%; TMPattern 66.8% / 60.6%; Shuf 55.7% / 50.1%. Energy-threshold-turner 268 vs 212, wall-bounce 585 vs 573.

Verdict

  • Learning: YES, decisively. The discrete-target TM predicts the GF bucket far above chance (46% vs 20%) and beats its shuffled control (ungated 18/18, p<0.0001). The "regression is a TM mismatch" diagnosis was correct.
  • Beats Linear: NO on the real surfers under bmPath (significantly worse, p=0.0075). Neutral under bmPoint. Matches/slightly beats Linear only on synthetic motion whose future is genuinely predictable.
  • Best configuration found: gated hard K=5, TM_CONF_MARGIN=0.25, TM_SHRINK=0.5 → 27.9% early / 22.0% overall (bmPath, real), +5.6 pp early / +1.7 pp overall vs the default TM gun, but 6.1 pp early / 2.3 pp overall behind Linear.

MEASURED vs INFERRED

MEASURED: every number in the tables above (pooled hits/shots, per-run distributions, paired sign tests, online classification accuracies). The position-ring label is our own recorded history at the base arrival tick; the shuffled control replaces only the label class with a uniform random draw.

INFERRED: that the residual loss on real surfers is because the linear lead is already the modal GF (label histogram is centred: real labels [3.8,6.6,17.6,6.6,3.5]×10⁵ for 5 classes) and the enemy's per-tick lateral reversal sign is not predictable enough from the 40 context bits to make a corrective excursion net-positive. Not directly measured.

What to try next (not done, time-boxed out)

  1. Radial target instead of angular. forecastRadialBlend work showed the dominant surfer error is range-holding (radial), not angle. A TM classifier over a RADIAL displacement bucket applied as an aim-distance correction targets the error the base actually has room to fix, and should matter most under bmPoint.
  2. Binary reversal with a two-candidate aim (brief candidate #1, unimplemented): predict "will the enemy reverse lateral direction before arrival?" and choose between the linear lead and a reversed lead. Same GF family, but a 2-class target is far more data-efficient; expected neutral given the GF result.
  3. Condition a genuinely weaker base. The measured wall says the deficit is the baseline; the Linear base leaves the TM no headroom. Feeding the TM the residual of forecastRadialBlend (a base that is worse on straight-liners but range-correct on surfers) is where a learned correction could plausibly pay.
  4. Richer context. 46% accuracy leaves room; the current context lacks the enemy's own recent GF history / segmentation that KNN/DecayGF exploit.