The user's goal was "a TM gun that can learn fast and generalize better".
Swept offline over the real DrussGT fixtures (no live battles) by coordinate
descent, one lever at a time, with a SHUFFLED-FEEDBACK CONTROL - a TM trained
on randomised targets. That control is what settles the question.
Final confirmation, 4 seeds each (~74,600 first-100-tick bullets per config):
config EARLY(first 100) OVERALL
Shuf_w3 (RANDOM feedback) 23.9% 20.0%
win3_s1.1 (best real TM found) 23.7% 20.2%
Shuf_w10 (RANDOM feedback) 23.1% 20.0%
win3_st100 (prior job's edit) 23.0% 20.1%
win3_off (TM correction ~= 0) 22.7% 20.2%
def_w10 (shipped default) 22.1% 20.3%
Linear (deterministic reference) 34.0% 24.3%
The best real config beats the default early (23.7% vs 22.1%, non-overlapping
per-seed ranges, z=+7.34, p=2e-13) - but its own SHUFFLED control scores 23.9%,
i.e. HIGHER, z=-0.91, p=0.37. Random targets do at least as well. So the early
gain is not learning.
Per-lever screens were flat: TM_N_CLAUSES 25/50/100/200 all 23.0% early,
completely flat; TM_N_STATES 4/32/100 all ~22-23% (unstable across seeds);
TM_S mildly monotonic (lower better early); TM_T flat; TM_WINDOW_SIZE 2/3/10
all within noise of each other and of the shuffled control.
Two further findings:
- The TM-off ablation (correction ~= 0) scores 22.7%/20.2%, essentially the
same as TM-on. The TM's correction is near-zero-mean noise; the gun's
one-shot internal linear baseline accounts for its accuracy.
- The TM gun is 10.3 pp behind Linear early and 4.1 pp behind overall. That
deficit is in the BASELINE MODEL (LinearGun iterates flight time; this gun
does not), not in the TM hyper-parameters. Tuning knobs cannot close it.
Conclusion: do not tune TM hyper-parameters further. Either the input
representation or the prediction target is what needs to change - the shuffled
control shows the TM is not extracting target information beyond its baseline.
Defaults left UNCHANGED (window=10/states=32/S=1.5/T=25/clauses=50); an
uncommitted prior edit (window=3/states=100) was reverted as unsupported.
Hyper-parameters are now compile-time overridable (-d:TM_WINDOW_SIZE=3 etc.)
so future sweeps need no gun edit.
NOT MEASURED: real hit rate vs DrussGT (offline only by design). The repo's own
docs/gun_rack_analysis.md 2 reports the virtual metric is a sign-unstable ranker
of real hit rate, so the comparison against "Linear 10.7% real" is not direct -
whether the TM is competitive live is INFERRED-unknown, not measured.
Guards: test_gun_harness 39/39, test_vbullet_metric, test_power_selection,
test_tsetlin_gun, test_tm_pattern_learning all green.