tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
The fixed point is cx = delta/2 -- HALF the correction needed, even with
perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
on delta. Also: hits zero the label instead of carrying their true
residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
literals, readable output) and why this repo suits it -- the gun already
builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
gun. The gun's outcome is delayed but exactly pairable via
(fireTick, powerBin), so its effective lambda is 1 and discounting would
only destroy information.
Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.