Files
SirRoboGarage/docs/research/tm-deb-assessment.md
T
SirStone 54e9757567 docs(research): TM learning tracks + note that the TM-DEB source was deleted
tm-learning-tracks.md covers three things, all marked [FACT]/[INFERENCE]/
[UNKNOWN]:
- Section A: the Tsetlin gun's label is measured against the wrong baseline.
  predX = linearX + cx, so rx = actual - predX = delta - cx, and inside
  tmLearnOne error = residual - predicted = (delta - cx) - cx = delta - 2cx.
  The fixed point is cx = delta/2 -- HALF the correction needed, even with
  perfect Granmo feedback. Fix: store linearX/linearY in TmTrace and train
  on delta. Also: hits zero the label instead of carrying their true
  residual, and the per-clause step is magnitude-blind.
- Section B: what a TM is actually good at (AND-clauses over binary
  literals, readable output) and why this repo suits it -- the gun already
  builds an 83-bit x 10-frame Gray-coded window (870 bits). Includes a
  falsifiable known-rule benchmark proposal.
- Section C: delayed-reward learning belongs to the MOVEMENT layer, not the
  gun. The gun's outcome is delayed but exactly pairable via
  (fireTick, powerBin), so its effective lambda is 1 and discounting would
  only destroy information.

Also records that docs/papers/tm-deb-paper.pdf was deleted by the user as
AI-generated and unverifiable, while the Granmo-based feedback diff in
tm-deb-assessment.md stands on its own.
2026-09-20 23:28:31 +02:00

272 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TM-DEB Paper Assessment — Adopt as a fix for the Tsetlin gun?
> **Editorial note — source document deleted.** The paper this file reviews,
> `docs/papers/tm-deb-paper.pdf`, was **deleted by the user** on the grounds that
> it is AI-generated and unverifiable (the concrete evidence is catalogued in §2:
> self-identifying "Generated by Gemini Notebook" provenance, a misattributed
> reference, an algorithm core rendered as literal `%`, and "Expected" rather
> than measured results). This assessment file is **kept** because its actionable
> content — the Granmo Table 2/3 feedback diff (§6) and the prioritised fix list
> (§6.5) — stands on its own from the *verified* sources
> (`docs/papers/1804.01508v15.pdf` and `common_libs/guns/tsetlin.nim`) and does
> not depend on the deleted document. Citations to the deleted PDF are retained
> only as the record of the review that was performed.
**Issue:** N/A (read-only review requested by orchestrator)
**Date:** 2026-09-20
**Paper under review:** `docs/papers/tm-deb-paper.pdf` — "Temporal Credit Assignment in Tsetlin Machines via Discounted Eligibility Experience Buffers (TM-DEB)"
**Sources used:**
- `docs/papers/1804.01508v15.pdf` (Granmo, "The Tsetlin Machine…", the authoritative source) — extracted to `/tmp/granmo.txt`
- `docs/papers/2609.06133v1.pdf` (Kumar et al., "Compressed Recurrent Feedback in Tsetlin Machines…") — extracted to `/tmp/rtm.txt`
- `common_libs/guns/tsetlin.nim` (read in full)
**Verification legend:** `[VERIFIED]` = I read the exact equation/line and/or ran the arithmetic myself. `[INFERRED]` = my reasoning from the code/paper, not measured at runtime. Per task constraints I did **not** build, run battles, or edit any `.nim` file.
---
## 1. Verdict (one paragraph)
**Do not adopt TM-DEB, and do not spend time on it.** The document is an AI-generated theoretical sketch (self-identifying header "Generated by Gemini Notebook", no author, a misattributed reference, four consecutive algorithm steps that are literally the character `%`, and a results table whose title says "Expected" and whose numbers contradict the very paper it claims to adapt). It does not fix the gun, because the gun's failure is **not** a delayed-credit-assignment problem — the gun already has exact `(fireTick, powerBin)` eligibility traces, so its effective credit-assignment `λ` is 1. The failure is a **bug in the Type I/Type II feedback rule inside `tmLearnOne`**: (i) Type I ignores the clause output and always pushes true literals toward *Include*, deleting Granmo's `(clause=0, literal=1) → 1/s toward Exclude` counteracting force, and (ii) Type II is dead code — its guard `cOut == 1` and its condition `lits[lit] == 0 and st > 0` are mutually exclusive, so it can never execute. With no counteracting force at all, the automata ratchet toward Include until the clause is hopelessly over-specific (~131 literals/clause), and no clause ever fires. TM-DEB changes only the *frequency* of updates (`γ^Δt`), not their *direction*, so it cannot move the fixed point: with the buggy rule it converges to the same saturated clause, only slower — and for the power-3 shots that matter most (`Δt` 60–90 ticks) it reduces the update probability to `5.8e-5 … 4.4e-7`, effectively deleting the exact feedback we now have. The correct fix is to implement Granmo's Table 2 and Table 3 faithfully.
---
## 2. Credibility assessment of the document
Strong, independently verifiable indicators that this is low-quality, AI-generated, unpublished material:
1. **Self-identifying provenance.** Page 1: "*Generated by Gemini Notebook*", author line `gemini-notebook@workspace.local`, date "September 19, 2026". `[VERIFIED]`
2. **Reference [2] is misattributed.** `[VERIFIED]`
- TM-DEB cites: `[2] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study," IEEE Transactions / ISTM, 2025.`
- `docs/papers/2609.06133v1.pdf` has exactly that title but is authored by **Ankit Kumar, Utkarsh Raj, Rishad Shafik, Sudip Roy** (`/tmp/rtm.txt` lines 1–6), arXiv:2609.06133v1. There is no "IEEE Transactions / ISTM".
- The author list in TM-DEB's `[2]` matches a *different* work that the RTM paper itself cites: `/tmp/rtm.txt` line 699 — `[3] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Recurrent Tsetlin Machine…"`. So TM-DEB appears to have fused one paper's title with another's author list.
3. **Algorithm Specification 2 is broken.** Steps 9–12 are literally `%` in **both** the `-raw` and `-layout` extractions (`/tmp/tm-deb.txt` lines 520–523; `/tmp/tm-deb-layout.txt` lines 332–336). `[VERIFIED]` This is the core of the paper — the actual automata update — and it is missing, so the algorithm is not reproducible.
4. **No experiment was run.** Table 1 is captioned "*Expected* Comparative Metric Framework". Section 7 is titled "*Empirical Verification Protocol & Benchmarks*" and says "we *define* a rigorous experimental protocol" — protocol language, no results, no seeds, no dataset sizes, no figures. `[VERIFIED]`
5. **Contradicts its own cited source.** The real RTM study reports `61.47±6.74%` / `62.94±9.92%` over 144 runs (`/tmp/rtm.txt` abstract). TM-DEB's Table 1 reports "Std RTM 52.1 ± 1.2". `[VERIFIED]` The numbers are not from that study.
6. **Other references.** `[1]` Granmo arXiv:1804.01508, `[3]` Tsetlin 1961, `[4]` Narendra & Thathachar 1989 are real. `[5]` (Tung & Kleinrock, "*Using Finite State Automata to Produce Self-Optimization…*") matches Granmo's reference `[15]` (`/tmp/granmo.txt` line 3335) and is real. `[VERIFIED]` So the only unverifiable/misattributed reference is `[2]`, but combined with items 1–5 the credibility is very low.
**Conclusion:** treat TM-DEB as an untrusted informal note, not as literature. Granmo's `1804.01508v15` is the authoritative source and is already in the repo.
---
## 3. T1 — Theorem-by-theorem verification
### (a) Theorem 2 (Eq. 11) — *"sign of expected payoff difference is discount-invariant"*
**Paper text** (`/tmp/tm-deb.txt`):
- Eq. 11: `Sign(E[R|α2,Δt] − E[R|α1,Δt]) = Sign(E[R|α2,0] − E[R|α1,0])`
- Eq. 13: `E[R|α2,Δt] = γ^Δt_d · E[R|α2,0]`; Eq. 14 same for `α1`
- Eq. 15: `ΔE[R]_Δt = γ^Δt_d · (E[R|α2,0] − E[R|α1,0])`
**My verdict: I AGREE with your critique, with one clarification.**
- **"Category error" — correct, under Granmo's own convention.** TM-DEB's Eq. 10 is `P_DEB(k,T) = γ^(T−k)·P_base(...)`. That scales the *probability that feedback is dispatched*, not the payoff values in Table 2/3. Granmo's `E[R|α]` (Lemma 1/2, `/tmp/granmo.txt` lines 1432–1565) is the expected reward-minus-penalty **conditional on the clause receiving feedback** — the resource-allocation factor `(T−clip(v,−T,T))/(2T)` is *not* folded into it. Under TM-DEB, conditional on dispatch the reward/penalty distribution is the *same table*, so the conditional expected payoff is **unchanged**: `E[R|α,Δt] = E[R|α,0]`, not `γ^Δt·E[R|α,0]`. The paper silently switched to an *unconditional per-step* expected payoff (which is indeed scaled by `γ^Δt` because no-update steps contribute 0). So Eq. 13–14 are not a valid inference under the paper's own cite to Granmo. `[VERIFIED against Granmo Lemmas 1–2]`
- **"Tautology" — correct, and this is the stronger point.** Even granting Eq. 13–14, Eq. 15 factors a *strictly positive scalar* out of a difference; a positive scalar can never flip a sign. So Theorem 2 proves sign-invariance for **any** positive multiplier, which is trivially true and says nothing about learning. The theorem is vacuous as a convergence guarantee.
- **What Theorem 2 ignores** (which is the actually harmful effect): `γ^Δt` scales the *learning rate* — the number of updates each automaton receives. Tsetlin automata need enough updates to grow their state space and converge (Granmo Lemma 3, `/tmp/granmo.txt` line ~1634). Reducing update probability by `γ^Δt` slows convergence and, for large `Δt`, starves the automaton of samples. Theorem 2 is silent on this trade-off.
### (b) Lemma 1 (Eq. 16–18) — *"SNR variance bound `σ²_base/(1−γ²)`"*
**Paper text:** Eq. 17 `σ²_DEB = Σ_{k=1}^{T} (γ^{T−k})² σ²_base = σ²_base Σ_{m=0}^{T−1} (γ²)^m`, Eq. 18 limit `σ²_base/(1−γ²)`.
**My verdict: I AGREE with both parts of your critique.**
- **Covariance dropped.** Eq. 17 uses `Var(Σ_k X_k) = Σ_k Var(X_k)`. Under TM-DEB every historical step's update is gated by the *same* terminal reward `R_T` (Algorithm Spec 2 requires one `R_T` and computes `y_k^eff` from it, `/tmp/tm-deb.txt` lines 508–515). The per-step feedback variables therefore share a common factor and are positively correlated; the covariance terms `2Σ_{k<l} γ^{…} Cov(X_k,X_l)` are silently dropped, so the bound is unjustified. Moreover the per-step variance under Bernoulli dispatch itself depends on `λ·P_base`, so it is not a fixed `σ²_base`. `[VERIFIED]`
- **Conflates two mechanisms.** Granmo's vanishing-SNR is a function of **team size `W`**: `SNR = 1/(W·p(1−p))` (Granmo Eq. 3, `/tmp/granmo.txt` lines 163–175). TM-DEB's Lemma 1 bounds a **temporal** accumulation over episode length `T` and never mentions `W`. A temporal discount does nothing about the `W·p(1−p)` team-noise term. Granmo's actual cure for vanishing SNR is *local* feedback from the literal and clause values plus resource allocation (Granmo Remark 2, `/tmp/granmo.txt` line ~2632, and Theorem 4), not a discount. `[VERIFIED]`
- Minor: the lemma is titled "SNR Variance Bound" but only bounds variance; it never bounds SNR (which would require the signal term `μ²`).
### (c) Theorem 3 — *"Generalized-ordinal-potential property is preserved"*
**Paper text:** "…every individual state transition in `A_current` aligns monotonically with increasing the expected episodic payoff `E[R_T]`. Hence, the ordinal potential property is preserved across temporal steps." (`/tmp/tm-deb.txt` lines 462–475)
**My verdict: I AGREE — it is an assertion, not a proof, and the premise is false.**
- It is two sentences ending in "Hence", with no case analysis. Compare Granmo's Theorem 4 (`/tmp/granmo.txt` lines 2635–2700), which is a real proof broken into the false-negative and false-positive cases and explicitly restricted to summation target `T=1`, positive clauses, and **single-step** classification accuracy. `[VERIFIED]`
- **The premise is false.** Granmo's alignment depends on local feedback derived from the *current input's* literal and clause values (Granmo Remark 2: "each single automaton gets feedback directly from the value of its literal and the value of its clause"). TM-DEB instead relabels *historical* inputs with a *terminal* target: Algorithm Spec 2 step 5 sets `y_k^eff = ŷ_k` if `R_T=+1` else `1−ŷ_k`. So a locally harmful decision that happened to precede eventual success is reinforced with Type I feedback. The historical `(X_k, C_k)` no longer determine whether the action at step `k` was correct, which is exactly the property Granmo's Theorem 4 relies on. The ordinal-potential correspondence with accuracy does not transfer. `[VERIFIED by reading both proofs]`
- Additional (missed by the paper): TM-DEB stores clause outputs `C_k` computed at step `k` but applies updates to the *final* automata state `A_current`; the stored clause outputs and the live automata are misaligned by construction. This is arguably a worse problem than the "checkpointing divergence" Theorem 1 warns about.
### (d) Algorithm Specification 2 incomplete — **CONFIRMED**
Steps 9–12 are literally `%` in both extractions: `/tmp/tm-deb.txt` lines 520–523 and `/tmp/tm-deb-layout.txt` lines 332–336. `[VERIFIED]` Not an extraction artifact.
**What is therefore unspecified:** after sampling `Random() ≤ λ·P_base` at step 8, the paper never says (i) whether the sampled event is Reward, Penalty, or Inaction (i.e., which cell of Granmo's Table 2/3), nor (ii) how that feedback maps to an automata state transition via `F(·,·)` (Eq. 2). In other words the **entire operational core** — the thing the paper claims to contribute — is absent. The algorithm cannot be reimplemented from the document.
### (e) Table 1 — **illustrative, not experimental**
**Verdict: CONFIRMED, these are illustrative/projected numbers.** Evidence `[VERIFIED]`:
- Caption: "*Expected* Comparative Metric Framework across Delayed Reward Benchmarks" (`/tmp/tm-deb.txt` line 552).
- Section 7 text: "we *define* a rigorous experimental protocol comparing TM-DEB…" (`/tmp/tm-deb.txt` line 533) — protocol language. There is no results section, no dataset description, no seeds, no compute, no figures.
- The only table is Table 1; the document ends at the Conclusion.
- The numbers disagree with the source: real RTM study `61.47±6.74%`/`62.94±9.92%` over 144 runs vs. Table 1 "Std RTM 52.1±1.2" (`/tmp/rtm.txt` abstract). `[VERIFIED]`
**Note:** The `±` values give a false impression of measured variance.
### (f) Reference [2] mismatch — **CONFIRMED**
- TM-DEB `[2]`: "C. Xu, S. Duan, R. Shafik, and A. Yakovlev, 'Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study,' IEEE Transactions / ISTM, 2025." (`/tmp/tm-deb.txt` lines 586–589)
- Actual `docs/papers/2609.06133v1.pdf`: Kumar, Raj, Shafik, Roy, arXiv:2609.06133v1 (Sep 2026), same title (`/tmp/rtm.txt` lines 1–6, 102).
- The author list in `[2]` is taken from a *different* paper cited by the RTM study (`/tmp/rtm.txt` line 699). `[VERIFIED]`
- No other reference is unverifiable: `[1]`/`[3]`/`[4]` are standard, and `[5]` matches Granmo's `[15]` (`/tmp/granmo.txt` line 3335). `[VERIFIED]`
### (g) Math errors you MISSED
**Correct as written `[VERIFIED]`:**
- **Eq. 2 (automaton transition)** is correct and matches Granmo Eq. 2: exclude range `1..N` → Penalty `+1`, Reward `−1`; include range `N+1..2N` → Penalty `−1`, Reward `+1`; boundaries saturate. (`/tmp/tm-deb.txt` lines ~112–130)
- **Eq. 3** `C_j(X) = ∧_{l_k∈L_j} l_k` — correct conjunction.
- **Eq. 4** `v = Σ C^1_j − Σ C^0_j`, `ŷ = 1[v≥0]` — matches Granmo Eq. 7.
- **Eq. 10** `P_DEB = γ^(T−k)·P_base` — well-formed probability (`P_base ≤ 1`, `γ ≤ 1`). Note it can only *reduce* the update count.
**Additional errors found:**
1. **Eq. 12 is not a faithful copy of Granmo's Lemma 2 (Include payoff).** TM-DEB Eq. 12:
`E[R|α2] = (s−1)/s·γ_e·P(X1_jk∩X1) − 1/s·γ_e·P(X1\X1_jk) − 1/s·(1−γ_e)·P(X0)`
Granmo's Include payoff (`/tmp/granmo.txt` lines 1540–1565) is:
`(s−1)/s·γ·P(X1_jk∩X1) + (s−1)/s·(1−γ)·P(X1_jk∩X0) − 1/s·γ·P(X1\X1_jk) − 1/s·(1−γ)·P(X0\X1_jk)`
TM-DEB **drops the positive `(1−γ_e)·P(X1_jk∩X0)` term** and **replaces `P(X0\X1_jk)` with `P(X0)`**. Since `P(X0) ≥ P(X0\X1_jk)`, the penalty is overstated. `[VERIFIED]`
2. **Buffer tuple inconsistency.** Eq. 9 defines `B_t = (t, X_t, C_t(X_t), a_t)` (4 fields, `/tmp/tm-deb.txt` line ~285), but Algorithm Spec 1 step 6 pushes `(t, X_t, C_t, v_t, ŷ_t)` (5 fields) and Algorithm Spec 2 step 3 reads `(k, X_k, C_k, v_k, ŷ_k)` (5 fields). The action `a_t` from Eq. 9 is never the one used (step 5 uses `ŷ_k`). `[VERIFIED]`
3. **Discount is direction-blind.** Eq. 10 multiplies the *whole* `P_base`, which aggregates both reward and penalty cells of Table 2/3. So `γ^Δt` cannot perform *credit assignment* in the directional sense — it is a pure forgetting/truncation knob applied identically to reinforcement and punishment. This is consistent with Theorem 2's vacuity but contradicts the paper's framing.
4. **Theorem 1's "checkpointing" target is a strawman**, and misses that TM-DEB itself replays stale clause outputs against a state that has since changed (see (c)). `[INFERRED]`
**Nothing in (g) changes the practical picture: TM-DEB adds no mechanism beyond what the gun already has, and its theory is unsound.**
---
## 4. T2 — Does TM-DEB address the actual failure (clause saturation)? **NO.**
**Mechanism.** TM-DEB's only operational additions are (a) an eligibility buffer of `(X_k, C_k)` tuples — which the gun **already has**, keyed exactly by `(fireTick, powerBin)` — and (b) multiplying the feedback-dispatch probability by `γ^(T−k)` (Eq. 10). It does **not** change the update rule (Type I/II) — step 7 still says "Compute base feedback prob. `P_base` via Eqns. (8)–(11)" of Granmo.
**Therefore:** the *direction* of every update is unchanged, and the saturation fixed point is unchanged. Discounting only slows the approach to the same saturated state. For the buggy rule in `tsetlin.nim`, the drift toward Include is monotone, so no amount of probability scaling can reduce the ~131 included literals.
**Concrete `γ` numbers for the paper's own `γ = 0.85` `[VERIFIED, computed]`:**
| `Δt` (ticks) | `γ^Δt` | meaning |
|---|---|---|
| 23 | `2.38e-2` | ~1 in 42 updates survives |
| 36 | `2.88e-3` | ~1 in 347 |
| 60 | `5.82e-5` | ~1 in 17,180 |
| 90 | `4.44e-7` | ~1 in 2.25 million |
Half-life: `Δt ≈ 4.27` ticks. `1/(1−γ²) = 3.60` (the paper's variance bound, i.e. only ~3.6 effective steps are remembered).
**Why this is exactly backwards for this gun.** The gun's virtual bullets resolve 23–90 ticks after firing (`bulletSpeed(power) = 20 − 3·power` in `common_libs/gun_harness/gun_interface.nim`; power-3 long shots reach the 90-tick regime, cf. the comment in `tsetlin.nim` about ~128 ticks). These long-range power-3 shots are precisely the ones whose prediction matters most and are the hardest to get right. TM-DEB would multiply their update probability by `5.8e-5 … 4.4e-7`, i.e. **delete** the exact feedback that the current `(fireTick, powerBin)` ring now captures cleanly (`trainedShots = 2141`, `traceMisses = 0`).
**Verdict:** TM-DEB **does not fix** the gun, and would **worsen** the long-range (`power-3`) case. Its discount is a liability here, not a feature.
---
## 5. T3 — What we already have vs. what TM-DEB adds
`tsetlin.nim` stores, per fired virtual bullet, a `TmTrace` containing the input vector `X_k` and the clause-output cache `C_k`, keyed exactly by `(fireTick, powerBin)`. When `onResult` fires, it retrieves the *exact* trace and applies the update. That is precisely TM-DEB's "experience buffer" — an eligibility trace — and the pairing is exact.
Therefore **our effective eligibility `λ = 1`**. TM-DEB's exponential discount `γ^(T−k)` exists to solve the classic RL temporal-credit-assignment problem: when you *cannot* identify which past action produced the reward, you must hedge by spreading credit over recent actions with a decay. In this gun, the virtual-bullet tracker *can* identify the responsible action exactly (the resolution carries `fireTick` and `powerBin`). Discounting would therefore discard the few, highly relevant long-delay samples in exchange for a recency prior we do not need. It is a workaround for missing information that we possess — hence **information-losing** in our case.
(One caveat: TM-DEB's credit assignment is *also* inexact in the other direction — it replays stale clause outputs against the final state. Our ring has the same stale-cache property, but at least it never throws away a correctly identified sample.)
---
## 6. T4 — Granmo's correct feedback vs. what `tsetlin.nim` implements
Reference: Granmo Table 2 (Type I) and Table 3 (Type II), `/tmp/granmo.txt` lines 705 and 723. In Granmo's transition (Eq. 2), for an automaton in the **Exclude** range a Reward moves to *more Exclude* and a Penalty moves *toward Include*; for the **Include** range it is reversed. The tables below are already collapsed to the resulting state move.
### 6.1 Granmo Table 2 (Type I) — collapsed
| clause `c` | literal `l_k` | automaton action | feedback probs | effective state move |
|---|---|---|---|---|
| 1 | 1 | Include | `P(Reward)=(s−1)/s`, `P(Inaction)=1/s` | **+1 (toward Include)** w.p. `(s−1)/s` |
| 1 | 1 | Exclude | `P(Penalty)=(s−1)/s`, `P(Inaction)=1/s` | **+1 (toward Include)** w.p. `(s−1)/s` |
| 0 | 1 | Include | `P(Penalty)=1/s`, `P(Inaction)=(s−1)/s` | **−1 (toward Exclude)** w.p. `1/s` |
| 0 | 1 | Exclude | `P(Reward)=1/s`, `P(Inaction)=(s−1)/s` | **−1 (toward Exclude)** w.p. `1/s` |
| 1 | 0 | Exclude | `P(Reward)=1/s`, `P(Inaction)=(s−1)/s` | **−1 (toward Exclude)** w.p. `1/s` |
| 0 | 0 | Include/Exclude | penalty/reward `1/s` | **−1 (toward Exclude)** w.p. `1/s` |
**Collapsed closed form (independent of current action):**
- `l_k == 1 and c == 1` → state `+1` w.p. `(s−1)/s`
- `l_k == 1 and c == 0` → state `−1` w.p. `1/s`
- `l_k == 0` (any `c`) → state `−1` w.p. `1/s`
### 6.2 Granmo Table 3 (Type II) — collapsed
The **only** non-Inaction cell is: `c == 1 and l_k == 0` → penalize the **Exclude** action → state **+1 (toward Include)** w.p. `1.0`. Everything else is Inaction. (This is what increases discrimination power by adding a zero-valued literal to a false-positive clause.)
### 6.3 What `tsetlin.nim` actually does
Type I branch (`tsetlin.nim`, `tmLearnOne`):
```nim
if (error > 0.0 and pol > 0.0) or (error < 0.0 and pol < 0.0):
for lit in 0..<TM_N_LITERALS:
if lits[lit] == 1'u8:
if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES)
else:
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
```
- **It never reads `cOut`.** For `l_k == 1` it *always* moves `+1` w.p. `(s−1)/s`, for *any* clause output.
- Correct for the `c == 1, l_k == 1` row only. **Wrong for `c == 0, l_k == 1`**, where the table requires `−1` w.p. `1/s`. That missing `−1` is the entire counteracting force (Granmo calls it "non-matching input"; `/tmp/granmo.txt` line ~640).
- `l_k == 0` → `−1` w.p. `1/s` is correct.
Type II branch:
```nim
else:
if cOut == 1'u8:
for lit in 0..<TM_N_LITERALS:
if lits[lit] == 0'u8:
... if st > 0: st = max(st - 1, -TM_N_STATES)
```
- **Dead code.** `tmEvalClause` returns `1` only if no included literal has value 0. So `cOut == 1` implies every `st > 0` literal has `lits[lit] == 1`. The guard `lits[lit] == 0 and st > 0` is therefore unsatisfiable and the body never executes. `[VERIFIED by reading `tmEvalClause` + `tmLearnOne`]`
- Even if it were reachable, it has the **wrong direction**: it penalizes an *included* false literal, whereas Granmo penalizes the *exclusion* of a false literal (i.e. increments an excluded false literal). And it is a no-op for the actual false-positive case.
Resource allocation: Granmo Eqns. 8–11 use `(T − clip(v,−T,T))/(2T)` (and `(T + clip(...))/(2T)`) to decide *whether* a clause gets feedback, distributing clauses across sub-patterns. `tsetlin.nim` instead uses `pFeedback = min(1, |error|/(2·80))` — a regression-error heuristic. `TM_T` (=25) is used only to normalize the vote for the output, not as a summation target. So the resource-allocation term is **absent**.
### 6.4 Most likely concrete reason for saturation at ~131 literals/clause
With Type I ignoring `c` (always `+1` for true literals w.p. `(s−1)/s`) and Type II dead, there is **no force pushing a clause back toward exclusion**. Each literal random-walks with drift `p_true − 1/s` per feedback event (`[INFERRED, matches Granmo's balance algebra]`), so any literal that is true more than `1/s` of the time marches to the Include boundary and stays there. `s = 1.5` gives threshold `1/s = 0.667`; larger `s` (as swept: 2.5/4.0/8.0) *lowers* the threshold and makes it worse — consistent with the observation that the `s` sweep did not help. The clause accumulates ~131 literals and fires with probability ~`∏ p(bit)`, effectively never. Hence `vote = 0`, `(cx,cy) = (0,0)`, and the gun is byte-identical to Linear.
### 6.5 Prioritised fix list (recommendations only — no source edits made)
1. **[CRITICAL] Make Type I condition on the clause output.** Replace the unconditional include-drift with the collapsed Table 2:
```nim
# cOut is the cached clause output for this clause at fire time
if lits[lit] == 1'u8:
if cOut == 1'u8:
# (c=1, lk=1): both actions rewarded/penalized toward Include
if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES)
else:
# (c=0, lk=1): both actions driven toward Exclude <-- MISSING TODAY
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
else:
# (lk=0, any c): driven toward Exclude
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
```
This single change restores the counteracting force and should stop the ratchet. `[INFERRED from Granmo Tables 2 + Eq. 2]`
2. **[CRITICAL] Repair Type II.** It must penalize the *exclusion* of a false literal when the clause fires:
```nim
else: # Type II
if cOut == 1'u8:
for lit in 0..<TM_N_LITERALS:
if lits[lit] == 0'u8:
let si = tmStateIdx(outIdx, c, lit)
if net.states[si] <= 0: # Exclude action -> Penalty -> toward Include
net.states[si] = int16(min(int(net.states[si]) + 1, TM_N_STATES))
```
Without this, the repaired Type I will keep clauses sparse but the TM cannot learn to *discriminate* (suppress false positives), so the correction will likely stay near 0 for the wrong reason. `[INFERRED]`
3. **[HIGH] Restore the summation-target / resource-allocation probability** from Granmo Eqns. 8–11 (`(T − clip(v,−T,T))/(2T)`), or an explicit regression analogue, instead of the pure `|error|` heuristic. `tsetlin.nim`'s `TM_T` is currently only a normaliser. This is the paper's "margin" mechanism and is what distributes clauses across sub-patterns.
4. **[MEDIUM] Re-examine the regression framing.** Granmo's TM is a binary classifier; this gun uses sign(error) as an implicit label and `|error|` as a dispatch probability. That adaptation is defensible (it mirrors a perceptron/sign update), but it should be validated *after* (1) and (2). A safer formulation is to predict the sign of the residual through the standard positive/negative clause machinery, and convert sign→pixels separately.
5. **[LOW] Optional hygiene:** initialise states at the Exclude boundary and prune all-Exclude clauses (Granmo Algorithm 1 line 27), and consider a state range/`N` that matches `s`.
**Validation protocol (recommended, not run):** before/after fix, log per-clause include counts over a 600-tick simulation (expect include count to fall from ~131 toward single digits) and confirm `(cx,cy) ≠ 0` on at least some ticks; then run the existing integration battle against `Linear` and `SittingDuck` and check the virtual-bullet hit count diverges from Linear's byte-for-byte. Do not touch `feedback`/trace plumbing — it is already correct.
---
## 7. When TM-DEB-style delayed credit assignment *would* be the right tool
Discounting an eligibility buffer is the right mechanism when the system **cannot** pair a past action to its outcome, so the only available signal is a sparse terminal reward. The gun does not have that property — `(fireTick, powerBin)` gives exact pairing.
A future case in this repo where it *would* apply: a **learned movement / positioning module** rewarded only by the round outcome (win/loss or final survival), where no per-tick target exists and the credit for a win cannot be attributed to a specific earlier turn or thrust. There, an eligibility trace with `γ < 1` (or a proper TD(λ)/REINFORCE estimator) is the standard and appropriate tool, because the responsible action is genuinely ambiguous. Note that even then, Granmo's *local* Type I/II machinery would still need to be correct; TM-DEB is a wrapper around it, not a substitute.
---
## 8. Bottom line for the orchestrator
- **Paper:** untrustworthy (AI-generated, misattributed ref, missing algorithm core, illustrative results). Do not cite or adopt.
- **Gun:** the real bug is in `common_libs/guns/tsetlin.nim` → `tmLearnOne`: Type I ignores `cOut` (deleting Granmo's `c=0` exclude-drift), and Type II is dead code with the wrong direction. Fixes 1 and 2 above are the actionable items.
- **TM-DEB:** does not fix, would worsen long-range (power-3) learning by deleting exactly attributable feedback via `γ^Δt` (`4.4e-7` at `Δt=90`).