diff --git a/docs/research/tm-deb-assessment.md b/docs/research/tm-deb-assessment.md new file mode 100644 index 0000000..96d5927 --- /dev/null +++ b/docs/research/tm-deb-assessment.md @@ -0,0 +1,259 @@ +# TM-DEB Paper Assessment — Adopt as a fix for the Tsetlin gun? + +**Issue:** N/A (read-only review requested by orchestrator) +**Date:** 2026-09-20 +**Paper under review:** `docs/papers/tm-deb-paper.pdf` — "Temporal Credit Assignment in Tsetlin Machines via Discounted Eligibility Experience Buffers (TM-DEB)" +**Sources used:** +- `docs/papers/1804.01508v15.pdf` (Granmo, "The Tsetlin Machine…", the authoritative source) — extracted to `/tmp/granmo.txt` +- `docs/papers/2609.06133v1.pdf` (Kumar et al., "Compressed Recurrent Feedback in Tsetlin Machines…") — extracted to `/tmp/rtm.txt` +- `common_libs/guns/tsetlin.nim` (read in full) + +**Verification legend:** `[VERIFIED]` = I read the exact equation/line and/or ran the arithmetic myself. `[INFERRED]` = my reasoning from the code/paper, not measured at runtime. Per task constraints I did **not** build, run battles, or edit any `.nim` file. + +--- + +## 1. Verdict (one paragraph) + +**Do not adopt TM-DEB, and do not spend time on it.** The document is an AI-generated theoretical sketch (self-identifying header "Generated by Gemini Notebook", no author, a misattributed reference, four consecutive algorithm steps that are literally the character `%`, and a results table whose title says "Expected" and whose numbers contradict the very paper it claims to adapt). It does not fix the gun, because the gun's failure is **not** a delayed-credit-assignment problem — the gun already has exact `(fireTick, powerBin)` eligibility traces, so its effective credit-assignment `λ` is 1. The failure is a **bug in the Type I/Type II feedback rule inside `tmLearnOne`**: (i) Type I ignores the clause output and always pushes true literals toward *Include*, deleting Granmo's `(clause=0, literal=1) → 1/s toward Exclude` counteracting force, and (ii) Type II is dead code — its guard `cOut == 1` and its condition `lits[lit] == 0 and st > 0` are mutually exclusive, so it can never execute. With no counteracting force at all, the automata ratchet toward Include until the clause is hopelessly over-specific (~131 literals/clause), and no clause ever fires. TM-DEB changes only the *frequency* of updates (`γ^Δt`), not their *direction*, so it cannot move the fixed point: with the buggy rule it converges to the same saturated clause, only slower — and for the power-3 shots that matter most (`Δt` 60–90 ticks) it reduces the update probability to `5.8e-5 … 4.4e-7`, effectively deleting the exact feedback we now have. The correct fix is to implement Granmo's Table 2 and Table 3 faithfully. + +--- + +## 2. Credibility assessment of the document + +Strong, independently verifiable indicators that this is low-quality, AI-generated, unpublished material: + +1. **Self-identifying provenance.** Page 1: "*Generated by Gemini Notebook*", author line `gemini-notebook@workspace.local`, date "September 19, 2026". `[VERIFIED]` +2. **Reference [2] is misattributed.** `[VERIFIED]` + - TM-DEB cites: `[2] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study," IEEE Transactions / ISTM, 2025.` + - `docs/papers/2609.06133v1.pdf` has exactly that title but is authored by **Ankit Kumar, Utkarsh Raj, Rishad Shafik, Sudip Roy** (`/tmp/rtm.txt` lines 1–6), arXiv:2609.06133v1. There is no "IEEE Transactions / ISTM". + - The author list in TM-DEB's `[2]` matches a *different* work that the RTM paper itself cites: `/tmp/rtm.txt` line 699 — `[3] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Recurrent Tsetlin Machine…"`. So TM-DEB appears to have fused one paper's title with another's author list. +3. **Algorithm Specification 2 is broken.** Steps 9–12 are literally `%` in **both** the `-raw` and `-layout` extractions (`/tmp/tm-deb.txt` lines 520–523; `/tmp/tm-deb-layout.txt` lines 332–336). `[VERIFIED]` This is the core of the paper — the actual automata update — and it is missing, so the algorithm is not reproducible. +4. **No experiment was run.** Table 1 is captioned "*Expected* Comparative Metric Framework". Section 7 is titled "*Empirical Verification Protocol & Benchmarks*" and says "we *define* a rigorous experimental protocol" — protocol language, no results, no seeds, no dataset sizes, no figures. `[VERIFIED]` +5. **Contradicts its own cited source.** The real RTM study reports `61.47±6.74%` / `62.94±9.92%` over 144 runs (`/tmp/rtm.txt` abstract). TM-DEB's Table 1 reports "Std RTM 52.1 ± 1.2". `[VERIFIED]` The numbers are not from that study. +6. **Other references.** `[1]` Granmo arXiv:1804.01508, `[3]` Tsetlin 1961, `[4]` Narendra & Thathachar 1989 are real. `[5]` (Tung & Kleinrock, "*Using Finite State Automata to Produce Self-Optimization…*") matches Granmo's reference `[15]` (`/tmp/granmo.txt` line 3335) and is real. `[VERIFIED]` So the only unverifiable/misattributed reference is `[2]`, but combined with items 1–5 the credibility is very low. + +**Conclusion:** treat TM-DEB as an untrusted informal note, not as literature. Granmo's `1804.01508v15` is the authoritative source and is already in the repo. + +--- + +## 3. T1 — Theorem-by-theorem verification + +### (a) Theorem 2 (Eq. 11) — *"sign of expected payoff difference is discount-invariant"* + +**Paper text** (`/tmp/tm-deb.txt`): +- Eq. 11: `Sign(E[R|α2,Δt] − E[R|α1,Δt]) = Sign(E[R|α2,0] − E[R|α1,0])` +- Eq. 13: `E[R|α2,Δt] = γ^Δt_d · E[R|α2,0]`; Eq. 14 same for `α1` +- Eq. 15: `ΔE[R]_Δt = γ^Δt_d · (E[R|α2,0] − E[R|α1,0])` + +**My verdict: I AGREE with your critique, with one clarification.** + +- **"Category error" — correct, under Granmo's own convention.** TM-DEB's Eq. 10 is `P_DEB(k,T) = γ^(T−k)·P_base(...)`. That scales the *probability that feedback is dispatched*, not the payoff values in Table 2/3. Granmo's `E[R|α]` (Lemma 1/2, `/tmp/granmo.txt` lines 1432–1565) is the expected reward-minus-penalty **conditional on the clause receiving feedback** — the resource-allocation factor `(T−clip(v,−T,T))/(2T)` is *not* folded into it. Under TM-DEB, conditional on dispatch the reward/penalty distribution is the *same table*, so the conditional expected payoff is **unchanged**: `E[R|α,Δt] = E[R|α,0]`, not `γ^Δt·E[R|α,0]`. The paper silently switched to an *unconditional per-step* expected payoff (which is indeed scaled by `γ^Δt` because no-update steps contribute 0). So Eq. 13–14 are not a valid inference under the paper's own cite to Granmo. `[VERIFIED against Granmo Lemmas 1–2]` +- **"Tautology" — correct, and this is the stronger point.** Even granting Eq. 13–14, Eq. 15 factors a *strictly positive scalar* out of a difference; a positive scalar can never flip a sign. So Theorem 2 proves sign-invariance for **any** positive multiplier, which is trivially true and says nothing about learning. The theorem is vacuous as a convergence guarantee. +- **What Theorem 2 ignores** (which is the actually harmful effect): `γ^Δt` scales the *learning rate* — the number of updates each automaton receives. Tsetlin automata need enough updates to grow their state space and converge (Granmo Lemma 3, `/tmp/granmo.txt` line ~1634). Reducing update probability by `γ^Δt` slows convergence and, for large `Δt`, starves the automaton of samples. Theorem 2 is silent on this trade-off. + +### (b) Lemma 1 (Eq. 16–18) — *"SNR variance bound `σ²_base/(1−γ²)`"* + +**Paper text:** Eq. 17 `σ²_DEB = Σ_{k=1}^{T} (γ^{T−k})² σ²_base = σ²_base Σ_{m=0}^{T−1} (γ²)^m`, Eq. 18 limit `σ²_base/(1−γ²)`. + +**My verdict: I AGREE with both parts of your critique.** +- **Covariance dropped.** Eq. 17 uses `Var(Σ_k X_k) = Σ_k Var(X_k)`. Under TM-DEB every historical step's update is gated by the *same* terminal reward `R_T` (Algorithm Spec 2 requires one `R_T` and computes `y_k^eff` from it, `/tmp/tm-deb.txt` lines 508–515). The per-step feedback variables therefore share a common factor and are positively correlated; the covariance terms `2Σ_{k 0.0 and pol > 0.0) or (error < 0.0 and pol < 0.0): + for lit in 0.. 0: st = max(st - 1, -TM_N_STATES) +``` +- **Dead code.** `tmEvalClause` returns `1` only if no included literal has value 0. So `cOut == 1` implies every `st > 0` literal has `lits[lit] == 1`. The guard `lits[lit] == 0 and st > 0` is therefore unsatisfiable and the body never executes. `[VERIFIED by reading `tmEvalClause` + `tmLearnOne`]` +- Even if it were reachable, it has the **wrong direction**: it penalizes an *included* false literal, whereas Granmo penalizes the *exclusion* of a false literal (i.e. increments an excluded false literal). And it is a no-op for the actual false-positive case. + +Resource allocation: Granmo Eqns. 8–11 use `(T − clip(v,−T,T))/(2T)` (and `(T + clip(...))/(2T)`) to decide *whether* a clause gets feedback, distributing clauses across sub-patterns. `tsetlin.nim` instead uses `pFeedback = min(1, |error|/(2·80))` — a regression-error heuristic. `TM_T` (=25) is used only to normalize the vote for the output, not as a summation target. So the resource-allocation term is **absent**. + +### 6.4 Most likely concrete reason for saturation at ~131 literals/clause + +With Type I ignoring `c` (always `+1` for true literals w.p. `(s−1)/s`) and Type II dead, there is **no force pushing a clause back toward exclusion**. Each literal random-walks with drift `p_true − 1/s` per feedback event (`[INFERRED, matches Granmo's balance algebra]`), so any literal that is true more than `1/s` of the time marches to the Include boundary and stays there. `s = 1.5` gives threshold `1/s = 0.667`; larger `s` (as swept: 2.5/4.0/8.0) *lowers* the threshold and makes it worse — consistent with the observation that the `s` sweep did not help. The clause accumulates ~131 literals and fires with probability ~`∏ p(bit)`, effectively never. Hence `vote = 0`, `(cx,cy) = (0,0)`, and the gun is byte-identical to Linear. + +### 6.5 Prioritised fix list (recommendations only — no source edits made) + +1. **[CRITICAL] Make Type I condition on the clause output.** Replace the unconditional include-drift with the collapsed Table 2: + ```nim + # cOut is the cached clause output for this clause at fire time + if lits[lit] == 1'u8: + if cOut == 1'u8: + # (c=1, lk=1): both actions rewarded/penalized toward Include + if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES) + else: + # (c=0, lk=1): both actions driven toward Exclude <-- MISSING TODAY + if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES) + else: + # (lk=0, any c): driven toward Exclude + if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES) + ``` + This single change restores the counteracting force and should stop the ratchet. `[INFERRED from Granmo Tables 2 + Eq. 2]` + +2. **[CRITICAL] Repair Type II.** It must penalize the *exclusion* of a false literal when the clause fires: + ```nim + else: # Type II + if cOut == 1'u8: + for lit in 0.. Penalty -> toward Include + net.states[si] = int16(min(int(net.states[si]) + 1, TM_N_STATES)) + ``` + Without this, the repaired Type I will keep clauses sparse but the TM cannot learn to *discriminate* (suppress false positives), so the correction will likely stay near 0 for the wrong reason. `[INFERRED]` + +3. **[HIGH] Restore the summation-target / resource-allocation probability** from Granmo Eqns. 8–11 (`(T − clip(v,−T,T))/(2T)`), or an explicit regression analogue, instead of the pure `|error|` heuristic. `tsetlin.nim`'s `TM_T` is currently only a normaliser. This is the paper's "margin" mechanism and is what distributes clauses across sub-patterns. + +4. **[MEDIUM] Re-examine the regression framing.** Granmo's TM is a binary classifier; this gun uses sign(error) as an implicit label and `|error|` as a dispatch probability. That adaptation is defensible (it mirrors a perceptron/sign update), but it should be validated *after* (1) and (2). A safer formulation is to predict the sign of the residual through the standard positive/negative clause machinery, and convert sign→pixels separately. + +5. **[LOW] Optional hygiene:** initialise states at the Exclude boundary and prune all-Exclude clauses (Granmo Algorithm 1 line 27), and consider a state range/`N` that matches `s`. + +**Validation protocol (recommended, not run):** before/after fix, log per-clause include counts over a 600-tick simulation (expect include count to fall from ~131 toward single digits) and confirm `(cx,cy) ≠ 0` on at least some ticks; then run the existing integration battle against `Linear` and `SittingDuck` and check the virtual-bullet hit count diverges from Linear's byte-for-byte. Do not touch `feedback`/trace plumbing — it is already correct. + +--- + +## 7. When TM-DEB-style delayed credit assignment *would* be the right tool + +Discounting an eligibility buffer is the right mechanism when the system **cannot** pair a past action to its outcome, so the only available signal is a sparse terminal reward. The gun does not have that property — `(fireTick, powerBin)` gives exact pairing. + +A future case in this repo where it *would* apply: a **learned movement / positioning module** rewarded only by the round outcome (win/loss or final survival), where no per-tick target exists and the credit for a win cannot be attributed to a specific earlier turn or thrust. There, an eligibility trace with `γ < 1` (or a proper TD(λ)/REINFORCE estimator) is the standard and appropriate tool, because the responsible action is genuinely ambiguous. Note that even then, Granmo's *local* Type I/II machinery would still need to be correct; TM-DEB is a wrapper around it, not a substitute. + +--- + +## 8. Bottom line for the orchestrator + +- **Paper:** untrustworthy (AI-generated, misattributed ref, missing algorithm core, illustrative results). Do not cite or adopt. +- **Gun:** the real bug is in `common_libs/guns/tsetlin.nim` → `tmLearnOne`: Type I ignores `cOut` (deleting Granmo's `c=0` exclude-drift), and Type II is dead code with the wrong direction. Fixes 1 and 2 above are the actionable items. +- **TM-DEB:** does not fix, would worsen long-range (power-3) learning by deleting exactly attributable feedback via `γ^Δt` (`4.4e-7` at `Δt=90`).