docs(research): assess TM-DEB paper against the Tsetlin gun's real failure
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by Gemini Notebook') solves temporal credit assignment under delayed reward, which is not our problem. Our gun's failure is clause saturation (~131 of 1740 literals included per clause -> conjunction fires with probability ~2^-131 -> correction identically 0), and TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave the fixed point untouched and additionally delete the long-range feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85, and those are the shots whose lead matters most. Credibility signals recorded in the doc: reference [2] misattributes authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%' placeholders so the automata update is simply absent; Table 1 is titled 'Expected' and reports never-measured accuracies; Eq. 12 is not a faithful copy of Granmo's Lemma 2. The audit also produced the actionable result: a line-by-line diff of Granmo Table 2/3 feedback against tmLearnOne, identifying why the automata saturate - Type I never conditions on the clause output so it omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T) resource allocation is missing entirely.
This commit is contained in:
@@ -0,0 +1,259 @@
|
||||
# TM-DEB Paper Assessment — Adopt as a fix for the Tsetlin gun?
|
||||
|
||||
**Issue:** N/A (read-only review requested by orchestrator)
|
||||
**Date:** 2026-09-20
|
||||
**Paper under review:** `docs/papers/tm-deb-paper.pdf` — "Temporal Credit Assignment in Tsetlin Machines via Discounted Eligibility Experience Buffers (TM-DEB)"
|
||||
**Sources used:**
|
||||
- `docs/papers/1804.01508v15.pdf` (Granmo, "The Tsetlin Machine…", the authoritative source) — extracted to `/tmp/granmo.txt`
|
||||
- `docs/papers/2609.06133v1.pdf` (Kumar et al., "Compressed Recurrent Feedback in Tsetlin Machines…") — extracted to `/tmp/rtm.txt`
|
||||
- `common_libs/guns/tsetlin.nim` (read in full)
|
||||
|
||||
**Verification legend:** `[VERIFIED]` = I read the exact equation/line and/or ran the arithmetic myself. `[INFERRED]` = my reasoning from the code/paper, not measured at runtime. Per task constraints I did **not** build, run battles, or edit any `.nim` file.
|
||||
|
||||
---
|
||||
|
||||
## 1. Verdict (one paragraph)
|
||||
|
||||
**Do not adopt TM-DEB, and do not spend time on it.** The document is an AI-generated theoretical sketch (self-identifying header "Generated by Gemini Notebook", no author, a misattributed reference, four consecutive algorithm steps that are literally the character `%`, and a results table whose title says "Expected" and whose numbers contradict the very paper it claims to adapt). It does not fix the gun, because the gun's failure is **not** a delayed-credit-assignment problem — the gun already has exact `(fireTick, powerBin)` eligibility traces, so its effective credit-assignment `λ` is 1. The failure is a **bug in the Type I/Type II feedback rule inside `tmLearnOne`**: (i) Type I ignores the clause output and always pushes true literals toward *Include*, deleting Granmo's `(clause=0, literal=1) → 1/s toward Exclude` counteracting force, and (ii) Type II is dead code — its guard `cOut == 1` and its condition `lits[lit] == 0 and st > 0` are mutually exclusive, so it can never execute. With no counteracting force at all, the automata ratchet toward Include until the clause is hopelessly over-specific (~131 literals/clause), and no clause ever fires. TM-DEB changes only the *frequency* of updates (`γ^Δt`), not their *direction*, so it cannot move the fixed point: with the buggy rule it converges to the same saturated clause, only slower — and for the power-3 shots that matter most (`Δt` 60–90 ticks) it reduces the update probability to `5.8e-5 … 4.4e-7`, effectively deleting the exact feedback we now have. The correct fix is to implement Granmo's Table 2 and Table 3 faithfully.
|
||||
|
||||
---
|
||||
|
||||
## 2. Credibility assessment of the document
|
||||
|
||||
Strong, independently verifiable indicators that this is low-quality, AI-generated, unpublished material:
|
||||
|
||||
1. **Self-identifying provenance.** Page 1: "*Generated by Gemini Notebook*", author line `gemini-notebook@workspace.local`, date "September 19, 2026". `[VERIFIED]`
|
||||
2. **Reference [2] is misattributed.** `[VERIFIED]`
|
||||
- TM-DEB cites: `[2] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study," IEEE Transactions / ISTM, 2025.`
|
||||
- `docs/papers/2609.06133v1.pdf` has exactly that title but is authored by **Ankit Kumar, Utkarsh Raj, Rishad Shafik, Sudip Roy** (`/tmp/rtm.txt` lines 1–6), arXiv:2609.06133v1. There is no "IEEE Transactions / ISTM".
|
||||
- The author list in TM-DEB's `[2]` matches a *different* work that the RTM paper itself cites: `/tmp/rtm.txt` line 699 — `[3] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Recurrent Tsetlin Machine…"`. So TM-DEB appears to have fused one paper's title with another's author list.
|
||||
3. **Algorithm Specification 2 is broken.** Steps 9–12 are literally `%` in **both** the `-raw` and `-layout` extractions (`/tmp/tm-deb.txt` lines 520–523; `/tmp/tm-deb-layout.txt` lines 332–336). `[VERIFIED]` This is the core of the paper — the actual automata update — and it is missing, so the algorithm is not reproducible.
|
||||
4. **No experiment was run.** Table 1 is captioned "*Expected* Comparative Metric Framework". Section 7 is titled "*Empirical Verification Protocol & Benchmarks*" and says "we *define* a rigorous experimental protocol" — protocol language, no results, no seeds, no dataset sizes, no figures. `[VERIFIED]`
|
||||
5. **Contradicts its own cited source.** The real RTM study reports `61.47±6.74%` / `62.94±9.92%` over 144 runs (`/tmp/rtm.txt` abstract). TM-DEB's Table 1 reports "Std RTM 52.1 ± 1.2". `[VERIFIED]` The numbers are not from that study.
|
||||
6. **Other references.** `[1]` Granmo arXiv:1804.01508, `[3]` Tsetlin 1961, `[4]` Narendra & Thathachar 1989 are real. `[5]` (Tung & Kleinrock, "*Using Finite State Automata to Produce Self-Optimization…*") matches Granmo's reference `[15]` (`/tmp/granmo.txt` line 3335) and is real. `[VERIFIED]` So the only unverifiable/misattributed reference is `[2]`, but combined with items 1–5 the credibility is very low.
|
||||
|
||||
**Conclusion:** treat TM-DEB as an untrusted informal note, not as literature. Granmo's `1804.01508v15` is the authoritative source and is already in the repo.
|
||||
|
||||
---
|
||||
|
||||
## 3. T1 — Theorem-by-theorem verification
|
||||
|
||||
### (a) Theorem 2 (Eq. 11) — *"sign of expected payoff difference is discount-invariant"*
|
||||
|
||||
**Paper text** (`/tmp/tm-deb.txt`):
|
||||
- Eq. 11: `Sign(E[R|α2,Δt] − E[R|α1,Δt]) = Sign(E[R|α2,0] − E[R|α1,0])`
|
||||
- Eq. 13: `E[R|α2,Δt] = γ^Δt_d · E[R|α2,0]`; Eq. 14 same for `α1`
|
||||
- Eq. 15: `ΔE[R]_Δt = γ^Δt_d · (E[R|α2,0] − E[R|α1,0])`
|
||||
|
||||
**My verdict: I AGREE with your critique, with one clarification.**
|
||||
|
||||
- **"Category error" — correct, under Granmo's own convention.** TM-DEB's Eq. 10 is `P_DEB(k,T) = γ^(T−k)·P_base(...)`. That scales the *probability that feedback is dispatched*, not the payoff values in Table 2/3. Granmo's `E[R|α]` (Lemma 1/2, `/tmp/granmo.txt` lines 1432–1565) is the expected reward-minus-penalty **conditional on the clause receiving feedback** — the resource-allocation factor `(T−clip(v,−T,T))/(2T)` is *not* folded into it. Under TM-DEB, conditional on dispatch the reward/penalty distribution is the *same table*, so the conditional expected payoff is **unchanged**: `E[R|α,Δt] = E[R|α,0]`, not `γ^Δt·E[R|α,0]`. The paper silently switched to an *unconditional per-step* expected payoff (which is indeed scaled by `γ^Δt` because no-update steps contribute 0). So Eq. 13–14 are not a valid inference under the paper's own cite to Granmo. `[VERIFIED against Granmo Lemmas 1–2]`
|
||||
- **"Tautology" — correct, and this is the stronger point.** Even granting Eq. 13–14, Eq. 15 factors a *strictly positive scalar* out of a difference; a positive scalar can never flip a sign. So Theorem 2 proves sign-invariance for **any** positive multiplier, which is trivially true and says nothing about learning. The theorem is vacuous as a convergence guarantee.
|
||||
- **What Theorem 2 ignores** (which is the actually harmful effect): `γ^Δt` scales the *learning rate* — the number of updates each automaton receives. Tsetlin automata need enough updates to grow their state space and converge (Granmo Lemma 3, `/tmp/granmo.txt` line ~1634). Reducing update probability by `γ^Δt` slows convergence and, for large `Δt`, starves the automaton of samples. Theorem 2 is silent on this trade-off.
|
||||
|
||||
### (b) Lemma 1 (Eq. 16–18) — *"SNR variance bound `σ²_base/(1−γ²)`"*
|
||||
|
||||
**Paper text:** Eq. 17 `σ²_DEB = Σ_{k=1}^{T} (γ^{T−k})² σ²_base = σ²_base Σ_{m=0}^{T−1} (γ²)^m`, Eq. 18 limit `σ²_base/(1−γ²)`.
|
||||
|
||||
**My verdict: I AGREE with both parts of your critique.**
|
||||
- **Covariance dropped.** Eq. 17 uses `Var(Σ_k X_k) = Σ_k Var(X_k)`. Under TM-DEB every historical step's update is gated by the *same* terminal reward `R_T` (Algorithm Spec 2 requires one `R_T` and computes `y_k^eff` from it, `/tmp/tm-deb.txt` lines 508–515). The per-step feedback variables therefore share a common factor and are positively correlated; the covariance terms `2Σ_{k<l} γ^{…} Cov(X_k,X_l)` are silently dropped, so the bound is unjustified. Moreover the per-step variance under Bernoulli dispatch itself depends on `λ·P_base`, so it is not a fixed `σ²_base`. `[VERIFIED]`
|
||||
- **Conflates two mechanisms.** Granmo's vanishing-SNR is a function of **team size `W`**: `SNR = 1/(W·p(1−p))` (Granmo Eq. 3, `/tmp/granmo.txt` lines 163–175). TM-DEB's Lemma 1 bounds a **temporal** accumulation over episode length `T` and never mentions `W`. A temporal discount does nothing about the `W·p(1−p)` team-noise term. Granmo's actual cure for vanishing SNR is *local* feedback from the literal and clause values plus resource allocation (Granmo Remark 2, `/tmp/granmo.txt` line ~2632, and Theorem 4), not a discount. `[VERIFIED]`
|
||||
- Minor: the lemma is titled "SNR Variance Bound" but only bounds variance; it never bounds SNR (which would require the signal term `μ²`).
|
||||
|
||||
### (c) Theorem 3 — *"Generalized-ordinal-potential property is preserved"*
|
||||
|
||||
**Paper text:** "…every individual state transition in `A_current` aligns monotonically with increasing the expected episodic payoff `E[R_T]`. Hence, the ordinal potential property is preserved across temporal steps." (`/tmp/tm-deb.txt` lines 462–475)
|
||||
|
||||
**My verdict: I AGREE — it is an assertion, not a proof, and the premise is false.**
|
||||
- It is two sentences ending in "Hence", with no case analysis. Compare Granmo's Theorem 4 (`/tmp/granmo.txt` lines 2635–2700), which is a real proof broken into the false-negative and false-positive cases and explicitly restricted to summation target `T=1`, positive clauses, and **single-step** classification accuracy. `[VERIFIED]`
|
||||
- **The premise is false.** Granmo's alignment depends on local feedback derived from the *current input's* literal and clause values (Granmo Remark 2: "each single automaton gets feedback directly from the value of its literal and the value of its clause"). TM-DEB instead relabels *historical* inputs with a *terminal* target: Algorithm Spec 2 step 5 sets `y_k^eff = ŷ_k` if `R_T=+1` else `1−ŷ_k`. So a locally harmful decision that happened to precede eventual success is reinforced with Type I feedback. The historical `(X_k, C_k)` no longer determine whether the action at step `k` was correct, which is exactly the property Granmo's Theorem 4 relies on. The ordinal-potential correspondence with accuracy does not transfer. `[VERIFIED by reading both proofs]`
|
||||
- Additional (missed by the paper): TM-DEB stores clause outputs `C_k` computed at step `k` but applies updates to the *final* automata state `A_current`; the stored clause outputs and the live automata are misaligned by construction. This is arguably a worse problem than the "checkpointing divergence" Theorem 1 warns about.
|
||||
|
||||
### (d) Algorithm Specification 2 incomplete — **CONFIRMED**
|
||||
|
||||
Steps 9–12 are literally `%` in both extractions: `/tmp/tm-deb.txt` lines 520–523 and `/tmp/tm-deb-layout.txt` lines 332–336. `[VERIFIED]` Not an extraction artifact.
|
||||
|
||||
**What is therefore unspecified:** after sampling `Random() ≤ λ·P_base` at step 8, the paper never says (i) whether the sampled event is Reward, Penalty, or Inaction (i.e., which cell of Granmo's Table 2/3), nor (ii) how that feedback maps to an automata state transition via `F(·,·)` (Eq. 2). In other words the **entire operational core** — the thing the paper claims to contribute — is absent. The algorithm cannot be reimplemented from the document.
|
||||
|
||||
### (e) Table 1 — **illustrative, not experimental**
|
||||
|
||||
**Verdict: CONFIRMED, these are illustrative/projected numbers.** Evidence `[VERIFIED]`:
|
||||
- Caption: "*Expected* Comparative Metric Framework across Delayed Reward Benchmarks" (`/tmp/tm-deb.txt` line 552).
|
||||
- Section 7 text: "we *define* a rigorous experimental protocol comparing TM-DEB…" (`/tmp/tm-deb.txt` line 533) — protocol language. There is no results section, no dataset description, no seeds, no compute, no figures.
|
||||
- The only table is Table 1; the document ends at the Conclusion.
|
||||
- The numbers disagree with the source: real RTM study `61.47±6.74%`/`62.94±9.92%` over 144 runs vs. Table 1 "Std RTM 52.1±1.2" (`/tmp/rtm.txt` abstract). `[VERIFIED]`
|
||||
|
||||
**Note:** The `±` values give a false impression of measured variance.
|
||||
|
||||
### (f) Reference [2] mismatch — **CONFIRMED**
|
||||
|
||||
- TM-DEB `[2]`: "C. Xu, S. Duan, R. Shafik, and A. Yakovlev, 'Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study,' IEEE Transactions / ISTM, 2025." (`/tmp/tm-deb.txt` lines 586–589)
|
||||
- Actual `docs/papers/2609.06133v1.pdf`: Kumar, Raj, Shafik, Roy, arXiv:2609.06133v1 (Sep 2026), same title (`/tmp/rtm.txt` lines 1–6, 102).
|
||||
- The author list in `[2]` is taken from a *different* paper cited by the RTM study (`/tmp/rtm.txt` line 699). `[VERIFIED]`
|
||||
- No other reference is unverifiable: `[1]`/`[3]`/`[4]` are standard, and `[5]` matches Granmo's `[15]` (`/tmp/granmo.txt` line 3335). `[VERIFIED]`
|
||||
|
||||
### (g) Math errors you MISSED
|
||||
|
||||
**Correct as written `[VERIFIED]`:**
|
||||
- **Eq. 2 (automaton transition)** is correct and matches Granmo Eq. 2: exclude range `1..N` → Penalty `+1`, Reward `−1`; include range `N+1..2N` → Penalty `−1`, Reward `+1`; boundaries saturate. (`/tmp/tm-deb.txt` lines ~112–130)
|
||||
- **Eq. 3** `C_j(X) = ∧_{l_k∈L_j} l_k` — correct conjunction.
|
||||
- **Eq. 4** `v = Σ C^1_j − Σ C^0_j`, `ŷ = 1[v≥0]` — matches Granmo Eq. 7.
|
||||
- **Eq. 10** `P_DEB = γ^(T−k)·P_base` — well-formed probability (`P_base ≤ 1`, `γ ≤ 1`). Note it can only *reduce* the update count.
|
||||
|
||||
**Additional errors found:**
|
||||
|
||||
1. **Eq. 12 is not a faithful copy of Granmo's Lemma 2 (Include payoff).** TM-DEB Eq. 12:
|
||||
`E[R|α2] = (s−1)/s·γ_e·P(X1_jk∩X1) − 1/s·γ_e·P(X1\X1_jk) − 1/s·(1−γ_e)·P(X0)`
|
||||
Granmo's Include payoff (`/tmp/granmo.txt` lines 1540–1565) is:
|
||||
`(s−1)/s·γ·P(X1_jk∩X1) + (s−1)/s·(1−γ)·P(X1_jk∩X0) − 1/s·γ·P(X1\X1_jk) − 1/s·(1−γ)·P(X0\X1_jk)`
|
||||
TM-DEB **drops the positive `(1−γ_e)·P(X1_jk∩X0)` term** and **replaces `P(X0\X1_jk)` with `P(X0)`**. Since `P(X0) ≥ P(X0\X1_jk)`, the penalty is overstated. `[VERIFIED]`
|
||||
2. **Buffer tuple inconsistency.** Eq. 9 defines `B_t = (t, X_t, C_t(X_t), a_t)` (4 fields, `/tmp/tm-deb.txt` line ~285), but Algorithm Spec 1 step 6 pushes `(t, X_t, C_t, v_t, ŷ_t)` (5 fields) and Algorithm Spec 2 step 3 reads `(k, X_k, C_k, v_k, ŷ_k)` (5 fields). The action `a_t` from Eq. 9 is never the one used (step 5 uses `ŷ_k`). `[VERIFIED]`
|
||||
3. **Discount is direction-blind.** Eq. 10 multiplies the *whole* `P_base`, which aggregates both reward and penalty cells of Table 2/3. So `γ^Δt` cannot perform *credit assignment* in the directional sense — it is a pure forgetting/truncation knob applied identically to reinforcement and punishment. This is consistent with Theorem 2's vacuity but contradicts the paper's framing.
|
||||
4. **Theorem 1's "checkpointing" target is a strawman**, and misses that TM-DEB itself replays stale clause outputs against a state that has since changed (see (c)). `[INFERRED]`
|
||||
|
||||
**Nothing in (g) changes the practical picture: TM-DEB adds no mechanism beyond what the gun already has, and its theory is unsound.**
|
||||
|
||||
---
|
||||
|
||||
## 4. T2 — Does TM-DEB address the actual failure (clause saturation)? **NO.**
|
||||
|
||||
**Mechanism.** TM-DEB's only operational additions are (a) an eligibility buffer of `(X_k, C_k)` tuples — which the gun **already has**, keyed exactly by `(fireTick, powerBin)` — and (b) multiplying the feedback-dispatch probability by `γ^(T−k)` (Eq. 10). It does **not** change the update rule (Type I/II) — step 7 still says "Compute base feedback prob. `P_base` via Eqns. (8)–(11)" of Granmo.
|
||||
|
||||
**Therefore:** the *direction* of every update is unchanged, and the saturation fixed point is unchanged. Discounting only slows the approach to the same saturated state. For the buggy rule in `tsetlin.nim`, the drift toward Include is monotone, so no amount of probability scaling can reduce the ~131 included literals.
|
||||
|
||||
**Concrete `γ` numbers for the paper's own `γ = 0.85` `[VERIFIED, computed]`:**
|
||||
|
||||
| `Δt` (ticks) | `γ^Δt` | meaning |
|
||||
|---|---|---|
|
||||
| 23 | `2.38e-2` | ~1 in 42 updates survives |
|
||||
| 36 | `2.88e-3` | ~1 in 347 |
|
||||
| 60 | `5.82e-5` | ~1 in 17,180 |
|
||||
| 90 | `4.44e-7` | ~1 in 2.25 million |
|
||||
|
||||
Half-life: `Δt ≈ 4.27` ticks. `1/(1−γ²) = 3.60` (the paper's variance bound, i.e. only ~3.6 effective steps are remembered).
|
||||
|
||||
**Why this is exactly backwards for this gun.** The gun's virtual bullets resolve 23–90 ticks after firing (`bulletSpeed(power) = 20 − 3·power` in `common_libs/gun_harness/gun_interface.nim`; power-3 long shots reach the 90-tick regime, cf. the comment in `tsetlin.nim` about ~128 ticks). These long-range power-3 shots are precisely the ones whose prediction matters most and are the hardest to get right. TM-DEB would multiply their update probability by `5.8e-5 … 4.4e-7`, i.e. **delete** the exact feedback that the current `(fireTick, powerBin)` ring now captures cleanly (`trainedShots = 2141`, `traceMisses = 0`).
|
||||
|
||||
**Verdict:** TM-DEB **does not fix** the gun, and would **worsen** the long-range (`power-3`) case. Its discount is a liability here, not a feature.
|
||||
|
||||
---
|
||||
|
||||
## 5. T3 — What we already have vs. what TM-DEB adds
|
||||
|
||||
`tsetlin.nim` stores, per fired virtual bullet, a `TmTrace` containing the input vector `X_k` and the clause-output cache `C_k`, keyed exactly by `(fireTick, powerBin)`. When `onResult` fires, it retrieves the *exact* trace and applies the update. That is precisely TM-DEB's "experience buffer" — an eligibility trace — and the pairing is exact.
|
||||
|
||||
Therefore **our effective eligibility `λ = 1`**. TM-DEB's exponential discount `γ^(T−k)` exists to solve the classic RL temporal-credit-assignment problem: when you *cannot* identify which past action produced the reward, you must hedge by spreading credit over recent actions with a decay. In this gun, the virtual-bullet tracker *can* identify the responsible action exactly (the resolution carries `fireTick` and `powerBin`). Discounting would therefore discard the few, highly relevant long-delay samples in exchange for a recency prior we do not need. It is a workaround for missing information that we possess — hence **information-losing** in our case.
|
||||
|
||||
(One caveat: TM-DEB's credit assignment is *also* inexact in the other direction — it replays stale clause outputs against the final state. Our ring has the same stale-cache property, but at least it never throws away a correctly identified sample.)
|
||||
|
||||
---
|
||||
|
||||
## 6. T4 — Granmo's correct feedback vs. what `tsetlin.nim` implements
|
||||
|
||||
Reference: Granmo Table 2 (Type I) and Table 3 (Type II), `/tmp/granmo.txt` lines 705 and 723. In Granmo's transition (Eq. 2), for an automaton in the **Exclude** range a Reward moves to *more Exclude* and a Penalty moves *toward Include*; for the **Include** range it is reversed. The tables below are already collapsed to the resulting state move.
|
||||
|
||||
### 6.1 Granmo Table 2 (Type I) — collapsed
|
||||
|
||||
| clause `c` | literal `l_k` | automaton action | feedback probs | effective state move |
|
||||
|---|---|---|---|---|
|
||||
| 1 | 1 | Include | `P(Reward)=(s−1)/s`, `P(Inaction)=1/s` | **+1 (toward Include)** w.p. `(s−1)/s` |
|
||||
| 1 | 1 | Exclude | `P(Penalty)=(s−1)/s`, `P(Inaction)=1/s` | **+1 (toward Include)** w.p. `(s−1)/s` |
|
||||
| 0 | 1 | Include | `P(Penalty)=1/s`, `P(Inaction)=(s−1)/s` | **−1 (toward Exclude)** w.p. `1/s` |
|
||||
| 0 | 1 | Exclude | `P(Reward)=1/s`, `P(Inaction)=(s−1)/s` | **−1 (toward Exclude)** w.p. `1/s` |
|
||||
| 1 | 0 | Exclude | `P(Reward)=1/s`, `P(Inaction)=(s−1)/s` | **−1 (toward Exclude)** w.p. `1/s` |
|
||||
| 0 | 0 | Include/Exclude | penalty/reward `1/s` | **−1 (toward Exclude)** w.p. `1/s` |
|
||||
|
||||
**Collapsed closed form (independent of current action):**
|
||||
- `l_k == 1 and c == 1` → state `+1` w.p. `(s−1)/s`
|
||||
- `l_k == 1 and c == 0` → state `−1` w.p. `1/s`
|
||||
- `l_k == 0` (any `c`) → state `−1` w.p. `1/s`
|
||||
|
||||
### 6.2 Granmo Table 3 (Type II) — collapsed
|
||||
|
||||
The **only** non-Inaction cell is: `c == 1 and l_k == 0` → penalize the **Exclude** action → state **+1 (toward Include)** w.p. `1.0`. Everything else is Inaction. (This is what increases discrimination power by adding a zero-valued literal to a false-positive clause.)
|
||||
|
||||
### 6.3 What `tsetlin.nim` actually does
|
||||
|
||||
Type I branch (`tsetlin.nim`, `tmLearnOne`):
|
||||
```nim
|
||||
if (error > 0.0 and pol > 0.0) or (error < 0.0 and pol < 0.0):
|
||||
for lit in 0..<TM_N_LITERALS:
|
||||
if lits[lit] == 1'u8:
|
||||
if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES)
|
||||
else:
|
||||
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
|
||||
```
|
||||
- **It never reads `cOut`.** For `l_k == 1` it *always* moves `+1` w.p. `(s−1)/s`, for *any* clause output.
|
||||
- Correct for the `c == 1, l_k == 1` row only. **Wrong for `c == 0, l_k == 1`**, where the table requires `−1` w.p. `1/s`. That missing `−1` is the entire counteracting force (Granmo calls it "non-matching input"; `/tmp/granmo.txt` line ~640).
|
||||
- `l_k == 0` → `−1` w.p. `1/s` is correct.
|
||||
|
||||
Type II branch:
|
||||
```nim
|
||||
else:
|
||||
if cOut == 1'u8:
|
||||
for lit in 0..<TM_N_LITERALS:
|
||||
if lits[lit] == 0'u8:
|
||||
... if st > 0: st = max(st - 1, -TM_N_STATES)
|
||||
```
|
||||
- **Dead code.** `tmEvalClause` returns `1` only if no included literal has value 0. So `cOut == 1` implies every `st > 0` literal has `lits[lit] == 1`. The guard `lits[lit] == 0 and st > 0` is therefore unsatisfiable and the body never executes. `[VERIFIED by reading `tmEvalClause` + `tmLearnOne`]`
|
||||
- Even if it were reachable, it has the **wrong direction**: it penalizes an *included* false literal, whereas Granmo penalizes the *exclusion* of a false literal (i.e. increments an excluded false literal). And it is a no-op for the actual false-positive case.
|
||||
|
||||
Resource allocation: Granmo Eqns. 8–11 use `(T − clip(v,−T,T))/(2T)` (and `(T + clip(...))/(2T)`) to decide *whether* a clause gets feedback, distributing clauses across sub-patterns. `tsetlin.nim` instead uses `pFeedback = min(1, |error|/(2·80))` — a regression-error heuristic. `TM_T` (=25) is used only to normalize the vote for the output, not as a summation target. So the resource-allocation term is **absent**.
|
||||
|
||||
### 6.4 Most likely concrete reason for saturation at ~131 literals/clause
|
||||
|
||||
With Type I ignoring `c` (always `+1` for true literals w.p. `(s−1)/s`) and Type II dead, there is **no force pushing a clause back toward exclusion**. Each literal random-walks with drift `p_true − 1/s` per feedback event (`[INFERRED, matches Granmo's balance algebra]`), so any literal that is true more than `1/s` of the time marches to the Include boundary and stays there. `s = 1.5` gives threshold `1/s = 0.667`; larger `s` (as swept: 2.5/4.0/8.0) *lowers* the threshold and makes it worse — consistent with the observation that the `s` sweep did not help. The clause accumulates ~131 literals and fires with probability ~`∏ p(bit)`, effectively never. Hence `vote = 0`, `(cx,cy) = (0,0)`, and the gun is byte-identical to Linear.
|
||||
|
||||
### 6.5 Prioritised fix list (recommendations only — no source edits made)
|
||||
|
||||
1. **[CRITICAL] Make Type I condition on the clause output.** Replace the unconditional include-drift with the collapsed Table 2:
|
||||
```nim
|
||||
# cOut is the cached clause output for this clause at fire time
|
||||
if lits[lit] == 1'u8:
|
||||
if cOut == 1'u8:
|
||||
# (c=1, lk=1): both actions rewarded/penalized toward Include
|
||||
if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES)
|
||||
else:
|
||||
# (c=0, lk=1): both actions driven toward Exclude <-- MISSING TODAY
|
||||
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
|
||||
else:
|
||||
# (lk=0, any c): driven toward Exclude
|
||||
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
|
||||
```
|
||||
This single change restores the counteracting force and should stop the ratchet. `[INFERRED from Granmo Tables 2 + Eq. 2]`
|
||||
|
||||
2. **[CRITICAL] Repair Type II.** It must penalize the *exclusion* of a false literal when the clause fires:
|
||||
```nim
|
||||
else: # Type II
|
||||
if cOut == 1'u8:
|
||||
for lit in 0..<TM_N_LITERALS:
|
||||
if lits[lit] == 0'u8:
|
||||
let si = tmStateIdx(outIdx, c, lit)
|
||||
if net.states[si] <= 0: # Exclude action -> Penalty -> toward Include
|
||||
net.states[si] = int16(min(int(net.states[si]) + 1, TM_N_STATES))
|
||||
```
|
||||
Without this, the repaired Type I will keep clauses sparse but the TM cannot learn to *discriminate* (suppress false positives), so the correction will likely stay near 0 for the wrong reason. `[INFERRED]`
|
||||
|
||||
3. **[HIGH] Restore the summation-target / resource-allocation probability** from Granmo Eqns. 8–11 (`(T − clip(v,−T,T))/(2T)`), or an explicit regression analogue, instead of the pure `|error|` heuristic. `tsetlin.nim`'s `TM_T` is currently only a normaliser. This is the paper's "margin" mechanism and is what distributes clauses across sub-patterns.
|
||||
|
||||
4. **[MEDIUM] Re-examine the regression framing.** Granmo's TM is a binary classifier; this gun uses sign(error) as an implicit label and `|error|` as a dispatch probability. That adaptation is defensible (it mirrors a perceptron/sign update), but it should be validated *after* (1) and (2). A safer formulation is to predict the sign of the residual through the standard positive/negative clause machinery, and convert sign→pixels separately.
|
||||
|
||||
5. **[LOW] Optional hygiene:** initialise states at the Exclude boundary and prune all-Exclude clauses (Granmo Algorithm 1 line 27), and consider a state range/`N` that matches `s`.
|
||||
|
||||
**Validation protocol (recommended, not run):** before/after fix, log per-clause include counts over a 600-tick simulation (expect include count to fall from ~131 toward single digits) and confirm `(cx,cy) ≠ 0` on at least some ticks; then run the existing integration battle against `Linear` and `SittingDuck` and check the virtual-bullet hit count diverges from Linear's byte-for-byte. Do not touch `feedback`/trace plumbing — it is already correct.
|
||||
|
||||
---
|
||||
|
||||
## 7. When TM-DEB-style delayed credit assignment *would* be the right tool
|
||||
|
||||
Discounting an eligibility buffer is the right mechanism when the system **cannot** pair a past action to its outcome, so the only available signal is a sparse terminal reward. The gun does not have that property — `(fireTick, powerBin)` gives exact pairing.
|
||||
|
||||
A future case in this repo where it *would* apply: a **learned movement / positioning module** rewarded only by the round outcome (win/loss or final survival), where no per-tick target exists and the credit for a win cannot be attributed to a specific earlier turn or thrust. There, an eligibility trace with `γ < 1` (or a proper TD(λ)/REINFORCE estimator) is the standard and appropriate tool, because the responsible action is genuinely ambiguous. Note that even then, Granmo's *local* Type I/II machinery would still need to be correct; TM-DEB is a wrapper around it, not a substitute.
|
||||
|
||||
---
|
||||
|
||||
## 8. Bottom line for the orchestrator
|
||||
|
||||
- **Paper:** untrustworthy (AI-generated, misattributed ref, missing algorithm core, illustrative results). Do not cite or adopt.
|
||||
- **Gun:** the real bug is in `common_libs/guns/tsetlin.nim` → `tmLearnOne`: Type I ignores `cOut` (deleting Granmo's `c=0` exclude-drift), and Type II is dead code with the wrong direction. Fixes 1 and 2 above are the actionable items.
|
||||
- **TM-DEB:** does not fix, would worsen long-range (power-3) learning by deleting exactly attributable feedback via `γ^Δt` (`4.4e-7` at `Δt=90`).
|
||||
Reference in New Issue
Block a user