Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by Gemini Notebook') solves temporal credit assignment under delayed reward, which is not our problem. Our gun's failure is clause saturation (~131 of 1740 literals included per clause -> conjunction fires with probability ~2^-131 -> correction identically 0), and TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave the fixed point untouched and additionally delete the long-range feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85, and those are the shots whose lead matters most. Credibility signals recorded in the doc: reference [2] misattributes authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%' placeholders so the automata update is simply absent; Table 1 is titled 'Expected' and reports never-measured accuracies; Eq. 12 is not a faithful copy of Granmo's Lemma 2. The audit also produced the actionable result: a line-by-line diff of Granmo Table 2/3 feedback against tmLearnOne, identifying why the automata saturate - Type I never conditions on the clause output so it omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T) resource allocation is missing entirely.
27 KiB
TM-DEB Paper Assessment — Adopt as a fix for the Tsetlin gun?
Issue: N/A (read-only review requested by orchestrator)
Date: 2026-09-20
Paper under review: docs/papers/tm-deb-paper.pdf — "Temporal Credit Assignment in Tsetlin Machines via Discounted Eligibility Experience Buffers (TM-DEB)"
Sources used:
docs/papers/1804.01508v15.pdf(Granmo, "The Tsetlin Machine…", the authoritative source) — extracted to/tmp/granmo.txtdocs/papers/2609.06133v1.pdf(Kumar et al., "Compressed Recurrent Feedback in Tsetlin Machines…") — extracted to/tmp/rtm.txtcommon_libs/guns/tsetlin.nim(read in full)
Verification legend: [VERIFIED] = I read the exact equation/line and/or ran the arithmetic myself. [INFERRED] = my reasoning from the code/paper, not measured at runtime. Per task constraints I did not build, run battles, or edit any .nim file.
1. Verdict (one paragraph)
Do not adopt TM-DEB, and do not spend time on it. The document is an AI-generated theoretical sketch (self-identifying header "Generated by Gemini Notebook", no author, a misattributed reference, four consecutive algorithm steps that are literally the character %, and a results table whose title says "Expected" and whose numbers contradict the very paper it claims to adapt). It does not fix the gun, because the gun's failure is not a delayed-credit-assignment problem — the gun already has exact (fireTick, powerBin) eligibility traces, so its effective credit-assignment λ is 1. The failure is a bug in the Type I/Type II feedback rule inside tmLearnOne: (i) Type I ignores the clause output and always pushes true literals toward Include, deleting Granmo's (clause=0, literal=1) → 1/s toward Exclude counteracting force, and (ii) Type II is dead code — its guard cOut == 1 and its condition lits[lit] == 0 and st > 0 are mutually exclusive, so it can never execute. With no counteracting force at all, the automata ratchet toward Include until the clause is hopelessly over-specific (~131 literals/clause), and no clause ever fires. TM-DEB changes only the frequency of updates (γ^Δt), not their direction, so it cannot move the fixed point: with the buggy rule it converges to the same saturated clause, only slower — and for the power-3 shots that matter most (Δt 60–90 ticks) it reduces the update probability to 5.8e-5 … 4.4e-7, effectively deleting the exact feedback we now have. The correct fix is to implement Granmo's Table 2 and Table 3 faithfully.
2. Credibility assessment of the document
Strong, independently verifiable indicators that this is low-quality, AI-generated, unpublished material:
- Self-identifying provenance. Page 1: "Generated by Gemini Notebook", author line
gemini-notebook@workspace.local, date "September 19, 2026".[VERIFIED] - Reference [2] is misattributed.
[VERIFIED]- TM-DEB cites:
[2] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study," IEEE Transactions / ISTM, 2025. docs/papers/2609.06133v1.pdfhas exactly that title but is authored by Ankit Kumar, Utkarsh Raj, Rishad Shafik, Sudip Roy (/tmp/rtm.txtlines 1–6), arXiv:2609.06133v1. There is no "IEEE Transactions / ISTM".- The author list in TM-DEB's
[2]matches a different work that the RTM paper itself cites:/tmp/rtm.txtline 699 —[3] C. Xu, S. Duan, R. Shafik, and A. Yakovlev, "Recurrent Tsetlin Machine…". So TM-DEB appears to have fused one paper's title with another's author list.
- TM-DEB cites:
- Algorithm Specification 2 is broken. Steps 9–12 are literally
%in both the-rawand-layoutextractions (/tmp/tm-deb.txtlines 520–523;/tmp/tm-deb-layout.txtlines 332–336).[VERIFIED]This is the core of the paper — the actual automata update — and it is missing, so the algorithm is not reproducible. - No experiment was run. Table 1 is captioned "Expected Comparative Metric Framework". Section 7 is titled "Empirical Verification Protocol & Benchmarks" and says "we define a rigorous experimental protocol" — protocol language, no results, no seeds, no dataset sizes, no figures.
[VERIFIED] - Contradicts its own cited source. The real RTM study reports
61.47±6.74%/62.94±9.92%over 144 runs (/tmp/rtm.txtabstract). TM-DEB's Table 1 reports "Std RTM 52.1 ± 1.2".[VERIFIED]The numbers are not from that study. - Other references.
[1]Granmo arXiv:1804.01508,[3]Tsetlin 1961,[4]Narendra & Thathachar 1989 are real.[5](Tung & Kleinrock, "Using Finite State Automata to Produce Self-Optimization…") matches Granmo's reference[15](/tmp/granmo.txtline 3335) and is real.[VERIFIED]So the only unverifiable/misattributed reference is[2], but combined with items 1–5 the credibility is very low.
Conclusion: treat TM-DEB as an untrusted informal note, not as literature. Granmo's 1804.01508v15 is the authoritative source and is already in the repo.
3. T1 — Theorem-by-theorem verification
(a) Theorem 2 (Eq. 11) — "sign of expected payoff difference is discount-invariant"
Paper text (/tmp/tm-deb.txt):
- Eq. 11:
Sign(E[R|α2,Δt] − E[R|α1,Δt]) = Sign(E[R|α2,0] − E[R|α1,0]) - Eq. 13:
E[R|α2,Δt] = γ^Δt_d · E[R|α2,0]; Eq. 14 same forα1 - Eq. 15:
ΔE[R]_Δt = γ^Δt_d · (E[R|α2,0] − E[R|α1,0])
My verdict: I AGREE with your critique, with one clarification.
- "Category error" — correct, under Granmo's own convention. TM-DEB's Eq. 10 is
P_DEB(k,T) = γ^(T−k)·P_base(...). That scales the probability that feedback is dispatched, not the payoff values in Table 2/3. Granmo'sE[R|α](Lemma 1/2,/tmp/granmo.txtlines 1432–1565) is the expected reward-minus-penalty conditional on the clause receiving feedback — the resource-allocation factor(T−clip(v,−T,T))/(2T)is not folded into it. Under TM-DEB, conditional on dispatch the reward/penalty distribution is the same table, so the conditional expected payoff is unchanged:E[R|α,Δt] = E[R|α,0], notγ^Δt·E[R|α,0]. The paper silently switched to an unconditional per-step expected payoff (which is indeed scaled byγ^Δtbecause no-update steps contribute 0). So Eq. 13–14 are not a valid inference under the paper's own cite to Granmo.[VERIFIED against Granmo Lemmas 1–2] - "Tautology" — correct, and this is the stronger point. Even granting Eq. 13–14, Eq. 15 factors a strictly positive scalar out of a difference; a positive scalar can never flip a sign. So Theorem 2 proves sign-invariance for any positive multiplier, which is trivially true and says nothing about learning. The theorem is vacuous as a convergence guarantee.
- What Theorem 2 ignores (which is the actually harmful effect):
γ^Δtscales the learning rate — the number of updates each automaton receives. Tsetlin automata need enough updates to grow their state space and converge (Granmo Lemma 3,/tmp/granmo.txtline ~1634). Reducing update probability byγ^Δtslows convergence and, for largeΔt, starves the automaton of samples. Theorem 2 is silent on this trade-off.
(b) Lemma 1 (Eq. 16–18) — "SNR variance bound σ²_base/(1−γ²)"
Paper text: Eq. 17 σ²_DEB = Σ_{k=1}^{T} (γ^{T−k})² σ²_base = σ²_base Σ_{m=0}^{T−1} (γ²)^m, Eq. 18 limit σ²_base/(1−γ²).
My verdict: I AGREE with both parts of your critique.
- Covariance dropped. Eq. 17 uses
Var(Σ_k X_k) = Σ_k Var(X_k). Under TM-DEB every historical step's update is gated by the same terminal rewardR_T(Algorithm Spec 2 requires oneR_Tand computesy_k^efffrom it,/tmp/tm-deb.txtlines 508–515). The per-step feedback variables therefore share a common factor and are positively correlated; the covariance terms2Σ_{k<l} γ^{…} Cov(X_k,X_l)are silently dropped, so the bound is unjustified. Moreover the per-step variance under Bernoulli dispatch itself depends onλ·P_base, so it is not a fixedσ²_base.[VERIFIED] - Conflates two mechanisms. Granmo's vanishing-SNR is a function of team size
W:SNR = 1/(W·p(1−p))(Granmo Eq. 3,/tmp/granmo.txtlines 163–175). TM-DEB's Lemma 1 bounds a temporal accumulation over episode lengthTand never mentionsW. A temporal discount does nothing about theW·p(1−p)team-noise term. Granmo's actual cure for vanishing SNR is local feedback from the literal and clause values plus resource allocation (Granmo Remark 2,/tmp/granmo.txtline ~2632, and Theorem 4), not a discount.[VERIFIED] - Minor: the lemma is titled "SNR Variance Bound" but only bounds variance; it never bounds SNR (which would require the signal term
μ²).
(c) Theorem 3 — "Generalized-ordinal-potential property is preserved"
Paper text: "…every individual state transition in A_current aligns monotonically with increasing the expected episodic payoff E[R_T]. Hence, the ordinal potential property is preserved across temporal steps." (/tmp/tm-deb.txt lines 462–475)
My verdict: I AGREE — it is an assertion, not a proof, and the premise is false.
- It is two sentences ending in "Hence", with no case analysis. Compare Granmo's Theorem 4 (
/tmp/granmo.txtlines 2635–2700), which is a real proof broken into the false-negative and false-positive cases and explicitly restricted to summation targetT=1, positive clauses, and single-step classification accuracy.[VERIFIED] - The premise is false. Granmo's alignment depends on local feedback derived from the current input's literal and clause values (Granmo Remark 2: "each single automaton gets feedback directly from the value of its literal and the value of its clause"). TM-DEB instead relabels historical inputs with a terminal target: Algorithm Spec 2 step 5 sets
y_k^eff = ŷ_kifR_T=+1else1−ŷ_k. So a locally harmful decision that happened to precede eventual success is reinforced with Type I feedback. The historical(X_k, C_k)no longer determine whether the action at stepkwas correct, which is exactly the property Granmo's Theorem 4 relies on. The ordinal-potential correspondence with accuracy does not transfer.[VERIFIED by reading both proofs] - Additional (missed by the paper): TM-DEB stores clause outputs
C_kcomputed at stepkbut applies updates to the final automata stateA_current; the stored clause outputs and the live automata are misaligned by construction. This is arguably a worse problem than the "checkpointing divergence" Theorem 1 warns about.
(d) Algorithm Specification 2 incomplete — CONFIRMED
Steps 9–12 are literally % in both extractions: /tmp/tm-deb.txt lines 520–523 and /tmp/tm-deb-layout.txt lines 332–336. [VERIFIED] Not an extraction artifact.
What is therefore unspecified: after sampling Random() ≤ λ·P_base at step 8, the paper never says (i) whether the sampled event is Reward, Penalty, or Inaction (i.e., which cell of Granmo's Table 2/3), nor (ii) how that feedback maps to an automata state transition via F(·,·) (Eq. 2). In other words the entire operational core — the thing the paper claims to contribute — is absent. The algorithm cannot be reimplemented from the document.
(e) Table 1 — illustrative, not experimental
Verdict: CONFIRMED, these are illustrative/projected numbers. Evidence [VERIFIED]:
- Caption: "Expected Comparative Metric Framework across Delayed Reward Benchmarks" (
/tmp/tm-deb.txtline 552). - Section 7 text: "we define a rigorous experimental protocol comparing TM-DEB…" (
/tmp/tm-deb.txtline 533) — protocol language. There is no results section, no dataset description, no seeds, no compute, no figures. - The only table is Table 1; the document ends at the Conclusion.
- The numbers disagree with the source: real RTM study
61.47±6.74%/62.94±9.92%over 144 runs vs. Table 1 "Std RTM 52.1±1.2" (/tmp/rtm.txtabstract).[VERIFIED]
Note: The ± values give a false impression of measured variance.
(f) Reference [2] mismatch — CONFIRMED
- TM-DEB
[2]: "C. Xu, S. Duan, R. Shafik, and A. Yakovlev, 'Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study,' IEEE Transactions / ISTM, 2025." (/tmp/tm-deb.txtlines 586–589) - Actual
docs/papers/2609.06133v1.pdf: Kumar, Raj, Shafik, Roy, arXiv:2609.06133v1 (Sep 2026), same title (/tmp/rtm.txtlines 1–6, 102). - The author list in
[2]is taken from a different paper cited by the RTM study (/tmp/rtm.txtline 699).[VERIFIED] - No other reference is unverifiable:
[1]/[3]/[4]are standard, and[5]matches Granmo's[15](/tmp/granmo.txtline 3335).[VERIFIED]
(g) Math errors you MISSED
Correct as written [VERIFIED]:
- Eq. 2 (automaton transition) is correct and matches Granmo Eq. 2: exclude range
1..N→ Penalty+1, Reward−1; include rangeN+1..2N→ Penalty−1, Reward+1; boundaries saturate. (/tmp/tm-deb.txtlines ~112–130) - Eq. 3
C_j(X) = ∧_{l_k∈L_j} l_k— correct conjunction. - Eq. 4
v = Σ C^1_j − Σ C^0_j,ŷ = 1[v≥0]— matches Granmo Eq. 7. - Eq. 10
P_DEB = γ^(T−k)·P_base— well-formed probability (P_base ≤ 1,γ ≤ 1). Note it can only reduce the update count.
Additional errors found:
- Eq. 12 is not a faithful copy of Granmo's Lemma 2 (Include payoff). TM-DEB Eq. 12:
E[R|α2] = (s−1)/s·γ_e·P(X1_jk∩X1) − 1/s·γ_e·P(X1\X1_jk) − 1/s·(1−γ_e)·P(X0)Granmo's Include payoff (/tmp/granmo.txtlines 1540–1565) is:(s−1)/s·γ·P(X1_jk∩X1) + (s−1)/s·(1−γ)·P(X1_jk∩X0) − 1/s·γ·P(X1\X1_jk) − 1/s·(1−γ)·P(X0\X1_jk)TM-DEB drops the positive(1−γ_e)·P(X1_jk∩X0)term and replacesP(X0\X1_jk)withP(X0). SinceP(X0) ≥ P(X0\X1_jk), the penalty is overstated.[VERIFIED] - Buffer tuple inconsistency. Eq. 9 defines
B_t = (t, X_t, C_t(X_t), a_t)(4 fields,/tmp/tm-deb.txtline ~285), but Algorithm Spec 1 step 6 pushes(t, X_t, C_t, v_t, ŷ_t)(5 fields) and Algorithm Spec 2 step 3 reads(k, X_k, C_k, v_k, ŷ_k)(5 fields). The actiona_tfrom Eq. 9 is never the one used (step 5 usesŷ_k).[VERIFIED] - Discount is direction-blind. Eq. 10 multiplies the whole
P_base, which aggregates both reward and penalty cells of Table 2/3. Soγ^Δtcannot perform credit assignment in the directional sense — it is a pure forgetting/truncation knob applied identically to reinforcement and punishment. This is consistent with Theorem 2's vacuity but contradicts the paper's framing. - Theorem 1's "checkpointing" target is a strawman, and misses that TM-DEB itself replays stale clause outputs against a state that has since changed (see (c)).
[INFERRED]
Nothing in (g) changes the practical picture: TM-DEB adds no mechanism beyond what the gun already has, and its theory is unsound.
4. T2 — Does TM-DEB address the actual failure (clause saturation)? NO.
Mechanism. TM-DEB's only operational additions are (a) an eligibility buffer of (X_k, C_k) tuples — which the gun already has, keyed exactly by (fireTick, powerBin) — and (b) multiplying the feedback-dispatch probability by γ^(T−k) (Eq. 10). It does not change the update rule (Type I/II) — step 7 still says "Compute base feedback prob. P_base via Eqns. (8)–(11)" of Granmo.
Therefore: the direction of every update is unchanged, and the saturation fixed point is unchanged. Discounting only slows the approach to the same saturated state. For the buggy rule in tsetlin.nim, the drift toward Include is monotone, so no amount of probability scaling can reduce the ~131 included literals.
Concrete γ numbers for the paper's own γ = 0.85 [VERIFIED, computed]:
Δt (ticks) |
γ^Δt |
meaning |
|---|---|---|
| 23 | 2.38e-2 |
~1 in 42 updates survives |
| 36 | 2.88e-3 |
~1 in 347 |
| 60 | 5.82e-5 |
~1 in 17,180 |
| 90 | 4.44e-7 |
~1 in 2.25 million |
Half-life: Δt ≈ 4.27 ticks. 1/(1−γ²) = 3.60 (the paper's variance bound, i.e. only ~3.6 effective steps are remembered).
Why this is exactly backwards for this gun. The gun's virtual bullets resolve 23–90 ticks after firing (bulletSpeed(power) = 20 − 3·power in common_libs/gun_harness/gun_interface.nim; power-3 long shots reach the 90-tick regime, cf. the comment in tsetlin.nim about ~128 ticks). These long-range power-3 shots are precisely the ones whose prediction matters most and are the hardest to get right. TM-DEB would multiply their update probability by 5.8e-5 … 4.4e-7, i.e. delete the exact feedback that the current (fireTick, powerBin) ring now captures cleanly (trainedShots = 2141, traceMisses = 0).
Verdict: TM-DEB does not fix the gun, and would worsen the long-range (power-3) case. Its discount is a liability here, not a feature.
5. T3 — What we already have vs. what TM-DEB adds
tsetlin.nim stores, per fired virtual bullet, a TmTrace containing the input vector X_k and the clause-output cache C_k, keyed exactly by (fireTick, powerBin). When onResult fires, it retrieves the exact trace and applies the update. That is precisely TM-DEB's "experience buffer" — an eligibility trace — and the pairing is exact.
Therefore our effective eligibility λ = 1. TM-DEB's exponential discount γ^(T−k) exists to solve the classic RL temporal-credit-assignment problem: when you cannot identify which past action produced the reward, you must hedge by spreading credit over recent actions with a decay. In this gun, the virtual-bullet tracker can identify the responsible action exactly (the resolution carries fireTick and powerBin). Discounting would therefore discard the few, highly relevant long-delay samples in exchange for a recency prior we do not need. It is a workaround for missing information that we possess — hence information-losing in our case.
(One caveat: TM-DEB's credit assignment is also inexact in the other direction — it replays stale clause outputs against the final state. Our ring has the same stale-cache property, but at least it never throws away a correctly identified sample.)
6. T4 — Granmo's correct feedback vs. what tsetlin.nim implements
Reference: Granmo Table 2 (Type I) and Table 3 (Type II), /tmp/granmo.txt lines 705 and 723. In Granmo's transition (Eq. 2), for an automaton in the Exclude range a Reward moves to more Exclude and a Penalty moves toward Include; for the Include range it is reversed. The tables below are already collapsed to the resulting state move.
6.1 Granmo Table 2 (Type I) — collapsed
clause c |
literal l_k |
automaton action | feedback probs | effective state move |
|---|---|---|---|---|
| 1 | 1 | Include | P(Reward)=(s−1)/s, P(Inaction)=1/s |
+1 (toward Include) w.p. (s−1)/s |
| 1 | 1 | Exclude | P(Penalty)=(s−1)/s, P(Inaction)=1/s |
+1 (toward Include) w.p. (s−1)/s |
| 0 | 1 | Include | P(Penalty)=1/s, P(Inaction)=(s−1)/s |
−1 (toward Exclude) w.p. 1/s |
| 0 | 1 | Exclude | P(Reward)=1/s, P(Inaction)=(s−1)/s |
−1 (toward Exclude) w.p. 1/s |
| 1 | 0 | Exclude | P(Reward)=1/s, P(Inaction)=(s−1)/s |
−1 (toward Exclude) w.p. 1/s |
| 0 | 0 | Include/Exclude | penalty/reward 1/s |
−1 (toward Exclude) w.p. 1/s |
Collapsed closed form (independent of current action):
l_k == 1 and c == 1→ state+1w.p.(s−1)/sl_k == 1 and c == 0→ state−1w.p.1/sl_k == 0(anyc) → state−1w.p.1/s
6.2 Granmo Table 3 (Type II) — collapsed
The only non-Inaction cell is: c == 1 and l_k == 0 → penalize the Exclude action → state +1 (toward Include) w.p. 1.0. Everything else is Inaction. (This is what increases discrimination power by adding a zero-valued literal to a false-positive clause.)
6.3 What tsetlin.nim actually does
Type I branch (tsetlin.nim, tmLearnOne):
if (error > 0.0 and pol > 0.0) or (error < 0.0 and pol < 0.0):
for lit in 0..<TM_N_LITERALS:
if lits[lit] == 1'u8:
if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES)
else:
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)
- It never reads
cOut. Forl_k == 1it always moves+1w.p.(s−1)/s, for any clause output. - Correct for the
c == 1, l_k == 1row only. Wrong forc == 0, l_k == 1, where the table requires−1w.p.1/s. That missing−1is the entire counteracting force (Granmo calls it "non-matching input";/tmp/granmo.txtline ~640). l_k == 0→−1w.p.1/sis correct.
Type II branch:
else:
if cOut == 1'u8:
for lit in 0..<TM_N_LITERALS:
if lits[lit] == 0'u8:
... if st > 0: st = max(st - 1, -TM_N_STATES)
- Dead code.
tmEvalClausereturns1only if no included literal has value 0. SocOut == 1implies everyst > 0literal haslits[lit] == 1. The guardlits[lit] == 0 and st > 0is therefore unsatisfiable and the body never executes.[VERIFIED by readingtmEvalClause+tmLearnOne] - Even if it were reachable, it has the wrong direction: it penalizes an included false literal, whereas Granmo penalizes the exclusion of a false literal (i.e. increments an excluded false literal). And it is a no-op for the actual false-positive case.
Resource allocation: Granmo Eqns. 8–11 use (T − clip(v,−T,T))/(2T) (and (T + clip(...))/(2T)) to decide whether a clause gets feedback, distributing clauses across sub-patterns. tsetlin.nim instead uses pFeedback = min(1, |error|/(2·80)) — a regression-error heuristic. TM_T (=25) is used only to normalize the vote for the output, not as a summation target. So the resource-allocation term is absent.
6.4 Most likely concrete reason for saturation at ~131 literals/clause
With Type I ignoring c (always +1 for true literals w.p. (s−1)/s) and Type II dead, there is no force pushing a clause back toward exclusion. Each literal random-walks with drift p_true − 1/s per feedback event ([INFERRED, matches Granmo's balance algebra]), so any literal that is true more than 1/s of the time marches to the Include boundary and stays there. s = 1.5 gives threshold 1/s = 0.667; larger s (as swept: 2.5/4.0/8.0) lowers the threshold and makes it worse — consistent with the observation that the s sweep did not help. The clause accumulates ~131 literals and fires with probability ~∏ p(bit), effectively never. Hence vote = 0, (cx,cy) = (0,0), and the gun is byte-identical to Linear.
6.5 Prioritised fix list (recommendations only — no source edits made)
-
[CRITICAL] Make Type I condition on the clause output. Replace the unconditional include-drift with the collapsed Table 2:
# cOut is the cached clause output for this clause at fire time if lits[lit] == 1'u8: if cOut == 1'u8: # (c=1, lk=1): both actions rewarded/penalized toward Include if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_N_STATES) else: # (c=0, lk=1): both actions driven toward Exclude <-- MISSING TODAY if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES) else: # (lk=0, any c): driven toward Exclude if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_N_STATES)This single change restores the counteracting force and should stop the ratchet.
[INFERRED from Granmo Tables 2 + Eq. 2] -
[CRITICAL] Repair Type II. It must penalize the exclusion of a false literal when the clause fires:
else: # Type II if cOut == 1'u8: for lit in 0..<TM_N_LITERALS: if lits[lit] == 0'u8: let si = tmStateIdx(outIdx, c, lit) if net.states[si] <= 0: # Exclude action -> Penalty -> toward Include net.states[si] = int16(min(int(net.states[si]) + 1, TM_N_STATES))Without this, the repaired Type I will keep clauses sparse but the TM cannot learn to discriminate (suppress false positives), so the correction will likely stay near 0 for the wrong reason.
[INFERRED] -
[HIGH] Restore the summation-target / resource-allocation probability from Granmo Eqns. 8–11 (
(T − clip(v,−T,T))/(2T)), or an explicit regression analogue, instead of the pure|error|heuristic.tsetlin.nim'sTM_Tis currently only a normaliser. This is the paper's "margin" mechanism and is what distributes clauses across sub-patterns. -
[MEDIUM] Re-examine the regression framing. Granmo's TM is a binary classifier; this gun uses sign(error) as an implicit label and
|error|as a dispatch probability. That adaptation is defensible (it mirrors a perceptron/sign update), but it should be validated after (1) and (2). A safer formulation is to predict the sign of the residual through the standard positive/negative clause machinery, and convert sign→pixels separately. -
[LOW] Optional hygiene: initialise states at the Exclude boundary and prune all-Exclude clauses (Granmo Algorithm 1 line 27), and consider a state range/
Nthat matchess.
Validation protocol (recommended, not run): before/after fix, log per-clause include counts over a 600-tick simulation (expect include count to fall from ~131 toward single digits) and confirm (cx,cy) ≠ 0 on at least some ticks; then run the existing integration battle against Linear and SittingDuck and check the virtual-bullet hit count diverges from Linear's byte-for-byte. Do not touch feedback/trace plumbing — it is already correct.
7. When TM-DEB-style delayed credit assignment would be the right tool
Discounting an eligibility buffer is the right mechanism when the system cannot pair a past action to its outcome, so the only available signal is a sparse terminal reward. The gun does not have that property — (fireTick, powerBin) gives exact pairing.
A future case in this repo where it would apply: a learned movement / positioning module rewarded only by the round outcome (win/loss or final survival), where no per-tick target exists and the credit for a win cannot be attributed to a specific earlier turn or thrust. There, an eligibility trace with γ < 1 (or a proper TD(λ)/REINFORCE estimator) is the standard and appropriate tool, because the responsible action is genuinely ambiguous. Note that even then, Granmo's local Type I/II machinery would still need to be correct; TM-DEB is a wrapper around it, not a substitute.
8. Bottom line for the orchestrator
- Paper: untrustworthy (AI-generated, misattributed ref, missing algorithm core, illustrative results). Do not cite or adopt.
- Gun: the real bug is in
common_libs/guns/tsetlin.nim→tmLearnOne: Type I ignorescOut(deleting Granmo'sc=0exclude-drift), and Type II is dead code with the wrong direction. Fixes 1 and 2 above are the actionable items. - TM-DEB: does not fix, would worsen long-range (power-3) learning by deleting exactly attributable feedback via
γ^Δt(4.4e-7atΔt=90).