669f9acd41
Verdict: the document (docs/papers/tm-deb-paper.pdf, 'Generated by Gemini Notebook') solves temporal credit assignment under delayed reward, which is not our problem. Our gun's failure is clause saturation (~131 of 1740 literals included per clause -> conjunction fires with probability ~2^-131 -> correction identically 0), and TM-DEB only scales update FREQUENCY by gamma^dt, so it would leave the fixed point untouched and additionally delete the long-range feedback we need: gamma^90 = 4.4e-7 at the paper's own gamma=0.85, and those are the shots whose lead matters most. Credibility signals recorded in the doc: reference [2] misattributes authors/venue/year; Algorithm Spec 2 steps 9-12 are literal '%' placeholders so the automata update is simply absent; Table 1 is titled 'Expected' and reports never-measured accuracies; Eq. 12 is not a faithful copy of Granmo's Lemma 2. The audit also produced the actionable result: a line-by-line diff of Granmo Table 2/3 feedback against tmLearnOne, identifying why the automata saturate - Type I never conditions on the clause output so it omits Granmo's c=0, lk=1 -> -1 w.p. 1/s counter-force (true literals ratchet toward Include), Type II's guard cOut==1 AND lits[lit]==0 is unsatisfiable and therefore dead code, and the (T - clip(v,-T,T))/(2T) resource allocation is missing entirely.