# tm_diag — the Tsetlin diagnostics kit First-class, offline diagnostics for a Tsetlin-Machine head. Answers the two questions that clause-reading alone cannot: > **Is the TM learning badly, or is the data bad?** Built before the new TM gun, so a bad design cannot hide behind the data and a bad dataset cannot hide behind the model. Everything is pure and offline — no battles, no Java, no harness. ## Files | file | contents | |---|---| | `feature_spec.nim` | `FeatureSpec`, `describe`, `describeClause`, `draftTMSpec()` (49-bit draft), `tmPatternSpec()` (40-bit shipped encoding) | | `tm_core.nim` | compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout | | `diagnostics.nim` | the eight groups (re-exports the two above) | Import everything with: ```nim import tm_diag/diagnostics ``` ## Task 1b — the `s` specificity knob (direction SETTLED, from the code) `TM_S` / `m.sValue` is the Type I specificity knob. It appears in ONE place in each core — the Type I feedback branch of `tmLearnDir`: `common_libs/guns/tm_pattern.nim` lines 258-265 (identical in `guns/tsetlin.nim` lines 224-231 and `tm_diag/tm_core.nim` lines 110-121): ```nim if pol * d > 0.0: # Type I (Table 2) collapsed to the resulting state move: # c=1, lk=1 -> +1 (toward Include) w.p. (s-1)/s # c=0, lk=1 -> -1 (toward Exclude) w.p. 1/s <- the missing counter-force # lk=0 -> -1 (toward Exclude) w.p. 1/s for lit in 0..=3`: the specificity mechanism is switched off. **Practical usable range: `s` in roughly [1.5, 5].** `s=1.0` is degenerate and must be avoided. The shipped guns sit inside the range: `tm_pattern` uses `3.0`, `tsetlin.nim` uses `1.5`. The clause-shape checker's healthy-band verdict is the tool that tells you whether a given `s` produced a sane geometry. ## Task 1 — named features / clause rendering ```nim let spec = draftTMSpec() # 49 bits, all one-hot spec.nBits # 49 spec.describe(45) # "lat DEAD-ON -18..+18" spec.describeLiteral(49 + 45) # "NOT lat DEAD-ON -18..+18" spec.describeClause(@[0, 45], 2) # "IF dist-wall<50 AND lat DEAD-ON -18..+18 THEN class=2" spec.describeClause(@[], 1) # "IF TRUE (empty clause) THEN class=1" ``` A block is added with `addBlock(name, count, bitNames?)`; a bit with no explicit name renders as `blockName[k]`, and a single-bit block renders as its name. `tmPatternSpec()` mirrors `guns/tm_pattern.nim`'s `tmBuildBits` exactly. **Its two last bits (`UNUSED-38/39`) are a real bug**: `tmBuildBits` writes only 38 raw bits into an `array[TM_NBITS=40, uint8]`, so bits 38 and 39 are always 0 and their negations always 1. The kit reports them as constant dead inputs. ## Task 2 — the data / clause / feature / accuracy / curve / ablation groups All functions take a trained `TmMachine` (or an externally supplied clause set) plus `seq[DiagSample]` where `DiagSample.lits` is the pos-then-neg literal vector and `DiagSample.label` the true class. ```nim # build samples from raw bits let s = makeSample(nBits, rawBits, label, order) # 1. pre-flight DATA checks let dc = dataChecks(labels, nClasses, threshold = 0.30) # dc.classCounts, dc.classShares, dc.majorityClass, dc.majorityShare, # dc.majorityAccuracy, dc.overThreshold, dc.flags let sc = shuffledLabelControl(tmplMachine, samples) # (acc, majority, ...) # 2. clause introspection let infos = clauseInfo(m, samples, spec) let summ = clauseSummary(infos) # empty / neverFired / length hist for c in topClauses(infos, 10): echo c.text, " votes=", c.votes for cb in clauseBalanceByClass(infos, nClasses): echo cb for cl in 0.. healthyHi` = `too long`, `< collapsedMax` = `collapsed`, between `collapsedMax` and `healthyLo` = `short`, in band = `healthy`, no non-empty clauses = `n/a`. The band is a PARAMETER of the call. * **per-block contribution** — `blockLengthContributions(m, spec)`: for each `FeatureSpec` block, the literals it contributes across clauses (`totalLits`, `meanPerClause`, `share`, `clausesUsing`, `perClauseMax`). A block that contributes to every clause is dominating; a block contributing ~0 is dead (the block-level view of the dead-INPUT list). * **coverage / concentration** — `clauseCoverage(m, samples)`: the fraction of clauses that fire on a sample (`meanFiringFraction`), the participation-ratio `effectiveClauses = (sum v)^2 / sum v^2` (if 3 clauses cast most votes it reads ~3, not the configured count) and `top3Share`. The one-line `healthLine` now appends `shape= ()` alongside the settledness / diversity / churn / disagreement verdicts. ## Task 5 — the AUTOMATA level (settledness / diversity / churn / disagreement) Task 2 reads the clauses. Task 5 reads the **automata inside them**: one automaton per input bit per clause, each holding a state that says how confident it is that its bit belongs in the clause. These are what tell us whether the TM is locking onto the enemy or just fidgeting, and whether the inertia `N` (the number of automata states) should go up or down. ### State conventions — from the ACTUAL code, not the textbook Derived from `tm_core.nim` and `guns/tm_pattern.nim` (`tmEval`, `tmLearnDir`, `tmNewTeam`, `resetMachine`): | property | value | |---|---| | state range | `[-nStates, nStates]` (int16); `nStates` = 64 for `tm_pattern`, 32 for `tsetlin.nim` | | initial value | `0` (both cores call it "the Exclude boundary") | | INCLUDE | `state > 0` (`tmEval` / `clauseLits`) | | EXCLUDE | `state <= 0` | | flip boundary | BETWEEN state `0` and state `1` — the middle of the range | | `commitment(st)` | `abs(st) / nStates` in `[0,1]`: 0 on the boundary, 1 at either extreme | So a flip is exactly a change in the predicate `state > 0`. ### The four metrics 1. **SETTLEDNESS** — per clause and overall, the mean `commitment` and the fraction of automata settled at/above a threshold (default `0.5`). High and **rising** as training proceeds is healthy; low means the clause is wavering noise. `settledness(m, threshold)`, `settlednessTrend(early, late)`. 2. **CLAUSE DIVERSITY** — mean pairwise **Jaccard** of the included-literal sets, within the same polarity. Low-to-moderate is healthy; ~1.0 = all clauses are one rule in 50 hats; ~0 = memorising ticks. Empty clauses carry no rule and are skipped by default. `clauseDiversity(m, skipEmpty = true)`, `jaccard(a,b)`. 3. **CHURN** — both levels, per training sample: the fraction of **automata** crossing the flip boundary, and the fraction of **clauses** whose included-literal set changed. High early and **falling** is healthy; flat-high = fidgeting; zero from the start = never learned. `churnTrace(tmpl, samples, epochs, seed, window, shuffle)`. `shuffle = false` replays the given (temporal) order, mirroring a live gun. 4. **VOTE DISAGREEMENT** — per class, over the samples where it casts a vote: the fraction of firing clauses whose polarity disagrees with the sign of the class's total vote. Low is healthy. `voteDisagreement(m, samples)`. ### The three readouts that make it readable - **state histogram** — the distribution of automata states across the range: `stateHistogram(m, nBins)`, `histogramText(h)`. - **per-input confidence table** — for each bit, the mean commitment of its automata (and a `constant` flag for zero-variance inputs): `perInputConfidence(m, spec, samples, threshold)`, `rankedInputConfidence(conf)`. - **one-line health summary** — `healthLine(ad)` gives `settledness / diversity / churn trend / disagreement / shape`, each with its own verdict word (`settling`, `fidgeting`, `frozen`, `coherent`, `healthy`, `too long`, ...), and `automataVerdict(ad)` reduces the trajectory to one word. `formatAutomataReport(ad)` prints everything, including the clause-shape report when `ad.shape` was computed. ### API ```nim import tm_diag/diagnostics let ad = automataDiagnostics(tmpl, samples, spec, epochs = 15, seed = 777, settleThreshold = 0.5, nHistBins = 9, window = 100, measureChurn = true) # ad.machine, ad.settledness, ad.diversity, ad.churn, ad.disagreement, # ad.histogram, ad.inputConfidence, ad.shape, ad.summary echo ad.summary # one line echo automataVerdict(ad) # "settling" | "fidgeting" | "collapsed" | ... echo formatAutomataReport(ad) # the full readout ``` For a machine you already have (e.g. the shipped gun's `exportTeams()`), call `settledness` / `clauseDiversity` / `stateHistogram` / `perInputConfidence` / `voteDisagreement` directly and `churnTrace` on a fresh copy. ### Validation — Case A (learnable) vs Case B (noise) `common_libs/tests/diag_automata_validation.nim` uses the planted rule from `diag_synthetic.nim` and the SAME inputs with shuffled labels. Measured (`nBits=49`, 3 classes, 40 clauses, `nStates=64`, 3000 samples, 15 epochs): | metric | Case A (learnable) | Case B (noise) | |---|---|---| | settledness mean | **0.970** | 0.719 | | settled fraction | 0.999 | 0.799 | | churn trend | **falling** (`2.5e-4` -> `2e-6`) | **flat** (`1.0e-3` -> `7.2e-4`) | | clause-change trend | falling (`0.011` -> `2.3e-4`) | flat (`0.069` -> `0.057`) | | diversity (overall Jaccard) | 0.176 | 0.014 | | disagreement | **0.003** | **0.298** | | verdict | settling | mixed (not settling) | Case A settledness RISES `0.475` (early prefix) -> `0.970` (full). The pair **separates learning from fidgeting**: the decisive signals are the churn trend (falling vs flat) and disagreement (0.003 vs 0.298). Settledness alone is NOT enough — on noise the automata still commit (0.719), just to the wrong thing. Case C (a forced-constant bit) is flagged: `perInputConfidence(...).constant` and `constantInputs` both surface it. Note a constant bit can show HIGH commitment (one literal is always 1), so the `constant` flag is what disambiguates. ### Inertia sweep (`N` = nStates, 10 epochs) The validation also sweeps `N` to see whether inertia moves the metrics: | N | Case A settled | A churn | A dis. | Case B settled | B churn | B dis. | |---|---|---|---|---|---|---| | 16 | 0.904 | falling | 0.003 | 0.720 | flat | 0.363 | | 32 | 0.943 | falling | 0.000 | 0.700 | flat | 0.388 | | 64 | 0.970 | falling | 0.003 | 0.715 | flat | 0.368 | | 128 | 0.971 | falling | 0.000 | 0.660 | flat | 0.304 | A and B are separated at EVERY N (churn falling vs flat, disagreement low vs high). Raising N only raises Case A's commitment (0.90 -> 0.97); it does NOT reduce noise-fitting in Case B. So **inertia is not the discriminator** the metrics identify — the churn trend and disagreement are. ### Real reading — the shipped `tm_pattern` GF head `diag_tm_pattern_offline.nim` over the committed DrussGT fixtures (automata read directly off the exported teams from `tr_drussgt_vs_modularbot`; churn measured on a `tm_core` temporal one-pass proxy, live-order): - **settledness 0.484** (`settling`), settled fraction 0.453 - **diversity 0.267** (`moderate`) — positive 0.170, negative 0.322 - **churn 0.094/100 per sample, FALLING** (0.00208 -> 0.00048); clause-change 5.31/100, falling - **disagreement 0.145** (`coherent`) - **verdict: settling** — not fidgeting, not collapsed - constant inputs flagged: bits 38/39 (the known never-written ones) plus 19/36/37 in this fixture - context: pooled warm accuracy 35.72% vs the 34.24% majority = **+1.48pp** **Clause-shape reading (group 8, same run):** `mean=19.17 median=16.00 p10=4.00 p90=44.80 std=13.89 min=1 max=57 nonEmpty=173 empty=27/200` -> verdict **`too long`**. Per polarity `posMean=19.80 / negMean=18.70`; per class `c0=11.4 c1=13.4 c2=31.2 c3=22.6 c4=21.6`. Per-block share is diffuse (no block dominates): wall-near 10.8%, distance-band 8.9%, flight-band 8.7%, radial-frac 7.6%, speed-band 7.0%, closing 6.8%, **UNUSED 6.4%** (the always-true negations of the never-written bits 38/39 are free padding), the rest 2-6%. Coverage: `firing/sample=48.53 (24.3%) effectiveClauses=97.85/200 top3Share=5.3%` — the voting is NOT concentrated on a few clauses. One-line health: ``` settledness=0.484 (settling) | diversity=0.267 (moderate) | churn=0.094/100 falling (settling) | disagreement=0.145 (coherent) | shape=19.17 (too long) ``` **Does `s` rescue this gun? MEASURED, no.** Recompiling the offline driver with `-d:TM_S_DEF=` (source untouched) retrains the gun end to end: | `s` | mean clause len | shape verdict | pooled warm acc | margin vs majority | |---|---|---|---|---| | 1.5 | 15.82 | too long | 32.22% | **-2.03pp** | | 2.0 | 14.99 | too long | 34.03% | **-0.22pp** | | 3.0 (shipped) | 19.17 | too long | 35.72% | **+1.48pp** | | 5.0 | 20.47 | too long | 34.72% | **+0.47pp** | Lowering `s` shrinks the clauses (15.8 at `s=1.5`) but makes accuracy WORSE; raising it pads them and also loses. The shipped `s=3.0` is the best of the four, and NO value gets the mean anywhere near the healthy 3-8 band. Combined with the settledness finding, the shape is consistent with **"no consistent short rule exists in this representation/target"** — the clauses are long, diffuse and padded, and the knob cannot fix a signal problem. The gun settles onto the within-battle labels but its settled rules barely beat the majority class, which points at the TARGET / representation rather than the inertia `N`. See the Task 5 report for the inertia discussion. ## The default-off real-gun hook `guns/tm_pattern.nim` gained only additive, default-off instrumentation: ```nim g.diagCapture = true # default false; no behaviour change when false # ... replay ... g.diagSamples # seq[TmDiagSample] (literal vector + label) g.exportTeams() # read-only GF clause teams g.exportRadTeams(); g.exportRevTeams() ``` To introspect an externally trained clause set: ```nim let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S, g.exportTeams()) ``` ## Running the demos / tests ```sh nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim # 48 pure unit checks nim c -r --path:common_libs common_libs/tests/test_tm_automata_diag.nim # 55 automata-metric unit checks nim c -r -d:release --path:common_libs common_libs/tests/test_tm_clause_shape.nim # 66 clause-shape unit + synthetic checks nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof (17 checks) nim c -r -d:release --path:common_libs common_libs/tests/diag_automata_validation.nim # Task 5 A/B/C proof nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim # Task 4 real reading + shape ``` See `common_libs/tests/diag_synthetic.nim` for the ground-truth validation, `common_libs/tests/test_tm_clause_shape.nim` for the clause-shape validation and `common_libs/tests/diag_tm_pattern_offline.nim` for the real reading.