# tm_diag — the Tsetlin diagnostics kit First-class, offline diagnostics for a Tsetlin-Machine head. Answers the two questions that clause-reading alone cannot: > **Is the TM learning badly, or is the data bad?** Built before the new TM gun, so a bad design cannot hide behind the data and a bad dataset cannot hide behind the model. Everything is pure and offline — no battles, no Java, no harness. ## Files | file | contents | |---|---| | `feature_spec.nim` | `FeatureSpec`, `describe`, `describeClause`, `draftTMSpec()` (49-bit draft), `tmPatternSpec()` (40-bit shipped encoding) | | `tm_core.nim` | compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout | | `diagnostics.nim` | the seven groups (re-exports the two above) | Import everything with: ```nim import tm_diag/diagnostics ``` ## Task 1 — named features / clause rendering ```nim let spec = draftTMSpec() # 49 bits, all one-hot spec.nBits # 49 spec.describe(45) # "lat DEAD-ON -18..+18" spec.describeLiteral(49 + 45) # "NOT lat DEAD-ON -18..+18" spec.describeClause(@[0, 45], 2) # "IF dist-wall<50 AND lat DEAD-ON -18..+18 THEN class=2" spec.describeClause(@[], 1) # "IF TRUE (empty clause) THEN class=1" ``` A block is added with `addBlock(name, count, bitNames?)`; a bit with no explicit name renders as `blockName[k]`, and a single-bit block renders as its name. `tmPatternSpec()` mirrors `guns/tm_pattern.nim`'s `tmBuildBits` exactly. **Its two last bits (`UNUSED-38/39`) are a real bug**: `tmBuildBits` writes only 38 raw bits into an `array[TM_NBITS=40, uint8]`, so bits 38 and 39 are always 0 and their negations always 1. The kit reports them as constant dead inputs. ## Task 2 — the six groups All functions take a trained `TmMachine` (or an externally supplied clause set) plus `seq[DiagSample]` where `DiagSample.lits` is the pos-then-neg literal vector and `DiagSample.label` the true class. ```nim # build samples from raw bits let s = makeSample(nBits, rawBits, label, order) # 1. pre-flight DATA checks let dc = dataChecks(labels, nClasses, threshold = 0.30) # dc.classCounts, dc.classShares, dc.majorityClass, dc.majorityShare, # dc.majorityAccuracy, dc.overThreshold, dc.flags let sc = shuffledLabelControl(tmplMachine, samples) # (acc, majority, ...) # 2. clause introspection let infos = clauseInfo(m, samples, spec) let summ = clauseSummary(infos) # empty / neverFired / length hist for c in topClauses(infos, 10): echo c.text, " votes=", c.votes for cb in clauseBalanceByClass(infos, nClasses): echo cb for cl in 0.. 0` (`tmEval` / `clauseLits`) | | EXCLUDE | `state <= 0` | | flip boundary | BETWEEN state `0` and state `1` — the middle of the range | | `commitment(st)` | `abs(st) / nStates` in `[0,1]`: 0 on the boundary, 1 at either extreme | So a flip is exactly a change in the predicate `state > 0`. ### The four metrics 1. **SETTLEDNESS** — per clause and overall, the mean `commitment` and the fraction of automata settled at/above a threshold (default `0.5`). High and **rising** as training proceeds is healthy; low means the clause is wavering noise. `settledness(m, threshold)`, `settlednessTrend(early, late)`. 2. **CLAUSE DIVERSITY** — mean pairwise **Jaccard** of the included-literal sets, within the same polarity. Low-to-moderate is healthy; ~1.0 = all clauses are one rule in 50 hats; ~0 = memorising ticks. Empty clauses carry no rule and are skipped by default. `clauseDiversity(m, skipEmpty = true)`, `jaccard(a,b)`. 3. **CHURN** — both levels, per training sample: the fraction of **automata** crossing the flip boundary, and the fraction of **clauses** whose included-literal set changed. High early and **falling** is healthy; flat-high = fidgeting; zero from the start = never learned. `churnTrace(tmpl, samples, epochs, seed, window, shuffle)`. `shuffle = false` replays the given (temporal) order, mirroring a live gun. 4. **VOTE DISAGREEMENT** — per class, over the samples where it casts a vote: the fraction of firing clauses whose polarity disagrees with the sign of the class's total vote. Low is healthy. `voteDisagreement(m, samples)`. ### The three readouts that make it readable - **state histogram** — the distribution of automata states across the range: `stateHistogram(m, nBins)`, `histogramText(h)`. - **per-input confidence table** — for each bit, the mean commitment of its automata (and a `constant` flag for zero-variance inputs): `perInputConfidence(m, spec, samples, threshold)`, `rankedInputConfidence(conf)`. - **one-line health summary** — `healthLine(ad)` gives `settledness / diversity / churn trend / disagreement`, each with its own verdict word (`settling`, `fidgeting`, `frozen`, `coherent`, ...), and `automataVerdict(ad)` reduces the trajectory to one word. `formatAutomataReport(ad)` prints everything. ### API ```nim import tm_diag/diagnostics let ad = automataDiagnostics(tmpl, samples, spec, epochs = 15, seed = 777, settleThreshold = 0.5, nHistBins = 9, window = 100, measureChurn = true) # ad.machine, ad.settledness, ad.diversity, ad.churn, ad.disagreement, # ad.histogram, ad.inputConfidence, ad.summary echo ad.summary # one line echo automataVerdict(ad) # "settling" | "fidgeting" | "collapsed" | ... echo formatAutomataReport(ad) # the full readout ``` For a machine you already have (e.g. the shipped gun's `exportTeams()`), call `settledness` / `clauseDiversity` / `stateHistogram` / `perInputConfidence` / `voteDisagreement` directly and `churnTrace` on a fresh copy. ### Validation — Case A (learnable) vs Case B (noise) `common_libs/tests/diag_automata_validation.nim` uses the planted rule from `diag_synthetic.nim` and the SAME inputs with shuffled labels. Measured (`nBits=49`, 3 classes, 40 clauses, `nStates=64`, 3000 samples, 15 epochs): | metric | Case A (learnable) | Case B (noise) | |---|---|---| | settledness mean | **0.970** | 0.719 | | settled fraction | 0.999 | 0.799 | | churn trend | **falling** (`2.5e-4` -> `2e-6`) | **flat** (`1.0e-3` -> `7.2e-4`) | | clause-change trend | falling (`0.011` -> `2.3e-4`) | flat (`0.069` -> `0.057`) | | diversity (overall Jaccard) | 0.176 | 0.014 | | disagreement | **0.003** | **0.298** | | verdict | settling | mixed (not settling) | Case A settledness RISES `0.475` (early prefix) -> `0.970` (full). The pair **separates learning from fidgeting**: the decisive signals are the churn trend (falling vs flat) and disagreement (0.003 vs 0.298). Settledness alone is NOT enough — on noise the automata still commit (0.719), just to the wrong thing. Case C (a forced-constant bit) is flagged: `perInputConfidence(...).constant` and `constantInputs` both surface it. Note a constant bit can show HIGH commitment (one literal is always 1), so the `constant` flag is what disambiguates. ### Inertia sweep (`N` = nStates, 10 epochs) The validation also sweeps `N` to see whether inertia moves the metrics: | N | Case A settled | A churn | A dis. | Case B settled | B churn | B dis. | |---|---|---|---|---|---|---| | 16 | 0.904 | falling | 0.003 | 0.720 | flat | 0.363 | | 32 | 0.943 | falling | 0.000 | 0.700 | flat | 0.388 | | 64 | 0.970 | falling | 0.003 | 0.715 | flat | 0.368 | | 128 | 0.971 | falling | 0.000 | 0.660 | flat | 0.304 | A and B are separated at EVERY N (churn falling vs flat, disagreement low vs high). Raising N only raises Case A's commitment (0.90 -> 0.97); it does NOT reduce noise-fitting in Case B. So **inertia is not the discriminator** the metrics identify — the churn trend and disagreement are. ### Real reading — the shipped `tm_pattern` GF head `diag_tm_pattern_offline.nim` over the committed DrussGT fixtures (automata read directly off the exported teams from `tr_drussgt_vs_modularbot`; churn measured on a `tm_core` temporal one-pass proxy, live-order): - **settledness 0.484** (`settling`), settled fraction 0.453 - **diversity 0.267** (`moderate`) — positive 0.170, negative 0.322 - **churn 0.094/100 per sample, FALLING** (0.00208 -> 0.00048); clause-change 5.31/100, falling - **disagreement 0.145** (`coherent`) - **verdict: settling** — not fidgeting, not collapsed - constant inputs flagged: bits 38/39 (the known never-written ones) plus 19/36/37 in this fixture - context: pooled warm accuracy 35.72% vs the 34.24% majority = **+1.48pp** The gun settles onto the within-battle labels but its settled rules barely beat the majority class, which points at the TARGET / representation rather than the inertia `N`. See the Task 5 report for the inertia discussion. ## The default-off real-gun hook `guns/tm_pattern.nim` gained only additive, default-off instrumentation: ```nim g.diagCapture = true # default false; no behaviour change when false # ... replay ... g.diagSamples # seq[TmDiagSample] (literal vector + label) g.exportTeams() # read-only GF clause teams g.exportRadTeams(); g.exportRevTeams() ``` To introspect an externally trained clause set: ```nim let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S, g.exportTeams()) ``` ## Running the demos / tests ```sh nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim # 48 pure unit checks nim c -r --path:common_libs common_libs/tests/test_tm_automata_diag.nim # 55 automata-metric unit checks nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof (17 checks) nim c -r -d:release --path:common_libs common_libs/tests/diag_automata_validation.nim # Task 5 A/B/C proof nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim # Task 4 real reading ``` See `common_libs/tests/diag_synthetic.nim` for the ground-truth validation and `common_libs/tests/diag_tm_pattern_offline.nim` for the real reading.