Files
SirRoboGarage/common_libs/tm_diag/README.md
T
SirStone f9f8d84671 TM diagnostics kit: VALIDATED (finds a known dead input), and it found a real bug
Built `common_libs/tm_diag/` as a first-class offline diagnostics kit for Tsetlin
work, BEFORE writing the new gun - because we hit two data problems tonight that no
amount of reading the TM's clauses would have revealed (a 38.8% majority answer,
and 36-58% mislabelled training samples).

WHAT IT PROVIDES
- `feature_spec.nim`: a NAMED feature container, so a learned clause prints as a
  sentence (`IF near-wall AND bullet-dead-on AND turn-left(t-2) THEN class=3`)
  instead of "feature 17". Includes the 49-bit draft spec from the design session
  and the shipped 40-bit encoding.
- `tm_core.nim`: a compact deterministic Granmo multiclass TM with an
  INTROSPECTABLE clause layout (mirrors the tm_pattern core).
- `diagnostics.nim`, six groups: (1) pre-flight DATA checks + shuffled-label
  control, (2) clause introspection (readable dump, per-clause vote counts, empty
  and never-fired clauses, length distribution, per-class balance), (3)
  per-feature contribution with an explicit DEAD-INPUT LIST and a ranked
  most-valuable list, (4) accuracy vs the majority baseline with per-class
  precision/recall and pred-majority share, (5) learning curve, (6) ablation hooks
  (drop a block / scramble a bit).

=== TASK 3: THE VALIDATION THAT GATES EVERYTHING - PASSED WITH NUMBERS ===
A diagnostic we never checked is worthless, so the kit was tested on a synthetic
set with a PLANTED RULE (class2 = A and B, class1 = A and not B, class0 = not A),
a deliberately IRRELEVANT block (US, 9 bits) and a PURE-NOISE bit (17).
- majority baseline 60.63% (class0); over-30% correctly flagged
- **the planted rule is recovered EXACTLY** via `necessaryLiterals`:
    class0 IF NOT dist-wall<50 | class1 IF dist-wall<50 AND NOT lat DEAD-ON |
    class2 IF dist-wall<50 AND lat DEAD-ON
- **DEAD-INPUT LIST = all 9 US bits AND the noise bit 17**, while the planted bits
  0 and 45 are correctly NOT listed
- top contributors: bit0 w=1241.7, bit45 w=583.3, then 49.8 - a 12-25x gap, so the
  relevant bits are unmistakable
- **ABLATION: drop WALLS -39.47pp, drop BULLETS -19.33pp, drop US 0.00pp**,
  scramble A -42.00pp, scramble the noise bit 0.00pp
- shuffled-label control 60.40% vs majority 60.63% = -0.23pp -> no leak
So the kit reliably finds a known dead input and a known relevant one.

=== TASK 4: THE REAL READING, AND A BUG IN THE SHIPPED GUN ===
`tm_pattern` GF head, 6 DrussGT fixtures, pooled 250,745 samples:
- label balance c2 = **34.4%** (majority-heavy, flagged); accuracy **35.72%** vs
  majority **34.24%** -> margin **+1.48pp**. On `tr_drussgt_vs_crazy` it is BELOW
  majority (33.81% vs 37.72%, -3.92pp).
- 200 clauses: **27 empty, 45 never fired**, mean length 19.17, max 57. The
  majority class is starved (class2: 22 non-empty, 18 empty, only 2 positive
  fired). Class4 fires 11-24-literal clauses -> memorisation signature.
- **REPRESENTATION BUG FOUND (reported, not silently fixed):** `tmBuildBits` writes
  only 38 raw bits into `var bits: array[TM_NBITS=40, uint8]` - bits 38 and 39 are
  NEVER ASSIGNED, so they are always 0 and their negated literals are always 1.
  The kit's `constantInputs` confirms 38/39 are constant, and **`UNUSED-38`/`39`
  rank #6 and #8 in the most-valuable-inputs list** - i.e. the model's
  highest-usage inputs are information-free. That is a representation bug, not a
  display artefact, and it is a concrete mechanism for part of the poor learning.

DRAFT ENCODING CHECKED: the 49-bit draft is arithmetically consistent
(4+4=8 walls, 6+3=9 us, 3+5+3+3+3+3=20 motion, 5+7=12 bullets = 49). No draft
inconsistency.

Guards: test_tm_diag 48 (new), diag_synthetic 17 (new), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, acceptance_offline_vs_online 12/12; tm_pattern_learning
passes. The tm_pattern hook is additive and default-OFF (no behaviour change).

NOT YET INCLUDED (the automata metrics discussed for the next step): per-clause
automata settledness (distance from the flip point), clause diversity (pairwise
overlap), literal-set churn over time, and cross-clause vote disagreement. The kit
has clause-level diagnostics but not the automata-state ones.
2026-09-22 21:24:27 +02:00

4.9 KiB

tm_diag — the Tsetlin diagnostics kit

First-class, offline diagnostics for a Tsetlin-Machine head. Answers the two questions that clause-reading alone cannot:

Is the TM learning badly, or is the data bad?

Built before the new TM gun, so a bad design cannot hide behind the data and a bad dataset cannot hide behind the model.

Everything is pure and offline — no battles, no Java, no harness.

Files

file contents
feature_spec.nim FeatureSpec, describe, describeClause, draftTMSpec() (49-bit draft), tmPatternSpec() (40-bit shipped encoding)
tm_core.nim compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout
diagnostics.nim the six groups (re-exports the two above)

Import everything with:

import tm_diag/diagnostics

Task 1 — named features / clause rendering

let spec = draftTMSpec()          # 49 bits, all one-hot
spec.nBits                        # 49
spec.describe(45)                 # "lat DEAD-ON -18..+18"
spec.describeLiteral(49 + 45)     # "NOT lat DEAD-ON -18..+18"
spec.describeClause(@[0, 45], 2)  # "IF dist-wall<50 AND lat DEAD-ON -18..+18 THEN class=2"
spec.describeClause(@[], 1)       # "IF TRUE (empty clause) THEN class=1"

A block is added with addBlock(name, count, bitNames?); a bit with no explicit name renders as blockName[k], and a single-bit block renders as its name.

tmPatternSpec() mirrors guns/tm_pattern.nim's tmBuildBits exactly. Its two last bits (UNUSED-38/39) are a real bug: tmBuildBits writes only 38 raw bits into an array[TM_NBITS=40, uint8], so bits 38 and 39 are always 0 and their negations always 1. The kit reports them as constant dead inputs.

Task 2 — the six groups

All functions take a trained TmMachine (or an externally supplied clause set) plus seq[DiagSample] where DiagSample.lits is the pos-then-neg literal vector and DiagSample.label the true class.

# build samples from raw bits
let s = makeSample(nBits, rawBits, label, order)

# 1. pre-flight DATA checks
let dc = dataChecks(labels, nClasses, threshold = 0.30)
#   dc.classCounts, dc.classShares, dc.majorityClass, dc.majorityShare,
#   dc.majorityAccuracy, dc.overThreshold, dc.flags
let sc = shuffledLabelControl(tmplMachine, samples)   # (acc, majority, ...)

# 2. clause introspection
let infos = clauseInfo(m, samples, spec)
let summ  = clauseSummary(infos)          # empty / neverFired / length hist
for c in topClauses(infos, 10): echo c.text, " votes=", c.votes
for cb in clauseBalanceByClass(infos, nClasses): echo cb
for cl in 0..<nClasses:                    # recover the class rule
  echo spec.describeClause(necessaryLiterals(m, samples, cl), cl)

# 3. per-feature contribution + DEAD-INPUT LIST
let contribs = featureContributions(m, samples, spec)
let dead     = deadInputs(contribs)        # weighted < 5% of top, or never used
let strict   = neverUsedInputs(contribs)   # appearances == 0
let consts   = constantInputs(samples, nBits)
for c in rankedInputs(contribs)[0..<10]: echo c.name, " ", c.weighted

# 4. accuracy diagnostics
let ad = accuracyDiagnostics(m, samples)
#   ad.acc, ad.majorityBaseline, ad.margin, ad.confusion,
#   ad.perClassRecall/Precision, ad.predMajorityShare

# 5. learning curve
let lc = learningCurve(tmplMachine, train, eval, nPoints = 10)
#   lc.points, lc.accs, lc.trend  ("flat" | "rising" | "rising-then-falling ...")

# 6. ablation hooks
let base = ablateBaseline(tmplMachine, train, eval)
for r in ablateDropAllBlocks(tmplMachine, train, eval, spec): echo r.name, r.delta
let rs = ablateScrambleFeature(tmplMachine, train, eval, spec, bit)

ablation retrains a fresh machine per variant (the strongest form of "does this input earn its bits"). Pass baselineAcc from ablateBaseline to avoid recomputing it per block.

The default-off real-gun hook

guns/tm_pattern.nim gained only additive, default-off instrumentation:

g.diagCapture = true            # default false; no behaviour change when false
# ... replay ...
g.diagSamples                   # seq[TmDiagSample] (literal vector + label)
g.exportTeams()                 # read-only GF clause teams
g.exportRadTeams(); g.exportRevTeams()

To introspect an externally trained clause set:

let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
                         g.exportTeams())

Running the demos / tests

nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim          # 48 pure unit checks
nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim        # Task 3 proof
nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim  # Task 4 real reading

See common_libs/tests/diag_synthetic.nim for the ground-truth validation and common_libs/tests/diag_tm_pattern_offline.nim for the real reading.