Files
SirRoboGarage/common_libs/tm_diag/README.md
T
SirStone ab8d383121 Automata metrics: settledness alone does NOT separate learning from fidgeting
Added the four automata-level metrics to the TM diagnostics kit (settledness,
clause diversity, churn, vote disagreement) plus a state histogram, a per-input
confidence table and a one-line health summary, and validated them on a
learnable-vs-noise pair.

STATE CONVENTIONS, read off OUR code rather than from memory:
  range [-nStates, nStates] as int16; nStates = 64 for tm_pattern, 32 for tsetlin
  initial value 0 = the Exclude boundary
  INCLUDE iff state > 0; EXCLUDE iff state <= 0
  flip boundary sits between state 0 and 1; commitment = abs(st)/nStates in [0,1]

=== THE GATE, AND A RESULT THAT MATTERS ===
Case A (learnable planted rule) vs Case B (shuffled labels), 49 bits, N=64:
  metric                    A (learnable)     B (shuffled)
  settledness mean              0.970            0.719
  churn flip/sample        0.000055 FALLING  0.000788 FLAT
  clause-change/sample       0.00263 falling   0.0595 flat
  diversity (Jaccard)           0.176            0.014
  disagreement                  0.003            0.298
  verdict                    settling        mixed (NOT settling)

**SETTLEDNESS ALONE DOES NOT WORK.** On noise the automata still COMMIT (0.719) -
they just commit to the wrong thing. The decisive separators are **churn TREND
(falling vs flat)** and **vote DISAGREEMENT (0.003 vs 0.298)**. Had we built only
the settledness metric - the one that seems most obvious - we would have been
misled. That is now recorded in the README.

INERTIA SWEEP: A vs B separate at N=16/32/64/128. **Raising N raises A's
commitment but does NOT reduce B's noise-fitting** - so more inertia does not
rescue a noise-fitting TM.

=== REAL READING ON THE SHIPPED GUN, AND THE INFERENCE IT SUPPORTS ===
tm_pattern GF head over the DrussGT fixtures: settledness 0.484 (settling),
diversity 0.267 (moderate), churn 0.094/100 FALLING, disagreement 0.145
(coherent). **VERDICT: SETTLING** - not fidgeting, not collapsed. Constant inputs
flagged: 38/39 (the known never-written bits) plus 19/36/37.
Context: pooled warm accuracy 35.72% vs 34.24% majority = +1.48pp.
So: **the old gun was NOT failing because of inertia or instability - it settled
properly and its settled rules still barely beat a lazy guess.** Its settledness
(0.484) is LOWER than both synthetic cases (0.97/0.72), which is the signature of
WEAK OR CONFLICTING SIGNAL rather than too much inertia.
CONCLUSION: **N and s are not the observed bottleneck. The target/representation
is.** That is exactly why the new design changes the target and the label
pipeline rather than sweeping knobs - and it means we should NOT spend effort on
an N/s sweep expecting it to fix anything.

Also adds `diag_automata_validation.nim` (Case A/B/C + inertia sweep) and
`test_tm_automata_diag.nim` (55 pure checks); `test_tm_diag` 48 and
`diag_synthetic` 17 still pass, plus all other guards. acceptance_offline_vs_online
was NOT run (it needs a live battle and there is no tm_diag dependency).

Caveat: churn on the real gun is a PROXY (a tm_core retrain over captured samples
in live order) because the live gun exposes no per-sample state trace; the other
metrics are read directly off the exported teams.
2026-09-22 21:40:27 +02:00

264 lines
12 KiB
Markdown

# tm_diag — the Tsetlin diagnostics kit
First-class, offline diagnostics for a Tsetlin-Machine head. Answers the two
questions that clause-reading alone cannot:
> **Is the TM learning badly, or is the data bad?**
Built before the new TM gun, so a bad design cannot hide behind the data and a
bad dataset cannot hide behind the model.
Everything is pure and offline — no battles, no Java, no harness.
## Files
| file | contents |
|---|---|
| `feature_spec.nim` | `FeatureSpec`, `describe`, `describeClause`, `draftTMSpec()` (49-bit draft), `tmPatternSpec()` (40-bit shipped encoding) |
| `tm_core.nim` | compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout |
| `diagnostics.nim` | the seven groups (re-exports the two above) |
Import everything with:
```nim
import tm_diag/diagnostics
```
## Task 1 — named features / clause rendering
```nim
let spec = draftTMSpec() # 49 bits, all one-hot
spec.nBits # 49
spec.describe(45) # "lat DEAD-ON -18..+18"
spec.describeLiteral(49 + 45) # "NOT lat DEAD-ON -18..+18"
spec.describeClause(@[0, 45], 2) # "IF dist-wall<50 AND lat DEAD-ON -18..+18 THEN class=2"
spec.describeClause(@[], 1) # "IF TRUE (empty clause) THEN class=1"
```
A block is added with `addBlock(name, count, bitNames?)`; a bit with no explicit
name renders as `blockName[k]`, and a single-bit block renders as its name.
`tmPatternSpec()` mirrors `guns/tm_pattern.nim`'s `tmBuildBits` exactly. **Its two
last bits (`UNUSED-38/39`) are a real bug**: `tmBuildBits` writes only 38 raw
bits into an `array[TM_NBITS=40, uint8]`, so bits 38 and 39 are always 0 and
their negations always 1. The kit reports them as constant dead inputs.
## Task 2 — the six groups
All functions take a trained `TmMachine` (or an externally supplied clause set)
plus `seq[DiagSample]` where `DiagSample.lits` is the pos-then-neg literal vector
and `DiagSample.label` the true class.
```nim
# build samples from raw bits
let s = makeSample(nBits, rawBits, label, order)
# 1. pre-flight DATA checks
let dc = dataChecks(labels, nClasses, threshold = 0.30)
# dc.classCounts, dc.classShares, dc.majorityClass, dc.majorityShare,
# dc.majorityAccuracy, dc.overThreshold, dc.flags
let sc = shuffledLabelControl(tmplMachine, samples) # (acc, majority, ...)
# 2. clause introspection
let infos = clauseInfo(m, samples, spec)
let summ = clauseSummary(infos) # empty / neverFired / length hist
for c in topClauses(infos, 10): echo c.text, " votes=", c.votes
for cb in clauseBalanceByClass(infos, nClasses): echo cb
for cl in 0..<nClasses: # recover the class rule
echo spec.describeClause(necessaryLiterals(m, samples, cl), cl)
# 3. per-feature contribution + DEAD-INPUT LIST
let contribs = featureContributions(m, samples, spec)
let dead = deadInputs(contribs) # weighted < 5% of top, or never used
let strict = neverUsedInputs(contribs) # appearances == 0
let consts = constantInputs(samples, nBits)
for c in rankedInputs(contribs)[0..<10]: echo c.name, " ", c.weighted
# 4. accuracy diagnostics
let ad = accuracyDiagnostics(m, samples)
# ad.acc, ad.majorityBaseline, ad.margin, ad.confusion,
# ad.perClassRecall/Precision, ad.predMajorityShare
# 5. learning curve
let lc = learningCurve(tmplMachine, train, eval, nPoints = 10)
# lc.points, lc.accs, lc.trend ("flat" | "rising" | "rising-then-falling ...")
# 6. ablation hooks
let base = ablateBaseline(tmplMachine, train, eval)
for r in ablateDropAllBlocks(tmplMachine, train, eval, spec): echo r.name, r.delta
let rs = ablateScrambleFeature(tmplMachine, train, eval, spec, bit)
```
`ablation` retrains a fresh machine per variant (the strongest form of "does this
input earn its bits"). Pass `baselineAcc` from `ablateBaseline` to avoid
recomputing it per block.
## Task 5 — the AUTOMATA level (settledness / diversity / churn / disagreement)
Task 2 reads the clauses. Task 5 reads the **automata inside them**: one
automaton per input bit per clause, each holding a state that says how
confident it is that its bit belongs in the clause. These are what tell us
whether the TM is locking onto the enemy or just fidgeting, and whether the
inertia `N` (the number of automata states) should go up or down.
### State conventions — from the ACTUAL code, not the textbook
Derived from `tm_core.nim` and `guns/tm_pattern.nim` (`tmEval`, `tmLearnDir`,
`tmNewTeam`, `resetMachine`):
| property | value |
|---|---|
| state range | `[-nStates, nStates]` (int16); `nStates` = 64 for `tm_pattern`, 32 for `tsetlin.nim` |
| initial value | `0` (both cores call it "the Exclude boundary") |
| INCLUDE | `state > 0` (`tmEval` / `clauseLits`) |
| EXCLUDE | `state <= 0` |
| flip boundary | BETWEEN state `0` and state `1` — the middle of the range |
| `commitment(st)` | `abs(st) / nStates` in `[0,1]`: 0 on the boundary, 1 at either extreme |
So a flip is exactly a change in the predicate `state > 0`.
### The four metrics
1. **SETTLEDNESS** — per clause and overall, the mean `commitment` and the
fraction of automata settled at/above a threshold (default `0.5`). High and
**rising** as training proceeds is healthy; low means the clause is wavering
noise. `settledness(m, threshold)`, `settlednessTrend(early, late)`.
2. **CLAUSE DIVERSITY** — mean pairwise **Jaccard** of the included-literal sets,
within the same polarity. Low-to-moderate is healthy; ~1.0 = all clauses are
one rule in 50 hats; ~0 = memorising ticks. Empty clauses carry no rule and
are skipped by default. `clauseDiversity(m, skipEmpty = true)`, `jaccard(a,b)`.
3. **CHURN** — both levels, per training sample: the fraction of **automata**
crossing the flip boundary, and the fraction of **clauses** whose
included-literal set changed. High early and **falling** is healthy; flat-high
= fidgeting; zero from the start = never learned. `churnTrace(tmpl, samples,
epochs, seed, window, shuffle)`. `shuffle = false` replays the given
(temporal) order, mirroring a live gun.
4. **VOTE DISAGREEMENT** — per class, over the samples where it casts a vote:
the fraction of firing clauses whose polarity disagrees with the sign of the
class's total vote. Low is healthy. `voteDisagreement(m, samples)`.
### The three readouts that make it readable
- **state histogram** — the distribution of automata states across the range:
`stateHistogram(m, nBins)`, `histogramText(h)`.
- **per-input confidence table** — for each bit, the mean commitment of its
automata (and a `constant` flag for zero-variance inputs):
`perInputConfidence(m, spec, samples, threshold)`,
`rankedInputConfidence(conf)`.
- **one-line health summary** — `healthLine(ad)` gives
`settledness / diversity / churn trend / disagreement`, each with its own
verdict word (`settling`, `fidgeting`, `frozen`, `coherent`, ...), and
`automataVerdict(ad)` reduces the trajectory to one word. `formatAutomataReport(ad)`
prints everything.
### API
```nim
import tm_diag/diagnostics
let ad = automataDiagnostics(tmpl, samples, spec,
epochs = 15, seed = 777,
settleThreshold = 0.5, nHistBins = 9,
window = 100, measureChurn = true)
# ad.machine, ad.settledness, ad.diversity, ad.churn, ad.disagreement,
# ad.histogram, ad.inputConfidence, ad.summary
echo ad.summary # one line
echo automataVerdict(ad) # "settling" | "fidgeting" | "collapsed" | ...
echo formatAutomataReport(ad) # the full readout
```
For a machine you already have (e.g. the shipped gun's `exportTeams()`), call
`settledness` / `clauseDiversity` / `stateHistogram` / `perInputConfidence` /
`voteDisagreement` directly and `churnTrace` on a fresh copy.
### Validation — Case A (learnable) vs Case B (noise)
`common_libs/tests/diag_automata_validation.nim` uses the planted rule from
`diag_synthetic.nim` and the SAME inputs with shuffled labels. Measured
(`nBits=49`, 3 classes, 40 clauses, `nStates=64`, 3000 samples, 15 epochs):
| metric | Case A (learnable) | Case B (noise) |
|---|---|---|
| settledness mean | **0.970** | 0.719 |
| settled fraction | 0.999 | 0.799 |
| churn trend | **falling** (`2.5e-4` -> `2e-6`) | **flat** (`1.0e-3` -> `7.2e-4`) |
| clause-change trend | falling (`0.011` -> `2.3e-4`) | flat (`0.069` -> `0.057`) |
| diversity (overall Jaccard) | 0.176 | 0.014 |
| disagreement | **0.003** | **0.298** |
| verdict | settling | mixed (not settling) |
Case A settledness RISES `0.475` (early prefix) -> `0.970` (full). The pair
**separates learning from fidgeting**: the decisive signals are the churn trend
(falling vs flat) and disagreement (0.003 vs 0.298). Settledness alone is NOT
enough — on noise the automata still commit (0.719), just to the wrong thing.
Case C (a forced-constant bit) is flagged: `perInputConfidence(...).constant`
and `constantInputs` both surface it. Note a constant bit can show HIGH
commitment (one literal is always 1), so the `constant` flag is what
disambiguates.
### Inertia sweep (`N` = nStates, 10 epochs)
The validation also sweeps `N` to see whether inertia moves the metrics:
| N | Case A settled | A churn | A dis. | Case B settled | B churn | B dis. |
|---|---|---|---|---|---|---|
| 16 | 0.904 | falling | 0.003 | 0.720 | flat | 0.363 |
| 32 | 0.943 | falling | 0.000 | 0.700 | flat | 0.388 |
| 64 | 0.970 | falling | 0.003 | 0.715 | flat | 0.368 |
| 128 | 0.971 | falling | 0.000 | 0.660 | flat | 0.304 |
A and B are separated at EVERY N (churn falling vs flat, disagreement low vs
high). Raising N only raises Case A's commitment (0.90 -> 0.97); it does NOT
reduce noise-fitting in Case B. So **inertia is not the discriminator** the
metrics identify — the churn trend and disagreement are.
### Real reading — the shipped `tm_pattern` GF head
`diag_tm_pattern_offline.nim` over the committed DrussGT fixtures (automata read
directly off the exported teams from `tr_drussgt_vs_modularbot`; churn measured
on a `tm_core` temporal one-pass proxy, live-order):
- **settledness 0.484** (`settling`), settled fraction 0.453
- **diversity 0.267** (`moderate`) — positive 0.170, negative 0.322
- **churn 0.094/100 per sample, FALLING** (0.00208 -> 0.00048); clause-change 5.31/100, falling
- **disagreement 0.145** (`coherent`)
- **verdict: settling** — not fidgeting, not collapsed
- constant inputs flagged: bits 38/39 (the known never-written ones) plus 19/36/37 in this fixture
- context: pooled warm accuracy 35.72% vs the 34.24% majority = **+1.48pp**
The gun settles onto the within-battle labels but its settled rules barely beat
the majority class, which points at the TARGET / representation rather than the
inertia `N`. See the Task 5 report for the inertia discussion.
## The default-off real-gun hook
`guns/tm_pattern.nim` gained only additive, default-off instrumentation:
```nim
g.diagCapture = true # default false; no behaviour change when false
# ... replay ...
g.diagSamples # seq[TmDiagSample] (literal vector + label)
g.exportTeams() # read-only GF clause teams
g.exportRadTeams(); g.exportRevTeams()
```
To introspect an externally trained clause set:
```nim
let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
g.exportTeams())
```
## Running the demos / tests
```sh
nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim # 48 pure unit checks
nim c -r --path:common_libs common_libs/tests/test_tm_automata_diag.nim # 55 automata-metric unit checks
nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof (17 checks)
nim c -r -d:release --path:common_libs common_libs/tests/diag_automata_validation.nim # Task 5 A/B/C proof
nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim # Task 4 real reading
```
See `common_libs/tests/diag_synthetic.nim` for the ground-truth validation and
`common_libs/tests/diag_tm_pattern_offline.nim` for the real reading.