ab8d383121
Added the four automata-level metrics to the TM diagnostics kit (settledness, clause diversity, churn, vote disagreement) plus a state histogram, a per-input confidence table and a one-line health summary, and validated them on a learnable-vs-noise pair. STATE CONVENTIONS, read off OUR code rather than from memory: range [-nStates, nStates] as int16; nStates = 64 for tm_pattern, 32 for tsetlin initial value 0 = the Exclude boundary INCLUDE iff state > 0; EXCLUDE iff state <= 0 flip boundary sits between state 0 and 1; commitment = abs(st)/nStates in [0,1] === THE GATE, AND A RESULT THAT MATTERS === Case A (learnable planted rule) vs Case B (shuffled labels), 49 bits, N=64: metric A (learnable) B (shuffled) settledness mean 0.970 0.719 churn flip/sample 0.000055 FALLING 0.000788 FLAT clause-change/sample 0.00263 falling 0.0595 flat diversity (Jaccard) 0.176 0.014 disagreement 0.003 0.298 verdict settling mixed (NOT settling) **SETTLEDNESS ALONE DOES NOT WORK.** On noise the automata still COMMIT (0.719) - they just commit to the wrong thing. The decisive separators are **churn TREND (falling vs flat)** and **vote DISAGREEMENT (0.003 vs 0.298)**. Had we built only the settledness metric - the one that seems most obvious - we would have been misled. That is now recorded in the README. INERTIA SWEEP: A vs B separate at N=16/32/64/128. **Raising N raises A's commitment but does NOT reduce B's noise-fitting** - so more inertia does not rescue a noise-fitting TM. === REAL READING ON THE SHIPPED GUN, AND THE INFERENCE IT SUPPORTS === tm_pattern GF head over the DrussGT fixtures: settledness 0.484 (settling), diversity 0.267 (moderate), churn 0.094/100 FALLING, disagreement 0.145 (coherent). **VERDICT: SETTLING** - not fidgeting, not collapsed. Constant inputs flagged: 38/39 (the known never-written bits) plus 19/36/37. Context: pooled warm accuracy 35.72% vs 34.24% majority = +1.48pp. So: **the old gun was NOT failing because of inertia or instability - it settled properly and its settled rules still barely beat a lazy guess.** Its settledness (0.484) is LOWER than both synthetic cases (0.97/0.72), which is the signature of WEAK OR CONFLICTING SIGNAL rather than too much inertia. CONCLUSION: **N and s are not the observed bottleneck. The target/representation is.** That is exactly why the new design changes the target and the label pipeline rather than sweeping knobs - and it means we should NOT spend effort on an N/s sweep expecting it to fix anything. Also adds `diag_automata_validation.nim` (Case A/B/C + inertia sweep) and `test_tm_automata_diag.nim` (55 pure checks); `test_tm_diag` 48 and `diag_synthetic` 17 still pass, plus all other guards. acceptance_offline_vs_online was NOT run (it needs a live battle and there is no tm_diag dependency). Caveat: churn on the real gun is a PROXY (a tm_core retrain over captured samples in live order) because the live gun exposes no per-sample state trace; the other metrics are read directly off the exported teams.
264 lines
12 KiB
Markdown
264 lines
12 KiB
Markdown
# tm_diag — the Tsetlin diagnostics kit
|
|
|
|
First-class, offline diagnostics for a Tsetlin-Machine head. Answers the two
|
|
questions that clause-reading alone cannot:
|
|
|
|
> **Is the TM learning badly, or is the data bad?**
|
|
|
|
Built before the new TM gun, so a bad design cannot hide behind the data and a
|
|
bad dataset cannot hide behind the model.
|
|
|
|
Everything is pure and offline — no battles, no Java, no harness.
|
|
|
|
## Files
|
|
|
|
| file | contents |
|
|
|---|---|
|
|
| `feature_spec.nim` | `FeatureSpec`, `describe`, `describeClause`, `draftTMSpec()` (49-bit draft), `tmPatternSpec()` (40-bit shipped encoding) |
|
|
| `tm_core.nim` | compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout |
|
|
| `diagnostics.nim` | the seven groups (re-exports the two above) |
|
|
|
|
Import everything with:
|
|
|
|
```nim
|
|
import tm_diag/diagnostics
|
|
```
|
|
|
|
## Task 1 — named features / clause rendering
|
|
|
|
```nim
|
|
let spec = draftTMSpec() # 49 bits, all one-hot
|
|
spec.nBits # 49
|
|
spec.describe(45) # "lat DEAD-ON -18..+18"
|
|
spec.describeLiteral(49 + 45) # "NOT lat DEAD-ON -18..+18"
|
|
spec.describeClause(@[0, 45], 2) # "IF dist-wall<50 AND lat DEAD-ON -18..+18 THEN class=2"
|
|
spec.describeClause(@[], 1) # "IF TRUE (empty clause) THEN class=1"
|
|
```
|
|
|
|
A block is added with `addBlock(name, count, bitNames?)`; a bit with no explicit
|
|
name renders as `blockName[k]`, and a single-bit block renders as its name.
|
|
|
|
`tmPatternSpec()` mirrors `guns/tm_pattern.nim`'s `tmBuildBits` exactly. **Its two
|
|
last bits (`UNUSED-38/39`) are a real bug**: `tmBuildBits` writes only 38 raw
|
|
bits into an `array[TM_NBITS=40, uint8]`, so bits 38 and 39 are always 0 and
|
|
their negations always 1. The kit reports them as constant dead inputs.
|
|
|
|
## Task 2 — the six groups
|
|
|
|
All functions take a trained `TmMachine` (or an externally supplied clause set)
|
|
plus `seq[DiagSample]` where `DiagSample.lits` is the pos-then-neg literal vector
|
|
and `DiagSample.label` the true class.
|
|
|
|
```nim
|
|
# build samples from raw bits
|
|
let s = makeSample(nBits, rawBits, label, order)
|
|
|
|
# 1. pre-flight DATA checks
|
|
let dc = dataChecks(labels, nClasses, threshold = 0.30)
|
|
# dc.classCounts, dc.classShares, dc.majorityClass, dc.majorityShare,
|
|
# dc.majorityAccuracy, dc.overThreshold, dc.flags
|
|
let sc = shuffledLabelControl(tmplMachine, samples) # (acc, majority, ...)
|
|
|
|
# 2. clause introspection
|
|
let infos = clauseInfo(m, samples, spec)
|
|
let summ = clauseSummary(infos) # empty / neverFired / length hist
|
|
for c in topClauses(infos, 10): echo c.text, " votes=", c.votes
|
|
for cb in clauseBalanceByClass(infos, nClasses): echo cb
|
|
for cl in 0..<nClasses: # recover the class rule
|
|
echo spec.describeClause(necessaryLiterals(m, samples, cl), cl)
|
|
|
|
# 3. per-feature contribution + DEAD-INPUT LIST
|
|
let contribs = featureContributions(m, samples, spec)
|
|
let dead = deadInputs(contribs) # weighted < 5% of top, or never used
|
|
let strict = neverUsedInputs(contribs) # appearances == 0
|
|
let consts = constantInputs(samples, nBits)
|
|
for c in rankedInputs(contribs)[0..<10]: echo c.name, " ", c.weighted
|
|
|
|
# 4. accuracy diagnostics
|
|
let ad = accuracyDiagnostics(m, samples)
|
|
# ad.acc, ad.majorityBaseline, ad.margin, ad.confusion,
|
|
# ad.perClassRecall/Precision, ad.predMajorityShare
|
|
|
|
# 5. learning curve
|
|
let lc = learningCurve(tmplMachine, train, eval, nPoints = 10)
|
|
# lc.points, lc.accs, lc.trend ("flat" | "rising" | "rising-then-falling ...")
|
|
|
|
# 6. ablation hooks
|
|
let base = ablateBaseline(tmplMachine, train, eval)
|
|
for r in ablateDropAllBlocks(tmplMachine, train, eval, spec): echo r.name, r.delta
|
|
let rs = ablateScrambleFeature(tmplMachine, train, eval, spec, bit)
|
|
```
|
|
|
|
`ablation` retrains a fresh machine per variant (the strongest form of "does this
|
|
input earn its bits"). Pass `baselineAcc` from `ablateBaseline` to avoid
|
|
recomputing it per block.
|
|
|
|
## Task 5 — the AUTOMATA level (settledness / diversity / churn / disagreement)
|
|
|
|
Task 2 reads the clauses. Task 5 reads the **automata inside them**: one
|
|
automaton per input bit per clause, each holding a state that says how
|
|
confident it is that its bit belongs in the clause. These are what tell us
|
|
whether the TM is locking onto the enemy or just fidgeting, and whether the
|
|
inertia `N` (the number of automata states) should go up or down.
|
|
|
|
### State conventions — from the ACTUAL code, not the textbook
|
|
|
|
Derived from `tm_core.nim` and `guns/tm_pattern.nim` (`tmEval`, `tmLearnDir`,
|
|
`tmNewTeam`, `resetMachine`):
|
|
|
|
| property | value |
|
|
|---|---|
|
|
| state range | `[-nStates, nStates]` (int16); `nStates` = 64 for `tm_pattern`, 32 for `tsetlin.nim` |
|
|
| initial value | `0` (both cores call it "the Exclude boundary") |
|
|
| INCLUDE | `state > 0` (`tmEval` / `clauseLits`) |
|
|
| EXCLUDE | `state <= 0` |
|
|
| flip boundary | BETWEEN state `0` and state `1` — the middle of the range |
|
|
| `commitment(st)` | `abs(st) / nStates` in `[0,1]`: 0 on the boundary, 1 at either extreme |
|
|
|
|
So a flip is exactly a change in the predicate `state > 0`.
|
|
|
|
### The four metrics
|
|
|
|
1. **SETTLEDNESS** — per clause and overall, the mean `commitment` and the
|
|
fraction of automata settled at/above a threshold (default `0.5`). High and
|
|
**rising** as training proceeds is healthy; low means the clause is wavering
|
|
noise. `settledness(m, threshold)`, `settlednessTrend(early, late)`.
|
|
2. **CLAUSE DIVERSITY** — mean pairwise **Jaccard** of the included-literal sets,
|
|
within the same polarity. Low-to-moderate is healthy; ~1.0 = all clauses are
|
|
one rule in 50 hats; ~0 = memorising ticks. Empty clauses carry no rule and
|
|
are skipped by default. `clauseDiversity(m, skipEmpty = true)`, `jaccard(a,b)`.
|
|
3. **CHURN** — both levels, per training sample: the fraction of **automata**
|
|
crossing the flip boundary, and the fraction of **clauses** whose
|
|
included-literal set changed. High early and **falling** is healthy; flat-high
|
|
= fidgeting; zero from the start = never learned. `churnTrace(tmpl, samples,
|
|
epochs, seed, window, shuffle)`. `shuffle = false` replays the given
|
|
(temporal) order, mirroring a live gun.
|
|
4. **VOTE DISAGREEMENT** — per class, over the samples where it casts a vote:
|
|
the fraction of firing clauses whose polarity disagrees with the sign of the
|
|
class's total vote. Low is healthy. `voteDisagreement(m, samples)`.
|
|
|
|
### The three readouts that make it readable
|
|
|
|
- **state histogram** — the distribution of automata states across the range:
|
|
`stateHistogram(m, nBins)`, `histogramText(h)`.
|
|
- **per-input confidence table** — for each bit, the mean commitment of its
|
|
automata (and a `constant` flag for zero-variance inputs):
|
|
`perInputConfidence(m, spec, samples, threshold)`,
|
|
`rankedInputConfidence(conf)`.
|
|
- **one-line health summary** — `healthLine(ad)` gives
|
|
`settledness / diversity / churn trend / disagreement`, each with its own
|
|
verdict word (`settling`, `fidgeting`, `frozen`, `coherent`, ...), and
|
|
`automataVerdict(ad)` reduces the trajectory to one word. `formatAutomataReport(ad)`
|
|
prints everything.
|
|
|
|
### API
|
|
|
|
```nim
|
|
import tm_diag/diagnostics
|
|
|
|
let ad = automataDiagnostics(tmpl, samples, spec,
|
|
epochs = 15, seed = 777,
|
|
settleThreshold = 0.5, nHistBins = 9,
|
|
window = 100, measureChurn = true)
|
|
# ad.machine, ad.settledness, ad.diversity, ad.churn, ad.disagreement,
|
|
# ad.histogram, ad.inputConfidence, ad.summary
|
|
echo ad.summary # one line
|
|
echo automataVerdict(ad) # "settling" | "fidgeting" | "collapsed" | ...
|
|
echo formatAutomataReport(ad) # the full readout
|
|
```
|
|
|
|
For a machine you already have (e.g. the shipped gun's `exportTeams()`), call
|
|
`settledness` / `clauseDiversity` / `stateHistogram` / `perInputConfidence` /
|
|
`voteDisagreement` directly and `churnTrace` on a fresh copy.
|
|
|
|
### Validation — Case A (learnable) vs Case B (noise)
|
|
|
|
`common_libs/tests/diag_automata_validation.nim` uses the planted rule from
|
|
`diag_synthetic.nim` and the SAME inputs with shuffled labels. Measured
|
|
(`nBits=49`, 3 classes, 40 clauses, `nStates=64`, 3000 samples, 15 epochs):
|
|
|
|
| metric | Case A (learnable) | Case B (noise) |
|
|
|---|---|---|
|
|
| settledness mean | **0.970** | 0.719 |
|
|
| settled fraction | 0.999 | 0.799 |
|
|
| churn trend | **falling** (`2.5e-4` -> `2e-6`) | **flat** (`1.0e-3` -> `7.2e-4`) |
|
|
| clause-change trend | falling (`0.011` -> `2.3e-4`) | flat (`0.069` -> `0.057`) |
|
|
| diversity (overall Jaccard) | 0.176 | 0.014 |
|
|
| disagreement | **0.003** | **0.298** |
|
|
| verdict | settling | mixed (not settling) |
|
|
|
|
Case A settledness RISES `0.475` (early prefix) -> `0.970` (full). The pair
|
|
**separates learning from fidgeting**: the decisive signals are the churn trend
|
|
(falling vs flat) and disagreement (0.003 vs 0.298). Settledness alone is NOT
|
|
enough — on noise the automata still commit (0.719), just to the wrong thing.
|
|
Case C (a forced-constant bit) is flagged: `perInputConfidence(...).constant`
|
|
and `constantInputs` both surface it. Note a constant bit can show HIGH
|
|
commitment (one literal is always 1), so the `constant` flag is what
|
|
disambiguates.
|
|
|
|
### Inertia sweep (`N` = nStates, 10 epochs)
|
|
|
|
The validation also sweeps `N` to see whether inertia moves the metrics:
|
|
|
|
| N | Case A settled | A churn | A dis. | Case B settled | B churn | B dis. |
|
|
|---|---|---|---|---|---|---|
|
|
| 16 | 0.904 | falling | 0.003 | 0.720 | flat | 0.363 |
|
|
| 32 | 0.943 | falling | 0.000 | 0.700 | flat | 0.388 |
|
|
| 64 | 0.970 | falling | 0.003 | 0.715 | flat | 0.368 |
|
|
| 128 | 0.971 | falling | 0.000 | 0.660 | flat | 0.304 |
|
|
|
|
A and B are separated at EVERY N (churn falling vs flat, disagreement low vs
|
|
high). Raising N only raises Case A's commitment (0.90 -> 0.97); it does NOT
|
|
reduce noise-fitting in Case B. So **inertia is not the discriminator** the
|
|
metrics identify — the churn trend and disagreement are.
|
|
|
|
### Real reading — the shipped `tm_pattern` GF head
|
|
|
|
`diag_tm_pattern_offline.nim` over the committed DrussGT fixtures (automata read
|
|
directly off the exported teams from `tr_drussgt_vs_modularbot`; churn measured
|
|
on a `tm_core` temporal one-pass proxy, live-order):
|
|
|
|
- **settledness 0.484** (`settling`), settled fraction 0.453
|
|
- **diversity 0.267** (`moderate`) — positive 0.170, negative 0.322
|
|
- **churn 0.094/100 per sample, FALLING** (0.00208 -> 0.00048); clause-change 5.31/100, falling
|
|
- **disagreement 0.145** (`coherent`)
|
|
- **verdict: settling** — not fidgeting, not collapsed
|
|
- constant inputs flagged: bits 38/39 (the known never-written ones) plus 19/36/37 in this fixture
|
|
- context: pooled warm accuracy 35.72% vs the 34.24% majority = **+1.48pp**
|
|
|
|
The gun settles onto the within-battle labels but its settled rules barely beat
|
|
the majority class, which points at the TARGET / representation rather than the
|
|
inertia `N`. See the Task 5 report for the inertia discussion.
|
|
|
|
## The default-off real-gun hook
|
|
|
|
`guns/tm_pattern.nim` gained only additive, default-off instrumentation:
|
|
|
|
```nim
|
|
g.diagCapture = true # default false; no behaviour change when false
|
|
# ... replay ...
|
|
g.diagSamples # seq[TmDiagSample] (literal vector + label)
|
|
g.exportTeams() # read-only GF clause teams
|
|
g.exportRadTeams(); g.exportRevTeams()
|
|
```
|
|
|
|
To introspect an externally trained clause set:
|
|
|
|
```nim
|
|
let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
|
|
g.exportTeams())
|
|
```
|
|
|
|
## Running the demos / tests
|
|
|
|
```sh
|
|
nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim # 48 pure unit checks
|
|
nim c -r --path:common_libs common_libs/tests/test_tm_automata_diag.nim # 55 automata-metric unit checks
|
|
nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof (17 checks)
|
|
nim c -r -d:release --path:common_libs common_libs/tests/diag_automata_validation.nim # Task 5 A/B/C proof
|
|
nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim # Task 4 real reading
|
|
```
|
|
|
|
See `common_libs/tests/diag_synthetic.nim` for the ground-truth validation and
|
|
`common_libs/tests/diag_tm_pattern_offline.nim` for the real reading.
|