=== TASK 1: THE DIRECTION QUESTION, ANSWERED WITH A DEMONSTRATION ===
I told the user `s` controls clause length but refused to claim the DIRECTION,
because I had seen it described both ways. It is now read out of our own code -
one site per core, in the Type I branch of `tmLearnDir` (`guns/tm_pattern.nim:258`,
`guns/tsetlin.nim:224`, `tm_diag/tm_core.nim:110`):
if pol * d > 0.0:
if lits[lit] == 1:
if cOut == 1: if rand < (s-1)/s: st += 1 # toward Include, w.p. (s-1)/s
else: if rand < 1/s: st -= 1 # toward Exclude, w.p. 1/s
else: if rand < 1/s: st -= 1 # toward Exclude, w.p. 1/s
=> **HIGHER `s` GIVES LONGER CLAUSES.** The include step runs w.p. (s-1)/s
(rising with s); both exclude steps run w.p. 1/s (falling with s).
DEMONSTRATED (49-bit draft, planted 2-literal rule, 3000 train / 1500 eval):
s=1.0 len 1.51 acc 100% s=2.0 len 1.82 acc 100% s=5.0 len 3.02 acc 100%
s=1.5 len 1.42 acc 100% s=3.0 len 2.19 acc 100% s=10 len 4.04 acc 99.7%
s=20 len 5.17 acc 94.5%
WHY s=1.0 DEGENERATES: (s-1)/s = 0 so the include step NEVER fires while 1/s = 1
so BOTH exclude steps always fire - Type I can only remove literals, so a clause
can grow only through the Type II penalty. (On random labels that leaves 43/120
non-empty clauses vs 120/120 at s>=3.)
USABLE RANGE ~[1.5, 5]. tm_pattern uses 3.0; tsetlin uses 1.5.
=== TASK 4: WOULD `s` HELP THE SHIPPED GUN? NO - MEASURED ===
Recompiling the offline driver with -d:TM_S_DEF=<v> (source untouched) retrains
the gun end to end:
s mean len verdict warm acc margin vs majority
1.5 15.82 too long 32.22% -2.03pp
2.0 14.99 too long 34.03% -0.22pp
3.0* 19.17 too long 35.72% +1.48pp (*shipped)
5.0 20.47 too long 34.72% +0.47pp
Lowering `s` shrinks the clauses and makes accuracy WORSE; raising it pads them
and also loses. The shipped 3.0 is the best of the four, and **no value comes
near the healthy 3-8 band.** Combined with the settledness finding, the shape is
consistent with "NO CONSISTENT SHORT RULE EXISTS in this representation/target".
So the bottleneck is the SIGNAL - now confirmed from a THIRD independent angle
(settledness, churn trend, and clause shape). This is the measurement behind the
decision not to spend effort sweeping N or s.
=== TASK 3: AN HONEST CORRECTION TO MY OWN HYPOTHESIS ===
I predicted that random labels would produce `too long` clauses (the TM padding).
MEASURED: on this encoding noise reads as **short / `collapsed`** (mean 1.88,
median 2.0, acc 33.3%) - the TM FAILS TO COMMIT rather than padding. So "too
long" is not the noise signature, which means the shipped gun's 19.17 mean is not
explained by label noise. Worth knowing.
Adds diagnostic group 8: the clause-shape checker - full length distribution
(min/median/p10/p90/std), per-polarity and per-class breakdowns, a
`clauseShapeVerdict` against a parameterised healthy band (default 3-8),
per-BLOCK length contributions, and clause coverage (mean firing clauses,
effectiveClauses = participation ratio, top3Share). `healthLine` now appends
`shape=<mean> (<verdict>)`.
Validation: `test_tm_clause_shape` 66 checks. A planted 2-literal rule reads
`healthy` with the literals recovered exactly; per-block correctly names the
planted blocks (WALLS 43.0%, BULLETS 28.1%) and buries an irrelevant block (3.4%,
below its uniform 8.3% share); random labels read `collapsed`.
REAL READING, shipped gun: mean 19.17 / median 16.00 / p90 44.80 / max 57,
173 non-empty of 200, 27 empty => **`too long`**; coverage firing/sample 48.53
(24.3%), effectiveClauses 97.85/200, top3Share 5.3% (voting NOT concentrated);
per-block is diffuse with no dominator, EXCEPT **UNUSED 6.4%** - the always-true
negations of the never-written bits 38/39 acting as FREE PADDING, the same bug the
kit found earlier now visible as clause bloat.
Guards: test_tm_clause_shape 66 (new), test_tm_diag 48, test_tm_automata_diag 55,
diag_synthetic 17, diag_automata_validation 11, test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12. acceptance_offline_vs_online not run (needs a live
battle; no tm_diag dependency).
19 KiB
tm_diag — the Tsetlin diagnostics kit
First-class, offline diagnostics for a Tsetlin-Machine head. Answers the two questions that clause-reading alone cannot:
Is the TM learning badly, or is the data bad?
Built before the new TM gun, so a bad design cannot hide behind the data and a bad dataset cannot hide behind the model.
Everything is pure and offline — no battles, no Java, no harness.
Files
| file | contents |
|---|---|
feature_spec.nim |
FeatureSpec, describe, describeClause, draftTMSpec() (49-bit draft), tmPatternSpec() (40-bit shipped encoding) |
tm_core.nim |
compact deterministic Granmo Table 2/3 multiclass TM (mirrors the tm_pattern core), introspectable clause layout |
diagnostics.nim |
the eight groups (re-exports the two above) |
Import everything with:
import tm_diag/diagnostics
Task 1b — the s specificity knob (direction SETTLED, from the code)
TM_S / m.sValue is the Type I specificity knob. It appears in ONE place in
each core — the Type I feedback branch of tmLearnDir:
common_libs/guns/tm_pattern.nim lines 258-265 (identical in
guns/tsetlin.nim lines 224-231 and tm_diag/tm_core.nim lines 110-121):
if pol * d > 0.0:
# Type I (Table 2) collapsed to the resulting state move:
# c=1, lk=1 -> +1 (toward Include) w.p. (s-1)/s
# c=0, lk=1 -> -1 (toward Exclude) w.p. 1/s <- the missing counter-force
# lk=0 -> -1 (toward Exclude) w.p. 1/s
for lit in 0..<TM_NLITS:
var st = int(team[base + lit])
if lits[lit] == 1'u8:
if cOut == 1'u8:
if rand(1.0) < (TM_S - 1.0) / TM_S: st = min(st + 1, TM_NSTATES)
else:
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_NSTATES)
else:
if rand(1.0) < 1.0 / TM_S: st = max(st - 1, -TM_NSTATES)
team[base + lit] = int16(st)
Direction: HIGHER s makes clauses LONGER. The include step runs with
probability (s-1)/s (rising with s), while both exclude steps run with
probability 1/s (falling with s). So a larger s reinforces inclusions more
often and removes literals less often.
MEASURED on the 49-bit draft encoding, the planted 2-literal rule
(diag_synthetic dataset, 3000 train / 1500 eval, 40 clauses/class, N=64,
25 epochs):
s |
mean clause len | max | eval acc |
|---|---|---|---|
| 1.0 | 1.51 | 2 | 100.0% |
| 1.5 | 1.42 | 2 | 100.0% |
| 2.0 | 1.82 | 3 | 100.0% |
| 3.0 | 2.19 | 3 | 100.0% |
| 5.0 | 3.02 | 4 | 100.0% |
| 10.0 | 4.04 | 5 | 99.7% |
| 20.0 | 5.17 | 7 | 94.5% |
Why s = 1.0 is degenerate: at s = 1.0, (s-1)/s = 0.0 so the include
step is never taken (rand(1.0) < 0.0 is never true) and 1/s = 1.0 so BOTH
exclude steps are always taken. Type I can only ever REMOVE literals — a clause
can never grow through the normal path, only via the Type II penalty
(cOut==1 + wrong direction, which includes ABSENT literals). On the easy
synthetic the two-literal rule is still reachable through Type II (100%), but on
random labels s=1.0 leaves only 43/120 non-empty clauses vs 120/120 at s>=3:
the specificity mechanism is switched off.
Practical usable range: s in roughly [1.5, 5]. s=1.0 is degenerate and
must be avoided. The shipped guns sit inside the range: tm_pattern uses 3.0,
tsetlin.nim uses 1.5. The clause-shape checker's healthy-band verdict is the
tool that tells you whether a given s produced a sane geometry.
Task 1 — named features / clause rendering
let spec = draftTMSpec() # 49 bits, all one-hot
spec.nBits # 49
spec.describe(45) # "lat DEAD-ON -18..+18"
spec.describeLiteral(49 + 45) # "NOT lat DEAD-ON -18..+18"
spec.describeClause(@[0, 45], 2) # "IF dist-wall<50 AND lat DEAD-ON -18..+18 THEN class=2"
spec.describeClause(@[], 1) # "IF TRUE (empty clause) THEN class=1"
A block is added with addBlock(name, count, bitNames?); a bit with no explicit
name renders as blockName[k], and a single-bit block renders as its name.
tmPatternSpec() mirrors guns/tm_pattern.nim's tmBuildBits exactly. Its two
last bits (UNUSED-38/39) are a real bug: tmBuildBits writes only 38 raw
bits into an array[TM_NBITS=40, uint8], so bits 38 and 39 are always 0 and
their negations always 1. The kit reports them as constant dead inputs.
Task 2 — the data / clause / feature / accuracy / curve / ablation groups
All functions take a trained TmMachine (or an externally supplied clause set)
plus seq[DiagSample] where DiagSample.lits is the pos-then-neg literal vector
and DiagSample.label the true class.
# build samples from raw bits
let s = makeSample(nBits, rawBits, label, order)
# 1. pre-flight DATA checks
let dc = dataChecks(labels, nClasses, threshold = 0.30)
# dc.classCounts, dc.classShares, dc.majorityClass, dc.majorityShare,
# dc.majorityAccuracy, dc.overThreshold, dc.flags
let sc = shuffledLabelControl(tmplMachine, samples) # (acc, majority, ...)
# 2. clause introspection
let infos = clauseInfo(m, samples, spec)
let summ = clauseSummary(infos) # empty / neverFired / length hist
for c in topClauses(infos, 10): echo c.text, " votes=", c.votes
for cb in clauseBalanceByClass(infos, nClasses): echo cb
for cl in 0..<nClasses: # recover the class rule
echo spec.describeClause(necessaryLiterals(m, samples, cl), cl)
# 3. per-feature contribution + DEAD-INPUT LIST
let contribs = featureContributions(m, samples, spec)
let dead = deadInputs(contribs) # weighted < 5% of top, or never used
let strict = neverUsedInputs(contribs) # appearances == 0
let consts = constantInputs(samples, nBits)
for c in rankedInputs(contribs)[0..<10]: echo c.name, " ", c.weighted
# 4. accuracy diagnostics
let ad = accuracyDiagnostics(m, samples)
# ad.acc, ad.majorityBaseline, ad.margin, ad.confusion,
# ad.perClassRecall/Precision, ad.predMajorityShare
# 5. learning curve
let lc = learningCurve(tmplMachine, train, eval, nPoints = 10)
# lc.points, lc.accs, lc.trend ("flat" | "rising" | "rising-then-falling ...")
# 6. ablation hooks
let base = ablateBaseline(tmplMachine, train, eval)
for r in ablateDropAllBlocks(tmplMachine, train, eval, spec): echo r.name, r.delta
let rs = ablateScrambleFeature(tmplMachine, train, eval, spec, bit)
ablation retrains a fresh machine per variant (the strongest form of "does this
input earn its bits"). Pass baselineAcc from ablateBaseline to avoid
recomputing it per block.
Task 2b / 4 — the CLAUSE-SHAPE checker (group 8)
The automata metrics say whether the TM is settling. The clause-shape checker says whether what it settled ON is sane: a rule needing ~19 conditions to fire is almost certainly fitting noise.
import tm_diag/diagnostics
let d = clauseShapeDiagnostics(m, samples, spec,
healthyLo = 3.0, healthyHi = 8.0,
collapsedMax = 2.0)
d.summary.meanLength # 19.17 for the shipped gun
d.summary.medianLength # 16.0
d.summary.p10Length # 4.0
d.summary.p90Length # 44.8
d.summary.maxLength # 57
d.summary.lengthHist # index = length, value = count
d.summary.posMeanLength / .negMeanLength
d.summary.perClassMeanLength / .perClassMedianLength / .perClassHist
d.verdict # "healthy" | "too long" | "collapsed" | "short" | "n/a"
d.blocks # per-FeatureSpec-block literal contribution
d.coverage # firing fraction + effective clause count
echo shapeLine(d) # one line, same style as healthLine
echo formatClauseShapeReport(d) # full readout
- length distribution —
meanLength,medianLength,p10Length,p90Length,minLength,maxLength,stdLengthandlengthHist, plus per-polarity (posMeanLength/negMeanLength) and per-class (perClassMeanLength/perClassMedianLength/perClassHist) breakdowns.lengthStats(lengths)andpercentile(sorted, p)are exported standalone. - healthy-band verdict —
clauseShapeVerdict(summ, healthyLo=3, healthyHi=8, collapsedMax=2):> healthyHi=too long,< collapsedMax=collapsed, betweencollapsedMaxandhealthyLo=short, in band =healthy, no non-empty clauses =n/a. The band is a PARAMETER of the call. - per-block contribution —
blockLengthContributions(m, spec): for eachFeatureSpecblock, the literals it contributes across clauses (totalLits,meanPerClause,share,clausesUsing,perClauseMax). A block that contributes to every clause is dominating; a block contributing ~0 is dead (the block-level view of the dead-INPUT list). - coverage / concentration —
clauseCoverage(m, samples): the fraction of clauses that fire on a sample (meanFiringFraction), the participation-ratioeffectiveClauses = (sum v)^2 / sum v^2(if 3 clauses cast most votes it reads ~3, not the configured count) andtop3Share.
The one-line healthLine now appends shape=<mean> (<verdict>) alongside the
settledness / diversity / churn / disagreement verdicts.
Task 5 — the AUTOMATA level (settledness / diversity / churn / disagreement)
Task 2 reads the clauses. Task 5 reads the automata inside them: one
automaton per input bit per clause, each holding a state that says how
confident it is that its bit belongs in the clause. These are what tell us
whether the TM is locking onto the enemy or just fidgeting, and whether the
inertia N (the number of automata states) should go up or down.
State conventions — from the ACTUAL code, not the textbook
Derived from tm_core.nim and guns/tm_pattern.nim (tmEval, tmLearnDir,
tmNewTeam, resetMachine):
| property | value |
|---|---|
| state range | [-nStates, nStates] (int16); nStates = 64 for tm_pattern, 32 for tsetlin.nim |
| initial value | 0 (both cores call it "the Exclude boundary") |
| INCLUDE | state > 0 (tmEval / clauseLits) |
| EXCLUDE | state <= 0 |
| flip boundary | BETWEEN state 0 and state 1 — the middle of the range |
commitment(st) |
abs(st) / nStates in [0,1]: 0 on the boundary, 1 at either extreme |
So a flip is exactly a change in the predicate state > 0.
The four metrics
- SETTLEDNESS — per clause and overall, the mean
commitmentand the fraction of automata settled at/above a threshold (default0.5). High and rising as training proceeds is healthy; low means the clause is wavering noise.settledness(m, threshold),settlednessTrend(early, late). - CLAUSE DIVERSITY — mean pairwise Jaccard of the included-literal sets,
within the same polarity. Low-to-moderate is healthy; ~1.0 = all clauses are
one rule in 50 hats; ~0 = memorising ticks. Empty clauses carry no rule and
are skipped by default.
clauseDiversity(m, skipEmpty = true),jaccard(a,b). - CHURN — both levels, per training sample: the fraction of automata
crossing the flip boundary, and the fraction of clauses whose
included-literal set changed. High early and falling is healthy; flat-high
= fidgeting; zero from the start = never learned.
churnTrace(tmpl, samples, epochs, seed, window, shuffle).shuffle = falsereplays the given (temporal) order, mirroring a live gun. - VOTE DISAGREEMENT — per class, over the samples where it casts a vote:
the fraction of firing clauses whose polarity disagrees with the sign of the
class's total vote. Low is healthy.
voteDisagreement(m, samples).
The three readouts that make it readable
- state histogram — the distribution of automata states across the range:
stateHistogram(m, nBins),histogramText(h). - per-input confidence table — for each bit, the mean commitment of its
automata (and a
constantflag for zero-variance inputs):perInputConfidence(m, spec, samples, threshold),rankedInputConfidence(conf). - one-line health summary —
healthLine(ad)givessettledness / diversity / churn trend / disagreement / shape, each with its own verdict word (settling,fidgeting,frozen,coherent,healthy,too long, ...), andautomataVerdict(ad)reduces the trajectory to one word.formatAutomataReport(ad)prints everything, including the clause-shape report whenad.shapewas computed.
API
import tm_diag/diagnostics
let ad = automataDiagnostics(tmpl, samples, spec,
epochs = 15, seed = 777,
settleThreshold = 0.5, nHistBins = 9,
window = 100, measureChurn = true)
# ad.machine, ad.settledness, ad.diversity, ad.churn, ad.disagreement,
# ad.histogram, ad.inputConfidence, ad.shape, ad.summary
echo ad.summary # one line
echo automataVerdict(ad) # "settling" | "fidgeting" | "collapsed" | ...
echo formatAutomataReport(ad) # the full readout
For a machine you already have (e.g. the shipped gun's exportTeams()), call
settledness / clauseDiversity / stateHistogram / perInputConfidence /
voteDisagreement directly and churnTrace on a fresh copy.
Validation — Case A (learnable) vs Case B (noise)
common_libs/tests/diag_automata_validation.nim uses the planted rule from
diag_synthetic.nim and the SAME inputs with shuffled labels. Measured
(nBits=49, 3 classes, 40 clauses, nStates=64, 3000 samples, 15 epochs):
| metric | Case A (learnable) | Case B (noise) |
|---|---|---|
| settledness mean | 0.970 | 0.719 |
| settled fraction | 0.999 | 0.799 |
| churn trend | falling (2.5e-4 -> 2e-6) |
flat (1.0e-3 -> 7.2e-4) |
| clause-change trend | falling (0.011 -> 2.3e-4) |
flat (0.069 -> 0.057) |
| diversity (overall Jaccard) | 0.176 | 0.014 |
| disagreement | 0.003 | 0.298 |
| verdict | settling | mixed (not settling) |
Case A settledness RISES 0.475 (early prefix) -> 0.970 (full). The pair
separates learning from fidgeting: the decisive signals are the churn trend
(falling vs flat) and disagreement (0.003 vs 0.298). Settledness alone is NOT
enough — on noise the automata still commit (0.719), just to the wrong thing.
Case C (a forced-constant bit) is flagged: perInputConfidence(...).constant
and constantInputs both surface it. Note a constant bit can show HIGH
commitment (one literal is always 1), so the constant flag is what
disambiguates.
Inertia sweep (N = nStates, 10 epochs)
The validation also sweeps N to see whether inertia moves the metrics:
| N | Case A settled | A churn | A dis. | Case B settled | B churn | B dis. |
|---|---|---|---|---|---|---|
| 16 | 0.904 | falling | 0.003 | 0.720 | flat | 0.363 |
| 32 | 0.943 | falling | 0.000 | 0.700 | flat | 0.388 |
| 64 | 0.970 | falling | 0.003 | 0.715 | flat | 0.368 |
| 128 | 0.971 | falling | 0.000 | 0.660 | flat | 0.304 |
A and B are separated at EVERY N (churn falling vs flat, disagreement low vs high). Raising N only raises Case A's commitment (0.90 -> 0.97); it does NOT reduce noise-fitting in Case B. So inertia is not the discriminator the metrics identify — the churn trend and disagreement are.
Real reading — the shipped tm_pattern GF head
diag_tm_pattern_offline.nim over the committed DrussGT fixtures (automata read
directly off the exported teams from tr_drussgt_vs_modularbot; churn measured
on a tm_core temporal one-pass proxy, live-order):
- settledness 0.484 (
settling), settled fraction 0.453 - diversity 0.267 (
moderate) — positive 0.170, negative 0.322 - churn 0.094/100 per sample, FALLING (0.00208 -> 0.00048); clause-change 5.31/100, falling
- disagreement 0.145 (
coherent) - verdict: settling — not fidgeting, not collapsed
- constant inputs flagged: bits 38/39 (the known never-written ones) plus 19/36/37 in this fixture
- context: pooled warm accuracy 35.72% vs the 34.24% majority = +1.48pp
Clause-shape reading (group 8, same run): mean=19.17 median=16.00 p10=4.00 p90=44.80 std=13.89 min=1 max=57 nonEmpty=173 empty=27/200 -> verdict
too long. Per polarity posMean=19.80 / negMean=18.70; per class
c0=11.4 c1=13.4 c2=31.2 c3=22.6 c4=21.6. Per-block share is diffuse (no block
dominates): wall-near 10.8%, distance-band 8.9%, flight-band 8.7%, radial-frac
7.6%, speed-band 7.0%, closing 6.8%, UNUSED 6.4% (the always-true negations of
the never-written bits 38/39 are free padding), the rest 2-6%. Coverage:
firing/sample=48.53 (24.3%) effectiveClauses=97.85/200 top3Share=5.3% — the
voting is NOT concentrated on a few clauses. One-line health:
settledness=0.484 (settling) | diversity=0.267 (moderate) | churn=0.094/100 falling (settling) | disagreement=0.145 (coherent) | shape=19.17 (too long)
Does s rescue this gun? MEASURED, no. Recompiling the offline driver with
-d:TM_S_DEF=<v> (source untouched) retrains the gun end to end:
s |
mean clause len | shape verdict | pooled warm acc | margin vs majority |
|---|---|---|---|---|
| 1.5 | 15.82 | too long | 32.22% | -2.03pp |
| 2.0 | 14.99 | too long | 34.03% | -0.22pp |
| 3.0 (shipped) | 19.17 | too long | 35.72% | +1.48pp |
| 5.0 | 20.47 | too long | 34.72% | +0.47pp |
Lowering s shrinks the clauses (15.8 at s=1.5) but makes accuracy WORSE;
raising it pads them and also loses. The shipped s=3.0 is the best of the four,
and NO value gets the mean anywhere near the healthy 3-8 band. Combined with the
settledness finding, the shape is consistent with "no consistent short rule
exists in this representation/target" — the clauses are long, diffuse and
padded, and the knob cannot fix a signal problem.
The gun settles onto the within-battle labels but its settled rules barely beat
the majority class, which points at the TARGET / representation rather than the
inertia N. See the Task 5 report for the inertia discussion.
The default-off real-gun hook
guns/tm_pattern.nim gained only additive, default-off instrumentation:
g.diagCapture = true # default false; no behaviour change when false
# ... replay ...
g.diagSamples # seq[TmDiagSample] (literal vector + label)
g.exportTeams() # read-only GF clause teams
g.exportRadTeams(); g.exportRevTeams()
To introspect an externally trained clause set:
let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
g.exportTeams())
Running the demos / tests
nim c -r -d:release --path:common_libs common_libs/tests/test_tm_diag.nim # 48 pure unit checks
nim c -r --path:common_libs common_libs/tests/test_tm_automata_diag.nim # 55 automata-metric unit checks
nim c -r -d:release --path:common_libs common_libs/tests/test_tm_clause_shape.nim # 66 clause-shape unit + synthetic checks
nim c -r -d:release --path:common_libs common_libs/tests/diag_synthetic.nim # Task 3 proof (17 checks)
nim c -r -d:release --path:common_libs common_libs/tests/diag_automata_validation.nim # Task 5 A/B/C proof
nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim # Task 4 real reading + shape
See common_libs/tests/diag_synthetic.nim for the ground-truth validation,
common_libs/tests/test_tm_clause_shape.nim for the clause-shape validation and
common_libs/tests/diag_tm_pattern_offline.nim for the real reading.