Files
SirStone f9f8d84671 TM diagnostics kit: VALIDATED (finds a known dead input), and it found a real bug
Built `common_libs/tm_diag/` as a first-class offline diagnostics kit for Tsetlin
work, BEFORE writing the new gun - because we hit two data problems tonight that no
amount of reading the TM's clauses would have revealed (a 38.8% majority answer,
and 36-58% mislabelled training samples).

WHAT IT PROVIDES
- `feature_spec.nim`: a NAMED feature container, so a learned clause prints as a
  sentence (`IF near-wall AND bullet-dead-on AND turn-left(t-2) THEN class=3`)
  instead of "feature 17". Includes the 49-bit draft spec from the design session
  and the shipped 40-bit encoding.
- `tm_core.nim`: a compact deterministic Granmo multiclass TM with an
  INTROSPECTABLE clause layout (mirrors the tm_pattern core).
- `diagnostics.nim`, six groups: (1) pre-flight DATA checks + shuffled-label
  control, (2) clause introspection (readable dump, per-clause vote counts, empty
  and never-fired clauses, length distribution, per-class balance), (3)
  per-feature contribution with an explicit DEAD-INPUT LIST and a ranked
  most-valuable list, (4) accuracy vs the majority baseline with per-class
  precision/recall and pred-majority share, (5) learning curve, (6) ablation hooks
  (drop a block / scramble a bit).

=== TASK 3: THE VALIDATION THAT GATES EVERYTHING - PASSED WITH NUMBERS ===
A diagnostic we never checked is worthless, so the kit was tested on a synthetic
set with a PLANTED RULE (class2 = A and B, class1 = A and not B, class0 = not A),
a deliberately IRRELEVANT block (US, 9 bits) and a PURE-NOISE bit (17).
- majority baseline 60.63% (class0); over-30% correctly flagged
- **the planted rule is recovered EXACTLY** via `necessaryLiterals`:
    class0 IF NOT dist-wall<50 | class1 IF dist-wall<50 AND NOT lat DEAD-ON |
    class2 IF dist-wall<50 AND lat DEAD-ON
- **DEAD-INPUT LIST = all 9 US bits AND the noise bit 17**, while the planted bits
  0 and 45 are correctly NOT listed
- top contributors: bit0 w=1241.7, bit45 w=583.3, then 49.8 - a 12-25x gap, so the
  relevant bits are unmistakable
- **ABLATION: drop WALLS -39.47pp, drop BULLETS -19.33pp, drop US 0.00pp**,
  scramble A -42.00pp, scramble the noise bit 0.00pp
- shuffled-label control 60.40% vs majority 60.63% = -0.23pp -> no leak
So the kit reliably finds a known dead input and a known relevant one.

=== TASK 4: THE REAL READING, AND A BUG IN THE SHIPPED GUN ===
`tm_pattern` GF head, 6 DrussGT fixtures, pooled 250,745 samples:
- label balance c2 = **34.4%** (majority-heavy, flagged); accuracy **35.72%** vs
  majority **34.24%** -> margin **+1.48pp**. On `tr_drussgt_vs_crazy` it is BELOW
  majority (33.81% vs 37.72%, -3.92pp).
- 200 clauses: **27 empty, 45 never fired**, mean length 19.17, max 57. The
  majority class is starved (class2: 22 non-empty, 18 empty, only 2 positive
  fired). Class4 fires 11-24-literal clauses -> memorisation signature.
- **REPRESENTATION BUG FOUND (reported, not silently fixed):** `tmBuildBits` writes
  only 38 raw bits into `var bits: array[TM_NBITS=40, uint8]` - bits 38 and 39 are
  NEVER ASSIGNED, so they are always 0 and their negated literals are always 1.
  The kit's `constantInputs` confirms 38/39 are constant, and **`UNUSED-38`/`39`
  rank #6 and #8 in the most-valuable-inputs list** - i.e. the model's
  highest-usage inputs are information-free. That is a representation bug, not a
  display artefact, and it is a concrete mechanism for part of the poor learning.

DRAFT ENCODING CHECKED: the 49-bit draft is arithmetically consistent
(4+4=8 walls, 6+3=9 us, 3+5+3+3+3+3=20 motion, 5+7=12 bullets = 49). No draft
inconsistency.

Guards: test_tm_diag 48 (new), diag_synthetic 17 (new), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, acceptance_offline_vs_online 12/12; tm_pattern_learning
passes. The tm_pattern hook is additive and default-OFF (no behaviour change).

NOT YET INCLUDED (the automata metrics discussed for the next step): per-clause
automata settledness (distance from the flip point), clause diversity (pairwise
overlap), literal-set churn over time, and cross-clause vote disagreement. The kit
has clause-level diagnostics but not the automata-state ones.
2026-09-22 21:24:27 +02:00

202 lines
9.1 KiB
Nim

## Task 3 — VALIDATE THE DIAGNOSTICS AGAINST A KNOWN GROUND TRUTH.
##
## A synthetic dataset is generated from a planted decision list over the DRAFT
## 49-bit encoding:
## A = bit 0 ("dist-wall<50", WALLS block)
## B = bit 45 ("lat DEAD-ON -18..+18", BULLETS block)
## class 2 = A AND B ; class 1 = A AND NOT B ; class 0 = NOT A
## Deliberately irrelevant: the whole US block (bits 8..16) plus one pure-noise
## MOTION bit (bit 17).
##
## The kit must recover the rule, flag the dead inputs, show ~0 ablation delta
## for the irrelevant block, and sit at the majority baseline under shuffled
## labels. Every number printed here is MEASURED.
##
## Run: nim c -r --path:common_libs common_libs/tests/diag_synthetic.nim
import std/[random, strformat, strutils, algorithm]
import tm_diag/diagnostics
const
NBits = 49
BitA = 0 ## dist-wall<50
BitB = 45 ## bullet lateral DEAD-ON
NoiseBit = 17
NClasses = 3
var failures = 0
proc check(name: string, ok: bool) =
if ok: echo "PASS: ", name
else: echo "FAIL: ", name; inc failures
proc genDataset(n, seed: int): seq[DiagSample] =
var rng = initRand(seed)
for i in 0..<n:
var raw = newSeq[int](NBits)
for b in 0..<NBits:
raw[b] = (if rng.rand(1.0) < 0.5: 1 else: 0)
let a = if rng.rand(1.0) < 0.4: 1 else: 0
let bb = if rng.rand(1.0) < 0.5: 1 else: 0
raw[BitA] = a
raw[BitB] = bb
let label =
if a == 1 and bb == 1: 2
elif a == 1: 1
else: 0
result.add makeSample(NBits, raw, label, i)
when isMainModule:
let spec = draftTMSpec()
let train = genDataset(3000, 1)
let eval = genDataset(1500, 2)
echo &"# dataset: train={train.len} eval={eval.len} nBits={spec.nBits} rule=A&B->2 | A&!B->1 | !A->0"
# ── GROUP 1: data checks ──
var labels = newSeq[int](train.len)
for i, s in train: labels[i] = s.label
let dc = dataChecks(labels, NClasses)
echo "\n## GROUP 1 data checks"
echo &"# n={dc.n} counts={dc.classCounts} shares={dc.classShares} " &
&"majority=class{dc.majorityClass} share={dc.majorityShare*100:.2f}% " &
&"overThreshold={dc.overThreshold}"
for f in dc.flags: echo "# FLAG ", f
# manual majority for the check
var manualMaj = 0
for c in 0..<NClasses:
if dc.classCounts[c] > dc.classCounts[manualMaj]: manualMaj = c
check "majority class computed from the counts",
dc.majorityClass == manualMaj and
abs(dc.majorityShare - dc.classCounts[manualMaj].float / dc.n.float) < 1e-9
# ── train the reference TM ──
let tmpl = newMachine(NBits, NClasses, nClauses = 40, nStates = 64,
sValue = 3.0, seed = 1)
let m = trainModel(tmpl, train, epochs = 25, seed = 777)
let trainAcc = evalAcc(m, train)
let testAcc = evalAcc(m, eval)
echo &"\n## trained reference TM: trainAcc={trainAcc*100:.2f}% testAcc={testAcc*100:.2f}%"
# ── GROUP 2: clause introspection ──
let infos = clauseInfo(m, eval, spec)
let summ = clauseSummary(infos)
echo "\n## GROUP 2 clause introspection"
echo &"# totalClauses={summ.totalClauses} empty={summ.emptyClauses} " &
&"nonEmpty={summ.nonEmpty} fired>=1={summ.firedAtLeastOnce} " &
&"neverFired={summ.neverFired} posFired={summ.posFired} negFired={summ.negFired} " &
&"meanLen={summ.meanLength:.2f} maxLen={summ.maxLength}"
echo "# lengthHist (idx=len): ", summ.lengthHist
echo "## top firing POSITIVE clauses per class (rule-encoding; polarity annotated)"
for cls in 0..<NClasses:
var clsInfos: seq[ClauseInfo]
for c in infos:
if c.cls == cls: clsInfos.add c
for c in topClausesByPolarity(clsInfos, 1, 3):
echo &"# class{c.cls} votes={c.votes:<5} len={c.length} pol={c.polarity:+d} {c.text}"
# planted rule recovery: the literals common to every firing positive clause
# of a class are exactly the planted rule (the TM pads clauses with junk).
proc hasRule(cls: int, wantPositive: set[uint8], wantNeg: set[uint8]): bool =
var pos, neg: set[uint8]
for l in necessaryLiterals(m, eval, cls):
if l < NBits: pos.incl uint8(l)
else: neg.incl uint8(l - NBits)
pos == wantPositive and neg == wantNeg
echo "# recovered necessary literals per class (intersection of firing positive clauses):"
for c in 0..<NClasses:
echo &"# class{c}: {spec.describeClause(necessaryLiterals(m, eval, c), c)}"
check "clause introspection recovers class2 = (A AND B)",
hasRule(2, {uint8(BitA), uint8(BitB)}, {})
check "clause introspection recovers class1 = (A AND NOT B)",
hasRule(1, {uint8(BitA)}, {uint8(BitB)})
check "clause introspection recovers class0 = (NOT A)",
hasRule(0, {}, {uint8(BitA)})
# ── GROUP 3: per-feature contribution + dead inputs ──
let contribs = featureContributions(m, eval, spec)
let dead = deadInputs(contribs)
let neverUsed = neverUsedInputs(contribs)
let consts = constantInputs(eval, NBits)
let ranked = rankedInputs(contribs)
echo "\n## GROUP 3 per-feature contribution"
echo "# top-8 ranked inputs:"
for i in 0..<min(8, ranked.len):
echo &"# {ranked[i].bit:>2} {ranked[i].name:<22} appearances={ranked[i].appearances} weighted={ranked[i].weighted:.1f}"
echo "# DEAD-INPUT LIST (weighted < 5% of top, plus never-in-voting-clause): ", dead
echo "# never-used (strict appearances==0): ", neverUsed
echo "# constant inputs: ", consts
check "the planted relevant bits are the top-2 contributors",
ranked[0].bit in {BitA, BitB} and ranked[1].bit in {BitA, BitB}
# the whole US block must be dead
var usDead = true
for b in 8..16:
if b notin dead: usDead = false
check "the whole irrelevant US block is on the dead-input list", usDead
check "the pure-noise MOTION bit is on the dead-input list", NoiseBit in dead
check "no planted relevant bit is called dead",
BitA notin dead and BitB notin dead
# ── GROUP 4: accuracy diagnostics ──
let ad = accuracyDiagnostics(m, eval)
echo "\n## GROUP 4 accuracy diagnostics"
echo &"# n={ad.n} correct={ad.correct} acc={ad.acc*100:.2f}% " &
&"majority=class{ad.majorityClass} baseline={ad.majorityBaseline*100:.2f}% " &
&"margin={ad.margin*100:+.2f}pp predMajorityShare={ad.predMajorityShare*100:.2f}%"
echo "# confusion [true][pred]:"
for c in 0..<NClasses:
echo &"# true {c}: {ad.confusion[c]}"
for c in 0..<NClasses:
echo &"# class{c}: recall={ad.perClassRecall[c]*100:.1f}% precision={ad.perClassPrecision[c]*100:.1f}%"
check "accuracy beats the majority baseline by a large margin", ad.margin > 0.3
check "uniform-ish predictions (not majority-spam): predMajorityShare < 0.7",
ad.predMajorityShare < 0.7
# ── GROUP 5: learning curve ──
let lc = learningCurve(tmpl, train, eval, nPoints = 10, seed = 31)
echo "\n## GROUP 5 learning curve"
echo &"# trend={lc.trend}"
for i in 0..<lc.points.len:
echo &"# samples={lc.points[i]:<5} acc={lc.accs[i]*100:.2f}%"
check "the learning curve is rising (not flat)",
lc.trend.startsWith("rising") and not lc.trend.startsWith("rising-then")
# ── GROUP 6: ablation ──
echo "\n## GROUP 6 ablation (retrain from scratch per variant)"
let base = ablateBaseline(tmpl, train, eval, epochs = 15, seed = 777)
echo &"# baseline acc={base.baselineAcc*100:.2f}%"
# drop every block
var usDelta = 0.0
var wallsDelta = 0.0
var bulletsDelta = 0.0
for i, b in spec.blocks:
let r = ablateDropBlock(tmpl, train, eval, spec, i, epochs = 15, seed = 777,
baselineAcc = base.baselineAcc)
echo &"# drop-block {b.name:<28} acc={r.ablatedAcc*100:.2f}% delta={r.delta*100:+.2f}pp"
if b.name == "dist-from-us": usDelta = r.delta
if b.name == "dist-to-nearest-wall": wallsDelta = r.delta
if b.name == "bullet-lateral-offset": bulletsDelta = r.delta
check "dropping the irrelevant US block costs ~0 accuracy", abs(usDelta) < 0.03
check "dropping the WALLS block (holds A) costs real accuracy", wallsDelta < -0.10
check "dropping the BULLETS block (holds B) costs real accuracy", bulletsDelta < -0.05
# scramble ONE relevant bit vs ONE noise bit
let scA = ablateScrambleFeature(tmpl, train, eval, spec, BitA, epochs = 15, seed = 777, baselineAcc = base.baselineAcc)
let scNoise = ablateScrambleFeature(tmpl, train, eval, spec, NoiseBit, epochs = 15, seed = 777, baselineAcc = base.baselineAcc)
echo &"# scramble {scA.name:<20} delta={scA.delta*100:+.2f}pp"
echo &"# scramble {scNoise.name:<20} delta={scNoise.delta*100:+.2f}pp"
check "scrambling relevant bit A costs real accuracy", scA.delta < -0.10
check "scrambling the noise bit costs ~0 accuracy", abs(scNoise.delta) < 0.03
# ── GROUP 1b: shuffled-label control ──
let sc = shuffledLabelControl(tmpl, train, epochs = 10, seed = 4242)
echo "\n## GROUP 1 shuffled-label control"
echo &"# shuffled-label accuracy={sc.acc*100:.2f}% majority={sc.majority*100:.2f}% " &
&"n={sc.n} gap={(sc.acc-sc.majority)*100:+.2f}pp"
check "shuffled-label control sits at the majority baseline (|gap| < 4pp)",
abs(sc.acc - sc.majority) < 0.04
echo ""
if failures > 0:
echo &"{failures} check(s) FAILED"
quit(1)
echo "All synthetic diagnostic validation checks passed."