TM diagnostics kit: VALIDATED (finds a known dead input), and it found a real bug

Built `common_libs/tm_diag/` as a first-class offline diagnostics kit for Tsetlin
work, BEFORE writing the new gun - because we hit two data problems tonight that no
amount of reading the TM's clauses would have revealed (a 38.8% majority answer,
and 36-58% mislabelled training samples).

WHAT IT PROVIDES
- `feature_spec.nim`: a NAMED feature container, so a learned clause prints as a
  sentence (`IF near-wall AND bullet-dead-on AND turn-left(t-2) THEN class=3`)
  instead of "feature 17". Includes the 49-bit draft spec from the design session
  and the shipped 40-bit encoding.
- `tm_core.nim`: a compact deterministic Granmo multiclass TM with an
  INTROSPECTABLE clause layout (mirrors the tm_pattern core).
- `diagnostics.nim`, six groups: (1) pre-flight DATA checks + shuffled-label
  control, (2) clause introspection (readable dump, per-clause vote counts, empty
  and never-fired clauses, length distribution, per-class balance), (3)
  per-feature contribution with an explicit DEAD-INPUT LIST and a ranked
  most-valuable list, (4) accuracy vs the majority baseline with per-class
  precision/recall and pred-majority share, (5) learning curve, (6) ablation hooks
  (drop a block / scramble a bit).

=== TASK 3: THE VALIDATION THAT GATES EVERYTHING - PASSED WITH NUMBERS ===
A diagnostic we never checked is worthless, so the kit was tested on a synthetic
set with a PLANTED RULE (class2 = A and B, class1 = A and not B, class0 = not A),
a deliberately IRRELEVANT block (US, 9 bits) and a PURE-NOISE bit (17).
- majority baseline 60.63% (class0); over-30% correctly flagged
- **the planted rule is recovered EXACTLY** via `necessaryLiterals`:
    class0 IF NOT dist-wall<50 | class1 IF dist-wall<50 AND NOT lat DEAD-ON |
    class2 IF dist-wall<50 AND lat DEAD-ON
- **DEAD-INPUT LIST = all 9 US bits AND the noise bit 17**, while the planted bits
  0 and 45 are correctly NOT listed
- top contributors: bit0 w=1241.7, bit45 w=583.3, then 49.8 - a 12-25x gap, so the
  relevant bits are unmistakable
- **ABLATION: drop WALLS -39.47pp, drop BULLETS -19.33pp, drop US 0.00pp**,
  scramble A -42.00pp, scramble the noise bit 0.00pp
- shuffled-label control 60.40% vs majority 60.63% = -0.23pp -> no leak
So the kit reliably finds a known dead input and a known relevant one.

=== TASK 4: THE REAL READING, AND A BUG IN THE SHIPPED GUN ===
`tm_pattern` GF head, 6 DrussGT fixtures, pooled 250,745 samples:
- label balance c2 = **34.4%** (majority-heavy, flagged); accuracy **35.72%** vs
  majority **34.24%** -> margin **+1.48pp**. On `tr_drussgt_vs_crazy` it is BELOW
  majority (33.81% vs 37.72%, -3.92pp).
- 200 clauses: **27 empty, 45 never fired**, mean length 19.17, max 57. The
  majority class is starved (class2: 22 non-empty, 18 empty, only 2 positive
  fired). Class4 fires 11-24-literal clauses -> memorisation signature.
- **REPRESENTATION BUG FOUND (reported, not silently fixed):** `tmBuildBits` writes
  only 38 raw bits into `var bits: array[TM_NBITS=40, uint8]` - bits 38 and 39 are
  NEVER ASSIGNED, so they are always 0 and their negated literals are always 1.
  The kit's `constantInputs` confirms 38/39 are constant, and **`UNUSED-38`/`39`
  rank #6 and #8 in the most-valuable-inputs list** - i.e. the model's
  highest-usage inputs are information-free. That is a representation bug, not a
  display artefact, and it is a concrete mechanism for part of the poor learning.

DRAFT ENCODING CHECKED: the 49-bit draft is arithmetically consistent
(4+4=8 walls, 6+3=9 us, 3+5+3+3+3+3=20 motion, 5+7=12 bullets = 49). No draft
inconsistency.

Guards: test_tm_diag 48 (new), diag_synthetic 17 (new), test_gun_harness 39,
test_vbullet_metric 11, test_power_selection 3, test_adaptive_radar 41,
test_tfil_ring_weights 24, test_power_policy 26, test_ram_decision 40,
test_rack_membership 48, test_selector_tiebreak 19, test_tm_pattern_registration 20,
test_vbullet_admit_gate 12, acceptance_offline_vs_online 12/12; tm_pattern_learning
passes. The tm_pattern hook is additive and default-OFF (no behaviour change).

NOT YET INCLUDED (the automata metrics discussed for the next step): per-clause
automata settledness (distance from the flip point), clause diversity (pairwise
overlap), literal-set churn over time, and cross-clause vote disagreement. The kit
has clause-level diagnostics but not the automata-state ones.
This commit is contained in:
2026-09-22 21:24:27 +02:00
parent b0654d18eb
commit f9f8d84671
8 changed files with 1608 additions and 0 deletions
@@ -0,0 +1,181 @@
## TASK 4 — the tm_diag kit against the EXISTING guns/tm_pattern.nim GF head.
##
## Replays the committed DrussGT fixtures through the offline gun range with a
## FRESH, diagCapture-enabled TM per fixture (the live semantics: cold every
## battle, overfit within the battle), then reports:
## * pooled label balance + majority baseline vs the gun's own warm accuracy;
## * the DEAD-INPUT LIST for the actual 40-bit tm_pattern feature set;
## * the top firing positive clauses and the recovered necessary literals.
##
## Offline only. Fixtures are READ-ONLY.
## Run: nim c -r -d:release --path:common_libs common_libs/tests/diag_tm_pattern_offline.nim
## [--fixtures=a,b,c] [--maxsamples=N]
import std/[os, strformat, strutils, algorithm, random]
import gun_harness/offline_range
import guns/tm_pattern
import tm_diag/diagnostics
const repoRoot = currentSourcePath().parentDir.parentDir.parentDir
const fixturesDir = repoRoot / "tools" / "fixtures"
proc toDiag(s: TmDiagSample): DiagSample =
result.lits = newSeq[uint8](TM_NLITS)
for i in 0..<TM_NLITS: result.lits[i] = s.lits[i]
result.label = s.label
result.order = s.order
proc printPooledConfusion(cm: array[TM_CLASSES, array[TM_CLASSES, int]]) =
var total = 0
var correct = 0
var rowSum, colSum: array[TM_CLASSES, int]
for c in 0..<TM_CLASSES:
for p in 0..<TM_CLASSES:
total += cm[c][p]
if c == p: correct += cm[c][p]
rowSum[c] += cm[c][p]
colSum[p] += cm[c][p]
var maj = 0
for c in 1..<TM_CLASSES:
if rowSum[c] > rowSum[maj]: maj = c
let majShare = if total > 0: rowSum[maj].float / total.float else: 0.0
let acc = if total > 0: correct.float / total.float else: 0.0
echo &"# pooled warm accuracy = {correct}/{total} = {acc*100:.2f}%"
echo &"# majority class = {maj} share/baseline = {rowSum[maj]}/{total} = {majShare*100:.2f}%"
echo &"# margin = {(acc-majShare)*100:+.2f}pp pred-majority share = " &
&"{colSum[maj].float/max(1,total).float*100:.2f}%"
for c in 0..<TM_CLASSES:
let rec = if rowSum[c] > 0: cm[c][c].float / rowSum[c].float else: 0.0
let prec = if colSum[c] > 0: cm[c][c].float / colSum[c].float else: 0.0
echo &"# class{c}: trueN={rowSum[c]:<6} predN={colSum[c]:<6} TP={cm[c][c]:<6} " &
&"recall={rec*100:5.1f}% precision={prec*100:5.1f}%"
proc main() =
var names = @["drussgt_vs_crazy", "drussgt_vs_spinbot", "drussgt_vs_drussgt",
"tr_drussgt_vs_crazy", "tr_drussgt_vs_spinbot",
"tr_drussgt_vs_modularbot"]
var maxSamples = 40000
for i in 1..paramCount():
let a = paramStr(i)
if a.startsWith("--fixtures="):
names = a[11..^1].split(',')
elif a.startsWith("--maxsamples="):
maxSamples = parseInt(a[13..^1])
let spec = tmPatternSpec()
echo &"# tm_pattern GF head over DrussGT fixtures (spec nBits={spec.nBits}, " &
&"TM_CLASSES={TM_CLASSES}, TM_NCLAUSES={TM_NCLAUSES})"
echo "# fixture,samples,captured,classTotal,labelHist"
var pooledLabels: seq[int]
var pooledConfusion: array[TM_CLASSES, array[TM_CLASSES, int]]
var pooledClassCorrect, pooledClassTotal = 0
var pooledLabelHist: array[TM_CLASSES, int]
var bestGun: TmPatternGun
var bestSamples: seq[DiagSample]
var bestName = ""
var bestClassTotal = -1
for name in names:
let path = fixturesDir / (name & ".jsonl")
if not fileExists(path):
echo &"# SKIP missing fixture {path}"
continue
let fx = loadFixture(path)
let g = new(TmPatternGun)
g[] = initTmPatternGun()
g[].targetMode = tmGF
g[].diagCapture = true
randomize(1234)
let drv = GunDriver(
name: "TMPatGF",
predictCb: proc(state: WorldState, bs: float): GunPrediction = g[].predict(state, bs),
resultCb: proc(e: FeedbackEvent) = g[].onResult(e),
readyCb: proc(): bool = true)
discard replayFixture(fx, @[drv], fx.enemyId)
var hs: string
for c in 0..<TM_CLASSES:
inc pooledLabelHist[c], g[].labelHist[c]
hs.add $g[].labelHist[c] & ","
for p in 0..<TM_CLASSES:
pooledConfusion[c][p] += g[].confusion[c][p]
pooledClassCorrect += g[].classCorrect
pooledClassTotal += g[].classTotal
echo &"# {name},{fx.states.len},{g[].diagSamples.len},{g[].classTotal},[{hs}]"
# Keep the fixture with the most warm labelled samples for clause
# introspection (the model is fresh per fixture, so we cannot pool teams).
var thisSamples: seq[DiagSample]
let cap = min(g[].diagSamples.len, maxSamples)
thisSamples = newSeq[DiagSample](cap)
for i in 0..<cap: thisSamples[i] = toDiag(g[].diagSamples[i])
if g[].classTotal > bestClassTotal and thisSamples.len > 0:
bestClassTotal = g[].classTotal
bestGun = g[]
bestSamples = thisSamples
bestName = name
# ── pooled label balance + majority baseline vs real accuracy ──
for c in 0..<TM_CLASSES:
for _ in 0..<pooledLabelHist[c]: pooledLabels.add c
let dc = dataChecks(pooledLabels, TM_CLASSES)
echo "\n## POOLED label balance (all resolved bullets, warm+cold)"
echo &"# n={dc.n} counts={dc.classCounts}"
var sh = ""
for c in 0..<TM_CLASSES: sh.add &"c{c}={dc.classShares[c]*100:.1f}% "
echo "# shares: ", sh
echo &"# majority = class{dc.majorityClass} @ {dc.majorityShare*100:.2f}% overThreshold={dc.overThreshold}"
for f in dc.flags: echo "# ", f
echo "\n## POOLED accuracy (gun's own warm confusion)"
echo &"# cross-check: gun classCorrect/classTotal = {pooledClassCorrect}/{pooledClassTotal}"
printPooledConfusion(pooledConfusion)
if bestSamples.len == 0:
echo "\n# no captured samples; aborting introspection"
return
echo &"\n## CLAUSE / INPUT INTROSPECTION on fixture '{bestName}' " &
&"(bestGun: {bestSamples.len} captured samples, {bestClassTotal} warm)"
let m = machineFromTeams(TM_NBITS, TM_CLASSES, TM_NCLAUSES, TM_NSTATES, TM_S,
bestGun.exportTeams())
let infos = clauseInfo(m, bestSamples, spec)
let summ = clauseSummary(infos)
echo &"# clauses: total={summ.totalClauses} empty={summ.emptyClauses} " &
&"nonEmpty={summ.nonEmpty} fired>=1={summ.firedAtLeastOnce} " &
&"neverFired={summ.neverFired} posFired={summ.posFired} negFired={summ.negFired}"
echo &"# clause length: mean={summ.meanLength:.2f} max={summ.maxLength} hist={summ.lengthHist}"
for cb in clauseBalanceByClass(infos, TM_CLASSES):
echo &"# class{cb.cls}: nonEmpty={cb.nonEmpty} empty={cb.empty} " &
&"posFired={cb.posFired} negFired={cb.negFired} neverFired={cb.neverFired}"
let contribs = featureContributions(m, bestSamples, spec)
let ranked = rankedInputs(contribs)
let dead = deadInputs(contribs)
let never = neverUsedInputs(contribs)
let consts = constantInputs(bestSamples, TM_NBITS)
echo "\n## INPUT VALUE RANKING (top 15 by weighted vote-share)"
for i in 0..<min(15, ranked.len):
echo &"# {ranked[i].bit:>2} {ranked[i].name:<26} appearances={ranked[i].appearances:<8} weighted={ranked[i].weighted:.1f}"
echo "\n## DEAD-INPUT LIST (weighted < 5% of top, or never in a voting clause):"
echo "# ", dead
echo "## never-used (strict appearances==0): ", never
echo "## constant inputs (zero variance in this fixture): ", consts
echo "\n## TOP FIRING POSITIVE CLAUSES per class"
for cls in 0..<TM_CLASSES:
var clsInfos: seq[ClauseInfo]
for c in infos:
if c.cls == cls: clsInfos.add c
let tc = topClausesByPolarity(clsInfos, 1, 4)
if tc.len == 0:
echo &"# class{cls}: (no firing positive clauses)"
for c in tc:
echo &"# class{cls} votes={c.votes:<7} len={c.length} {c.text}"
echo &"# -> necessary literals: {spec.describeClause(necessaryLiterals(m, bestSamples, cls), cls)}"
when isMainModule:
main()